EDBT 2026 Demo / reviewers in the wild / expert
Ling Huang 0001
dblp:90/3799-1
· DBLP profile ↗
35ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-authorDatabases, data management, data science and information retrieval · 9Computer networks · 6 · 1 first-authorSecurity and privacy · 6Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
8 papers |
Cryptographic primitives and cryptanalysis · 37% Privacy and data protection · 14% Cryptographic protocols and secure computation · 12% | |
| Databases, data mining, and information retrieval
7 papers |
Data mining · 84% Web and social media mining · 10% Information retrieval · 6% | |
| Artificial intelligence
6 papers |
Information extraction and text analysis · 45% Generative modeling · 22% Learning theory · 11% | |
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Performance modeling and evaluation · 46% Distributed systems · 29% Embedded and real-time systems · 20% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Theoretical computer science
2 papers |
Graph algorithms and graph theory · 67% Mathematical optimization · 33% | |
| Software engineering, system software, and programming languages
4 papers |
Software maintenance and evolution · 86% Empirical software engineering · 14% |
Topics — the 30 heaviest of 65, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
anomaly detection |
0.8 | 2 | 2020 | Modeling Heterogeneous Statistical Patterns in High-dimensional Data by Adversarial Distributions: An Unsupervised Generative Framework · WWW 2020 No Place to Hide: Catching Fraudulent Entities in Tensors · WWW 2019 |
Cryptographic primitives and cryptanalysis › functional encryption
hidden vector encryption |
0.5 | 2 | 2016 | CloudKeyBank: Privacy and owner authorization enforced key management framework · ICDE 2016 CloudKeyBank: Privacy and Owner Authorization Enforced Key Management Framework · IEEE Trans. Knowl. Data Eng. 2015 |
Cryptographic protocols and secure computation
key management |
0.5 | 2 | 2016 | CloudKeyBank: Privacy and owner authorization enforced key management framework · ICDE 2016 CloudKeyBank: Privacy and Owner Authorization Enforced Key Management Framework · IEEE Trans. Knowl. Data Eng. 2015 |
Cryptographic primitives and cryptanalysis
proxy re-encryption |
0.5 | 2 | 2016 | CloudKeyBank: Privacy and owner authorization enforced key management framework · ICDE 2016 CloudKeyBank: Privacy and Owner Authorization Enforced Key Management Framework · IEEE Trans. Knowl. Data Eng. 2015 |
Machine learning › Generative modeling
generative model |
0.4 | 1 | 2020 | Modeling Heterogeneous Statistical Patterns in High-dimensional Data by Adversarial Distributions: An Unsupervised Generative Framework · WWW 2020 |
Data mining › anomaly detection
unsupervised anomaly detection |
0.4 | 1 | 2020 | Modeling Heterogeneous Statistical Patterns in High-dimensional Data by Adversarial Distributions: An Unsupervised Generative Framework · WWW 2020 |
Visualization and visual analytics › visual analytics
fraud detection |
0.4 | 1 | 2020 | FDHelper: Assist Unsupervised Fraud Detection Experts with Interactive Feature Selection and Evaluation · CHI 2020 |
Visualization and visual analytics
visual analytics |
0.4 | 1 | 2020 | FDHelper: Assist Unsupervised Fraud Detection Experts with Interactive Feature Selection and Evaluation · CHI 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.4 | 1 | 2019 | DIAG-NRE: A Neural Pattern Diagnosis Framework for Distantly Supervised Neural Relation Extraction · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.4 | 1 | 2019 | DIAG-NRE: A Neural Pattern Diagnosis Framework for Distantly Supervised Neural Relation Extraction · ACL (1) 2019 |
Data mining › anomaly detection
dense block detection |
0.4 | 1 | 2019 | No Place to Hide: Catching Fraudulent Entities in Tensors · WWW 2019 |
Data mining › anomaly detection
fraud detection |
0.4 | 1 | 2019 | No Place to Hide: Catching Fraudulent Entities in Tensors · WWW 2019 |
Data mining › multidimensional data analysis › multiway data analysis
tensor analysis |
0.4 | 1 | 2019 | No Place to Hide: Catching Fraudulent Entities in Tensors · WWW 2019 |
Graph algorithms and graph theory
dense subgraph discovery |
0.4 | 1 | 2019 | No Place to Hide: Catching Fraudulent Entities in Tensors · WWW 2019 |
Usable security › security operations
abuse detection |
0.3 | 1 | 2018 | Do Not Pull My Data for Resale: Protecting Data Providers Using Data Retrieval Pattern Analysis · SIGIR 2018 |
Systems and software security › vulnerability analysis
API misuse detection |
0.3 | 1 | 2018 | Do Not Pull My Data for Resale: Protecting Data Providers Using Data Retrieval Pattern Analysis · SIGIR 2018 |
Software maintenance and evolution
log analysis |
0.3 | 3 | 2010 | Detecting Large-Scale System Problems by Mining Console Logs · ICML 2010 Detecting large-scale system problems by mining console logs · SOSP 2009 Online System Problem Detection by Mining Patterns of Console Logs · ICDM 2009 |
Performance modeling and evaluation
workload characterization |
0.3 | 2 | 2013 | Mantis: Automatic Performance Prediction for Smartphone Applications · USENIX ATC 2013 Predicting Execution Time of Computer Programs Using Sparse Polynomial Regression · NIPS 2010 |
Cryptographic primitives and cryptanalysis › proxy re-encryption
conditional proxy re-encryption |
0.2 | 1 | 2016 | CloudKeyBank: Privacy and owner authorization enforced key management framework · ICDE 2016 |
Cryptographic primitives and cryptanalysis
searchable encryption |
0.2 | 1 | 2016 | CloudKeyBank: Privacy and owner authorization enforced key management framework · ICDE 2016 |
Web and social media mining › user identity linkage
author linking |
0.2 | 1 | 2015 | What You Submit Is Who You Are: A Multimodal Approach for Deanonymizing Scientific Publications · IEEE Trans. Inf. Forensics Secur. 2015 |
Information retrieval › text analysis
stylometry |
0.2 | 1 | 2015 | What You Submit Is Who You Are: A Multimodal Approach for Deanonymizing Scientific Publications · IEEE Trans. Inf. Forensics Secur. 2015 |
Privacy and data protection
anonymity |
0.2 | 1 | 2015 | What You Submit Is Who You Are: A Multimodal Approach for Deanonymizing Scientific Publications · IEEE Trans. Inf. Forensics Secur. 2015 |
Embedded and real-time systems
mobile devices |
0.2 | 1 | 2015 | Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile Devices · IEEE Trans. Mob. Comput. 2015 |
Performance modeling and evaluation › performance prediction
resource usage prediction |
0.2 | 1 | 2015 | Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile Devices · IEEE Trans. Mob. Comput. 2015 |
Machine learning › Learning theory › margin-based learning
large margin classification |
0.2 | 1 | 2014 | Large-Margin Convex Polytope Machine · NIPS 2014 |
Mathematical optimization › continuous optimization
convex optimization |
0.2 | 1 | 2014 | Large-Margin Convex Polytope Machine · NIPS 2014 |
Data mining
clustering |
0.2 | 2 | 2009 | Fast approximate spectral clustering · KDD 2009 Spectral Clustering with Perturbed Data · NIPS 2008 |
Data mining › clustering
spectral clustering |
0.2 | 2 | 2009 | Fast approximate spectral clustering · KDD 2009 Spectral Clustering with Perturbed Data · NIPS 2008 |
Network security › intrusion detection and prevention
intrusion detection |
0.2 | 2 | 2009 | ANTIDOTE: understanding and defending against poisoning of anomaly detectors · Internet Measurement Conference 2009 Communication-Efficient Online Detection of Network-Wide Anomalies · INFOCOM 2007 |
Methods — techniques the papers use, named apart from their topics
machine learning · 0.9user study · 0.9interactive visualization · 0.9generative framework · 0.9entropy-based distance metric · 0.9case study · 0.9parallel densest subgraph algorithm · 0.8information sharing graph · 0.8proxy re-encryption · 0.7hidden vector encryption · 0.7ensemble classifier · 0.7multiclass classification · 0.4multi-label classification · 0.4adversarial distributions · 0.4adversarial distribution · 0.4neural pattern diagnosis · 0.4human-in-the-loop · 0.4feature engineering · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | FDHelper: Assist Unsupervised Fraud Detection Experts with Interactive Feature Selection and EvaluationabstractOnline fraud is the well-known dark side of the modern Internet. Unsupervised fraud detection algorithms are widely used to address this problem. However, selecting features, adjusting hyperparameters, evaluating the algorithms, and eliminating false positives all require human expert involvement. In this work, we design and implement an end-to-end interactive visualization system, FDHelper, based on the deep understanding of the mechanism of the black market and fraud detection algorithms. We identify a workflow based on experience from both fraud detection algorithm experts and domain experts. Using a multi-granularity three-layer visualization map embedding an entropy-based distance metric ColDis, analysts can interactively select different feature sets, refine fraud detection algorithms, tune parameters and evaluate the detection result in near real-time. We demonstrate the effectiveness and significance of FDHelper through two case studies with state-of-the-art fraud detection algorithms, interviews with domain experts and algorithm experts, and a user study with eight first-time end users. Jiao Sun, Yin Li 0008, Charley Chen, Jihae Lee, Zhongping Zhang, Ling Huang 0001, Lei Shi 0002, Wei Xu 0005 |
CHI | 7 |
| 2020 | Modeling Heterogeneous Statistical Patterns in High-dimensional Data by Adversarial Distributions: An Unsupervised Generative FrameworkabstractSince the label collecting is prohibitive and time-consuming, unsupervised methods are preferred in applications such as fraud detection. Meanwhile, such applications usually require modeling the intrinsic clusters in high-dimensional data, which usually displays heterogeneous statistical patterns as the patterns of different clusters may appear in different dimensions. Existing methods propose to model the data clusters on selected dimensions, yet globally omitting any dimension may damage the pattern of certain clusters. To address the above issues, we propose a novel unsupervised generative framework called FIRD, which utilizes adversarial distributions to fit and disentangle the heterogeneous statistical patterns. When applying to discrete spaces, FIRD effectively distinguishes the synchronized fraudsters from normal users. Besides, FIRD also provides superior performance on anomaly detection datasets compared with SOTA anomaly detection methods (over 5% average AUC improvement). The significant experiment results on various datasets verify that the proposed method can better model the heterogeneous statistical patterns in high-dimensional data and benefit downstream applications. Wenhao Zheng 0001, Charley Chen, Kevin Gao, Yao Hu 0002, Ling Huang 0001, Wei Xu 0005 |
WWW | 6 |
| 2019 | DIAG-NRE: A Neural Pattern Diagnosis Framework for Distantly Supervised Neural Relation ExtractionabstractPattern-based labeling methods have achieved promising results in alleviating the inevitable labeling noises of distantly supervised neural relation extraction.However, these methods require significant expert labor to write relation-specific patterns, which makes them too sophisticated to generalize quickly.To ease the labor-intensive workload of pattern writing and enable the quick generalization to new relation types, we propose a neural pattern diagnosis framework, DIAG-NRE, that can automatically summarize and refine highquality relational patterns from noise data with human experts in the loop.To demonstrate the effectiveness of DIAG-NRE, we apply it to two real-world datasets and present both significant and interpretable improvements over state-of-the-art methods. Shun Zheng 0001, Xu Han 0007, Yankai Lin 0001, Ling Huang 0001, Zhiyuan Liu 0001, Wei Xu 0005 |
ACL (1) | 6 |
| 2019 | No Place to Hide: Catching Fraudulent Entities in TensorsabstractMany approaches focus on detecting dense blocks in the tensor of multimodal data to prevent fraudulent entities (e.g., accounts, links) from retweet boosting, hashtag hijacking, link advertising, etc. However, no existing method is effective to find the dense block if it only possesses high density on a subset of all dimensions in tensors. In this paper, we novelly identify dense-block detection with dense-subgraph mining, by modeling a tensor into a weighted graph without any density information lost. Based on the weighted graph, which we call information sharing graph (ISG), we propose an algorithm for finding multiple densest subgraphs, D-Spot, that is faster (up to 11x faster than the state-of-the-art algorithm) and can be computed in parallel. In an N-dimensional tensor, the entity group found by the ISG+D-Spot is at least 1/2 of the optimum with respect to density, compared with the 1/N guarantee ensured by competing methods. We use nine datasets to demonstrate that ISG+D-Spot becomes new state-of-the-art dense-block detection method in terms of accuracy specifically for fraud detection. Yikun Ban, Ling Huang 0001, Yitao Duan, Xue (Steve) Liu, Wei Xu 0005 |
WWW | 3 |
| 2018 | FraudVis: Understanding Unsupervised Fraud Detection AlgorithmsabstractDiscovering fraud user behaviors is vital to keeping online websites healthy. Fraudsters usually exhibit grouping behaviors, and researchers have effectively leveraged this behavior to design unsupervised algorithms to detect fraud user groups. In this work, we propose a visualization system, FraudVis, to visually analyze the unsupervised fraud detection algorithms from temporal, intra-group correlation, inter-group correlation, feature selection, and the individual user perspectives. FraudVis helps domain experts better understand the algorithm output and the detected fraud behaviors. Meanwhile, FraudVis also helps algorithm experts to fine-tune the algorithm design through the visual comparison. By using the visualization system, we solve two real-world cases of fraud detection, one for a social video website and another for an e-commerce website. The results on both cases demonstrate the effectiveness of FraudVis in understanding unsupervised fraud detection algorithms. Jiao Sun, Qixin Zhu, Zhifei Liu, Jihae Lee, Zhigang Su, Lei Shi 0002, Ling Huang 0001, Wei Xu 0005 |
PacificVis | 8 |
| 2018 | Do Not Pull My Data for Resale: Protecting Data Providers Using Data Retrieval Pattern AnalysisabstractData providers have a profound contribution to many fields such as finance, economy, and academia by serving people with both web-based and API-based query service of specialized data. Among the data users, there are data resellers who abuse the query APIs to retrieve and resell the data to make a profit, which harms the data provider's interests and causes copyright infringement. In this work, we define the "anti-data-reselling" problem and propose a new systematic method that combines feature engineering and machine learning models to provide a solution. We apply our method to a real query log of over 9,000 users with limited labels provided by a large financial data provider and get reasonable results, insightful observations, and real deployments. Guosai Wang, Shiyang Xiang, Yitao Duan, Ling Huang 0001, Wei Xu 0005 |
SIGIR | 4 |
| 2016 | Reviewer Integration and Performance Measurement for Malware Detection
Brad Miller 0002, Alex Kantchelian, Michael Carl Tschantz, Sadia Afroz 0001, Rekha Bachwani, Riyaz Faizullabhoy, Ling Huang 0001, Vaishaal Shankar, Tony Wu 0002, George Yiu, Anthony D. Joseph, J. D. Tygar |
DIMVA | 7 |
| 2016 | CloudKeyBank: Privacy and owner authorization enforced key management frameworkabstractOutsourcing keys (including passwords and data encryption keys) to professional password managers (honest-butcurious service providers) is attracting more and more attention from the researchers and users in the era of cloud computing. However, existing solutions in traditional data outsourcing scenario are unable to simultaneously meet the following three security requirements for keys outsourcing: (1) Confidentiality and privacy of keys; (2) Search privacy on identity attributes tied to keys; (3) Owner controllable authorization over his/her shared keys. In this paper, we propose CloudKeyBank, the first unified key management framework that addresses all the three goals above. To implement CloudKeyBank efficiently, we propose a new cryptographic primitive named Searchable Conditional Proxy Re-Encryption (SC-PRE) which combines the techniques of Hidden Vector Encryption (HVE) and Proxy Re-Encryption (PRE) seamlessly. Xiuxia Tian, Ling Huang 0001, Tony Wu 0002, Xiaoling Wang 0004, Aoying Zhou |
ICDE | 2 |
| 2015 | DualAcE: fine-grained dual access control enforcement with multi-privacy guarantee in DaaSabstractAbstract Database as a service (DaaS), a new paradigm of software as a service based on cloud computing, is attracting more and more enterprises (data owners) to delegate their database management to a professional third party (database service provider) such as Amazon Web Services and Rackspace. Data owners in DaaS lose control of their sensitive data, which are stored in the delegated database and managed by the untrusted database service provider. Therefore, many encryption‐based approaches including attribute‐based encryption were proposed to implement fine‐grained access control in DaaS scenarios. However, most of the proposed access control enforcement approaches only support one or two of the following privacy guarantees: data privacy, policy privacy and key privacy. In this paper, we first propose a novel concept of DualAcE: a flexible fine‐grained dual access control enforcement mechanism in DaaS by efficiently combining the ciphertext‐policy attribute‐set‐based encryption with database service provider re‐encryption into a DaaS paradigm. The proposed mechanism has implemented dual access control enforcement with multi‐privacy guarantee: data privacy in delegated database, policy privacy in delegated authorization table and key privacy in key distribution process.We describe the security and efficiency analysis through cryptography theory and experimental results. Copyright © 2014 John Wiley & Sons, Ltd. Xiuxia Tian, Ling Huang 0001, Chaofeng Sha, Xiaoling Wang 0004 |
Secur. Commun. Networks | 2 |
| 2015 | What You Submit Is Who You Are: A Multimodal Approach for Deanonymizing Scientific PublicationsabstractThe peer-review system of most academic conferences relies on the anonymity of both the authors and reviewers of submissions. In particular, with respect to the authors, the anonymity requirement is heavily disputed and pros and cons are discussed exclusively on a qualitative level. In this paper, we contribute a quantitative argument to this discussion by showing that it is possible for a machine to reveal the identity of authors of scientific publications with high accuracy. We attack the anonymity of authors using statistical analysis of multiple heterogeneous aspects of a paper, such as its citations, its writing style, and its content. We apply several multilabel, multiclass machine learning methods to model the patterns exhibited in each feature category for individual authors and combine them to a single ensemble classifier to deanonymize authors with high accuracy. To the best of our knowledge, this is the first approach that exploits multiple categories of discriminative features and uses multiple, partially complementing classifiers in a single, focused attack on the anonymity of the authors of an academic publication. We evaluate our author identification framework, deAnon, based on a real-world data set of 3894 papers. From these papers, we target 1405 productive authors that each have at least three publications in our data set. Our approach returns a ranking of probable authors for anonymous papers, an ordering for guessing the authors of a paper. In our experiments, following this ranking, the first guess corresponds to one of the authors of a paper in 39.7% of the cases, and at least one of the authors is among the top 10 guesses in 65.6% of all cases. Thus, deAnon significantly outperforms current state-of-the-art techniques for automatic deanonymization. Mathias Payer, Ling Huang 0001, Neil Zhenqiang Gong, Kevin Borgolte, Mario Frank 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | CloudKeyBank: Privacy and Owner Authorization Enforced Key Management FrameworkabstractExplosive growth in the number of passwords for Web based applications and encryption keys for outsourced data storage well exceeds the management limit of users. Therefore, outsourcing keys (including passwords and data encryption keys) to professional password managers (honest-but-curious service providers) is attracting the attention of many users. However, existing solutions in a traditional data outsourcing scenario are unable to simultaneously meet the following three security requirements for keys outsourcing: (1) Confidentiality and privacy of keys; (2) Search privacy on identity attributes tied to keys; (3) Owner controllable authorization over his/her shared keys. In this paper, we propose CloudKeyBank, the first unified key management framework that addresses all the three goals above. Under our framework, the key owner can perform privacy and controllable authorization enforced encryption with minimum information leakage. To implement CloudKeyBank efficiently, we propose a new cryptographic primitive named Searchable Conditional Proxy Re-Encryption (SC-PRE) which combines the techniques of Hidden Vector Encryption (HVE) and Proxy Re-Encryption (PRE) seamlessly, and propose a concrete SC-PRE scheme based on existing HVE and PRE schemes. Our experimental results and security analysis show the efficiency and security goals are well achieved. Xiuxia Tian, Ling Huang 0001, Tony Wu 0002, Xiaoling Wang 0004, Aoying Zhou |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile DevicesabstractWe present Mantis, a framework for predicting the computational resource consumption (CRC) of Android applications on given inputs accurately, and efficiently. A key insight underlying Mantis is that program codes often contain features that correlate with performance and these features can be automatically computed efficiently. Mantis synergistically combines techniques from program analysis and machine learning. It constructs concise CRC models by choosing from many program execution features only a handful that are most correlated with the program's CRC metric yet can be evaluated efficiently from the program's input. We apply program slicing to reduce evaluation time of a feature and automatically generate executable code snippets for efficiently evaluating features. Our evaluation shows that Mantis predicts four CRC metrics of seven Android apps with estimation error in the range of 0-11.1 percent by executing predictor code spending at most 1.3 percent of their execution time on Galaxy Nexus. Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
IEEE Trans. Mob. Comput. | 7 |
| 2014 | Large-Margin Convex Polytope Machine
Alex Kantchelian, Michael Carl Tschantz, Ling Huang 0001, Peter L. Bartlett, Anthony D. Joseph, J. D. Tygar |
NIPS | 3 |
| 2014 | I Know Why You Went to the Clinic: Risks and Realization of HTTPS Traffic Analysis
Brad Miller 0002, Ling Huang 0001, Anthony D. Joseph, J. D. Tygar |
Privacy Enhancing Technologies | 2 |
| 2014 | Joint Link Prediction and Attribute Inference Using a Social-Attribute NetworkabstractThe effects of social influence and homophily suggest that both network structure and node-attribute information should inform the tasks of link prediction and node-attribute inference. Recently, Yin et al. [2010a, 2010b] proposed an attribute-augmented social network model, which we callSocial-Attribute Network(SAN), to integrate network structure and node attributes to perform both link prediction and attribute inference. They focused on generalizing the random walk with a restart algorithm to the SAN framework and showed improved performance. In this article, we extend the SAN framework with several leading supervised and unsupervised link-prediction algorithms and demonstrate performance improvement for each algorithm on both link prediction and attribute inference. Moreover, we make the novel observation that attribute inference can help inform link prediction, that is, link-prediction accuracy is further improved by first inferring missing attributes. We comprehensively evaluate these algorithms and compare them with other existing algorithms using a novel, large-scale Google+ dataset, which we make publicly available (http://www.cs.berkeley.edu/~stevgong/gplus.html). Neil Zhenqiang Gong, Ameet Talwalkar, Lester Mackey, Ling Huang 0001, Richard Shin, Emil Stefanov, Elaine Shi, Dawn Song |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2013 | Mantis: Automatic Performance Prediction for Smartphone Applications
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
USENIX ATC | 7 |
| 2012 | Juxtapp: A Scalable System for Detecting Code Reuse among Android Applications
Steve Hanna, Ling Huang 0001, Edward XueJun Wu, Saung Li, Dawn Song |
DIMVA | 2 |
| 2012 | Evolution of social-attribute networks: measurements, modeling, and implications using google+abstractUnderstanding social network structure and evolution has important implications for many aspects of network and system design including provisioning, bootstrapping trust and reputation systems via social networks, and defenses against Sybil attacks. Several recent results suggest that augmenting the social network structure with user attributes (e.g., location, employer, communities of interest) can provide a more fine-grained understanding of social networks. However, there have been few studies to provide a systematic understanding of these effects at scale. Neil Zhenqiang Gong, Wenchang Xu, Ling Huang 0001, Prateek Mittal, Emil Stefanov, Vyas Sekar, Dawn Song |
Internet Measurement Conference | 3 |
| 2012 | Query Strategies for Evading Convex-Inducing Classifiers
Blaine Nelson, Benjamin I. P. Rubinstein, Ling Huang 0001, Anthony D. Joseph, Steven J. Lee, Satish Rao, J. D. Tygar |
J. Mach. Learn. Res. | 3 |
| 2010 | An Analysis of the Convergence of Graph Laplacians
Daniel Ting, Ling Huang 0001, Michael I. Jordan |
ICML | 2 |
| 2010 | Detecting Large-Scale System Problems by Mining Console Logs
Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan |
ICML | 2 |
| 2010 | Predicting Execution Time of Computer Programs Using Sparse Polynomial RegressionabstractPredicting the execution time of computer programs is an important but challenging problem in the community of computer systems. Existing methods require experts to perform detailed analysis of program code in order to construct predictors or select important features. We recently developed a new system to automatically extract a large number of features from program execution on sample inputs, on which prediction models can be constructed without expert knowledge. In this paper we study the construction of predictive models for this problem. We propose the SPORE (Sparse POlynomial REgression) methodology to build accurate prediction models of program performance using feature data collected from program execution on sample inputs. Our two SPORE algorithms are able to build relationships between responses (e.g., the execution time of a computer program) and features, and select a few from hundreds of the retrieved features to construct an explicitly sparse and non-linear model to predict the response variable. The compact and explicitly polynomial form of the estimated model could reveal important insights into the computer program (e.g., features and their non-linear combinations that dominate the execution time), enabling a better understanding of the program’s behavior. Our evaluation on three widely used computer programs shows that SPORE methods can give accurate prediction with relative error less than 7% by using a moderate number of training data samples. In addition, we compare SPORE algorithms to state-of-the-art sparse regression algorithms, and show that SPORE methods, motivated by real applications, outperform the other methods in terms of both interpretability and prediction accuracy. Ling Huang 0001, Jinzhu Jia, Bin Yu 0001, Byung-Gon Chun, Petros Maniatis, Mayur Naik |
NIPS | 1 |
| 2009 | Online System Problem Detection by Mining Patterns of Console LogsabstractWe describe a novel application of using data mining and statistical learning methods to automatically monitor and detect abnormal execution traces from console logs in an online setting. Different from existing solutions, we use a two stage detection system. The first stage uses frequent pattern mining and distribution estimation techniques to capture the dominant patterns (both frequent sequences and time duration). The second stage use principal component analysis based anomaly detection technique to identify actual problems. Using real system data from a 203-node Hadoop cluster, we show that we can not only achieve highly accurate and fast problem detection, but also help operators better understand execution patterns in their system. Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan |
ICDM | 2 |
| 2009 | ANTIDOTE: understanding and defending against poisoning of anomaly detectorsabstractStatistical machine learning techniques have recently garnered increased popularity as a means to improve network design and security. For intrusion detection, such methods build a model for normal behavior from training data and detect attacks as deviations from that model. This process invites adversaries to manipulate the training data so that the learned model fails to detect subsequent attacks. Benjamin I. P. Rubinstein, Blaine Nelson, Ling Huang 0001, Anthony D. Joseph, Shing-hon Lau, Satish Rao, Nina Taft, J. D. Tygar |
Internet Measurement Conference | 3 |
| 2009 | Fast approximate spectral clusteringabstractSpectral clustering refers to a flexible class of clustering procedures that can produce high-quality clusterings on small data sets but which has limited applicability to large-scale problems due to its computational complexity of O(n^3), with n the number of data points. We extend the range of spectral clustering by developing a general framework for fast approximate spectral clustering in which a distortion-minimizing local transformation is first applied to the data. This framework is based on a theoretical analysis that provides a statistical characterization of the effect of local distortion on the mis-clustering rate. We develop two concrete instances of our general framework, one based on local k-means clustering (KASP) and one based on random projection trees (RASP). Extensive experiments show that these algorithms can achieve significant speedups with little degradation in clustering accuracy. Specifically, our algorithms outperform k-means by a large margin in terms of accuracy, and run several times faster than approximate spectral clustering based on the Nystrom method, with comparable accuracy and significantly smaller memory footprint. Remarkably, our algorithms make it possible for a single machine to spectral cluster data sets with a million observations within several minutes. Donghui Yan, Ling Huang 0001, Michael I. Jordan |
KDD | 2 |
| 2009 | Detecting large-scale system problems by mining console logsabstractSurprisingly, console logs rarely help operators detect problems in large-scale datacenter services, for they often consist of the voluminous intermixing of messages from many software components written by independent developers. We propose a general methodology to mine this rich source of information to automatically detect system runtime problems. We first parse console logs by combining source code analysis with information retrieval to create composite features. We then analyze these features using machine learning to detect operational problems. We show that our method enables analyses that are impossible with previous methods because of its superior ability to create sophisticated features. We also show how to distill the results of our analysis to an operator-friendly one-page decision tree showing the critical messages associated with the detected problems. We validate our approach using the Darkstar online game server and the Hadoop File System, where we detect numerous real problems with high accuracy and few false positives. In the Hadoop case, we are able to analyze 24 million lines of console logs in 3 minutes. Our methodology works on textual console logs of any size and requires no changes to the service software, no human input, and no knowledge of the software's internals. Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan |
SOSP | 2 |
| 2008 | Spectral Clustering with Perturbed DataabstractSpectral clustering is useful for a wide-ranging set of applications in areas such as biological data analysis, image processing and data mining. However, the computational and/or communication resources required by the method in processing large-scale data sets are often prohibitively high, and practitioners are often required to perturb the original data in various ways (quantization, downsampling, etc) before invoking a spectral algorithm. In this paper, we use stochastic perturbation theory to study the effects of data perturbation on the performance of spectral clustering. We show that the error under perturbation of spectral clustering is closely related to the perturbation of the eigenvectors of the Laplacian matrix. From this result we derive approximate upper bounds on the clustering error. We show that this bound is tight empirically across a wide range of problems, suggesting that it can be used in practical settings to determine the amount of data reduction allowed in order to meet a specification of permitted loss in clustering performance. Ling Huang 0001, Donghui Yan, Michael I. Jordan, Nina Taft |
NIPS | 1 |
| 2008 | Support Vector Machines, Data Reduction, and Approximate Kernel Matrices
XuanLong Nguyen, Ling Huang 0001, Anthony D. Joseph |
ECML/PKDD (2) | 2 |
| 2008 | Evading Anomaly Detection through Variance Injection Attacks on PCA
Benjamin I. P. Rubinstein, Blaine Nelson, Ling Huang 0001, Anthony D. Joseph, Shing-hon Lau, Nina Taft, J. D. Tygar |
RAID | 3 |
| 2007 | Communication-Efficient Tracking of Distributed Cumulative TriggersabstractIn recent work, we proposed D-Trigger, a framework for tracking a global condition over a large network that allows us to detect anomalies while only collecting a very limited amount of data from distributed monitors. In this paper, we expand our previous work by designing a new class of queries (conditions) that can be tracked for anomaly violations. We show how security violations can be detected over a time window of any size. This is important because security operators do not know in advance the window of time in which measurements should be made to detect anomalies. We also present an algorithm that determines how each machine should filter its time series measurements before back-hauling them to a central operations center. Our filters are computed analytically such that upper bounds on false positive and missed detection rates are guaranteed. In our evaluation, we show that botnet detection can be carried out successfully over a distributed set of machines, while simultaneously filtering out 80 to 90% of the measurement data. Ling Huang 0001, Minos N. Garofalakis, Anthony D. Joseph, Nina Taft |
ICDCS | 1 |
| 2007 | Communication-Efficient Online Detection of Network-Wide AnomaliesabstractThere has been growing interest in building large-scale distributed monitoring systems for sensor, enterprise, and ISP networks. Recent work has proposed using principal component analysis (PCA) over global traffic matrix statistics to effectively isolate network-wide anomalies. To allow such a PCA-based anomaly detection scheme to scale, we propose a novel approximation scheme that dramatically reduces the burden on the production network. Our scheme avoids the expensive step of centralizing all the data by performing intelligent filtering at the distributed monitors. This filtering reduces monitoring bandwidth overheads, but can result in the anomaly detector making incorrect decisions based on a perturbed view of the global data set. We employ stochastic matrix perturbation theory to bound such errors. Our algorithm selects the filtering parameters at local monitors such that the errors made by the detector are guaranteed to lie below a user-specified upper bound. Our algorithm thus allows network operators to explicitly balance the tradeoff between detection accuracy and the amount of data communicated over the network. In addition, our approach enables real-time detection because we exploit continuous monitoring at the distributed monitors. Experiments with traffic data from Abilene backbone network demonstrate that our methods yield significant communication benefits while simultaneously achieving high detection accuracy. Ling Huang 0001, XuanLong Nguyen, Minos N. Garofalakis, Joseph M. Hellerstein, Michael I. Jordan, Anthony D. Joseph, Nina Taft |
INFOCOM | 1 |
| 2006 | In-Network PCA and Anomaly DetectionabstractWe consider the problem of network anomaly detection in large distributed systems. In this setting, Principal Component Analysis (PCA) has been proposed as a method for discover- ing anomalies by continuously tracking the projection of the data onto a residual subspace. This method was shown to work well empirically in highly aggregated networks, that is, those with a limited number of large nodes and at coarse time scales. This approach, how- ever, has scalability limitations. To overcome these limitations, we develop a PCA-based anomaly detector in which adaptive local data (cid:2)lters send to a coordinator just enough data to enable accurate global detection. Our method is based on a stochastic matrix perturba- tion analysis that characterizes the tradeoff between the accuracy of anomaly detection and the amount of data communicated over the network. Ling Huang 0001, XuanLong Nguyen, Minos N. Garofalakis, Michael I. Jordan, Anthony D. Joseph, Nina Taft |
NIPS | 1 |
| 2004 | Tapestry: a resilient global-scale overlay for service deploymentabstractWe present Tapestry, a peer-to-peer overlay routing infrastructure offering efficient, scalable, location-independent routing of messages directly to nearby copies of an object or service using only localized resources. Tapestry supports a generic decentralized object location and routing applications programming interface using a self-repairing, soft-state-based routing layer. The paper presents the Tapestry architecture, algorithms, and implementation. It explores the behavior of a Tapestry deployment on PlanetLab, a global testbed of approximately 100 machines. Experimental results show that Tapestry exhibits stable behavior and performance as an overlay, despite the instability of the underlying network layers. Several widely distributed applications have been implemented on Tapestry, illustrating its utility as a deployment infrastructure. Ben Y. Zhao, Ling Huang 0001, Jeremy Stribling, Sean C. Rhea, Anthony D. Joseph, John Kubiatowicz |
IEEE J. Sel. Areas Commun. | 2 |
| 2003 | Exploiting Routing Redundancy via Structured Peer-to-Peer OverlaysabstractStructured peer-to-peer overlays provide a natural infrastructure for resilient routing via efficient fault detection and precomputation of backup paths. These overlays can respond to faults in a few hundred milliseconds by rapidly shifting between alternate routes. In this paper, we present two adaptive mechanisms for structured overlays and illustrate their operation in the context of Tapestry, a fault-resilient overlay from Berkeley. We also describe a transparent, protocol-independent traffic redirection mechanism that tunnels legacy application traffic through overlays. Our measurements of a Tapestry prototype show it to be a highly responsive routing service, effective at circumventing a range of failures while incurring reasonable cost in maintenance bandwidth and additional routing latency. Ben Y. Zhao, Ling Huang 0001, Jeremy Stribling, Anthony D. Joseph, John Kubiatowicz |
ICNP | 2 |
| 2003 | Approximate Object Location and Spam Filtering on Peer-to-Peer Systems
Li Zhuang, Ben Y. Zhao, Ling Huang 0001, Anthony D. Joseph, John Kubiatowicz |
Middleware | 4 |