VLDB 2026 Research / reviewers in the wild / expert
Han Hee Song
dblp:52/2373
· DBLP profile ↗
20ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 56% Web and social media mining · 29% Knowledge graphs · 13% | |
| Computer networks
8 papers |
Network measurement and analytics · 48% Network management and operations · 13% Vehicular, aerial and satellite networks · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 100% | |
| Network and information security
1 paper |
Privacy and data protection · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia systems and quality of experience · 100% |
Topics — the 22 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Medical and health informatics › telemedicine
remote patient monitoring |
0.4 | 1 | 2019 | Developing Measures of Cognitive Impairment in the Real World from Consumer-Grade Multimodal Sensor Streams · KDD 2019 |
Web and social media mining
app categorization |
0.2 | 1 | 2016 | Macro-scale mobile app market analysis using customized hierarchical categorization · INFOCOM 2016 |
Data mining › text mining › text classification
hierarchical classification |
0.2 | 1 | 2016 | Macro-scale mobile app market analysis using customized hierarchical categorization · INFOCOM 2016 |
Data mining › text mining
text classification |
0.2 | 1 | 2016 | Macro-scale mobile app market analysis using customized hierarchical categorization · INFOCOM 2016 |
Data mining
distance estimation |
0.2 | 2 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 Scalable proximity estimation and link prediction in online social networks · Internet Measurement Conference 2009 |
Web and social media mining
social network analysis |
0.2 | 2 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 Scalable proximity estimation and link prediction in online social networks · Internet Measurement Conference 2009 |
Network measurement and analytics
traffic analysis |
0.2 | 1 | 2013 | Mosaic: quantifying privacy leakage in mobile networks · SIGCOMM 2013 |
Privacy and data protection › information leakage
privacy leakage |
0.2 | 1 | 2013 | Mosaic: quantifying privacy leakage in mobile networks · SIGCOMM 2013 |
Data mining › structured data mining
graph mining |
0.1 | 1 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 |
Knowledge graphs
link prediction |
0.1 | 1 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 |
Multimedia systems and quality of experience › quality of experience
qoe measurement |
0.1 | 1 | 2011 | Q-score: proactive service quality assessment in a large IPTV system · Internet Measurement Conference 2011 |
Vehicular, aerial and satellite networks › data delivery
vehicular content distribution |
0.1 | 1 | 2010 | Enabling high-bandwidth vehicular content distribution · CoNEXT 2010 |
Routing and switching › routing
delay-tolerant network routing |
0.1 | 1 | 2008 | Incentive-aware routing in DTNs · ICNP 2008 |
Knowledge graphs
ontology |
0.1 | 1 | 2016 | Macro-scale mobile app market analysis using customized hierarchical categorization · INFOCOM 2016 |
Network measurement and analytics › latency measurement
latency estimation |
0.0 | 1 | 2004 | An algebraic approach to practical and scalable overlay network monitoring · SIGCOMM 2004 |
Physical-layer communications › channel modeling › path loss modeling
path loss estimation |
0.0 | 1 | 2004 | An algebraic approach to practical and scalable overlay network monitoring · SIGCOMM 2004 |
Graph data management
graph embedding |
0.0 | 1 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 |
Data mining › dimensionality reduction
spectral embedding |
0.0 | 1 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 |
Network measurement and analytics › network performance measurement
quality of service monitoring |
0.0 | 1 | 2011 | Q-score: proactive service quality assessment in a large IPTV system · Internet Measurement Conference 2011 |
Data mining › structured data mining › graph mining
scalable graph mining |
0.0 | 1 | 2009 | Scalable proximity estimation and link prediction in online social networks · Internet Measurement Conference 2009 |
Network measurement and analytics
measurement infrastructure |
0.0 | 1 | 2009 | NetQuest: a flexible framework for large-scale network measurement · IEEE/ACM Trans. Netw. 2009 |
Graph algorithms and graph theory › network analysis
link prediction |
0.0 | 1 | 2009 | Scalable proximity estimation and link prediction in online social networks · Internet Measurement Conference 2009 |
Methods — techniques the papers use, named apart from their topics
time alignment · 0.8imputation · 0.8feature engineering · 0.8traffic analysis · 0.5proactive monitoring · 0.2performance indicator selection · 0.2hierarchical classification · 0.2customized class hierarchy induction · 0.2incremental update algorithm · 0.2approximation · 0.2dimensionality reduction · 0.1clustered spectral graph embedding · 0.1measurement · 0.1tit-for-tat · 0.1game theory · 0.1algorithm evaluation · 0.1algebraic modeling · 0.1linear independence · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | NEMA: Automatic Integration of Large Network Management DatabasesabstractNetwork management, whether for malfunction analysis, failure prediction, performance monitoring and improvement, generally involves large amounts of data from different sources. To effectively integrate and manage these sources, automatically finding semantic matches among their schemas or ontologies is crucial. Existing approaches on database matching mainly fall into two categories. One focuses on the schema-level matching based on schema properties such as field names, data types, constraints and schema structures. Network management databases contain massive tables (e.g., network products, incidents, security alert and logs) from different departments and groups with nonuniform field names and schema characteristics. It is not reliable to match them by those schema properties. The other category is based on the instance-level matching using general string similarity techniques, which are not applicable for the matching of large network management databases. In this article, we develop a matching technique for large NEtwork MAnagement databases (NEMA) deploying instance-level matching for effective data integration and connection. We design matching metrics and scores for both numerical and non-numerical fields and propose algorithms for matching these fields. The effectiveness and efficiency of NEMA are evaluated by conducting experiments based on ground truth field pairs in large network management databases. Our measurement on large databases with 1,458 fields, each of which contains over 10 million records, reveals that NEMA can achieve accuracy of 95%. We further compare with several other existing algorithms, and show that NEMA outperforms them by 7%-15% in numerical matching and achieves the best trade-off for non-numerical matching. Fubao Wu, Han Hee Song, Jiangtao Yin, Lixin Gao 0001, Mario Baldi, Narendra Anand |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Developing Measures of Cognitive Impairment in the Real World from Consumer-Grade Multimodal Sensor StreamsabstractThe ubiquity and remarkable technological progress of wearable consumer devices and mobile-computing platforms (smart phone, smart watch, tablet), along with the multitude of sensor modalities available, have enabled continuous monitoring of patients and their daily activities. Such rich, longitudinal information can be mined for physiological and behavioral signatures of cognitive impairment and provide new avenues for detecting MCI in a timely and cost-effective manner. In this work, we present a platform for remote and unobtrusive monitoring of symptoms related to cognitive impairment using several consumer-grade smart devices. We demonstrate how the platform has been used to collect a total of 16TB of data during the Lilly Exploratory Digital Assessment Study, a 12-week feasibility study which monitored 31 people with cognitive impairment and 82 without cognitive impairment in free living conditions. We describe how careful data unification, time-alignment, and imputation techniques can handle missing data rates inherent in real-world settings and ultimately show utility of these disparate data in differentiating symptomatics from healthy controls based on features computed purely from device data. Richard J. Chen, Filip Jankovic, Nikki Marinsek, Luca Foschini 0002, Lampros Kourtis, Alessio Signorini, Melissa Pugh, Roy Yaari, Vera Maljkovic, Marc Sunga, Han Hee Song, Hyun Joon Jung, Belle L. Tseng, Andrew Trister |
KDD | 12 |
| 2018 | AWESoME: Big Data for Automatic Web Service Management in SDNabstractSoftware defined network (SDN) has enabled consistent and programmable management in computer networks. However, the explosion of cloud services and content delivery networks (CDNs)-coupled with the momentum of encryption-challenges the simple per-flow management and calls for a more comprehensive approach for managing Web traffic. We propose a new approach based on a “per service” management concept, which allows to identify and prioritize all traffic of important Web services, while segregating others, even if they are running on the same cloud platform, or served by the same CDN. We design and evaluate AWESoME, automatic Web service manager, a novel SDN application to address the above problem. On the one hand, it leverages big data algorithms to automatically build models describing the traffic of thousands of Web services. On the other hand, it uses the models to install rules in SDN switches to steer all flows related to the originating services. Using traffic traces from volunteers and operational networks, we provide extensive experimental results to show that AWESoME associates flows to the corresponding Web service in real-time and with high accuracy. AWESoME introduces a negligible load on the SDN controller and installs a limited number of rules on switches, hence scaling well in realistic deployments. Finally, for easy reproducibility, we release ground truth traces and scripts implementing AWESoME core components. Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2017 | On the most representative summaries of network user activities
Joshua Stein, Han Hee Song, Mario Baldi, Jun Li 0001 |
Comput. Networks | 2 |
| 2016 | Joining user profiles across online social networks: From the perspective of an adversaryabstractBeing the anchor points for building social relationships in the cyberspace, online social networks (OSNs) play an integral part of modern peoples life. Since different OSNs are designed to address specific social needs, people take part in multiple OSNs to cover different facets of their life. While the fragmented pieces of information about a user in each OSN may be of limited use, serious privacy issues arise if a sophisticated adversary pieces information together from multiple OSNs. To this end, we undertake the role of such an adversary and demonstrate the possibility of “splicing” user profiles across multiple OSNs and present associated security risks to users. In doing so, we develop a scalable and systematic profile joining scheme, Splicer, that focuses on various aspects of profile attributes by simultaneously performing exact, quasi-perfect and partial matches between pairs of profiles. From our evaluations on three real OSN data, Splicer not only handles large-scale OSN profiles efficiently by saving 87% computation time compared to all-pair profile comparisons, but also far exceeds the recall of generic distance measure based approach at the same precision level by 33%. Finally, we quantify the amount of information “lift” attributed to joining of OSNs, where on average 22% additional profile attributes can be added to 24% of users. Qiang Ma 0006, Han Hee Song, S. Muthukrishnan 0001, Antonio Nucci |
ASONAM | 2 |
| 2016 | WHAT: A big data approach for accounting of modern web servicesabstractHTTP(S) has become the main means to access the Internet. The web is a tangle, with (i) multiple services and applications co-located on the same infrastructure and (ii) several websites, services and applications embedding objects from CDN, ads and tracking platforms. Traditional solutions for traffic classification and metering fall short in providing visibility in users' activities. Service providers and corporate network administrators are left with huge amounts of measurements, which cannot immediately reveal the real impact of each web service on the network. Such visibility is key to dimension the network, charge users and policy traffic. This paper introduces the Web Helper Accounting Tool (WHAT), a system to uncover the overall traffic produced by specific web services. WHAT combines big data and machine learning approaches to process large volumes of network flow measurements and learn how to group traffic due to pre-defined services of interest. Our evaluation demonstrates WHAT effectiveness in enabling accurate accounting of the traffic associated to each service. WHAT illustrates the power of machine learning when applied to large datasets of network measurements, and allows network administrators to regain the lost visibility on network usage. Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi |
IEEE BigData | 4 |
| 2016 | Macro-scale mobile app market analysis using customized hierarchical categorizationabstractThanks to the widespread use of smart devices, recent years have witnessed the proliferation of mobile apps available on online stores such as Apple iTunes and Google Play. As the number of new mobile apps continues to grow at a rapid pace, automatic classification of the apps has become an increasingly important problem to facilitate browsing, searching, and recommending them. This paper presents a framework that automatically labels apps with a richer and more detailed categorization and uses the labeled apps to study the app market. Leveraging a fine-grained, hierarchical ontology as a guide, we developed a framework not only to label the apps with fine-grained categorical information but also to induce a customized class hierarchy optimized for mobile app classification. With the classification accuracy of 93%, large-scale categorization conducted with our framework on 168,000 Google Play apps discovers novel inter-class relationships among categories of Google Play market. Han Hee Song, Mario Baldi, Pang-Ning Tan |
INFOCOM | 2 |
| 2014 | Toward the most representative summaries of network user activitiesabstractA summary of a user's Internet activities, such as web visits, can closely reflect their interests and preferences. However, automating the summarization process is not trivial as it should strike a good balance between generality and specificity, while there is no gold standard for doing so. In our approach to summarizing user information, we introduce two scoring mechanisms that cooperatively optimize for polarizing criteria. Having mapped user activity information onto a category tree, the scoring mechanisms highlight the most representative tree node; the node provides an aggregated view, i.e., a summary, of the activities most representative of the user. We evaluate our approach by summarizing web activity on the network of a large cellular service provider to devise interests of individual users as well as user groups. Joshua Stein, Han Hee Song, Mario Baldi, Jun Li 0001 |
IWQoS | 2 |
| 2014 | On Understanding User Interests through Heterogeneous Data Sources
Samamon Khemmarat, Sabyasachi Saha, Han Hee Song, Mario Baldi, Lixin Gao 0001 |
PAM | 3 |
| 2013 | TUCAN: Twitter user centric ANalyzerabstractTwitter has attracted millions of users that generate a humongous flow of information at constant pace. The research community has thus started proposing tools to extract meaningful information from tweets. In this paper, we take a different angle from the mainstream of previous works: we explicitly target the analysis of the timeline of tweets from "single users". We define a framework - named TUCAN - to compare information offered by the target users over time, and to pinpoint recurrent topics or topics of interest. First, tweets belonging to the same time window are aggregated into "bird songs". Several filtering procedures can be selected to remove stop-words and reduce noise. Then, each pair of bird songs is compared using a similarity score to automatically highlight the most common terms, thus highlighting recurrent or persistent topics. TUCAN can be naturally applied to compare bird song pairs generated from timelines of different users. Luigi Grimaudo, Han Hee Song, Mario Baldi, Marco Mellia, Maurizio M. Munafò |
ASONAM | 2 |
| 2013 | Mosaic: quantifying privacy leakage in mobile networksabstractWith the proliferation of online social networking (OSN) and mobile devices, preserving user privacy has become a great challenge. While prior studies have directly focused on OSN services, we call attention to the privacy leakage in mobile network data. This concern is motivated by two factors. First, the prevalence of OSN usage leaves identifiable digital footprints that can be traced back to users in the real-world. Second, the association between users and their mobile devices makes it easier to associate traffic to its owners. These pose a serious threat to user privacy as they enable an adversary to attribute significant portions of data traffic including the ones with NO identity leaks to network users' true identities. To demonstrate its feasibility, we develop the Tessellation methodology. By applying Tessellation on traffic from a cellular service provider (CSP), we show that up to 50% of the traffic can be attributed to the names of users. In addition to revealing the user identity, the reconstructed profile, dubbed as "mosaic," associates personal information such as political views, browsing habits, and favorite apps to the users. We conclude by discussing approaches for preventing and mitigating the alarming leakage of sensitive user information. Ning Xia, Han Hee Song, Marios Iliofotou, Antonio Nucci, Zhi-Li Zhang, Aleksandar Kuzmanovic |
SIGCOMM | 2 |
| 2012 | Clustered embedding of massive social networksabstractThe explosive growth of social networks has created numerous exciting research opportunities. A central concept in the analysis of social networks is a proximity measure, which captures the closeness or similarity between nodes in the network. Despite much research on proximity measures, there is a lack of techniques to efficiently and accurately compute proximity measures for large-scale social networks. In this paper, we embed the original massive social graph into a much smaller graph, using a novel dimensionality reduction technique termed Clustered Spectral Graph Embedding. We show that the embedded graph captures the essential clustering and spectral structure of the original graph and allow a wide range of analysis to be performed on massive social graphs. Applying the clustered embedding to proximity measurement of social networks, we develop accurate, scalable, and flexible solutions to three important social network analysis tasks: proximity estimation, missing link inference, and link prediction. We demonstrate the effectiveness of our solutions to the tasks in the context of large real-world social network datasets: Flickr, LiveJournal, and MySpace with up to 2 million nodes and 90 million links. Han Hee Song, Berkant Savas, Tae Won Cho, Vacha Dave, Zhengdong Lu, Inderjit S. Dhillon, Yin Zhang 0001, Lili Qiu |
SIGMETRICS | 1 |
| 2011 | Q-score: proactive service quality assessment in a large IPTV systemabstractIn large-scale IPTV systems, it is essential to maintain high service quality while providing a wider variety of service features than typical traditional TV. Thus service quality assessment systems are of paramount importance as they monitor the user-perceived service quality and alert when issues occurs. For IPTV systems, however, there is no simple metric to represent user-perceived service quality and Quality of Experience (QoE). Moreover, there is only limited user feedback, often in the form of noisy and delayed customer calls. Therefore, we aim to approximate the QoE through a selected set of performance indicators in a proactive (i.e., detect issues before customers reports to call centers) and scalable fashion. Han Hee Song, Zihui Ge, Ajay Mahimkar, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Andrea Basso 0001, Min Chen 0010 |
Internet Measurement Conference | 1 |
| 2010 | Enabling high-bandwidth vehicular content distributionabstractWe present VCD, a novel system for enabling high-bandwidth content distribution in vehicular networks. In VCD, a vehicle opportunistically communicates with nearby access points (APs) to download the content of interest. To fully take advantage of such transient contact with APs, we proactively push content to the APs that the vehicles will likely visit in the near future. In this way, vehicles can enjoy the full wireless capacity instead of being bottle-necked by the Internet connectivity, which is either slow or even unavailable. We develop a new algorithm for predicting the APs that will soon be visited by the vehicles. We then develop a replication scheme that leverages the synergy among (i) Internet connectivity (which is persistent but has limited coverage and low bandwidth), (ii) local wireless connectivity (which has high bandwidth but transient duration), (iii) vehicular relay connectivity (which has high bandwidth but high delay), and (iv) mesh connectivity among APs (which has high bandwidth but low coverage). We demonstrate the effectiveness of VCD system using trace-driven simulation and Emulab emulation based on real taxi traces. We further deploy VCD in two vehicular networks: one using 802.11b and the other using 802.11n, to demonstrate its effectiveness. Upendra Shevade, Yi-Chao Chen 0001, Lili Qiu, Yin Zhang 0001, Vinoth Chandar, Mi Kyung Han, Han Hee Song, Yousuk Seung |
CoNEXT | 7 |
| 2010 | Detecting the performance impact of upgrades in large operational networksabstractNetworks continue to change to support new applications, improve reliability and performance and reduce the operational cost. The changes are made to the network in the form of upgrades such as software or hardware upgrades, new network or service features and network configuration changes. It is crucial to monitor the network when upgrades are made because they can have a significant impact on network performance and if not monitored may lead to unexpected consequences in operational networks. This can be achieved manually for a small number of devices, but does not scale to large networks with hundreds or thousands of routers and extremely large number of different upgrades made on a regular basis. Ajay Mahimkar, Han Hee Song, Zihui Ge, Aman Shaikh, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Joanne Emmons |
SIGCOMM | 2 |
| 2009 | Scalable proximity estimation and link prediction in online social networksabstractProximity measures quantify the closeness or similarity between nodes in a social network and form the basis of a range of applications in social sciences, business, information technology, computer networks, and cyber security. It is challenging to estimate proximity measures in online social networks due to their massive scale (with millions of users) and dynamic nature (with hundreds of thousands of new nodes and millions of edges added daily). To address this challenge, we develop two novel methods to efficiently and accurately approximate a large family of proximity measures. We also propose a novel incremental update algorithm to enable near real-time proximity estimation in highly dynamic social networks. Evaluation based on a large amount of real data collected in five popular online social networks shows that our methods are accurate and can easily scale to networks with millions of nodes. Han Hee Song, Tae Won Cho, Vacha Dave, Yin Zhang 0001, Lili Qiu |
Internet Measurement Conference | 1 |
| 2009 | NetQuest: a flexible framework for large-scale network measurement
Han Hee Song, Lili Qiu, Yin Zhang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2008 | Incentive-aware routing in DTNsabstractDisruption tolerant networks (DTNs) are a class of networks in which no contemporaneous path may exist between the source and destination at a given time. In such a network, routing takes place with the help of relay nodes and in a store-and-forward fashion. If the nodes in a DTN are controlled by rational entities, such as people or organizations, the nodes can be expected to behave selfishly and attempt to maximize their utilities and conserve their resources. Since routing is an inherently cooperative activity, system operation will be critically impaired unless cooperation is somehow incentivized. The lack of end-to-end paths, high variation in network conditions, and long feedback delay in DTNs imply that existing solutions for mobile ad-hoc networks do not apply to DTNs. In this paper, we propose the use of pair-wise tit-for-tat (TFT) as a simple, robust and practical incentive mechanism for DTNs. Existing TFT mechanisms often face bootstrapping problems or suffer from exploitation. We propose a TFT mechanism that incorporates generosity and contrition to address these issues. We then develop an incentive-aware routing protocol that allows selfish nodes to maximize their own performance while conforming to TFT constraints. For comparison, we also develop techniques to optimize the system-wide performance when all nodes are cooperative. Using both synthetic and real DTN traces, we show that without an incentive mechanism, the delivery ratio among selfish nodes can be as low as 20% as what is achieved under full cooperation; in contrast, with TFT as a basis of cooperation among selfish nodes, the delivery ratio increases to 60% or higher as under full cooperation. We also address the practical challenges involved in implementing the TFT mechanism. To our knowledge, this is the first practical incentive-aware routing scheme for DTNs. Upendra Shevade, Han Hee Song, Lili Qiu, Yin Zhang 0001 |
ICNP | 2 |
| 2007 | Algebra-based scalable overlay network monitoring: algorithms, evaluation, and applications
Yan Chen 0004, David Bindel, Han Hee Song, Randy H. Katz |
IEEE/ACM Trans. Netw. | 3 |
| 2004 | An algebraic approach to practical and scalable overlay network monitoringabstractOverlay network monitoring enables distributed Internet applications to detect and recover from path outages and periods of degraded performance within seconds. For an overlay network with n end hosts, existing systems either require O(n2) measurements, and thus lack scalability, or can only estimate the latency but not congestion or failures. Our earlier extended abstract [1] briefly proposes an algebraic approach that selectively monitors k linearly independent paths that can fully describe all the O(n2) paths. The loss rates and latency of these k paths can be used to estimate the loss rates and latency of all other paths. Our scheme only assumes knowledge of the underlying IP topology, with links dynamically varying between lossy and normal.In this paper, we improve, implement and extensively evaluate such a monitoring system. We further make the following contributions: i) scalability analysis indicating that for reasonably large n (e.g., 100), the growth of k is bounded as O(n log n), ii) efficient adaptation algorithms for topology changes, such as the addition or removal of end hosts and routing changes, iii) measurement load balancing schemes, and iv) topology measurement error handling. Both simulation and Internet experiments demonstrate we obtain highly accurate path loss rate estimation while adapting to topology changes within seconds and handling topology errors. Yan Chen 0004, David Bindel, Han Hee Song, Randy H. Katz |
SIGCOMM | 3 |