VLDB 2026 Research / reviewers in the wild / expert
Yaoping Ruan
dblp:25/3546
· DBLP profile ↗
22ranked-venue papers
5as first author
1since 2021 · last 2026
0000-0002-3929-4195ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-authorComputer networks · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Cloud and datacenter computing · 70% Distributed systems · 13% Processor architecture and microarchitecture · 11% | |
| Software engineering, system software, and programming languages
4 papers |
Operating systems · 51% Software testing · 49% | |
| Computer networks
2 papers |
Network measurement and analytics · 51% Internet architecture and protocols · 40% Transport protocols and congestion control · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Distributed and cloud data management · 100% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 10 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed and cloud data management
mapreduce |
0.1 | 1 | 2011 | Sedic: privacy-aware data intensive computing on hybrid clouds · CCS 2011 |
Cloud and datacenter computing › cloud deployment
hybrid cloud |
0.1 | 1 | 2011 | Sedic: privacy-aware data intensive computing on hybrid clouds · CCS 2011 |
Cloud and datacenter computing
application migration |
0.1 | 1 | 2010 | Splitter: a proxy-based approach for post-migration testing of web applications · EuroSys 2010 |
Cloud and datacenter computing
virtualization |
0.1 | 1 | 2010 | Splitter: a proxy-based approach for post-migration testing of web applications · EuroSys 2010 |
Internet architecture and protocols
domain name system |
0.1 | 1 | 2006 | How DNS Misnaming Distorts Internet Topology Mapping · USENIX ATC, General Track 2006 |
Network measurement and analytics
internet topology mapping |
0.1 | 1 | 2006 | How DNS Misnaming Distorts Internet Topology Mapping · USENIX ATC, General Track 2006 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.1 | 1 | 2005 | Evaluating the impact of simultaneous multithreading on network servers using real hardware · SIGMETRICS 2005 |
Privacy and data protection › privacy-preserving data sharing
privacy-preserving data outsourcing |
0.0 | 1 | 2011 | Sedic: privacy-aware data intensive computing on hybrid clouds · CCS 2011 |
Operating systems
network stack |
0.0 | 1 | 2006 | Understanding and Addressing Blocking-Induced Network Server Latency · USENIX ATC, General Track 2006 |
Performance modeling and evaluation › profiling
microarchitectural profiling |
0.0 | 1 | 2005 | Evaluating the impact of simultaneous multithreading on network servers using real hardware · SIGMETRICS 2005 |
Methods — techniques the papers use, named apart from their topics
reduction structure transformation · 0.4map task scheduling · 0.4data replication · 0.4test case generation · 0.2proxy-based testing · 0.2performance measurement · 0.1kernel modification · 0.1workload characterization · 0.1hardware performance counters · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ForumSeeker: Fusion Retrieval of Online Technical Forums for Effective Troubleshooting
Youyang Kim, Yaoping Ruan, Young-Kyoon Suh, Liqiang Wang 0001, Byung-Chul Tak |
FASE | 2 |
| 2017 | Mapping client messages to a unified data model with mixture feature embedding convolutional neural networkabstractData mapping among different data standards in health institutes is often a necessity when data exchanges occur among different institutes. However, no matter rule-based approaches or traditional machine learning methods, none of these methods have achieved satisfactory results yet. In this work, we propose a deep learning method, mixture feature embedding convolutional neural network (MfeCNN), to convert the data mapping to a multiple classification problem. Multi-modal features were extracted from different semantic space with a medical NLP package and powerful feature embeddings were generated by MfeCNN. Classes as many as ten were classified simultaneously by a fully-connected soft-max layer based on multi-view embedding. Experimental results show that our proposed MfeCNN achieved best results than traditional state-of-the-art machine learning models and also much better results than the convolutional neural network of only using bag-of-words as inputs. Dingcheng Li, Peini Liu, Ming Huang 0006, Yu Gu 0001, Daniel Dean, Jingmin Xu, Hui Lei 0001, Yaoping Ruan |
BIBM | 11 |
| 2017 | Engineering Scalable, Secure, Multi-Tenant Cloud for Healthcare DataabstractCloud-based analytics allow for inexpensive processing of large amount of data. However, processing protected health information (PHI) in cloud is a challenging task due to strict regulations (e.g., HIPAA) requiring features (e.g., data isolation) which most cloud-based platforms do not currently support in their offerings. This makes it difficult to leverage many technologies well suited to the cloud (e.g., Apache Spark)to process PHI. To address this issue, we have developed the Watson Health Cloud (WHC), a cloud-based platform for the storage and analysis of large amount of PHI. The WHC enables all the features necessary to store and process PHI, with little customization needed by the end-user. This paper describes the lessons learned from developing a cloud platform for PHI. Specifically, we discuss the architecture and implementation challenges we faced throughout development. We hope the insights gained from our experiences help others when designing frameworks and applications which process PHI. Daniel Joseph Dean, Rohit Ranchal, Anca Sailer, Shakil Khan, Kirk A. Beaty, Senthil Bakthavachalam, Yichong Yu, Yaoping Ruan, Paul Bastide 0001 |
SERVICES | 9 |
| 2016 | LOGAN: Problem Diagnosis in the Cloud Using Log-Based Reference ModelsabstractProblem diagnosis is one crucial aspect in the cloud operation that is becoming increasingly challenging. On the one hand, the volume of logs generated in today's cloud is overwhelmingly large. On the other hand, cloud architecture becomes more distributed and complex, which makes it more difficult to troubleshoot failures. In order to address these challenges, we have developed a tool, called LOGAN, that enables operators to quickly identify the log entries that potentially lead to the root cause of a problem. It constructs behavioral reference models from logs that represent the normal patterns. When problem occurs, our tool enables operators to inspect the divergence of current logs from the reference model and highlight logs likely to contain the hints to the root cause. To support these capabilities we have designed and developed several mechanisms. First, we developed log correlation algorithms using various IDs embedded in logs to help identify and isolate log entries that belong to the failed request. Second, we provide efficient log comparison to help understand the differences between different executions. Finally we designed mechanisms to highlight critical log entries that are likely to contain information pertaining to the root cause of the problem. We have implemented the proposed approach in a popular cloud management system, OpenStack, and through case studies, we demonstrate this tool can help operators perform problem diagnosis quickly and effectively. Byung-Chul Tak, Shu Tao, Yaoping Ruan |
IC2E | 5 |
| 2016 | SalesExplorer: Exploring sales opportunities from white-space customers in the enterprise market
Dongsheng Li 0002, Yaoping Ruan, Qin Lv |
Knowl. Based Syst. | 2 |
| 2015 | Measuring enterprise network usage pattern & deploying passive optical LANsabstractRecent advances in the manufacturing and commercialization of passive optical components are now extending the capabilities of fiber to edge and campus networks. This paper presents a comparison between Passive Optical LAN (POL) and copper-based LAN solution, and demonstrate the benefits of PON such as reduced infrastructure footprint and cost, reduced power requirements, future-proof bandwidth, greener infrastructure, safer, higher security and higher reliability. Yaoping Ruan, Nikos Anerousis, Mudhakar Srivatsa, Jin Xiao 0005, R. Todd Christner, Luis Farrolas, John Short |
IM | 1 |
| 2014 | Managing risk in multi-node automation of endpoint managementabstractEndpoint management, including patching, health checking, configuration etc., is a key function for data center and cloud management. Managing multiple nodes through automation tools or scripts significantly increases efficiency. However, the risks of adverse impact due to excessive privilege or human error may propagate to a large pool of endpoints and lead to massive service disruptions and SLA (Service Level Agreement) violations. In this paper, we present a system that proactively and systematically manages the risk throughout the lifecycle phases of automation. We present a prototype implementation consisting of an authorization mechanism that guarantees the right level of eligibility and privilege of accessing the automation content (during the deployment stage), and an execution validator that controls the risk of human error which may cause massive damage to the infrastructure (during execution of the automation content). Our current implementation has been deployed to more than a dozen customer environments and achieved an efficiency gain of 58% with high execution accuracy. Sai Zeng, Constantin Adam, Frederick Wu, Shang Guo, Yaoping Ruan, Cashchakanithara Venugopal, Rajeev Puri |
NOMS | 5 |
| 2013 | Assessing service deployment readiness using enterprise crowdsourcing
Maja Vukovic, Jim Laredo, Yaoping Ruan, Milton Hernandez, Sriram Rajagopal |
IM | 3 |
| 2012 | Privileged identity management in enterprise service-hosting environmentsabstractIAM needs will only grow as devices, servers, and end points continue to increase . Current schemes are not sustainable as the number of IDs will explode. Environment is heterogeneous, and constantly adding new systems including Cloud. Our solution offers a platform where a user gets an individual user ID on a system - but only if they need it, when they need it, for only as long as they need it . Reusable ID scheme reduces the number of IDs in the system yielding cost savings on lifecycle management activities, improved security compliance . A compliance readiness platform can be enabled to prevent, flag, or monitor questionable access in or near real-time . Provide easily accessible logs to prove compliance policies. Kumar Bhaskaran, Milton Hernandez, Jim Laredo, Laura Luan, Yaoping Ruan, Maja Vukovic, Paul Driscoll, Alan Skinner, Girish Verma, Prema Vivekanandan, Leanne Chen, Gregory Gaskill |
NOMS | 5 |
| 2012 | Integrated user activity monitoring for regulatory servicesabstractRegulations such as FFIEC [5] and HIPAA [6] require activities of system administration to be captured and reviewed regularly. In IT service delivery environment, system maintenance activities are usually performed by the service provider whose system administrators access customer environment based on problem and change ticket being assigned. Mattias Marder, Kumar Bhaskaran, Milton Hernandez, Jim Laredo, Daniela Rosu 0001, Yaoping Ruan, Paul Driscoll, Alan Skinner |
NOMS | 6 |
| 2012 | Lightweight searchable screen video recordingabstractCommand logging of maintenance and operation activities of modern computer systems has become an integral component of customer and audit requirements. In recent years, this logging has usually been achieved via desktop video recording. However, the conventional approach of video recording requires high computation overhead, high network bandwidth, and a large storage size. Searching through video files is also a challenge. In this paper, we present a lossy, but text text-preserving, compression scheme that meets these challenges by creating a sparse bitonal image suitable for optical character recognition (OCR). Using our system for auditing, the bitonal image gets stored on a server. Due to the mechanism's text-preserving compression, we can apply OCR off-line to create annotations of each video frame, making the output searchable. Compared to state-of-the-art compression of raw video, our approach can reduce file size by 50-80%, while using CPU and memory resources similar to other methods. Mattias Marder, Amir Geva, Yaoping Ruan |
VCIP | 3 |
| 2011 | Sedic: privacy-aware data intensive computing on hybrid cloudsabstractThe emergence of cost-effective cloud services offers organizations great opportunity to reduce their cost and increase productivity. This development, however, is hampered by privacy concerns: a significant amount of organizational computing workload at least partially involves sensitive data and therefore cannot be directly outsourced to the public cloud. The scale of these computing tasks also renders existing secure outsourcing techniques less applicable. A natural solution is to split a task, keeping the computation on the private data within an organization's private cloud while moving the rest to the public commercial cloud. However, this hybrid cloud computing is not supported by today's data-intensive computing frameworks, MapReduce in particular, which forces the users to manually split their computing tasks. In this paper, we present a suite of new techniques that make such privacy-aware data-intensive computing possible. Our system, called Sedic, leverages the special features of MapReduce to automatically partition a computing job according to the security levels of the data it works on, and arrange the computation across a hybrid cloud. Specifically, we modified MapReduce's distributed file system to strategically replicate data, moving sanitized data blocks to the public cloud. Over this data placement, map tasks are carefully scheduled to outsource as much workload to the public cloud as possible, given sensitive data always stay on the private cloud. To minimize inter-cloud communication, our approach also automatically analyzes and transforms the reduction structure of a submitted job to aggregate the map outcomes within the public cloud before sending the result back to the private cloud for the final reduction. This also allows the users to interact with our system in the same way they work with MapReduce, and directly run their legacy code in our framework. We implemented Sedic on Hadoop and evaluated it using both real and synthesized computing jobs on a large-scale cloud test-bed. The study shows that our techniques effectively protect sensitive user data, offload a large amount of computation to the public cloud and also fully preserve the scalability of MapReduce. Kehuan Zhang, Xiao-yong Zhou, Yangyi Chen, XiaoFeng Wang 0001, Yaoping Ruan |
CCS | 5 |
| 2011 | ETree: Effective and Efficient Event Modeling for Real-Time Online Social Media NetworksabstractOutline social media networks (OSMNs) such as Twitter provide great opportunities for public engagement and event information dissemination. Event-related discussions occur in real time and at the worldwide scale. However, these discussions are in the form of short, unstructured messages and dynamically woven into daily chats and status updates. Compared with traditional news articles, the rich and diverse user-generated content raises unique new challenges for tracking and analyzing events. Effective and efficient event modeling is thus essential for real-time information-intensive OSMNs. In this work, we propose ETree, an effective and efficient event modeling solution for social media network sites. Targeting the unique challenges of this problem, ETree consists of three key components: (1) an n-gram based content analysis technique for identifying core information blocks from a large number of short messages, (2) an incremental and hierarchical modeling technique for identifying and constructing event theme structures at different granularities, and (3) an enhanced temporal analysis technique for identifying inherent causalities between information blocks. Detailed evaluation using 3.5 million tweets over a 5-month period demonstrates that ETree can efficiently generate high-quality event structures and identify inherent causal relationships with high accuracy. Hansu Gu, Xing Xie 0003, Qin Lv, Yaoping Ruan, Li Shang 0001 |
Web Intelligence | 4 |
| 2010 | Splitter: a proxy-based approach for post-migration testing of web applicationsabstractThe benefits of virtualized IT environments, such as compute clouds, have drawn interested enterprises to migrate their applications onto new platforms to gain the advantages of reduced hardware and energy costs, increased flexibility and deployment speed, and reduced management complexity. However, the process of migrating a complex application takes a considerable amount of effort, particularly when performing post-migration testing to verify that the application still functions correctly in the target environment. The traditional approach of test case generation and execution can take weeks and synthetic test cases may not adequately reflect actual application usage. Xiaoning Ding, Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Brian Peterson, Xiaodong Zhang 0001 |
EuroSys | 3 |
| 2009 | Building end-to-end management analytics for enterprise data centersabstractThe complexity of modern data centers has evolved significantly in recent years. One typically is comprised of a large number and types of middleware and applications that are hosted in a heterogeneous pool of both physical and virtual servers, connected by a complex web of virtual and physical networks. Therefore, to manage everything in a data center, system administrators usually need a plethora of management tools since one tool often manages only one type of devices. The boundaries between the different management tools can limit productivity of system administrators on their daily tasks as each tool only offers a partial view of the entire managed environment. As a result, advanced analytics such as impact analysis and problem determination are generally not achievable using the traditional management tools as they require a holistic view of the entire data center. In this paper, we describe an integrated management system for applications, servers, network and storage devices called DataGraph. Our system integrates data across heterogeneous point products and agents for management and monitoring to enable the above mentioned management analytics capabilities. A common data model is introduced to federate data collected by the different tools in multiple database repositories so no modifications are needed to existing management tools. A common integrated web user interface is implemented to facilitate management tasks that would otherwise require invoking multiple tools. We deployed this tool in a lab environment and demonstrated these analytics capabilities through several case studies. Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Ramani Routray, Chung-Hao Tan, Sandeep Gopisetty |
Integrated Network Management | 2 |
| 2008 | Automatic Software Fault Diagnosis by Exploiting Application Signatures
Xiaoning Ding, Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Xiaodong Zhang 0001 |
LISA | 3 |
| 2007 | PDA: A Tool for Automated Problem Determination
Hai Huang 0002, Raymond B. Jennings III, Yaoping Ruan, Ramendra K. Sahoo, Sambit Sahu, Anees Shaikh |
LISA | 3 |
| 2006 | Understanding and Addressing Blocking-Induced Network Server Latency
Yaoping Ruan, Vivek S. Pai |
USENIX ATC, General Track | 1 |
| 2006 | How DNS Misnaming Distorts Internet Topology Mapping
Ming Zhang 0005, Yaoping Ruan, Vivek S. Pai, Jennifer Rexford |
USENIX ATC, General Track | 2 |
| 2005 | Evaluating the impact of simultaneous multithreading on network servers using real hardwareabstractThis paper examines the performance of simultaneous multithreading (SMT) for network servers using actual hardware, multiple network server applications, and several workloads. Using three versions of the Intel Xeon processor with Hyper-Threading, we perform macroscopic analysis as well as microarchitectural measurements to understand the origins of the performance bottlenecks for SMT processors in these environments. The results of our evaluation suggest that the current SMT support in the Xeon is application and workload sensitive, and may not yield significant benefits for network servers.In general, we find that enabling SMT on real hardware usually produces only slight performance gains, and can sometimes lead to performance loss. In the uniprocessor case, previous studies appear to have neglected the OS overhead in switching from a uniprocessor kernel to an SMT-enabled kernel. The performance loss associated with such support is comparable to the gains provided by SMT. In the 2-way multiprocessor case, the higher number of memory references from SMT often causes the memory system to become the bottleneck, offsetting any processor utilization gains. This effect is compounded by the growing gap between processor speeds and memory latency. In trying to understand the large gains shown by simulation studies, we find that while the general trends for microarchitectural behavior agree with real hardware, differences in sizing assumptions and performance models yield much more optimistic benefits for SMT than we observe. Yaoping Ruan, Vivek S. Pai, Erich M. Nahum, John M. Tracey |
SIGMETRICS | 1 |
| 2004 | The origins of network server latency & the myth of connection schedulingabstractWe investigate the origins of server-induced latency to understand how to improve latency optimization techniques. Using the Flash Web server [4], we analyze latency behavior under various loads. Despite latency profiles that suggest standard queuing delays, we find that most latency actually originates from negative interactions between the application and the locking and blocking mechanisms in the kernel. Modifying the server and kernel to avoid these problems yields both qualitative and quantitative changes in the latency profiles -- latency drops by more than an order of magnitude, and the effective service discipline also improves.We find our modifications also mitigate service burstiness in the application, reducing the event queue lengths dramatically and eliminating any benefit from application-level connection scheduling. We identify one remaining source of unfairness, related to competition in the networking stack. We show that adjusting the TCP congestion window size addresses this problem, reducing latency by an additional factor of three. Yaoping Ruan, Vivek S. Pai |
SIGMETRICS | 1 |
| 2004 | Making the "Box" Transparent: System Call Performance as a First-Class Result
Yaoping Ruan, Vivek S. Pai |
USENIX ATC, General Track | 1 |