VLDB 2026 Research / reviewers in the wild / expert
Qi Zhao 0006
dblp:05/490-6
· DBLP profile ↗
8ranked-venue papers
5as first author
0since 2021 · last 2009
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
7 papers |
Network measurement and analytics · 49% Internet of things and sensor networks · 22% Network management and operations · 19% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 50% Data mining · 50% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Network and information security
1 paper |
Network security · 100% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Network management and operations › fault management
fault diagnosis |
0.1 | 1 | 2009 | Towards automated performance diagnosis in a large IPTV network · SIGCOMM 2009 |
Network management and operations › performance management
performance diagnosis |
0.1 | 1 | 2009 | Towards automated performance diagnosis in a large IPTV network · SIGCOMM 2009 |
Network measurement and analytics
workload characterization |
0.1 | 1 | 2009 | Modeling user activities in a large IPTV system · Internet Measurement Conference 2009 |
Internet of things and sensor networks › time synchronization
clock skew and offset estimation |
0.1 | 1 | 2008 | ACES: adaptive clock estimation and synchronization using Kalman filtering · MobiCom 2008 |
Internet of things and sensor networks
time synchronization |
0.1 | 1 | 2008 | ACES: adaptive clock estimation and synchronization using Kalman filtering · MobiCom 2008 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.1 | 1 | 2006 | Finding global icebergs over distributed data sets · PODS 2006 |
Query processing and optimization › aggregate query processing
iceberg query |
0.1 | 1 | 2006 | Finding global icebergs over distributed data sets · PODS 2006 |
Network measurement and analytics
traffic measurement |
0.1 | 1 | 2006 | Detection of Super Sources and Destinations in High-Speed Networks: Algorithms, Analysis and Evaluation · IEEE J. Sel. Areas Commun. 2006 |
Network measurement and analytics
heavy hitter detection |
0.1 | 1 | 2005 | Joint Data Streaming and Sampling Techniques for Detection of Super Sources and Destinations · Internet Measurement Conference 2005 |
Network measurement and analytics
streaming data |
0.1 | 1 | 2005 | Data streaming algorithms for accurate and efficient measurement of traffic and flow matrices · SIGMETRICS 2005 |
Network measurement and analytics
traffic matrix estimation |
0.1 | 1 | 2005 | Data streaming algorithms for accurate and efficient measurement of traffic and flow matrices · SIGMETRICS 2005 |
Network measurement and analytics › traffic measurement
traffic monitoring |
0.1 | 1 | 2005 | Joint Data Streaming and Sampling Techniques for Detection of Super Sources and Destinations · Internet Measurement Conference 2005 |
Internet architecture and protocols
packet scheduling |
0.0 | 1 | 2004 | On the Computational Complexity of Maintaining GPS Clock in Packet Scheduling · INFOCOM 2004 |
Mathematical optimization › combinatorial optimization
scheduling complexity |
0.0 | 1 | 2004 | On the Computational Complexity of Maintaining GPS Clock in Packet Scheduling · INFOCOM 2004 |
Internet of things and sensor networks
resource-constrained networks |
0.0 | 1 | 2008 | ACES: adaptive clock estimation and synchronization using Kalman filtering · MobiCom 2008 |
Internet of things and sensor networks
wireless sensor network |
0.0 | 1 | 2008 | ACES: adaptive clock estimation and synchronization using Kalman filtering · MobiCom 2008 |
Network security › intrusion detection and prevention › intrusion detection
anomaly detection |
0.0 | 1 | 2006 | Detection of Super Sources and Destinations in High-Speed Networks: Algorithms, Analysis and Evaluation · IEEE J. Sel. Areas Commun. 2006 |
Network security
traffic analysis |
0.0 | 1 | 2006 | Detection of Super Sources and Destinations in High-Speed Networks: Algorithms, Analysis and Evaluation · IEEE J. Sel. Areas Commun. 2006 |
Routing and switching
traffic engineering |
0.0 | 1 | 2005 | Data streaming algorithms for accurate and efficient measurement of traffic and flow matrices · SIGMETRICS 2005 |
Internet architecture and protocols
quality of service |
0.0 | 1 | 2004 | On the Computational Complexity of Maintaining GPS Clock in Packet Scheduling · INFOCOM 2004 |
Methods — techniques the papers use, named apart from their topics
data streaming · 0.2sampling · 0.2hash-based flow sampling · 0.1workload generation · 0.1trace analysis · 0.1statistical modeling · 0.1statistical data mining · 0.1multi-resolution analysis · 0.1kalman filtering · 0.1counting sketch · 0.1bayesian statistics · 0.1computational complexity analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Modeling user activities in a large IPTV systemabstractInternet Protocol Television (IPTV) has emerged as a new delivery method for TV. In contrast with native broadcast in traditional cable and satellite TV system, video streams in IPTV are encoded in IP packets and distributed using IP unicast and multicast. This new architecture has been strategically embraced by ISPs across the globe, recognizing the opportunity for new services and its potential toward a more interactive style of TV watching experience in the future. Since user activities such as channel switches in IPTV impose workload beyond local TV or set-top box (different from broadcast TV systems), it becomes essential to characterize and model the aggregate user activities in an IPTV network to support various system design and performance evaluation functions such as network capacity planning. In this work, we perform an in-depth study on several intrinsic characteristics of IPTV user activities by analyzing the real data collected from an operational nation-wide IPTV system. We further generalize the findings and develop a series of models for capturing both the probability distribution and time-dynamics of user activities. We then combine theses models to design an IPTV user activity workload generation tool called SIMUL WATCH, which takes a small number of input parameters and generates synthetic workload traces that mimic a set of real users watching IPTV. We validate all the models and the prototype of SIMUL WATCH using the real traces. In particular, we show that SIMUL WATCH can estimate the unicast and multicast traffic accurately, proving itself as a useful tool in driving the performance study in IPTV systems. Tongqing Qiu, Zihui Ge, Seungjoon Lee, Jia Wang 0001, Jun (Jim) Xu, Qi Zhao 0006 |
Internet Measurement Conference | 6 |
| 2009 | Towards automated performance diagnosis in a large IPTV networkabstractIPTV is increasingly being deployed and offered as a commercial service to residential broadband customers. Compared with traditional ISP networks, an IPTV distribution network (i) typically adopts a hierarchical instead of mesh-like structure, (ii) imposes more stringent requirements on both reliability and performance, (iii) has different distribution protocols (which make heavy use of IP multicast) and traffic patterns, and (iv) faces more serious scalability challenges in managing millions of network elements. These unique characteristics impose tremendous challenges in the effective management of IPTV network and service. In this paper, we focus on characterizing and troubleshooting performance issues in one of the largest IPTV networks in North America. We collect a large amount of measurement data from a wide range of sources, including device usage and error logs, user activity logs, video quality alarms, and customer trouble tickets. We develop a novel diagnosis tool called Giza that is specifically tailored to the enormous scale and hierarchical structure of the IPTV network. Giza applies multi-resolution data analysis to quickly detect and localize regions in the IPTV distribution hierarchy that are experiencing serious performance problems. Giza then uses several statistical data mining techniques to troubleshoot the identified problems and diagnose their root causes. Validation against operational experiences demonstrates the effectiveness of Giza in detecting important performance issues and identifying interesting dependencies. The methodology and algorithms in Giza promise to be of great use in IPTV network operations. Ajay Mahimkar, Zihui Ge, Aman Shaikh, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Qi Zhao 0006 |
SIGCOMM | 7 |
| 2008 | ACES: adaptive clock estimation and synchronization using Kalman filteringabstractClock synchronization across a network is essential for a large number of applications ranging from wired network measurements to data fusion in sensor networks. Earlier techniques are either limited to undesirable accuracy or rely on specific hardware characteristics that may not be available for certain systems. In this work, we examine the clock synchronization problem in resource-constrained networks such as wireless sensor networks where nodes have limited energy and bandwidth, and also lack the high accuracy oscillators or programmable network interfaces some previous protocols depend on. This paper derives a general model for clock offset and skew and demonstrates its applicability. We design efficient algorithms based on this model to achieve high synchronization accuracy given limited resources. These algorithms apply the Kalman filter to track the clock offset and skew, and adaptively adjust the synchronization interval so that the desired error bounds are achieved. We demonstrate the performance advantages of our schemes through extensive simulations obeying real-world constraints. Benjamin R. Hamilton, Xiaoli Ma, Qi Zhao 0006, Jun (Jim) Xu |
MobiCom | 3 |
| 2006 | Finding global icebergs over distributed data setsabstractFinding icebergs – items whose frequency of occurrence is above a certain threshold – is an important problem with a wide range of applications. Most of the existing work focuses on iceberg queries at a single node. However, in many real-life applications, data sets are distributed across a large number of nodes. Two naïve approaches might be considered. In the first, each node ships its entire data set to a central server, and the central server uses single-node algorithms to find icebergs. But it may incur prohibitive communication overhead. In the second, each node submits local icebergs, and the central server combines local icebergs to find global icebergs. But it may fail because in many important applications, globally frequent items may not be frequent at any node. In this work, we propose two novel schemes that provide accurate and efficient solutions to this problem: a sampling-based scheme and a counting-sketch-based scheme. In particular, the latter scheme incurs a communication cost at least an order of magnitude smaller than the naïve scheme of shipping all data, yet is able to achieve very high accuracy. Through rigorous theoretical and experimental analysis we establish the statistical properties of our proposed algorithms, including their accuracy bounds. Qi Zhao 0006, Mitsunori Ogihara, Haixun Wang, Jun (Jim) Xu |
PODS | 1 |
| 2006 | Detection of Super Sources and Destinations in High-Speed Networks: Algorithms, Analysis and EvaluationabstractDetecting the sources or destinations that have communicated with a large number of distinct destinations or sources (i.e., large “fan-out” or “fan-in”) during a small time interval is an important problem in network measurement and security. Previous detection approaches are not able to deliver the desired accuracy at high link speeds (10–40 Gb/s). In this work, we propose two novel algorithms that provide accurate and efficient solutions to this problem. Their designs are based on the insight that sampling and data streaming are often suitable for capturing different and complementary regions of the information spectrum, and a close collaboration between them is an excellent way to recover the complete information. Our first solution builds on the standard hash-based flow sampling algorithm. Its main innovation is that the sampled traffic is further filtered by a data streaming module which allows for much higher sampling rate (hence, much higher accuracy) than achievable with standard hash-based flow sampling. Our second solution is more sophisticated but offers higher accuracy. It combines the power of data streaming in efficiently estimating quantities (e.g., fan-out) associated with a given identity, and the power of sampling in collecting a list of candidate identities. The performance of both solutions are evaluated using both mathematical analysis and trace-driven experiments on real-world Internet traffic. Qi Zhao 0006, Jun (Jim) Xu, Abhishek Kumar 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2005 | Joint Data Streaming and Sampling Techniques for Detection of Super Sources and Destinations
Qi Zhao 0006, Abhishek Kumar 0003, Jun (Jim) Xu |
Internet Measurement Conference | 1 |
| 2005 | Data streaming algorithms for accurate and efficient measurement of traffic and flow matricesabstractThe traffic volume between origin/destination (OD) pairs in a network, known as traffic matrix, is essential for efficient network provisioning and traffic engineering. Existing approaches of estimating the traffic matrix, based on statistical inference and/or packet sampling, usually cannot achieve very high estimation accuracy. In this work, we take a brand new approach in attacking this problem. We propose a novel data streaming algorithm that can process traffic stream at very high speed (e.g., 40 Gbps) and produce traffic digests that are orders of magnitude smaller than the traffic stream. By correlating the digests collected at any OD pair using Bayesian statistics, the volume of traffic flowing between the OD pair can be accurately determined. We also establish principles and techniques for optimally combining this streaming method with sampling, when sampling is necessary due to stringent resource constraints. In addition, we propose another data streaming algorithm that estimates flow matrix, a finer-grained characterization than traffic matrix. Flow matrix is concerned with not only the total traffic between an OD pair (traffic matrix), but also how it splits into flows of various sizes. Through rigorous theoretical analysis and extensive synthetic experiments on real Internet traffic, we demonstrate that these two algorithms can produce very accurate estimation of traffic matrix and flow matrix respectively. Qi Zhao 0006, Abhishek Kumar 0003, Jia Wang 0001, Jun (Jim) Xu |
SIGMETRICS | 1 |
| 2004 | On the Computational Complexity of Maintaining GPS Clock in Packet SchedulingabstractPacket scheduling is an important mechanism to provide QoS guarantees in data networks. A scheduling algorithm generally consists of two functions: one estimates how the GPS (general processor sharing) clock progresses with respect to the real time; the other decides the order of serving packets based on an estimation of their GPS start/finish times. In this work, we answer important open questions concerning the computational complexity of performing the first function. We systematically study the complexity of computing the GPS virtual start/finish times of the packets, which has long been believed to be /spl Omega/(n) per packet but has never been either proved or refuted. We also answer several other related questions such as "whether the complexity can be lower if the only thing that needs to be computed is the relative order of the GPS finish times of the packets rather than their exact values?". Qi Zhao 0006, Jun (Jim) Xu |
INFOCOM | 1 |