EDBT 2026 Demo / reviewers in the wild / expert
Tom Z. J. Fu
dblp:89/6622 · also Tom Zhengjia, Tom Zhengjia Fu
· DBLP profile ↗
38ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0003-3312-2402ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 16 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 6 · 1 since 2021Systems, architecture and hardware · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Cloud and datacenter computing · 49% Parallel and multicore computing · 20% Storage systems · 13% | |
| Artificial intelligence
4 papers |
Transfer learning and domain adaptation · 35% Time series and sequential data · 17% Representation and self-supervised learning · 17% | |
| Databases, data mining, and information retrieval
8 papers |
Data stream processing · 60% Spatial and temporal data management · 30% Indexing and storage engines · 9% | |
| Computer networks
9 papers |
Content delivery and video streaming · 51% Software-defined and programmable networks · 23% Network optimization and economics · 19% |
Topics — the 30 heaviest of 49, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
causal representation learning |
0.8 | 1 | 2024 | Granger causal representation learning for groups of time series · Sci. China Inf. Sci. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality |
0.8 | 1 | 2024 | Granger causal representation learning for groups of time series · Sci. China Inf. Sci. 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
semi-supervised domain adaptation |
0.8 | 1 | 2024 | Transferable Time-Series Forecasting Under Causal Conditional Shift · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
time series domain adaptation |
0.8 | 1 | 2024 | Transferable Time-Series Forecasting Under Causal Conditional Shift · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.8 | 1 | 2024 | Transferable Time-Series Forecasting Under Causal Conditional Shift · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Content delivery and video streaming › video-on-demand
peer-to-peer video-on-demand |
0.7 | 5 | 2015 | A Unifying Model and Analysis of P2P VoD Replication and Scheduling · IEEE/ACM Trans. Netw. 2015 On Replication Algorithm in P2P VoD · IEEE/ACM Trans. Netw. 2013 A unifying model and analysis of P2P VoD replication and scheduling · INFOCOM 2012 |
Cloud and datacenter computing › resource allocation
bandwidth allocation |
0.6 | 2 | 2018 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream Analytics · ICNP 2018 Impacts of task placement and bandwidth allocation on stream analytics · ICNP 2017 |
Cloud and datacenter computing › big data analytics
stream analytics |
0.6 | 2 | 2018 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream Analytics · ICNP 2018 Impacts of task placement and bandwidth allocation on stream analytics · ICNP 2017 |
Data stream processing
streaming analytics |
0.4 | 2 | 2019 | DRS: Auto-Scaling for Real-Time Stream Analytics · IEEE/ACM Trans. Netw. 2017 On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream Analytics · IEEE J. Sel. Areas Commun. 2019 |
Computer vision › Video understanding and tracking › multi-camera tracking
multi-target multi-camera tracking |
0.4 | 1 | 2019 | Interactive Multi-camera Soccer Video Analysis System · ACM Multimedia 2019 |
Data stream processing › stream processing systems
elastic stream processing |
0.4 | 1 | 2019 | Elasticutor: Rapid Elasticity for Realtime Stateful Stream Processing · SIGMOD Conference 2019 |
Data stream processing › stream processing systems
stateful stream processing |
0.4 | 1 | 2019 | Elasticutor: Rapid Elasticity for Realtime Stateful Stream Processing · SIGMOD Conference 2019 |
Multimedia analysis and retrieval
sports video analysis |
0.4 | 1 | 2019 | Interactive Multi-camera Soccer Video Analysis System · ACM Multimedia 2019 |
Network optimization and economics › resource allocation › bandwidth allocation
dynamic bandwidth allocation |
0.4 | 1 | 2019 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream Analytics · IEEE J. Sel. Areas Commun. 2019 |
Software-defined and programmable networks
SDN control plane |
0.4 | 1 | 2019 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream Analytics · IEEE J. Sel. Areas Commun. 2019 |
Parallel and multicore computing › task scheduling
dynamic scheduling |
0.4 | 1 | 2019 | Elasticutor: Rapid Elasticity for Realtime Stateful Stream Processing · SIGMOD Conference 2019 |
Data stream processing › streaming data retrieval
stream indexing |
0.3 | 1 | 2018 | Waterwheel: Realtime Indexing and Temporal Range Query Processing over Massive Data Streams · ICDE 2018 |
Spatial and temporal data management › temporal query processing
temporal range query |
0.3 | 1 | 2018 | Waterwheel: Realtime Indexing and Temporal Range Query Processing over Massive Data Streams · ICDE 2018 |
Storage systems
cross-layer optimization |
0.3 | 1 | 2018 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream Analytics · ICNP 2018 |
Distributed systems › distributed data processing
distributed indexing |
0.3 | 1 | 2018 | Waterwheel: Realtime Indexing and Temporal Range Query Processing over Massive Data Streams · ICDE 2018 |
Storage systems
indexing |
0.3 | 1 | 2018 | Waterwheel: Realtime Indexing and Temporal Range Query Processing over Massive Data Streams · ICDE 2018 |
Indexing and storage engines
distributed indexing |
0.3 | 1 | 2017 | DITIR: Distributed Index for High Throughput Trajectory Insertion and Real-time Temporal Range Query · Proc. VLDB Endow. 2017 |
Spatial and temporal data management
spatio-temporal indexing |
0.3 | 1 | 2017 | DITIR: Distributed Index for High Throughput Trajectory Insertion and Real-time Temporal Range Query · Proc. VLDB Endow. 2017 |
Spatial and temporal data management › spatio-temporal indexing
trajectory indexing |
0.3 | 1 | 2017 | DITIR: Distributed Index for High Throughput Trajectory Insertion and Real-time Temporal Range Query · Proc. VLDB Endow. 2017 |
Cloud and datacenter computing
autoscaling |
0.3 | 1 | 2017 | DRS: Auto-Scaling for Real-Time Stream Analytics · IEEE/ACM Trans. Netw. 2017 |
Cloud and datacenter computing › resource management
cloud resource management |
0.3 | 1 | 2017 | DRS: Auto-Scaling for Real-Time Stream Analytics · IEEE/ACM Trans. Netw. 2017 |
Parallel and multicore computing
task allocation |
0.3 | 1 | 2017 | Impacts of task placement and bandwidth allocation on stream analytics · ICNP 2017 |
Robotics › Motion planning and robot control › robot control
trajectory tracking |
0.2 | 1 | 2015 | LiveTraj: Real-Time Trajectory Tracking over Live Video Streams · ACM Multimedia 2015 |
Cloud and datacenter computing
cloud platform |
0.2 | 1 | 2015 | LiveTraj: Real-Time Trajectory Tracking over Live Video Streams · ACM Multimedia 2015 |
Cloud and datacenter computing › resource management › cloud resource management
elastic resource management |
0.2 | 1 | 2015 | LiveTraj: Real-Time Trajectory Tracking over Live Video Streams · ACM Multimedia 2015 |
Methods — techniques the papers use, named apart from their topics
granger causality · 1.5optimization formulation · 1.2trajectory computation · 1.1parallel processing · 1.1cross-layer optimization · 1.1bandwidth sharing algorithms · 1.1simulation · 0.9variational inference · 0.8representation learning · 0.8model-based scheduling · 0.8intra-executor load balancing · 0.8causal structure discovery · 0.8heuristic routing · 0.7dynamic bandwidth adjustment · 0.7load balancing · 0.6data partitioning · 0.3measurement-driven analysis · 0.3heuristics · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Granger causal representation learning for groups of time series
Ruichu Cai, Yunjin Wu, Xiaokai Huang, Wei Chen 0103, Tom Z. J. Fu, Zhifeng Hao 0004 |
Sci. China Inf. Sci. | 5 |
| 2024 | Transferable Time-Series Forecasting Under Causal Conditional ShiftabstractThis paper focuses on the problem of semi-supervised domain adaptation for time-series forecasting, which is underexplored in literature, despite being often encountered in practice. Existing methods on time-series domain adaptation mainly follow the paradigm designed for static data, which cannot handle domain-specific complex conditional dependencies raised by data offset, time lags, and variant data distributions. In order to address these challenges, we analyze variational conditional dependencies in time-series data and find that the causal structures are usually stable among domains, and further raise the causal conditional shift assumption. Enlightened by this assumption, we consider the causal generation process for time-series data and propose an end-to-end model for the semi-supervised domain adaptation problem on time-series forecasting. Our method can not only discover the Granger-Causal structures among cross-domain data but also address the cross-domain time-series forecasting problem with accurate and interpretable predicted results. We further theoretically analyze the superiority of the proposed method, where the generalization error on the target domain is bounded by the empirical risks and by the discrepancy between the causal structures from different domains. Experimental results on both synthetic and real data demonstrate the effectiveness of our method for the semi-supervised domain adaptation method on time-series forecasting. Zijian Li 0001, Ruichu Cai, Tom Z. J. Fu, Zhifeng Hao 0004, Kun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | DiffPerf: Toward Performance Differentiation and Optimization With SDN ImplementationabstractThe continuous growth of Internet traffic, especially video content, presents challenges for access providers (APs) who must upgrade their infrastructure to meet increasing demands. Ensuring a high-quality experience (QoE) for end-users and finding ways to monetize network resources are key concerns. Guaranteeing QoE is complex, as it depends not only on link capacity but also on competing traffic flows and shared network data plane buffers. To address these challenges, we proposeDiffPerf, an in-network, online, and dynamic allocation system.DiffPerfoperates at both macroscopic and microscopic levels. At the macroscopic level, it elastically allocates bandwidth to performance-centric service classes defined by APs to accommodate different performance requirements. At the microscopic level,DiffPerfemploys a lightweight data-driven algorithm to statistically differentiate and isolate traffic flows within each class, improving their performance. We implementedDiffPerfprototypes using SDN-based technology, one with OpenDaylight and OpenFlow hardware switches, and the other with programmable Intel Tofino switches. Our evaluation focused on on-demand video streaming. The results demonstrate thatDiffPerfoffers APs a range of allocation choices while ensuring strong performance isolation. Additionally,DiffPerfimproves fairness and enhances overall user-perceived QoE within each class. Notably,DiffPerfconserves bandwidth and delivers a QoE improvement approximately$4.6\times $higher than TCP BBR, the most popular congestion control mechanism on the Internet. Walid Aljoby, Xin Wang 0040, Dinil Mon Divakaran, Tom Z. J. Fu, Richard T. B. Ma, Khaled A. Harras |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2021 | DiffPerf: An In-Network Performance Optimization for Improving User-Perceived QoEabstractContinuing the current trend, Internet traffic is expected to grow significantly over the coming years, with video traffic consuming the biggest share. Despite numerous optimizations of the transport congestion control, and the switch butter sizing and management algorithms; however, the complex interaction among all of them still leads to uncertain user performance and thus degrades user-perceived quality, under various network and traffic conditions. The culprit is the difficulty to dynamically control the amount of bandwidth allocated to each of the competing flows under bottleneck due to the algorithms lack of visibility of butter content where the flows reside. We address this bandwidth allocation problem by proposing DiffPerf, an in-network system that relies on a lightweight learning algorithm to statistically differentiate and isolate user flows to help them achieve better performance in an online and dynamic manner. We built two SDN-based prototypes of DiffPerf; one on OpenDaylight with OpenFlow Brocade switch and the other with programmable data plane Barefoot Tofino switch. We evaluate it from an application perspective for ABR video streaming as it accounts for a majority of the Internet traffic. Our evaluations demonstrate the practicality and flexibility that DiffPerf assists users in achieving better fairness and improving overall user-perceived quality. On average DiffPerf yields a quality improvement of about $4.6\times$ and $1.2\times$ higher than TCP BBR and TCP CUBIC, respectively. Walid Aljoby, Xin Wang 0040, Dinil Mon Divakaran, Tom Z. J. Fu, Richard T. B. Ma |
NetSoft | 4 |
| 2021 | CrowdSR: enabling high-quality video ingest in crowdsourced livecast via super-resolutionabstractThe prevalence of personal devices motivates the rapid development of crowdsourced livecast in recent years. However, there exists huge diversity of upstream bandwidth among amateur broadcasters. Moreover, the highest video quality that can be streamed is limited by the hardware configuration of broadcaster devices (e.g., 540p for low-end mobile devices). The above factors pose significant challenges to the ingestion of high-resolution live video streams, and result in poor quality-of-experience (QoE) for viewers. In this paper, we propose a novel live video ingest approach called CrowdSR for crowdsourced livecast. CrowdSR can transform a low-resolution video stream uploaded by weak devices into a high-resolution video stream via super-resolution, and then deliver the stream to viewers. CrowdSR can exploit crowdsourced high-resolution video patches from similar broadcasters to speedup model training. Different from previous work, our approach does not require any modification at the client side, and thus is more practical and easy to implement. Finally, we implement and evaluate CrowdSR by conducting a series of real-world experiments. The results show that CrowdSR significantly outperforms the baseline approaches by 0.42-1.09 dB in terms of PSNR and 0.006-0.014 in terms of SSIM. Zhenxiao Luo, Miao Hu 0001, Yipeng Zhou, Tom Z. J. Fu, Di Wu 0001 |
NOSSDAV | 6 |
| 2021 | Vibra: neural adaptive streaming of VBR-encoded videosabstractVariable Bitrate (VBR) video encoding can provide much high quality-to-bits ratio compared to the widely adopted Constant Bitrate (CBR) encoding, and thus receives significant attentions by content providers in recent years. However, it is challenging to design efficient adaptive bitrate algorithms for VBR-encoded videos due to the sharply fluctuating chunk size and the resulting bitrate burstiness. In this paper, we propose a neural adaptive streaming framework called Vibra for VBR-encoded videos, which can well accommodate the high fluctuation of video chunk sizes and improve the quality-of-experience (QoE) of end users significantly. Our framework takes the characteristics of VBR-encoded videos into account, and adopts the technique of deep reinforcement learning to train a model for bitrate adaptation. We also conduct extensive trace-driven experiments, and the results show that Vibra outperforms the state-of-the-art ABR algorithms with an improvement of 8.17% -- 29.21% in terms of the average QoE. Gangqiang Zhou, Run Wu, Miao Hu 0001, Yipeng Zhou, Tom Z. J. Fu, Di Wu 0001 |
NOSSDAV | 5 |
| 2021 | Causal Mechanism Transfer Network for Time Series Domain Adaptation in Mechanical SystemsabstractData-driven models are becoming essential parts in modern mechanical systems, commonly used to capture the behavior of various equipment and varying environmental characteristics. Despite the advantages of these data-driven models on excellent adaptivity to high dynamics and aging equipment, they are usually hungry for massive labels, mostly contributed by human engineers at a high cost. Fortunately, domain adaptation enhances the model generalization by utilizing the labeled source data and the unlabeled target data. However, the mainstream domain adaptation methods cannot achieve ideal performance on time series data, since they assume that the conditional distributions are equal. This assumption works well in the static data but is inapplicable for the time series data. Even the first-order Markov dependence assumption requires the dependence between any two consecutive time steps. In this article, we assume that the causal mechanism is invariant and present our Causal Mechanism Transfer Network (CMTN) for time series domain adaptation. By capturing causal mechanisms of time series data, CMTN allows the data-driven models to exploit existing data and labels from similar systems, such that the resulting model on a new system is highly reliable even with limited data. We report our empirical results and lessons learned from two real-world case studies, on chiller plant energy optimization and boiler fault detection, which outperform the existing state-of-the-art method. Zijian Li 0001, Ruichu Cai, Hong Wei Ng, Marianne Winslett, Tom Z. J. Fu |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | Interactive Multi-camera Soccer Video Analysis SystemabstractAutomatic sports video analysis is an active field of research, and accurate player & ball tracking is essential for soccer video analysis and visualization. However, the variations over frames and the scarceness of large-scale well-annotated datasets make it difficult to perform supervised learning using pre-trained models, especially for Multi-Camera Multi-Target Tracking (MCMT). In this paper, we introduce an end-to-end system for multi-camera soccer video analysis that makes heavy use of parallel processing for optimization of the processing workflow. The proposed thread-level parallelism speeds up our system by more than 15 times while maintaining the level of accuracy. The system tracks the trajectories of the ball and the players in a world coordinate system based on soccer videos captured by a set of synchronized cameras. Based on these trajectories, various player-, ball-, and team-related statistics are computed, and the resulting data and visualizations can be interactively explored by the user. Yunjin Wu, Ziyuan Zhao, Shengqiang Zhang, Lulu Yao, Tom Z. J. Fu |
ACM Multimedia | 6 |
| 2019 | Elasticutor: Rapid Elasticity for Realtime Stateful Stream ProcessingabstractElasticity is highly desirable for stream systems to guarantee low latency against workload dynamics, such as surges in arrival rate and fluctuations in data distribution. Existing systems achieve elasticity using a resource-centric approach that repartitions keys across the parallel instances, i.e., executors, to balance the workload and scale operators. However, such operator-level repartitioning requires global synchronization and prohibits rapid elasticity. We propose an executor-centric approach that avoids operator-level key repartitioning and implements executors as the building blocks of elasticity. By this new approach, we design the Elasticutor framework with two level of optimizations: i) a novel implementation of executors, i.e., elastic executors, that perform elastic multi-core execution via efficient intra-executor load balancing and executor scaling and ii) a global model-based scheduler that dynamically allocates CPU cores to executors based on the instantaneous workloads. We implemented a prototype of Elasticutor and conducted extensive experiments. We show that Elasticutor doubles the throughput and achieves up to two orders of magnitude lower latency than previous methods for dynamic workloads of real-world applications. Tom Z. J. Fu, Richard T. B. Ma, Marianne Winslett |
SIGMOD Conference | 2 |
| 2019 | Inter-domain routing bottlenecks and their aggravation
Xia Yin 0001, Xingang Shi, Jiong He, Tom Z. J. Fu, Marianne Winslett |
Comput. Networks | 6 |
| 2019 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream AnalyticsabstractData communication in cloud-based distributed stream data analytics often involves a collection of parallel and pipelined TCP flows. As the standard TCP congestion control mechanism and its variants are designed for achieving “fairness” among competing flows and are agnostic to the application layer contexts, the bandwidth allocation among a set of TCP flows traversing bottleneck links often leads to sub-optimal application-layer performance measures, e.g., stream processing throughput or average tuple complete latency. Motivated by this and enabled by the rapid development of the software-defined networking (SDN) techniques, in this paper, we re-investigate the design space of the bandwidth allocation problem and propose a cross-layer framework which utilizes the instantaneous information obtained from the application layer and provides on-the-fly and dynamic bandwidth adjustment algorithms for assisting the stream analytics applications achieving better performance during the runtime. We implement a prototype cross-layer bandwidth allocation framework based on a popular open-source distributed stream processing platform, Apache Storm, together with the OpenDaylight controller, and carry out extensive experiments with real-world analytical workloads on top of a local cluster consisting of ten workstations interconnected by a SDN-enabled fat-tree like testbed. The experiment results clearly validate the effectiveness and efficiency of our proposed framework and algorithms. Finally, we leverage the proposed cross-layer SDN framework and introduce an exemplary mechanism for bandwidth sharing and performance reasoning among multiple active applications and show a case of a point solution on how to approximate application-level fairness. Walid Aljoby, Xin Wang 0040, Tom Z. J. Fu, Richard T. B. Ma |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | HaaS: Cloud-Based Real-Time Data Analytics with Heterogeneity-Aware SchedulingabstractReal-time data analytics has become increasingly important in modern times as many organizations and companies are generating and analyzing high volume of data constantly. Despite of the impressive technical development, it remains a challenging job to analyze the stream data effectively and efficiently because traditional hardware and software lack specific designs and optimizations for those emerging requirements. In this paper, we discuss our experience on real-time data analytics, with our in-house processing framework HaaS. HaaS is designed to exploit existing data analytics tools and libraries as well as distributed computing technologies to embrace heterogeneous computation resources in the cloud. HaaS utilizes hierarchical clustering to partition physical topology of clusters weighted with task topology information into densely connected sub-graphs. HaaS is also equipped with a heterogeneity-aware scheduling algorithm to facilitate holistic optimization over multiple running tasks with various service level agreements. To the best of our knowledge, HaaS is the first ever streaming analytical framework providing users with flexible and optimized usage with CPUs, GPUs and FPGAs in the cloud. Users with stream processing tasks can easily enjoy remarkable advantages of CPUs, GPUs and FPGAs in throughput, power consumption and monetary cost over others. In our empirical evaluations with highly diversified workloads, HaaS saves over 18% on power consumption and 24% on monetary cost over existing system design architecture, while the overall throughput of HaaS remains no lower than 90% of the theoretical limit. Jiong He, Yao Chen 0008, Tom Z. J. Fu, Marianne Winslett, Liang You |
ICDCS | 3 |
| 2018 | Waterwheel: Realtime Indexing and Temporal Range Query Processing over Massive Data StreamsabstractMassive data streams from sensors in Internet of Things (IoT) and smart devices with Global Positioning System (GPS) are now flooding to database systems for further processing and analysis. The capability of real-time retrieval from both fresh and historical data turns out to be the key enabler to the real world applications in smart manufacturing and smart city utilizing these data streams. In this paper, we present a simple and effective distributed solution to achieve millions of tuple insertions per second and ad-hoc temporal range query processing in milliseconds. To this end, we propose a new data partitioning scheme that takes advantage of the workload characteristics and avoids expensive global data merging. Furthermore, to resolve the throughput bottleneck, we adopt a template-based index method to skip unnecessary index structure adjustments over the relatively stable distribution of incoming tuples. To parallelize data insertion and query processing, we propose an efficient dispatching mechanism and effective load balancing strategies to fully utilize computational resources in a workload-aware manner. On both synthetic and real workloads, our solution consistently outperforms state-of-the-art open-source systems by at least an order of magnitude. Ruichu Cai, Tom Z. J. Fu, Jiong He, Zijie Lu, Marianne Winslett |
ICDE | 3 |
| 2018 | On SDN-Enabled Online and Dynamic Bandwidth Allocation for Stream AnalyticsabstractData communication in cloud-based distributed stream data analytics often involves a collection of parallel and pipelined TCP flows. As the standard TCP congestion control mechanism is designed for achieving "fairness" among competing flows and is agnostic to the application layer contexts, the bandwidth allocation among a set of TCP flows traversing bottleneck links often leads to sub-optimal application-layer performance measures, e.g., stream processing throughput or average tuple complete latency. Motivated by this and enabled by the rapid development of the Software-Defined Networking (SDN) techniques, in this paper, we re-investigate the design space of the bandwidth allocation problem and propose a cross-layer framework which utilizes the additional information obtained from the application layer and provides on-the-fly and dynamic bandwidth adjustment algorithms for helping the stream analytics applications achieving better performance during the runtime. We implement a prototype cross-layer bandwidth allocation framework based on a popular open-source distributed stream processing platform, Apache Storm, together with the OpenDaylight controller, and carry out extensive experiments with real-world analytical workloads on top of a local cluster consisting of 10 workstations interconnected by a SDN-enabled switch. The experiment results clearly validate the effectiveness and efficiency of our proposed framework and algorithms. Walid Aljoby, Xin Wang 0040, Tom Z. J. Fu, Richard T. B. Ma |
ICNP | 3 |
| 2018 | A component-driven distributed framework for real-time video dehazing
Meihua Wang, Jiaming Mai, Yun Liang 0003, Ruichu Cai, Tom Z. J. Fu |
Multim. Tools Appl. | 5 |
| 2018 | Distributed Stream Rebalance for Stateful Operator Under Workload VarianceabstractKey-based workload partitioning is now commonly used in parallel stream processing, enabling effective key-value tuple distribution over worker threads in a logical operator. While randomized hashing on the keys is capable of balancing the workload for key-based partitioning when the keys generally follow a static distribution, it is likely to generate poor balancing performance when workload variance occurs on the incoming data stream. This paper presents a new key-based workload partitioning framework, with practical algorithms to support dynamic workload assignment for stateful operators. The framework combines hash-based and explicit key-based routing strategies for workload distribution, which specifies the destination worker threads for a handful of keys and assigns the other keys with the hashing function. We formulate the rebalance operation as an optimization problem, with multiple objectives on minimizing state migration costs, controlling the size of the routing table and breaking workload imbalance among the worker threads. Despite of the NP-hardness nature behind the optimization formulation, we carefully investigate and justify the heuristics behind key (re)routing and state migration, to facilitate fast response to workload variance with ignorable cost to the normal processing in the distributed system. Empirical studies on synthetic data and real-world stream applications validate the usefulness of our proposals and prove the huge advantage of our approaches over state-of-the-art solutions in the literature. Junhua Fang, Rong Zhang 0002, Tom Z. J. Fu, Aoying Zhou, Xiaofang Zhou 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Parallel Stream Processing Against Workload Skewness and VarianceabstractKey-based workload partitioning is a common strategy used in parallel stream processing engines, enabling effective key-value tuple distribution over worker threads in a logical operator. It is likely to generate poor balancing performance when workload variance occurs on the incoming data stream. This paper presents a new key-based workload partitioning framework, with practical algorithms to support dynamic workload assignment for stateful operators. The framework combines hash-based and explicit key-based routing strategies for workload distribution, which specifies the destination worker threads for a handful of keys and assigns the other keys with the hash function. When short-term distribution fluctuations occur to the incoming data stream, the system adaptively updates the routing table containing the chosen keys, in order to rebalance the workload with minimal migration overhead within the stateful operator. We formulate the rebalance operation as an optimization problem, with multiple objectives on minimizing state migration costs, controlling the size of the routing table and breaking workload imbalance among worker threads. Despite of the NP-hardness nature behind the optimization formulation, we carefully investigate and justify the heuristics behind key (re)routing and state migration, to facilitate fast response to workload variance with ignorable cost to the normal processing in the distributed system. Empirical studies on synthetic data and real-world stream applications validate the usefulness of our proposals. Junhua Fang, Rong Zhang 0002, Tom Z. J. Fu, Aoying Zhou, Junhua Zhu |
HPDC | 3 |
| 2017 | Distributed Publish/Subscribe Query Processing on the Spatio-Textual Data StreamabstractHuge amount of data with both space and text information, e.g., geo-tagged tweets, is flooding on the Internet. Such spatio-textual data stream contains valuable information for millions of users with various interests on different keywords and locations. Publish/subscribe systems enable efficient and effective information distribution by allowing users to register continuous queries with both spatial and textual constraints. However, the explosive growth of data scale and user base has posed challenges to the existing centralized publish/subscribe systems for spatiotextual data streams. In this paper, we propose our distributed publish/subscribe system, called PS2Stream, which digests a massive spatio-textual data stream and directs the stream to target users with registered interests. Compared with existing systems, PS2Stream achieves a better workload distribution in terms of both minimizing the total amount of workload and balancing the load of workers. To achieve this, we propose a new workload distribution algorithm considering both space and text properties of the data. Additionally, PS2Stream supports dynamic load adjustments to adapt to the change of the workload, which makes PS2Stream adaptive. Extensive empirical evaluation, on commercial cloud computing platform with real data, validates the superiority of our system design and advantages of our techniques on system performance improvement. Zhida Chen, Gao Cong, Tom Z. J. Fu, Lisi Chen 0001 |
ICDE | 4 |
| 2017 | Impacts of task placement and bandwidth allocation on stream analyticsabstractWe consider data intensive cloud-based stream analytics where data transmission through the underlying communication network is the cause of the performance bottleneck. Two key inter-related problems are investigated: task placement and bandwidth allocation. We seek to answer the following questions. How does task placement make impact on the application-level throughput? Does a careful bandwidth allocation among data flows traversing a bottleneck link results in better performance? In this paper, we address these questions by conducting measurement-driven analysis in a SDN-enabled computer cluster running stream processing applications on top of Apache Storm. The results reveal (i) how tasks are assigned to computing nodes make large difference in application level performance; (ii) under certain task placement, a proper bandwidth allocation helps further improve the performance as compared to the default TCP mechanism; and (iii) task placement and bandwidth allocation are collaboratively making effects in overall performance. Walid Aljoby, Tom Z. J. Fu, Richard T. B. Ma |
ICNP | 2 |
| 2017 | DITIR: Distributed Index for High Throughput Trajectory Insertion and Real-time Temporal Range QueryabstractThe prosperity of mobile social network and location-based services, e.g., Uber, is backing the explosive growth of spatial temporal streams on the Internet. It raises new challenges to the underlying data store system, which is supposed to support extremely high-throughput trajectory insertion and low-latency querying with spatial and temporal constraints. State-of-the-art solutions, e.g., HBase, do not render satisfactory performance, due to the high overhead on index update. In this demonstration, we present DITIR, our new system prototype tailored to efficiently processing temporal and spacial queries over historical data as well as latest updates. Our system provides better performance guarantee, by physically partitioning the incoming data tuples on their arrivals and exploiting a template-based insertion schema, to reach the desired ingestion throughput. Load balancing mechanism is also introduced to DITIR, by using which the system is capable of achieving reliable performance against workload dynamics. Our demonstration shows that DITIR supports over 1 million tuple insertions in a second, when running on a 10-node cluster. It also significantly outperforms HBase by 7 times on ingestion throughput and 5 times faster on query latency. Ruichu Cai, Zijie Lu, Tom Z. J. Fu, Marianne Winslett |
Proc. VLDB Endow. | 5 |
| 2017 | DRS: Auto-Scaling for Real-Time Stream AnalyticsabstractIn a stream data analytics system, input data arrive continuously and trigger the processing and updating of analytics results. We focus on applications with real-time constraints, in which, any data unit must be completely processed within a given time duration. To handle fast data, it is common to place the stream data analytics system on top of a cloud infrastructure. Because stream properties, such as arrival rates can fluctuate unpredictably, cloud resources must be dynamically provisioned and scheduled accordingly to ensure real-time responses. It is essential, for existing systems or future developments, to possess the ability of scaling resources dynamically according to the instantaneous workload, in order to avoid wasting resources or failing in delivering the correct analytics results on time. Motivated by this, we propose DRS, a dynamic resource scaling framework for cloud-based stream data analytics systems. DRS overcomes three fundamental challenges: 1) how to model the relationship between the provisioned resources and the application performance, 2) where to best place resources, and 3) how to measure the system load with minimal overhead. In particular, DRS includes an accurate performance model based on the theory of Jackson open queueing networks and is capable of handling arbitrary operator topologies, possibly with loops, splits, and joins. Extensive experiments with real data show that DRS is capable of detecting sub-optimal resource allocation and making quick and effective resource adjustment. Tom Z. J. Fu, Jianbing Ding, Richard T. B. Ma, Marianne Winslett, Yin Yang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | Cost-Effective Stream Join Algorithm on Cloud SystemabstractMatrix-based scheme (Join-Matrix) can prefectly support distributed stream joins, especially for arbitrary join predicates, because it guarantees any tuples from two streams to meet with each other. However,the dynamics and unpredictability features of stream require quick actions on scheme changing. Otherwise, they may lead to degradation of system throughputs and increament of processing latency with the waste of system resources, such as CPUs and Memories. Since Join-Matrix model has the fixed processing architecture with replicated data, these kinds of adverseness will be magnified. Therefore, it is urgent to find a solution that preserves advantages of Join-Matrix model and promises a good usage to computation resources when it meets scheme changing. In this paper, we propose a cost-effective stream join algorithm, which ensures the adaptability of Join-Matrix but with lower resources consumption. Specifically, a varietal matrix generation algorithm is proposed to generate an irregular matrix scheme for assigning the minimal number of tasks; a lightweight migration algorithm is designed to ensure state migration at a low cost; a complete load balance process framework is described to guarantee the correctness during the scheme changing. We conduct extensive experiments to compare our method with baseline systems on both benchmarks and real-workloads, and explain the results in detail. Junhua Fang, Rong Zhang 0002, Tom Z. J. Fu, Aoying Zhou |
CIKM | 4 |
| 2015 | DRS: Dynamic Resource Scheduling for Real-Time Analytics over Fast StreamsabstractIn a data stream management system (DSMS), users register continuous queries, and receive result updates as data arrive and expire. We focus on applications with real-time constraints, in which the user must receive each result update within a given period after the update occurs. To handle fast data, the DSMS is commonly placed on top of a cloud infrastructure. Because stream properties such as arrival rates can fluctuate unpredictably, cloud resources must be dynamically provisioned and scheduled accordingly to ensure real-time response. It is essential, for the existing systems or future developments, to possess the ability of scheduling resources dynamically according to the current workload, in order to avoid wasting resources, or failing in delivering correct results on time. Motivated by this, we propose DRS, a novel dynamic resource scheduler for cloud-based DSMSs. DRS overcomes three fundamental challenges: (a) how to model the relationship between the provisioned resources and query response time (b) where to best place resources, and (c) how to measure system load with minimal overhead. In particular, DRS includes an accurate performance model based on the theory of Jackson open queueing networks and is capable of handling arbitrary operator topologies, possibly with loops, splits and joins. Extensive experiments with real data confirm that DRS achieves real-time response with close to optimal resource consumption. Tom Z. J. Fu, Jianbing Ding, Richard T. B. Ma, Marianne Winslett, Yin Yang 0001 |
ICDCS | 1 |
| 2015 | LiveTraj: Real-Time Trajectory Tracking over Live Video StreamsabstractWe present LiveTraj, a novel system for tracking trajectories in a live video stream in real time, backed by a cloud platform. Although trajectory tracking is a well-studied topic in computer vision, so far most attention has been devoted to improving the accuracy of trajectory tracking, rather than the efficiency. To our knowledge, LiveTraj is the first that achieves real-time efficiency in trajectory tracking, which can be a key enabler in many important applications such as video surveillance, action recognition and robotics. LiveTraj is based on a state-of-the-art approach to (offline) trajectory tracking; its main innovation is to adapt this base solution to run on an elastic cloud platform to achieve real-time tracking speed at an affordable cost. The video demo shows the offline base solution and LiveTraj side by side, both running on a video stream containing human actions. Besides demonstrating the real-time efficiency of LiveTraj, our video demo also exhibits important system parameters to the audience such as latency and cloud resource usage for different components of the system. Further, if the conference venue provides sufficiently fast Internet connection to our cloud platform, we also plan to demonstrate LiveTraj on-site, during which we will show LiveTraj identifying and tracking trajectories from a live video stream captured by a camera. Tom Z. J. Fu, Jianbing Ding, Richard T. B. Ma, Marianne Winslett, Yin Yang 0001, Yong Pei, Bingbing Ni |
ACM Multimedia | 1 |
| 2015 | Turbocharged Video Distribution via P2PabstractThere are two types of P2P systems satisfying two different user demands: 1) file downloading and 2) video-on-demand (VoD) streaming. An example of file downloading is the original BitTorrent, and examples for VoD streaming include various commercial P2P-based VoD streaming systems such as that offered by PPLive. We have a hypothesis - by combining a type: 1) system and 2) system as a single P2P system, both the file downloading users and the streaming users of the same video will benefit in performance. The reasoning is that at any moment, only a subset of the file downloading peers can provide good service to VoD streaming peers and the VoD streaming peers are only good at providing service to a different subset of the file downloading peers. The former subset is the set of peers close to completing the downloading of the video file; whereas the latter subset is the set of peers starting to download a video. In this paper, we propose a novel design for a mesh-based video distribution system without depending on video replication on streaming peers. We produce simple back-of-the-envelop analysis to show its effectiveness. Then, we further validate our design and compare it with other designs through simulation and experiments in practical networking environment by implementing a prototype. Yipeng Zhou, Liang Chen 0009, Tom Z. J. Fu, Dah-Ming Chiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | A Unifying Model and Analysis of P2P VoD Replication and SchedulingabstractWe consider a peer-to-peer (P2P)-assisted video-on-demand (VoD) system where each peer can store a relatively small number of movies to offload the server when these movies are requested. User requests are stochastic based on some movie popularity distribution. The problem is how to replicate (or place) content at peer storage to minimize the server load. Several variations of this replication problem have been studied recently with somewhat different conclusions. In this paper, we first point out and explain that the main difference between these studies is in how they model the scheduling of peers to serve user requests, and show that these different scheduling assumptions will lead to different “optimal” replication strategies. We then propose a unifying request scheduling model, parameterized by the maximum number of peers that can be used to serve a single request. This scheduling is called Fair Sharing with Bounded Degree (FSBD). Based on this unifying model, we can compare the different replication strategies for different degree bounds and see how and why different replication strategies are favored depending on the degree. We also propose a simple (primarily) distributed replication algorithm and show that this algorithm is able to adapt itself to work well for different degrees in scheduling. Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
IEEE/ACM Trans. Netw. | 2 |
| 2013 | An Adaptive Cloud Downloading ServiceabstractVideo content downloading using the P2P approach is scalable, but does not always give good performance. Recently, subscription-based premium services have emerged, referred to as cloud downloading. In this service, the cloud storage and server caches user-interested content and updates the cache based on user downloading requests. If a requested video is not in the cache, the request is held in a waiting state until the cache is updated. We call this design server mode. An alternative design is to let the cloud server serve all downloading requests as soon as they arrive, behaving as a helper peer. We call this design helper mode. Our model and analysis show that both these designs are useful for certain operating regimes. The helper mode is good at handling a high request rate, while the server mode is good at scaling with video population size. We design an adaptive algorithm (AMS) to select the service mode automatically. Intuitively, AMS switches service mode from server mode to helper mode when too many peers request blocked movies, and vice versa. The ability of AMS to achieve good performance in different operating regimes is validated by simulation . Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
IEEE Trans. Multim. | 2 |
| 2013 | On Replication Algorithm in P2P VoDabstractTraditional video-on-demand (VoD) systems rely purely on servers to stream video content to clients, which does not scale. In recent years, peer-to-peer assisted VoD (P2P VoD) has proven to be practical and effective. In P2P VoD, each peer contributes some storage to store videos (or segments of videos) to help the video server. Assuming peers have sufficient bandwidth for the given video playback rate, a fundamental question is what is the relationship between the storage capacity (at each peer), the number of videos, the number of peers, and the resultant off-loading of video server bandwidth. In this paper, we use a simple statistical model to derive this relationship. We propose and analyze a generic replication algorithm Random with Load Balancing (RLB) that balances the service to all movies for both deterministic and random (but stationary) demand models and both homogeneous and heterogeneous peers (in upload bandwidth). We use simulation to validate our results for sensitivity analysis and for comparisons to other popular replication algorithms. This study leads to several fundamental insights for P2P VoD system design in practice. Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
IEEE/ACM Trans. Netw. | 2 |
| 2012 | A unifying model and analysis of P2P VoD replication and schedulingabstractWe consider a P2P-assisted Video-on-Demand (VoD) system where each peer can store a relatively small number of movies to offload the server when these movies are requested. User requests are stochastic based on some movie popularity distribution. The problem is how to replicate (or place) content at peer storage to minimize the server load. Several variation of this replication problem have been studied recently with somewhat different conclusions. In this paper, we first point out that the main difference between these studies is in how they model the scheduling of peers to serve user requests, and show that these different scheduling assumptions will lead to different “optimal” replication strategies. We then propose a unifying request scheduling model, parameterized by the maximum number of peers that can be used to serve a single request. This scheduling is called Fair Sharing with Bounded Out-Degree (FSBD). Based on this unifying model, we can compare the different replication strategies for different out-degree bounds and see how and why different replication strategies are favored depending on the out-degree. We also propose a new simple, adaptive, and essentially distributed replication algorithm, and show that this algorithm is able to adapt itself to work well for different out-degree in scheduling. Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
INFOCOM | 2 |
| 2012 | Division-of-labor between server and P2P for streaming VoDabstractWe consider a P2P-assisted content storage and delivery system to support a streaming Video-on-Demand (VoD) service. In this system, the peers are part of the service provider (e.g. set-top boxes) with limited storage space. Servers with ample storage and bandwidth are deployed to guarantee the availability and quality, but it is desirable to minimize the server utilization to reduce costs. Based on experience of implementing a deployed P2P VoD system, it was suggested in [1] that a movie's availability should be proportional to the movie's popularity. Based on further refinement, it is observed [2] that performance can be further improved by more (than proportional) availability for cold movies in P2P system. In this paper, we show that as the number of movies becomes large and there is some skewness in movie popularity, then one cannot expect the P2P part of the system to reduce server load as well as provide availability to all movies at the same time. It is a trade-off between coverage of movies and streaming throughput provided by the P2P system. If the goal is to minimize server load, under some reasonable conditions, we show that it is best to store and replicate only the hottest K* movies in the P2P part of the system. We also study the relationship between the skewness of the movie popularity distribution, P2P resources and the value of K*. Finally, we use simulation to validate our results. Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
IWQoS | 2 |
| 2012 | Server-assisted adaptive video replication for P2P VoD
Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
Signal Process. Image Commun. | 2 |
| 2011 | Statistical modeling and analysis of P2P replication to support VoD serviceabstractTraditional Video-on-Demand (VoD) systems reply purely on servers to stream video content to clients, which does not scale. In recent years, Peer-to-peer assisted VoD (P2P VoD) has proven to be practical and effective. In P2P VoD, each peer contributes some storage to store videos (or segments of videos) to help the video server. Assuming peers have sufficient bandwidth for the given video playback rate, a fundamental question is what is the relationship between the storage capacity (at each peer), the number of videos, the number of peers and the resultant off-loading of video server bandwidth. In this paper, we use a simple statistical model to derive this relationship. We propose and analyze a generic replication algorithm RLB which balances the service to all movies, for both deterministic and random demand models, and both homogeneous and heterogeneous peers (in upload bandwidth). We use simulation to validate our results, for sensitivity analysis and for comparisons with other popular replication algorithms. This study leads to several fundamental insights for design P2P VoD systems in practice. Yipeng Zhou, Tom Z. J. Fu, Dah-Ming Chiu |
INFOCOM | 2 |
| 2011 | Perceptual quality assessment on B-D tradeoff of P2P assisted layered video streamingabstractIn this paper, we study the impact of the image quality (hence, the video streaming bit-rate) and the chunk-level impairment (hence the playback discontinuity) on the perceptual assessment of the viewing experience of P2P layered video streaming services. Through subjective QoE experiments, we have obtained several interesting findings and useful insights. The results clearly reveal the tradeoffs between streaming bit-rate (B) and discontinuity (D) on the MOS under certain network condition. It is also showed that B-D tradeoff pattern is video content dependant. Based on these observations, a heuristic model is proposed to derive a group of MOS contours. This type of contours have practical usage and can help selecting suitable video streaming bit-rate (layer) for playback so as to maximize the perceptual QoE of the end users. Tom Z. J. Fu, Dah-Ming Chiu, Zhibin Lei |
VCIP | 2 |
| 2011 | Design and evaluation of load balancing algorithms in P2P streaming protocols
Tom Z. J. Fu, Dah-Ming Chiu |
Comput. Networks | 2 |
| 2010 | Designing QoE experiments to evaluate peer-to-peer streaming applicationsabstractQuality of Experience (QoE) refers to subjective criteria for evaluating multimedia content. Methods have been devised to study the design of Voice over IP systems and video codecs. In recent years, due to more abundant network bandwidth, it has become quite popular to watch video streamed over the Internet, whether by clientserver method or through a peer-to-peer (P2P) network. In this paper, we describe our experience in conducting QoE studies of P2P streaming using a chunk-level model. Instead of considering fine-grained network service impairments such as bit errors, packet losses or delays, we focus on chunk level delays. We carry out some preliminary QoE experiments on low-bit rate, low-frame rate and low-resolution video (3L-video) sequences. We apply the chunk-level model to help improve the design of the P2P streaming algorithms, and the design of video players that playback network streamed video. Tom Z. J. Fu, Dah-Ming Chiu, Zhibin Lei |
VCIP | 1 |
| 2009 | PBS: Periodic Behavioral Spectrum of P2P Applications
Tom Z. J. Fu, Xingang Shi, Dah-Ming Chiu, John C. S. Lui |
PAM | 1 |
| 2008 | Challenges, design and analysis of a large-scale p2p-vod systemabstractP2P file downloading and streaming have already become very popular Internet applications. These systems dramatically reduce the server loading, and provide a platform for scalable content distribution, as long as there is interest for the content. P2P-based video-on-demand (P2P-VoD) is a new challenge for the P2P technology. Unlike streaming live content, P2P-VoD has less synchrony in the users sharing video content, therefore it is much more difficult to alleviate the server loading and at the same time maintaining the streaming performance. To compensate, a small storage is contributed by every peer, and new mechanisms for coordinating content replication, content discovery, and peer scheduling are carefully designed. In this paper, we describe and discuss the challenges and the architectural design issues of a large-scale P2P-VoD system based on the experiences of a real system deployed by PPLive. The system is also designed and instrumented with monitoring capability to measure both system and component specific performance metrics (for design improvements) as well as user satisfaction. After analyzing a large amount of collected data, we present a number of results on user behavior, various system performance metrics, including user satisfaction, and discuss what we observe based on the system design. The study of a real life system provides valuable insights for the future development of P2P-VoD technology. Tom Z. J. Fu, Dah-Ming Chiu, John C. S. Lui |
SIGCOMM | 2 |
| 2007 | Performance metrics and configuration strategies for group network communicationabstractThere is an increasing number of group-based multimedia applications over the Internet, for example, voice conference or multi-player games. For these applications, it is often necessary to select a strategy to distribute the multimedia streams or mixing the multimedia stream data so as to provide better quality of service (QoS) guarantees. However, there is no appropriate metrics to evaluate the QoS of a group multimedia session, despite abundant literature on how to evaluate the QoS for two-party communication (e.g. MOS, E-Model). In this paper, we propose a new measure which is called the group mean opinion score (GMOS). To leverage on existing work, our definition of GMOS is based on two-party MOS, hence, it can be estimated via measurement of network parameters and fitting these data into the E-Model. We conduct large scale experiments using the latest SKYPE conference software. We first calibrate the GMOS based on the subjective scores of our experiments, then for individual conference sessions, we check whether our approach can pick a server configuration strategy to achieve the best GMOS. The study shows our proposed methodology is very promising and the potential of applying to other group-based applications. Tom Z. J. Fu, Dah-Ming Chiu, John C. S. Lui |
IWQoS | 1 |