EDBT 2026 Demo / reviewers in the wild / expert
Zhenxiao Luo
dblp:66/2693
· DBLP profile ↗
16ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-4349-2194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAILS: A Synchronous Accessible Immersive Online Learning System for Young LearnersabstractOnline distance learning emerged as a prominent means of education during the pandemic and is expected to continue as a long-lasting trend, enabling access to top education resources without the constraints of physical distance. However, conferencing software predominantly used for synchronous online teaching and learning is not suitable or sufficient for young students and teachers. The learning experiences of these students heavily rely on hands-on activities and interactions with peers and teachers that are challenging to replicate. Additionally, teachers face difficulties in effectively monitoring student progress. To address these challenges, we propose a Synchronous Accessible Immersive Online Learning System (SAILS). It offers an immersive and interactive platform for young learners by integrating real-life school activities with a virtual learning environment. Using the system, teachers can easily organize classes, assess students' work. We evaluated the effectiveness of SAILS through a user study involving 40 young learners from a kindergarten and several primary schools, along with their teachers. The results demonstrated a strong preference for our system among the participants, highlighting its ability to provide a more immersive and engaging online learning experience. Yuran Sun, Zhuoying Zhang, Zhenxiao Luo, Alan William Dougherty, Man Ho Yip, Yi-King Choi, Chuan Wu 0001 |
MMSys | 3 |
| 2024 | SDSR: Optimizing Metaverse Video Streaming via Saliency-Driven Dynamic Super-ResolutionabstractMetaverse (especially 360-degree) video streaming allows broadcasting virtual events in the metaverse to a broad audience. To reduce the huge bandwidth consumption, quite a few super-resolution (SR)-enhanced 360-degree video streaming systems have been proposed. However, there is very limited work to investigate how the granularity of SR model affects the system performance, and how to choose a proper SR model for different video contents under diverse environmental conditions. In this paper, we first conduct a dedicated measurement study to unveil the impact of different granularities of SR models. It is found that the scene of a video largely determines the effectiveness of SR models in different granularities. Based on our observations, we propose a novel 360-degree video streaming framework with saliency-driven dynamic super-resolution, called SDSR. To maximize user QoE, we formally formulate an optimization problem and adopt the model predictive control (MPC) theory for bitrate adaptation and SR model selection. To improve the effectiveness of SR model, we leverage the saliency information, which well reflects users’ view interests, for model training. In addition, we reuse an SR model for similar chunks based on temporal redundancy of a video. Finally, we conduct extensive experiments on real traces and the results show that SDSR outperforms the state-of-the-art algorithms with an improvement up to 32.78% in terms of the average QoE. Baili Chai, Zhenxiao Luo, Miao Hu 0001, Yipeng Zhou, Di Wu 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2024 | Goku: A Schemaless Time Series Database for Large Scale Monitoring at PinterestabstractEngineers rely heavily on observability tools to monitor their business and system metrics and set up alerting on it. A reliable and efficient monitoring system is very important for development velocity. In this paper, we introduce Goku, a time series database (TSDB) we built from the ground up at Pinterest. Over the years, we have studied user patterns and common requests to constantly evolve Goku to store and serve the use cases at Pinterest with high efficiency and reduced costs. At its core, Goku uses tiered storage to store new and frequently queried metrics data in memory while leveraging solid state drive (SSD) and hard disks (HDD) for older data. Goku aggregates metrics data at write time while also rolling up datapoints with lower time granularity to low latency to certain use cases. Goku also supports modifying configurations on metrics data like time to live (TTL), rollup granularity, backfilling capability, etc. Using multiple replicas and AWS S3 as backup, Goku is highly available and fault tolerant. Monil Mukesh Sanghavi, Ming-May Hu, Zhenxiao Luo, Kapil Bajaj |
Proc. VLDB Endow. | 3 |
| 2024 | AutoFL: A Bayesian Game Approach for Autonomous Client Participation in Federated Edge LearningabstractGiven that devices (i.e., clients) participating in federated edge learning (FEL) are autonomous and resource-constrained in nature, it is critical to design effective incentive mechanisms to encourage client participation so as to improve the performance of FEL. In this article, we aim to boost the FEL training efficiency by answering how much compute resource should clients autonomously contribute to maximize their utilities. To this end, we develop AutoFL, an autonomous client participation decision framework for federated learning at the network edge without assuming that each client possesses complete information. We first model the problem of autonomous client participation as a Bayesian game with incomplete information, where each player in the game is associated with a set of types according to network conditions. We optimize an individual client's decision based on the dynamics of the population estimated following the Bayes rule. We prove that AutoFL can converge to a unique Bayesian Nash equilibrium point. Empirical results on three real datasets show that AutoFL achieves a higher model accuracy with only 15.5-24.5% model aggregation time per global training round, and its energy cost saving on mobile devices is 82.2-86.8% compared to the state-of-the-art algorithms. Moreover, we can achieve a 2.75-3.2x long-term fairness compared to classical solutions. Miao Hu 0001, Wenzhuo Yang, Zhenxiao Luo, Xuezheng Liu, Yipeng Zhou, Xu Chen 0004, Di Wu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | LiveSR: Enabling Universal HD Live Video Streaming With Crowdsourced Online LearningabstractThe high-definition (HD) live video streaming has gained significant popularity due to the rapid growth of 4 G/5 G and social media. However, for devices with constrained bandwidth, they still have no sufficient bandwidth to support HD live video streaming. In this paper, we propose a neural-enhanced HD live video streaming framework calledLiveSRto provide universal HD live video streaming for both bandwidth-constrained and bandwidth-rich devices. For bandwidth-constrained devices, LiveSR delivers low-quality video streams and then boosts video quality at the device side with super-resolution (SR) techniques. The difficulty lies in how to train the SR model with low cost and conduct quality enhancement in real time. To address these challenges, we design a crowdsourced online training method by exploiting computation resources and HD video data on bandwidth-rich devices in the same video channel. We also propose an imitation learning-based decision making algorithm to make downloading decisions for video chunks and SR models under limited bandwidth. We implement and evaluate our proposed LiveSR framework using real network traces, and the experiment results show that LiveSR outperforms all the other baseline approaches, with 65.5% improvement in terms of the average QoE and 5.7% in terms of video quality (i.e., PSNR), and the achieved frame rate can be as high as 30 frames per second. Zhenxiao Luo, Miao Hu 0001, Yipeng Zhou, Di Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Masked360: Enabling Robust 360-degree Video Streaming with Ultra Low Bandwidth Consumptionabstract360-degree video streaming has gained tremendous growth over the past years. However, the delivery of 360-degree videos over the Internet still suffers from the scarcity of network bandwidth and adverse network conditions (e.g., packet loss, delay). In this paper, we propose a practical neural-enhanced 360-degree video streaming framework called Masked360, which can significantly reduce bandwidth consumption and achieve robustness against packet loss. In Masked360, instead of transmitting the complete video frame, the video server only transmits a masked low-resolution version of each video frame to reduce bandwidth significantly. When delivering masked video frames, the video server also sends a lightweight neural network model called MaskedEncoder to clients. Upon receiving masked frames, the client can reconstruct the original 360-degree video frames and start playback. To further improve the quality of video streaming, we also propose a set of optimization techniques, such as complexity-based patch selection, quarter masking strategy, redundant patch transmission and enhanced model training methods. In addition to bandwidth savings, Masked360 is also robust to packet loss during the transmission, because packet losses can be concealed by the reconstruction operation performed by the MaskedEncoder. Finally, we implement the whole Masked360 framework and evaluate its performance using real datasets. The experimental results show that Masked360 can achieve 4K 360-degree video streaming with bandwidth as low as 2.4 Mbps. Besides, video quality of Masked360 is also improved significantly, with an improvement of 5.24-16.61% in terms of PSNR and 4.74-16.15% in terms of SSIM compared to other baselines. Zhenxiao Luo, Baili Chai, Miao Hu 0001, Di Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | Taming Hybrid-Cloud Fast and Scalable Graph Analytics at TwitterabstractWe have witnessed a boosted demand for graph analytics at Twitter in recent years, and graph analytics has become one of the key parts of Twitter's large-scale data analytics and machine learning for driving engagement, serving the most relevant content, and promoting healthier conversations. However, infrastructure for graph analytics has historically not been an area of investment at Twitter, resulting in a long timeline and huge engineering effort for each project to deal with graphs at the Twitter scale. How do we build a unified graph analytics user experience to fulfill modern data analytics on various graph scales spanning from thousands to hundreds of billions of vertices and edges? To bring fast and scalable graph analytics capability into production, we investigate the challenges we are facing in largescale graph analytics at Twitter and propose a unified graph analytics platform for efficient, scalable, and reliable graph analytics across on-premises and cloud to fulfill the requirements of diverse graph use cases and challenging scales. We also conduct quantitative benchmarking on Twitter's production-level graph use cases between popular graph analytics frameworks to certify our solution. Chunxu Tang, Zhenxiao Luo, Mainak Ghosh, Anneliese Lu, Ruchin Kabra, Nikhil Kantibhai Navadiya, Prachi Mishra, Prateek Mukhedkar, Vrushali Channapattan |
GLOBECOM | 3 |
| 2022 | From Batch Processing to Real Time Analytics: Running Presto® at ScaleabstractPresto is an open source distributed query engine used widely at Facebook, Uber, Twitter, Pinterest, and many other internet companies. Since open sourced in 2013, the Presto community has made several rounds of design and implementations, to support a variety of use cases, including interactive analytics, real time reporting and dashboard, ETL workloads, A/B testing, monitoring and alerts, etc. In this paper, we'd like to introduce some of the most important features and performance improvements the open source Presto community made in recent years, which enables companies running Presto at scale, supporting millions of queries per day, with hundreds of thousands of machines. Specifically, how Presto provides unified SQL on heterogeneous storage systems without data copy; how Presto deals with complex data, including nested columnar data and schema evolution; How Presto supports geospatial queries efficiently, and how file list cache works in Presto. We also talk about cluster federation, and Presto on cloud. Experimental results and our production experience could help others running interactive SQL systems at scale. Zhenxiao Luo, Lu Niu, Venki Korukanti, Yutian Sun, Masha Basmanova, Beinan Wang, Devesh Agrawal, Chunxu Tang, Girish Baliga, Maosong Fu |
ICDE | 1 |
| 2022 | Otus: A Gaze Model-based Privacy Control Framework for Eye Tracking ApplicationsabstractEye tracking techniques have been widely adopted by a wide range of devices (e.g., AR/VR headsets, smartphones) to enhance user experiences. However, eye gaze data is private in nature, which can reveal users’ psychological and physiological features. Privacy protection techniques can be incorporated to preserve privacy of eye tracking information. Yet, most existing solutions based on Differential Privacy (DP) mechanisms cannot well protect privacy for individual users without sacrificing user experience. In this paper, we are among the first to propose a novel gaze model-based privacy control framework called Otus for eye tracking applications, which incorporates local DP (LDP) mechanisms to preserve user privacy and improves user experience in the meanwhile. First, we conduct a measurement study on real traces to illustrate that direct noise injection on raw gaze trajectories can significantly lower the utility of gaze data. To preserve utility and privacy simultaneously, Otus injects noises in two steps: (1) Extracting model features from raw data to depict gaze trajectories on individual users; (2) Adding LDP noises into model features so as to protect privacy. On one hand, established models can be used to recover user gaze data in order to improve service quality of eye tracking applications. On the other hand, we only need to add LDP noises to distort a small number of model parameters rather than every point on a trajectory to preserve privacy, which has less impact on the utility of gaze data given the same privacy budget. By applying the tile view graph model in step (1), we illustrate the entire workflow of Otus and prove its privacy protection level. For evaluation, we conduct extensive experiments using real gaze traces and the results show that Otus can effectively protect privacy for individual users without significantly compromising gaze data utility. Miao Hu 0001, Zhenxiao Luo, Yipeng Zhou, Xuezheng Liu, Di Wu 0001 |
INFOCOM | 2 |
| 2022 | Revisiting super-resolution for internet video streamingabstractRecent advancements of neural-enhanced techniques, especially super-resolution (SR), show great potential in revolutionizing the landscape of Internet video delivery. However, there are still quite a few key questions (e.g., how to choose a proper resolution configuration for training samples, how to set the training patch size, how to perform the best patch selection, how to set the update frequency of SR model) that have not been well investigated and understood. In this paper, we perform a dedicated measurement study to revisit super-resolution techniques for Internet video streaming. Our measurements are based on real-world video datasets, and the results provide a number of important insights: (1) It is possible that the SR model trained with low-resolution patches (e.g., (540p, 1080p) pairs) can achieve almost the same performance as that trained with high-resolution patches (e.g., (1080p, 2160p) pairs); (2) Compared to the saliency of training patches, the size of training patches has little impact on the performance of trained SR model; (3) The improvement of video quality brought by more frequent SR model update is not very significant. We also discuss the implications of our findings for system design, and we believe that our work is essential for paving the way for the success of future neural-enhanced video streaming systems. Zhenxiao Luo, Miao Hu 0001, Di Wu 0001, Youlong Cao |
NOSSDAV | 2 |
| 2021 | A Performance Evaluation of Spark GraphFrames for Fast and Scalable Graph Analytics at TwitterabstractGraph analytics demand is emerging rapidly and has become one of the key parts of Twitter machine learning for driving engagement, serving most relevant content, and promoting healthier conversations. However, due to lack of infrastructure support for graph analytics, we are suffering from a long timeline and huge engineering effort for each project to deal with graphs at the Twitter scale, which blocks us from fast iteration. To bring fast and scalable graph analytics capability into Twitter, we adopted Spark GraphFrames and conducted a performance evaluation for a typical graph analytics use case -Connected Component - on one of the largest graphs at Twitter. Compared to our existing Scalding solution, Spark GraphFrames achieved more than 33x speedup in running time. Zhenxiao Luo, Chunxu Tang, Mainak Ghosh, Anneliese Lu, Ruchin Kabra, Nikhil Kantibhai Navadiya, Prachi Mishra, Prateek Mukhedkar, Vrushali Channapattan |
IEEE BigData | 2 |
| 2021 | Forecasting SQL Query Cost at TwitterabstractWith the advent of the Big Data era, it is usually computationally expensive to calculate the resource usages of a SQL query with traditional DBMS approaches. Can we estimate the cost of each query more efficiently without any computation in a SQL engine kernel? Can machine learning techniques help to estimate SQL query resource utilization? The answers are yes. We propose a SQL query cost predictor service, which employs machine learning techniques to train models from historical query request logs and rapidly forecasts the CPU and memory resource usages of online queries without any computation in a SQL engine. At Twitter, infrastructure engineers are maintaining a large-scale SQL federation system across on-premises and cloud data centers for serving ad-hoc queries. The proposed service can help to improve query scheduling by relieving the issue of imbalanced online analytical processing (OLAP) workloads in the SQL engine clusters. It can also assist in enabling preemptive scaling. Additionally, the proposed approach uses plain SQL statements for the model training and online prediction, indicating it is both hardware and software-agnostic. The method can be generalized to broader SQL systems and heterogeneous environments. The models can achieve 97.9% accuracy for CPU usage prediction and 97% accuracy for memory usage prediction. Chunxu Tang, Beinan Wang, Zhenxiao Luo, Shajan Dasan, Maosong Fu, Mainak Ghosh, Ruchin Kabra, Nikhil Kantibhai Navadiya, Da Cheng, Fred Dai, Vrushali Channapattan, Prachi Mishra |
IC2E | 3 |
| 2021 | CrowdSR: enabling high-quality video ingest in crowdsourced livecast via super-resolutionabstractThe prevalence of personal devices motivates the rapid development of crowdsourced livecast in recent years. However, there exists huge diversity of upstream bandwidth among amateur broadcasters. Moreover, the highest video quality that can be streamed is limited by the hardware configuration of broadcaster devices (e.g., 540p for low-end mobile devices). The above factors pose significant challenges to the ingestion of high-resolution live video streams, and result in poor quality-of-experience (QoE) for viewers. In this paper, we propose a novel live video ingest approach called CrowdSR for crowdsourced livecast. CrowdSR can transform a low-resolution video stream uploaded by weak devices into a high-resolution video stream via super-resolution, and then deliver the stream to viewers. CrowdSR can exploit crowdsourced high-resolution video patches from similar broadcasters to speedup model training. Different from previous work, our approach does not require any modification at the client side, and thus is more practical and easy to implement. Finally, we implement and evaluate CrowdSR by conducting a series of real-world experiments. The results show that CrowdSR significantly outperforms the baseline approaches by 0.42-1.09 dB in terms of PSNR and 0.006-0.014 in terms of SSIM. Zhenxiao Luo, Miao Hu 0001, Yipeng Zhou, Tom Z. J. Fu, Di Wu 0001 |
NOSSDAV | 1 |
| 2020 | SR360: boosting 360-degree video streaming with super-resolutionabstract360-degree videos have gained increasing popularity due to its capability to provide users with immersive viewing experience. Given the limited network bandwidth, it is a common approach to only stream video tiles in the user's Field-of-View (FoV) with high quality. However, it is difficult to perform accurate FoV prediction due to diverse user behaviors and time-varying network conditions. In this paper, we re-design the 360-degree video streaming systems by leveraging the technique of super-resolution (SR). The basic idea of our proposed SR360 framework is to utilize abundant computation resources on the user devices to trade off a reduction of network bandwidth. In the SR360 framework, a video tile with low resolution can be boosted to a video tile with high resolution using SR techniques at the client side. We adopt the theory of deep reinforcement learning (DRL) to make a set of decisions jointly, including user FoV prediction, bitrate allocation and SR enhancement. By conducting extensive trace-driven evaluations, we compare the performance of our proposed SR360 with other state-of-the-art methods and the results show that SR360 significantly outperforms other methods by at least 30% on average under different QoE metrics. Miao Hu 0001, Zhenxiao Luo, Di Wu 0001 |
NOSSDAV | 3 |
| 2009 | Identification of ambiguous queries in web search
Ruihua Song, Zhenxiao Luo, Jian-Yun Nie, Yong Yu 0001, Hsiao-Wuen Hon |
Inf. Process. Manag. | 2 |
| 2007 | Identifying ambiguous queries in web searchabstractIt is widely believed that some queries submitted to search engines are by nature ambiguous (e.g., java, apple). However, few studies have investigated the questions of "how many queries are ambiguous?" and "how can we automatically identify an ambiguous query?" This paper deals with these issues. First, we construct the taxonomy of query ambiguity, and ask human annotators to manually classify queries based upon it. From manually labeled results, we find that query ambiguity is to some extent predictable. We then use a supervised learning approach to automatically classify queries as being ambiguous or not. Experimental results show that we can correctly identify 87% of labeled queries. Finally, we estimate that about 16% of queries in a real search log are ambiguous. Ruihua Song, Zhenxiao Luo, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon |
WWW | 2 |