VLDB 2026 Research / reviewers in the wild / expert
Qile Wang
dblp:256/2523
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Elastic Scheduling for Mix-Flow in Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) is the most promising network infrastructure for various time-critical applications in Industry 4.0. However, industry applications generate a mix of time-triggered (TT) and event-triggered (ET) flows. Scheduling such mix-flows is a key challenge for TSN. Though current TSN scheduling mechanisms commonly provide deterministic transmission for TT flows with stringent latency requirements, they cannot flexibly accommodate ET flows, which are usually generated by emergency events. In this paper, we propose Elastic Backoff (EBO), a systematic solution for scheduling mix-flows in an elastic way. Our key insight is that the network resources should be reasonably allocated for ET flows while minimally impacting TT flows. To this end, we incorporate the elasticity into the TSN scheduling to make resource reservations for ET flows without hurting TT flows. We conduct extensive experiments on both testbeds and simulations. The evaluation results show that, compared with the state-of-the-art designs, EBO improves the schedulability of ET flows by up to 7.5×, while still ensuring the deterministic transmission of TT flows. Jiawei Huang 0001, Shengwen Zhou, Hui Li 0120, Yijun Li 0002, Qile Wang, Jishu Tian, Kengchang Chen |
ICDCS | 8 |
| 2025 | Leveraging Large Language Models for Review Classification and Rating Estimation of Mental Health ApplicationsabstractLarge Language Models (LLMs) can analyze large datasets semantically. However, research on applying LLMs for mental health text classification is relatively new and developing. Existing methods often use supervised, deep, and reinforcement learning, which rely heavily on fine-tuning and reward models. To investigate whether LLMs can assist in recommending mental health apps based on user reviews, our study collected approximately 200k user reviews from 73 mental health mobile applications. We instructed selected LLMs to classify individual reviews into 1-5 star ratings, subsequently averaging these results to derive an overall rating for each app reflecting current user feedback. While the best supervised learning method in our experiments achieved an F1-Score of 0.79 which required significantly more human effort, the GPT-4 and Gemini 1.5 Pro delivered a strong ‘out-of-the-box’ performance with an overall F1-Score of 0.76. We provide further statistical comparisons and discussions of the performance of these models for the text classification task. Using a crowdsourcing platform to determine agreement levels, we observed that human ratings align closely with GPT ratings. In addition, we analyze specific features and concerns highlighted in mental health app reviews. Alongside our analysis, we make our data available for further experimentation and benchmarking. Qile Wang, Moath Erqsous, Prerana Khatiwada, Abhishek Karwankar, Fatimah Mohammad Alhassan, Aishwarya Chandrasekaran, Benita Abraham, Faith Lovell, Andrew Anh Ngo, Matthew Louis Mauriello |
ICWSM | 1 |
| 2025 | LATA: A Pilot Study on LLM-Assisted Thematic Analysis of Online Social Network Data Generation ExperiencesabstractLarge Language Models (LLMs) have gained attention in research and industry, aiming to streamline processes and enhance text analysis performance. Thematic Analysis (TA), a prevalent qualitative method for analyzing interview content, often requires at least two human experts to review and analyze data. This study demonstrates the feasibility of LLM-Assisted Thematic Analysis (LATA) using GPT-4 and Gemini. Specifically, we conducted semi-structured interviews with 14 researchers to gather insights on their experiences generating and analyzing Online Social Network (OSN) communications datasets. Following Braun and Clarke's six-phase TA framework with an inductive approach, we initially analyzed our interview transcripts with human experts. Subsequently, we iteratively designed prompts to guide LLMs through a similar process. We compare and discuss the manually analyzed outcomes with responses generated by LLMs and achieve a cosine similarity score up to 0.76, demonstrating a promising prospect for LATA. Additionally, the study delves into researchers' experiences navigating the complexities of collecting and analyzing OSN data, offering recommendations for future research and application designers. Qile Wang, Moath Erqsous, Kenneth E. Barner, Matthew Louis Mauriello |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2025 | Progress-Aware Transmission Protocol for Efficient In-Network Aggregation in Distributed Machine LearningabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortunately, the phase of gradient aggregation has become the performance bottleneck for data-parallel DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. Besides, we reveal that existing INA solutions cannot provide the fairness performance among multiple jobs with varying number of workers. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. Moreover, to ensure the fair throughput among multiple jobs, we dynamically adjust the aggregator allocation for each job by tuning the number of hash operations. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 5 |
| 2024 | Achieving High Efficiency for Datacenter Multicast using Skewed Bloom FilterabstractMulticast serves as an important approach for one-to-many communication in data center networks. To reduce overhead and improve scalability, bloom filters are employed in current multicast approaches to store forwarding ports of switches. However, the well-known false positive issue of bloom filter incurs wrong forwarding behaviors and redundant traffic in multicast tree, degrading transmission efficiency and increasing the risk of data leakage. Inspired by the fact that, given the same false positive ratio, the switch in the upper layers of multicast tree generates more redundant traffic, we propose RSBF, a fine-grained and resource-aware multicast approach using skewed bloom filters. Specifically, RSBF maintains multiple bloom filters corresponding to different layers of multicast tree, and allocates more ample space to the bloom filter of the upper layer switches, thereby reducing the overall redundant traffic. The test results of large-scale simulation demonstrate that RSBF reduces both redundant traffic and header overhead by up to 64% and 49% compared with the state-of-the-art approaches, respectively. Jiawei Huang 0001, Hui Li 0120, Qile Wang, Sitan Li, Zhidong He, Wanchun Jiang |
ICPP | 6 |
| 2024 | Achieving Efficient Scheduling based on Accurate Measurement of Small Flows in Data CenterabstractIn modern data centers, many flow scheduling schemes are proposed to accelerate data transfer and improve user experience. However, these schemes assume ideally the prior knowledge of the flow size information, which, unfortunately, is hard to obtain without modifying data center applications. The sketch-based approaches measure the flow size at switch with a compact memory structure, high throughput, and acceptable accuracy loss. However, existing sketches commonly focus on large or specific flows, while most flows in data center networks are small, resulting in missing or overestimated size information about small flows. We propose Strainer Sketch, which enables accurate and fast measurement of small flows with small memory and flexible deployment in a variety of scheduling algorithms. Specifically, Strainer Sketch uses the hierarchical structure to mitigate hash collisions between large and small flows, and the probabilistic counting algorithm to mitigate overestimation due to hash collisions between small flows. Furthermore, we propose a packet scheduling algorithm SW-PIFO, which provides the flow discrimination for a huge number of small flows by using a limited number of queues. Through the testbed experiments and simulations of typical data center applications, we show that our scheme reduces the small flow completion time (FCT) by up to 56.7 <?TeX $\%$?> Math 1 compared with flow scheduling using classic sketches. Jiawei Huang 0001, Qile Wang, Yijun Li 0002, Sitan Li, Jingling Liu, Min Zhan, Jianxin Wang 0001 |
ICPP | 2 |
| 2024 | P2Sketch: Finding Persistent Items in Data Streams Based on Periodic ArrivalabstractFinding persistent items provides indispensable information for data stream tasks. However, accurately identifying persistent items becomes very challenging with the increasing volume of data streams in memory-constrained environments. Existing solutions for finding persistent items often rely solely on estimating the persistence of items to make replacement decisions, requiring sufficiently large memory for acceptable performance. However, persistent items are frequently erroneously replaced in scenarios with numerous non-persistent items, leading to suboptimal accuracy. To address this issue, we reveal that periodically arriving items provide another useful feature for finding persistent items. We further propose P2Sketch that selectively replaces non-persistent items and preserves persistent items based on multi-dimensional statistics of estimated persistence and periodicity of items. Specifically, P2Sketch leverages the characteristics of persistence and periodic arrival to replace stored items selectively. When multiple candidates map to the same bucket, we replace items with longer periodic intervals and smaller estimated persistence to ensure more persistent items are protected. Experimental results show that P2Sketch significantly improves the F1 score by 1.15x and reduces the ARE by 2.66x under the condition of 50KB of memory compared with the state-of-the-art solutions. Jiawei Huang 0001, Qile Wang, Hui Li 0120, Sitan Li |
IPCCC | 4 |
| 2023 | MEB: an Efficient and Accurate Multicast using Bloom Filter with Customized Hash FunctionabstractMulticast is widely used to support a huge range of applications with one-to-many or many-to-many communication patterns. However, multicast systems do not scale due to considerable state and communication overheads. Some stateful multicast approaches require maintaining the state of each multicast session at switches, thus incurring large memory overhead. Some stateless ones utilize Bloom filter (BF) to encode multicast tree into the packet header to minimize communication overhead, but potentially suffer from the substantial false positive due to the probabilistic nature of Bloom filter. In this paper, we propose a stateless multicast scheme MEB, which uses Bloom filter to achieve large-scale multicast communication with low error, small overhead and high scalability. Specifically, to control the rate of false positive, MEB elaborately selects the hash functions for Bloom filters when constructing the packet header at the sender side, and makes forwarding decision according to packet header at the switch with negligible overhead. We compare MEB against the state-of-the-art multicast system in large-scale simulations. The test results show that MEB reduces the traffic overhead by up to 70% with small error rate. Jiawei Huang 0001, Qile Wang, Jingling Liu, Shengwen Zhou, Zhidong He |
APNet | 3 |
| 2023 | PA-ATP: Progress-Aware Transmission Protocol for In-Network AggregationabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortu-nately, the phase of gradient aggregation has become the performance bottleneck for DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
ICNP | 5 |