VLDB 2026 Research / reviewers in the wild / expert
Yiyang Shao
dblp:157/9649
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlexSpark: Robust and Efficient Multi-Device Collaborative Inference over Wireless Network
Yiyang Shao, Shuihai Hu, Xinle Du, Jingbin Zhou |
APNet | 1 |
| 2025 | Learning Optimal Multimodal Information Bottleneck RepresentationsabstractLeveraging high-quality joint representations from multimodal data can greatly enhance model performance in various machine-learning based applications. Recent multimodal learning methods, based on the multimodal information bottleneck (MIB) principle, aim to generate optimal MIB with maximal task-relevant information and minimal superfluous information via regularization. However, these methods often set regularization weights in an ad hoc manner and overlook imbalanced task-relevant information across modalities, limiting their ability to achieve optimal MIB. To address this gap, we propose a novel multimodal learning framework, Optimal Multimodal Information Bottleneck (OMIB), whose optimization objective guarantees the achievability of optimal MIB by setting the regularization weight within a theoretically derived bound. OMIB further addresses imbalanced task-relevant information by dynamically adjusting regularization weights per modality, ensuring the inclusion of all task-relevant information. Moreover, we establish a solid information-theoretical foundation for OMIB’s optimization and implement it under the variational approximation framework for computational efficiency. Finally, we empirically validate the OMIB’s theoretical properties on synthetic data and demonstrate its superiority over the state-of-the-art benchmark methods in various downstream tasks. Yiyang Shao, Jun Wang 0018 |
ICML | 2 |
| 2025 | Filtering Discomforting Recommendations with Large Language ModelsabstractPersonalized algorithms can inadvertently expose users to discomforting recommendations, potentially triggering negative consequences. The subjectivity of discomfort and the black-box nature of these algorithms make it challenging to effectively identify and filter such content. To address this, we first conducted a formative study to understand users' practices and expectations regarding discomforting recommendation filtering. Then, we designed a Large Language Model (LLM)-based tool named DiscomfortFilter, which constructs an editable preference profile for a user and helps the user express filtering needs through conversation to mask discomforting preferences within the profile. Based on the edited profile, DiscomfortFilter facilitates the discomforting recommendations filtering in a plug-and-play manner, maintaining flexibility and transparency. The constructed preference profile improves LLM reasoning and simplifies user alignment, enabling a 3.8B open-source LLM to rival top commercial models in an offline proxy task. A one-week user study with 24 participants demonstrated the effectiveness of DiscomfortFilter, while also highlighting its potential impact on platform recommendation outcomes. We conclude by discussing the ongoing challenges, highlighting its relevance to broader research, assessing stakeholder impact, and outlining future research directions. Jiahao Liu 0009, Yiyang Shao, Peng Zhang 0060, Dongsheng Li 0002, Hansu Gu, Chao Chen 0016, Longzhi Du, Tun Lu, Ning Gu 0001 |
WWW | 2 |
| 2025 | SoulSearch: applying heuristic optimization to enhance text-to-image generation with personalized human-LMM collaboration
Yubo Shu, Peng Zhang 0060, Hansu Gu, Yaqiong Li, Yiyang Shao, Tun Lu, Ning Gu 0001 |
Sci. China Inf. Sci. | 6 |
| 2025 | DeMod: A Holistic Tool with Explainable Detection and Personalized Modification for Toxicity CensorshipabstractAlthough there have been automated approaches and tools supporting toxicity censorship for social posts, most of them focus on detection. Toxicity censorship is a complex process, wherein detection is just an initial task and a user can have further needs such as rationale understanding and content modification. For this problem, we conduct a need-finding study to investigate people's diverse needs in toxicity censorship and then build a ChatGPT-based censorship tool named DeMod accordingly. DeMod is equipped with the features of explainable De tection and personalized Mod ification, providing fine-grained detection results, detailed explanations, and personalized modification suggestions. We also implemented the tool and recruited 35 Weibo users for evaluation. The results suggest DeMod's multiple strengths like the richness of functionality, the accuracy of censorship, and ease of use. Based on the findings, we further propose several insights into the design of content censorship systems. Yaqiong Li, Peng Zhang 0060, Hansu Gu, Tun Lu, Siyuan Qiao, Yubo Shu, Yiyang Shao, Ning Gu 0001 |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2024 | Revisiting Congestion Control for WiFi NetworksabstractWiFi networks, widely utilized by wireless devices, have become increasingly complex and congested environments, leading to noticeable delays, jitter, and throughput degradation for end-to-end network flows in today’s Internet. Through detailed experimental observations, we identified the TCP victim problem in WiFi, where TCP erroneously detects congestion and interacts with WiFi routers, resulting in a significant throughput decrease for certain hosts. In this paper, we introduce Cupid, a novel congestion control algorithm that relies on receiver-side WiFi physical layer measurements. Cupid accurately assesses congestion by measuring parameters such as airtime utilization, concurrency, and rates, allowing for precise rate adjustments. Our results demonstrate that whether employed alone or alongside other congestion controls, Cupid effectively mitigates the TCP victim problem. Furthermore, when used independently, Cupid can also reduce latency. Xinle Du, Yiyang Shao, Wei Wang 0505, Shuihai Hu, Jingbin Zhou, Kun Tan 0003 |
APNet | 3 |
| 2024 | Performant TCP over Wi-Fi DirectabstractWi-Fi Direct has been serving a progressively wide range of applications such as device-to-device file sharing, face-to-face interactive gaming, and wireless projection. However, when TCP meets Wi-Fi Direct, we find that two independent control loops exist, i.e., the transport-layer control loop and the link-layer control loop. First, these functionally redundant loops result in spectrum inefficiency. Second, the lack of effective information interaction between layers results in local optimal. To tackle these issues, this paper proposes Wi-Fi Direct TCP (WDTCP), a performant TCP that provides a full protocol design of the acknowledgment de-redundancy and explicit-capacity-based congestion control. WDTCP tightly couples the two control loops by capturing the WiFi Direct’s key feature of one-hop communication. Evaluation results demonstrate that WDTCP can maximize bandwidth utilization while keeping low latency. For instance, compared to legacy TCP, WDTCP improves throughput by up to 49.2% and reduces average and 95th latency by up to 32.4% and 50.7%, respectively. Hanlin Huang, Ke Xu 0002, Xinle Du, Yiyang Shao, Tong Li 0014 |
IWQoS | 4 |
| 2022 | Muses: Enabling Lightweight Learning-Based Congestion Control for Mobile DevicesabstractVarious congestion control (CC) algorithms have been designed to target specific scenarios. To automate this process, researchers have begun to use machine learning to automatically control the congestion window. These, however, often rely on heavyweight learning models (e.g., neural networks). This can make them unsuitable for resource-constrained mobile devices. On the other hand, lightweight models (e.g., decision trees) are often incapable of reflecting the complexity of diverse mobile wireless environments. To address this, we present Muses, a learning-based approach for generating lightweight congestion control algorithms. Muses relies on imitation learning to train a universal (heavy) LSTM model, which is then used to extract (lightweight) decision tree models that are each targeted at an individual environment. Muses then dynamically selects the most appropriate decision tree on a per-flow basis. We show that Muses can generate high throughput policies across a diverse set of environments, and it is sufficiently light to operate on mobile devices. Zhiren Zhong, Wei Wang 0334, Yiyang Shao, Zhenyu Li 0001, Hongtao Guan, Gareth Tyson, Gaogang Xie, Kai Zheng 0003 |
INFOCOM | 3 |
| 2022 | A multi-flexible video summarization scheme using property-constraint decision tree
Xiaoyu Teng, Xiaolin Gui, Yiyang Shao, Jianglei Tong, Tianjiao Du, Huijun Dai |
Neurocomputing | 4 |
| 2021 | Leveraging Domain Knowledge for Robust Deep Reinforcement Learning in NetworkingabstractThe past few years has witnessed a surge of interest towards deep reinforcement learning (Deep RL) in computer networks. With extraordinary ability of feature extraction, Deep RL has the potential to re-engineer the fundamental resource allocation problems in networking without relying on pre-programmed models or assumptions about dynamic environments. However, such black-box systems suffer from poor robustness, showing high performance variance and poor tail performance. In this work, we propose a unified Teacher-Student learning framework that harnesses rich domain knowledge to improve robustness. The domain-specific algorithms, less performant but more trustable than Deep RL, play the role of teachers providing advice at critical states; the student neural network is steered to maximize the expected reward as usual and mimic the teacher's advice meanwhile. The Teacher-Student method comprises of three modules where the confidence check module locates wrong decisions and risky decisions, the reward shaping module designs a new updating function to incentive the learning of student network, and the prioritized experience replay module to effectively utilize the advised actions. We further implement our Teacher-Student framework in existing video streaming (Pensieve), load balancing (DeepLB) and TCP congestion control (Aurora). Experimental results manifest that the proposed approach reduces the performance standard deviation of DeepLB by 37%; it improves the 90th, 95th and 99th tail performance of Pensieve by 7.6%, 8.8%, 10.7% respectively; and it accelerates the rate of growth of Aurora by 2x at the initial stage, and achieves a more stable performance in dynamic environments. Ying Zheng 0004, Qingyang Duan, Lixiang Lin, Yiyang Shao, Wei Wang 0334, Xin Wang 0002, Yuedong Xu 0001 |
INFOCOM | 5 |
| 2017 | Chord: Checkpoint-based scheduling using hybrid waiting list in shared clusters
Yiyang Shao, Weidong Bao 0001, Xiaomin Zhu 0001, Wenhua Xiao, Jian Wang 0105 |
J. Syst. Softw. | 1 |
| 2017 | Student cluster competition: ParConnect reproducibility task report
Ying Hao Tan, Yiyang Shao, Bu-Sung Lee |
Parallel Comput. | 2 |
| 2016 | CHIME: A Checkpoint-Based Approach to Improving the Performance of Shared ClustersabstractDue to the limitation of resources, preemption frequently occurs in almost all the commercial cloud platforms, such as Google cluster and Amazon cluster. Since preemption can ensure that once the system is in heavy workload, high-priority tasks will be executed primarily and at the same time, some low-priority tasks will be killed immediately. Then when more resources are available, the killed tasks will restart to execute. Especially, during the peak time, some low-priority tasks could possibly be preempted and restarted repeatedly resulting in much more consuming precious resources including CPU cores, RAM and hard drives. Thanks to the checkpoint technology, it provides an efficient solution to addressing the preemption issue. But checkpoint technology has limitations, e.g., making checkpoint frequently will add redundant overhead to the cluster and cause I/O congestion. In this paper, by leveraging checkpoint technology, we designed a novel approach to improving the performance of shared clusters. Specifically, by checking the occupancy of resources periodically, making decisions to checkpoint or not and checkpointing for certain tasks, our method can reduce unnecessary checkpoints and exalt the performance of the whole cloud, especially tasks with low-priority. Extensive simulation experiments injecting tasks following the Google cloud trace logs were conducted to validate the superiority of our approach by comparing it with some baselines. Yiyang Shao, Xiaomin Zhu 0001, Weidong Bao 0001, Wen Zhou 0013, Wenhua Xiao |
ICPADS | 1 |