VLDB 2026 Research / reviewers in the wild / expert
Yukai Miao
dblp:204/0178
· DBLP profile ↗
12ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-6401-5979ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeraClass: Towards Open-World Network Flow Classification via Traffic-Language Mapping
Ni Jin, Libin Liu 0001, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Xiuting Xu, Baojiang Cui |
IWQoS | 3 |
| 2026 | Example Generalizing Network Configuration Synthesizer via Graph-Informed Large Language Models
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao, Liyu Ma |
IEEE Trans. Netw. | 4 |
| 2025 | Towards Automatic Network Diagram ComprehensionabstractNetwork Diagram Comprehension (NDC) is a vital task for networking professionals, offering essential insights into network topology and configurations. However, NDC remains a labor-intensive process heavily reliant on human expertise, with existing tools falling short in addressing this challenge. It is critical to develop an Automatic NDC (ANDC) system that ensures high faithfulness and completeness in information extraction while supporting practical, end-to-end NDC applications. Moreover, a comprehensive dataset and benchmark are necessary to systematically evaluate and drive the progress of ANDC.In this work, we introduce Layered Extractor of Network Diagrams (LEND), the first ANDC system designed to comprehensively and faithfully extract and utilize information from network diagrams. LEND employs a three-stage pipeline: (1) a layer extractor to decompose diagrams and identify key elements with a denoising cascade, (2) an inter-layer combiner to reconstruct entity relations with positional and domain knowledge, and (3) a task-specific interpreter for networking applications.To support this effort, we develop two extensive NDC datasets comprising over 4,000 network diagrams and icons from diverse sources, along with the first benchmark to evaluate ANDC systems across three distinct metrics. Empirical experiments demonstrate that LEND outperforms existing methods by achieving at 1.21– 5.10× better faithfulness and completeness, and improves its capability as a NetOps engineer by 30.5% on the Cisco Certified Network Associate (CCNA) exam. Yanyu Ren, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yu Bai 0021 |
ICNP | 2 |
| 2025 | Transcending Cost-Quality Tradeoff in Agent Serving via Session-AwarenessabstractLarge Language Model (LLM) agents are capable of task execution across various domains by autonomously interacting with environments and refining LLM responses based on feedback.
However, existing model serving systems are not optimized for the unique demands of serving agents. Compared to classic model serving, agent serving has different characteristics:
predictable request pattern, increasing quality requirement, and unique prompt formatting. We identify a key problem for agent serving: LLM serving systems lack session-awareness. They neither perform effective KV cache management nor precisely select the cheapest yet competent model in each round.
This leads to a cost-quality tradeoff, and we identify an opportunity to surpass it in an agent serving system.
To this end, we introduce AgServe for AGile AGent SERVing.
AgServe features a session-aware server that boosts KV cache reuse via Estimated-Time-of-Arrival-based eviction and in-place positional embedding calibration, a quality-aware client that performs session-aware model cascading through real-time quality assessment, and a dynamic resource scheduler that maximizes GPU utilization.
With AgServe, we allow agents to select and upgrade models during the session lifetime, and to achieve similar quality at much lower costs, effectively transcending the tradeoff. Extensive experiments on real testbeds demonstrate that AgServe (1) achieves comparable response quality to GPT-4o at a 16.5\% cost. (2) delivers 1.8$\times$ improvement in quality relative to the tradeoff curve. Yanyu Ren, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yukai Miao, Yu Bai 0021 |
NeurIPS | 6 |
| 2025 | CEGS: Configuration Example Generalizing Synthesizer
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao |
NSDI | 4 |
| 2025 | Resolving Packets from Counters: Enabling Multi-scale Network Traffic Super Resolution via Composable Large Traffic Model
Xizheng Wang, Libin Liu 0001, Li Chen 0008, Dan Li 0001, Yukai Miao, Yu Bai 0021 |
NSDI | 5 |
| 2024 | ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option Combinations with Only Documentation via Large Language ModelabstractVulnerabilities related to option combinations pose a significant challenge in software security testing due to their vast search space. Previous research primarily addressed this challenge through mutation or filtering techniques, which inefficiently treated all option combinations as having equal potential for vulnerabilities, thus wasting considerable time on non-vulnerable targets and resulting in low testing efficiency. In this paper, we utilize carefully designed prompt engineering to drive the large language model (LLM) to predict high-risk option combinations (i.e., more likely to contain vulnerabilities) and perform fuzz testing automatically without human intervention. We developed a tool called ProphetFuzz and evaluated it on a dataset comprising 52 programs collected from three related studies. The entire experiment consumed 10.44 CPU years. ProphetFuzz successfully predicted 1748 high-risk option combinations at an average cost of only \8.69 per program. Results show that after 72 hours of fuzzing, ProphetFuzz discovered 364 unique vulnerabilities associated with 12.30% of the predicted high-risk option combinations, which was 32.85% higher than that found by state-of-the-art in the same timeframe. Additionally, using ProphetFuzz, we conducted persistent fuzzing on the latest versions of these programs, uncovering 140 vulnerabilities, with 93 confirmed by developers and 21 awarded CVE numbers. Dawei Wang 0021, Geng Zhou, Li Chen 0008, Dan Li 0001, Yukai Miao |
CCS | 5 |
| 2024 | BClean: A Bayesian Data Cleaning SystemabstractThere is a considerable body of work on data cleaning which employs various principles to rectify erroneous data and transform a dirty dataset into a cleaner one. One of prevalent approaches is probabilistic methods, including Bayesian methods. However, existing probabilistic methods often assume a simplistic distribution (e.g., Gaussian distribution), which is frequently under-fitted in practice, or they necessitate experts to provide a complex prior distribution (e.g., via a programming language). This requirement is both labor-intensive and costly, rendering these methods less suitable for real-world applications. In this paper, we propose BClean, a Bayesian Cleaning system that features automatic Bayesian network construction and user interaction. We recast the data cleaning problem as a Bayesian inference that fully exploits the relationships between attributes in the observed dataset and any prior information provided by users. To this end, we present an automatic Bayesian network construction method that extends a structure learning-based functional dependency discovery method with similarity functions to capture the relationships between attributes. Furthermore, our system allows users to modify the generated Bayesian network in order to specify prior information or correct inaccuracies identified by the automatic generation process. We also design an effective scoring model (called the compensative scoring model) necessary for the Bayesian inference. To enhance the efficiency of data cleaning, we propose several approximation strategies for the Bayesian inference, including graph partitioning, domain pruning, and pre-detection. By evaluating on both real-world and synthetic datasets, we demonstrate that BClean is capable of achieving an F-measure of up to 0.9 in data cleaning, outperforming existing Bayesian methods by 2% and other data cleaning methods by 15%. Jianbin Qin, Sifan Huang, Yaoshu Wang, Yukai Miao, Rui Mao 0001, Makoto Onizuka, Chuan Xiao 0001 |
ICDE | 6 |
| 2022 | Software-defined network assimilation: bridging the last mile towards centralized network configuration management with NAssimabstractOn-boarding new devices into an existing SDN network is a pain for network operations (NetOps) teams, because much expert effort is required to bridge the gap between the configuration models of the new devices and the unified data model in the SDN controller. In this work, we present an assistant framework NAssim, to help NetOps accelerate the process of assimilating a new device into a SDN network. Our solution features a unified parser framework to parse diverse device user manuals into preliminary configuration models, a rigorous validator that confirm the correctness of the models via formal syntax analysis, model hierarchy validation and empirical data validation, and a deep-learning-based mapping algorithm that uses state-of-the-art neural language processing techniques to produce human-comprehensible recommended mapping between the validated configuration model and the one in the SDN controller. In all, NAssim liberates the NetOps from most tedious tasks by learning directly from devices' manuals to produce data models which are comprehensible by both the SDN controller and human experts. Our evaluation shows, NAssim can accelerate the assimilation process by 9.1x. In this process, we also identify and correct 243 errors in four mainstream vendors' device manuals, and release a validated and expert-curated dataset of parsed manual corpus for future research. Huangxun Chen, Yukai Miao, Li Chen 0008, Haifeng Sun 0001, Hong Xu 0001, Libin Liu 0001, Gong Zhang 0001, Wei Wang 0011 |
SIGCOMM | 2 |
| 2021 | Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionabstractBERT has set a new state-of-the-art performance on entity resolution (ER) task, largely owed to fine-tuning pre-trained language models and the deep pair-wise interaction. Albeit being remarkably effective, it comes with a steep increase in computational cost, as the deep-interaction requires to exhaustively compute every tuple pair to search for co-references. For ER task, it is often prohibitively expensive due to the large cardinality to be matched. To tackle this, we introduce a siamese network structure that independently encodes tuples using BERT but delays the pair-wise interaction via an enhanced alignment network. This siamese structure enables a dedicated blocking module to quickly filter out obviously dissimilar tuple pairs, and thus drastically reduces the cardinality of fine-grained matching. Further, the blocking and entity matching are integrated into a multi-task learning framework for facilitating both tasks. Extensive experiments on multiple datasets demonstrate that our model significantly outperforms state-of-the-art models (including BERT) in both efficiency and effectiveness. Bing Li 0002, Yukai Miao, Yaoshu Wang, Yifang Sun, Wei Wang 0011 |
AAAI | 2 |
| 2020 | A Recurrent Model for Collective Entity Linking with Adaptive FeaturesabstractThe vast amount of web data enables us to build knowledge bases with unprecedented quality and coverage. Named Entity Disambiguation (NED) is an important task that automatically resolves ambiguous mentions in free text to correct target entries in the knowledge base. Traditional machine learning based methods for NED were outperformed and made obsolete by the state-of-the-art deep learning based models. However, deep learning models are more complex, requiring large amount of training data and lengthy training and parameter tuning time. In this paper, we revisit traditional machine learning techniques and propose a light-weight, tuneable and time-efficient method without using deep learning or deep learning generated features. We propose novel adaptive features that focus on extracting discriminative features to better model similarities between candidate entities and the mention's context. We learn a local ranking model based on traditional and the new adaptive features based on the learning-to-rank framework. While arriving at linking decisions individually via the local model, our method also takes into consideration the correlation between decisions by running multiple recurrent global models, which can be deemed as a learned local search method. Our method attains performances comparable to the state-of-the-art deep learning-based methods on NED benchmark datasets while being significantly faster to train. Xiaoling Zhou, Yukai Miao, Wei Wang 0011, Jianbin Qin |
AAAI | 2 |
| 2017 | Graph Summarization for Entity Relatedness VisualizationabstractIn modern search engines, Knowledge Graphs have become a key component for knowledge discovery. When a user searches for an entity, the existing systems usually provide a list of related entities, but they do not necessarily give explanations of how they are related. However, with the help of knowledge graphs, we can generate relatedness graphs between any pair of existing entities. Existing methods of this problem are either graph-based or list-based, but they all have some limitations when dealing with large complex relatedness graphs of two related entity. In this work, we investigate how to summarize the relatedness graphs and how to use the summarized graphs to assistant the users to retrieve target information. We also implemented our approach in an online query system and performed experiments and evaluations on it. The results show that our method produces much better result than previous work. Yukai Miao, Jianbin Qin, Wei Wang 0011 |
SIGIR | 1 |