VLDB 2026 Research / reviewers in the wild / expert
Mingyu Qiao
dblp:342/5702
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 1 heaviest of 1, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.0 | 1 | 2026 | Not All Data are What You Need: A Data-Efficient Training Method Using Heterogeneous Hardware · IEEE Trans. Knowl. Data Eng. 2026 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Not All Pretrained Representation has The Sweet Danger of Sugar: Robust and Trustworthy Representation Learning for Encrypted Malicious Traffic IdentificationabstractIdentifying encrypted malicious traffic is a key challenge in network security. Although pre-training techniques based on self-supervised learning are a trend towards reducing dependence on labeled data, existing mainstream methods that directly apply the architecture of natural language processing and masked language modeling tasks have been shown to be "high sugar" by recent research. They may rely on specific "shortcuts" rather than robust traffic behavioral representations to achieve inflated performance. How to extract more robust and generalizable features from covert and sparse encrypted malicious traffic is a major challenge. Therefore, we propose SUGARLESS, a robust representation learning for encrypted malicious traffic identification. To depart from the BERT-based paradigm, we first propose a spatial-temporal contrastive learning as the pre-training task that aligns temporal and spatial modal features without using NLP-style objectives, encouraging the model to learn cross-modal traffic correlations between modalities. We also propose a traffic-specific prompt-tuning mechanism to bridge the gap between pre-training and downstream tasks. Meanwhile, we develop a spatial-temporal feature fusion module to maintain this alignment during fine-tuning. Experiments on two public malware traffic datasets show that SUGARLESS achieves the best precision, recall, and F1-score with competitive accuracy, improves the average F1-score by 9.26% over YaTC, and is more robust under packet loss and reordering. Mingyu Qiao, Zulong Diao, Guangxing Zhang, Wanhua Li, Haiyang Jiang 0001, Zhenyu Li 0001, Gaogang Xie |
APNet | 1 |
| 2026 | Not All Flows are Worth N Packets: Robust Encrypted Traffic Classification via Dynamic Patch-level Feature LearningabstractNetwork traffic classification is significant for modern network security. The widespread use of encryption protocols, such as TLS, has resulted in fewer identifiable features of traffic, rendering encrypted traffic classification a challenging task. However, existing flow-level classification methods typically rely on the features of the first N packets. This fixed-length paradigm, truncating long flows and padding short flows with invalid data, is difficult to adapt to the dynamic distribution of traffic length, which limits robustness in interference environments such as packet loss. Considering this limitation, we propose DART, a robust encrypted traffic classification method via dynamic patch-level feature learning. We first design a robust feature representation method to improve robustness against interference, which divides a flow into multiple sub-flows named patches to extract patch-level features, replacing the interference-prone packet-by-packet sequence and significantly reducing the length of the feature sequence. We also propose a sample-adaptive dynamic inference mechanism in response to traffic heterogeneity. The input scale is adaptively selected based on traffic complexity, and simple samples are allowed to exit early in shallow layers, achieving a balance between accuracy and efficiency. The experimental results on two public malware traffic datasets demonstrate that DART exhibits excellent anti-interference robustness and improves classification accuracy by 6.59%-30.51%, while maintaining high classification efficiency compared to existing mainstream methods. Mingyu Qiao, Zulong Diao, Guangxing Zhang, Haiyang Jiang 0001, Zhenyu Li 0001, Gaogang Xie |
APNet | 1 |
| 2026 | Task-Aware Network Traffic Label Reuse With Schema-Gated EvidenceabstractNetwork traffic labels are built from heterogeneous evidence: capture metadata, port rules, service names, model predictions, and expert decisions. These sources are not interchangeable across tasks: HTTP can support an application label but cannot prove attack behavior. We propose TaskLabel, a schema-gated framework that makes traffic-label reuse traceable and reviewable. TaskLabel stores each label as a schema-bound artifact with traffic unit, target schema, confidence, provenance, evidence requirement, and reuse boundary. Its compatibility gate admits weak-source votes only when source semantics match the target dataset schema; otherwise, the cue remains provenance and is routed for review. In a retrospective pilot on 1652 held-out RT-IoT2022 flows, non-model cues reach 64.6% of flows, but direct reuse disagrees with held-out target labels on 32.5% of reached flows. Schema-agnostic weak fusion falls to 58.5% macro-F1, whereas TaskLabel blocks incompatible votes and preserves the feature-only target predictor at 93.3% macro-F1 while routing a 16.4% queue that captures an oracle-estimated 86.8% of target-label errors. Ming You, Guangxing Zhang, Mingyu Qiao, Haiyang Jiang 0001, Jinsheng Zhao |
APNet | 4 |
| 2026 | Not All Data are What You Need: A Data-Efficient Training Method Using Heterogeneous Hardware
Zulong Diao, Mingyu Qiao, Xin Wang 0001, Guangxing Zhang, Wei Liang 0005, Jianguo Chen 0001, Changhua Pei, Yanbiao Li 0001, Zhenyu Li 0001, Gaogang Xie |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | EC-GCN: A encrypted traffic classification framework based on multi-scale graph convolution networks
Zulong Diao, Gaogang Xie, Xin Wang 0001, Xuying Meng, Guangxing Zhang, Kun Xie 0001, Mingyu Qiao |
Comput. Networks | 8 |