VLDB 2026 Research / reviewers in the wild / expert
Liukun He
dblp:333/1496
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
stream mining |
1.0 | 1 | 2026 | Identifying Hierarchical Super Spreaders in a Data Stream by Hot-Separated and Mergeable Sketch · INFOCOM 2026 |
Data stream processing
super spreader identification |
1.0 | 1 | 2026 | Identifying Hierarchical Super Spreaders in a Data Stream by Hot-Separated and Mergeable Sketch · INFOCOM 2026 |
Network measurement and analytics
sketch-based measurement |
1.0 | 1 | 2026 | Identifying Hierarchical Super Spreaders in a Data Stream by Hot-Separated and Mergeable Sketch · INFOCOM 2026 |
Network measurement and analytics
traffic measurement |
1.0 | 1 | 2026 | Identifying Hierarchical Super Spreaders in a Data Stream by Hot-Separated and Mergeable Sketch · INFOCOM 2026 |
Methods — techniques the papers use, named apart from their topics
sketch data structure · 2.0hot-separated mergeable sketch · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Online Estimation of Weighted Reachability Between Vertex Pairs with Edge Similarity Decay in Temporal Interaction Graphs
Liukun He, Qingjun Xiao |
DASFAA (6) | 1 |
| 2026 | Identifying Hierarchical Super Spreaders in a Data Stream by Hot-Separated and Mergeable Sketch
Qingjun Xiao, Liukun He |
INFOCOM | 3 |
| 2026 | Identifying influential vertices with top-K largest temporal katz centralities in a streaming graph using constant memoryabstractIn graph theory, Katz centrality is a widely used measure to quantify the influence of a vertex within a network. Unlike simple degree centrality, it considers not only immediate neighbors but also multi-hop directed paths, weighting them by length and inter-hop time delays to reflect indirect influence. Traditional computation assumes a non-temporal graph, where edges lack timestamps, and computes using recursive multiplication of the graph’s adjacency matrix. However, this approach becomes impractical for graphs with billions of vertices due to high time/memory demands, and it overlooks the temporal information available on edges. By contrast, we focus on the problem of processing a streaming temporal graph , where edges are timestamped and arrive sequentially at a central analyzer. We want to online estimate the temporal Katz centrality of each vertex, and identify the top- K influential vertices with the largest Katz centralities. To solve this problem, we propose two solutions: TAS-TKC and ATAS-TKC. Both algorithms are designed to operate with a small constant-memory footprint, but ATAS-TKC builds on TAS-TKC to further improve performance. TAS-TKC avoids memory growth with the number of vertices by using a time-adaptive Count-Min sketch that allows all vertices to share memory for centrality estimation. It also maintains the top- K influential vertices using a min-heap-based tracker updated on each edge arrival. Additionally, ATAS-TKC enhances this design in two key ways: First, it achieves O ( 1 ) lookup time for the top- K vertex tracker by augmenting the min-heap with an auxiliary hash table. Second, it improves estimation accuracy by placing the top- K tracker as a prefilter before the sketch. In the prefilter, the top- K influential vertices are allocated dedicated memory and bypass the sketch entirely, avoiding estimation errors caused by memory sharing. For these two proposed solutions, we have conducted extensive evaluation based on real-world graph datasets. The results show that the average estimation error for all vertices is smaller than 4 % and the identification precision for the top- K vertices can be larger than 97 % when given only 400 KB memory to process a million-vertex graph dataset. Qingjun Xiao, Liukun He, Qifan Zhang 0006 |
Expert Syst. Appl. | 4 |
| 2026 | LogRMC: Robust Log Anomaly Detection in IoT Systems via Multifield Fusion, Contrastive Learning, and Pseudo-Label RefinementabstractIn large-scale Internet of Things (IoT) environments, the reliability and security of complex cloud-edge ecosystems critically depend on accurate anomaly detection from system logs. However, most existing methods rely primarily on log templates, often ignoring the diagnostic signals embedded in heterogeneous fields (e.g., timestamps, components, and severity levels). Furthermore, they struggle to learn discriminative representations for rare anomalies due to severe class imbalance, causing models to become biased toward normal patterns. Finally, these approaches often fail to capture the latent structure of unlabeled data, limiting their ability to generalize to unseen anomalies. These limitations render them unreliable in practice. To address the above challenges, we propose LogRMC, a semi-supervised framework for robust log anomaly detection through multi-field log feature fusion, contrastive representation learning, and pseudo-label refinement. First, LogRMC extracts features from diverse log fields and fuses them into unified contextual embeddings that encode inter-field dependencies. Next, it applies contrastive learning on normal samples to learn discriminative and stable representations, thereby improving the separation of normal and anomalous sequence embeddings. To exploit unlabeled data, LogRMC further introduces a cluster-refined pseudo-labeling strategy: it performs density-based clustering on contrastively learned embeddings and applies density ratio estimation within each cluster to generate confidence-aware pseudo-labels. Finally, an attention-based GRU network is trained jointly on labeled and pseudo-labeled log sequences. Extensive experiments on public log datasets demonstrate that LogRMC outperforms state-of-the-art unsupervised and semi-supervised methods, and achieves performance comparable to fully supervised approaches. Qingjun Xiao, Liukun He |
IEEE Internet Things J. | 4 |
| 2023 | FlowMFD: Characterisation and classification of tor traffic using MFD chromatographic features and spatial-temporal modellingabstractAbstract Tor traffic tracking is valuable for combating cybercrime as it provides insights into the traffic active on the Tor network. Tor‐based application traffic classification is one of the tracking methods, which can effectively classify Tor application services. However, it is not effective in classifying specific applications due to more complicated traffic patterns in the spatial and temporal dimensions. As a solution, the authors propose FlowMFD, a novel Tor‐based application traffic classification approach using amount‐frequency‐direction (MFD) chromatographic features and spatial‐temporal modelling. Expressly, FlowMFD mines the interaction pattern between Tor applications and servers by analysing the time series features (TSFs) of different size packets. Then MFD chromatographic features (MFDCF) are designed to represent the pattern. Those features integrate multiple low‐dimensional TSFs into a single plane and retain most pattern information. In addition, FlowMFD utilises a cascaded model with a two‐dimensional convolutional neural network (2D‐CNN) and a bidirectional gated recurrent unit to capture spatial‐temporal dependencies between MFDCF. The authors evaluate FlowMFD under the public ISCXTor2016 dataset and the self‐collected dataset, where we achieve an accuracy of 92.1% (4.2%↑) and 88.3% (4.5%↑), respectively, outperforming state‐of‐the‐art comparison methods. Liukun He, Liangmin Wang 0001, Keyang Cheng |
IET Inf. Secur. | 1 |