VLDB 2026 Research / reviewers in the wild / expert
Tianyu Cui
dblp:197/3968
· DBLP profile ↗
24ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 8 since 2021Computer networks · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM AgentsabstractXiangyu Wu, Yuwei Hu, Tianyu Cui, Yueying Tian, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Yang Yang, Jianfeng Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianyu Cui, Yueying Tian, Weihua Luo, Kaifu Zhang, Yang Yang 0074, Jianfeng Lu 0003 |
ACL (1) | 3 |
| 2026 | ATOPOS: Dynamic Path Exploration with Adaptive Probe Construction for Extensive and Efficient Network Topology Discovery
Yaochen Ren, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Tianyu Cui, Junzheng Shi |
INFOCOM | 6 |
| 2026 | Robust LLM-Based Website Fingerprinting under Dynamic Real-World ConditionsabstractWebsite Fingerprinting (WF) attacks aim to infer the websites visited by Tor users by analyzing patterns in encrypted network traffic. However, most existing WF attacks are evaluated on traffic collected in controlled environments with fixed configurations, failing to reflect the complexity and variability of real-world conditions. In practice, traffic is far more dynamic and diverse due to heterogeneous network conditions, the large number of subpages within individual websites, and continuous evolution of website content. These factors increase intra-class variability and induce temporal feature drift, which ultimately degrades the long-term effectiveness of existing attacks. In this paper, we propose TraVerse, an LLM-based representation learning framework designed to achieve robust WF attacks under real-world conditions. TraVerse applies architectural adaptation and large-scale fine-tuning on diverse unlabeled traffic to learn generalizable and resilient representations that remain effective in dynamic and evolving environments. Furthermore, TraVerse integrates a lightweight classifier atop the LLM-derived representations, enabling accurate website identification and efficient few-shot adaptation with minimal model updates. We prototype TraVerse and conduct comprehensive evaluations using real-user traffic. Experimental results show that TraVerse improves Accuracy@3 by an average of 176.3% and weighted F1 by 343.3% over state-of-the-art baselines, while maintaining strong performance throughout a three-month longitudinal evaluation. Xinhao Deng 0001, Tianyu Cui, Ke Xu 0002, Qi Li 0002 |
WWW | 3 |
| 2025 | Bridging the Gap: LLM-Powered Transfer Learning for Log Anomaly Detection in New Software SystemsabstractFor large IT companies, maintaining numerous software systems presents considerable complexity. Logs are invaluable for depicting the state of systems, making log-based anomaly detection crucial for ensuring system reliability. Existing methods require extensive log data for training, hindering their rapid deployment for new systems. Cross-system log anomaly detection methods attempt to transfer knowledge from mature systems to new ones but often struggle with syntax differences and system-specific knowledge, which hinders their effectiveness. To address these issues, this paper proposes LogSynergy, a novel transfer learning-based log anomaly detection framework. LogSynergy employs (1) LLM-based event interpretation (LEI) to standardize log syntax across different systems, and (2) system-unified feature extraction (SUFE) to disentangle system-specific features from system-unified features. These bridge the gap among different systems and enhance LogSynergy's generalizability. LogSynergy has been deployed in the production environment of a top-tier global Internet Service Provider (ISP), where it was evaluated on three real-world datasets. Additionally, we conducted evaluations on three public datasets. The results demonstrate that LogSynergy significantly outperforms existing methods. It achieves F1-scores over 89% on the real-world datasets and over 83% on the public datasets, using only 5000 labeled log sequences from the new system. These results underscore LogSynergy's effectiveness in rapidly deploying anomaly detection models for new systems. The code of LogSynergy has been open-sourced at https://github.com/DDUtian/LogSynergy Yicheng Sui, Tianyu Cui, Tong Xiao 0002, Chenghao He, Shenglin Zhang, Yongqian Sun, Dan Pei |
ICDE | 3 |
| 2025 | InfoSEM: A Deep Generative Model with Informative Priors for Gene Regulatory Network InferenceabstractInferring Gene Regulatory Networks (GRNs) from gene expression data is crucial for understanding biological processes. While supervised models are reported to achieve high performance for this task, they rely on costly ground truth (GT) labels and risk learning gene-specific biases—such as class imbalances of GT interactions—rather than true regulatory mechanisms. To address these issues, we introduce InfoSEM, an unsupervised generative model that leverages textual gene embeddings as informative priors, improving GRN inference without GT labels. InfoSEM can also integrate GT labels as an additional prior when available, avoiding biases and further enhancing performance. Additionally, we propose a biologically motivated benchmarking framework that better reflects real-world applications such as biomarker discovery and reveals learned biases of existing supervised methods. InfoSEM outperforms existing models by 38.5% across four datasets using textual embeddings prior and further boosts performance by 11.1% when integrating labeled data as priors. Tianyu Cui, Song-Jun Xu, Artem Moskalev, Shuwei Li, Tommaso Mansi, Mangal Prakash, Rui Liao |
ICML | 1 |
| 2025 | Geometric Hyena Networks for Large-scale Equivariant LearningabstractProcessing global geometric context while preserving equivariance is crucial when modeling biological, chemical, and physical systems. Yet, this is challenging due to the computational demands of equivariance and global context at scale. Standard methods such as equivariant self-attention suffer from quadratic complexity, while local methods such as distance-based message passing sacrifice global information. Inspired by the recent success of state-space and long-convolutional models, we introduce Geometric Hyena, the first equivariant long-convolutional model for geometric systems. Geometric Hyena captures global geometric context at sub-quadratic complexity while maintaining equivariance to rotations and translations. Evaluated on all-atom property prediction of large RNA molecules and full protein molecular dynamics, Geometric Hyena outperforms existing equivariant models while requiring significantly less memory and compute that equivariant self-attention. Notably, our model processes the geometric context of $30k$ tokens $20 \times$ faster than the equivariant transformer and allows $72 \times$ longer context within the same budget. Artem Moskalev, Mangal Prakash, Tianyu Cui, Rui Liao, Tommaso Mansi |
ICML | 4 |
| 2025 | IPv6 Prefix Target Generation through Pattern and Distribution Learning using Vision-Transformer and Guided-Diffusion
Yaochen Ren, Gaopeng Gou, Chengshang Hou, Tianyu Cui, Zhen Li 0011, Gang Xiong 0001, Chang Liu 0049 |
INFOCOM | 4 |
| 2025 | AetherLog: Log-based Root Cause Analysis by Integrating Large Language Models with Knowledge GraphsabstractLog-based fault root cause analysis (RCA) is paramount for ensuring the reliability of large-scale software systems. While small language model (SLM)-based methods offer efficiency and ease of deployment, their limited generalization across diverse fault scenarios often hinders their effectiveness. Conversely, large language model (LLM)-based methods demonstrate strong semantic understanding but can suffer from inaccuracies and hallucinations due to a lack of domain-specific knowledge. To overcome these limitations, we present AetherLog, a novel RCA framework synergistically integrating LLMs with knowledge graphs (KGs). In an offline phase, AetherLog employs LLMs to extract fault-relevant entities and relations, constructing a compact and semantically aligned KG through embedding-based clustering and normalization. During online analysis, the framework leverages an LLM to summarize fault logs and extract pertinent entities. Subsequently, it retrieves semantically similar entities from the KG to enrich the context and formulates context-enhanced prompts, leading to more accurate RCA. Extensive experiments conducted on two real-world datasets demonstrate that AetherLog consistently surpasses state-of-the-art baselines, achieving F1-scores of 0.93 and 0.97. These results represent significant improvements of 6% and 8% over the best existing methods, respectively, demonstrating AetherLog’s effectiveness and generalizability in log-based fault RCA. Tianyu Cui, Ruowei Fu, Changchang Liu, Yuhe Ji, Wenwei Gu, Shenglin Zhang, Yongqian Sun, Dan Pei |
ISSRE | 1 |
| 2025 | 6RIS: IPv6 Address Correlation Attacks on TLS Encrypted Traffic Using Joint Representation of Interaction and Sequential BehaviorabstractIPv6 address correlation attacks determine whether two temporary addresses belong to the same user, compromising user privacy. Particularly, existing works have shown that methods based on TLS traffic analysis can be used to perform correlation attacks. However, they suffer from inaccurate differentiation of complex user behaviors and low correlation efficiency, leading to limitations in practical applications. In this paper, we propose a 6RIS model to improve IPv6 address correlation attacks on TLS-encrypted traffic. 6RIS learns the joint representation of interaction and sequential behavior from traffic, which is used to construct a KD-Tree for efficient correlation. Statistical aggregation and semantic preference modules are designed to extract generalized features from complex interaction behavior. To model sequential behavior, we utilize a sequence learning module to capture service dependencies, enhancing behavior representation. Experiments on a real-world IPv6 dataset show that 6RIS ($\mathbf{9 1. 8 6 \%}$TPR,$\mathbf{0. 8 3 \%}$FPR) outperforms state-of-theart methods. The correlation efficiency of 6RIS improves by at least 57 % compared to existing methods. Additionally, we further confirm through 6RIS that persistent session IDs in TLS session resumption can directly expose IPv6 temporary addresses to correlation attacks. Yang Li 0002, Chang Liu 0049, Gaopeng Gou, Tianyu Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
IWQoS | 4 |
| 2025 | LogEval: A comprehensive benchmark suite for LLMs in log analysis
Tianyu Cui, Shiyu Ma, Tong Xiao 0002, Shimin Tao, Yilun Liu 0001, Shenglin Zhang, Duoming Lin, Changchang Liu, Yuzhe Cai, Weibin Meng, Yongqian Sun, Dan Pei |
Empir. Softw. Eng. | 1 |
| 2024 | Harmonizing Generalization and Personalization in Federated Prompt LearningabstractFederated Prompt Learning (FPL) incorporates large pre-trained Vision-Language models (VLM) into federated learning through prompt tuning. The transferable representations and remarkable generalization capacity of VLM make them highly compatible with the integration of federated learning. Addressing data heterogeneity in federated learning requires personalization, but excessive focus on it across clients could compromise the model’s ability to generalize effectively. To preserve the impressive generalization capability of VLM, it is crucial to strike a balance between personalization and generalization in FPL. To tackle this challenge, we proposed Federated Prompt Learning with CLIP Generalization and low-rank Personalization (FedPGP), which employs pre-trained CLIP to provide knowledge-guidance on the global prompt for improved generalization and incorporates a low-rank adaptation term to personalize the global prompt. Further, FedPGP integrates a prompt-wise contrastive loss to achieve knowledge guidance and personalized adaptation simultaneously, enabling a harmonious balance between personalization and generalization in FPL. We conduct extensive experiments on various datasets to explore base-to-novel generalization in both category-level and domain-level scenarios with heterogeneous data, showing the superiority of FedPGP in balancing generalization and personalization. Tianyu Cui, Jingya Wang 0001, Ye Shi 0001 |
ICML | 1 |
| 2024 | Measuring Encrypted DNS Service with TLS1.3 Support over IPv6abstractThe Encrypted Domain Name System (DNS) and Encrypted Server Name Indication (ESNI) are recently proposed to enhance network security and privacy protection; we refer to these schemes collectively as domain name encryption technologies. Previous research has shown that the destination IP address accessed by the user cannot be associated with common web services such as websites because a large number of websites are hosted through cloud or CDN over IPv6. However, encrypted DNS, as an internet infrastructure service, is typically deployed independently by the service provider rather than hosted through cloud or CDN. In this paper, we propose a method to discover the unique service provider of encrypted DNS resolvers on a large-scale encrypted traffic with TLS1.3 support over IPv6. The model utilizes a Siamese Network to determine whether two IPv6 resolver addresses belong to the same service provider of encrypted DNS, even if the DNS query is protected by ESNI. Through a comprehensive analysis of two real-world datasets, which include encrypted DNS data and common web data, we find that the implementation of TLS1.3, especially ESNI, does not impact the association of encrypted DNS server addresses. Our model achieves an accuracy rate of 95.29%. Liang Jiao, Wen-Xiu Zhang, Tianyu Cui, Yujia Zhu, Qingyun Liu 0001 |
ISCC | 4 |
| 2023 | Incorporating functional summary information in Bayesian neural networks using a Dirichlet process likelihood approachabstractBayesian neural networks (BNNs) can account for both aleatoric and epistemic uncertainty. However, in BNNs the priors are often specified over the weights which rarely reflects true prior knowledge in large and complex neural network architectures. We present a simple approach to incorporate prior knowledge in BNNs based on external summary information about the predicted classification probabilities for a given dataset. The available summary information is incorporated as augmented data and modeled with a Dirichlet process, and we derive the corresponding Summary Evidence Lower BOund. The approach is founded on Bayesian principles, and all hyperparameters have a proper probabilistic interpretation. We show how the method can inform the model about task difficulty and class imbalance. Extensive experiments show that, with negligible computational overhead, our method parallels and in many cases outperforms popular alternatives in accuracy, uncertainty calibration, and robustness against corruptions with both balanced and imbalanced data. Vishnu Raj, Tianyu Cui, Markus Heinonen, Pekka Marttinen |
AISTATS | 2 |
| 2023 | Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy ContinuationabstractDeep Equilibrium Models (DEQs) and Neural Ordinary Differential Equations (Neural ODEs) are two branches of implicit models that have achieved remarkable success owing to their superior performance and low memory consumption. While both are implicit models, DEQs and Neural ODEs are derived from different mathematical formulations. Inspired by homotopy continuation, we establish a connection between these two models and illustrate that they are actually two sides of the same coin. Homotopy continuation is a classical method of solving nonlinear equations based on a corresponding ODE. Given this connection, we proposed a new implicit model called HomoODE that inherits the property of high accuracy from DEQs and the property of stability from Neural ODEs. Unlike DEQs, which explicitly solve an equilibrium-point-finding problem via Newton's methods in the forward pass, HomoODE solves the equilibrium-point-finding problem implicitly using a modified Neural ODE via homotopy continuation. Further, we developed an acceleration method for HomoODE with a shared learnable initial point. It is worth noting that our model also provides a better understanding of why Augmented Neural ODEs work as long as the augmented part is regarded as the equilibrium point to find. Comprehensive experiments with several image classification tasks demonstrate that HomoODE surpasses existing implicit models in terms of both accuracy and memory consumption. Shutong Ding, Tianyu Cui, Jingya Wang 0001, Ye Shi 0001 |
NeurIPS | 2 |
| 2023 | LncReader: identification of dual functional long noncoding RNAs using a multi-head self-attention mechanismabstractLong noncoding ribonucleic acids (RNAs; LncRNAs) endowed with both protein-coding and noncoding functions are referred to as 'dual functional lncRNAs'. Recently, dual functional lncRNAs have been intensively studied and identified as involved in various fundamental cellular processes. However, apart from time-consuming and cell-type-specific experiments, there is virtually no in silico method for predicting the identity of dual functional lncRNAs. Here, we developed a deep-learning model with a multi-head self-attention mechanism, LncReader, to identify dual functional lncRNAs. Our data demonstrated that LncReader showed multiple advantages compared to various classical machine learning methods using benchmark datasets from our previously reported cncRNAdb project. Moreover, to obtain independent in-house datasets for robust testing, mass spectrometry proteomics combined with RNA-seq and Ribo-seq were applied in four leukaemia cell lines, which further confirmed that LncReader achieved the best performance compared to other tools. Therefore, LncReader provides an accurate and practical tool that enables fast dual functional lncRNA identification. Bohao Zou, Manman He, Yongfei Hu, Yiying Dou, Tianyu Cui, Puwen Tan, Shaobin Li, Shuan Rao, Sixi Liu, Kaican Cai, Dong Wang 0011 |
Briefings Bioinform. | 6 |
| 2023 | DELFMUT: duplex sequencing-oriented depth estimation model for stable detection of low-frequency mutationsabstractDuplex sequencing technology has been widely used in the detection of low-frequency mutations in circulating tumor deoxyribonucleic acid (DNA), but how to determine the sequencing depth and other experimental parameters to ensure the stable detection of low-frequency mutations is still an urgent problem to be solved. The mutation detection rules of duplex sequencing constrain not only the number of mutated templates but also the number of mutation-supportive reads corresponding to each forward and reverse strand of the mutated templates. To tackle this problem, we proposed a Depth Estimation model for stable detection of Low-Frequency MUTations in duplex sequencing (DELFMUT), which models the identity correspondence and quantitative relationships between templates and reads using the zero-truncated negative binomial distribution without considering the sequences composed of bases. The results of DELFMUT were verified by real duplex sequencing data. In the case of known mutation frequency and mutation detection rule, DELFMUT can recommend the combinations of DNA input and sequencing depth to guarantee the stable detection of mutations, and it has a great application value in guiding the experimental parameter setting of duplex sequencing technology. Guiying Wu, Tianyu Cui, Zicong Jiao, Liyan Ji, Jiayin Wang 0002, Xuefeng Xia, Huan Fang 0005, Yanfang Guan |
Briefings Bioinform. | 4 |
| 2022 | Deconfounded Representation Similarity for Comparison of Neural NetworksabstractSimilarity metrics such as representational similarity analysis (RSA) and centered kernel alignment (CKA) have been used to understand neural networks by comparing their layer-wise representations. However, these metrics are confounded by the population structure of data items in the input space, leading to inconsistent conclusions about the \emph{functional} similarity between neural networks, such as spuriously high similarity of completely random neural networks and inconsistent domain relations in transfer learning. We introduce a simple and generally applicable fix to adjust for the confounder with covariate adjustment regression, which improves the ability of CKA and RSA to reveal functional similarity and also retains the intuitive invariance properties of the original similarity measures. We show that deconfounding the similarity metrics increases the resolution of detecting functionally similar neural networks across domains. Moreover, in real-world applications, deconfounding improves the consistency between CKA and domain similarity in transfer learning, and increases the correlation between CKA and model out-of-distribution accuracy similarity. Tianyu Cui, Pekka Marttinen, Samuel Kaski |
NeurIPS | 1 |
| 2022 | GALG: Linking Addresses in Tracking Ecosystem Using Graph Autoencoder with Link Generation
Tianyu Cui, Gang Xiong 0001, Chang Liu 0049, Junzheng Shi, Peipei Fu, Gaopeng Gou |
ECML/PKDD (6) | 1 |
| 2021 | 6GAN: IPv6 Multi-Pattern Target Generation via Generative Adversarial Nets with Reinforcement LearningabstractGlobal IPv6 scanning has always been a challenge for researchers because of the limited network speed and computational power. Target generation algorithms are recently proposed to overcome the problem for Internet assessments by predicting a candidate set to scan. However, IPv6 custom address configuration emerges diverse addressing patterns discouraging algorithmic inference. Widespread IPv6 alias could also mislead the algorithm to discover aliased regions rather than valid host targets. In this paper, we introduce 6GAN, a novel architecture built with Generative Adversarial Net (GAN) and reinforcement learning for multi-pattern target generation. 6GAN forces multiple generators to train with a multi-class discriminator and an alias detector to generate non-aliased active targets with different addressing pattern types. The rewards from the discriminator and the alias detector help supervise the address sequence decision-making process. After adversarial training, 6GAN's generators could keep a strong imitating ability for each pattern and 6GAN's discriminator obtains outstanding pattern discrimination ability with a 0.966 accuracy. Experiments indicate that our work outperformed the state-of-the-art target generation algorithms by reaching a higher-quality candidate set. Tianyu Cui, Gaopeng Gou, Gang Xiong 0001, Chang Liu 0049, Peipei Fu, Zhen Li 0011 |
INFOCOM | 1 |
| 2021 | SiamHAN: IPv6 Address Correlation Attacks on TLS Encrypted Traffic via Siamese Heterogeneous Graph Attention Network
Tianyu Cui, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui, Chang Liu 0049 |
USENIX Security Symposium | 1 |
| 2020 | Learning Global Pairwise Interactions with Bayesian Neural NetworksabstractEstimating global pairwise interaction effects, i.e., the difference between the joint effect and the sum of marginal effects of two input features, with uncertainty properly quantified, is centrally important in science applications. We propose a non-parametric probabilistic method for detecting interaction effects of unknown form. First, the relationship between the features and the output is modelled using a Bayesian neural network, capable of representing complex interactions and principled uncertainty. Second, interaction effects and their uncertainty are estimated from the trained model. For the second step, we propose an intuitive global interaction measure: Bayesian Group Expected Hessian (GEH), which aggregates information of local interactions as captured by the Hessian. GEH provides a natural trade-off between type I and type II error and, moreover, comes with theoretical guarantees ensuring that the estimated interaction effects and their uncertainty can be improved by training a more accurate BNN. The method empirically outperforms available non-probabilistic alternatives on simulated and real-world data. Finally, we demonstrate its ability to detect interpretable interactions between higher-level features (at deeper layers of the neural network). Tianyu Cui, Pekka Marttinen, Samuel Kaski |
ECAI | 1 |
| 2020 | 6GCVAE: Gated Convolutional Variational Autoencoder for IPv6 Target Generation
Tianyu Cui, Gaopeng Gou, Gang Xiong 0001 |
PAKDD (1) | 1 |
| 2020 | 6VecLM: Language Modeling in Vector Space for IPv6 Target Generation
Tianyu Cui, Gang Xiong 0001, Gaopeng Gou, Junzheng Shi |
ECML/PKDD (4) | 1 |
| 2019 | A Comprehensive Study of Accelerating IPv6 DeploymentabstractSince the lack of IPv6 network development, China is currently accelerating IPv6 deployment. In this scenario, traffic and network structure show a huge shift. However, due to the long-term prosperity, we are ignorant of the problems behind such outbreak of traffic and performance improvement events in accelerating deployment. IPv6 development in some regions will still face similar challenges in the future. To contribute to solving this problem, in this paper, we produce a new measurement framework and implement a 5-month passive measurement on the IPv6 network during the accelerating deployment in China. We combine 6 global-scale datasets to form the normal status of IPv6 network, which is against to the accelerating status formed by the passive traffic. Moreover, we compare with the traffic during World IPv6 Day 2011 and Launch 2012 to discuss the common nature of accelerating deployment. Finally, the results indicate that the IPv6 accelerating deployment is often accompanied by an unbalanced network status. It exposes unresolved security issues including the challenge of user privacy and inappropriate access methods. According to the investigation, we point the future IPv6 development after accelerating deployment. Tianyu Cui, Chang Liu 0049, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001 |
IPCCC | 1 |