VLDB 2026 Research / reviewers in the wild / expert
Xiaoxuan Fan
dblp:338/8591
· DBLP profile ↗
18ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Model Selection for Multivariate Time Series ForecastingabstractAccurate multivariate time series forecasting (MTSF) is critical for intelligent web services in Web of Things. When confronted with unseen multivariate time series (MTS), the industry typically invests significant time and resources in training multiple models to identify the optimal model for deployment. This paper proposes a novel, efficient, and scalable MTSF model selection method that directly selects suitable MTSF methods based on data characteristics without extensive model training. Model selection is a core component of AutoML, which has made significant progress in recent years. However, existing methods incur high operational costs and cannot be directly applied to MTSF tasks. Moreover, there is a lack of a comprehensive and cohesive public time series library for MTSF model selection. To address these challenges, we compile the first large heterogeneous labeled MTSF model selection dataset, called the ModelPile, which covers 41 mainstream datasets across 11 domains. We then propose AutoMTSF, a large model-enabled model selection method that transforms the MTSF model selection problem into a time series classification problem and utilizes the ModelPile to unlock large-scale multi-dataset training. AutoMTSF first uses the pre-trained large model to encode raw MTS. Given the coarse-grained limitations of large model encoding, Recursive Temporal Pattern Feature (RTPF) is proposed to capture both fine-grained and global temporal feature evolution, thereby effectively mapping data characteristics to the MTSF method space. Experiments comparing AutoMTSF with 2 baselines, 17 MTSF methods, and 4 large time series models show that AutoMTSF outperforms state-of-the-art methods while maintaining comparable execution time. This work represents a critical step in validating the accuracy and efficiency of large model-enabled classification for MTSF. Xiaoxuan Fan, Xianjun Deng, Qiankun Zhang 0001, Wei Xiang 0005, Shenghao Liu, Lingzhi Yi |
WWW | 1 |
| 2026 | Long-Term Traffic Forecasting via Spatial-Temporal Wavelet Attention Network for Mobile IoT-Enabled Transportation SystemsabstractAccurate long-term traffic forecasting improves traffic efficiency and safety in Intelligent Transportation Systems (ITS). However, previous methods typically overlook the complex mixed characteristics. So they struggle to capture the intricate features of newly added data effectively under distribution shifts, which exacerbates the complexity of spatial-temporal variations. Moreover, these methods fail to model global and local dynamic spatial correlations effectively, efficiently, and comprehensively. To mitigate these issues, this paper proposes an innovative Spatial-Temporal Wavelet Attention Network (STWAN). STWAN first decomposes traffic data into stable trends and fluctuating events, effectively dealing with the adverse effects of distribution shifts. Then, spatial-temporal encoder captures component-specific temporal variations and extracts dynamic global-local spatial correlations comprehensively with linear computational complexity. Additionally, transformer attention and forecasting decoder model the latent patterns, while trend-event fusion module integrates essential information for accurate forecasts. Comprehensive experiments across two real-world traffic forecasting tasks indicate that STWAN significantly outperforms state-of-the-art methods in terms of accuracy and robustness, attaining a maximum reduction of 5.54% in MAE and showcasing its robustness and broad applicability in long-term forecasting. Xianjun Deng, Shenghao Liu, Xiaoxuan Fan, Lingzhi Yi, Chenlu Zhu, Weiwei Chen 0004, Haipeng Dai 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingabstractContrastive Language-Image Pre-training (CLIP) models as a milestone of modern multimodal intelligence, its gener-alization mechanism grasped massive research interests in the community. While existing studies limited in the scope of pre-training knowledge, hardly underpinned its generalization to countless open-world concepts absent from the pre-training regime. This paper dives into such Out-of-Pre-training (OOP) generalization problem from a holistic perspective. We propose LAION-Beyond benchmark to isolate the evaluation of OOP concepts from pre-training knowledge, with regards to OpenCLIP and its reproducible variants derived from LAION datasets. Empirical analysis evidences that despite image features of OOP concepts born with significant category margins, their zero-shot transfer significantly fails due to the poor image-text alignment. To this, we elaborate the "name-tuning" methodology with its theoretical merits in terms of OOP generalization, then propose few-shot name learning (FSNL) and zero-shot name learning (ZSNL) algorithms to achieve OOP generalization in a data-efficient manner. LAION-Beyond dataset and codes: http://m-huangx.github.io/laion_beyond/. Ziliang Chen 0001, Xiaoxuan Fan, Keze Wang, Yuyu Zhou, Quanlong Guan, Liang Lin 0004 |
CVPR | 3 |
| 2025 | MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have exhibited remarkable progress. However, deficiencies remain compared to human intelligence, such as hallucination and shallow pattern matching. In this work, we aim to evaluate a fundamental yet underexplored intelligence: association, a cornerstone of human cognition for creative thinking and knowledge integration. Current benchmarks, often limited to closed-ended tasks, fail to capture the complexity of open-ended association reasoning vital for real-world applications. To address this, we present MM-OPERA, a systematic benchmark with 11,497 instances across two open-ended tasks: Remote-Item Association (RIA) and In-Context Association (ICA), aligning association intelligence evaluation with human psychometric principles. It challenges LVLMs to resemble the spirit of divergent thinking and convergent associative reasoning through free-form responses and explicit reasoning paths. We deploy tailored LLM-as-a-Judge strategies to evaluate open-ended outputs, applying process-reward-informed judgment to dissect reasoning with precision. Extensive empirical studies on state-of-the-art LVLMs, including sensitivity analysis of task instances, validity analysis of LLM-as-a-Judge strategies, and diversity analysis across abilities, domains, languages, cultures, etc., provide a comprehensive and nuanced understanding of the limitations of current LVLMs in associative reasoning, paving the way for more human-like and general-purpose AI. The dataset and code are available at https://github.com/MM-OPERA-Bench/MM-OPERA. Zimeng Huang, Jinxin Ke, Xiaoxuan Fan, Yang Liu 0084, Liu Zhonghan, Zedi Wang, Junteng Dai, Haoyi Jiang, Yuyu Zhou, Keze Wang, Ziliang Chen 0001 |
NeurIPS | 3 |
| 2025 | A collaborative adversarial framework: Distribution characteristics-guided alignment mechanism for fault diagnosis of machines considering domain shift
Xiaoxuan Fan, Lixiang Duan, Mingyu Shen |
Adv. Eng. Informatics | 1 |
| 2025 | A multi-scale graph-guided dynamic enhanced alignment network for mechanical fault diagnosis considering domain shift and data imbalance
Xiaoxuan Fan, Lixiang Duan |
Neurocomputing | 1 |
| 2025 | Area Coverage Reliability Evaluation for Collaborative Intelligence and Meta-Computing of Decentralized Industrial Internet of ThingsabstractIndustrial Internet of Things (IIoT) is evolving toward decentralization and autonomous operation. Nodes in decentralized IIoT collaboratively sense data and communicate to participate in meta-computing and provide diverse intelligent services. Area coverage reliability evaluates the information collaborative intelligence sensing and communication capabilities of nodes in decentralized IIoT for a target area. It serves as a crucial performance evaluation metric for the meta-computing and intelligent services of decentralized IIoT. Existing area coverage reliability evaluation algorithms do not consider the dynamic node states in decentralized IIoT, and overlook the capability of nodes to collaborate in meta-computing. To address these issues, this article proposes a novel confident information coverage (CIC)-based area coverage reliability metric (CACREL), which comprehensively considers node multistate, node collaboration, network coverage rate, and network connectivity. To effectively evaluate CACREL, a CIC-based coverage reliability evaluation algorithm (CACR) is proposed. Specifically, CACR transforms complex networks into grid networks and utilizes a coverage table (CT) to describe different coverage states of each grid, which reduces the complexity of the area coverage reliability evaluation. Additionally, CACR employs a reliable path algorithm to converge the CT of each grid to the sink node based on grid connectivity. Simulation results demonstrate that the proposed CACR can accurately evaluate the area coverage reliability of decentralized IIoT and enhance its service performance. Chenlu Zhu, Xiaoxuan Fan, Xianjun Deng, Shenghao Liu, Duan Yin, Hanjun Gao, Laurence T. Yang |
IEEE Internet Things J. | 2 |
| 2025 | Graph-Empowered Multidimensional Target Full-Coverage Reliability for Internet of EverythingabstractWireless sensor network plays a crucial role in sensing everything in Internet of Everything (IoE) applications. Network reliability, which measures the ability of the network to satisfy specific requirements, is one of the core factors influencing the quality of service of the network and a vital support for ensuring the normal operation of IoE applications. Existing reliability evaluation methods are mainly based on minimum cutsets or paths, which are inefficient and not suitable for large-scale networks. Furthermore, most work either focuses on coverage functionality or connectivity functionality, lacking energy awareness. To address these limitations, this article proposes a multidimensional target full-coverage reliability (TFCR). TFCR comprehensively considers various factors affecting network reliability. To evaluate TFCR, a graph-empowered confident information coverage (CIC) and signal-to-interference and noise ratio (SINR)-based energy-aware reliability algorithm (CSERA) is proposed. This algorithm evaluates network coverage based on the CIC model. Additionally, graph neural networks and the SINR-based fade tail connectivity (FTC) model are used to evaluate network connectivity functionality. CSERA balances computational accuracy and efficiency, providing reliability evaluation values within an acceptable margin of error. Extensive simulations and comparative experiments from multiple perspectives demonstrate the superiority of the proposed method CSERA over existing approaches. Chenlu Zhu, Wujie Zheng, Xiaoxuan Fan, Xianjun Deng, Shenghao Liu, Lingzhi Yi, Wei Xi 0003, Young-Sik Jeong |
IEEE Internet Things J. | 3 |
| 2025 | Co-Designed Communication and Computing for Data Reliability in Industrial Cyber-Physical Systems With Cloud-Fog AutomationabstractThe Cloud-Fog Automation is a newly proposed digital industrial automation architecture aimed at accelerating the integration and collaboration of communication, computing, and control towards next-generation cyber-physical systems (CPSs). Data reliability is one of the key considerations for achieving Cloud-Fog Automation. Sensor nodes serve as infrastructures for data collection within industrial CPSs and are essential for maintaining ultra-high data reliability. However, the underlying sensor nodes communicate frequently, are damage-prone and difficult to identify, which dramatically shortens the network lifetime and poses great challenges to data reliability. Motivated by this fact, this paper co-designs communication architecture, algorithms, and computing models for next-generation industrial CPSs with Cloud-Fog Automation to ensure data reliability and functional security. First, a four-layer energy-efficient communication architecture is proposed and a cluster head computing algorithm based on double deep Q-learning (CH-DDQ) is designed inside the architecture. Besides, a 2-stage hyBrid fault detection scheme (2-Brain) is proposed for underlying sensor nodes. 2-Brain first incorporates the Obstacle Triple Jump Protocol (OTP) and OTP packets to improve hard fault detection performance. Then, an unsupervised sensor reading soft fault detection model (SR-SFD) based on contrastive learning, momentum, and tensor is adopted to learn discriminative representations of sensor readings and identify soft faults. Simulations and a case study in the nuclear power industry manifest CH-DDQ improves the network lifetime by 5.4%~484.3% compared to three peer methods, and OTP performs better than baselines by 33.1% on average. Additionally, SR-SFD exhibits high efficiency in sensor soft fault detection and other application scenarios. Xiaoxuan Fan, Xianjun Deng, Shenghao Liu, Chenlu Zhu, Xinlei Zhou, Lingzhi Yi, Jong Hyuk Park 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | Improving Ethereum Mixing Address Linking With Tensor Computation, Neighbor Data Utilization, and Asymmetric Information ModelingabstractDue to the strong untraceability of mixing services, numerous criminals exploit these services to engage in illicit activities, posing a significant threat to the blockchain ecosystem. This paper addresses the challenge of linking transaction addresses in Tornado Cash, a popular mixing service on Ethereum. While existing state-of-the-art solutions like MixBroker attempt to address this problem, two fundamental limitations persist: insufficient utilization of neighbor information and neglect of address information asymmetry. To address these gaps, a novel framework termed “MixLinker” is proposed, which enhances neighbor information utilization and models information asymmetry. Specifically, a Normalized Adjusted Personal PageRank (NAPPR) module is designed to prioritize significant neighbor nodes while mitigating interference from super and irrelevant addresses. Additionally, tensors are employed to model transactions, capturing rich interaction features related to transaction attributes. Based on historical transaction sequences, Tensor Long Short-Term Memory (TLSTM) is used to obtain high-quality initial input features for the Graph Neural Network (GNN) module, enabling effective learning of nonlinear dynamics. To ensure symmetric output results and model asymmetric information, a temporal-aware symmetry classifier is constructed that leverages asymmetric information through permutation operations and an order-aware classifier. Extensive experiments demonstrate that MixLinker outperforms other methods, validating the effectiveness of the proposed approach and confirming the two underlying motivations. Shuilong Wang, Laurence T. Yang, Debin Liu, Ruonan Zhao, Xianjun Deng, Cannian Zou, Xiaoxuan Fan |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Contrastive Learning-Based Speech Spoofing Detection for Multimedia Security in Edge IntelligenceabstractAI-empowered edge computing has given rise to a new paradigm and effectively facilitated the promotion and development of multimedia applications. The speech assistant is one of the significant services provided by multimedia applications, which aims to offer intelligent interactive experiences between humans and machines. However, malicious attackers may exploit spoofed speeches to deceive speech assistants, posing great challenges to the security of multimedia applications. The limited resources of multimedia terminal devices hinder their ability to effectively load speech spoofing detection models. Furthermore, processing and analyzing speech in the cloud can result in poor real-time performance and potential privacy risks. Existing speech spoofing detection methods rely heavily on annotated data and exhibit poor generalization capabilities for unseen spoofed speeches. To address these challenges, this article first proposes the Coordinate Attention Network (CA2Net) that consists of coordinate attention blocks and Res2Net blocks. CA2Net can simultaneously extract temporal and spectral speech feature information and represent multi-scale speech features at a granularity level. Besides, a contrastive learning-based speech spoofing detection framework named GEMINI is proposed. GEMINI can be effectively deployed on edge nodes and autonomously learn speech features with strong generalization capabilities. GEMINI first performs data augmentation on speech signals and extracts conventional acoustic features to enhance the feature robustness. Subsequently, GEMINI utilizes the proposed CA2Net to further explore the discriminative speech features. Then, a tensor-based multi-attention comparison model is employed to maximize the consistency between speech contexts. GEMINI continuously updates CA2Net with contrastive learning, which enables CA2Net to effectively represent speech signals and accurately detect spoofed speeches. Extensive experiments on the ASVspoof2019 dataset show that GEMINI reduces the Equal Error Rate and tandem Detection Cost Function by up to 96.75% and 96.35% in the physical access scenario, and by up to 86.62% and 87.71% in the logical access scenario compared to peer methods. Xianjun Deng, Shenghao Liu, Xiaoxuan Fan, Yongling Huang, Yuanyuan He 0002, Celimuge Wu, Jong Hyuk Park 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Effective Multivariate Voice Liveness Detection System for Internet of Things SecurityabstractVoice assistants, as crucial components of the Internet of Things (IoT), are vulnerable to voice spoofing attacks and pose great threats to the security of IoT. Passive liveness detection distinguishes between genuine and spoofing voices by analyzing the collected voice, eliminating the need for deploying additional sensors. This method plays a crucial role in detecting spoofing speeches and ensuring the security of the IoT. However, current passive liveness detection methods typically require users to adopt specific gestures. Meanwhile, these methods are often designed for specific attacks and cannot accommodate multivariate attacks. To address these challenges, this paper proposes an efficient and robust liveness feature called VoiceID, which utilizes the inherent vocal cord vibrations and voiced language to authenticate the collected voice. The VoiceID is defined as the set of maximum magnitude-peak frequency bins in the magnitude spectrum of each frame for voice. VoiceID can be combined with existing acoustic features to compensate for the granularity gap in extracting fine-grained features and distinguishing between genuine and spoofing voices. Furthermore, to leverage VoiceID, this paper proposes a solid fake voice liveness detection system named SFSys and elaborates on a series of acoustic features that can work with VoiceID. Extensive experiments on authoritative ASVspoof 2019 and ASVspoof 2021 datasets reveal that VoiceID reduces the equal error rate and the minimum tandem decision cost function of the existing acoustic features by at most 6.19% and 0.2479. Moreover, SFSys outperforms existing voice liveness detection schemes and exhibits robustness in various advanced spoofing attack environments. Xiaoxuan Fan, Xianjun Deng, Shibo He, Shenghao Liu, Lingzhi Yi, Jing Wang 0036, Laurence T. Yang |
IEEE Trans. Netw. | 1 |
| 2024 | Idempotence-Constrained Representation Learning: Balancing Sensitive Information Elimination and Feature RobustnessabstractMachine learning models often inadvertently learn and perpetuate biases present in training data, particularly concerning sensitive attributes like gender and race. While existing fair representation learning approaches attempt to address this issue, they face challenges in balancing information preservation with bias elimination. This paper addresses the challenge of algorithmic bias in machine learning models by proposing a novel idempotence-constrained fair representation learning framework. We introduce a two-stage architecture that effectively eliminates both explicit and implicit bias while maintaining model performance. The first stage employs an idempotent encoder to remove explicit bias and enhance feature robustness through adaptive adversarial training, while the second stage utilizes a privacy filtering unit to eliminate implicit bias and task-irrelevant features. Our framework optimizes four key objectives: task performance, information preservation, sensitive information elimination, and representation stability. Theoretical analysis demonstrates convergence guarantees and representation stability of our approach. Extensive experiments on benchmark datasets, including UCI Adult and Heritage Health, show that our method achieves state-of-the-art performance in both fairness and accuracy metrics. Xiaoxuan Fan, Bochuan Song, Yuxue Yang |
IEEE Big Data | 1 |
| 2024 | Structured Intention Generation with Multimodal Graph Transformers: The MMIntent-LLM FrameworkabstractIn the task of answering questions related to electricity knowledge, accurately understanding user intentions is fundamental to building an effective reasoning process. Questions in this domain often cover complex issues such as troubleshooting, electricity bill consultation, and electricity safety. User queries frequently involve multiple aspects and varying levels of information. Traditional single-modal methods struggle to fully address these diverse needs. Intention understanding plays a crucial role in constructing a clear and logical reasoning chain by accurately identifying the user’s core requirements. Through in-depth analysis of user intentions, the system can better organize and reason with multimodal information such as text, voice, and images—leading to logically coherent and accurate responses. Intention understanding not only enhances the accuracy of the reasoning process but also significantly improves the response efficiency and service quality of the electricity knowledge question-and-answer system. Current approaches to intention analysis primarily rely on classification-based methods, which limit the flexibility and richness of intent understanding. To broaden the scope and depth of user intent recognition, we introduce MMIntent-LLM, a novel framework that combines T5 with a Graph Transformer to generate structured intent representations from multimodal social media content. Our approach introduces three key innovations: (1) a structured intention reasoning framework based on ATOMIC, which provides a systematic method for decomposing intent generation into interpretable components; (2) a graph-based alignment mechanism for multimodal data (text and images) that ensures semantic consistency across modalities; and (3) an adaptive fine-tuning strategy that effectively transfers knowledge from pre-trained language models to the specific intent generation task. Extensive experiments on our multimodal social intention dataset show that MMIntent-LLM achieves state-of-the-art performance, improving the average BERT score by 8.7% and human evaluation scores by 12.3% compared to baseline methods. Bochuan Song, Xiaoxuan Fan, Quanye Jia |
IEEE Big Data | 2 |
| 2024 | PT-Tuning: Bridging the Gap between Time Series Masked Reconstruction and Forecasting via Prompt Token Tuning
Jinrui Gan, Xiaoxuan Fan, Chuanxian Luo, Guangxin Jiang, Yucheng Qian, Changwei Zhao |
DASFAA (2) | 3 |
| 2024 | Tensor-Based Confident Information Coverage Reliability of Hybrid Internet of ThingsabstractThe widespread applications of the Hybrid Internet of Things (HIoT) have put forward higher requirements for network reliability. Coverage reliability is one of the important metrics of reliability, and reliable coverage ensures network data perception and transmission to improve the Quality of Service (QoS). In this article, we define Confident Information Coverage Reliability (CICR) based on the Confident Information Coverage Model (CIC), which comprehensively considers sensor multistate, sensor energy, coverage rate, and connectivity robustness to evaluate coverage reliability. Furthermore, a Tensor-based Confident Information Coverage Reliability Algorithm (T-CICR) is proposed based on tensor modeling to evaluateCICR. The algorithm uses a tensor-based Markov model to predict sensor multistate. Three tensors of coverage rate, sensor multistate, and sensor energy are constructed to provide unified representations. Simulation results show that our proposed algorithm can significantly improve coverage reliability in terms of duty cycle, coverage rate requirement, sensing range, Root Mean Square Error (RMSE) threshold, connectivity robustness requirement, and link reliability. Xiaoxuan Fan, Xianjun Deng, Yunzhi Xia, Lingzhi Yi, Laurence T. Yang, Chenlu Zhu |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Tensor Network-Based Entropy Coding For Learned Image CompressionabstractEntropy coding is fundamental for reducing the coding redundancy in image compression. However, existing entropy models for learned image compression are restricted by independent or autoregressive modeling based on presumed distribution generated from the family of Gaussian functions. In this paper, we propose a novel tensor network-based entropy model that can explicitly infer the joint distribution of image representation for learned image compression. We utilize a tree tensor network (TTN) to enable exact computation for probabilities of image representation. Specifically, we produce efficient tensor representation for entropy modeling based on bit planes slicing with gray code. Furthermore, we leverage tensor contraction to accurately calculate the partition function and jointly predict the entire bit plane for entropy coding. To our best knowledge, this paper is the first to directly infer the joint distribution of image representation in learned image compression without any presumed condition. Experimental results demonstrate that the proposed method is competitive with the state-of-the-art in both lossless and lossy image compression. Xiaoxuan Fan, Wen Fei, Wenrui Dai, Junni Zou, Hongkai Xiong |
PCS | 1 |
| 2022 | Coverage Reliability of IoT Intrusion Detection System based on Attack-Defense Game DesignabstractThe emergence of new applications of Internet of Things (IoT) makes its security and reliability become one of the most concerning issues and requires more breakthroughs. To ensure reliable operation of IoT, network reliability measures are essential for quantifying the performance of such networks. In this paper, we focus on the problem of coverage reliability of IoT intrusion detection systems based on Attack-Defense Game Design. A comprehensive coverage reliability algorithm is proposed based on Monte Carlo simulations. The algorithm employs Byzantine attack and defense ideas to determine network node attributes and uses confident information model to calculate the network coverage area. Furthermore, we propose a system reliability metric based on the analytic hierarchy process method, which takes advantage of node attributes, network coverage and connectivity. The metric is used to compare algorithms in simulated experiments, and a series of simulation comparisons illustrate the superiority and usability of the proposed approach. Xiaoxuan Fan, Yunzhi Xia, Chenlu Zhu, Shenghao Liu, Lingzhi Yi |
TrustCom | 2 |