Xintong Song

dblp:193/0670 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0005-8086-2237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BufferNAS: Buffer pool sampling in neural architecture search
Hongzhi Wang 0001, Chunnan Wang, Xintong Song, Fei Geng
Inf. Sci.3
2026 PiKGL: Leveraging Pruned Knowledge Graphs for Explainable Stance Detection
abstract
Abstract Stance detection on social media plays a vital role in understanding public opinion on contentious topics. While prior work leverages external knowledge sources like Wikipedia to enrich limited target information, it primarily introduces conceptual content, neglecting the interpretability potential of knowledge and often leading to the incorporation of irrelevant or redundant information that hinders stance prediction performance. To address this, we introduce PiKGL, a Pruned interpretable Knowledge Graph Learning framework for explainable stance detection. Specifically, we first extract event triplets and topics to obtain real-world knowledge, which is then used to construct an interpretable knowledge graph. To ensure precision and minimize noise, we introduce a retrieval-guided pruning strategy that incorporates commonsense knowledge, filtering redundant information of the interpretable knowledge graph. Finally, the pruned knowledge graph is injected into a large language model to jointly model textual, target, and commonsense for improved stance comprehension. Experimental results conducted on three public datasets demonstrate our PiKGL achieves state-of-the-art performance on stance detection.
Jingjie Lin, Zhixin Bai, Xintong Song, Qianlong Wang 0001, Min Yang 0007, Jing Li 0049, Ruifeng Xu 0001
Trans. Assoc. Comput. Linguistics4
2025 Meta-Learning Based CTR Algorithm Selection and Hyperparameter Optimization
abstract
The existing Click-Through Rate (CTR) algorithms have their own advantages and are sensitive to hyperparameters. Quickly obtaining a high-performance CTR model for the new task can bring good application effects. However, ordinary users fail to do so due to the lack of domain knowledge. In this paper, we remedy this deficiency by proposing AutoCTR, an efficient meta-learning based Combined Algorithm Selection and Hyperparameter Optimization (CASH) algorithm, to help non-expert users quickly find the best CTR model. In AutoCTR, we introduce the meta-learning technique to make full use of the meta-information w.r.t. CTR to guide for the new CTR task. Specifically, we utilize the meta-information to learn characteristics and representations of CTR algorithms with different settings. We use these meta experiences combined with few evaluation information on the target CTR dataset to efficiently exploring the huge CTR CASH search space for the new task. The CTR model representation method has significant influence on the quality of the learned meta experiences. To further enhance the experiences quality, we also design a Graph Neural Network (GNN) based embedding learning method. This method can link different CTR models through their components, and thus quickly learning higher-quality model representations. Extensive experimental results show that AutoCTR can quickly select suitable CTR models for different CTR tasks. Compared with the existing CASH algorithms, which ignore meta-information or rely on a huge amount of meta-information, AutoCTR is more reasonable and efficient.
Chunnan Wang, Xiang Chen 0019, Xintong Song, Tianyu Mu, Hongzhi Wang 0001
ICDE4
2025 Bridging Time Gaps: Temporal Logic Relations for Enhancing Temporal Reasoning in Large Language Models
abstract
The understanding and cognition of time are the basis for large language models to understand the world. Although large language models (LLMs) have demonstrated strong capabilities in multiple reasoning tasks, they still have significant deficiencies in temporal reasoning, mainly due to the diversity of temporal expressions and the lack of temporal logic reasoning capabilities. In this study, we propose a novel Temporal Chain of Thought framework(TempCoT) to improve the performance of LLM in temporal reasoning tasks through a three-stage reasoning strategy. First, TempCoT explicitly extracts time constraints to ensure the accuracy of time references during reasoning. Second, a semantic retrieval mechanism is introduced to dynamically obtain key temporal facts to enhance the integrity and reliability of information. Finally, an explicit temporal logic reasoning module is constructed based on point algebra to improve the consistency and interpretability of reasoning. Experimental results show that TempCoT significantly improves the temporal reasoning performance of five different LLMs and shows stronger robustness on complex temporal tasks.
Xintong Song, Bin Liang 0004, Chenhua Zhang, Ruifeng Xu 0001
SIGIR1
2025 EO-Shield: A Shield-Based Protection Scheme Against Both Invasive and Non-Invasive Attacks
abstract
Smart devices, especially Internet-connected devices, typically incorporate security protocols and cryptographic algorithms to ensure the control flow integrity and information security. However, various types of attacks try to tamper with these devices, including invasive and non-invasive. Chip-level shields have been proven effective against invasive attacks, but the potential of shields as a protection mechanism against side-channel analysis (SCA) attacks remains under-explored. To bridge this gap, we propose a shield-based multi-functional protection scheme, named EO-Shield, capable of simultaneously thwarting invasive and non-invasive attacks. EO-Shield is implemented using the chip’s top metal layer and includes an Information Leakage Obfuscation Module (ILOM) underneath. This module generates its protection patterns based on the operating conditions of the circuit that need to be protected, thus reducing the correlation between electromagnetic (EM) emanations and cryptographic data. Additionally, we introduce a simulation technique to test the protection efficacy of EO-Shield at the layout level, utilizing commercial Electronic Design Automation (EDA) tools and the EMSim/EMSim+ tool. Simulation experiments demonstrate that the ILOM decreases the signal-to-noise (SNR) ratio to below 0.6 and improves the difficulty of SCA attacks by more than 100 times. Compared to existing single-function protection methods against physical attacks, EO-Shield leverages the EM protection potential of shields to offer multi-functional protection.
Ya Gao 0007, Qizhi Zhang 0001, Xintong Song, Haocheng Ma, Jiaji He 0001, Yiqiang Zhao
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 Research on APT group classification method based on graph attention networks
Yazhou Du, Weiwu Ren, Xintong Song
J. Supercomput.3
2024 EMSim+: Accelerating Electromagnetic Security Evaluation With Generative Adversarial Network and Transfer Learning
abstract
Electromagnetic side-channel analysis (EM SCA) attack poses a serious threat to integrated circuits (ICs), necessitating timely vulnerability detection before deployment to enhance EM side-channel security. Various EM simulation methods have emerged for analyzing EM side-channel leakage, providing sufficiently accurate results. However, these simulator-based methods still face two principal challenges in the design process of high security chips. Firstly, the large volume of measurement data required for a single security evaluation results in substantial time overhead. Secondly, design iterations lead to repetitive security evaluations, thus increasing the evaluation cost. In this paper, we propose EMSim+ which includes two efficient and accurate layout-level EM side-channel leakage evaluation frameworks named EMSim+GAN and EMSim+GAN+TL to mitigate the above challenges, respectively. EMSim+GAN integrates a Generative Adversarial Network (GAN) model that utilizes the chip’s cell current and power grid information to predict EM emanations quickly. EMSim+GAN+TL further incorporates transfer learning (TL) within the framework, leveraging the experience of existing designs to reduce the training datasets for new designs and achieve the target accuracy. We compare the simulation results of EMSim+ with the state-of-the-art EM simulation tool, EMSim as well as silicon measurements. Experimental results not only prove the high efficiency and high simulation accuracy of EMSim+, but also verify its generalization ability across different designs and technology nodes.
Ya Gao 0007, Haocheng Ma, Qizhi Zhang 0001, Xintong Song, Yier Jin, Jiaji He 0001, Yiqiang Zhao
IEEE Trans. Inf. Forensics Secur.4
2023 TransFusion Model Fusion Mechanism Based on Transformer for Traffic Flow Prediction
abstract
In recent years, the problem of traffic congestion has become a hot topic. Accurate traffic flow prediction methods have received extensive attention from many researchers all over the world. Although many methods proposed at present have achieved good results in the field of traffic flow prediction, most of them only consider the static characteristic of traffic data, but do not consider the dynamic characteristic of traffic data. The factors that affect traffic flow prediction are changeable, and they will change over time. In response to this dynamic characteristic, the authors propose a model fusion mechanism based on transformer (TransFusion). The authors adopt two basic forecasting models (TCN and LSTM) as the underlying architectures. In view of the performance of different models on the traffic data at different times, the authors design a model fusion mechanism to assign dynamic weights to basic models at different times. Experiments on three datasets have proved that TransFusion has a significant improvement compared with basic models.
Xintong Song, Donghua Yang, Hongzhi Wang 0001, Bo Zheng 0012
J. Database Manag.1
2023 ADOps: An Anomaly Detection Pipeline in Structured Logs
abstract
Anomaly detection has been extensively implemented in industry. The reality is that an application may have numerous scenarios where anomalies need to be monitored. However, the complete process of anomaly detection will take much time, including data acquisition, data processing, model training, and model deployment. In particular, some simple scenarios do not require building complex anomaly detection models. This results in a waste of resources. To solve these problems, we build an anomaly detection pipeline(ADOps) to modularize each step. For simple anomaly detection scenarios, no programming is required and new anomaly detection tasks can be created by simply modifying the configuration file. In addition, it can also improve the development efficiency of complex anomaly detection models. We show how users create anomaly detection tasks on the anomaly detection pipeline and how engineers use it to develop anomaly detection models.
Xintong Song, Yusen Zhu, Jianfei Wu, Bai Liu 0002, Hongkang Wei
Proc. VLDB Endow.1
2023 Side Channel Security Oriented Evaluation and Protection on Hardware Implementations of Kyber
abstract
The emergence of quantum computing and its impact on current cryptographic algorithms has triggered the migration to post-quantum cryptography (PQC). Among the PQC candidates, CRYSTALS-Kyber is a key encapsulation mechanism (KEM) that stands out from the National Institute of Standards and Technology (NIST) standardization project. While software implementations of Kyber have been developed and evaluated recently, Kyber’s hardware implementations especially those designed with parallel architecture, are rarely discussed. To help better understand Kyber hardware designs and their security against side-channel analysis (SCA) attacks, in this paper, we first adapt the two most recent Kyber hardware designs for FPGA implementations. We then perform SCA attacks against these hardware designs with different architectures, i.e., parallelization and pipelining. Our experimental results show that Kyber designs on FPGA boards are vulnerable to SCA attacks including electromagnetic (EM) and power side channels. An attacker only needs$27 \sim 1,600$power traces or$60 \sim 2,680$EM traces to recover the decryption key successfully. Furthermore, we propose two first-order IND-CPA Kyber decapsulation masking protected designs, and then we evaluate their securities and overheads. The experimental results demonstrate that the side channel security of masked Kyber designs has increased by more than 10x.
Yiqiang Zhao, Shijian Pan, Haocheng Ma, Ya Gao 0007, Xintong Song, Jiaji He 0001, Yier Jin
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 CO-AutoML: An Optimizable Automated Machine Learning System
Chunnan Wang, Hongzhi Wang 0001, Xintong Song, Yuhao Bao, Bo Zheng 0012
DASFAA (3)4
2020 Fair and Efficient Caching Algorithms and Strategies for Peer Data Sharing in Pervasive Edge Computing Environments
abstract
Edge devices with sensing, storage, and communication resources (e.g., smartphones, tablets, connected vehicles, and IoT nodes) are increasingly penetrating our daily lives. Many novel applications can be created through sharing data among nearby peer edge devices. In such applications, caching data at some edge devices can greatly improve data availability, retrieval robustness, and delivery latency. In this paper, we study the unique problem of caching fairness in edge computing environments. Due to the heterogeneity of peer edge devices, load balance is a critical issue that affects the fairness in caching. We propose fairness metrics to characterize this issue and formulate the caching fairness problem as an integer linear programming problem, which is shown as the summation of multiple Connected Facility Location (ConFL) problems. We provide an approximation algorithm by leveraging an existing ConFL approximation algorithm, and prove that it preserves a 6.55 approximation ratio. We further develop a distributed algorithm where devices exchange data reachability information and identify popular candidates as caching nodes. Finally, we update the fairness metric and apply it to algorithms for making continuous caching decisions overtime. Our extensive evaluation results show that compared with existing caching algorithms for wireless networks, our proposed algorithms significantly improve the data caching fairness while keeping the contention induced latency comparable to the best existing algorithms.
Yaodong Huang, Xintong Song, Fan Ye 0003, Yuanyuan Yang 0001, Xiaoming Li 0001
IEEE Trans. Mob. Comput.2
2017 Fair Caching Algorithms for Peer Data Sharing in Pervasive Edge Computing Environments
abstract
Edge devices (e.g., smartphones, tablets, connected vehicles, IoT nodes) with sensing, storage and communication resources are increasingly penetrating our environments. Many novel applications can be created when nearby peer edge devices share data. Caching can greatly improve the data availability, retrieval robustness and latency. In this paper, we study the unique issue of caching fairness in edge environment. Due to distinct ownership of peer devices, caching load balance is critical. We consider fairness metrics and formulate an integer linear programming problem, which is shown as summation of multiple Connected Facility Location (ConFL) problems. We propose an approximation algorithm leveraging an existing ConFL approximation algorithm, and prove that it preserves a 6.55 approximation ratio. We further develop a distributed algorithm where devices exchange data reachability and identify popular candidates as caching nodes. Extensive evaluation shows that compared with existing wireless network caching algorithms, our algorithms significantly improve data caching fairness, while keeping the contention induced latency similar to the best existing algorithms.
Yaodong Huang, Xintong Song, Fan Ye 0003, Yuanyuan Yang 0001, Xiaoming Li 0001
ICDCS2
2017 Content Centric Peer Data Sharing in Pervasive Edge Computing Environments
abstract
The proliferation and daily congregation of modern mobile devices have created abundant opportunities for peer edge devices to share valuable data with each other. The short contact durations, relatively small sharing sizes, and uncertain data availability, demand agile, light weight peer based data sharing. In this paper, we propose Peer Data Sharing (PDS) that enables edge devices to discover which data exist in nearby peers, and retrieve interested data robustly and efficiently. PDS uses novel lingering queries, mixedcast and en-route message rewriting techniques to minimize redundant transmissions and maximize opportunistic overhearing thus caching in data discovery and retrieval. Extensive evaluations based on an Android prototype show that PDS discovers and retrieves almost 100% data in tens of seconds, and remains robust despite wireless contention, simultaneous consumer requests and user mobility.
Xintong Song, Yaodong Huang, Qian Zhou 0008, Fan Ye 0003, Yuanyuan Yang 0001, Xiaoming Li 0001
ICDCS1
2017 Understanding Smartphone Sensor and App Data for Enhancing the Security of Secret Questions
abstract
Many web applications provide secondary authentication methods, i.e., secret questions (or password recovery questions), to reset the account password when a user's login fails. However, the answers to many such secret questions can be easily guessed by an acquaintance or exposed to a stranger that has access to public online tools (e.g., online social networks); moreover, a user may forget her/his answers long after creating the secret questions. Today's prevalence of smartphones has granted us new opportunities to observe and understand how the personal data collected by smartphone sensors and apps can help create personalized secret questions without violating the users’ privacy concerns. In this paper, we present aSecret-Question based Authenticationsystem, called “Secret-QA”, that creates a set of secret questions on basic of people's smartphone usage. We develop a prototype on Android smartphones, and evaluate the security of the secret questions by asking the acquaintance/stranger who participates in our user study to guess the answers with and without the help of online tools; meanwhile, we observe the questions’ reliability by asking participants to answer their own questions. Our experimental results reveal that the secret questions related to motion sensors, calendar, app installment, and part of legacy app usage history (e.g., phone calls) have the best memorability for users as well as the highest robustness to attacks.
Kaigui Bian, Tong Zhao 0001, Xintong Song, Jung-Min Park 0001, Xiaoming Li 0001, Fan Ye 0003, Wei Yan 0007
IEEE Trans. Mob. Comput.4
2016 Holistic Reality Examination on Practical Challenges in a Mobile CrowdSensing Application
abstract
Despite significant research efforts and great advances on Mobile CrowdSensing (MCS), building MCS applications remains difficult. In this paper, we develop and run Dining Halls on Live (DHOL), a campus dining population density monitoring system over several months. We make a holistic reality examination, discover key technical and practical difficulties, develop effective solutions and share our experiences and insights. We find two main obstacles on data fusion and incentive design: insufficient data quantity/quality and ``irrational'' user behavior. We develop effective methods by combining historical and real time data, and allocating a given budget among users to address them. We also conduct a detailed user survey to identify reasons behind interesting discoveries, important practical difficulties in acquiring sufficient users and location data, and share our experiences dealing with them. Our main insight is that insufficient data quantity/quality and ``irrational'' user behavior demand practical yet effective data fusion and incentive mechanisms, and one must provide values to users to acquire and retain a large user base.
Xintong Song, Fan Ye 0003, Xiaoming Li 0001, Yuanyuan Yang 0001
GLOBECOM1