Bin Shi 0003

dblp:63/4724-3 · DBLP profile ↗
← Back
49ranked-venue papers
6as first author
39since 2021 · last 2026
0000-0001-8272-9361ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 1 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Databases, data management, data science and information retrieval · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Security and privacy · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 since 2021Computer networks · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 Scope Delineation Before Localization: A Two-Stage Framework for Enhancing Failure Attribution in Multi-Agent Systems
abstract
Large language models (LLMs) are seeing growing adoption in multi-agent systems. In these systems, efficient failure attribution is critical for ensuring robustness and interpretability. Current LLM-based attribution methods often face challenges with lengthy logs and lacking expert knowledge. Drawing inspiration from human debugging strategies, we propose an automated failure attribution framework, Scope Delineation Before Localization, which operates in two key stages: (1) identifying the failure scope and (2) pinpointing the failure step. By decoupling failure attribution into the two stages, our approach alleviates the reasoning workload of LLMs, enabling more precise failure attribution. To support scope delineation, we further introduce two strategies: Stepwise Scope Delineation and Expertise-Assisted Scope Delineation. Experiments on the Who&When dataset validate the efficacy of our two-stage framework, demonstrating substantial improvements over prior methods (up to 24.27% on step-level accuracy).
Bo Dong 0001, Bin Shi 0003
AAAI6
2026 Enhancing Pre-training Data Detection in LLMs Through Discriminative and Symmetric Prefix Selection
abstract
The rapid development of large language models (LLMs) has relied on access to high-quality, large-scale datasets, yet growing concerns around data privacy and security have spurred substantial research into pre-training data detection. While state-of-the-art (SOTA) methods such as RECALL and CON-RECALL leverage auxiliary prefixes to enhance detection performance, their dependence on individual prefixes introduces notable instability across varying prefix conditions. To address this, we first conduct a theoretical analysis to assess the impact of prefixes on existing prefix-based methods. Building on the analysis, we propose a novel prefix selection method to identify optimal prefixes. Specifically, our method derives two key criteria Discriminability and Symmetry. These criteria serve to quantify the effectiveness of prefixes in detecting pre-training data, enabling precise selection of high-performing candidate prefixes. Experiments on the WikiMIA dataset demonstrate that our method consistently improves the performance of RECALL and CON-RECALL, achieving gains of up to 21.1% in AUC scores while significantly enhancing robustness.
Bo Dong 0001, Bin Shi 0003
AAAI5
2026 TransTS: an adaptive post-hoc method for probability calibration under label noise
Yuefei Wu, Bin Shi 0003, Bo Dong 0001
Sci. China Inf. Sci.2
2026 A guard against ambiguous sentiment for multimodal aspect-level sentiment classification
Yanjing Wang 0008, Bin Shi 0003, Kaihao Zhang, Bo Dong 0001
Inf. Process. Manag.3
2026 RADAR: Relation-assisted dual-graph aligning recognition for grounded multimodal named entity recognition
Zai Zhang 0002, Bin Shi 0003, Bo Dong 0001
Inf. Process. Manag.2
2026 Revisiting weakly supervised tabular anomaly detection from a cell-level perspective
Zhen Peng 0005, Xujing Jia, Qika Lin, Bin Shi 0003
Neural Networks6
2025 Revisiting Graph Contrastive Learning on Anomaly Detection: A Structural Imbalance Perspective
abstract
The superiority of graph contrastive learning (GCL) has prompted its application to anomaly detection tasks for more powerful risk warning systems. Unfortunately, existing GCL-based models tend to excessively prioritize overall detection performance while neglecting robustness to structural imbalance, which can be problematic for many real-world networks following power-law degree distributions. Particularly, GCL-based methods may fail to capture tail anomalies (abnormal nodes with low degrees). This raises concerns about the security and robustness of current anomaly detection algorithms and therefore hinders their applicability in a variety of realistic high-risk scenarios. To the best of our knowledge, research on the robustness of graph anomaly detection to structural imbalance has received little scrutiny. To address the above issues, this paper presents a novel GCL-based framework named AD-GCL. It devises the neighbor pruning strategy to filter noisy edges for head nodes and facilitate the detection of genuine tail nodes by aligning from head nodes to forged tail nodes. Moreover, AD-GCL actively explores potential neighbors to enlarge the receptive field of tail nodes through anomaly-guided neighbor completion. We further introduce intra- and inter-view consistency loss of the original and augmentation graph for enhanced representation. The performance evaluation of the whole, head, and tail nodes on multiple datasets validates the comprehensive superiority of the proposed AD-GCL in detecting both head anomalies and tail anomalies.
Yiming Xu 0001, Zhen Peng 0005, Bin Shi 0003, Xu Hua, Bo Dong 0001, Song Wang 0013, Chen Chen 0022
AAAI3
2025 VERO: Verification and Zero-Shot Feedback Acquisition for Few-Shot Multimodal Aspect-Level Sentiment Classification
abstract
Deep learning approaches for multimodal aspect-level sentiment classification (MALSC) often require extensive data, which is costly and time-consuming to obtain. To mitigate this, current methods typically fine-tune small-scale pretrained models like BERT and BART with few-shot examples. While these models have shown success, Large Vision-Language Models (LVLMs) offer significant advantages due to their greater capacity and ability to understand nuanced language in both zero-shot and few-shot settings. However, there is limited work on fine-tuning LVLMs for MALSC. A major challenge lies in selecting few-shot examples that effectively capture the underlying patterns in data for these LVLMs. To bridge this research gap, we propose an acquisition function designed to select challenging samples for the few-shot learning of LVLMs for MALSC. We compare our approach, Verification and ZERO-shot feedback acquisition (VERO), with diverse acquisition functions for few-shot learning in MALSC. Our experiments show that VERO outperforms prior methods, achieving an F1 score improvement of up to 6.07% on MALSC benchmark datasets.
Bin Shi 0003, Samuel Mensah, Bo Dong 0001
AAAI3
2025 Out-of-Distribution Generalization on Graphs via Progressive Inference
abstract
The development and evaluation of graph neural networks (GNNs) generally follow the independent and identically distributed (i.i.d.) assumption. Yet this assumption is often untenable in practice due to the uncontrollable data generation mechanism. In particular, when the data distribution shows a significant shift, most GNNs would fail to produce reliable predictions and may even make decisions randomly. One of the most promising solutions to improve the model generalization is to pick out causal invariant parts in the input graph. Nonetheless, we observe a significant distribution gap between the causal parts learned by existing methods and the ground-truth, leading to undesirable performance. In response to the above issues, this paper presents GPro, a model that learns graph causal invariance with progressive inference. Specifically, the complicated graph causal invariant learning is decomposed into multiple intermediate inference steps from easy to hard, and the perception of GPro is continuously strengthened through a progressive inference process to extract causal features that are stable to distribution shifts. We also enlarge the training distribution by creating counterfactual samples to enhance the capability of the GPro in capturing the causal invariant parts. Extensive experiments demonstrate that our proposed GPro outperforms the state-of-the-art methods by 4.91% on average. For datasets with more severe distribution shifts, the performance improvement can be up to 6.86%.
Yiming Xu 0001, Bin Shi 0003, Zhen Peng 0005, Huixiang Liu, Bo Dong 0001, Chen Chen 0022
AAAI2
2025 Generation, Validation, and Selection: A Proof-by-Contradiction Reasoning Chain Data Synthesis Method for Enhancing Indirect Reasoning in Large Language Models
Bin Shi 0003, Bo Dong 0001
IEEE Big Data3
2025 Text-Attributed Graph Anomaly Detection via Multi-Scale Cross- and Uni-Modal Contrastive Learning
abstract
The widespread application of graph data in various high-risk scenarios has increased attention to graph anomaly detection (GAD). Faced with real-world graphs that often carry node descriptions in the form of raw text sequences, termed text-attributed graphs (TAGs), existing graph anomaly detection pipelines typically involve shallow embedding techniques to encode such textual information into features, and then rely on complex self-supervised tasks within the graph domain to detect anomalies. However, this text encoding process is separated from the anomaly detection training objective in the graph domain, making it difficult to ensure that the extracted textual features focus on GAD-relevant information, seriously constraining the detection capability. How to seamlessly integrate raw text and graph topology to unleash the vast potential of cross-modal data in TAGs for anomaly detection poses a challenging issue. This paper presents a novel end-to-end paradigm for text-attributed graph anomaly detection, named CMUCL. We simultaneously model data from both text and graph structures, and jointly train text and graph encoders by leveraging cross-modal and uni-modal multi-scale consistency to uncover potential anomaly-related information. Accordingly, we design an anomaly score estimator based on inconsistency mining to derive node-specific anomaly scores. Considering the lack of benchmark datasets tailored for anomaly detection on TAGs, we release 8 datasets to facilitate future research. Extensive evaluations show that CMUCL significantly advances in text-attributed graph anomaly detection, delivering an 11.13% increase in average accuracy (AP) over the suboptimal.
Yiming Xu 0001, Xu Hua, Zhen Peng 0005, Bin Shi 0003, Jiarun Chen, Xingbo Fu, Song Wang 0013, Bo Dong 0001
ECAI4
2025 Court of LLMs: Evidence-Augmented Generation via Multi-LLM Collaboration for Text-Attributed Graph Anomaly Detection
abstract
The natural combination of intricate topological structures and rich textual information in text-attributed graphs (TAGs) opens up a novel perspective for graph anomaly detection (GAD). However, existing GAD methods primarily focus on designing complex optimization objectives within the graph domain, overlooking the complementary value of the textual modality, whose features are often encoded by shallow embedding techniques, such as bag-of-words or skip-gram, so that semantic context related to anomalies may be missed. To unleash the enormous potential of textual modality, large language models (LLMs) have emerged as promising alternatives due to their strong semantic understanding and reasoning capabilities. Nevertheless, their application to TAG anomaly detection remains nascent, and they struggle to encode high-order structural information inherent in graphs due to input length constraints. For high-quality anomaly detection in TAGs, we propose CoLL, a novel framework that combines LLMs and graph neural networks (GNNs) to leverage their complementary strengths. CoLL employs multi-LLM collaboration for evidence-augmented generation to capture anomaly-relevant contexts while delivering human-readable rationales for detected anomalies. Moreover, CoLL integrates a GNN equipped with a gating mechanism to adaptively fuse textual features with evidence while preserving high-order topological information. Extensive experiments demonstrate the superiority of CoLL, achieving an average improvement of 13.37% in AP. This study opens a new avenue for incorporating LLMs in advancing GAD.
Yiming Xu 0001, Jiarun Chen, Zhen Peng 0005, Zihan Chen 0002, Qika Lin, Bin Shi 0003, Bo Dong 0001
ACM Multimedia7
2025 NI-GDBA: Non-Intrusive Distributed Backdoor Attack Based on Adaptive Perturbation on Federated Graph Learning
abstract
Federated Graph Learning (FedGL) is an emerging Federated Learning (FL) framework that learns the graph data from various clients to train better Graph Neural Networks(GNNs) model. Owing to concerns regarding the security of such framework, numerous studies have attempted to execute backdoor attacks on FedGL, with a particular focus on distributed backdoor attacks. However, all existing methods posting distributed backdoor attack on FedGL only focus on injecting distributed backdoor triggers into the training data of each malicious client, which will cause model performance degradation on original task and is not always effective when confronted with robust federated learning defense algorithms, leading to low success rate of attack. What's more, the backdoor signals introduced by the malicious clients may be smoothed out by other clean signals from the honest clients, which potentially undermining the performance of the attack.
Ken Li, Bin Shi 0003, Jiazhe Wei, Bo Dong 0001
WWW2
2025 Uncertainty-aware large language model response length perception
Bin Shi 0003, Bo Dong 0001
Sci. China Inf. Sci.1
2025 TED: related party transaction guided tax evasion detection on heterogeneous graph
Yiming Xu 0001, Bin Shi 0003, Bo Dong 0001, Hua Wei 0001
Data Min. Knowl. Discov.2
2025 ALM-PU: positive and unlabeled learning with constrained optimization
Jiazhe Wei, Yuefei Wu, Bin Shi 0003, Ken Li, Bo Dong 0001
Mach. Learn.3
2025 Management Actions Enhanced Traffic Prediction Under Sparse Data
abstract
Accurate traffic prediction is crucial for real-time traffic management and a variety of subsequent applications. In the domain of traffic prediction, while the latest data-driven studies have achieved satisfactory results, they often overlook the significant problem of data sparsity (i.e., only part of the traffic system is observed) in real-world scenarios which can lead to insufficient data and erroneous predictions. Conventional strategies to counteract sparse data in traffic prediction problems typically involve a two-step process: initially imputing missing data, then predicting traffic state. A critical flaw in these existing methods is their assumption that traffic on adjacent roads equally influences unobserved roads, disregarding the impact of traffic management actions. In response to this challenge, we proposeTMA-GNN, an approach that integrates traffic management actions into traditional data-driven models, treating traffic signal control as governing rules. This approach constrains the attention mechanism between roads, enabling real-time updates of road connectivity and effectively capturing diverse vehicle behavior patterns such as emerging from a residential area or entering a parking lot. These enhancements not only address the limitations of ignoring traffic strategies but also challenge the conventional assumption that vehicles always travel strictly downstream from upstream segments, thereby improving traffic state estimations on both observed and unobserved roads. Nonetheless, in certain extreme scenarios (e.g., when a major residential area is connected by a single road), the unobserved road may be unpredictable and lead to inaccurate imputations. To address this,TMA-GNNintroduces a complementary strategy named EERA for evaluating the robustness of imputed values. This method effectively identifies roads with the highest uncertainty and prioritizes them for sensor installation, bridging data gaps on such unobserved roads. We conducted experiments on both synthetic and real-world datasets to validate the efficacy of theTMA-GNNframework and its components. Additionally, we carried out a series of case studies demonstrating the practical utility of EERA in real-world scenarios.
Zhiming Liang, Bin Shi 0003, Bo Dong 0001, Hua Wei 0001
IEEE Trans. Intell. Transp. Syst.2
2025 End-to-End Abnormal Subgraph Detection via Subgraph-Level Contrastive Learning
abstract
Abnormal subgraph (AS) detection plays a significant role in ensuring the security of many high-impact domains. Unlike node anomaly detection, identifying subgraph anomalies is extremely challenging due to the exponentially large subgraph space caused by various combinations of nodes and edges. Moreover, in the absence of supervisory signals, how to quantify the abnormality of subgraphs poses another pressing challenge. Traditional methods typically rely on handcrafted subgraph anomaly measures, making it hard to handle potential unknown anomalies with limited prior knowledge. Recent deep learning-based techniques are predominantly designed to discover individual node anomalies, which could be suboptimal for AS detection due to the inconsideration of collaborative behaviors between nodes in the subgraph. In fact, existing studies have put very little effort into this task, and even dedicated performance evaluation metrics are not yet available. To address the above challenges and promote related research, in this article, we propose a end-to-end unsupervised subgraph anomaly detection framework (EndSubG), which jointly models subgraph partition and AS detection as a whole instead of treating them as two separate stages. Specifically, EndSubG uncovers potential AS boundaries that violate the Homophily assumption by modeling the edge existence probability, then achieves anomaly-aware graph embedding and subgraph partition based on the refined topology. By forming a coarsened subgraph network, EndSubG picks out subgraph anomalies by learning the "subgraph-vicinity" matching patterns. Additionally, we design an evaluation metric weighted normalized mutual information centered on AS (AS-WNMI) specifically for subgraph anomaly detection, which is a variant of vanilla NMI and quantifies detection performance from both subgraph partition and anomaly recognition. The experimental results on synthetic and real-world datasets corroborate the superiority of end-to-end unsupervised subgraph anomaly detection framework (EndSubG) in terms of area under the curve (AUC), average precision (AP), and AS-WNMI. We also provide an intuitive analysis of the detected subgraphs through visualization for better understanding.
Zhen Peng 0005, Qika Lin, Bin Shi 0003, Chen Chen 0022, Bo Dong 0001, Chao Shen 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 RR-PU: A Synergistic Two-Stage Positive and Unlabeled Learning Framework for Robust Tax Evasion Detection
abstract
Tax evasion, an unlawful practice in which taxpayers deliberately conceal information to avoid paying tax liabilities, poses significant challenges for tax authorities. Effective tax evasion detection is critical for assisting tax authorities in mitigating tax revenue loss. Recently, machine-learning-based methods, particularly those employing positive and unlabeled (PU) learning, have been adopted for tax evasion detection, achieving notable success. However, these methods exhibit two major practical limitations. First, their success heavily relies on the strong assumption that the label frequency (the fraction of identified taxpayers among tax evaders) is known in advance. Second, although some methods attempt to estimate label frequency using approaches like Mixture Proportion Estimation (MPE) without making any assumptions, they subsequently construct a classifier based on the error-prone label frequency obtained from the previous estimation. This two-stage approach may not be optimal, as it neglects error accumulation in classifier training resulting from the estimation bias in the first stage. To address these limitations, we propose a novel PU learning-based tax evasion detection framework called RR-PU, which can revise the bias in a two-stage synergistic manner. Specifically, RR-PU refines the label frequency initialization by leveraging a regrouping technique to fortify the MPE perspective. Subsequently, we integrate a trainable slack variable to fine-tune the initial label frequency, concurrently optimizing this variable and the classifier to eliminate latent bias in the initial stage. Experimental results on three real-world tax datasets demonstrate that RR-PU outperforms state-of-the-art methods in tax evasion detection tasks.
Shuzhi Cao, Jianfei Ruan, Bo Dong 0001, Bin Shi 0003
AAAI4
2024 The Evidence Contraction Issue in Deep Evidential Regression: Discussion and Solution
abstract
Deep Evidential Regression (DER) places a prior on the original Gaussian likelihood and treats learning as an evidence acquisition process to quantify uncertainty. For the validity of the evidence theory, DER requires specialized activation functions to ensure that the prior parameters remain non-negative. However, such constraints will trigger evidence contraction, causing sub-optimal performance. In this paper, we analyse DER theoretically, revealing the intrinsic limitations for sub-optimal performance: the non-negativity constraints on the Normal Inverse-Gamma (NIG) prior parameter trigger the evidence contraction under the specialized activation function, which hinders the optimization of DER performance. On this basis, we design a Non-saturating Uncertainty Regularization term, which effectively ensures that the performance is further optimized in the right direction. Experiments on real-world datasets show that our proposed approach improves the performance of DER while maintaining the ability to quantify uncertainty.
Yuefei Wu, Bin Shi 0003, Bo Dong 0001, Hua Wei 0001
AAAI2
2024 Estimating Noisy Class Posterior with Part-level Labels for Noisy Label Learning
abstract
In noisy label learning, estimating noisy class posteriors plays a fundamental role for developing consistent classifiers, as it forms the basis for estimating clean class posteriors and the transition matrix. Existing methods typically learn noisy class posteriors by training a classification model with noisy labels. However, when labels are incorrect, these models may be misled to overemphasize the feature parts that do not reflect the instance characteristics, resulting in significant errors in estimating noisy class posteriors. To address this issue, this paper proposes to augment the supervised information with part-level labels, encouraging the model to focus on and integrate richer information from various parts. Specifically, our method first partitions features into distinct parts by cropping instances, yielding part-level labels associated with these various parts. Subsequently, we introduce a novel single-to-multiple transition matrix to model the relationship between the noisy and part-level labels, which incorporates part-level labels into a classifier-consistent framework. Utilizing this framework with part-level labels, we can learn the noisy class posteriors more precisely by guiding the model to integrate information from various parts, ultimately improving the classification performance. Our method is theoretically sound, while experiments show that it is empirically effective in synthetic and real-world noisy benchmarks.
Rui Zhao 0028, Bin Shi 0003, Jianfei Ruan, Tianze Pan, Bo Dong 0001
CVPR2
2024 Tackling Instance-Dependent Label Noise with Class Rebalance and Geometric Regularization
abstract
In label-noise learning, accurately identifying the transition matrix is crucial for developing statistically consistent classifiers. This task is complicated by instance-dependent noise, which introduces identifiability challenges in the absence of stringent assumptions. Existing methods use neural networks to estimate the transition matrix by initially extracting confident clean instances. However, this extraction process is hindered by severe inter-class imbalance and a bias toward selecting unambiguous intra-class instances, leading to a distorted understanding of noise patterns. To tackle these challenges, our paper introduces a Class Rebalance and Geometric Regularization-based Framework (CRGR). CRGR employs a smoothed, noise-tolerant reweighting mechanism to equilibrate inter-class representation, thereby mitigating the risk of model overfitting to dominant classes. Additionally, recognizing that instances with similar characteristics often exhibit parallel noise patterns, we propose that the transition matrix should mirror the similarity of the feature space. This insight promotes the inclusion of ambiguous instances in training, serving as a form of geometric regularization. Such a strategy enhances the model's ability to navigate diverse noise patterns and strengthens its generalization capabilities. By addressing both inter-class and intra-class biases, CRGR offers a more balanced and robust classification model. Extensive experiments on both synthetic and real-world datasets demonstrate CRGR's superiority over existing state-of-the-art methods, significantly boosting classification accuracy and showcasing its effectiveness in handling instance-dependent noise.
Shuzhi Cao, Jianfei Ruan, Bo Dong 0001, Bin Shi 0003
KDD4
2024 SimMix: Local similarity-aware data augmentation for time series
Pin Liu, Rui Wang 0118, Bin Shi 0003
Expert Syst. Appl.7
2024 Libsignal: an open library for traffic signal control
abstract
This paper introduces a library for cross-simulator comparison of reinforcement learning models in traffic signal control tasks. This library is developed to implement recent state-of-the-art reinforcement learning models with extensible interfaces and unified cross-simulator evaluation metrics. It supports commonly-used simulators in traffic signal control tasks, including Simulation of Urban MObility(SUMO) and CityFlow, and multiple benchmark datasets for fair comparisons. We conducted experiments to validate our implementation of the models and to calibrate the simulators so that the experiments from one simulator could be referential to the other. Based on the validated models and calibrated environments, this paper compares and reports the performance of current state-of-the-art RL algorithms across different datasets and simulators. This is the first time that these methods have been compared fairly under the same datasets with different simulators.
Xiaoliang Lei, Longchao Da, Bin Shi 0003, Hua Wei 0001
Mach. Learn.4
2024 Learning dynamic graph representations through timespan view contrasts
Yiming Xu 0001, Zhen Peng 0005, Bin Shi 0003, Xu Hua, Bo Dong 0001
Neural Networks3
2023 Rethinking Sentiment Analysis under Uncertainty
abstract
Sentiment Analysis (SA) is a fundamental task in natural language processing, which is widely used in public decision-making. Recently, deep learning have demonstrated great potential to deal with this task. However, prior works have mostly treated SA as a deterministic classification problem, and meanwhile, without quantifying the predictive uncertainty. This presents a serious problem in the SA, different annotator, due to the differences in beliefs, values, and experiences, may have different perspectives on how to label the text sentiment. Such situation will lead to inevitable data uncertainty and make the deterministic classification models feel puzzle to make decision. To address this issue, we propose a new SA paradigm with the consideration of uncertainty and conduct an expensive empirical study. Specifically, we treat SA as the regression task and introduce uncertainty quantification to obtain confidence intervals for predictions, which enables the risk assessment ability of the model and can improve the credibility of SA-aids decision-making. Experiments on five datasets show that our proposed new paradigm effectively quantifies uncertainty in SA while remaining competitive performance to point estimation, in addition to being capable of Out-Of-Distribution~(OOD) detection.
Yuefei Wu, Bin Shi 0003, Jiarun Chen, Bo Dong 0001, Hua Wei 0001
CIKM2
2023 CLDG: Contrastive Learning on Dynamic Graphs
abstract
The graph with complex annotations is the most potent data type, whose constantly evolving motivates further exploration of the unsupervised dynamic graph representation. One of the representative paradigms is graph contrastive learning. It constructs self-supervised signals by maximizing the mutual information between the statistic graph’s augmentation views. However, the semantics and labels may change within the augmentation process, causing a significant performance drop in downstream tasks. This drawback becomes greatly magnified on dynamic graphs. To address this problem, we designed a simple yet effective framework named CLDG. Firstly, we elaborate that dynamic graphs have temporal translation invariance at different levels. Then, we proposed a sampling layer to extract the temporally-persistent signals. It will encourage the node to maintain consistent local and global representations, i.e., temporal translation invariance under the timespan views. The extensive experiments demonstrate the effectiveness and efficiency of the method on seven datasets by outperforming eight unsupervised state-of-the-art baselines and showing competitiveness against four semi-supervised methods. Compared with the existing dynamic graph method, the number of model parameters and training time is reduced by an average of 2,001.86 times and 130.31 times on seven datasets, respectively. The code and data are available at: https://github.com/yimingxu24/CLDG.
Yiming Xu 0001, Bin Shi 0003, Bo Dong 0001, Haoyi Zhou
ICDE2
2023 Uncertainty-aware Traffic Prediction under Missing Data
abstract
Traffic prediction is a crucial topic because of its broad scope of applications in the transportation domain. Though recent studies have achieved promising results, most of them cannot adequately deal with positions with no historical data, which is common due to limited resources in real life. Apart from this, the lack of uncertainty measurements also makes current models unable to manage risks, especially for the downstream tasks involving decision-making. Inspired by the previous inductive graph neural network, we proposed an uncertainty-aware framework to 1) extend prediction to locations with no historical records and significantly extend spatial coverage of prediction while reducing sensor deployment and 2) generate probabilistic prediction with uncertainty quantification to help the risk management. The experiment results show that our method achieved promising results on prediction tasks, and the uncertainty quantification gives consistent results that highly correlate with the locations with and without historical data. We also show that our model could help support sensor deployment tasks in the transportation field to achieve higher accuracy with a limited sensor deployment budget.
Junxian Li 0001, Zhiming Liang, Guanjie Zheng, Bin Shi 0003, Hua Wei 0001
ICDM5
2023 Reinforcement Learning Approaches for Traffic Signal Control under Missing Data
abstract
The emergence of reinforcement learning (RL) methods in traffic signal control (TSC) tasks has achieved promising results. Most RL approaches require the observation of the environment for the agent to decide which action is optimal for a long-term reward. However, in real-world urban scenarios, missing observation of traffic states may frequently occur due to the lack of sensors, which makes existing RL methods inapplicable on road networks with missing observation. In this work, we aim to control the traffic signals in a real-world setting, where some of the intersections in the road network are not installed with sensors and thus with no direct observations around them. To the best of our knowledge, we are the first to use RL methods to tackle the TSC problem in this real-world setting. Specifically, we propose two solutions: 1) imputes the traffic states to enable adaptive control. 2) imputes both states and rewards to enable adaptive control and the training of RL agents. Through extensive experiments on both synthetic and real-world road network traffic, we reveal that our method outperforms conventional approaches and performs consistently with different missing rates. We also investigate how missing data influences the performance of our model.
Junxian Li 0001, Bin Shi 0003, Hua Wei 0001
IJCAI3
2023 NerCo: A Contrastive Learning Based Two-Stage Chinese NER Method
abstract
Sequence labeling serves as the most commonly used scheme for Chinese named entity recognition(NER). However, traditional sequence labeling methods classify tokens within an entity into different classes according to their positions. As a result, different tokens in the same entity may be learned with representations that are isolated and unrelated in target representation space, which could finally negatively affect the subsequent performance of token classification. In this paper, we point out and define this problem as Entity Representation Segmentation in Label-semantics. And then we present NerCo: Named entity recognition with Contrastive learning, a novel NER framework which can better exploit labeled data and avoid the above problem. Following the pretrain-finetune paradigm, NerCo firstly guides the encoder to learn powerful label-semantics based representations by gathering the encoded token representations of the same Semantic Class while pushing apart that of different. Subsequently, NerCo finetunes the learned encoder for final entity prediction. Extensive experiments on several datasets demonstrate that our framework can consistently improve the baseline and achieve state-of-the-art performance.
Zai Zhang 0002, Bin Shi 0003, Haokun Zhang, Huang Xu 0002, Yuefei Wu, Bo Dong 0001
IJCAI2
2023 MDP: Privacy-Preserving GNN Based on Matrix Decomposition and Differential Privacy
abstract
In recent years, graph neural networks (GNN) have developed rapidly in various fields, but the high computational consumption of its model training often discourages some graph owners who want to train GNN models but lack computing power. Therefore, these data owners often cooperate with external calculators during the model training process, which will raise critical severe privacy concerns. Protecting private information in graph, however, is difficult due to the complex graph structure consisting of node features and edges. To solve this problem, we propose a new privacy-preserving GNN named MDP based on matrix decomposition and differential privacy (DP), which allows external calculators train GNN models without knowing the original data. Specifically, we first introduce the concept of topological secret sharing (TSS), and design a novel matrix decomposition method named eigenvalue selection (ES) according to TSS, which can preserve the message passing ability of adjacency matrix while hiding edge information. We evaluate the feasibility and performance of our model through extensive experiments, which demonstrates that MDP model achieves accuracy comparable to the original model, with practically affordable overhead.
Wanghan Xu, Bin Shi 0003, Jiqiang Zhang, Zhiyuan Feng, Tianze Pan, Bo Dong 0001
JCC2
2023 An edge feature aware heterogeneous graph neural network model to support tax evasion detection
Bin Shi 0003, Bo Dong 0001, Yiming Xu 0001
Expert Syst. Appl.1
2023 Exploring Interactive and Contrastive Relations for Nested Named Entity Recognition
abstract
Nested named entities (nested NEs) refer to the situation where one named entity is included or nested within another named entity, which cannot be recognized by the traditional sequence labeling methods. Recently, span-based methods have become the mainstream methods for nested Named Entity Recognition (nested NER). The fundamental concept behind this method is to enumerate nearly all potential spans as entity mentions and subsequently classify them. However, span-based methods independently classify spans without considering the semantic relations among them, which negatively impacts the span representation. To address the issue, we propose a novel deep learning architecture for nested NER that explores interactive and contrastive relations among spans. Specifically, we design a scale transformation mechanism that embeds geometric information into span representations, which enhances the model's ability to encode interactive relations between spans. Additionally, we introduce a supervised contrastive learning loss that pulls apart highly overlapping spans in the embedding space to encode the contrastive relations. Experiments show that our method achieves state-of-the-art or competitive performance on three publicly nested NER datasets, thus validating its effectiveness.
Yuefei Wu, Guangtao Wang, Yanping Chen 0010, Wei Wu 0069, Zai Zhang 0002, Bin Shi 0003, Bo Dong 0001
IEEE ACM Trans. Audio Speech Lang. Process.7
2023 Tax Evasion Detection With FBNE-PU Algorithm Based on PnCGCN and PU Learning
abstract
Tax evasion is an illegal activity in which individuals or entities avoid paying their true tax liability. It has always been a crucial issue for both governments and academic researchers to efficiently detect tax evasion. Recent research has proposed the use of machine learning technology to detect tax evasion and has shown good results in some specific areas. Regrettably, there are still two major obstacles to detect tax evasion. First, it is hard to extract powerful features because of the complexity of tax data. Second, due to the complicated process of tax auditing, labeled data are limited. Such obstacles motivate the contributions of this work. In this paper, we propose a novel tax evasion detection framework named FBNE-PU, a multi-stage method to detect tax evasion in real-life scenarios. In this paper, we perform an in-depth analysis of the characteristics of the transaction network and propose a novel network embedding algorithm, the PnCGCN. It significantly improves detection performance by extracting powerful features from basic features and the tax-related transaction network. Moreover, we utilize nnPU to assign pseudo labels for unlabeled data. Finally, a MLP is trained as the decision function. Experiments on three real-world datasets demonstrate that our method significantly outperforms the comparison methods in the tax evasion detection task.
Yuda Gao, Bin Shi 0003, Bo Dong 0001, Lingyun Mi
IEEE Trans. Knowl. Data Eng.2
2022 Adaptive Shapelets Preservation for Time Series Augmentation
abstract
Time series augmentation is an essential technique in training deep learning models for time series, especially achieving remarkable results in tackling the overfitting problems. However, existing methods fail to specifically protect discriminative features that contribute significantly to the classification results, so these features may be destroyed during the augmentation process. This leads to lower fidelity of augmented time series, which ultimately interferes with classification decisions during inference. To address this issue, we propose an adaptive shapelets preservation approach, named ASP. First, we exploit the saliency map to detect shapelets on the original time series that contain discriminative features. Second, we preserve them during augmentation and assign a proprietary label to each time series. It improves the fidelity of augmented time series and the confidence of their labels, thereby avoiding the risk of interfering with classification decisions. Experimental results on 128 datasets of the UCR2018 archive show that our method ASP outperforms that without augmentation on 98 datasets, and helps the classifier achieve the average accuracy improvement from 71.26% to 75.46%, which is far better than the state-of-the-art approaches.
Pin Liu, Xiaohui Guo, Bin Shi 0003, Tianyu Wo, Xudong Liu 0001
IJCNN4
2022 Modeling Network-level Traffic Flow Transitions on Sparse Data
abstract
Modeling how network-level traffic flow changes in the urban environment is useful for decision-making in transportation, public safety and urban planning. The traffic flow system can be viewed as a dynamic process that transits between states (e.g., traffic volumes on each road segment) over time. In the real-world traffic system with traffic operation actions like traffic signal control or reversible lane changing, the system's state is influenced by both the historical states and the actions of traffic operations. In this paper, we consider the problem of modeling network-level traffic flow under a real-world setting, where the available data is sparse (i.e., only part of the traffic system is observed). We present DTIGNN, an approach that can predict network-level traffic flows from sparse data. DTIGNN models the traffic system as a dynamic graph influenced by traffic signals, learns the transition models grounded by fundamental transition equations from transportation, and predicts future traffic states with imputation in the process. Through comprehensive experiments, we demonstrate that our method outperforms state-of-the-art methods and can better support decision-making in transportation.
Xiaoliang Lei, Bin Shi 0003, Hua Wei 0001
KDD3
2022 Non-salient region erasure for time series augmentation
Pin Liu, Xiaohui Guo, Bin Shi 0003, Rui Wang 0118, Tianyu Wo, Xudong Liu 0001
Frontiers Comput. Sci.3
2022 Be United in Actions: Taking Live Snapshots of Heterogeneous Edge-Cloud Collaborative Cluster With Low Overhead
abstract
Failure recovery is one of the most essential problems in the Internet of Things (IoT) systems, especially in crucial scenarios, such as traffic control and healthcare. Meanwhile, with the ever-increasing demand of IoT applications and for latency and security considerations, more and more IoT applications are migrated to large clusters that consist of both cloud and edge servers. However, with the scale of edge–cloud collaborative clusters continue to expand, the risk of system errors and failures is also increasing. The conventional snapshot/rollback method is a powerful way for solving this problem and it is widely used in cloud computing scenarios. But when transplanting to edge–cloud collaborative clusters with the nature of distribution and heterogeneity, it will introduce serious network interruption and guest performance impact. Therefore, in this article, to address the above problems, we propose a duration-aware cluster snapshot system, named Phalanx, which can take live snapshots of edge–cloud collaborative clusters with low performance overhead. In Phalanx, we use the low-overhead precopy model and first propose a virtual machine (VM) snapshot duration prediction method that can accurately predict the snapshot duration of each single VM. Then, based on the prediction results, we coordinate the snapshot process to ensure the whole cluster has a consistency-friendly schedule, thereby solving the network interruption problems and finally, minimizing the adverse performance impact to the guest IoT applications. We implement the prototype of Phalanx on QEMU/KVM platform and conduct several experiments. The experimental results show that Phalanx offers negligible network interruption while incurring 10.68%–20.9% less performance impact over existing solutions.
Bin Shi 0003, Bo Dong 0001
IEEE Internet Things J.1
2022 Memory/Disk Operation Aware Lightweight VM Live Migration
abstract
Live virtual machine migration technique allows migrating an entire OS with running applications from one physical host to another, while keeping all services available without interruption. It provides a flexible and powerful way to balance system load, save power, and tolerate faults in data centers. Meanwhile, with the stringent requirements of latency, scalability, and availability, an increasing number of applications are deployed across distributed data-centers. However, existing live migration approaches still suffer from long downtime and serious performance degradation in cross data-center scenes due to the mass of dirty retransmission, which limits the ability of cross data-center scheduling. In this paper, we propose a system named Memory/disk operation aware Lightweight VM Live Migration across data-centers with low performance impact (MLLM). It significantly improves the cross data-center migration performance by reducing the amount of dirty data in the migration process. In MLLM, we predict disk read workingset (i.e., more frequently read contents) and memory write workingset (i.e., more frequently write contents) based on the access sequence traces. And then we adjust the migration models and data transfer sequence by the workingset information. We further proposed an improved algorithm for workingset estimation. Moreover, we discussed the potential use of machine learning (ML) to enhance the performance of the VM migration and also propose a two-level hierarchical network model to make the ML-based prediction more efficient. We implement MLLM and its improved versions on the QEMU/KVM platform and conduct several experiments. The experimental results show that 1) MLLM averagely reduces 62.9% of total migration time and 36.0% service downtime over existing methods; 2) The improved workingset estimation algorithm reduces 9.32% memory pre-copy time on average over the original algorithm.
Bin Shi 0003, Haiying Shen, Bo Dong 0001
IEEE/ACM Trans. Netw.1
2020 RVAE-ABFA : Robust Anomaly Detection for HighDimensional Data Using Variational Autoencoder
abstract
The curse of dimensionality is a fundamental difficulty in anomaly detection for high dimensional data. To deal with this problem, the autoencoder based approach is an elegant solution. However, existing works require a clean training dataset that is not always guaranteed in real scenarios. In this paper, we propose a novel anomaly detection method named RVAE-ABFA (robust variational autoencoder with attention based feature adaptation for high dimensional data anomaly detection), which significantly improves the anomaly detection performance when training data is contaminated. Rather than only utilize reconstruction error, we take the learned low dimensional embeddings generated by variational autoencoder into consideration. In RVAE-ABFA, the learned low dimensional embeddings are helpful to detect anomalies in contaminated data because of the ability of variational inference. We also propose an ABFA (attention based feature adaptation) mechanism to adjust the weights of low dimensional embeddings and reconstruction error. Furthermore, we adopt the adversarial training criterion to perform variational inference by the adversarial network named RAAE-ABFA (robust adversarial autoencoder with attention based feature adaptation for high dimensional data anomaly detection) in which we can generate extra samples when training data is not enough. Experimental results on several benchmark datasets show that the proposed method significantly outperforms state-of-the-art unsupervised anomaly detection methods and is more robust when training data is contaminated.
Yuda Gao, Bin Shi 0003, Bo Dong 0001, Yan Chen 0031, Lingyun Mi, Zhiping Huang
COMPSAC2
2020 TTED-PU: A Transferable Tax Evasion Detection Method Based on Positive and Unlabeled Learning
abstract
Tax evasion usually refers to taxpayers making false declarations in order to reduce their tax obligations. One of the most common types of tax evasion is to lower the declared taxable amount. This kind of behavior will lead to the loss of tax revenues and damage the fairness of taxation. One of the main roles of the tax authorities is to conduct tax evasion testing through efficient auditing methods. At present, by using machine learning technology along with large amounts of labeled data, tax evasion detection models have achieved good results in specific areas. However, it is a long and costly process for tax experts to label large amounts of data. Since, the data distribution characteristics vary from region to region, models cannot be used across regions. In this paper, we propose a new method called a transferable tax evasion detection method based on positive and unlabeled learning (TTED-PU), which uses only semi-supervised techniques to detect tax evasion in the source domain. In addition, we use the idea of transfer to adapt to the domain to predict tax evasion behavior on the target domain where labeled tax data are unavailable. We evaluate our method on real-world tax data set. The experimental results show that our model can detect tax evasion in both the source and target domains.
Fa Zhang 0002, Bin Shi 0003, Bo Dong 0001, Xiangting Ji
COMPSAC2
2020 A Tax Evasion Detection Method Based on Positive and Unlabeled Learning with Network Embedding Features
Lingyun Mi, Bo Dong 0001, Bin Shi 0003
ICONIP (2)3
2019 Memory/Disk Operation Aware Lightweight VM Live Migration Across Data-centers with Low Performance Impact
abstract
Live virtual machine migration technique allows migrating an entire OS with running applications from one physical host to another, while keeping all services available without interruption. It provides a flexible and powerful way to balance system load, save power and tolerant faults in data centers. Meanwhile, with the stringent requirements of latency, scalability, and availability, an increasing number of applications are deployed across distributed cloud data-centers. However, existing live migration approaches still suffer from long downtime and serious performance degradation in cross data-center scenes due to the mass of dirty retransmission, which limits the ability of cross data-center scheduling. In this paper, we propose a system named Memory/disk operation aware Lightweight VM Live Migration across data-centers with low performance impact (MLLM). It significantly improves the cross data-center migration performance by reducing the amount of dirty data in the migration process. In MLLM, we predict disk read workingset (i.e., more frequently read contents) and memory write workingset (i.e., more frequently write contents) based on the access sequence trace. And then we adjust the migration models and data transfer sequence based on the workingset information. We also present two optimizing methods to filter unused blocks and to de-duplicate data content by a hot data cache, thereby greatly decreasing the amount of data to be transferred. We implement MLLM on the QEMU/KVM platform and conduct several real-world experiments. The experimental results show that our method averagely reduces 67.0% of total migration time and 41.6% service downtime over existing methods.
Bin Shi 0003, Haiying Shen
INFOCOM1
2018 ShadowMonitor: An Effective In-VM Monitoring Framework with Hardware-Enforced Isolation
Bin Shi 0003, Lei Cui 0003, Bo Li 0005, Xudong Liu 0001, Zhiyu Hao, Haiying Shen
RAID1
2016 CloudAuditor: A Cloud Auditing Framework Based on Nested Virtualization
abstract
Recent years witness the successful adoption of Cloud computing. However, security remains the top concern for cloud users. The fundamental issue is that cloud providers cannot convince cloud users the trustworthiness of cloud platforms. In this paper, we propose a cloud auditing framework, named CloudAuditor, to examine the behaviors of cloud platforms. By leveraging nested virtualization technology, CloudAuditor could identify the stealthy memory and disk access from cloud platforms to users' virtual machines and can support the mainstream IaaS platforms such as VMware, Xen and KVM. We evaluate the effectiveness and efficiency of CloudAuditor through comprehensive experiments. The results show that CloudAuditor can identify the suspicious behaviors of cloud platforms with acceptable performance overhead.
Bin Shi 0003, Bo Li 0005
CSCloud4
2016 A Remote Backup Approach for Virtual Machine Images
abstract
Recent years witness the successful application of Cloud computing. Virtualization plays a key role in cloud computing and greatly facilitates application deployment and migration. Tenants' applications are hosted by virtual machines. The security and safety of user applications receive much attention from academia and industry. However, fault tolerance and availability issues of cloud applications are overlooked. In this paper, we focus on the high availability issue of virtual machines. We propose a remote backup approach, named LiveRB, for saving the running states of virtual machines in an online manner. The backup process operates in background and is transparent to the applications hosted in virtual machines. Live migration technique is used to save the running states of virtual machines. A virtual block device is designed to cache I/O operations in memory and save incremental virtual disk data of the virtual machine to a remote server. We implement LiveRB on KVM virtualization platform. We evaluate the effectiveness and efficiency of LiveRB through comprehensive experiments. The results show that LiveRB can lively backup a virtual machine to a remote server with only slight performance penalty.
Bin Shi 0003, Bo Li 0005
CSCloud4
2016 Controlling a Car Through OBD Injection
abstract
Internet of vehicles(IOV) is an application of Internet of things in Intelligent Transport System, and has attracted high attention of researchers. IOV brings network connectivity to traditional vehicles, while also introduces security risks. This paper presents experimental analysis on the security of vehicles with Internet connections and propose an approach to Controlling a Car Through OBD Injection. In the experiments, we successfully penetrated several types of cars in a wireless way. We also put out a multi-level safety model of cars, which divides cars into different groups and gives analysis and explanations of each group. All of these things are done for indicating a point of view that traditional cars are not safe enough on information security. It is surely risky to put a car without the ability to resist the attack of informational ways into the Internet of vehicles.
Yu Zhang 0097, Binbin Ge, Bin Shi 0003, Bo Li 0005
CSCloud4
2015 PARS: A Page-Aware Replication System for Efficiently Storing Virtual Machine Snapshots
abstract
Virtual machine (VM) snapshot enhances the system availability by saving the running state into stable storage during failure-free execution and rolling back to the snapshot point upon failures. Unfortunately, the snapshot state may be lost due to disk failures, so that the VM fails to be recovered. The popular distributed file systems employ replication technique to tolerate disk failures by placing redundant copies across disperse disks. However, unless user-specific personalization is provided, these systems consider the data in the file as of same importance and create identical copies of the entire file, leading to non-trivial additional storage overhead.
Lei Cui 0003, Tianyu Wo, Bo Li 0005, Jianxin Li 0002, Bin Shi 0003, Jinpeng Huai
VEE5
2014 FENet: An SDN-based scheme for virtual network management
abstract
Virtual networking is vital to efficient resource management in Clouds, and it is in fact one of the main services provided by many Cloud Computing platforms. Virtual network management needs to meet specific requirements, including tenant isolation and adaption to virtual machines' lifecycle. Most of the existing schemes for virtual network management are based on the use of overlay networks in order to achieve a desirable degree of flexibility. However, these schemes suffer from a common limit, i.e. relatively high performance penalty due to a complicated forwarding process. We address this performance concern by developing a new management scheme, FENet, which makes use of Software-Defined Networks (SDN) to create virtual networks and manage them via the SDN controller programs. We present the design of an SDN controller, with the definition of flow entry rules based on the OpenFlow protocol and the specification of a routing algorithm. The results from our experimental evaluation show that our SDN-based prototype can control virtual network interconnections and tenant isolation appropriately. FENet achieves about 30% better network performance than the management scheme based on OpenVPN and lower latency in comparison with the traditional bridging scheme.
Tianyu Wo, Lei Cui 0003, Bin Shi 0003, Jie Xu 0007
ICPADS4