EDBT 2026 Demo / reviewers in the wild / expert
Xibin Zhao
dblp:62/5754
· DBLP profile ↗
98ranked-venue papers
4as first author
54since 2021 · last 2026
0000-0002-6168-7016ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 2 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 1 first-author · 16 since 2021Computer networks · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 6 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Systems, architecture and hardware · 7 · 3 since 2021Security and privacy · 7 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Domain Generalization and Subpopulation Shift for Generalizable Graph-Level Anomaly DetectionabstractGraph-level anomaly detection (GLAD), which identifies rare or atypical graphs within a graph set, is crucial for applications such as image analysis, industrial defect inspection and fraud detection. However, existing GLAD approaches typically rely on the in-distribution hypothesis while lacking generalization capability for out-of-distribution (OOD) scenarios (e.g., different graph sizes), which largely limits the application in the real world. For the first time, we formulate the OOD generalization problem for GLAD, where testing graph data exhibit significant distributional shifts from training data. To tackle two common types of distributional shifts, domain generalization and subpopulation shift, we propose the Fine-Grained Subpopulation Graph-Level Anomaly Detection (FGS-GLAD). First, we propose a Graph Information Bottleneck-based Anomaly Detection Module (GIB4AD) that implements graph reverse distillation and graph information bottleneck on the graph to enhance task-relevant feature extraction for domain generalization. Second, We propose a Fine-Grained Subpoulation Inference Module (FGSI) to predict fine-grained subpopulations and focus on critical inter-subpopulation features through a supervised contrastive mechanism. Experiments on seven benchmark datasets and ten baselines demonstrate our model's superiority in handling domain generalization and subpopulation shift. Xiaoxiang Li, Xihe Xie, Hai Wan, Xibin Zhao |
AAAI | 4 |
| 2026 | Interpretable and Robust Behavior Abstraction via Environment-Disentangled Heterogeneous GraphabstractTo identify the root causes of attacks, behavior abstraction (BA) converts audit logs into multiple behavior graphs and finds similar ones, which has proven effective in bridging the semantic gap and reducing manual workload. Existing works fail to achieve both interpretability and generalization, while also exhibiting limited robustness when facing adversarial attacks. In this paper, we give the first attempt at interpretable and robust behavior abstraction and propose a novel method called Environment-Disentangled Heterogeneous Graph Neural Network (EDHGNN). Motivated by Information Bottleneck (IB) principle, we propose a Heterogeneous Subgraph Disentanglement (HSD) module to disentangle label-relevant and environmental subgraphs through single optimization. We also introduce an Adapted Graph-Level Attention (AGLA) module to extract minimal sufficient representations from label-relevant subgraphs, a Label-Guided Graph Reconstructor (LGGR) to maximize environmental information coverage via reconstruction, and a Relevance Discriminator (RD) to enhance disentanglement quality. Additionally, we construct a new dataset contains ground-truth explanations and 4,160 behavior graphs. Extensive experiments demonstrate that EDHGNN outperforms the state-of-the-art methods in terms of interpretability and robustness against adversarial attacks. Zhibin Ni, Hai Wan, Xibin Zhao |
AAAI | 3 |
| 2026 | Disentangled Generation-Based Prototypical Alignment for Few-Shot Unsupervised Domain Adaptation in Graph-Level Anomaly DetectionabstractGraph-Level Anomaly Detection (GLAD) seeks to identify anomalous graphs within graph datasets, which has significant applications across diverse real-world fields. Most existing GLAD methods are trained in an unsupervised manner due to high costs for labeling, resulting in sub-optimal performance when compared to supervised methods. To fill this gap, we propose a Disentangled Generation-Based Prototypical Alignment (DGPA) method that extends graph-level anomaly detection to Few-Shot Unsupervised Domain Adaptation (FUDA) setting, aiming to identify anomalous graphs from a set of unlabeled graphs (target domain) by using partially labeled graphs from a different but related domain (source domain), which fulfills the practical requirement of transferring anomaly knowledge. This is specifically achieved through a dedicated Disentangled Sample Generation module, which addresses label scarcity by generating faithful samples with disentangled representation learning grounded in Information Bottleneck principle, along with a Graph-based Prototypical Self-Supervision module, which alleviates domain shift by encoding and aligning semantic structures in the shared latent space across domains in a self-supervised manner. Extensive experiments on five benchmark datasets reveal the effectiveness of our proposed DGPA. Zhibin Ni, Chenghao Zhang 0006, Hai Wan, Xibin Zhao |
AAAI | 4 |
| 2026 | Advanced Global Wildfire Activity Modeling with Hierarchical Graph ODE
Fan Xu 0009, Wei Gong 0001, Hao Wu 0094, Lilan Peng, Nan Wang 0015, Qingsong Wen, Xian Wu 0001, Kun Wang 0056, Xibin Zhao |
KDD (1) | 9 |
| 2026 | APIECHO: Training-Less Anomaly Detection via Intra-API Behavioral Comparison for Web Applications
Yihao Peng, Yiming Wu 0009, Du Wu, Shouling Ji, Hai Wan, Xibin Zhao |
SP | 6 |
| 2026 | HyperDetector: Advanced Persistent Threat Detection via Hypergraph Neural Networks with Enhanced Global PerceptionabstractAdvanced Persistent Threats (APTs) represent sophisticated cyberattacks that evade detection through stealthy, multistage operations, posing severe risks to critical infrastructure and organizational security. Due to their ability to effectively capture contextual information of attack behaviors, provenance graphs have emerged as a promising approach for APT detection. However, traditional binary edges in provenance graphs fail to represent the collaborative nature of APT attacks, where multiple entities coordinate in single operations, and local graph structures cannot capture the long-range dependencies across attack stages. To address these challenges, we propose HyperDetector, a novel hypergraph-based method for APT detection. First, we introduce hypergraph representation for provenance data, where hyperedges naturally connect multiple entities involved in system events, preserving the higher-order relational structures that characterize APT behaviors. Second, we employ block self-attention mechanisms that enable global reasoning across distant hypergraph regions, effectively linking dispersed attack indicators throughout the system. Through the synergistic integration of these approaches, HyperDetector achieves comprehensive understanding of both localized multi-entity collaborative behaviors and system-wide attack propagation patterns. Extensive evaluations across multiple prominent datasets demonstrate that HyperDetector outperforms state-of-the-art methods, showcasing its effectiveness for robust and holistic APT detection. Additionally, we make our code and datasets publicly available to facilitate reproducibility and foster further research in this critical area. Ziyue Wu, Nan Wang 0013, Jiqiang Liu, Hairong Dong 0001, Xibin Zhao |
WWW | 5 |
| 2026 | ERINYES: Request-Level Provenance Analysis for Serverless AttacksabstractThe serverless architecture has attracted significant attention due to its cost-effectiveness and ease of management. However, the serverless framework increases the attack surface of applications, resulting in frequent security incidents. Consequently, conducting comprehensive attack investigation and analysis in serverless applications has become critically important. Current serverless investigation methods face challenges such as dependency explosion (DE), incomplete information records, and lack of user transparency. These challenges lead to inadequate visibility of application interaction behaviors, complicating effective attack investigation and analysis. To mitigate these issues, this paper introducesErinyes, a solution that facilitates request-level attack investigation and analysis in serverless environments through the construction of provenance graph.Erinyesimproves the visibility of serverless applications via three core components. The partition enabling module effectively partitions function operations based on incoming requests; the log collection module is responsible for aggregating audit and network logs pertinent to function operations; and the provenance graph builder consolidates and parses the collected logs into a comprehensive provenance graph.Erinyeshas been evaluated on the OpenFaaS platform in 5 distinct attack scenarios, achieving an average accuracy of 99.6% in execution partition, a completeness of 100% in provenance graph, and an average runtime overhead of 7.05%. Hao Xi, Hai Wan, Xibin Zhao, Mohsen Guizani |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Knowledge is Power: A Knowledge Graph-Based Approach for Mobile Malware Traceability Analysis
Yao Zhang 0019, Guangquan Xu, Xiaohong Li 0001, Sen Chen 0001, Zhenchang Xing, Yude Bai, Yongqiang Lyu 0001, Wei Gong 0001, Xibin Zhao |
IEEE Trans. Mob. Comput. | 10 |
| 2026 | Multimodal Anomaly Detection for Microservice Systems via Grassmann Manifolds-Based Graph FusionabstractMicroservice architecture decomposes complex applications into multiple small, independent services, enabling independent deployment and scaling. However, microservice systems introduce complex and dynamic interactions between service instances. When a service instance fails, it can significantly degrade the performance of the entire system, harm the user experience, and cause substantial economic losses. Therefore, effective anomaly detection for service instances is crucial to ensure system reliability. Metrics, logs, and traces provide complementary insights into microservice operations, and recent approaches have attempted to jointly exploit these modalities for anomaly detection. However, the varying data volumes per interval and heterogeneous data types across modalities, combined with the complex and elusive relationships between them, make effective representation and learning challenging. To address these limitations, we propose MGFusion, an unsupervised, end-to-end multimodal anomaly detection method for microservice systems using Grassmann manifolds-based graph fusion. MGFusion employs a robust method to unify metrics, logs, and traces into time series, enabling consistent processing across modalities. It leverages multiple graph structure learning (GSL) techniques to explore multi-view relationships between the modalities. To mitigate the impact of data noise, we further introduce prior knowledge based on the Adamic-Adar Index (AAI) and employ a Grassmann manifolds-based graph fusion method to combine multiple basic graph structures. Finally, anomaly detection is achieved using Diffusion Convolutional Recurrent Neural Network (DCRNN) predictors and anomaly score calculation. Extensive experiments on a public dataset demonstrate that MGFusion effectively fuses multimodal data and captures complex relationships, significantly improving the accuracy and efficiency of anomaly detection. Compared to the state-of-the-art unsupervised multimodal anomaly detection method, AnoFusion, MGFusion achieves an improvement in the F1 score ranging from 11.1% to 15.3%. Shiming He, Keyao Feng, Kaixuan Meng, Kun Xie 0001, Xibin Zhao |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person RetrievalabstractVideo surveillance systems are crucial components for ensuring public safety and management in smart city. As a fundamental task in video surveillance, text-to-image person retrieval aims to retrieve the target person from an image gallery that best matches the given text description. Most existing text-to-image person retrieval methods are trained in a supervised manner that requires sufficient labeled data in the target domain. However, it is common in practice that only unlabeled data is available in the target domain due to the difficulty and cost of data annotation, which limits the generalization of existing methods in practical application scenarios. To address this issue, we propose a novel unsupervised domain adaptation method, termed Graph-Based Cross-Domain Knowledge Distillation (GCKD), to learn the cross-modal feature representation for text-to-image person retrieval in a cross-dataset scenario. The proposed GCKD method consists of two main components. Firstly, a graph-based multi-modal propagation module is designed to bridge the cross-domain correlation among the visual and textual samples. Secondly, a contrastive momentum knowledge distillation module is proposed to learn the cross-modal feature representation using the online knowledge distillation strategy. By jointly optimizing the two modules, the proposed method is able to achieve efficient performance for cross-dataset text-to-image person retrieval. Extensive experiments on three publicly available text-to-image person retrieval datasets demonstrate the effectiveness of the proposed GCKD method, which consistently outperforms the state-of-the-art baselines. Bingjun Luo, Xibin Zhao |
AAAI | 5 |
| 2025 | Robust Heterogeneous Graph Classification for Molecular Property Prediction with Information BottleneckabstractHeterogeneous Graph Neural Networks (HGNNs) have achieved state-of-the-art performance in classifying molecular graphs, capitalizing on their ability to capture rich semantics. However, HGNNs for molecule property prediction exhibit significant susceptibility to adversarial attacks—a challenge that prior research has entirely overlooked. To fill this gap, this paper introduces the first study focused on robust graph-level representation learning tailored for heterogeneous molecular graphs. To achieve this goal, we propose a comprehensive Robust Heterogeneous Graph Classification (RHGC) framework grounded in the Information Bottleneck principle, which aims to identify the most informative and least noisy heterogeneous subgraphs to derive robust, holistic representations. This is specifically accomplished through a dedicated Node Semantic Purifier, which enhances node-level and semantic-level robustness by eliminating label-irrelevant interference using graph stochastic attention and the Hilbert-Schmidt Independence Criterion, along with a Global Graph Disentanglement method, which improves graph-level robustness by addressing information leak. Experiments on three molecular benchmarks demonstrate that RHGC enhances accuracy by an average of 5.06% under all three attack settings and meanwhile by 4.33% on clean data. Zhibin Ni, Hai Wan, Xibin Zhao |
AAAI | 4 |
| 2025 | Adaptive Gaussian Mixture Model with Hierarchical Propagation for One-Class Graph Fraud DetectionabstractExisting graph fraud detection (GFD) methods have made remarkable progress with well-labeled training samples. However, in real applications, adequate training data may be unavailable due to the high cost of manual annotation and the scarcity of fraud samples. Therefore, we explore the one-class graph fraud detection task for the first time, which trains the model only on normal data and can detect fraud samples during inference. This task faces two main challenges: heterogeneity discrepancy in different relationships and diverse distribution of the normal data. To address the above challenges, we propose a novel one-class GFD method named OC-GFD. We design a hierarchical message propagation mechanism that learns both global features and local features under different relationships to accurately extract node representations from the GNN model. Subsequently, we integrate our model with an adaptive Gaussian mixture module to capture the diverse distribution of normal samples, enhancing the characterization of subtle differences in normal behavior and improving fraud detection accuracy. Experimental results show that OC-GFD outperforms state-of-the-art graph fraud detection and one-class classification approaches on Yelp and Amazon datasets in the one-class scenario. Code is available at https://github.com/THSS-GAD/OC-GFD. Xiaoxiang Li, Zhibin Ni, Hai Wan, Xibin Zhao |
ICME | 7 |
| 2025 | Detecting and Characterizing APT Attacks in the Open WorldabstractThe Intrusion Detection System (IDS) is an essential component of cybersecurity for Advanced Persistent Threat (APT) defense. A successful APT attack is a series of tactics aimed at achieving specific goals. Due to the versatility of these tactics, IDS must respond to numerous novel and previously unobserved attacks. However, traditional IDS systems are ineffective in defending against unknown attacks, as they assume that training and real data belong to the same distribution. To tackle this problem, we introduce OpenSentinel, which leverages a deep open set recognition method to effectively detect unknown attacks and pinpoint them to specific APT stages. With a specially designed log modeling approach and a neural network model, OpenSentinel generates human-readable reports to characterize attacks and facilitate further analysis for security experts. We validate the detection performance of OpenSentinel in two experimental environments with over 100 scenarios. Qualitative and quantitative results demonstrate that our method achieves an accuracy of over 90% and remains robust when facing real-world attacks. Meanwhile, we developed a benchmark APT attack dataset with well-defined stages named BeATT&CKed, which can be used for future research. Hao Xi, Yibin Han, Xiaoxiang Li, Jingwei Song, Hai Wan, Xibin Zhao |
ICPADS | 7 |
| 2025 | Breaking the Discretization Barrier of Continuous Physics Simulation LearningabstractThe modeling of complicated time-evolving physical dynamics from partial observations is a long-standing challenge. Particularly, observations can be sparsely distributed in a seemingly random or unstructured manner, making it difficult to capture highly nonlinear features in a variety of scientific and engineering problems. However, existing data-driven approaches are often constrained by fixed spatial and temporal discretization. While some researchers attempt to achieve spatio-temporal continuity by designing novel strategies, they either overly rely on traditional numerical methods or fail to truly overcome the limitations imposed by discretization. To address these, we propose CoPS, a purely data-driven methods, to effectively model continuous physics simulation from partial observations. Specifically, we employ multiplicative filter network to fuse and encode spatial information with the corresponding observations. Then we customize geometric grids and use message-passing mechanism to map features from original spatial domain to the customized grids. Subsequently, CoPS models continuous-time dynamics by designing multi-scale graph ODEs, while introducing a Markov-based neural auto-correction module to assist and constrain the continuous extrapolations. Comprehensive experiments demonstrate that CoPS advances the state-of-the-art methods in space-time continuous modeling across various scenarios. The source code is available at~\url{https://github.com/Sunxkissed/CoPS}. Fan Xu 0009, Hao Wu 0094, Nan Wang 0015, Lilan Peng, Kun Wang 0042, Wei Gong 0001, Xibin Zhao |
NeurIPS | 7 |
| 2025 | AutoLabel: Automated Fine-Grained Log Labeling for Cyber Attack Dataset Generation
Yihao Peng, Tongxin Zhang, Jieshao Lai, Hai Wan, Xibin Zhao |
USENIX Security Symposium | 7 |
| 2025 | FG-CIBGC: A Unified Framework for Fine-Grained and Class-Incremental Behavior Graph ClassificationabstractLearning-based Behavior Graph Classification (BGC) is widely used in Internet infrastructure for partitioning and identifying similar behavior graphs, yet its real-world application faces notable challenges. The challenges are: (i) fine-grained emerging behavior graphs, and (ii) incremental model adaptations. To tackle these issues, we propose to (i) mine semantics in multi-source logs using Large Language Models (LLMs) under In-Context Learning (ICL), and (ii) bridge the gap between Out-Of-Distribution (OOD) detection and class-incremental graph learning. Based on these ideas, we develop the first unified framework termed as Fine-Grained and Class-Incremental Behavior Graph Classification (FG-CIBGC ). It consists of two novel modules, i.e., gPartition and gAdapt, that are used for partitioning fine-grained graphs and performing unknown class detection and adaptation, respectively. To validate FG-CIBGC, we introduce a new benchmark, including a 4,992-graph, 32-class dataset from 8 attack scenarios and a novel Edge Intersection over Union (EIoU) metric. Extensive experiments show FG-CIBGC outperforms baselines on fine-grained class-incremental BGC task and generates behavior graphs which enhance downstream tasks. Zhibin Ni, Pan Fan, Shengzhuo Dai, Hai Wan, Xibin Zhao |
WWW | 6 |
| 2025 | Channel pruning on frequency response
Lin Bie, Chenggang Yan 0001, Xibin Zhao, Yue Gao 0002 |
Sci. China Inf. Sci. | 5 |
| 2025 | Factor-wise disentangled contrastive learning for cross-domain few-shot molecular property prediction
Zhibin Ni, Chenghao Zhang 0006, Hai Wan, Xibin Zhao |
Frontiers Comput. Sci. | 4 |
| 2025 | Cross-Modal 3D Shape Retrieval via Heterogeneous Dynamic Graph RepresentationabstractCross-modal 3D shape retrieval is a crucial and widely applied task in the field of 3D vision. Its goal is to construct retrieval representations capable of measuring the similarity between instances of different 3D modalities. However, existing methods face challenges due to the performance bottlenecks of single-modal representation extractors and the modality gap across 3D modalities. To tackle these issues, we propose a Heterogeneous Dynamic Graph Representation (HDGR) network, which incorporates context-dependent dynamic relations within a heterogeneous framework. By capturing correlations among diverse 3D objects, HDGR overcomes the limitations of ambiguous representations obtained solely from instances. Within the context of varying mini-batches, dynamic graphs are constructed to capture proximal intra-modal relations, and dynamic bipartite graphs represent implicit cross-modal relations, effectively addressing the two challenges above. Subsequently, message passing and aggregation are performed using Dynamic Graph Convolution (DGConv) and Dynamic Bipartite Graph Convolution (DBConv), enhancing features through heterogeneous dynamic relation learning. Finally, intra-modal, cross-modal, and self-transformed features are redistributed and integrated into a heterogeneous dynamic representation for cross-modal 3D shape retrieval. HDGR establishes a stable, context-enhanced, structure-aware 3D shape representation by capturing heterogeneous inter-object relationships and adapting to varying contextual dynamics. Extensive experiments conducted on the ModelNet10, ModelNet40, and real-world ABO datasets demonstrate the state-of-the-art performance of HDGR in cross-modal and intra-modal retrieval tasks. Moreover, under the supervision of robust loss functions, HDGR achieves remarkable cross-modal retrieval against label noise on the 3D MNIST dataset. The comprehensive experimental results highlight the effectiveness and efficiency of HDGR on cross-modal 3D shape retrieval. Yue Dai 0003, Yifan Feng 0001, Nan Ma 0012, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Filter Pruning by High-Order Spectral ClusteringabstractLarge amount of redundancy is widely present in convolutional neural networks (CNNs). Identifying the redundancy in the network and removing the redundant filters is an effective way to compress the CNN model size with a minimal reduction in performance. However, most of the existing redundancy-based pruning methods only consider the distance information between two filters, which can only model simple correlations between filters. Moreover, we point out that distance-based pruning methods are not applicable for high-dimensional features in CNN models by our experimental observations and analysis. To tackle this issue, we propose a new pruning strategy based on high-order spectral clustering. In this approach, we use hypergraph structure to construct complex correlations among filters, and obtain high-order information among filters by hypergraph structure learning. Finally, based on the high-order information, we can perform better clustering on the filters and remove the redundant filters in each cluster. Experiments on various CNN models and datasets demonstrate that our proposed method outperforms the recent state-of-the-art works. For example, with ResNet50, we achieve a 57.1% FLOPs reduction with no accuracy drop on ImageNet, which is the first to achieve lossless pruning with such a high compression ratio. Yubo Zhang 0006, Lin Bie, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Real-Time Cross-Domain Gesture and User Identification via COTS WiFiabstractWiFi-based gesture recognition has emerged as a promising alternative to computer vision, enabling seamless integration and enhanced interaction in human-computer interaction systems. Simultaneously identifying users during gesture recognition is vital for improving security and personalization. However, existing WiFi-based dual-task recognition approaches often rely on handcrafted features, which hinder precision and introduce delays in cross-domain scenarios. To address these challenges, we propose WiDual, a real-time system for cross-domain gesture recognition and user identification using WiFi signals. By integrating spatial and channel attention mechanisms, WiDual adaptively extracts crucial features for dual-task recognition. The system employs Channel State Information (CSI) visualization to convert WiFi signals into images, facilitating efficient feature extraction and minimizing information loss and latency. Furthermore, a collaborative module fuses gesture and user identity features, enhancing recognition performance. Experimental evaluations on a public dataset with six gestures and six users across diverse environments demonstrate WiDual's effectiveness. It achieves 96% accuracy in cross-domain gesture recognition and 91.27% in user identification. Compared to state-of-the-art methods, WiDual improves user identification accuracy by 26%, gesture recognition by 8%, and reduces processing time sixfold, showcasing its potential for real-time applications. Chenhong Cao, Miaoling Dai, Wei Gong 0001, Xibin Zhao |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Multi-Energy Guided Image Translation with Stochastic Differential Equations for Near-Infrared Facial Expression RecognitionabstractIllumination variation has been a long-term challenge in real-world facial expression recognition (FER). Under uncontrolled or non-visible light conditions, near-infrared (NIR) can provide a simple and alternative solution to obtain high-quality images and supplement the geometric and texture details that are missing in the visible (VIS) domain. Due to the lack of large-scale NIR facial expression datasets, directly extending VIS FER methods to the NIR spectrum may be ineffective. Additionally, previous heterogeneous image synthesis methods are restricted by low controllability without prior task knowledge. To tackle these issues, we present the first approach, called for NIR-FER Stochastic Differential Equations (NFER-SDE), that transforms face expression appearance between heterogeneous modalities to the overfitting problem on small-scale NIR data. NFER-SDE can take the whole VIS source image as input and, together with domain-specific knowledge, guide the preservation of modality-invariant information in the high-frequency content of the image. Extensive experiments and ablation studies show that NFER-SDE significantly improves the performance of NIR FER and achieves state-of-the-art results on the only two available NIR FER datasets, Oulu-CASIA and Large-HFE. Bingjun Luo, Xibin Zhao, Yue Gao 0002 |
AAAI | 5 |
| 2024 | Hypergraph-Guided Disentangled Spectrum Transformer Networks for Near-Infrared Facial Expression RecognitionabstractWith the strong robusticity on illumination variations, near-infrared (NIR) can be an effective and essential complement to visible (VIS) facial expression recognition in low lighting or complete darkness conditions. However, facial expression recognition (FER) from NIR images presents a more challenging problem than traditional FER due to the limitations imposed by the data scale and the difficulty of extracting discriminative features from incomplete visible lighting contents. In this paper, we give the first attempt at deep NIR facial expression recognition and propose a novel method called near-infrared facial expression transformer (NFER-Former). Specifically, to make full use of the abundant label information in the field of VIS, we introduce a Self-Attention Orthogonal Decomposition mechanism that disentangles the expression information and spectrum information from the input image, so that the expression features can be extracted without the interference of spectrum variation. We also propose a Hypergraph-Guided Feature Embedding method that models some key facial behaviors and learns the structure of the complex correlations between them, thereby alleviating the interference of inter-class similarity. Additionally, we construct a large NIR-VIS Facial Expression dataset that includes 360 subjects to better validate the efficiency of NFER-Former. Extensive experiments and ablation studies show that NFER-Former significantly improves the performance of NIR FER and achieves state-of-the-art results on the only two available NIR FER datasets, Oulu-CASIA and Large-HFE. Bingjun Luo, Xibin Zhao, Yue Gao 0002 |
AAAI | 5 |
| 2024 | Revisiting Graph-Based Fraud Detection in Sight of Heterophily and SpectrumabstractGraph-based fraud detection (GFD) can be regarded as a challenging semi-supervised node binary classification task. In recent years, Graph Neural Networks (GNN) have been widely applied to GFD, characterizing the anomalous possibility of a node by aggregating neighbor information. However, fraud graphs are inherently heterophilic, thus most of GNNs perform poorly due to their assumption of homophily. In addition, due to the existence of heterophily and class imbalance problem, the existing models do not fully utilize the precious node label information. To address the above issues, this paper proposes a semi-supervised GNN-based fraud detector SEC-GFD. This detector includes a hybrid filtering module and a local environmental constraint module, the two modules are utilized to solve heterophily and label utilization problem respectively. The first module starts from the perspective of the spectral domain, and solves the heterophily problem to a certain extent. Specifically, it divides the spectrum into various mixed-frequency bands based on the correlation between spectrum energy distribution and heterophily. Then in order to make full use of the node label information, a local environmental constraint module is adaptively designed. The comprehensive experimental results on four real-world fraud detection datasets denote that SEC-GFD outperforms other competitive graph-based fraud detectors. We release our code at https://github.com/Sunxkissed/SEC-GFD. Fan Xu 0009, Nan Wang 0015, Hao Wu 0094, Xuezhi Wen, Xibin Zhao, Hai Wan |
AAAI | 5 |
| 2024 | DSFM: Enhancing Functional Code Clone Detection with Deep Subtree InteractionsabstractFunctional code clone detection is important for software maintenance. In recent years, deep learning techniques are introduced to improve the performance of functional code clone detectors. By representing each code snippet as a vector containing its program semantics, syntactically dissimilar functional clones are detected. However, existing deep learning-based approaches attach too much importance to code feature learning, hoping to project all recognizable knowledge of a code snippet into a single vector. We argue that these deep learning-based approaches can be enhanced by considering the characteristics of syntactic code clone detection, where we need to compare the contents of the source code (e.g., intersection of tokens, similar flow graphs, and similar subtrees) to obtain code clones. In this paper, we propose a novel deep learning-based approach named DSFM, which incorporates comparisons between code snippets for detecting functional code clones. Specifically, we improve the typical deep clone detectors with deep subtree interactions that compare every two subtrees extracted abstract syntax trees (ASTs) of two code snippets, thereby introducing more fine-grained semantic similarity. By conducting extensive experiments on three widely-used datasets, GCJ, OJClone, and BigCloneBench, we demonstrate the great potential of deep subtree interactions in code clone detection task. The proposed DSFM outperforms the state-of-the-art approaches, including two traditional approaches, two unsupervised and four supervised deep learning-based baselines. Shaohua Qiang, Dinghong Song, Min Zhou 0001, Hai Wan, Xibin Zhao, Ping Luo 0004, Hongyu Zhang 0002 |
ICSE | 6 |
| 2024 | Trident: Detecting SQL Injection Attacks via Abstract Syntax Tree-based Neural NetworkabstractSQL injection attacks have posed a significant threat to web applications for decades. They obfuscate malicious codes into natural SQL statements so as to steal sensitive data, making them difficult to detect. Generally, malicious signals can be identified by using the contextual information of SQL statements. Such contextual information, however, is not always easily captured. Due to the fact that SQL as a formal language is highly structured, two tokens that are spatially far away may be semantically very close. An effective approach thus should take the structural feature of SQL statements into account when modeling their contextual information. Min Zhou 0001, Hai Wan, Xibin Zhao |
ASE | 5 |
| 2024 | GLADformer: A Mixed Perspective for Graph-Level Anomaly Detection
Fan Xu 0009, Nan Wang 0015, Hao Wu 0094, Xuezhi Wen, Dalin Zhang 0003, Siyang Lu, Binyong Li, Wei Gong 0001, Hai Wan, Xibin Zhao |
ECML/PKDD (6) | 10 |
| 2024 | Fairness based on anomaly score and adaptive weight in network attack detection
Xuezhi Wen, Meiqi Gao, Nan Wang 0015, Jiahui Ma, Dalin Zhang 0003, Xibin Zhao, Jiqiang Liu |
Inf. Sci. | 6 |
| 2024 | The Last Mile of Attack Investigation: Audit Log Analysis Toward Software Vulnerability LocationabstractCyberattacks have caused significant damage and losses in various domains. While existing attack investigations against cyberattacks focus on identifying compromised system entities and reconstructing attack stories, there is a lack of information that security analysts can use to locate software vulnerabilities and thus fix them. In this paper, we present AiVl, a novel software vulnerability location method to push the attack investigation further. AiVl relies on logs collected by the default built-in system auditing tool and program binaries within the system. Given a sequence of malicious log entries obtained through traditional attack investigations, AiVl can identify the functions responsible for generating these logs and trace the corresponding function call paths, namely the location of vulnerabilities in the source code. To achieve this, AiVl proposes an accurate, concise, and complete specific-domain program modeling that constructs all system call flows by static-dynamic techniques from the binary, and develops effective matching-based algorithms between the log sequences and program models. To evaluate the effectiveness of AiVl, we conduct experiments on 18 real-world attack scenarios and an APT, covering comprehensive categories of vulnerabilities and program execution classes. The results show that compared to actual vulnerability remediation reports, AiVl achieves a 100% precision and an average recall of 90%. Besides, the runtime overhead is reasonable, averaging at 7%. Changhua Chen, Tingzhen Yan, Chenxuan Shi, Hao Xi, Zhirui Fan, Hai Wan, Xibin Zhao |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Enhancing Single-Frame Supervision for Better Temporal Action LocalizationabstractTemporal action localization aims to identify the boundaries and categories of actions in videos, such as scoring a goal in a football match. Single-frame supervision has emerged as a labor-efficient way to train action localizers as it requires only one annotated frame per action. However, it often suffers from poor performance due to the lack of precise boundary annotations. To address this issue, we propose a visual analysis method that aligns similar actions and then propagates a few user-provided annotations (e.g., boundaries, category labels) to similar actions via the generated alignments. Our method models the alignment between actions as a heaviest path problem and the annotation propagation as a quadratic optimization problem. As the automatically generated alignments may not accurately match the associated actions and could produce inaccurate localization results, we develop a storyline visualization to explain the localization results of actions and their alignments. This visualization facilitates users in correcting wrong localization results and misalignments. The corrections are then used to improve the localization results of other actions. The effectiveness of our method in improving localization performance is demonstrated through quantitative evaluation and a case study. Changjian Chen, Jiashu Chen, Weikai Yang, Haoze Wang, Johannes Knittel, Xibin Zhao, Steffen Koch 0001, Thomas Ertl, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Learning Deep Hierarchical Features with Spatial Regularization for One-Class Facial Expression RecognitionabstractExisting methods on facial expression recognition (FER) are mainly trained in the setting when multi-class data is available. However, to detect the alien expressions that are absent during training, this type of methods cannot work. To address this problem, we develop a Hierarchical Spatial One Class Facial Expression Recognition Network (HS-OCFER) which can construct the decision boundary of a given expression class (called normal class) by training on only one-class data. Specifically, HS-OCFER consists of three novel components. First, hierarchical bottleneck modules are proposed to enrich the representation power of the model and extract detailed feature hierarchy from different levels. Second, multi-scale spatial regularization with facial geometric information is employed to guide the feature extraction towards emotional facial representations and prevent the model from overfitting extraneous disturbing factors. Third, compact intra-class variation is adopted to separate the normal class from alien classes in the decision space. Extensive evaluations on 4 typical FER datasets from both laboratory and wild scenarios show that our method consistently outperforms state-of-the-art One-Class Classification (OCC) approaches. Bingjun Luo, Sicheng Zhao, Xibin Zhao, Yue Gao 0002 |
AAAI | 6 |
| 2023 | Fairness with adaptive weight in network attack detectionabstractNetwork attacks aim to exploit vulnerabilities inherent in network protocols, which is widely used in many real-world applications. In the process of network anomaly detection, most methods train the model by minimizing the average empirical risk of all samples. However, due to the uneven distribution of samples from different protocols, detection models tend to be biased against minority protocols groups. To address this issue, we propose an adaptive weight assignment method for network attack detection, which emphasizes more on error-prone samples in prediction and enhances adequate representation of minority groups for fairness. We conduct experiments on two widely used datasets, KDD and NSLKDD. According to the results, our method achieves better performance than state-of-the-art methods for classification and regression tasks, and is robust to label noise in the test dataset. Xuezhi Wen, Nan Wang 0015, Yuanlin Sun, Fan Xu 0009, Dalin Zhang 0003, Xibin Zhao |
ICPADS | 6 |
| 2023 | Few-shot Message-Enhanced Contrastive Learning for Graph Anomaly DetectionabstractGraph anomaly detection plays a crucial role in identifying exceptional instances in graph data that deviate significantly from the majority. It has gained substantial attention in various domains of information security, including network intrusion, financial fraud, and malicious comments, et al. Existing methods are primarily developed in an unsupervised manner due to the challenge in obtaining labeled data. For lack of guidance from prior knowledge in unsupervised manner, the identified anomalies may prove to be data noise or individual data instances. In real-world scenarios, a limited batch of labeled anomalies can be captured, making it crucial to investigate the few-shot problem in graph anomaly detection. Taking advantage of this potential, we propose a novel few-shot Graph Anomaly Detection model called FMGAD (Few-shot Message-Enhanced Contrastive-based Graph Anomaly Detector). FMGAD leverages a self-supervised contrastive learning strategy within and across views to capture intrinsic and transferable structural representations. Furthermore, we propose the Deep-GNN message-enhanced reconstruction module, which extensively exploits the few-shot label information and enables long-range propagation to disseminate supervision signals to deeper unlabeled nodes. This module in turn assists in the training of self-supervised contrastive learning. Comprehensive experimental results on six real-world datasets demonstrate that FMGAD can achieve better performance than other state-of-the-art methods, regardless of artificially injected anomalies or domain-organic anomalies. Fan Xu 0009, Nan Wang 0015, Xuezhi Wen, Meiqi Gao, Chaoqun Guo, Xibin Zhao |
ICPADS | 6 |
| 2023 | Variance-Aware Bi-Attention Expression Transformer for Open-Set Facial Expression Recognition in the WildabstractDespite the great accomplishments of facial expression recognition (FER) models in closed-set scenarios, they still lack open-world robustness when it comes to handling unknown samples. To address the demands of operating in an open environment, open-set FER models should improve their performance in rejecting unknown samples while maintaining their efficiency in recognizing known expressions. With this goal in mind, we propose an open-set FER framework named Variance-Aware Bi-Attention Expression Transformer (VBExT), which enhances conventional closed-set FER models with open-world robustness for unknown samples. Specifically, to make full use of the expression representation capabilities of learned features, we introduce a bi-attention feature augmentation mechanism that learns the important regions and integrates the hierarchical features extracted by the emotional CNN backbone. We also propose a variance-aware distribution modeling method that adapts to the diverse distribution of different expression classes in the open environment, thereby enhancing the detection ability of unknown expressions. Additionally, we have constructed a Fine-Grained Light Facial Expression dataset that includes 30 different light brightnesses to better validate the efficiency of VBExT. Extensive experiments and ablation studies show that VBExT significantly improves the performance of open-set FER and achieves state-of-the-art results on CFEE (lab, basic), RAF-DB (wild, basic+compound), and FGL-FE (multiple light brightnesses, basic). Bingjun Luo, Jinghang Tan, Xibin Zhao, Yue Gao 0002 |
ACM Multimedia | 5 |
| 2023 | xASTNN: Improved Code Representations for Industrial PracticeabstractThe application of deep learning techniques in software engineering becomes increasingly popular. One key problem is developing high-quality and easy-to-use source code representations for code-related tasks. The research community has acquired impressive results in recent years. However, due to the deployment difficulties and performance bottlenecks, seldom these approaches are applied to the industry. In this paper, we present xASTNN, an eXtreme Abstract Syntax Tree (AST)-based Neural Network for source code representation, aiming to push this technique to industrial practice. The proposed xASTNN has three advantages. First, xASTNN is completely based on widely-used ASTs and does not require complicated data pre-processing, making it applicable to various programming languages and practical scenarios. Second, three closely-related designs are proposed to guarantee the effectiveness of xASTNN, including statement subtree sequence for code naturalness, gated recursive unit for syntactical information, and gated recurrent unit for sequential information. Third, a dynamic batching algorithm is introduced to significantly reduce the time complexity of xASTNN. Two code comprehension downstream tasks, code classification and code clone detection, are adopted for evaluation. The results demonstrate that our xASTNN can improve the state-of-the-art while being faster than the baselines. Min Zhou 0001, Xibin Zhao, Yang Chen 0001, Hongyu Zhang 0002 |
ESEC/SIGSOFT FSE | 3 |
| 2023 | Exploring Global and Local Information for Anomaly Detection with Normal SamplesabstractAnomaly detection aims to detect data that do not conform to regular patterns, and such data is also called outliers. The anomalies to be detected are often tiny in proportion, containing crucial information, and are suitable for application scenes like intrusion detection, fraud detection, fault diagnosis, e-commerce platforms, et al. However, in many realistic scenarios, only the samples following normal behavior are observed, while we can hardly obtain any anomaly information. To address such problem, we propose an anomaly detection method GALDetector which is combined of global and local information based on observed normal samples. The proposed method can be divided into a three-stage method. Firstly, the global similar normal scores and the local sparsity scores of unlabeled samples are computed separately. Secondly, potential anomaly samples are separated from the unlabeled samples corresponding to these two scores and corresponding weights are assigned to the selected samples. Finally, a weighted anomaly detector is trained by loads of samples, then the detector is utilized to identify else anomalies. To evaluate the effectiveness of the proposed method, we conducted experiments on three categories of real-world datasets from diverse domains, and experimental results show that our method achieves better performance when compared with other state-of-the-art methods. Fan Xu 0009, Nan Wang 0015, Xibin Zhao |
SMC | 3 |
| 2023 | TeSec: Accurate Server-side Attack Investigation for Web ApplicationsabstractThe user interface (UI) of web applications is usually the entry point of web attacks against enterprises and organizations. Finding the UI elements utilized by the intruders is of great importance both for attack interception and web application fixing. Current attack investigation methods targeting web UI either provide rough analysis results or have poor performance in high concurrency scenarios, which leads to heavy manual analysis work. In this paper, we propose TeSec, an accurate attack investigation method for web UI applications. TeSec makes use of two kinds of correlations. The first one, built from annotated audit log partitioned by PID/TID and delimiter-logs, captures the correspondence between audit log entries and web requests. The second one, modeled by an Aho-Corasick automaton built during system testing period, captures the correspondence between requests and the UI elements/events. Leveraging these two correlations, TeSec can accurately and automatically locate the UI elements/events (i.e., the root cause of the alarm) from an alarm, even in high concurrency scenarios. Furthermore, TeSec only needs to be deployed in the server and does not need to collect logs from the client-side browsers. We evaluate TeSec on 12 web applications. The experimental results show that the matching accuracy between UI events/elements and the alarm is above 99.6%. And security analysts only need to check no more than 2 UI elements on average for each individual forensics analysis. The maximum overhead of average response time and audit log space overhead are low (4.3% and 4.6% respectively). Yihao Peng, Yilun Sun, Xuancheng Zhang, Hai Wan, Xibin Zhao |
SP | 6 |
| 2023 | Structure Evolution on Manifold for Graph LearningabstractGraph has been widely used in various applications, while how to optimize the graph is still an open question. In this paper, we propose a framework to optimize the graph structure via structure evolution on graph manifold. We first define the graph manifold and search the best graph structure on this manifold. Concretely, associated with the data features and the prediction results of a given task, we define a graph energy to measure how the graph fits the graph manifold from an initial graph structure. The graph structure then evolves by minimizing the graph energy. In this process, the graph structure can be evolved on the graph manifold corresponding to the update of the prediction results. Alternatively iterating these two processes, both the graph structure and the prediction results can be updated until converge. It achieves the suitable structure for graph learning without searching all hyperparameters. To evaluate the performance of the proposed method, we have conducted experiments on eight datasets and compared with the recent state-of-the-art methods. Experiment results demonstrate that our method outperforms the state-of-the-art methods in both transductive and inductive settings. Hai Wan, Xinwei Zhang 0012, Yubo Zhang 0006, Xibin Zhao, Shihui Ying, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Graph Learning on Millions of Data in Seconds: Label Propagation Acceleration on Graph Using Data DistributionabstractGraph-based semi-supervised learning methods have been used in a wide range of real-world applications, e.g., from social relationship mining to multimedia classification and retrieval. However, existing methods are limited along with high computational complexity or not facilitating incremental learning, which may not be powerful to deal with large-scale data, whose scale may continuously increase, in real world. This paper proposes a new method called Data Distribution Based Graph Learning (DDGL) for semi-supervised learning on large-scale data. This method can achieve a fast and effective label propagation and supports incremental learning. The key motivation is to propagate the labels along smaller-scale data distribution model parameters, rather than directly dealing with the raw data as previous methods, which accelerate the data propagation significantly. It also improves the prediction accuracy since the loss of structure information can be alleviated in this way. To enable incremental learning, we propose an adaptive graph updating strategy which can update the model when there is distribution bias between new data and the already seen data. We have conducted comprehensive experiments on multiple datasets with sample sizes increasing from seven thousand to five million. Experimental results on the classification task on large-scale data demonstrate that our proposed DDGL method improves the classification accuracy by a large margin while consuming much less time compared to state-of-the-art methods. Yubo Zhang 0006, Shuyi Ji, Changqing Zou, Xibin Zhao, Shihui Ying, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Cost-Sensitive Hypergraph Learning With F-Measure OptimizationabstractThe imbalanced issue among data is common in many machine-learning applications, where samples from one or more classes are rare. To address this issue, many imbalanced machine-learning methods have been proposed. Most of these methods rely on cost-sensitive learning. However, we note that it is infeasible to determine the precise cost values even with great domain knowledge for those cost-sensitive machine-learning methods. So in this method, due to the superiority of F-measure on evaluating the performance of imbalanced data classification, we employ F-measure to calculate the cost information and propose a cost-sensitive hypergraph learning method with F-measure optimization to solve the imbalanced issue. In this method, we employ the hypergraph structure to explore the high-order relationships among the imbalanced data. Based on the constructed hypergraph structure, we optimize the cost value with F-measure and further conduct cost-sensitive hypergraph learning with the optimized cost information. The comprehensive experiments validate the effectiveness of the proposed method. Nan Wang 0015, Ruozhou Liang, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Cybern. | 3 |
| 2023 | Exploring High-Order Spatio-Temporal Correlations From Skeleton for Person Re-IdentificationabstractPerson re- identification (Re-ID) has become a hot research topic due to its widespread applications. Conducting person Re-ID in video sequences is a practical requirement, in which the crucial challenge is how to pursue a robust video representation based on spatial and temporal features. However, most of the previous methods only consider how to integrate part-level features in the spatio-temporal range, while how to model and generate the part-correlations is little exploited. In this paper, we propose a skeleton-based dynamic hypergraph framework, namely Skeletal Temporal Dynamic Hypergraph Neural Network (ST-DHGNN) for person Re-ID, which resorts to modeling the high-order correlations among various body parts based on a time series of skeletal information. Specifically, multi-shape and multi-scale patches are heuristically cropped from feature maps, constituting spatial representations in different frames. A joint-centered hypergraph and a bone-centered hypergraph are constructed in parallel from multiple body parts (i.e., head, trunk, and legs) with spatio-temporal multi-granularity in the entire video sequence, in which the graph vertices representing regional features and hyperedges denoting relationships. Dynamic hypergraph propagation containing the re- planning module and the hyperedge elimination module is proposed to better integrate features among vertices. Feature aggregation and attention mechanisms are also adopted to obtain a better video representation for person Re-ID. Experiments show that the proposed method performs significantly better than the state-of-the-art on three video-based person Re-ID datasets, including iLIDS-VID, PRID-2011, and MARS. Jiaxuan Lu, Hai Wan, Xibin Zhao, Nan Ma 0012, Yue Gao 0002 |
IEEE Trans. Image Process. | 4 |
| 2023 | Knowledge Conditioned Variational Learning for One-Class Facial Expression RecognitionabstractThe openness of application scenarios and the difficulties of data collection make it impossible to prepare all kinds of expressions for training. Hence, detecting expression absent during the training (called alien expression) is important to enhance the robustness of the recognition system. So in this paper, we propose a facial expression recognition (FER) model, named OneExpressNet, to quantify the probability that a test expression sample belongs to the distribution of training data. The proposed model is based on variational auto-encoder and enjoys several merits. First, different from conventional one class classification protocol, OneExpressNet transfers the useful knowledge from the related domain as a constraint condition of the target distribution. By doing so, OneExpressNet will pay more attention to the descriptive region for FER. Second, features from both source and target tasks will aggregate after constructing a skip connection between the encoder and decoder. Finally, to further separate alien expression from training expression, empirical compact variation loss is jointly optimized, so that training expression will concentrate on the compact manifold of feature space. The experimental results show that our method can achieve state-of-the-art results in one class facial expression recognition on small-scale lab-controlled datasets including CFEE and KDEF, and large-scale in-the-wild datasets including RAF-DB and ExpW. Bingjun Luo, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Image Process. | 5 |
| 2022 | Grow and Merge: A Unified Framework for Continuous Categories DiscoveryabstractAlthough a number of studies are devoted to novel category discovery, most of them assume a static setting where both labeled and unlabeled data are given at once for finding new categories. In this work, we focus on the application scenarios where unlabeled data are continuously fed into the category discovery system. We refer to it as the {\bf Continuous Category Discovery} ({\bf CCD}) problem, which is significantly more challenging than the static setting. A common challenge faced by novel category discovery is that different sets of features are needed for classification and category discovery: class discriminative features are preferred for classification, while rich and diverse features are more suitable for new category mining. This challenge becomes more severe for dynamic setting as the system is asked to deliver good performance for known classes over time, and at the same time continuously discover new classes from unlabeled data. To address this challenge, we develop a framework of {\bf Grow and Merge} ({\bf GM}) that works by alternating between a growing phase and a merge phase: in the growing phase, it increases the diversity of features through a continuous self-supervised learning for effective category mining, and in the merging phase, it merges the grown model with a static one to ensure satisfying performance for known classes. Our extensive studies verify that the proposed GM framework is significantly more effective than the state-of-the-art approaches for continuous category discovery. Xinwei Zhang 0012, Jianwen Jiang, Yutong Feng, Zhi-Fan Wu, Xibin Zhao, Hai Wan, Mingqian Tang, Rong Jin 0001, Yue Gao 0002 |
NeurIPS | 5 |
| 2022 | SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu |
Comput. Graph. | 3 |
| 2022 | Search-based cost-sensitive hypergraph learning for anomaly detection
Nan Wang 0015, Yubo Zhang 0006, Xibin Zhao, Yingli Zheng, Boya Zhou, Yue Gao 0002 |
Inf. Sci. | 3 |
| 2022 | Hypergraph Learning: Methods and PracticesabstractHypergraph learning is a technique for conducting learning on a hypergraph structure. In recent years, hypergraph learning has attracted increasing attention due to its flexibility and capability in modeling complex data correlation. In this paper, we first systematically review existing literature regarding hypergraph generation, including distance-based, representation-based, attribute-based, and network-based approaches. Then, we introduce the existing learning methods on a hypergraph, including transductive hypergraph learning, inductive hypergraph learning, hypergraph structure updating, and multi-modal hypergraph learning. After that, we present a tensor-based dynamic hypergraph representation and learning framework that can effectively describe high-order correlation in a hypergraph. To study the effectiveness and efficiency of hypergraph generation and learning methods, we conduct comprehensive evaluations on several typical applications, including object and action recognition, Microblog sentiment prediction, and clustering. In addition, we contribute a hypergraph learning development toolkit called THU-HyperG. Yue Gao 0002, Zizhao Zhang 0003, Haojie Lin, Xibin Zhao, Shaoyi Du, Changqing Zou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | View-Aware Geometry-Structure Joint Learning for Single-View 3D Shape ReconstructionabstractReconstructing a 3D shape from a single-view image using deep learning has become increasingly popular recently. Most existing methods only focus on reconstructing the 3D shape geometry based on image constraints. The lack of explicit modeling of structure relations among shape parts yields low-quality reconstruction results for structure-rich man-made shapes. In addition, conventional 2D-3D joint embedding architecture for image-based 3D shape reconstruction often omits the specific view information from the given image, which may lead to degraded geometry and structure reconstruction. We address these problems by introducing VGSNet, an encoder-decoder architecture for view-aware joint geometry and structure learning. The key idea is to jointly learn a multimodal feature representation of 2D image, 3D shape geometry and structure so that both geometry and structure details can be reconstructed from a single-view image. To this end, we explicitly represent 3D shape structures as part relations and employ image supervision to guide the geometry and structure reconstruction. Trained with pairs of view-aligned images and 3D shapes, the VGSNet implicitly encodes the view-aware shape information in the latent feature space. Qualitative and quantitative comparisons with the state-of-the-art baseline methods as well as ablation studies demonstrate the effectiveness of the VGSNet for structure-aware single-view 3D shape reconstruction. Xuancheng Zhang, Rui Ma 0011, Changqing Zou, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Kullback-Leibler Divergence Metric LearningabstractThe Kullback-Leibler divergence (KLD), which is widely used to measure the similarity between two distributions, plays an important role in many applications. In this article, we address the KLD metric-learning task, which aims at learning the best KLD-type metric from the distributions of datasets. Concretely, first, we extend the conventional KLD by introducing a linear mapping and obtain the best KLD to well express the similarity of data distributions by optimizing such a linear mapping. It improves the expressivity of data distribution, which means it makes the distributions in the same class close and those in different classes far away. Then, the KLD metric learning is modeled by a minimization problem on the manifold of all positive-definite matrices. To deal with this optimization task, we develop an intrinsic steepest descent method, which preserves the manifold structure of the metric in the iteration. Finally, we apply the proposed method along with ten popular metric-learning approaches on the tasks of 3-D object classification and document classification. The experimental results illustrate that our proposed method outperforms all other methods. Shuyi Ji, Zizhao Zhang 0003, Shihui Ying, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Cybern. | 5 |
| 2021 | View-Guided Point Cloud CompletionabstractThis paper presents a view-guided solution for the task of point cloud completion. Unlike most existing methods directly inferring the missing points using shape priors, we address this task by introducing ViPC (view-guided point cloud completion) that takes the missing crucial global structure information from an extra single-view image. By leveraging a framework that sequentially performs effective cross-modality and cross-level fusions, our method achieves significantly superior results over typical existing solutions on a new large-scale dataset we collect for the view-guided point cloud completion task. Xuancheng Zhang, Yutong Feng, Siqi Li 0001, Changqing Zou, Hai Wan, Xibin Zhao, Yandong Guo, Yue Gao 0002 |
CVPR | 6 |
| 2021 | TTDeep: Time-Triggered Scheduling for Real-Time Ethernet via Deep Reinforcement LearningabstractSchedule scheme is essential for real-time Ethernet. Due to the inevitable change of network configurations, the solution requires to be incrementally scheduled in a timely manner. Solver-based methods are time-consuming, while handcrafted scheduling heuristics require domain knowledge and professional expertise, and their application scenarios are usually limited. Instead of designing heuristic strategy manually, we propose TTDeep, a deep reinforcement learning schedule framework, to incrementally schedule Time-Triggered (TT) flows and adapt to various topologies. Our novel framework includes 3 key designs: a period layer to capture the periodical transmission nature of TT flows, the graph neural network to extract and represent topology features, and a 3-step selection paradigm to alleviate the huge action search space issue. Comprehensive experiments show that TTDeep can schedule TT flows much faster than solver-based methods and schedule nearly twice more TT flows on average compared to handcrafted heuristics. Hongyu Jia, Chunmeng Zhong, Hai Wan, Xibin Zhao |
GLOBECOM | 5 |
| 2021 | A Boolean Network Tomography based Method for Deterministic Multi-point Fault DetectionabstractTime-sensitive networking (TSN) is a deterministic network, where failures will seriously affect the quality of the real-time data transmission service. Hence, fault localization is critical in TSN. In the application domain of TSN, an ideal link failure detection method needs to have properties including execution time deterministic, low detection traffic overhead, multi-point fault detection capability for arbitrary typologies. However, current state-of-the-art fault detection algorithms can not achieve these properties simultaneously. In this paper, we propose a Boolean network tomography-based multi-point fault detection method that leverages the deterministic transmission mechanism of TSN to guarantee that the detection time is upper bounded. The method includes an offline preparation phase and an online operation phase. In the preparation phase, we use a detection flow generation algorithm to get a set of detection flows that can identify up to$K$faults with as few paths as possible. In the operation phase, detection packets are realized as time-sensitive flows routing along the detection paths. They are sent periodically to the controller which performs a topology discovery algorithm to infer the faulty links based on the arrival status of each detection packet. Comprehensive experiments are conducted and the results show that compared with existing fault detection methods, the proposed algorithm can accurately identify multiple faulty links in deterministic time, and the generated detection paths set needs fewer paths than the prior methods. Sukun Zhang, Hai Wan, Xibin Zhao |
GLOBECOM | 3 |
| 2021 | DRLS: A Deep Reinforcement Learning Based Scheduler for Time-Triggered EthernetabstractTime-triggered (TT) communication has long been studied in various industrial domains. The most challenging task of TT communication is to find a feasible schedule table. Network changes are inevitable due to the topology dynamics, varying data transmission requirements, etc. Once changes occur, the schedule table needs to be re-calculated in a timely manner. Solver-based methods and heuristic-based methods were proposed to solve this problem. However, solver-based methods employ integer linear programming (ILP) or satisfiability modulo theories (SMT) which have high computational complexity. On the other hand, heuristic-based methods are fast, but they need to be handcrafted based on the application characteristics. Thus, these methods are not general enough to work in complex scenarios especially in large networks.In this paper we propose DRLS – Deep Reinforcement Learning based TT Scheduling method. DRLS first trains an application or network specific scheduling agent offline. Then, the agent can be used for online scheduling of TT flows. However, off-the-shelf reinforcement learning techniques cannot handle the TT scheduling problem with typical complexity and scale. DRLS provides novel solutions to this challenge, including three key innovations: new representations for TT network adapted to various topologies, proper deep neural network (DNN) structures to capture network characteristics, and scalable reinforcement learning (RL) models to handle online TT scheduling. Comprehensive experiments have been conducted to compare the performance of DRLS and other methods (heuristics-based methods such as HLS, LS, HLD + LD, LS + LD, and ILP-based method). The results show that DRLS can not only adapt to specific network topologies, but also have better performance: runs much faster than ILP solver-based methods, and schedules about 23.9% more flows than traditional handcrafted heuristic-based methods. Chunmeng Zhong, Hongyu Jia, Hai Wan, Xibin Zhao |
ICCCN | 4 |
| 2021 | CPS-Based Self-Adaptive Collaborative Control for Smart Production-Logistics SystemsabstractDiscrete manufacturing systems are characterized by dynamics and uncertainty of operations and behavior due to exceptions in production-logistics synchronization. To deal with this problem, a self-adaptive collaborative control (SCC) mode is proposed for smart production-logistics systems to enhance the capability of intelligence, flexibility, and resilience. By leveraging cyber-physical systems (CPSs) and industrial Internet of Things (IIoT), real-time status data are collected and processed to perform decision making and optimization. Hybrid automata is used to model the dynamic behavior of physical manufacturing resources, such as machines and vehicles in shop floors. Three levels of collaborative control granularity, including nodal SCC, local SCC, and global SCC, are introduced to address different degrees of exceptions. Collaborative optimization problems are solved using analytical target cascading (ATC). A proof of concept simulation based on a Chinese aero-engine manufacturer validates the applicability and efficiency of the proposed method, showing reductions in waiting time, makespan, and energy consumption with reasonable computational time. This article potentially enables manufacturers to implement CPS and IIoT in manufacturing environments and build up smart, flexible, and resilient production-logistics systems. Zhengang Guo, Yingfeng Zhang, Xibin Zhao |
IEEE Trans. Cybern. | 3 |
| 2021 | Flow Scheduling for Conflict-Free Network Updates in Time-Sensitive Software-Defined NetworksabstractThe digital transformation of industry requires industrial control networks provide high flexibility and determinacy. Time-sensitive software-defined networking that combines time-sensitive networking and software-defined networking is a new network paradigm which provides both real-time transmission feature and network flexibility. During network updates, the transmission consistency needs to be maintained. However, previous mechanisms mostly target on the proper schedule transition, which cannot guarantee no frame loss and also introduces extra update overhead. The article proposes a novel flow schedule generation model which guarantees no frame loss during network updates even with the basic two-phase update mechanism and introduces no extra update overhead. Two algorithms are designed for the model to adapt to different application scenarios: the offline algorithm poses better schedulability, whereas the online one consumes less time with slightly decreased schedulability. The experiments on two real-world industrial networks demonstrate our mechanism achieves zero frame loss without extra update overhead compared to existing methods, and the online algorithm saves 40% execution time with at most 10% schedulability decrease when the bandwidth utilization is less than 50%. Zaiyu Pang, Zonghui Li, Sukun Zhang, Yanfen Xu, Hai Wan, Xibin Zhao |
IEEE Trans. Ind. Informatics | 7 |
| 2020 | Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringabstractUnderstanding questions and finding clues for answers are the key for video question answering. Compared with image question answering, video question answering (Video QA) requires to find the clues accurately on both spatial and temporal dimension simultaneously, and thus is more challenging. However, the relationship between spatio-temporal information and question still has not been well utilized in most existing methods for Video QA. To tackle this problem, we propose a Question-Guided Spatio-Temporal Contextual Attention Network (QueST) method. In QueST, we divide the semantic features generated from question into two separate parts: the spatial part and the temporal part, respectively guiding the process of constructing the contextual attention on spatial and temporal dimension. Under the guidance of the corresponding contextual attention, visual features can be better exploited on both spatial and temporal dimensions. To evaluate the effectiveness of the proposed method, experiments are conducted on TGIF-QA dataset, MSRVTT-QA dataset and MSVD-QA dataset. Experimental results and comparisons with the state-of-the-art methods have shown that our method can achieve superior performance. Jianwen Jiang, Haojie Lin, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2020 | Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionabstractThis paper presents an end-to-end 3D convolutional network named attention-based multi-modal fusion network (AMFNet) for the semantic scene completion (SSC) task of inferring the occupancy and semantic labels of a volumetric 3D scene from single-view RGB-D images. Compared with previous methods which use only the semantic features extracted from RGB-D images, the proposed AMFNet learns to perform effective 3D scene completion and semantic segmentation simultaneously via leveraging the experience of inferring 2D semantic segmentation from RGB-D images as well as the reliable depth cues in spatial dimension. It is achieved by employing a multi-modal fusion architecture boosted from 2D semantic segmentation and a 3D semantic completion network empowered by residual attention blocks. We validate our method on both the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset and the results show that our method respectively achieves the gains of 2.5% and 2.6% on the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset against the state-of-the-art method. Siqi Li 0001, Changqing Zou, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2020 | Hypergraph Label Propagation NetworkabstractIn recent years, with the explosion of information on the Internet, there has been a large amount of data produced, and analyzing these data is useful and has been widely employed in real world applications. Since data labeling is costly, lots of research has focused on how to efficiently label data through semi-supervised learning. Among the methods, graph and hypergraph based label propagation algorithms have been a widely used method. However, traditional hypergraph learning methods may suffer from their high computational cost. In this paper, we propose a Hypergraph Label Propagation Network (HLPN) which combines hypergraph-based label propagation and deep neural networks in order to optimize the feature embedding for optimal hypergraph learning through an end-to-end architecture. The proposed method is more effective and also efficient for data labeling compared with traditional hypergraph learning methods. We verify the effectiveness of our proposed HLPN method on a real-world microblog dataset gathered from Sina Weibo. Experiments demonstrate that the proposed method can significantly outperform the state-of-the-art methods and alternative approaches. Yubo Zhang 0006, Nan Wang 0015, Changqing Zou, Hai Wan, Xibin Zhao, Yue Gao 0002 |
AAAI | 6 |
| 2020 | Speeding up Very Fast Decision Tree with Low Computational CostabstractVery Fast Decision Tree (VFDT) is one of the most widely used online decision tree induction algorithms, and it provides high classification accuracy with theoretical guarantees. In VFDT, the split-attempt operation is essential for leaf-split. It is computation-intensive since it computes the heuristic measure of all attributes of a leaf. To reduce split-attempts, VFDT tries to split at constant intervals (for example, every 200 examples). However, this mechanism introduces split-delay for split can only happen at fixed intervals, which slows down the growth of VFDT and finally lowers accuracy. To address this problem, we first devise an online incremental algorithm that computes the heuristic measure of an attribute with a much lower computational cost. Then a subset of attributes is carefully selected to find a potential split timing using this algorithm. A split-attempt will be carried out once the timing is verified. By the whole process, computational cost and split-delay are lowered significantly. Comprehensive experiments are conducted using multiple synthetic and real datasets. Compared with state-of-the-art algorithms, our method reduces split-attempts by about 5 to 10 times on average with much lower split-delay, which makes our algorithm run faster and more accurate. Hongyu Jia, Hai Wan, Xibin Zhao |
IJCAI | 7 |
| 2020 | Dual Channel Hypergraph Collaborative FilteringabstractCollaborative filtering (CF) is one of the most popular and important recommendation methodologies in the heart of numerous recommender systems today. Although widely adopted, existing CF-based methods, ranging from matrix factorization to the emerging graph-based methods, suffer inferior performance especially when the data for training are very limited. In this paper, we first pinpoint the root causes of such deficiency and observe two main disadvantages that stem from the inherent designs of existing CF-based methods, i.e., 1) inflexible modeling of users and items and 2) insufficient modeling of high-order correlations among the subjects. Under such circumstances, we propose a dual channel hypergraph collaborative filtering (DHCF) framework to tackle the above issues. First, a dual channel learning strategy, which holistically leverages the divide-and-conquer strategy, is introduced to learn the representation of users and items so that these two types of data can be elegantly interconnected while still maintaining their specific properties. Second, the hypergraph structure is employed for modeling users and items with explicit hybrid high-order correlations. The jump hypergraph convolution (JHConv) method is proposed to support the explicit and efficient embedding propagation of high-order correlations. Comprehensive experiments on two public benchmarks and two new real-world datasets demonstrate that DHCF can achieve significant and consistent improvements against other state-of-the-art methods. Shuyi Ji, Yifan Feng 0001, Rongrong Ji, Xibin Zhao, Wanwan Tang, Yue Gao 0002 |
KDD | 4 |
| 2020 | IExpressNet: Facial Expression Recognition with Incremental ClassesabstractExisting methods on facial expression recognition (FER) are mainly trained in the setting when all expression classes are fixed in advance. However, in real applications, expression classes are becoming increasingly fine-grained and incremental. To deal with sequential expression classes, we can fine-tune or re-train these models, but this often results in poor performance or large computing resources consumption. To address these problems, we develop an Incremental Facial Expression Recognition Network (IExpressNet), which can learn a competitive multi-class classifier at any time with a lower requirement of computing resources. Specifically, IExpressNet consists of two novel components. First, we construct an exemplar set by dynamically selecting representative samples from old expression classes. Then, the exemplar set and new expression classes samples constitute the training set. Second, we design a novel center-expression-distilled loss. As for facial expression in the wild, center-expression-distilled loss enhances the discriminative power of the deeply learned features and prevents catastrophic forgetting. Extensive experiments are conducted on two large-scale FER datasets in the wild, RAF-DB and AffectNet. The results demonstrate the superiority of the proposed method as compared to state-of-the-art incremental learning approaches. Bingjun Luo, Sicheng Zhao, Shihui Ying, Xibin Zhao, Yue Gao 0002 |
ACM Multimedia | 5 |
| 2020 | Time-Triggered Switch-Memory-Switch Architecture for Time-Sensitive Networking SwitchesabstractTime-sensitive networking (TSN) is a set of extended standards for the IEEE 802.3 Ethernet under development by the IEEE 802.1 TSN task group. TSN depends on two key components, scheduling and fault tolerance, to provide realtime and reliable transmission. There is a strong motivation to replace the widely used field-buses with TSNs in industrial networking applications. However, industrial network devices are typical application-specific embedded systems with limited memory resources. Time-sensitive (TS) transmission certainly prefers on-chip memory, which is even more scarce for embedded systems. As a result, it is critical for TSNs to develop memory-efficient switching techniques with scalable schedulability and elegant fault-tolerance support. This paper proposes a time-triggered switch-memory-switch (SMS) architecture for memory-efficient TSN switches. First, based on the SMS shared memory, our architecture makes it possible to statically schedule memory allocation with full utilization for TS traffic and the remaining memory for other traffic. Compared with perport memory, the shared memory achieves a ratio of (nn/n!) (≈ (en/√(2πn)), n → ∞), where n is the port number, in the feasible solution space under memory constraints and thus significantly improves scheduling memory ability and flexibility. Moreover, we develop a fault-tolerance scheme for reliable transmission. It facilitates a memory-efficient implementation of the popular multiline redundancy in industrial networks. The scheme is validated by five classes of memory conflicts and a case study on two-line redundancy. Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Model-Based Adaptation of Mixed-Criticality Multiservice Systems for Extreme Physical EnvironmentsabstractAn increasingly important trend in the design of industry-strength embedded systems is the integration of multiple services with varying criticality levels into a common computing platform. Such systems are characterized as mixed-criticality multiservice systems (MCMSs). An MCMS has to survive in rigorous environments posed by industry-level requirements. Such survival, however, is becoming continuously more challenging due to the growing system complexity and integrating more and more services. While existing works typically target reliability-driven design optimization to improve the system robustness rather than deal with the surviving problem of the system in extreme physical environments, this paper addresses the problem by enabling the service capability transitions of an MCMS to adapt to the environments. This paper proposes a service capability model to capture the importance of functional modules for the criticality of different services. A model-based service-capability transition mechanism is designed to automatically identify the maximum allowed service capability under a given physical environment. A case study of the proposed techniques was performed on an industrial Ethernet switch which is a typical MCMS, to validate the capability of adaptation to high and low temperatures. The experimental results demonstrate the significant potential of our approach to improve system survivability under extreme physical environments. Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Online Scheduling for Dynamic VM Migration in Multicast Time-Sensitive NetworksabstractWith the development of hardware virtualization and cloud computing, modern industry has a tendency to upgrade from the traditional industrial networks to virtual machine (VM) based networks. To provide firm latency guarantees for control messages in these networks, the time-sensitive network (TSN) is a promising technology due to its determinacy for real-time applications. However, TSN faces the challenge of providing a rapid response to dynamic transmission requirement changes incurred by VM migrations. In this paper, we proposed an online scheduling approach to deal with dynamic VM migrations in multicast TSN. In this approach, we devise a novel online scheduling framework [minimal distance tree (MDT) construction - heuristic breadth first search] containing an offline scheduling phase and an online rescheduling phase. While the offline phase introduces a MDT to increase reusable scheduling results, the online phase proposes a heuristic scheduling approach to reuse the results of the offline phase as much as possible to accelerate the rescheduling process. Experiments show that our framework can provide a rapid response to dynamic VM migrations compared with the existing approaches where the amount of control data does not exceed 50% of the bandwidth. Qinghan Yu, Hai Wan, Xibin Zhao, Yue Gao 0002, Ming Gu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Implementation of Differential Tag Sampling for COTS RFID SystemsabstractTag inventory is one of the most fundamental tasks for RFID systems. However, the Framed Slotted Aloha (FSA) protocol specified in the C1G2 standard is of low time-efficiency, because it needs to collect all tags in the system. To improve time-efficiency, research communities proposed a batch of sampling-based approaches, in which the reader only needs to collect a small set of sampled tags instead of all. Although time-efficiency has been improved, existing sampling-based approaches still have two common limitations. First, all tags in the system are assumed to have the same sampling probability. It is unfair that tags attached to differential items (e.g., different values) have the same chance to be sampled and collected. Second, all existing sampling-based approaches stay in theory level and cannot be deployed on Commercial Off-The-Shelf (COTS) RFID devices, because the C1G2 standard does not support the sampling function at all. To deal with the above two limitations, this paper studies the new problem of differential tag sampling-letting each RFID tag be identified with a given sampling probability. In this paper, we use the COTS RFID devices including Impinj Speedway R420 reader and Monza 4QT tags to implement the Differential Tag Sampling (DTS) operation. Then, we apply probabilistic analytics on the collected tag data to address some practically important problems such as Multi-category Tag Cardinality Estimation (MTCE), and Value-based Missing Tag Detection (VMTD). Although the analytics results are not 100 percent accurate, the deviation in the results can be controlled below a small threshold and DTS can significantly improve the time-efficiency. DTS can be easily deployed on the COTS RFID systems, because it is totally compliant with the C1G2 standard. Extensive experiments demonstrate that DTS is able to let each tag take the given sampling probability to be sampled and identified. Moreover, the proposed DTS protocol can significantly reduce the execution time of MTCE and VMTD by nearly 70 percent than the FSA protocol. Xin Xie 0001, Xiulong Liu 0001, Xibin Zhao, Weilian Xue, Bin Xiao 0001, Heng Qi, Keqiu Li, Jie Wu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2019 | MeshNet: Mesh Neural Network for 3D Shape RepresentationabstractMesh is an important and powerful type of data for 3D shapes and widely studied in the field of computer vision and computer graphics. Regarding the task of 3D shape representation, there have been extensive research efforts concentrating on how to represent 3D shapes well using volumetric grid, multi-view and point cloud. However, there is little effort on using mesh data in recent years, due to the complexity and irregularity of mesh data. In this paper, we propose a mesh neural network, named MeshNet, to learn 3D shape representation from mesh data. In this method, face-unit and feature splitting are introduced, and a general architecture with available and effective blocks are proposed. In this way, MeshNet is able to solve the complexity and irregularity problem of mesh and conduct 3D shape representation well. We have applied the proposed MeshNet method in the applications of 3D shape classification and retrieval. Experimental results and comparisons with the state-of-the-art methods demonstrate that the proposed MeshNet can achieve satisfying 3D shape classification and retrieval performance, which indicates the effectiveness of the proposed method on 3D shape representation. Yutong Feng, Yifan Feng 0001, Haoxuan You, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2019 | DeepCCFV: Camera Constraint-Free Multi-View Convolutional Neural Network for 3D Object Retrievalabstract3D object retrieval has a compelling demand in the field of computer vision with the rapid development of 3D vision technology and increasing applications of 3D objects. 3D objects can be described in different ways such as voxel, point cloud, and multi-view. Among them, multi-view based approaches proposed in recent years show promising results. Most of them require a fixed predefined camera position setting which provides a complete and uniform sampling of views for objects in the training stage. However, this causes heavy over-fitting problems which make the models failed to generalize well in free camera setting applications, particularly when insufficient views are provided. Experiments show the performance drastically drops when the number of views reduces, hindering these methods from practical applications. In this paper, we investigate the over-fitting issue and remove the constraint of the camera setting. First, two basic feature augmentation strategies Dropout and Dropview are introduced to solve the over-fitting issue, and a more precise and more efficient method named DropMax is proposed after analyzing the drawback of the basic ones. Then, by reducing the over-fitting issue, a camera constraint-free multi-view convolutional neural network named DeepCCFV is constructed. Extensive experiments on both single-modal and cross-modal cases demonstrate the effectiveness of the proposed method in free camera settings comparing with existing state-of-theart 3D object retrieval methods. Zhengyue Huang, Zhehui Zhao, Hengguang Zhou, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2019 | MLVCNN: Multi-Loop-View Convolutional Neural Network for 3D Shape Retrievalabstract3D shape retrieval has attracted much attention and become a hot topic in computer vision field recently.With the development of deep learning, 3D shape retrieval has also made great progress and many view-based methods have been introduced in recent years. However, how to represent 3D shapes better is still a challenging problem. At the same time, the intrinsic hierarchical associations among views still have not been well utilized. In order to tackle these problems, in this paper, we propose a multi-loop-view convolutional neural network (MLVCNN) framework for 3D shape retrieval. In this method, multiple groups of views are extracted from different loop directions first. Given these multiple loop views, the proposed MLVCNN framework introduces a hierarchical view-loop-shape architecture, i.e., the view level, the loop level, and the shape level, to conduct 3D shape representation from different scales. In the view-level, a convolutional neural network is first trained to extract view features. Then, the proposed Loop Normalization and LSTM are utilized for each loop of view to generate the loop-level features, which considering the intrinsic associations of the different views in the same loop. Finally, all the loop-level descriptors are combined into a shape-level descriptor for 3D shape representation, which is used for 3D shape retrieval. Our proposed method has been evaluated on the public 3D shape benchmark, i.e., ModelNet40. Experiments and comparisons with the state-of-the-art methods show that the proposed MLVCNN method can achieve significant performance improvement on 3D shape retrieval tasks. Our MLVCNN outperforms the state-of-the-art methods by the mAP of 4.84% in 3D shape retrieval task. We have also evaluated the performance of the proposed method on the 3D shape classification task where MLVCNN also achieves superior performance compared with recent methods. Jianwen Jiang, Di Bao, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2019 | PVRNet: Point-View Relation Neural Network for 3D Shape RecognitionabstractThree-dimensional (3D) shape recognition has drawn much research attention in the field of computer vision. The advances of deep learning encourage various deep models for 3D feature representation. For point cloud and multi-view data, two popular 3D data modalities, different models are proposed with remarkable performance. However the relation between point cloud and views has been rarely investigated. In this paper, we introduce Point-View Relation Network (PVRNet), an effective network designed to well fuse the view features and the point cloud feature with a proposed relation score module. More specifically, based on the relation score module, the point-single-view fusion feature is first extracted by fusing the point cloud feature and each single view feature with point-singe-view relation, then the pointmulti- view fusion feature is extracted by fusing the point cloud feature and the features of different number of views with point-multi-view relation. Finally, the point-single-view fusion feature and point-multi-view fusion feature are further combined together to achieve a unified representation for a 3D shape. Our proposed PVRNet has been evaluated on ModelNet40 dataset for 3D shape classification and retrieval. Experimental results indicate our model can achieve significant performance improvement compared with the state-of-the-art models. Haoxuan You, Yifan Feng 0001, Xibin Zhao, Changqing Zou, Rongrong Ji, Yue Gao 0002 |
AAAI | 3 |
| 2019 | Emotion Recognition by Edge-Weighted Hypergraph Neural NetworkabstractOver the past decade, increasing research efforts have been concentrated on emotion recognition from physiological signals due to their capability on emotion information representation. Existing works mainly focus on exploring the relationship between stimulus and subjects, while ignoring the effects of latent correlations among different subjects, which are important for personalized emotion recognition. To tackle this issue, we aim to conduct emotion recognition using multi-modal physiological signals through an edge-weighted hyper-graph neural network, in which complex relationship among subjects is formulated using hypergraph for each modality respectively. In our model, the differences in significance of influence which various samples leave on the classification can be better represented. The major contribution of this network lies in its concern that the associate strengths between various samples are different, which have different impact on the training result. The hyperedge between the vertices with closer correlation should be assigned a larger weight. Reversely, the looser relation, the minor weight. To evaluate the proposed method, experiments have been conducted on the DEAP dataset and ASCERTAIN dataset. Experimental results and comparison with state-of-the-art methods show that the proposed method can achieve better performance. Jingzhi Shao, Yuxuan Wei, Yifan Feng 0001, Xibin Zhao |
ICIP | 5 |
| 2019 | Emotion Recognition from Physiological Signals using Multi-Hypergraph Neural NetworksabstractEmotion recognition from physiological signals is an effective way to discern the inner state of users. Existing works are lack in the exploration of latent correlation among multiple physiological signals and relationship among different subjects. To tackle this issue, we propose to recognize emotion from physiological signals using multi-hypergraph neural networks (MHGNN). In this method, the correlation among different subjects is formulated in the multi-hypergraph structure, where each type of physiological signal is used to generate one hypergraph. In each hypergraph, the hyperedges are used to represent the connections among the vertices (subject, stimuli). Thus, the emotion recognition task is modeled as classifying each vertex in the multi-hypergraph. Experimental results and comparisons with the state-of-the-art methods in the DEAP dataset demonstrate the superior performance of our method. The comparative experiments based on available biological knowledge verify that MHGNN can depict the real biological response process in a much more precise way. Xibin Zhao, Han Hu 0003, Yue Gao 0002 |
ICME | 2 |
| 2019 | Adaptive Scheduling for Multicluster Time-Triggered Train Communication NetworksabstractThe execution time of conventional incremental off-line schedule approaches for time-triggered networks increases rapidly when networks become larger. When the traffic in a network changes, they need to reschedule all influenced flows once again. Traffic changes at the cluster level involve many data flows. An incremental scheduler cannot react quickly to such changes. We propose an algorithm based on mixed integer linear programming and counterexample guided methodology. Our algorithm can generate adaptive schedule for cluster-level changes of the system. The adaptive schedule can react quickly to the changes during runtime. Our algorithm enhances the incremental schedulers. It allows schedulers to react to changing at both the flow level and the cluster level. Experiments show that our approach is effective. In the scenarios of coupling train consists, our algorithm can generate the schedule table of the train network within a few seconds. Ningchen Wang, Qinghan Yu, Hai Wan, Xibin Zhao |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Correntropy-Induced Robust Low-Rank HypergraphabstractHypergraph learning has been widely exploited in various image processing applications, due to its advantages in modeling the high-order information. Its efficacy highly depends on building an informative hypergraph structure to accurately and robustly formulate the underlying data correlation. However, the existing hypergraph learning methods are sensitive to non- Gaussian noise, which hurts the corresponding performance. In this paper, we present a noise-resistant hypergraph learning model, which provides superior robustness against various non- Gaussian noises. In particular, our model adopts low-rank representation to construct a hypergraph, which captures the globally linear data structure as well as preserving the grouping effect of highly-correlated data. We further introduce a correntropyinduced local metric to measure the reconstruction errors, which is particularly robust to non-Gaussian noises. Finally, the Frobenious-norm based regularization is proposed to combine with the low-rank regularizer, which enables our model to regularize the singular values of the coefficient matrix. By such, the non-zero coefficients are selected to generate a hyperedge set as well as the hyperedge weights. We have evaluated the proposed hypergraph model in the tasks of image clustering and semi-supervised image classification. Quantitatively, our scheme significantly enhances the performance of the state-of-the-art hypergraph models on several benchmark datasets. Taisong Jin, Rongrong Ji, Yue Gao 0002, Xiaoshuai Sun, Xibin Zhao, Dacheng Tao |
IEEE Trans. Image Process. | 5 |
| 2019 | Hypergraph-Induced Convolutional Networks for Visual ClassificationabstractAt present, convolutional neural networks (CNNs) have become popular in visual classification tasks because of their superior performance. However, CNN-based methods do not consider the correlation of visual data to be classified. Recently, graph convolutional networks (GCNs) have mitigated this problem by modeling the pairwise relationship in visual data. Real-world tasks of visual classification typically must address numerous complex relationships in the data, which are not fit for the modeling of the graph structure using GCNs. Therefore, it is vital to explore the underlying correlation of visual data. Regarding this issue, we propose a framework called the hypergraph-induced convolutional network to explore the high-order correlation in visual data during deep neural networks. First, a hypergraph structure is constructed to formulate the relationship in visual data. Then, the high-order correlation is optimized by a learning process based on the constructed hypergraph. The classification tasks are performed by considering the high-order correlation in the data. Thus, the convolution of the hypergraph-induced convolutional network is based on the corresponding high-order relationship, and the optimization on the network uses each data and considers the high-order correlation of the data. To evaluate the proposed hypergraph-induced convolutional network framework, we have conducted experiments on three visual data sets: the National Taiwan University 3-D model data set, Princeton Shape Benchmark, and multiview RGB-depth object data set. The experimental results and comparison in all data sets demonstrate the effectiveness of our proposed hypergraph-induced convolutional network compared with the state-of-the-art methods. Heyuan Shi, Yubo Zhang 0006, Zizhao Zhang 0003, Nan Ma 0012, Xibin Zhao, Yue Gao 0002, Jia-Guang Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | An Enhanced Reconfiguration for Deterministic Transmission in Time-Triggered NetworksabstractThe emerging momentum of digital transformation of industry, i.e. Industry 4.0, poses strong demands for integrating industrial control networks, and Ethernet to enable the real-time Internet of Things (RT-IoT). Time-triggered (TT) networks provide a cost-efficient integrated solution while RT-IoT arouses the reconfiguration challenges: the network has to be flexible enough to adapt to changes and yet provides deterministic transmission persistently during network reconfiguration. Software defined network benefits the flexible industrial control by configuring the rules handling frames. However, previous reconfiguration mechanisms are mostly oriented to the context of data centers and wide area networks and thus do not consider the deterministic transmission in TT networks. This paper focuses on the reconfiguration (i.e., updates) for the deterministic transmission. To minimize the overhead during updates, namely the minimum number of loss frames and the minimum duration time of updates, we first establish an update theory based on the dependence relationship derived by the conflicts during updates. In addition then the reconfiguration problem is modeled with the dependence graph built by the relationship. On such a basis, we present a reconfiguration mechanism and its implementation to solve the problem. Finally, we evaluate the proposed reconfiguration mechanism in two real industrial network topologies. The experimental results demonstrate that compared with previous methods, our mechanism significantly reduces the number of loss frames and achieves zero loss in almost all cases. Zonghui Li, Hai Wan, Zaiyu Pang, Qiubo Chen, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2019 | Fast RFID Sensory Data Collection: Trade-off Between Computation and Communication CostsabstractThis paper studies the important sensory data collection problem in the sensor-augmented RFID systems, which is to quickly and accurately collect sensory data from a predefined set of target tags with the coexistence of unexpected tags. The existing RFID data collection schemes suffer from either low time-efficiency due to tag-collisions or serious data corruption issue due to interference of unexpected tags. To overcome these limitations, we propose the hierarchical-hashing data collection (HDC) protocol, which can not only significantly improve the utilization of RFID wireless communication channel by establishing bijective mapping between k target tags and the first k slots in time frame, but also effectively filter out the serious interference of unexpected tags. Although HDC has attractive advantages, the theoretical analysis reveals that the computation cost involved in it is as huge as O(k2k), where k is normally large in practice. By making some modifications to the basic HDC protocol, we propose the multi-framed hierarchical-hashing data collection (MHDC) protocol to effectively reduce the involved computation complexity. Unlike HDC that only issues a single time frame, MHDC uses multiple time frames to collaboratively collect sensory data from the k target tags. It can be understood as that a big computation task is disintegrated into multiple small pieces and then shared by multiple time frames. As a result, the computation cost involved in MHDC is reduced to O(k2n), where n ≪ k is the expected number of target tags that each time frame handles. Theoretical analysis is given to jointly consider the communication cost and computation cost thereby maximizing the overall time-efficiency of MHDC. Extensive simulation results reveal that the proposed MHDC protocol can correctly collect all sensory data and is always about more than 2× faster than the state-of-the-art RFID sensory data collection protocols. Xiulong Liu 0001, Jiannong Cao 0001, Yanni Yang 0003, Wenyu Qu, Xibin Zhao, Keqiu Li, Didi Yao |
IEEE/ACM Trans. Netw. | 5 |
| 2018 | Energy-Efficient Automatic Train Driving by Learning Driving PatternsabstractRailway is regarded as the most sustainable means of modern transportation. With the fast-growing of fleet size and the railway mileage, the energy consumption of trains is becoming a serious concern globally. The nature of railway offers a unique opportunity to optimize the energy efficiency of locomotives by taking advantage of the undulating terrains along a route. The derivation of an energy-optimal train driving solution, however, proves to be a significant challenge due to the high dimension, nonlinearity, complex constraints, and time-varying characteristic of the problem. An optimized solution can only be attained by considering both the complex environmental conditions of a given route and the inherent characteristics of a locomotive. To tackle the problem, this paper employs a high-order correlation learning method for online generation of the energy optimized train driving solutions. Based on the driving data of experienced human drivers, a hypergraph model is used to learn the optimal embedding from the specified features for the decision of a driving operation. First, we design a feature set capturing the driving status. Next all the training data are formulated as a hypergraph and an inductive learning process is conducted to obtain the embedding matrix. The hypergraph model can be used for real-time generation of driving operation. We also proposed a reinforcement updating scheme, which offers the capability of sustainable enhancement on the hypergraph model in industrial applications. The learned model can be used to determine an optimized driving operation in real-time tested on the Hardware-in-Loop platform. Validation experiments proved that the energy consumption of the proposed solution is around 10% lower than that of average human drivers. Jin Huang 0002, Yue Gao 0002, Xibin Zhao, Yangdong Deng, Ming Gu 0001 |
AAAI | 4 |
| 2018 | EMD Metric LearningabstractEarth Mover's Distance (EMD), targeting at measuring the many-to-many distances, has shown its superiority and been widely applied in computer vision tasks, such as object recognition, hyperspectral image classification and gesture recognition. However, there is still little effort concentrated on optimizing the EMD metric towards better matching performance. To tackle this issue, we propose an EMD metric learning algorithm in this paper. In our method, the objective is to learn a discriminative distance metric for EMD ground distance matrix generation which can better measure the similarity between compared subjects. More specifically, given a group of labeled data from different categories, we first select a subset of training data and then optimize the metric for ground distance matrix generation. Here, both the EMD metric and the EMD flow-network are alternatively optimized until a steady EMD value can be achieved. This method is able to generate a discriminative ground distance matrix which can further improve the EMD distance measurement. We then apply our EMD metric learning method on two tasks, i.e., multi-view object classification and document classification. The experimental results have shown better performance of our proposed EMD metric learning method compared with the traditional EMD method and the state-of-the-art methods. It is noted that the proposed EMD metric learning method can be also used in other applications. Zizhao Zhang 0003, Yubo Zhang 0006, Xibin Zhao, Yue Gao 0002 |
AAAI | 3 |
| 2018 | Hypergraph Learning With Cost Interval OptimizationabstractIn many classification tasks, the misclassification costs of different categories usually vary significantly. Under such circumstances, it is essential to identify the importance of different categories and thus assign different misclassification losses in many applications, such as medical diagnosis, saliency detection and software defect prediction. However, we note that it is infeasible to determine the accurate cost value without great domain knowledge. In most common cases, we may just have the information that which category is more important than the other categories, i.e., the identification of defect-prone softwares is more important than that of defect-free. To tackle these issues, in this paper, we propose a hypergraph learning method with cost interval optimization, which is able to handle cost interval when data is formulated using the high-order relationships. In this way, data correlations are modeled by a hypergraph structure, which has the merit to exploit the underlying relationships behind the data. With a cost-sensitive hypergraph structure, in order to improve the performance of the classifier without precise cost value, we further introduce cost interval optimization to hypergraph learning. In this process, the optimization on cost interval achieves better performance instead of choosing uncertain fixed cost in the learning process. To evaluate the effectiveness of the proposed method, we have conducted experiments on two groups of dataset, i.e., the NASA Metrics Data Program (NASA) dataset and UCI Machine Learning Repository (UCI) dataset. Experimental results and comparisons with state-of-the-art methods have exhibited better performance of our proposed method. Xibin Zhao, Nan Wang 0015, Heyuan Shi, Hai Wan, Jin Huang 0002, Yue Gao 0002 |
AAAI | 1 |
| 2018 | GVCNN: Group-View Convolutional Neural Networks for 3D Shape Recognitionabstract3D shape recognition has attracted much attention recently. Its recent advances advocate the usage of deep features and achieve the state-of-the-art performance. However, existing deep features for 3D shape recognition are restricted to a view-to-shape setting, which learns the shape descriptor from the view-level feature directly. Despite the exciting progress on view-based 3D shape description, the intrinsic hierarchical correlation and discriminability among views have not been well exploited, which is important for 3D shape representation. To tackle this issue, in this paper, we propose a group-view convolutional neural network (GVCNN) framework for hierarchical correlation modeling towards discriminative 3D shape description. The proposed GVCNN framework is composed of a hierarchical view-group-shape architecture, i.e., from the view level, the group level and the shape level, which are organized using a grouping strategy. Concretely, we first use an expanded CNN to extract a view level descriptor. Then, a grouping module is introduced to estimate the content discrimination of each view, based on which all views can be splitted into different groups according to their discriminative level. A group level description can be further generated by pooling from view descriptors. Finally, all group level descriptors are combined into the shape level descriptor according to their discriminative weights. Experimental results and comparison with state-of-the-art methods show that our proposed GVCNN method can achieve a significant performance gain on both the 3D shape classification and retrieval tasks. Yifan Feng 0001, Zizhao Zhang 0003, Xibin Zhao, Rongrong Ji, Yue Gao 0002 |
CVPR | 3 |
| 2018 | Iterative Metric Learning for Imbalance Data ClassificationabstractIn many classification applications, the amount of data from different categories usually vary significantly, such as software defect predication and medical diagnosis. Under such circumstances, it is essential to propose a proper method to solve the imbalance issue among the data. However, most of the existing methods mainly focus on improving the performance of classifiers rather than searching for an appropriate way to find an effective data space for classification. In this paper, we propose a method named Iterative Metric Learning (IML) to explore the correlations among imbalance data and construct an effective data space for classification. Given the imbalance training data, it is important to select a subset of training samples for each testing data. Thus, we aim to find a more stable neighborhood for testing data using the iterative metric learning strategy. To evaluate the effectiveness of the proposed method, we have conducted experiments on two groups of dataset, i.e., the NASA Metrics Data Program (NASA) dataset and UCI Machine Learning Repository (UCI) dataset. Experimental results and comparisons with state-of-the-art methods have exhibited better performance of our proposed method. Nan Wang 0015, Xibin Zhao, Yue Gao 0002 |
IJCAI | 2 |
| 2018 | Multi-model induced network for participatory-sensing-based classification tasks in intelligent and connected transportation systems
Heyuan Shi, Xibin Zhao, Hai Wan, Huihui Wang 0001, Jian Dong 0001, Anfeng Liu |
Comput. Networks | 2 |
| 2018 | IPAD: Intensity Potential for Adaptive De-QuantizationabstractDisplay devices at bit depth of 10 or higher have been mature but the mainstream media source is still at bit depth of eight. To accommodate the gap, the most economic solution is to render source at low bit depth for high bit-depth display, which is essentially the procedure of de-quantization. Traditional methods, such as zero-padding or bit replication, introduce annoying false contour artifacts. To better estimate the least-significant bits, later works use filtering or interpolation approaches, which exploit only limited neighbor information, cannot thoroughly remove the false contours. In this paper, we propose a novel intensity potential (IP) field to model the complicated relationships among pixels. The potential value decreases as the spatial distance to the field source increases and the potentials from different field sources are additive. Based on the proposed IP field, an adaptive de-quantization procedure is then proposed to convert low-bit-depth images to high-bit-depth ones. To the best of our knowledge, this is the first attempt to apply potential field for natural images. The proposed potential field preserves local consistency and models the complicated contexts well. Extensive experiments on natural, synthetic, and high-dynamic range image data sets validate the efficiency of the proposed IP field. Significant improvements have been achieved over the state-of-the-art methods on both the peak signal-to-noise ratio and the structural similarity. Jing Liu 0002, Guangtao Zhai, Anan Liu, Xiaokang Yang 0001, Xibin Zhao, Chang Wen Chen |
IEEE Trans. Image Process. | 5 |
| 2018 | Inductive Multi-Hypergraph Learning and Its Application on View-Based 3D Object ClassificationabstractThe wide 3D applications have led to increasing amount of 3D object data, and thus effective 3D object classification technique has become an urgent requirement. One important and challenging task for 3D object classification is how to formulate the 3D data correlation and exploit it. Most of the previous works focus on learning optimal pairwise distance metric for object comparison, which may lose the global correlation among 3D objects. Recently, a transductive hypergraph learning has been investigated for classification, which can jointly explore the correlation among multiple objects, including both the labeled and unlabeled data. Although these methods have shown better performance, they are still limited due to 1) a considerable amount of testing data may not be available in practice and 2) the high computational cost to test new coming data. To handle this problem, considering the multi-modal representations of 3D objects in practice, we propose an inductive multi-hypergraph learning algorithm, which targets on learning an optimal projection for the multi-modal training data. In this method, all the training data are formulated in multi-hypergraph based on the features, and the inductive learning is conducted to learn the projection matrices and the optimal multi-hypergraph combination weights simultaneously. Different from the transductive learning on hypergraph, the high cost training process is off-line, and the testing process is very efficient for the inductive learning on hypergraph. We have conducted experiments on two 3D benchmarks, i.e., the NTU and the ModelNet40 data sets, and compared the proposed algorithm with the state-of-the-art methods and traditional transductive multi-hypergraph learning methods. Experimental results have demonstrated that the proposed method can achieve effective and efficient classification performance. We also note that the proposed method is a general framework and has the potential to be applied in other applications in practice. Zizhao Zhang 0003, Haojie Lin, Xibin Zhao, Rongrong Ji, Yue Gao 0002 |
IEEE Trans. Image Process. | 3 |
| 2018 | Fast Identification of Blocked RFID TagsabstractThe widely used RFID systems are vulnerable to the denial-of-service (DoS) attacks launched by malicious blocker tags. This paper studies how to quickly and completely identify the valid RFID tags that are blocked. The existing work that can seemingly address this problem suffers from either low time-efficiency or serious false positives. This paper proposes a hybrid approach that consists of two complementary component protocols, namelyAloha Filtering(AF) andPoll&Listen(PL).AFis fast but inaccurate, whilePLis accurate but slow. Taking the merit of each protocol, our hybrid approach is to first repeat the fastAFfor multiple rounds to quickly filter out the target tags that are definitely not blocked. Then, on the size-reduced remaining set that just contains a small number of suspicious tags, we invoke the accuratePLto verify the intactness of each suspicious tag with 100 percent confidence. We optimize the round count ofAFthat trades off between the time costs ofAFandPLto minimize the total time ofAF+PL. As required in the optimization process, we need to know the size of the blocked tag set and that of the unknown tag set, which, however, are not known in advance. To estimate these two set sizes, we propose a supplementary protocol calledSimultaneous Estimation of the Blocked tag size and the Unknown tag size(SEBU). The key advantages of our approach over the prior art are four-fold. First, unlike the detection protocol that just discovers the existence of blocking attacks, our approach exactly identifies all the blocked target tags. Second, our approach is compliant with the C1G2 standard, and does not require any modifications to be made to the commercial RFID tags. It only needs to be installed on readers as a software module. Third, our approach does not involve any false positives. Finally, our approach significantly reduces the execution time when compared with the state-of-the-art schemes that can completely identify the blocked tags. Xiulong Liu 0001, Xin Xie 0001, Xibin Zhao, Kun Wang 0005, Keqiu Li, Alex X. Liu, Song Guo 0001, Jie Wu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2018 | Beyond Pairwise Matching: Person Reidentification via High-Order Relevance LearningabstractPerson reidentification has attracted extensive research efforts in recent years. It is challenging due to the varied visual appearance from illumination, view angle, background, and possible occlusions, leading to the difficulties when measuring the relevance, i.e., similarities, between probe and gallery images. Existing methods mainly focus on pairwise distance metric learning for person reidentification. In practice, pairwise image matching may limit the data for comparison (just the probe and one gallery subject) and yet lead to suboptimal results. The correlation among gallery data can be also helpful for the person reidentification task. In this paper, we propose to investigate the high-order correlation among the probe and gallery data, not the pairwise matching, to jointly learn the relevance of gallery data to the probe. Recalling recent progresses on feature representation in person reidentification, it is difficult to select the best feature and each type of feature can benefit person description from different aspects. Under such circumstances, we propose a multihypergraph joint learning algorithm to learn the relevance in corporation with multiple features of the imaging data. More specifically, one hypergraph is constructed using one type of feature and multiple hypergraphs can be generated accordingly. Then, the learning process is conducted on the multihypergraph structure, and the identity of a probe is determined by its relevance to each gallery data. The merit of the proposed scheme is twofold. First, different from pairwise image matching, the proposed method jointly explores the relationships among different images. Second, multimodal data, i.e., different features, can be formulated in the multihypergraph structure, which can convey more information in the learning process and can be easily extended. We note that the proposed method is a general framework to incorporate with any combination of features, and thus is flexible in practice. Experimental results and comparisons with the state-of-the-art methods on three public benchmarking data sets demonstrate the superiority of the proposed method. Xibin Zhao, Nan Wang 0015, Yubo Zhang 0006, Shaoyi Du, Yue Gao 0002, Jia-Guang Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Human experience knowledge induction based intelligent train drivingabstractAs the most sustainable means of modern transportation, the railway trains are eagerly approaching autonomous driving due to their congenital advantages on operating environments compare to, e.g., road traffics. The intelligent automatic train driving aims at train control with a goal of energy efficiency, punctuality and safety. The derivation of an optimized train driving solution by taking advantage of the undulating terrains along a route, however, proves to be a significant challenge due to the high dimension, nonlinearity, complex constraints, and time-varying characteristic of the problem. To tackle the problem, we propose a two-level human driving experience learning framework and employ the fuzzy rule induction method for online generation of the optimized driving solutions. Based on the records of experienced human drivers, a FURIA model was built to learn the driving rules indicating the correlation between the specified features to the decision of a driving sequence. The fuzzy rules can generally find the best-match driving operation under certain running circumstances. The learned model can be used to determine an optimized driving operation in real-time. Validation experiments show that the energy consumption of the proposed solution is around 8.93% lower than that of average human drivers. Jin Huang 0002, Yangdong Deng, Xibin Zhao, Ming Gu 0001 |
ICIS | 4 |
| 2017 | Infer Precise Program Invariant Using Abstract Interpretation with Recurrence SolvingabstractProgram invariant is formal description of properties that should hold at certain program location in every valid execution. It is very useful for program analysis and verification. In this paper, we introduce an abstraction interpretation approach for generating program invariant efficiently and precisely. A polynomial interval domain is proposed for representing abstract state and precise loop effect is summarized by recognizing and solving recurrence relations. Our method has implemented and its effectiveness is shown in various kinds of cases. Experiment results show that our approach generates more accurate program invariants quickly. Zhenpeng Fang, Xibin Zhao |
COMPSAC (1) | 2 |
| 2017 | Vertex-Weighted Hypergraph Learning for Multi-View Object Classificationabstract3D object classification with multi-view representation has become very popular, thanks to the progress on computer techniques and graphic hardware, and attracted much research attention in recent years. Regarding this task, there are mainly two challenging issues, i.e., the complex correlation among multiple views and the possible imbalance data issue. In this work, we propose to employ the hypergraph structure to formulate the relationship among 3D objects, taking the advantage of hypergraph on high-order correlation modelling. However, traditional hypergraph learning method may suffer from the imbalance data issue. To this end, we propose a vertex-weighted hypergraph learning algorithm for multi-view 3D object classification, introducing an updated hypergraph structure. In our method, the correlation among different objects is formulated in a hypergraph structure and each object (vertex) is associated with a corresponding weight, weighting the importance of each sample in the learning process. The learning process is conducted on the vertex-weighted hypergraph and the estimated object relevance is employed for object classification. The proposed method has been evaluated on two public benchmarks, i.e., the NTU and the PSB datasets. Experimental results and comparison with the state-of-the-art methods and recent deep learning method demonstrate the effectiveness of our proposed method. Lifan Su, Yue Gao 0002, Xibin Zhao, Hai Wan, Ming Gu 0001, Jia-Guang Sun 0001 |
IJCAI | 3 |
| 2017 | Handling scheduling uncertainties through traffic shaping in Time-Triggered train networksabstractWhile trains traditionally relied on field bus to support real-time control applications, next-generation trains are moving toward Ethernet as an integrated, high-bandwidth communication infrastructure for real-time control and best-effort consumer traffic. Time-Triggered Ethernet (TT-Ethernet) is a promising technology for train networks because of its capability to achieve deterministic latencies for real-time applications based on pre-computed transmission schedules. However, the deterministic scheduling approach of TT-Ethernet faces significant challenges in handling scheduling uncertainties caused by switch failures and legacy end devices in train networks. Due to the physical constraints on trains, train networks deal with switch failures by bypassing failed switches using a short circuiting mechanism. Unfortunately, this mechanism incurs scheduling errors as frames bypassing the failed switch may arrive ahead of the pre-computed schedule, resulting in early, unexpected, and out of order arrivals. Furthermore, as trains evolve from traditional communication technologies to TT-Ethernet, the network must support legacy end devices that may generate frames at times unknown to the TT-Ethernet. We propose a novel traffic shaping approach to deal with scheduling uncertainties in TT-Ethernet. The traffic shaper of a TT-Ethernet switch buffers early frames and then releases them at their pre-scheduled arrive time. Furthermore, we devise an efficient buffer management method for the traffic shaper in face of fault scenarios. Finally, we use the traffic shaper to integrate legacy devices into TT-Ethernet. We have implemented the traffic shaping approach in a 24-port TT-Ethernet switch specifically designed for train networks. Experiments show the traffic shaping strategy can effectively deal with scheduling uncertainties incurred by switch failures and legacy devices. Qinghan Yu, Xibin Zhao, Hai Wan, Yue Gao 0002, Chenyang Lu 0001, Ming Gu 0001 |
IWQoS | 2 |
| 2017 | Representative band selection for hyperspectral image classification
Ronglu Yang, Lifan Su, Xibin Zhao, Hai Wan, Jia-Guang Sun 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Event Classification in Microblogs via Social TrackingabstractSocial media websites have become important information sharing platforms. The rapid development of social media platforms has led to increasingly large-scale social media data, which has shown remarkable societal and marketing values. There are needs to extract important events in live social media streams. However, microblogs event classification is challenging due to two facts, i.e., the short/conversational nature and the incompatible meanings between the text and the corresponding image in social posts, and the rapidly evolving contents. In this article, we propose to conduct event classification via deep learning and social tracking. First, we introduce a Multi-modal Multi-instance Deep Network (M 2 DN) for microblogs classification, which is able to handle the weakly labeled microblogs data oriented from the incompatible meanings inside microblogs. Besides predicting each microblogs as predefined events, we propose to employ social tracking to extract social-related auxiliary information to enrich the testing samples. We extract a set of candidate-relevant microblogs in a short time window by using social connections, such as related users and geographical locations. All these selected microblogs and the testing data are formulated in a Markov Random Field model. The inference on the Markov Random Field is conducted to update the classification results of the testing microblogs. This method is evaluated on the Brand-Social-Net dataset for classification of 20 events. Experimental results and comparison with the state of the arts show that the proposed method can achieve better performance for the event classification task. Yue Gao 0002, Hanwang Zhang, Xibin Zhao, Shuicheng Yan |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2014 | Optimal robust control for generalized fuzzy dynamical systems: A novel use on fuzzy uncertaintiesabstractA novel approach for optimal robust control of a class of generalized fuzzy dynamical systems is proposed. This is a novel use of fuzzy uncertainty in doing dynamical system control. The system may have nonlinear nominal terms and the other terms with uncertainty, including unknown parameters and input disturbances. The Fuzzy sets theory is creatively employed in presenting the system parameter and input uncertainty, and then the control structure is deterministic (versus if-then rule-based as is typical in Mamdani-type fuzzy control). The desired controlled system performance is also deterministic, with guaranteed performances of uniform boundedness and uniform ultimate boundedness. Fuzzy informations on the uncertainties are used in searching optimal control gain under a proposed LQG-like quadratic cost index. The control gain design problem is formulated as a constrained optimization problem with the solution be proved to be always existed and unique. Systematic procedure is summarized for such control design. Jin Huang 0002, Jia-Guang Sun 0001, Xibin Zhao, Ming Gu 0001 |
CICA | 3 |
| 2014 | Towards Accurate Object Localization with SmartphonesabstractIn this study, we explore the possibility of locating remote objects via cameras together with built-in inertial sensors of off-the-shelf smartphones. Our solution, CamLoc, enables a user taking two photos of an object using a smartphone at a fixed location and immediately knowing the location of the object in global coordinates, thus facilitating myriad location-based services. Such usage is user-friendly but error prone. We devise several techniques to mitigate the errors caused by cheap and noisy sensors, upgrading the positioning accuracy to an applicable level. We prototype CamLoc on Android OS, and evaluate its performance across different scenarios with various building densities. Experiment results show that our system achieves 89 percent and 72 percent physical location mapping accuracy in rural and downtown areas, respectively, which is competitive with existing solutions. Longfei Shangguan, Zimu Zhou, Zheng Yang 0002, Kebin Liu 0001, Zhenjiang Li 0001, Xibin Zhao, Yunhao Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2012 | Integrated Importance Measure of Component States Based on Loss of System PerformanceabstractThis paper mainly focuses on the integrated importance measure (IIM) of component states based on loss of system performance. To describe the impact of each component state, we first introduce the performance function of the multi-state system. Then, we present the definition of IIM of component states. We demonstrate its corresponding physical meaning, and then analyze the relationships between IIM and Griffith importance, Wu importance, and Natvig importance. Secondly, we present the evaluation method of IIM for multi-state systems. Thirdly, the characteristics of IIM of component states are discussed. Finally, we demonstrate a numerical example, and an application to an offshore oil and gas production system for IIM to verify the proposed method. The results show that 1) the IIM of component states concerns not only the probability distributions and transition intensities of the states of the object component, but also the change in the system performance under the change of the state distribution of the object component; and 2) IIM can be used to identify the key state of a component that affects the system performance most. Shubin Si, Hongyan Dui, Xibin Zhao, Shenggui Zhang, Shudong Sun |
IEEE Trans. Reliab. | 3 |
| 2011 | Self-diagnosis for large scale wireless sensor networksabstractExisting approaches to diagnosing sensor networks are generally sink-based, which rely on actively pulling state information from all sensor nodes so as to conduct centralized analysis. However, the sink-based diagnosis tools incur huge communication overhead to the traffic sensitive sensor networks. Also, due to the unreliable wireless communications, sink often obtains incomplete and sometimes suspicious information, leading to highly inaccurate judgments. Even worse, we observe that it is always more difficult to obtain state information from the problematic or critical regions. To address the above issues, we present the concept of self-diagnosis, which encourages each single sensor to join the fault decision process. We design a series of novel fault detectors through which multiple nodes can cooperate with each other in a diagnosis task. The fault detectors encode the diagnosis process to state transitions. Each sensor can participate in the fault diagnosis by transiting the detector's current state to a new one based on local evidences and then pass the fault detector to other nodes. Having sufficient evidences, the fault detector achieves the Accept state and outputs the final diagnosis report. We examine the performance of our self-diagnosis tool called TinyD2 on a 100 nodes testbed. Kebin Liu 0001, Qiang Ma 0007, Xibin Zhao, Yunhao Liu 0001 |
INFOCOM | 3 |
| 2006 | Minimal Threshold Closure
Xibin Zhao, Kwok-Yan Lam, Guiming Luo, Siu Leung Chung, Ming Gu 0001 |
ESORICS | 1 |
| 2005 | Secure Anonymous Communication with Conditional Traceability
Zhaofeng Ma, Xibin Zhao, Zhi Guo, Ming Gu 0001, Jia-Guang Sun 0001 |
NPC | 2 |
| 2004 | Authorization Mechanisms for Virtual Organizations in Distributed Computing Systems
Xibin Zhao, Kwok-Yan Lam, Siu Leung Chung, Ming Gu 0001, Jia-Guang Sun 0001 |
ACISP | 1 |