EDBT 2026 Demo / reviewers in the wild / expert
Siqi Wang 0001
dblp:145/2904-1
· DBLP profile ↗
45ranked-venue papers
12as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attention to Threat-Relevant Objects: Reasoning Detection in Autonomous Driving via Multimodal Large Language ModelsabstractPerceiving threats is an innate human instinct. During driving, humans naturally focus their attention on objects that pose real potential risks. Motivated by this observation, we shift the focus from traditional class-based detection to a novel task termed threat-oriented reasoning detection in autonomous driving. This task aims to localize threat objects and reason about their threat levels from a driver-centric perspective. To support this task, we build a benchmark comprising diverse corner-case scenarios, annotated by multiple experienced drivers to reflect human-aligned threat cognition. Given the reasoning demands of this task, we then explore the capabilities of multi-modal large language models (MLLMs) and introduce two methods based on whether the MLLM supports object detection: 1) For MLLMs lacking detection capability, we introduce ThreatCoT, a plug-and-play training-free method that combines chain-of-thought (CoT) with a visual expert toolchain to support step-by-step reasoning. 2) For MLLMs with detection support, we introduce ThreatReasoner, an end-to-end reinforcement learning (RL)-based method built on the GRPO algorithm, which enables per-object reasoning through a fully unsupervised reward strategy. Both quantitative and qualitative experiments show that our methods can effectively unlock the new capabilities of MLLM in threat-oriented reasoning detection. Yu-Lin He, Wei Chen 0009, Xinbiao Gan, Siqi Wang 0001, Haotian Wang 0001, Yusong Tan |
AAAI | 4 |
| 2026 | SCAD: A self-constrained solution to automate context-guided zero-shot image anomaly detection
Siqi Wang 0001, Guangpu Wang, Xinwang Liu 0002, Jie Liu 0002, Jiyuan Liu 0003, Siwei Wang 0001 |
Neural Networks | 1 |
| 2026 | Communication-Efficient Federated Multi-View ClusteringabstractFederated multi-view clustering is an emerging machine learning paradigm that groups the data with each view distributed on an isolated client while preserving their privacies. Although recent researches have proposed a few feasible solutions, they are severely limited by two drawbacks. In specific, the clients are required to share their data representations at each iteration of model training, leading to heavy communication overhead. On the other hand, existing researches handle large-scale data by employing the matrix factorization and neural network encoding techniques, failing to utilize their similarity information sufficiently. To address these issues, we propose a communication-efficient federated multi-view clustering framework by approximating the data representation with pseudo-label and centroid matrix, where the latter two are shared in model training. Meanwhile, the framework is instanced by incorporating linear kernel function to consider the data pairwise similarities. Note that, corresponding linear kernels are not required to compute explicitly, making the resultant method able to be optimized in linear complexity to the number of samples. Nevertheless, the proposed method is evaluated on benchmark datasets. It not only achieves inspiring results (26.84% accuracy improvement on average, 2.9$_\times$×-2153$_\times$× computation speedup and 98.4% communication overhead reduction at most) compared with existing federated multi-view clustering methods, but also outperforms centralized multi-view clustering approaches on performance and computation efficiency. Jiyuan Liu 0003, Xinwang Liu 0002, Siqi Wang 0001, Xinhang Wan, Dongsheng Li 0001, Kai Lu 0001, Kunlun He |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Effective video anomaly detection by step-constrained diffusion model
Junhua Xi, Siqi Wang 0001, Zhiping Cai, Kefeng Deng, Kaijun Ren |
Pattern Recognit. | 3 |
| 2026 | "Stones From Other Hills Can Polish Jade": Zero-Shot Anomaly Synthesis via Cross-Domain Anomaly InjectionabstractIndustrial image anomaly detection (IAD) is a pivotal topic with huge value. Due to the nature of anomalies, real anomalies in a specific modern industrial domain (i.e., domain-specific anomalies) are usually too rare to collect, which severely hinders IAD. Thus, zero-shot anomaly synthesis (ZSAS), which synthesizes pseudo anomaly images without any domain-specific anomaly, emerges as a vital technique for IAD. However, existing solutions are either unable to synthesize authentic pseudo anomalies, or require cumbersome training. Thus, we focus on ZSAS and propose a brand-new paradigm that can realize both authentic and training-free ZSAS. It is based on a chronically-ignored fact: Although domain-specific anomalies are rare, real anomalies from other domains (i.e., cross-domain anomalies) are actually abundant and directly applicable to ZSAS. Specifically, our new ZSAS paradigm makes three-fold contributions: First, we propose a novel method named Cross-domain Anomaly Injection (CAI), which directly exploits cross-domain anomalies to enable highly authentic ZSAS in a training-free manner. Second, to supply CAI with sufficient cross-domain anomalies, we build the first Domain-agnostic Anomaly Dataset (DAAD) within our best knowledge, which provides ZSAS with abundant real anomaly patterns. Third, we propose a CAI-guided Diffusion Mechanism, which can further break the quantity limit of real anomalies and enable unlimited anomaly synthesis. Our head-to-head comparison with existing ZSAS solutions justifies the superior performance of our paradigm for IAD and demonstrates it as an effective and pragmatic ZSAS solution. Siqi Wang 0001, Yuanze Hu, Xinwang Liu 0002, Siwei Wang 0001, Guangpu Wang, Chuanfu Xu, Jie Liu 0002, Ping Chen 0004 |
IEEE Trans. Image Process. | 1 |
| 2026 | Towards Mitigation of False Negatives in Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) aims to retrieve semantically related images from a large gallery given a text query. Most existing TIReID methods train the model on an ideal assumption that positive (negative) samples are semantically correlated (uncorrelated) on the visual and textual modal. However, we observe that incorrect annotation is unavoidable and ambiguous textual description is often taken as the input in practice, which leads to false negatives that ruin the feature aligning between modals for deviated identification. This work presents a novel False Negative Mitigation (FNM) method by identifying potential false negatives through distribution differences and mitigating them with dedicated loss. Specifically, the false negative mitigation loss is designed to adaptively adjust the optimization margin for potential false negatives in the latent space. Moreover, a Dual-Level Feature Representation method is introduced to leverage both global and local features during the mitigation, whereas a Momentum Contractive (MoC) module is plugged to enrich the training data for accurate similarity distribution estimation. We conduct extensive experiments on three public benchmark datasets and the results demonstrate that our FNM method achieves State-of-The-Art (SoTA) performance at all evaluation metrics. Ruigeng Zeng, Wentao Ma 0003, Tongqing Zhou, Siqi Wang 0001, Xinjun Mao, Jie Liu 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic DisentanglementabstractOccupancy prediction plays a pivotal role in autonomous driving (AD) due to its capabilities of fine-grained 3D perception and general object recognition. However, existing methods often incur high computational costs, which conflict with AD's real-time demand. To this end, we redirect the focus from accuracy only to both accuracy and efficiency. By conducting a head-to-head comparison of existing methods, we find it challenging to balance accuracy and efficiency. We identify a core issue for this challenge: the strong coupling between geometry and semantics. Specifically, the predicted geometric structure (e.g., depth) guides the projection of 2D image features into 3D voxel space, which significantly affects feature discriminability and subsequent semantic learning. To address this issue, we focus on two key aspects: model design and learning strategies. 1) For model design, we propose a dual-branch network that disentangles the representation of geometry and semantics. The voxel branch utilizes a novel re-parameterized large-kernel 3D convolution to refine geometric structure efficiently, while the BEV branch employs temporal fusion and BEV encoding for efficient semantic learning. 2) For learning strategies, we propose to separate geometric learning from semantic learning by the mixup of ground-truth and predicted depths. Our method achieves 39.4% mIoU at 20 FPS on Occ3D-nuScenes, showcasing a state-of-the-art balance between accuracy and efficiency. Yu-Lin He, Wei Chen 0009, Siqi Wang 0001, Tianci Xun, Yusong Tan |
AAAI | 3 |
| 2025 | PG3D-ViT: A Prompt-Guided 3D Vision Transformer for Medical Image Classificationabstract3D medical image classification is challenging due to small, subtle lesions and substantial irrelevant context, which often mislead deep models. Inspired by the top-down diagnos-tic process of clinicians—first identifying anatomical context, then locating anomalies—we propose Prompt-Guided 3D Vision Transformer (PG 3D- ViT), a framework that simulates clinical reasoning through prompt-driven attention. To address limited 3D training data, PG3D- ViT leverages 2D masked auto encoder (MAE) pretraining to learn transferable image features. Through the prompt generation module, consistency difference analysis is performed between normal and abnormal samples to extract anatomical structure and global spatial prompt information related to the lesion context. These prompts are injected as query into a cross-attention mechanism, guiding the model to focus on lesion-relevant regions across the 3D volume. Evaluated on 7 public datasets spanning multiple modalities and pathologies, PG3D-ViT achieves a 1.88% average AUC improvement over state-of-the-art methods. The attention map visualizations demonstrate that the model can accurately localize lesion regions, validating the effectiveness of the clinical prompting mechanism in enhancing both the performance and interpretability of 3D medical image classification. The code is available at the provided link11https://github.comJUMED-P/PG3D-ViT Jue Gong, Ke Zuo, Siqi Wang 0001, Xiaoguang Mao, Jie Liu 0002 |
ICDM | 4 |
| 2025 | UP-Bern: A Unified Progressive Transition Framework for Biomedical Entity Recognition and NormalizationabstractAccurate recognition of biomedical entities (e.g., diseases, drugs, proteins, and genes) from literature, along with their normalization using standardized biomedical vocabularies, is crucial for facilitating various downstream tasks and boosting significant advancements in further biomedical research. However, most current studies still struggle with boundary inconsistencies stemming from separate decoding stages. Additionally, such models often neglect valuable semantic information encapsulated in the vocabulary's candidate concepts (e.g., surface forms of text), which is vital for effective entity normalization. In this paper, we propose a unified progressive transition framework named UP-Bern, which progressively constructs biomedical entity recognition and normalization outputs through an action sequence prediction process. This framework promotes joint modeling through shared input representations and concurrently optimizes outputs within a unified search space via state transitions. Moreover, we integrate an attention mechanism to fully harness the surface form information of candidate concepts, thereby enhancing the accuracy of entity normalization. We evaluated the proposed approach on four public biomedical datasets against 11 established methods, demonstrating consistent and notable improvements in performance. Canqun Yang, Siqi Wang 0001 |
ICDM | 3 |
| 2025 | Recalling Unknowns Without Losing Precision: An Effective Solution to Large Model-Guided Open World Object DetectionabstractOpen World Object Detection (OWOD) aims to adapt object detection to an open-world environment, so as to detect unknown objects and learn knowledge incrementally. Existing OWOD methods typically leverage training sets with a relatively small number of known objects. Due to the absence of generic object knowledge, they fail to comprehensively perceive objects beyond the scope of training sets. Recent advancements in large vision models (LVMs), trained on extensive large-scale data, offer a promising opportunity to harness rich generic knowledge for the fundamental advancement of OWOD. Motivated by Segment Anything Model (SAM), a prominent LVM lauded for its exceptional ability to segment generic objects, we first demonstrate the possibility to employ SAM for OWOD and establish the very first SAM-Guided OWOD baseline solution. Subsequently, we identify and address two fundamental challenges in SAM-Guided OWOD and propose a pioneering SAM-Guided Robust Open-world Detector (SGROD) method, which can significantly improve the recall of unknown objects without losing the precision on known objects. Specifically, the two challenges in SAM-Guided OWOD include: 1) Noisy labels caused by the class-agnostic nature of SAM; 2) Precision degradation on known objects when more unknown objects are recalled. For the first problem, we propose a dynamic label assignment (DLA) method that adaptively selects confident labels from SAM during training, evidently reducing the noise impact. For the second problem, we introduce cross-layer learning (CLL) and SAM-based negative sampling (SNS), which enable SGROD to avoid precision loss by learning robust decision boundaries of objectness and classification. Experiments on public datasets show that SGROD not only improves the recall of unknown objects by a large margin (~20%), but also preserves highly-competitive precision on known objects. The program codes are available at https://github.com/harrylin-hyl/SGROD. Yu-Lin He, Wei Chen 0009, Siqi Wang 0001, Tianrui Liu 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | SAMCL: Subgraph-Aligned Multiview Contrastive Learning for Graph Anomaly DetectionabstractGraph anomaly detection (GAD) has gained increasing attention in various attribute graph applications, i.e., social communication and financial fraud transaction networks. Recently, graph contrastive learning (GCL)-based methods have been widely adopted as the mainstream for GAD with remarkable success. However, existing GCL strategies in GAD mainly focus on node-node and node-subgraph contrast and fail to explore subgraph-subgraph level comparison. Furthermore, the different sizes or component node indices of the sampled subgraph pairs may cause the "nonaligned" issue, making it difficult to accurately measure the similarity of subgraph pairs. In this article, we propose a novel subgraph-aligned multiview contrastive approach for graph anomaly detection, named SAMCL, which fills the subgraph-subgraph contrastive-level blank for GAD tasks. Specifically, we first generate the multiview augmented subgraphs by capturing different neighbors of target nodes forming contrasting subgraph pairs. Then, to fulfill the nonaligned subgraph pair contrast, we propose a subgraph-aligned strategy that estimates similarities with the Earth mover's distance (EMD) of both considering the node embedding distributions and typology awareness. With the newly established similarity measure for subgraphs, we conduct the interview subgraph-aligned contrastive learning module to better detect changes for nodes with different local subgraphs. Moreover, we conduct intraview node-subgraph contrastive learning to supplement richer information on abnormalities. Finally, we also employ the node reconstruction task for the masked subgraph to measure the local change of the target node. Finally, the anomaly score for each node is jointly calculated by these three modules. Extensive experiments conducted on benchmark datasets verify the effectiveness of our approach compared to existing state-of-the-art (SOTA) methods with significant performance gains (up to 6.36% improvement on ACM). Our code can be verified at https://github.com/hujingtao/SAMCL. Jingtao Hu, Bin Xiao 0002, Hu Jin 0005, Jingcan Duan, Siwei Wang 0001, Zhao Lv, Siqi Wang 0001, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Sniffing Threatening Open-World Objects in Autonomous Driving by Open-Vocabulary ModelsabstractAutonomous driving (AD) is a typical application that requires effectively exploiting multimedia information. For AD, it is critical to ensure safety by detecting unknown objects in an open world, driving the demand for open world object detection (OWOD). However, existing OWOD methods treat generic objects beyond known classes in the train set as unknown objects and prioritize recall in evaluation. This encourages excessive false positives and endangers safety of AD. To address this issue, we restrict the definition of unknown objects to threatening objects in AD, and introduce a new evaluation protocol, which is built upon a new metric named U-ARecall, to alleviate biased evaluation caused by neglecting false positives. Under the new evaluation protocol, we re-evaluate existing OWOD methods and discover that they typically perform poorly in AD. Then, we propose a novel OWOD paradigm for AD based on fine-tuning foundational open-vocabulary models (OVMs), as they can exploit rich linguistic and visual prior knowledge for OWOD. Following this new paradigm, we propose a brand-new OWOD solution, which effectively addresses two core challenges of fine-tuning OVMs via two novel techniques: 1) the maintenance of open-world generic knowledge by a dual-branch architecture; 2) the acquisition of scenario-specific knowledge by the visual-oriented contrastive learning scheme. Besides, a dual-branch prediction fusion module is proposed to avoid post-processing and hand-crafted heuristics. Extensive experiments show that our proposed method not only surpasses classic OWOD methods in unknown object detection by a large margin (∼× U-ARecall), but also notably outperforms OVMs without fine-tuning in known object detection (∼ 20% K-mAP). Our codes are available at https://github.com/harrylin-hyl/AD-OWOD. Yu-Lin He, Siqi Wang 0001, Wei Chen 0009, Tianci Xun, Yusong Tan |
ACM Multimedia | 2 |
| 2024 | Multiview Deep Anomaly Detection: A Systematic ExplorationabstractAnomaly detection (AD), which models a given normal class and distinguishes it from the rest of abnormal classes, has been a long-standing topic with ubiquitous applications. As modern scenarios often deal with massive high-dimensional complex data spawned by multiple sources, it is natural to consider AD from the perspective of multiview deep learning. However, it has not been formally discussed by the literature and remains underexplored. Motivated by this blank, this article makes fourfold contributions: First, to the best of our knowledge, this is the first work that formally identifies and formulates the multiview deep AD problem. Second, we take recent advances in relevant areas into account and systematically devise various baseline solutions, which lays the foundation for multiview deep AD research. Third, to remedy the problem that limited benchmark datasets are available for multiview deep AD, we extensively collect the existing public data and process them into more than 30 multiview benchmark datasets via multiple means, so as to provide a better evaluation platform for multiview deep AD. Finally, by comprehensively evaluating the devised solutions on different types of multiview deep AD benchmark datasets, we conduct a thorough analysis on the effectiveness of the designed baselines and hopefully provide other researchers with beneficial guidance and insight into the new multiview deep AD topic. Siqi Wang 0001, Jiyuan Liu 0003, Xinwang Liu 0002, Sihang Zhou 0001, En Zhu, Yuexiang Yang, Jianping Yin, Wenjing Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | DrugProtKGE: Weakly Supervised Knowledge Graph Embedding for Highly-Effective Drug-Protein Interaction RepresentationabstractWith the exponential growth of biomedical knowledge in unstructured text repositories such as PubMed, it is imminent to establish a knowledge graph-style, efficient searchable and targeted database that can support the need of information retrieval from researchers and clinicians. To mine knowledge from graph databases, most previous methods view a triple in a graph (see Fig. 1) as the basic processing unit and embed the triplet element (i.e. drugs/chemicals, proteins/genes and their interaction) as separated embedding matrices, which cannot capture the semantic correlation among triple elements. To remedy the loss of semantic correlation caused by disjoint embeddings, we propose a novel approach to learn triple embeddings by combining entities and interactions into a unified representation. Furthermore, traditional methods usually learn triple embeddings from scratch, which cannot take advantage of the rich domain knowledge embedded in pre-trained models, and is also another significant reason for the fact that they cannot distinguish the differences implied by the same entity in the multi-interaction triples. In this paper, we propose a novel fine-tuning based approach to learn better triple embeddings by creating weakly supervised signals from pre-trained knowledge graph embeddings. The method automatically samples triples from knowledge graphs and estimates their pairwise similarity from pre-trained embedding models. The triples are then fed pairwise into a Siamese-like neural architecture, where the triple representation is fine-tuned in the manner bootstrapped by triple similarity scores. Finally, we demonstrate that triple embeddings learned with our method can be readily applied to several downstream applications (e.g. triple classification and triple clustering). We evaluated the proposed method on two open-source drug-protein knowledge graphs constructed from PubMed abstracts, as provided by BioCreative. Our method achieves consistent improvement in both triple classification and triple clustering tasks when compared to other state-of-the-art triple embedding methods, with an average 35% improvement of F1 score for the multi-interaction triples. Siqi Wang 0001, Xi Yang 0020, Xinyuan Qiu, Chengkun Wu, Yingbo Cui 0001, Canqun Yang |
BIBM | 2 |
| 2023 | E$^{3}$3Outlier: a Self-Supervised Framework for Unsupervised Deep Outlier DetectionabstractExisting unsupervised outlier detection (OD) solutions face a grave challenge with surging visual data like images. Although deep neural networks (DNNs) prove successful for visual data, deep OD remains difficult due to OD’s unsupervised nature. This paper proposes a novel framework namedE$^{3}$Outlierthat can performeffective andend-to-end deep outlier removal. Its core idea is to introduceself-supervisioninto deep OD. Specifically, our major solution is to adopt a discriminative learning paradigm that creates multiple pseudo classes from given unlabeled data by various data operations, which enables us to apply prevalent discriminative DNNs (e.g. ResNet) to the unsupervised OD problem. Then, with theoretical and empirical demonstration, we argue that inlier priority, a property that encourages DNN to prioritize inliers during self-supervised learning, makes it possible to perform end-to-end OD. Meanwhile, unlike frequently-used outlierness measures (e.g. density, proximity) in previous OD methods, we explore network uncertainty and validate it as a highly effective outlierness measure, while two practical score refinement strategies are also designed to improve OD performance. Finally, in addition to the discriminative learning paradigm above, we also explore the solutions that exploit other learning paradigms (i.e. generative learning and contrastive learning) to introduce self-supervision forE$^{3}$Outlier. Such extendibility not only brings further performance gain on relatively difficult datasets, but also enablesE$^{3}$Outlierto be applied to other OD applications like video abnormal event detection. Extensive experiments demonstrate thatE$^{3}$Outliercan considerably outperform state-of-the-art counterparts by 10%-30% AUROC. Demo codes are available athttps://github.com/demonzyj56/E3Outlier. Siqi Wang 0001, Yijie Zeng, Zhen Cheng 0004, Xinwang Liu 0002, Sihang Zhou 0001, En Zhu, Marius Kloft, Jianping Yin, Qing Liao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Video Anomaly Detection via Visual Cloze TestsabstractAlthough great progress has been sparked in video anomaly detection (VAD) by deep neural networks (DNNs), existing solutions still fall short in two aspects: (1) The extraction of video events cannot be both precise and comprehensive. (2) The semantics and temporal context are under-explored. To tackle above issues, we are inspired by cloze tests in language education and propose a novel approach namedVisual Cloze Completion(VCC), which conducts VAD by completingvisual cloze tests(VCTs). Specifically, VCC first localizes each video event and encloses it into a spatio-temporal cube (STC). To realize both precise and comprehensive event extraction, appearance and motion are used as complementary cues to mark the object region associated with each event. For each marked region, a normalized patch sequence is extracted from several neighboring frames and stacked into a STC. With each patch and the patch sequence of a STC regarded as a visual “word” and “sentence” respectively, we deliberately erase a certain “word” (patch) to yield a VCT. Then, the VCT is completed by training DNNs to infer the erased patch and its optical flow via video semantics. Meanwhile, VCC fully exploits temporal context by alternatively erasing each patch in temporal context and creating multiple VCTs. Furthermore, we propose localization-level, event-level, model-level and decision-level solutions to enhance VCC, which can further exploit VCC’s potential and produce significant VAD performance improvement. Extensive experiments demonstrate that VCC achieves highly competitive VAD performance. Siqi Wang 0001, Zhiping Cai, Xinwang Liu 0002, En Zhu, Jianping Yin |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Multiple Kernel Clustering With Compressed Subspace AlignmentabstractMultiple kernel clustering (MKC) has recently achieved remarkable progress in fusing multisource information to boost the clustering performance. However, the$\mathcal {O}({n}^{2})$memory consumption and$\mathcal {O}({n}^{3})$computational complexity prohibit these methods from being applied into median- or large-scale applications, where$n$denotes the number of samples. To address these issues, we carefully redesign the formulation of subspace segmentation-based MKC, which reduces the memory and computational complexity to$\mathcal {O}({n})$and$\mathcal {O}({n}^{2})$, respectively. The proposed algorithm adopts a novel sampling strategy to enhance the performance and accelerate the speed of MKC. Specifically, we first mathematically model the sampling process and then learn it simultaneously during the procedure of information fusion. By this way, the generated anchor point set can better serve data reconstruction across different views, leading to improved discriminative capability of the reconstruction matrix and boosted clustering performance. Although the integrated sampling process makes the proposed algorithm less efficient than the linear complexity algorithms, the elaborate formulation makes our algorithm straightforward for parallelization. Through the acceleration of GPU and multicore techniques, our algorithm achieves superior performance against the compared state-of-the-art methods on six datasets with comparable time cost to the linear complexity algorithms. Sihang Zhou 0001, Qiyuan Ou, Xinwang Liu 0002, Siqi Wang 0001, Luyan Liu, Siwei Wang 0001, En Zhu, Jianping Yin, Xin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Deep Anomaly Discovery from Unlabeled Videos via Normality Advantage and Self-Paced RefinementabstractWhile classic video anomaly detection (VAD) requires labeled normal videos for training, emerging unsupervised VAD (UVAD) aims to discover anomalies directly from fully unlabeled videos. However, existing UVAD methods still rely on shallow models to perform detection or initialization, and they are evidently inferior to classic VAD methods. This paper proposes a full deep neural network (DNN) based solution that can realize highly effective UVAD. First, we, for the first time, point out that deep reconstruction can be surprisingly effective for UVAD, which inspires us to unveil a property named “normality advantage”, i.e., normal events will enjoy lower reconstruction loss when DNN learns to reconstruct unlabeled videos. With this property, we propose Localization based Reconstruction (LBR) as a strong UVAD baseline and a solid foundation of our solution. Second, we propose a novel self-paced refinement (SPR) scheme, which is synthesized into LBR to conduct UVAD. Unlike ordinary self-paced learning that injects more samples in an easy-to-hard manner, the proposed SPR scheme gradually drops samples so that suspicious anomalies can be removed from the learning process. In this way, SPR consolidates normality advantage and enables better UVAD in a more proactive way. Finally, we further design a variant solution that explicitly takes the motion cues into account. The solution evidently enhances the UVAD performance, and it sometimes even surpasses the best classic VAD methods. Experiments show that our solution not only significantly outperforms existing UVAD methods by a wide margin (5% to 9% AUROC), but also enables UVAD to catch up with the mainstream performance of classic VAD. Siqi Wang 0001, Zhiping Cai, Xinwang Liu 0002, Chuanfu Xu, Chengkun Wu |
CVPR | 2 |
| 2022 | Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly DetectionabstractAnomaly detection in multivariate time series data is challenging due to complex temporal and feature correlations. This paper proposes a novel unsupervised multi-scale stacked spatial-temporal graph attention network for multivariate time series anomaly detection (STGAT-MAD). The core of our framework is to coherently capture the feature and temporal correlations among multivariate time-series data by stackable STGAT networks. Meanwhile, a multi-scale input network is exploited to capture the temporal correlations in different time-scales. Besides, a new dataset derived from a real-world wind farm is built and released for multivariate time series anomaly detection. Experiments on the proprietary dataset and three public datasets show that our method significantly outperforms existing baseline approaches, and provides interpretability for anomaly location. Siqi Wang 0001, Xiandong Ma, Chengkun Wu, Canqun Yang, Detian Zeng, Shi-Lin Wang |
ICASSP | 2 |
| 2022 | Detecting Anomalous Events from Unlabeled Videos via Temporal Masked Auto-EncodingabstractUnsupervised video anomaly detection (UVAD) intends to discern anomalous events from fully unlabeled videos. However, existing UVAD methods suffer from poor performance. Inspired by recent masked autoencoder (MAE) [1], we propose Temporal Masked Auto-Encoding (TMAE) as an effective end-to-end UVAD method. Specifically, we first denote video events by spatial-temporal cubes (STCs), which are built by temporally consecutive foreground patches from unlabeled videos. Then, half of patches in an STC are masked along the temporal dimension, while a vision transformer (ViT) is trained to exploit unmasked patches to predict masked patches. The rare and unusual nature of anomaly will result in a poorer prediction for anomalous events, which enables us to discriminate anomalies from unlabeled videos and compute the anomaly scores. Furthermore, to utilize motion clues in videos, we also propose to apply TMAE on optical flow, which can further boost performance. Experiments show that TMAE significantly outperforms existing UVAD methods by a notable margin (3.9%–6.6% AUC). Jingtao Hu, Siqi Wang 0001, En Zhu, Zhiping Cai, Xinzhong Zhu |
ICME | 3 |
| 2022 | Effective Video Abnormal Event Detection by Learning A Consistency-Aware High-Level Feature ExtractorabstractWith pure normal training videos, video abnormal event detection (VAD) aims to build a normality model, and then detect abnormal events that deviate from this model. Despite of some progress, existing VAD methods typically train the normality model by a low-level learning objective (e.g. pixel-wise reconstruction/prediction), which often overlooks the high-level semantics in videos. To better exploit high-level semantics for VAD, we propose a novel paradigm that performs VAD by learning a Consistency-Aware high-level Feature Extractor (CAFE). Specifically, with a pre-trained deep neural network (DNN) as teacher network, we first feed raw video events into the teacher network and extract the outputs of multiple hidden layers as their high-level features, which contain rich high-level semantics. Guided by high-level features extracted from normal training videos, we train a student network to be the high-level feature extractor of normal events, so as to explicitly consider high-level semantics in training. For inference, a video event can be viewed as normal if the student extractor produces similar high-level features to the teacher network. Second, based on the fact that consecutive video frames usually enjoy minor differences, we propose a consistency-aware scheme that requires high-level features extracted from neighboring frames to be consistent. Our consistency-aware scheme not only encourages the student extractor to ignore low-level differences and capture more high-level semantics, but also enables better anomaly scoring. Last, we also design a generic framework that can bridge high-level and low-level learning in VAD to further ameliorate VAD performance. By flexibly embedding one or more low-level learning objectives into CAFE, the framework makes it possible to combine the strengths of both high-level and low-level learning. The proposed method attains state-of-the-art results on commonly-used benchmark datasets. Siqi Wang 0001, Zhiping Cai, Xinwang Liu 0002, Chengkun Wu |
ACM Multimedia | 2 |
| 2022 | FlowDNN: a physics-informed deep neural network for fast and accurate flow predictionabstractfor flow-related design optimization problems, e.g., aircraft and automobile aerodynamic design, computational fluid dynamics (CFD) simulations are commonly used to predict flow fields and analyze performance. While important, CFD simulations are a resource-demanding and time-consuming iterative process. The expensive simulation overhead limits the opportunities for large design space exploration and prevents interactive design. In this paper, we propose FlowDNN, a novel deep neural network (DNN) to efficiently learn flow representations from CFD results. FlowDNN saves computational time by directly predicting the expected flow fields based on given flow conditions and geometry shapes. FlowDNN is the first DNN that incorporates the underlying physical conservation laws of fluid dynamics with a carefully designed attention mechanism for steady flow prediction. This approach not only improves the prediction accuracy, but also preserves the physical consistency of the predicted flow fields, which is essential for CFD. Various metrics are derived to evaluate FlowDNN with respect to the whole flow fields or regions of interest (RoIs) (e.g., boundary layers where flow quantities change rapidly). Experiments show that FlowDNN significantly outperforms alternative methods with faster inference and more accurate results. It speeds up a graphics processing unit (GPU) accelerated CFD solver by more than 14 000×, while keeping the prediction error under 5%. Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Siqi Wang 0001, Shizhao Chen, Jianbin Fang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | A Self-Supervised Gait Encoding Approach With Locality-Awareness for 3D Skeleton Based Person Re-IdentificationabstractPerson re-identification (Re-ID) via gait features within 3D skeleton sequences is a newly-emerging topic with several advantages. Existing solutions either rely on hand-crafted descriptors or supervised gait representation learning. This paper proposes a self-supervised gait encoding approach that can leverage unlabeled skeleton data to learn gait representations for person Re-ID. Specifically, we first create self-supervision by learning to reconstruct unlabeled skeleton sequences reversely, which involves richer high-level semantics to obtain better gait representations. Other pretext tasks are also explored to further improve self-supervised learning. Second, inspired by the fact that motion's continuity endows adjacent skeletons in one skeleton sequence and temporally consecutive skeleton sequences with higher correlations (referred as locality in 3D skeleton data), we propose a locality-aware attention mechanism and a locality-aware contrastive learning scheme, which aim to preserve locality-awareness on intra-sequence level and inter-sequence level respectively during self-supervised learning. Last, with context vectors learned by our locality-aware attention mechanism and contrastive learning scheme, a novel feature named Constrastive Attention-based Gait Encodings (CAGEs) is designed to represent gait effectively. Empirical evaluations show that our approach significantly outperforms skeleton-based counterparts by 15-40 percent Rank-1 accuracy, and it even achieves superior performance to numerous multi-modal methods with extra RGB or depth information. Our codes are available at https://github.com/Kali-Hac/Locality-Awareness-SGE. Haocong Rao, Siqi Wang 0001, Xiping Hu, Mingkui Tan, Yi Guo 0007, Jun Cheng 0002, Xinwang Liu 0002, Bin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | MOOD 2020: A Public Benchmark for Out-of-Distribution Detection and Localization on Medical ImagesabstractDetecting Out-of-Distribution (OoD) data is one of the greatest challenges in safe and robust deployment of machine learning algorithms in medicine. When the algorithms encounter cases that deviate from the distribution of the training data, they often produce incorrect and over-confident predictions. OoD detection algorithms aim to catch erroneous predictions in advance by analysing the data distribution and detecting potential instances of failure. Moreover, flagging OoD cases may support human readers in identifying incidental findings. Due to the increased interest in OoD algorithms, benchmarks for different domains have recently been established. In the medical imaging domain, for which reliable predictions are often essential, an open benchmark has been missing. We introduce the Medical-Out-Of-Distribution-Analysis-Challenge (MOOD) as an open, fair, and unbiased benchmark for OoD methods in the medical imaging domain. The analysis of the submitted algorithms shows that performance has a strong positive correlation with the perceived difficulty, and that all algorithms show a high variance for different anomalies, making it yet hard to recommend them for clinical practice. We also see a strong correlation between challenge ranking and performance on a simple toy test set, indicating that this might be a valuable addition as a proxy dataset during anomaly detection algorithm development. David Zimmerer, Peter M. Full, Fabian Isensee, Paul F. Jaeger, Tim Adler, Jens Petersen, Gregor Köhler, Tobias Roß, Annika Reinke, Antanas Kascenas, Bjørn Sand Jensen, Alison O'Neil, Jeremy Tan, Benjamin Hou, James Batten, Huaqi Qiu, Bernhard Kainz, Nina Shvetsova, Irina Fedulova, Dmitry V. Dylov, Baolun Yu, Jianyang Zhai, Jingtao Hu, Runxuan Si, Sihang Zhou 0001, Siqi Wang 0001, Xuerun Chen, Yang Zhao 0003, Sergio Naval Marimont, Giacomo Tarroni, Victor Saase, Lena Maier-Hein, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 26 |
| 2021 | One-pass Multi-view Clustering for Large-scale DataabstractExisting non-negative matrix factorization based multi-view clustering algorithms compute multiple coefficient matrices respect to different data views, and learn a common consensus concurrently. The final partition is always obtained from the consensus with classical clustering techniques, such as k-means. However, the non-negativity constraint prevents from obtaining a more discriminative embedding. Meanwhile, this two-step procedure fails to unify multi-view matrix factorization with partition generation closely, resulting in unpromising performance. Therefore, we propose an one-pass multi-view clustering algorithm by removing the non-negativity constraint and jointly optimize the aforementioned two steps. In this way, the generated partition can guide multi-view matrix factorization to produce more purposive coefficient matrix which, as a feedback, improves the quality of partition. To solve the resultant optimization problem, we design an alternate strategy which is guaranteed to be convergent theoretically. Moreover, the proposed algorithm is free of parameter and of linear complexity, making it practical in applications. In addition, the proposed algorithm is compared with recent advances in literature on benchmarks, demonstrating its effectiveness, superiority and efficiency. Jiyuan Liu 0003, Xinwang Liu 0002, Yuexiang Yang, Li Liu 0002, Siqi Wang 0001, Weixuan Liang, Jiangyong Shi |
ICCV | 5 |
| 2021 | Self-Representation Subspace Clustering for Incomplete Multi-view DataabstractIncomplete multi-view clustering is an important research topic in multimedia where partial data entries of one or more views are missing. Current subspace clustering approaches mostly employ matrix factorization on the observed feature matrices to address this issue. Meanwhile, self-representation technique is left unexplored, since it explicitly relies on full data entries to construct the coefficient matrix, which is contradictory to the incomplete data setting. However, it is widely observed that self-representation subspace method enjoys a better clustering performance over the factorization based one. Therefore, we adapt it to incomplete data by jointly performing data imputation and self-representation learning. To the best of our knowledge, this is the first attempt in incomplete multi-view clustering literature. Besides, the proposed method is carefully compared with current advances in experiment with respect to different missing ratios, verifying its effectiveness. Jiyuan Liu 0003, Xinwang Liu 0002, Yi Zhang 0104, Pei Zhang 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Weixuan Liang, Siqi Wang 0001, Yuexiang Yang |
ACM Multimedia | 9 |
| 2021 | Improved autoencoder for unsupervised anomaly detectionabstractDeep autoencoder-based methods are the majority of deep anomaly detection. An autoencoder learning on training data is assumed to produce higher reconstruction error for the anomalous samples than the normal samples and thus can distinguish anomalies from normal data. However, this assumption does not always hold in practice, especially in unsupervised anomaly detection, where the training data is anomaly contaminated. We observe that the autoencoder generalizes so well on the training data that it can reconstruct both the normal data and the anomalous data well, leading to poor anomaly detection performance. Besides, we find that anomaly detection performance is not stable when using reconstruction error as anomaly score, which is unacceptable in the unsupervised scenario. Because there are no labels to guide on selecting a proper model. To mitigate these drawbacks for autoencoder-based anomaly detection methods, we propose an Improved AutoEncoder for unsupervised Anomaly Detection (IAEAD). Specifically, we manipulate feature space to make normal data points closer using anomaly detection-based loss as guidance. Different from previous methods, by integrating the anomaly detection-based loss and autoencoder's reconstruction loss, IAEAD can jointly optimize for anomaly detection tasks and learn representations that preserve the local data structure to avoid feature distortion. Experiments on five image data sets empirically validate the effectiveness and stability of our method. Zhen Cheng 0004, Siwei Wang 0001, Pei Zhang 0008, Siqi Wang 0001, Xinwang Liu 0002, En Zhu |
Int. J. Intell. Syst. | 4 |
| 2021 | Facial Expression Recognition Using Frequency Neural NetworkabstractFacial expression recognition has become a newly-emerging topic in recent decades, which has important value in the field of human-computer interaction. In this paper, we present a deep learning based approach, named frequency neural network (FreNet), for facial expression recognition. Different from convolutional neural network in spatial domain, FreNet inherits the advantages of processing image in frequency domain, such as efficient computation and spatial redundancy elimination. First, we propose the learnable multiplication kernel and construct multiple multiplication layers to learn features in frequency domain. Second, a summarization layer is proposed following multiplication layers to further yield high-level features. Third, based on the property of discrete cosine transform (DCT), we utilize multiplication layers and summarization layer to construct the Basic-FreNet, which can yield high-level features on the widely used DCT feature. Finally, to further achieve better performance on Basic-FreNet, we propose the Block-FreNet in which the weight-shared multiplication kernel is designed for feature learning and the block sub-sampling is designed for dimension reduction. The experimental results show that the Block-FreNet not only achieves superior performance, but also greatly reduces the computational cost. To our best knowledge, the proposed approach is the first attempt to fill in the blank of frequency based deep learning model for facial expression recognition. Xingming Zhang 0001, Xiping Hu, Siqi Wang 0001, Haoxiang Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2020 | Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-IdentificationabstractGait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supervised learning paradigms. Unlike previous methods, we for the first time propose a generic gait encoding approach that can utilize unlabeled skeleton data to learn gait representations in a self-supervised manner. Specifically, we first propose to introduce self-supervision by learning to reconstruct input skeleton sequences in reverse order, which facilitates learning richer high-level semantics and better gait representations. Second, inspired by the fact that motion's continuity endows temporally adjacent skeletons with higher correlations (“locality”), we propose a locality-aware attention mechanism that encourages learning larger attention weights for temporally adjacent skeletons when reconstructing current skeleton, so as to learn locality when encoding gait. Finally, we propose Attention-based Gait Encodings (AGEs), which are built using context vectors learned by locality-aware attention, as final gait representations. AGEs are directly utilized to realize effective person Re-ID. Our approach typically improves existing skeleton-based methods by 10-20% Rank-1 accuracy, and it achieves comparable or even superior performance to multi-modal methods with extra RGB or depth information. Haocong Rao, Siqi Wang 0001, Xiping Hu, Mingkui Tan, Huang Da, Jun Cheng 0002, Bin Hu 0001 |
IJCAI | 2 |
| 2020 | Cloze Test Helps: Effective Video Anomaly Detection via Learning to Complete Video EventsabstractAs a vital topic in media content interpretation, video anomaly detection (VAD) has made fruitful progress via deep neural network (DNN). However, existing methods usually follow a reconstruction or frame prediction routine. They suffer from two gaps: (1) They cannot localize video activities in a both precise and comprehensive manner. (2) They lack sufficient abilities to utilize high-level semantics and temporal context information. Inspired by frequently-used cloze test in language study, we propose a brand-new VAD solution named Video Event Completion (VEC) to bridge gaps above: First, we propose a novel pipeline to achieve both precise and comprehensive enclosure of video activities. Appearance and motion are exploited as mutually complimentary cues to localize regions of interest (RoIs). A normalized spatio-temporal cube (STC) is built from each RoI as a video event, which lays the foundation of VEC and serves as a basic processing unit. Second, we encourage DNN to capture high-level semantics by solving a visual cloze test. To build such a visual cloze test, a certain patch of STC is erased to yield an incomplete event (IE). The DNN learns to restore the original video event from the IE by inferring the missing patch. Third, to incorporate richer motion dynamics, another DNN is trained to infer erased patches' optical flow. Finally, two ensemble strategies using different types of IE and modalities are proposed to boost VAD performance, so as to fully exploit the temporal context and modality information for VAD. VEC can consistently outperform state-of-the-art methods by a notable margin (typically 1.5%-5% AUROC) on commonly-used VAD benchmarks. Our codes and results can be verified at github.com/yuguangnudt/VEC_VAD Siqi Wang 0001, Zhiping Cai, En Zhu, Chuanfu Xu, Jianping Yin, Marius Kloft |
ACM Multimedia | 2 |
| 2020 | Unsupervised Online Anomaly Detection With Parameter Adaptation for KPI Abrupt ChangesabstractIT companies need to monitor various Key Performance Indicators (KPIs) and detect anomalies in real time to ensure the quality and reliability of Internet-based services. However, due to the diversity of KPIs, the ambiguity and scarcity of anomalies and the lack of labels, anomaly detection for various KPIs has been a great challenge. Existing KPI anomaly detection methods have not explored the properties of anomalies in KPIs in detail to our best knowledge. Therefore, we explore anomalies in KPIs and recognize a common and important form of anomalies namedabrupt changes, which often indicate potential failures in the relevant services. Forabrupt changesin various KPIs, we proposeDDCOL, an unsupervised online anomaly detection algorithm with parameter adaptation from the perspective of anomalies for the first time. We propose three techniques: high order${D}$ifference extraction and combination,${D}$ensity-based${C}$lustering with parameter adaptation and${O}\text{n}{L}$ine detection with subsampling (DDCOL). Compared with traditional statistical methods and unsupervised learning methods, extensive experimental results and analysis on a large number of public KPIs show the competitive performance ofDDCOLand the significance ofabrupt changes. Furthermore, we provide an interpretation for the promising results, which shows thatDDCOLcan be robust to KPI expected concept drifts, and obtain a good feature distribution of normal data in KPIs. Zhiping Cai, Siqi Wang 0001, Haiwen Chen, Fang Liu 0002, Anfeng Liu |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Robustness Can Be Cheap: A Highly Efficient Approach to Discover Outliers under High Outlier RatiosabstractEfficient detection of outliers from massive data with a high outlier ratio is challenging but not explicitly discussed yet. In such a case, existing methods either suffer from poor robustness or require expensive computations. This paper proposes a Low-rank based Efficient Outlier Detection (LEOD) framework to achieve favorable robustness against high outlier ratios with much cheaper computations. Specifically, it is worth highlighting the following aspects of LEOD: (1) Our framework exploits the low-rank structure embedded in the similarity matrix and considers inliers/outliers equally based on this low-rank structure, which facilitates us to encourage satisfying robustness with low computational cost later; (2) A novel re-weighting algorithm is derived as a new general solution to the constrained eigenvalue problem, which is a major bottleneck for the optimization process. Instead of the high space and time complexity (O((2n)2)/O((2n)3)) required by the classic solution, our algorithm enjoys O(n) space complexity and a faster optimization speed in the experiments; (3) A new alternative formulation is proposed for further acceleration of the solution process, where a cheap closed-form solution can be obtained. Experiments show that LEOD achieves strong robustness under an outlier ratio from 20% to 60%, while it is at most 100 times more memory efficient and 1000 times faster than its previous counterpart that attains comparable performance. The codes of LEOD are publicly available at https://github.com/demonzyj56/LEOD. Siqi Wang 0001, En Zhu, Xiping Hu, Xinwang Liu 0002, Qiang Liu 0004, Jianping Yin, Fei Wang 0001 |
AAAI | 1 |
| 2019 | Two-stage Unsupervised Video Anomaly Detection using Low-rank based Unsupervised One-class Learning with Ridge RegressionabstractVideo anomaly detection is a valuable but challenging task, especially in the field of surveillance videos for public safety. Almost all existing methods tackle the problem under the supervised setting and only a few attempts are conducted on the unsupervised learning. To avoid the cost of labeling training videos, this paper proposes to discriminate anomaly by a novel two-stage framework in a fully unsupervised manner. Unlike previous unsupervised approaches using local change detection to discover abnormality, our method enjoys the global information from video context by considering the pair-wise similarity of all video events. In this way, our method formulates video anomaly detection as an extension of unsupervised one-class learning, which has not been explored in the literature of video anomaly detection. Specifically, our method consists of two stages: The first stage of our kernel-based method, named Low-rank based Unsupervised One-class Learning with Ridge Regression (LR-UOCL-RR), reformulates the optimization goal of UOCL with ridge regression to avoid expensive computation, which enables our method to handle massive unlabeled data from videos. In the second stage, the estimated normal video events from the first stage are fed into the one-class support vector machine to refine the profile around normal events and enhance the performance. The experimental results conducted on two challenging video benchmarks indicate that our method is considerably superior, up to 15:7% AUC gain, to the state-of-the-art methods in the unsupervised anomaly detection task and even better than several supervised approaches. Jingtao Hu, En Zhu, Siqi Wang 0001, Siwei Wang 0001, Xinwang Liu 0002, Jianping Yin |
IJCNN | 3 |
| 2019 | Effective End-to-end Unsupervised Outlier Detection via Inlier Priority of Discriminative NetworkabstractDespite the wide success of deep neural networks (DNN), little progress has been made on end-to-end unsupervised outlier detection (UOD) from high dimensional data like raw images. In this paper, we propose a framework named E^3Outlier, which can perform UOD in a both effective and end-to-end manner: First, instead of the commonly-used autoencoders in previous end-to-end UOD methods, E^3Outlier for the first time leverages a discriminative DNN for better representation learning, by using surrogate supervision to create multiple pseudo classes from original unlabelled data. Next, unlike classic UOD that utilizes data characteristics like density or proximity, we exploit a novel property named inlier priority to enable end-to-end UOD by discriminative DNN. We demonstrate theoretically and empirically that the intrinsic class imbalance of inliers/outliers will make the network prioritize minimizing inliers' loss when inliers/outliers are indiscriminately fed into the network for training, which enables us to differentiate outliers directly from DNN's outputs. Finally, based on inlier priority, we propose the negative entropy based score as a simple and effective outlierness measure. Extensive evaluations show that E^3Outlier significantly advances UOD performance by up to 30% AUROC against state-of-the-art counterparts, especially on relatively difficult benchmarks. Siqi Wang 0001, Yijie Zeng, Xinwang Liu 0002, En Zhu, Jianping Yin, Chuanfu Xu, Marius Kloft |
NeurIPS | 1 |
| 2018 | Chronic Poisoning against Machine Learning Based IDSs Using Edge Pattern DetectionabstractIn big data era, machine learning is one of fundamental techniques in intrusion detection systems (IDSs). Poisoning attack, which is one of the most recognized security threats towards machine learning- based IDSs, injects some adversarial samples into the training phase, inducing data drifting of training data and a significant performance decrease of target IDSs over testing data. In this paper, we adopt the Edge Pattern Detection (EPD) algorithm to design a novel poisoning method that attack against several machine learning algorithms used in IDSs. Specifically, we propose a boundary pattern detection algorithm to efficiently generate the points that are near to abnormal data but considered to be normal ones by current classifiers. Then, we introduce a Batch-EPD Boundary Pattern (BEBP) detection algorithm to overcome the limitation of the number of edge pattern points generated by EPD and to obtain more useful adversarial samples. Based on BEBP, we further present a moderate but effective poisoning method called chronic poisoning attack. Extensive experiments on synthetic and three real network data sets demonstrate the performance of the proposed poisoning method against several well-known machine learning algorithms and a practical intrusion detection method named FMIFS-LSSVM-IDS. Pan Li 0006, Qiang Liu 0004, Siqi Wang 0001 |
ICC | 5 |
| 2018 | Octree-based Convolutional Autoencoder Extreme Learning Machine for 3D Shape ClassificationabstractWe introduce Octree-based Convolutional Autoencoder Extreme Learning Machine (OCA-ELM) for 3D shape classification. This approach combines Convolutional Autoencoder Extreme Learning Machine (CAE-ELM) with octreebased con- volution to generate feature maps from several types of geometric data, and extract discriminative features with Extreme Learning Machine Autoencoder (ELM-AE). The extracted features can then be used for various computer graphics applications, such as 3D shape classification. Compared with other 3D classification methods, the proposed OCA-ELM has superior classification performance. Experiments on ModelNet40 show that OCA-ELM outperforms state-of-the-art CNN-based methods and surpasses CAE-ELM in classification accuracy by 3.69%, demonstrating the effectiveness of our method. Jichao Chen, Yijie Zeng, Siqi Wang 0001, Soh Ling Min, Guang-Bin Huang |
IJCNN | 3 |
| 2018 | Detecting Abnormality without Knowing Normality: A Two-stage Approach for Unsupervised Video Abnormal Event DetectionabstractAbnormal event detection in video surveillance is a valuable but challenging problem. Most methods adopt a supervised setting that requires collecting videos with only normal events for training. However, very few attempts are made under unsupervised setting that detects abnormality without priorly knowing normal events. Existing unsupervised methods detect drastic local changes as abnormality, which overlooks the global spatio-temporal context. This paper proposes a novel unsupervised approach, which not only avoids manually specifying normality for training as supervised methods do, but also takes the whole spatio-temporal context into consideration. Our approach consists of two stages: First, normality estimation stage trains an autoencoder and estimates the normal events globally from the entire unlabeled videos by a self-adaptive reconstruction loss thresholding scheme. Second, normality modeling stage feeds the estimated normal events from the previous stage into one-class support vector machine to build a refined normality model, which can further exclude abnormal events and enhance abnormality detection performance. Experiments on various benchmark datasets reveal that our method is not only able to outperform existing unsupervised methods by a large margin (up to 14.2% AUC gain), but also favorably yields comparable or even superior performance to state-of-the-art supervised methods. Siqi Wang 0001, Yijie Zeng, Qiang Liu 0004, Chengzhang Zhu, En Zhu, Jianping Yin |
ACM Multimedia | 1 |
| 2018 | Video anomaly detection and localization by local motion based joint video representation and OCELM
Siqi Wang 0001, En Zhu, Jianping Yin, Fatih Porikli |
Neurocomputing | 1 |
| 2018 | Hyperparameter selection of one-class support vector machine by self-adaptive data shifting
Siqi Wang 0001, Qiang Liu 0004, En Zhu, Fatih Porikli, Jianping Yin |
Pattern Recognit. | 1 |
| 2018 | Incremental multiple kernel extreme learning machine and its application in Robo-advisors
Jingming Xue, Qiang Liu 0004, Miaomiao Li 0001, Xinwang Liu 0002, Yongkai Ye, Siqi Wang 0001, Jianping Yin |
Soft Comput. | 6 |
| 2017 | Complete three-phase detection framework for identifying abnormal cervical cellsabstractAutomatic identification of abnormal cervical cells, including feature representation, feature combination and classification strategy, is highly demanded in women's annual cervical cancer screenings. However, previous methods only deal with one or two of these three phases, and currently there is few complete framework for this problem. A novel three‐phrase boosting framework is proposed for the detection of abnormal cells from cervical smear images. First, the authors extract 160 dimensional features with respect to each cervical cell from three aspects, including cytology morphology, chromatin pathology and region intensity. In particular, 106 dimensional chromatin pathology features are newly adopted to describe the nucleus textural transformation. Second, an adaptive feature combination method is introduced to select the optimal feature patterns, which can combine all features using a reinforced margin‐based approach with the heuristic knowledge. Finally, a two‐stage classification strategy is presented to reduce erroneous classification abnormal cells using two different classifiers. Experimental results achieve state‐of‐the‐art performance and the proposed framework outperforms the other 16 compared detection methods. Kuan Li, Jianping Yin, Qiang Liu 0004, Siqi Wang 0001 |
IET Image Process. | 5 |
| 2017 | MST-GEN: An Efficient Parameter Selection Method for One-Class Extreme Learning MachineabstractOne-class classification (OCC) models a set of target data from one class to detect outliers. OCC approaches like one-class support vector machine (OCSVM) and support vector data description (SVDD) have wide practical applications. Recently, one-class extreme learning machine (OCELM), which inherits the fast learning speed of original ELM and achieves equivalent or higher data description performance than OCSVM and SVDD, is proposed as a promising alternative. However, OCELM faces the same thorny parameter selection problem as OCSVM and SVDD. It significantly affects the performance of OCELM and remains under-explored. This paper proposes minimal spanning tree (MST)-GEN, an automatic way to select proper parameters for OCELM. Specifically, we first build a n -round MST to model the structure and distribution of the given target set. With information from n -round MST, a controllable number of pseudo outliers are generated by edge pattern detection and a novel "repelling" process, which readily overcomes two fundamental problems in previous outlier generation methods: where and how many pseudo outliers should be generated. Unlike previous methods that only generate pseudo outliers, we further exploit n -round MST to generate pseudo target data, so as to avoid the time-consuming cross-validation process and accelerate the parameter selection. Extensive experiments on various datasets suggest that the proposed method can select parameters for OCELM in a highly efficient and accurate manner when compared with existing methods, which enables OCELM to achieve better OCC performance in OCC applications. Furthermore, our experiments show that MST-GEN can also be favorably applied to other prevalent OCC methods like OCSVM and SVDD. Siqi Wang 0001, Qiang Liu 0004, En Zhu, Jianping Yin |
IEEE Trans. Cybern. | 1 |
| 2016 | Anomaly detection in crowded scenes by SL-HOF descriptor and foreground classificationabstractWith the widespread use of surveillance cameras, massive video data analysis has become an extremely labor-intensive work. In this paper, we propose an efficient approach to detect video anomaly in crowded scenes based on Spatially Localized Histogram of Optical Flow (SL-HOF) descriptor and foreground classification. For motion description, the new SL-HOF descriptor can not only preserve classic HOF descriptor's favorable capability of characterizing the motion velocity and direction of foreground in crowded scene, but also depicts the spatial distribution of optical flow, which implicitly encodes the structure and local motion information of foreground objects in videos. SL-HOF is shown to significantly outperform other classic video descriptors. To further boost the performance of anomaly localization, we then introduce Robust PCA based foreground classification to discriminate anomalous foreground texture. Instead of computationally expensive approaches like l1-norm Sparse Coding, we adopt classic one-class SVM (OCSVM) to model normal video events and detect outliers (anomaly). Our experiments on the challenging UCSD datasets show our approach can achieve state-of-the-art results when compared to existing video anomaly detection methods. Siqi Wang 0001, En Zhu, Jianping Yin, Fatih Porikli |
ICPR | 1 |
| 2016 | Video anomaly detection based on ULGP-OF descriptor and one-class ELMabstractSmart video analysis is attracting increasing attention with the pervasive use of surveillance camera. In this paper, we address video anomaly detection by Uniform Local Gradient Pattern based Optical Flow (ULGP-OF) descriptor and one-class extreme learning machine (OCELM). Using the proposed ULGP-OF descriptor, we naturally combine the robust 2D image texture descriptor LGP with video optical flow to jointly descibe the texture and motion characteristics of video. ULGP-OF significantly outperforms other frequently-used classic video decriptors by a 6% to 10% EER reduction. As to normal video event modeling, the newly emergent ELM is introduced for the first time to tackle the unbearable training time incurred by massive training data from video streams. Compared to classic data description algorithms like one-class SVM (OCSVM) and sparse coding, OCELM can yield competitive results with a significant improvement in learning speed, which makes our approach more applicable to large-scale video analysis and easier for updating when video data are explosively generated in this day and age. Moreover, by adopting consistency-based criteria, only one parameter needs to be appointed for OCELM before training, which renders our approach much more parameter-free than other anomaly detection techniques like sparse coding. Experiments on UCSD ped1 and ped2 datasets demonstrate the effectiveness of our approach. Siqi Wang 0001, En Zhu, Jianping Yin |
IJCNN | 1 |
| 2016 | Random Fourier extreme learning machine with ℓ2, 1-norm regularization
Sihang Zhou 0001, Xinwang Liu 0002, Qiang Liu 0004, Siqi Wang 0001, Chengzhang Zhu, Jianping Yin |
Neurocomputing | 4 |