Xiao-Bo Jin

dblp:21/4209 · also Xiaobo Jin · DBLP profile ↗
← Back
41ranked-venue papers
9as first author
30since 2021 · last 2026
0000-0003-1671-1379ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Security and privacy · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A benchmark and method for photographed table reasoning
Xiaoqiang Kang, Xiaochen Zi, Xiao-Bo Jin, Kaizhu Huang, Qiufeng Wang 0001
Pattern Recognit.4
2026 StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors
abstract
Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well-trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models-crucial to medical image analysis-remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under closed-box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model-agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets (CMR dataset from UK Biobank, the SEG fundus dataset, the EchoNet echocardiography dataset, and the PraNet colonoscopy dataset) and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved attack success rates (ASR) above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores-significantly outperforming backdoor-based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at https://github.com/Qinkaiyu/StealthMark.
Qinkai Yu, Chong Zhang 0006, Gaojie Jin, Tianjin Huang, Wei Zhou 0021, Xiao-Bo Jin, Bo Huang 0012, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng
IEEE Trans. Image Process.7
2025 Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation
abstract
Solving tabular math word problems (TMWPs) has become a critical role in evaluating the mathematical reasoning ability of large language models (LLMs), where large-scale TMWP samples are commonly required for fine-tuning. Since the collection of high-quality TMWP datasets is costly and time-consuming, recent research has concentrated on automatic TMWP generation. However, current generated samples usually suffer from issues of either correctness or diversity. In this paper, we propose a Template-driven LLM-paraphrased (TeLL) framework for generating high-quality TMWP samples with diverse backgrounds and accurate tables, questions, answers, and solutions. To this end, we first extract templates from existing real samples to generate initial problems, ensuring correctness. Then, we adopt an LLM to extend templates and paraphrase problems, obtaining diverse TMWP samples. Furthermore, we find the reasoning annotation is important for solving TMWPs. Therefore, we propose to enrich each solution with illustrative reasoning steps. Through the proposed framework, we construct a high-quality dataset TabMWP-TeLL by adhering to the question types in the TabMWP dataset, and we conduct extensive experiments on a variety of LLMs to demonstrate the effectiveness of TabMWP-TeLL in improving TMWP-solving performance.
Xiaoqiang Kang, Xiao-Bo Jin, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001
AAAI3
2025 ShadowCraft-NeRF: Occlusion and Shadow Mitigation via SAM-Guided NeRF
Yushi Li, Yunyao Shen, Rong Chen 0003, Xiao-Bo Jin, Along Jin, Yu Han 0001
CASA6
2025 Can GRPO Boost Complex Multimodal Table Understanding?
abstract
Xiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu, Xiaobo Jin, Kaizhu Huang, Wei Wang, Yutao Yue, Xiaowei Huang, Qiufeng Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xiaoqiang Kang, Shengen Wu, Xiao-Bo Jin, Kaizhu Huang, Wei Wang 0042, Yutao Yue
EMNLP5
2025 DMFI: A Dual-Modality Log Analysis Framework for Insider Threat Detection with LoRA-Tuned Language Models
abstract
Insider threat detection (ITD) poses a persistent and high-impact challenge in cybersecurity due to the subtle, long-term, and context-dependent nature of malicious insider behaviors. Traditional models often struggle to capture semantic intent and complex behavior dynamics, while existing LLMbased solutions face limitations in prompt adaptability and modality coverage. To bridge this gap, we propose DMFI, a dual-modality framework that integrates semantic inference with behavior-aware fine-tuning. DMFI converts raw logs into two structured views: (1) a semantic view that processes content-rich artifacts (e.g., emails, https) using instruction-formatted prompts; and (2) a behavioral abstraction, constructed via a 4 W -guided (When-Where-What-Which) transformation to encode contextual action sequences. Two LoRA-enhanced LLMs are fine-tuned independently, and their outputs are fused via a lightweight MLP-based decision module. We further introduce DMFI-B, a discriminative adaptation strategy that separates normal and abnormal behavior representations, improving robustness under severe class imbalance. Experiments on CERT r4.2 and r5.2 datasets demonstrate that DMFI outperforms state-of-the-art methods in detection accuracy. Our approach combines the semantic reasoning power of LLMs with structured behavior modeling, offering a scalable and effective solution for real-world insider threat detection.
Kaichuan Kong, Dongjie Liu, Xiao-Bo Jin, Guanggang Geng, Zhiying Li 0003, Jian Weng 0001
ICDM3
2025 The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
Kaizhu Huang, Qiufeng Wang 0001, Xiao-Bo Jin
ICDM6
2025 ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.
Zihan Ye, Shreyank N. Gowda, Shiming Chen 0002, Xiaowei Huang 0001, Fahad Shahbaz Khan, Yaochu Jin, Kaizhu Huang, Xiao-Bo Jin
ICLR9
2025 Twin Progressive Generative Adversarial Network For High-Resolution Image Inpainting
abstract
Image inpainting aims to generate content for missing regions while maintaining visual coherence in the reconstructed images. Generative Adversarial Networks (GANs) have received increasing attention for their ability to repair high-resolution images. However, existing methods often overly rely on surrounding pixel information, overlooking noisy or incomplete information in boundary regions, which leads to blurred and unnatural results. To address these limitations, we propose a Twin Progressive Generative Adversarial Network (TP-GAN), which leverages global visual features from distant image contexts to reconstruct the overall structure and texture, improving the quality of high-resolution image inpainting. TP-GAN incorporates two generators and one discriminator, where the generators collaborate via exponential moving average optimization, focusing respectively on capturing fine details and global information. A progressive learning strategy is employed, starting with low-resolution restoration and gradually increasing resolution to simplify tasks and enhance adaptability to boundary regions. Extensive experimental evaluations on popular datasets demonstrate the superiority of TP-GAN.
Zhiying Li 0003, Zhaoxin Fan, Kaichuan Kong, Xiao-Bo Jin, Guanggang Geng
ICME5
2025 Performance is not All You Need: Sustainability Considerations for Algorithms
Chong Zhang 0006, Shreyank N. Gowda, Yushi Li, Xiao-Bo Jin
PRCV (12)6
2025 Log2Sig: Frequency-Aware Insider Threat Detection via Multivariate Behavioral Signal Decomposition
abstract
Insider threat detection presents a significant challenge due to the deceptive nature of malicious behaviors, which often resemble legitimate user operations. However, existing approaches typically model system logs as flat event sequences, thereby failing to capture the inherent frequency dynamics and multiscale disturbance patterns embedded in user behavior. To address these limitations, we propose Log2Sig, a robust anomaly detection framework that transforms user logs into multivariate behavioral frequency signals, introducing a novel representation of user behavior. Log2Sig employs Multivariate Variational Mode Decomposition (MVMD) to extract IMFs, which reveal behavioral fluctuations across multiple temporal scales. Based on this, the model further performs joint modeling of behavioral sequences and frequency-decomposed signals: the daily behavior sequences are encoded using a Mamba-Based temporal encoder to capture long-term dependencies, while the corresponding frequency components are linearly projected to match the encoder’s output dimension. These dual-view representations are then fused to construct a comprehensive user behavior profile, which is fed into a multilayer perceptron for precise anomaly detection. Experimental results on the CERT r4.2 and r5.2 datasets demonstrate that Log2Sig significantly outperforms state-of-the-art baselines in both accuracy and F1 score.
Kaichuan Kong, Dongjie Liu, Xiao-Bo Jin, Zhiying Li 0003, Guanggang Geng
TrustCom3
2025 Unveiling traffic paths: Explainable path signature feature-based encrypted traffic classification
abstract
Encryption technology ensures secure transmission for internet communications but poses significant challenges for effective encrypted traffic classification, which categorizes traffic into distinct groups, facilitating the process of monitoring network activities to uncover patterns and extract valuable information applicable in areas such as network management and anomaly detection. To this end, machine learning has emerged as a powerful technology for conducting encrypted traffic classification without compromising user data privacy. Machine learning-based classification demonstrates remarkable capabilities in processing vast amounts of data through sophisticated handcrafted features, with traffic path signature features representing the cutting edge of this field. This method shows stable performance improvements for common encrypted traffic types using only packet length information. However, it also yields a high dimensionality of path signature features, complicating the training of lightweight models and hindering further innovation due to a lack of model explainability. In this paper, we first propose leveraging feature selection to conduct feature dimensionality reduction, and then try to focus on the explanation of the model from both global and local perspectives. Performance comparisons indicate that our proposed method significantly reduces the number of path signature features while preserving classification performance, which enhances computational efficiency and meets the demand for lightweight models in various application scenarios. Furthermore, this significant reduction in the feature dimensionality allows for the interpretability of the model, which gives the user a clear understanding of the modeling decision-making process.
Kai-Chuan Kong, Xiao-Bo Jin, Guanggang Geng
Comput. Secur.3
2025 Progressive Enhancement Dehazing for object detection in extreme weather
Zhiying Li 0003, Junhao Wu 0003, Shuyuan Lin, Zheng Wang 0013, Xiao-Bo Jin, Guanggang Geng, Feiran Huang, Jian Weng 0001
Eng. Appl. Artif. Intell.5
2025 DPI-ITD: A Dual-Perspective Information-Driven Framework for Insider Threat Detection in IoT Systems
abstract
In Internet of Things (IoT) environments, insider threat detection has advanced with the integration of deep learning techniques, which can effectively model complex behaviors and heterogeneous data. However, the fragmented nature of IoT logs, behavioral redundancy, and the sparsity of insider actions increase detection complexity. While fine-grained behavior classification can improve accuracy, it also raises computational overhead, limiting applicability in resource-constrained scenarios. To address these challenges, we propose dual-perspective information-driven framework for insider threat detection (DPI-ITD), which combines user-centric and behavior-centric analyses to enhance detection efficiency and accuracy. DPI-ITD introduces a symbolic tagging strategy guided by tagging scores (TS), derived from user action diversity and behavioral context, to filter redundant fragments and focus on high-impact behaviors. It further incorporates an adaptive embedding mechanism based on GloVe, which dynamically adjusts the context window for rare but critical actions. Experiments on multiple closed and open behavioral datasets demonstrate DPI-ITD’s superior detection performance, scalability, and efficiency, confirming its suitability for lightweight deployment in real-world IoT security systems.
Kai-Chuan Kong, Xiao-Bo Jin, Dongjie Liu, Zhiquan Liu 0001, Guanggang Geng
IEEE Internet Things J.2
2025 Fine-grained recognition of citrus varieties via wavelet channel attention network
Fukai Zhang, Xiao-Bo Jin, Shan An, Qiang Lyu
Knowl. Based Syst.2
2024 Bridging the Projection Gap: Overcoming Projection Bias Through Parameterized Distance Learning
Chong Zhang 0006, Mingyu Jin, Qinkai Yu, Haochen Xue, Shreyank N. Gowda, Xiao-Bo Jin
ACCV (8)6
2024 Target-driven Attack for Large Language Models
abstract
Current large language models (LLM) provide a strong foundation for large-scale user-oriented natural language tasks. Many users can easily inject adversarial text or instructions through the user interface, thus causing LLM model security challenges like the language model not giving the correct answer. Although there is currently a large amount of research on black-box attacks, most of these black-box attacks use random and heuristic strategies. It is unclear how these strategies relate to the success rate of attacks and thus effectively improve model robustness. To solve this problem, we propose our target-driven black-box attack method to maximize the KL divergence between the conditional probabilities of the clean text and the attack text to redefine the attack’s goal. We transform the distance maximization problem into two convex optimization problems based on the attack goal to solve the attack text and estimate the covariance. Furthermore, the projected gradient descent algorithm solves the vector corresponding to the attack text. Our target-driven black-box attack approach includes two attack strategies: token manipulation and misinformation attack. Experimental results on multiple Large Language Models and datasets demonstrate the effectiveness of our attack method.
Chong Zhang 0006, Mingyu Jin, Dong Shu, Taowen Wang, Dongfang Liu, Xiao-Bo Jin
ECAI6
2024 Goal-Guided Generative Prompt Injection Attack on Large Language Models
abstract
Current large language models (LLMs) provide a strong foundation for large-scale user-oriented natural language tasks. Numerous users can easily inject adversarial text or instructions through the user interface, thus causing LLM model security challenges. Although there is much research on prompt injection attacks, most black-box attacks use heuristic strategies. It is unclear how these heuristic strategies relate to the success rate of attacks and thus effectively improve model robustness. To solve this problem, we redefine the goal of the attack: to maximize the KL divergence between the conditional probabilities of the clean text and the adversarial text. Furthermore, we prove that maximizing the KL divergence is equivalent to maximizing the Mahalanobis distance between the embedded representation$x$and$x^{\prime}$of the clean text and the adversarial text when the conditional probability is a Gaussian distribution and gives a quantitative relationship on$x$and$x^{\prime}$. Then we designed a simple and effective goal-guided generative prompt injection strategy (G2PIA) to find an injection text that satisfies specific constraints to achieve the optimal attack effect approximately. Notably, our attack method is a query-free black-box attack method with a low computational cost. Experimental results on seven LLM models and four datasets show the effectiveness of our attack method.
Chong Zhang 0006, Mingyu Jin, Qinkai Yu, Haochen Xue, Xiao-Bo Jin
ICDM6
2024 Multi-task Prompt Words Learning for Social Media Content Generation
abstract
The rapid development of the Internet has profoundly changed human life. Humans are increasingly expressing themselves and interacting with others on social media platforms. However, although artificial intelligence technology has been widely used in many aspects of life, its application in social media content creation is still blank. To solve this problem, we propose a new prompt word generation framework based on multi-modal information fusion, which combines multiple tasks including topic classification, sentiment analysis, scene recognition and keyword extraction to generate more comprehensive prompt words. Subsequently, we use a template containing a set of prompt words to guide ChatGPT to generate high-quality tweets. Furthermore, in the absence of effective and objective evaluation criteria in the field of content generation, we use the ChatGPT tool to evaluate the results generated by the algorithm, making large-scale evaluation of content generation algorithms possible. Evaluation results on extensive content generation demonstrate that our cue word generation framework generates higher quality content compared to manual methods and other cueing techniques, while topic classification, sentiment analysis, and scene recognition significantly enhance content clarity and its consistency with the image.
Haochen Xue, Chong Zhang 0006, Chenzhi Liu, Fangyu Wu 0001, Xiao-Bo Jin
IJCNN5
2024 STFT-TCAN: A TCN-attention based multivariate time series anomaly detection architecture with time-frequency analysis for cyber-industrial systems
abstract
Networks and industrial systems play a pivotal role in modern society, and their security has garnered increasing attention. Anomalies within industrial equipment may propagate through fault transmission, leading to a cascade of failures. Additionally, cyberattacks on equipment can result in significant losses. Therefore, in the realm of industrial and cyberspace domains, an effective multivariate time series anomaly detection system for monitoring equipment is instrumental in ensuring the healthy operation of the machinery. Nevertheless, detecting anomalies in numerous time series remains challenging, stemming from the absence of anomaly labels and the complexity of the data patterns. Existing algorithms predominantly concentrate on modeling within the time domain, falling short in fully leveraging the informative features present in frequency domain data, resulting in diminished detection performance. This paper introduces STFT-TCAN, a model for anomaly detection in time series that seamlessly integrates information from both time and frequency domains for extracting data features. Sliding windows and the Short Time Fourier Transform (STFT) are utilized to construct a frequency matrix, effectively amalgamating the characteristics of both time and frequency domains within the time series. Furthermore, the model employs Temporal Convolutional Networks (TCN) and Transformer attention mechanisms (which combined to form the TCAN module) to capture the features of multivariate time series, thereby resulting in heightened detection accuracy. The proposed model undergoes validation on six publicly available datasets, showcasing the superior performance of the STFT-TCAN model in comparison to current baseline methods. It adeptly extracts features from both frequency and time domains in sequential data, thereby achieving state-of-the-art performance in tasks related to anomaly detection in multivariate time series.
Fei-Fan Tu, Dongjie Liu, Zhiwei Yan, Xiao-Bo Jin, Guanggang Geng
Comput. Secur.4
2023 Context Does Matter: End-to-end Panoptic Narrative Grounding with Deformable Attention Refined Matching Network
abstract
Panoramic Narrative Grounding (PNG) is an emerging visual grounding task that aims to segment visual objects in images based on dense narrative captions. The current state-of-the-art methods first refine the representation of phrase by aggregating the most similar k image pixels, and then match the refined text representations with the pixels of the image feature map to generate segmentation results. However, simply aggregating sampled image features ignores the contextual information, which can lead to phrase-to-pixel mis-match. In this paper, we propose a novel learning framework called Deformable Attention Refined Matching Network (DRMN), whose main idea is to bring deformable attention in the iterative process of feature learning to incorporate essential context information of different scales of pixels. DRMN iteratively re-encodes pixels with the deformable attention network after updating the feature representation of the top-k most similar pixels. As such, DRMN can lead to accurate yet discriminative pixel representations, purify the top-k most similar pixels, and consequently alleviate the phrase-to-pixel mis-match substantially. Experimental results show that our novel design significantly improves the matching results between text phrases and image pixels. Concretely, DRMN achieves new state-of-the-art performance on the PNG benchmark with an average recall improvement 3.5%. The codes are available in: https://github.com/JaMesLiMers/DRMN.
Xiao-Bo Jin, Qiufeng Wang 0001, Kaizhu Huang
ICDM2
2023 WCANet: Wavelet Channel Attention Network for Citrus Variety Identification
abstract
The effective fine-grained identification of citrus varieties plays a vital role in the differential production management of citrus orchards. To our knowledge, there are few studies and publicly available datasets on fine-grained identification of citrus varieties. In this study, we propose Wavelet Channel Attention Network (WCANet) to solve the problem of fine-grained visual classification of citrus varieties and create a Citrus Variety Dataset (CVD) consisting of tree canopy images. WCANet combines global average pooling to extract global features and wavelet transform to capture local features, which greatly improves the capability of channel attention modules for multi-scale feature extraction. Experimental results demonstrate that the WCANet outperforms the state-of-the-art confidence estimation approaches on various benchmarks. Our code and dataset will be open-sourced at https://github.com/fightero/WCANet.
Fukai Zhang, Xiao-Bo Jin, Shan An, Qiang Lyu
ICIP2
2023 Image Blending Algorithm with Automatic Mask Generation
Haochen Xue, Mingyu Jin, Chong Zhang 0006, Qian Weng, Xiao-Bo Jin
ICONIP (8)6
2023 Rebalanced Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to identify unseen classes with zero samples during training. Broadly speaking, present ZSL methods usually adopt class-level semantic labels and compare them with instance-level semantic predictions to infer unseen classes. However, we find that such existing models mostly produce imbalanced semantic predictions, i.e. these models could perform precisely for some semantics, but may not for others. To address the drawback, we aim to introduce an imbalanced learning framework into ZSL. However, we find that imbalanced ZSL has two unique challenges: (1) Its imbalanced predictions are highly correlated with the value of semantic labels rather than the number of samples as typically considered in the traditional imbalanced learning; (2) Different semantics follow quite different error distributions between classes. To mitigate these issues, we first formalize ZSL as an imbalanced regression problem which offers empirical evidences to interpret how semantic labels lead to imbalanced semantic predictions. We then propose a re-weighted loss termed Re-balanced Mean-Squared Error (ReMSE), which tracks the mean and variance of error distributions, thus ensuring rebalanced learning across classes. As a major contribution, we conduct a series of analyses showing that ReMSE is theoretically well established. Extensive experiments demonstrate that the proposed method effectively alleviates the imbalance in semantic prediction and outperforms many state-of-the-art ZSL methods.
Zihan Ye, Guanyu Yang 0002, Xiao-Bo Jin, Youfa Liu, Kaizhu Huang
IEEE Trans. Image Process.3
2022 Towards Accurate Alignment and Sufficient Context in Scene Text Recognition
Yijie Hu, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Xiao-Bo Jin, Kaizhu Huang
ICONIP (3)5
2022 Sparse matrix factorization with L2, 1 norm for matrix completion
Xiao-Bo Jin, Jianyu Miao, Qiufeng Wang 0001, Guanggang Geng, Kaizhu Huang
Pattern Recognit.1
2022 Seeing Traffic Paths: Encrypted Traffic Classification With Path Signature Features
abstract
Although many network traffic protection methods have been developed to protect user privacy, encrypted traffic can still reveal sensitive user information with sophisticated analysis. In this paper, we propose ETC-PS, a novel encrypted traffic classification method with path signature. We first construct the traffic path with a session packet length sequence to represent the interactions between the client and the server. Then, path transformations are conducted to exhibit its structure and obtain different information. A multiscale path signature is finally computed as a kind of distinctive feature to train the traditional machine learning classifier, which achieves highly robust accuracy and low training overhead. Six publicly available datasets with different traffic types of HTTPS/1, HTTPS/2, QUIC, VPN, non-VPN, Tor, and non-Tor are used to conduct closed-world and open-world evaluations to verify the effectiveness of ETC-PS. The experimental results demonstrate that ETC-PS is superior to the state-of-the-art methods in terms of accuracy, f1 score, time complexity, and stability.
Guanggang Geng, Xiao-Bo Jin, Dongjie Liu, Jian Weng 0001
IEEE Trans. Inf. Forensics Secur.3
2022 Exploiting Attention-Consistency Loss For Spatial-Temporal Stream Action Recognition
abstract
Currently, many action recognition methods mostly consider the information from spatial streams. We propose a new perspective inspired by the human visual system to combine both spatial and temporal streams to measure their attention consistency. Specifically, a branch-independent convolutional neural network (CNN) based algorithm is developed with a novel attention-consistency loss metric, enabling the temporal stream to concentrate on consistent discriminative regions with the spatial stream in the same period. The consistency loss is further combined with the cross-entropy loss to enhance the visual attention consistency. We evaluate the proposed method for action recognition on two benchmark datasets: Kinetics400 and UCF101. Despite its apparent simplicity, our proposed framework with the attention consistency achieves better performance than most of the two-stream networks, i.e., 75.7% top-1 accuracy on Kinetics400 and 95.7% on UCF101, while reducing 7.1% computational cost compared with our baseline. Particularly, our proposed method can attain remarkable improvements on complex action classes, showing that our proposed network can act as a potential benchmark to handle complicated scenarios in industry 4.0 applications.
Xiao-Bo Jin, Qiufeng Wang 0001, Amir Hussain 0001, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.2
2021 An efficient multistage phishing website detection model based on the CASE feature framework: Aiming at the real web environment
abstract
Phishing has become a favorite method of hackers for committing data theft and continues to evolve. As long as phishing websites continue to operate, many more people and companies will suffer privacy leaks or financial losses. Therefore, the demand for fast and accurate phishing website detection grows stronger. However, the existing phishing detection methods do not fully analyze the features of phishing, and the performance and efficiency of the models only apply to certain limited datasets and need to be improved to be applied to the real web environment. This paper fully considers the social engineering principles of phishing, proposes a comprehensive and interpretable CASE feature framework and designs a multistage phishing detection model to effectively detect phishing sites, especially in the real web environment, where high efficiency and performance and extremely low false alarm rates are required. To fully verify the proposed method, two kinds of data experiments were carried out. One was the comparative experiments among different features and different detection models on CASE, which covers both classic machine learning and deep learning algorithms based on a constructed complex dataset. The other was a one-year phishing discovery experiment in the real web environment. The proposed method achieves better detection results under the premise of significantly shortening the execution time and works well in real phishing discovery, which proves its high practicability in reality.
Dongjie Liu, Guanggang Geng, Xiao-Bo Jin, Wei Wang 0083
Comput. Secur.3
2021 Unsupervised feature selection by non-convex regularized self-representation
Jianyu Miao, Yuan Ping 0003, Zhensong Chen 0001, Xiao-Bo Jin, Peijia Li, Lingfeng Niu
Expert Syst. Appl.4
2020 Multi-scale Attention Consistency for Multi-label Image Classification
Xiao-Bo Jin, Qiufeng Wang 0001, Kaizhu Huang
ICONIP (4)2
2019 Attentive Region Embedding Network for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of them study the discrimination power implied in local image regions (parts), which, in some sense, correspond to semantic attributes, have stronger discrimination than attributes, and can thus assist the semantic transfer between seen/unseen classes. In this paper, to discover (semantic) regions, we propose the attentive region embedding network (AREN), which is tailored to advance the ZSL task. Specifically, AREN is end-to-end trainable and consists of two network branches, i.e., the attentive region embedding (ARE) stream, and the attentive compressed second-order embedding (ACSE) stream. ARE is capable of discovering multiple part regions under the guidance of the attention and the compatibility loss. Moreover, a novel adaptive thresholding mechanism is proposed for suppressing redundant (such as background) attention regions. To further guarantee more stable semantic transfer from the perspective of second-order collaboration, ACSE is incorporated into the AREN. In the comprehensive evaluations on four benchmarks, our models achieve state-of-the-art performances under ZSL setting, and compelling results under generalized ZSL setting.
Guosen Xie, Li Liu 0004, Xiao-Bo Jin, Fan Zhu 0001, Zheng Zhang 0006, Jie Qin 0004, Yazhou Yao, Ling Shao 0001
CVPR3
2019 Stochastic Conjugate Gradient Algorithm With Variance Reduction
abstract
Conjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction1and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functions. We experimentally demonstrate that the CG with variance reduction algorithm converges faster than its counterparts for four learning models, which may be convex, nonconvex or nonsmooth. In addition, its area under the curve performance on six large-scale data sets is comparable to that of the LIBLINEAR solver for the L2-regularized L2-loss but with a significant improvement in computational efficiency.
Xiao-Bo Jin, Xu-Yao Zhang, Kaizhu Huang, Guanggang Geng
IEEE Trans. Neural Networks Learn. Syst.1
2018 Resource Allocation for Energy Efficiency Maximization of Layered Multicast in SCMA Networks
abstract
As a candidate non-orthogonal multiple access technology for 5G, sparse code multiple access (SCMA) can effectively improve spectrum efficiency. In this paper, we first study the layered multicast scheme in SCMA network, its system capacity is no longer limited by the worst channel quality user in the multicast group. Specifically, we formulate a network energy efficiency (EE) maximization problem subject to quality-of-service (QoS) requirements, codebook assignment, and power allocation. In order to reduce the computational complexity, we then propose a sub-optimization algorithm by separating codebook assignment and power allocation. Finally, the simulation results verify the superiority of the proposed algorithm in layered multicast based on SCMA network in terms of the network EE.
Xiao-Bo Jin, Xiaoxiang Wang
APCC1
2018 Approximately optimizing NDCG using pair-wise loss
Xiao-Bo Jin, Guanggang Geng, Guosen Xie, Kaizhu Huang
Inf. Sci.1
2017 Boosting the phishing detection performance by semantic analysis
abstract
Phishing is increasingly severe in recent years, which seriously threatens the privacy and property security of netizens. Phishing is essentially a counterfeiting of brands. In order to effectively cheat the victim, phishing sites are visually and semantically highly similar to real sites. In recent years, anti-phishing methods based on machine learning are mainstream anti-phishing methods. The effectiveness of the machine learning models hinges on the extracted statistical features. However, the extracted statistical features mainly focus on visual similarity, stealing information and third-party services, which ignore the semantic information of web pages. Therefore, we extract a series of semantic features through word2vec to better describe the features of phishing sites, and further fuse them with other multi-scale statistical features to construct a more robust phishing detection model. The experimental results on the actual data sets show that the majority of phishing websites are effectively identified by only mining the semantic features of word embeddings. The phishing detection models based on fusion features obtained the best detection results, which shows that semantic features and other statistical features have good complementarity. The proposed method provides a promising way for phishing detection in actual Internet environment, which boosts the phishing detection performance effectively.
Xiao-Bo Jin, Zhiwei Yan, Guanggang Geng
IEEE BigData3
2015 Combination of multiple bipartite ranking for multipartite web content quality evaluation
Xiao-Bo Jin, Guanggang Geng, Minghe Sun, Dexian Zhang
Neurocomputing1
2012 Multi-label learning vector quantization algorithm
Xiao-Bo Jin, Guanggang Geng, Dexian Zhang
ICPR1
2010 Multi-class AdaBoost with Hypothesis Margin
abstract
Most AdaBoost algorithms for multi-class problems have to decompose the multi-class classification into multiple binary problems, like the Adaboost.MH and the LogitBoost. This paper proposes a new multi-class AdaBoost algorithm based on hypothesis margin, called AdaBoost.HM, which directly combines multi-class weak classifiers. The hypothesis margin maximizes the output about the positive class meanwhile minimizes the maximal outputs about the negative classes. We discuss the upper bound of the training error about AdaBoost.HM and a previous multi-class learning algorithm AdaBoost.M1. Our experiments using feed forward neural networks as weak learners show that the proposed AdaBoost.HM yields higher classification accuracies than the AdaBoost.M1 and the AdaBoost.MH, and meanwhile, AdaBoost.HM is computationally efficient in training.
Xiao-Bo Jin, Xinwen Hou, Cheng-Lin Liu 0001
ICPR1
2010 Regularized margin-based conditional log-likelihood loss for prototype learning
Xiao-Bo Jin, Cheng-Lin Liu 0001, Xinwen Hou
Pattern Recognit.1
2008 Prototype learning with margin-based conditional log-likelihood loss
abstract
The classification performance of nearest prototype classifiers largely relies on the prototype learning algorithms, such as the learning vector quantization (LVQ) and the minimum classification error (MCE). This paper proposes a new prototype learning algorithm based on the minimization of a conditional log-likelihood loss (CLL), called log-likelihood of margin (LOGM). A regularization term is added to avoid over-fitting in training. The CLL loss in LOGM is a convex function of margin, and so, gives better convergence than the MCE algorithm. Our empirical study on a large suite of benchmark datasets demonstrates that the proposed algorithm yields higher accuracies than the MCE, the generalized LVQ (GLVQ), and the soft nearest prototype classifier (SNPC).
Xiao-Bo Jin, Cheng-Lin Liu 0001, Xinwen Hou
ICPR1