VLDB 2026 Research / reviewers in the wild / expert
Xiangjian He
dblp:75/2122 · also Sean He
· DBLP profile ↗
226ranked-venue papers
12as first author
62since 2021 · last 2026
0000-0001-8962-540XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 106 · 7 first-author · 25 since 2021Artificial intelligence and machine learning · 62 · 2 first-author · 24 since 2021Computer networks · 19 · 1 first-author · 7 since 2021Systems, architecture and hardware · 17 · 3 first-author · 1 since 2021Security and privacy · 17 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFFormer: Cross CNN-Transformer channel attention and spatial feature fusion for improved segmentation of heterogeneous medical images
Qing Xu 0014, Xiangjian He, Daokun Zhang, Ruili Wang 0001, Rong Qu, Guoping Qiu |
Expert Syst. Appl. | 3 |
| 2026 | SP-Det: Self-prompted dual-text fusion for generalized multi-label lesion detectionabstractAutomated lesion detection in chest X-rays has demonstrated significant potential for improving clinical diagnosis by precisely localizing pathological abnormalities. While recent promptable detection frameworks have achieved remarkable accuracy in target localization, existing methods typically rely on manual annotations as prompts, which are labor-intensive and impractical for clinical applications. To address this limitation, we propose SP-Det, a novel self-prompted detection framework that automatically generates rich textual context to guide multi-label lesion detection without requiring expert annotations. Specifically, we introduce an expert-free dual-text prompt generator (DTPG) that leverages two complementary textual modalities: semantic context prompts that capture global pathological patterns and disease beacon prompts that focus on disease-specific manifestations. Moreover, we devise a bidirectional feature enhancer (BFE) that synergistically integrates comprehensive diagnostic context with disease-specific embeddings to significantly improve feature representation and detection accuracy. Extensive experiments on two chest X-ray datasets with diverse thoracic disease categories demonstrate that our SP-Det framework outperforms state-of-the-art detection methods while completely eliminating the dependency on expert-annotated prompts compared to existing promptable architectures. Qing Xu 0014, Yanqian Wang, Xiangjian He, Yixuan Zhang 0006, Rong Qu, Wenting Duan, Zhen Chen 0013 |
Knowl. Based Syst. | 3 |
| 2026 | UDG-Prom: A unified dense-guided semantic prompting for cross-domain few-shot image segmentationabstract• MAF preserves low-level feature representations, while fusing global and local information to generate robust class-agnostic features. • TA2MP, as a unified feature transformation mechanism equipped with an automatic learnable prompt branch, reduces human reliance and disentangles domain- and class-specific information through contrastive learning. • UDG-Prom integrates the MAF and TA2MP modules to address the CD-FSS task with SAM. • Our model achieves competitive or superior performance compared to state-of-the-art methods on four CD-FSS benchmarks, and its strong generalization ability is comprehensively validated through evaluations on more difficult cross-domain datasets including CT-Lung (medical) and SUIM (underwater). Large Vision Models (LVMs), exemplified by SAM, contain powerful general knowledge from extensive pre-training, yet they often underperform in highly specialized domains. Building large models tailored for each domain is usually impractical due to the substantial cost of data collection and training. Therefore, a key challenge is how to tap into SAM’s strong knowledge base and transfer it effectively to new, domain-specific tasks, especially under Cross-Domain or Few-Shot constraints. Previous efforts have leveraged prior knowledge from foundation models for transfer learning; however, they typically target specific tasks and exhibit limited robustness in broader applications. To tackle this issue, we propose a Unified Dense-Guided Semantic Prompting framework (UDG-Prom), a new paradigm for Cross-Domain Few-Shot Segmentation (CD-FSS). First, a Multi-level Adaptation Framework (MAF) is used for integrated feature extraction as prior knowledge. Then, we incorporate a Task-Adaptive Auto Meta Prompt (TA 2 MP) module to enable the extraction of class-domain-agnostic features and generate high-quality, learnable visual prompts. By combining learnable prompts with a structured model and prototype disentanglement, this method retains SAM’s prior knowledge and effectively adapts to CD-FSS through category and domain cues. Extensive experiments on four benchmarks show that our model not only surpasses state-of-the-art CD-FSS approaches but also achieves a remarkable improvement in average accuracy. Xiangjian He, Xin Chen 0003, Jingxi Hu, LinLin Shen, Guoping Qiu |
Knowl. Based Syst. | 2 |
| 2026 | Reliable-Teacher: Uncertainty-Guided Collaborative Learning for Nighttime Object DetectionabstractNighttime object detection presents significant challenges due to the scarcity of large-scale, high-quality annotations across diverse nighttime scenarios. To circumvent the need for manual nighttime image annotation, researchers have explored Unsupervised Domain Adaptive Object Detection (UDA-OD), which transfers knowledge from labeled daytime datasets to unlabeled nighttime data through pseudo-labeling. While existing approaches have shown promising results, their effectiveness remains limited by the low quality of pseudo labels, restricting model adaptation to nighttime conditions. To address these limitations, we propose Reliable-Teacher, a novel mutual-learning framework that comprehensively leverages target domain knowledge through Uncertainty-Guided Collaborative Learning. Specifically, our approach consists of three key components: 1) A Collaborative Pseudo-Label Construction module that intelligently integrates reliable Teacher-generated pseudo-labels into Student proposals, significantly enhancing pseudo-label quality; 2) An Uncertainty-Guided Consistency Reasoning module that enforces inter-category consistency between Teacher and Student predictions at both anchor and bounding box levels; 3) A Reliability-Weighted Classification Loss that minimizes the influence of unreliable predictions to further enhance uncertainty-guided learning. Extensive experiments demonstrate that Reliable-Teacher significantly outperforms state-of-the-art methods, achieving performance gain of up to 3.1%, 2.2% and 1.7% mAP on BDD100K [1], SHIFT [2], and VisDrone [3] benchmarks, respectively. Upon acceptance, our code will be released to facilitate further research in this domain. Wenjing Jia, Jiaqi Xiao, Jinchang Ren, Di Yuan 0002, Qiguang Miao, Xiangjian He |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Action Recognition in RGB-D Videos via an End-to-End Cross-Modal Attention TransformerabstractRecognizing actions in RGB-D videos necessitates an in-depth understanding of spatial and temporal information from both RGB and depth modalities, as well as an efficient fusion of these data streams. Existing late fusion approaches often suffer from modality collapse due to high semantic similarity between modalities, where the fusion diminishes the distinct contributions of each modality and compromises overall effectiveness. In this context, we propose a novel End-to-end Cross-Modal Attention Transformer (E-CMAT) model. Our model processes RGB and depth inputs through two distinct expert encoders that utilize factorized spatio-temporal representations to capture dimension-independent features, enhancing motion understanding. The extracted RGB and depth tokens are subsequently fused via our innovative cross-modal attention, which is applied iteratively to ensure a robust integration that accentuates critical patterns and discrepancies between the modalities, thereby facilitating more precise action recognition. Our cross-modal attention utilizes one class token per expert for inter-expert information exchange, focusing attention on salient features and fostering synergistic predictions while reducing computational complexity from quadratic to linear. Furthermore, extensive experiments conducted on widely used benchmark datasets, such as NTU RGB-D 60, NTU RGB-D 120, and THU-READ, demonstrate the favorable performance of E-CMAT compared to state-of-the-art models. Yujun Ma, Benjia Zhou, Hong Zhang 0050, Xizheng Zhang, Xiangjian He, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | De-LightSAM: Modality-Decoupled Lightweight SAM for Generalizable Medical SegmentationabstractThe universality of deep neural networks across different modalities and their generalization capabilities to unseen domains play an essential role in medical image segmentation. The recent segment anything model (SAM) has demonstrated strong adaptability across diverse natural scenarios. However, the huge computational costs, demand for manual annotations as prompts and conflict-prone decoding process of SAM degrade its generalization capabilities in medical scenarios. To address these limitations, we propose a modality-decoupled lightweight SAM for domain-generalized medical image segmentation, named De-LightSAM. Specifically, we first devise a lightweight domain-controllable image encoder (DC-Encoder) that produces discriminative visual features for diverse modalities. Further, we introduce the self-patch prompt generator (SP-Generator) to automatically generate high-quality dense prompt embeddings for guiding segmentation decoding. Finally, we design the query-decoupled modality decoder (QM-Decoder) that leverages a one-to-one strategy to provide an independent decoding channel for every modality, preventing mutual knowledge interference of different modalities. Moreover, we design a multi-modal decoupled knowledge distillation (MDKD) strategy to leverage robust common knowledge to complement domain-specific medical feature representations. Extensive experiments indicate that De-LightSAM outperforms state-of-the-arts in diverse medical imaging segmentation tasks, displaying superior modality universality and generalization capabilities. Especially, De-LightSAM uses only 2.0% parameters compared to SAM-H. The source code is available at https://github.com/xq141839/De-LightSAM. Qing Xu 0014, Xiangjian He, Chenxin Li, Fiseha B. Tesema, Wenting Duan, Zhen Chen 0013, Rong Qu, Jonathan M. Garibaldi, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | AGT-Diff: Anatomy-Guided Multimodal Adaptive Diffusion for Temporally-Aware PET DenoisingabstractPositron emission tomography (PET) is essential in clinical imaging but limited by long scan times and radiation exposure. Enhancing low-dose PET (LDPET) images is vital for reducing patient burden while maintaining diagnostic reliability. Existing denoising methods either ignore anatomical information from CT or over-integrate it, leading to poor generalization and PET signal distortion. We propose AGT-Diff—Anatomy-Guided Multimodal Adaptive Diffusion for temporally-aware PET denoising. AGT-Diff introduces two key modules. The Anatomy-Guided Multimodal Adaptive Fusion (AGMA) module selectively incorporates CT-derived structural priors through learnable soft fusion, preserving PET-specific metabolic patterns. The Progressive Time-Aware Supervision (PTA) module aligns intermediateduration PET scans with diffusion steps via signal-to-noise-ratio scheduling, enabling realistic and duration-adaptive denoising. Extensive experiments on whole-body PET/CT datasets demonstrate that AGT-Diff surpasses state-of-the-art methods in quantitative accuracy and structural fidelity while reducing dependence on paired high-dose data. Its fast inference and adaptive multimodal design make AGT-Diff a practical and generalizable framework for clinical PET enhancement. Jingxi Hu, Xiangjian He, Zhanli Hu |
BIBM | 3 |
| 2025 | Beyond Human Labels: A Multi-Linguistic Auto-Generated Benchmark for Evaluating Large Language Models on Resume ParsingabstractEfficient resume parsing is critical for global hiring, yet the absence of dedicated benchmarks for evaluating large language models (LLMs) on multilingual, structure-rich resumes hinders progress. To address this, we introduce ResumeBench, the first privacy-compliant benchmark comprising 2,500 synthetic resumes spanning 50 templates, 30 career fields, and 5 languages. These resumes are generated through a human-in-the-loop pipeline that prioritizes realism, diversity, and privacy compliance, which are validated against real-world resumes. This paper evaluates 24 state-of-the-art LLMs on ResumeBench, revealing substantial variations in handling resume complexities. Specifically, top-performing models like GPT-4o exhibit challenges in cross-lingual structural alignment while smaller models show inconsistent scaling effects. Code-specialized LLMs underperform relative to generalists, while JSON outputs enhance schema compliance but fail to address semantic ambiguities. Our findings underscore the necessity for domain-specific optimization and hybrid training strategies to enhance structural and contextual reasoning in LLMs. Zijian Ling, Han Zhang 0027, Zhequn Wu, Xu Sun 0002, Xiangjian He |
EMNLP | 7 |
| 2025 | Swin-VasMamba: A Topologically Constrained Model For 3D Vascular SegmentationabstractAccurate 3D vascular segmentation is essential for diagnosing and treating vascular diseases. This task remains challenging due to the complexity of the 3D data and the morphological diversity of blood vessels. In recent years, state space models (SSMs) have received a great attention for its good performance while preserving global receptive field and consuming less computing resources and time. Inspired by this, we propose a model called Swin-VasMamba for 3D vascular segmentation. It consists of a network called CMU-Net and a topologically constrained loss function called dsh loss. We compare our model with several other advanced segmentation models based on CNN, Transformer and Mamba. The results show that Swin-VasMamba achieves a state-of-the-art performance, with the highest Dice coefficient of 0.880, the lowest 95th-percentile of Hausdorff Distance (HD95) of 0.673, and the lowest Average Surface Distance (ASD) of 0.159 on a benchmark dataset. Xiangjian He, Qing Xu 0014, Xin Chen 0003, Shoujun Zhou |
ICASSP | 3 |
| 2025 | WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageabstractRecent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy. Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Xiangjian He, Xiaohan Xing, Sen Yang 0006, LinLin Shen |
ICCV | 6 |
| 2025 | 🤖 WSI-Agents: A Collaborative Multi-agent System for Multi-modal Whole Slide Image Analysis
Xinheng Lyu, Yuci Liang, Wenting Chen, Meidan Ding, Guolin Huang, Daokun Zhang, Xiangjian He, LinLin Shen |
MICCAI (5) | 8 |
| 2025 | EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001 |
MICCAI (13) | 10 |
| 2025 | CaDGS: Modeling Inter-Gaussian Mutual Information for Dynamic Novel View SynthesisabstractDynamic novel view synthesis (NVS) aims to render time-varying scenes from arbitrary viewpoints, balancing rendering quality and computational efficiency. While recent 4D Gaussian Splatting approaches offer promising real-time performance, they fundamentally overlook critical interdependence between Gaussians by modeling deformations independently. Our information-theoretic analysis reveals substantial mutual information across the Gaussian field, manifesting as appearance-preserving radiance coherence and motion-consistent deformation propagation. This finding establishes that rendering quality emerges from coordinated transformation rather than independent processing. We propose Correlation-aware Dynamic Gaussian Splatting (CaDGS) with our novel Gaussian Correlation Tensor Projection (GCTP) method, which efficiently transforms the complex O(n3) mutual information tensor into a dual-channel O(n2) spatial matrix, preserving the critical topological structure of Gaussian interactions. Combined with our Spatio-Temporal Deformation Consistency (STDC) learning, which enforces volumetric coherence through tensor-guided regularization across multiple scales, CaDGS prevents geometric distortions and texture inconsistencies common in previous approaches. Experimental results demonstrate state-of-the-art performance, achieving 32.4 PSNR on the Neu3D dataset with fewer Gaussians while maintaining rendering speeds of 323 FPS at 1353 × 1014 resolution. Yunlong Zhao 0003, Xiaoheng Deng, Zhuohua Qiu, Chang Xu 0002, Xiangjian He, Shan You, Xiu Su |
ACM Multimedia | 6 |
| 2025 | LiteSpiralGCN: Lightweight 3D hand mesh reconstruction via spiral graph convolution
Yiteng Wang, Minqi Li, Kaibing Zhang, Xiangjian He |
Appl. Intell. | 4 |
| 2025 | Towards empathic medical conversation in Narrative Medicine: A visualization approach based on intelligence augmentationabstractEmpathic medical conversation is central to patient-centered care within Narrative Medicine. However, difficulties, such as physicians’ limited empathic capabilities and lack of time, impede the practice. Research on real-time, on-site empathic medical exchanges has been limited in exploring technology to assist and enhance physicians’ capabilities. This paper proposed the Empathic Opportunity Perception and Distinction (EOPD) framework for building physician-AI collaboration based on Intelligence Augmentation (IA) for empathic conversations. The EOPD integrates two multi-modal machine learning (ML) models based on facial and verbal cues, presenting a physician-AI interaction framework and three distinctive visualization components: emotional reference, opportunity reminding and keyword collection, and situation understanding. To assess EOPD's effectiveness and gauge physicians’ and patients’ receptiveness, a prototype system named EMVIS ( EM otional VIS ualization ) was designed and developed. Results from the study demonstrated improvements in physicians’ empathy efforts and perceived empathy performance when using EMVIS, particularly for junior physicians. Physicians and patients held positive attitudes towards EMVIS, with patients expressing a high expectation that EMVIS would improve the physician-patient relationship. The research showed the efficacy of the multi-modal ML models in supporting complex affective empathy and EMVIS in facilitating and complementing empathy concerns. It highlighted the tailored support to junior and senior physicians and emphasized physician-AI collaboration to maintain user autonomy and mitigate potential biases. Future research should explore extensive system applications, tailor visual and interactive support for physicians, and implement adaptive and reflective ML models to improve the effectiveness and efficiency of empathy communications. Effie Lai-Chong Law, Xu Sun 0002, Weili Yang, Xiangjian He, Glyn Lawson, Huizhong Zheng, Qingfeng Wang 0002, Xiaoru Yuan |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | MDNet: Multimodal Cooperative Perception via Spatial Alignment of Modal Decision-MakingabstractThrough Internet of Things (IoT) communication technology, collaborative perception enhances a vehicle’s capacity to discern its surroundings while driving by integrating and synchronizing sensor data from multiple agents. With the advancement of cooperative perception techniques in single-modality methods, there has been a growing trend toward integrating multimodal data from heterogeneous sensors in recent years. However, due to the data heterogeneity inherent in diverse sensors, Bird’s Eye View (BEV) maps generated from different types of sensors may exhibit local discrepancies in the spatial representation of entity positions. Furthermore, individual agents may produce uncertain and flawed feature representations in real noisy environments. The influence of this indeterminacy exacerbates the issue of local inconsistency, leading to misalignment of the detected target during BEV alignment and fusion, thereby reducing detection accuracy. To address these problems, we propose a modal decision-making spatial alignment cooperative perception network (MDNet). First, the network generates BEV feature maps through dense depth image supervision for voxel feature extraction and model-guided selective feature fusion. Subsequently, we achieve enhanced accuracy in object detection by performing spatial alignment of BEV representations generated from two distinct sensors, both globally and locally within the spatial domain. Besides, we employ a cascaded centralized pyramid strategy during the message fusion stage, facilitating flexible sampling across horizontal and vertical spatial dimensions, promoting deep interaction among multiple agents. We conduct quantitative and qualitative experiments on the public OPV2V and DAIR-V2X-C benchmarks, and our proposed MDNet exhibits superior performance and stronger robustness in the 3-D object detection task, providing more precise target detection results. Junyang He, Xiaoheng Deng, Jinsong Gui, Tao Zhang 0010, Xiangjian He |
IEEE Internet Things J. | 5 |
| 2025 | HFA-UNet: hybrid and full attention UNet for thyroid nodule segmentationabstractUltrasound imaging is the most commonly used method for screening thyroid nodules due to its low cost and non-invasive nature. Thyroid nodule lesions have variable shapes, rich aspect ratios, unclear boundaries, calcified nodule-induced acoustic shadows, and noise interference, causing challenges in accurate segmentation. Recent methods ignore various scale features and details in different resolutions of images, leading to redundant or missing feature information and then affecting the segmentation performance. In this paper, we introduce a hybrid and full attention UNet model for ultrasound thyroid nodule segmentation. Self, spatial and channel attention are combined in a U-Net-like structure to extract global and local features simultaneously. A novel full attention multi-scale fusion stage is designed to enhance boundary features while suppressing noise features. At the same time, the model dynamically adjusts the number of skip connections corresponding to images of different resolutions to better utilize multi-scale features and detailed information. We evaluate our model on DDTI, TN3K and Stanford Cine-Clip datasets, including internal validation and cross-dataset testing. The results show that our proposed model for internal validation in the DDTI dataset increases the Dice score and mean intersection over union by 2.36 % and 1.04 % compared to the state-of-the-art model. In the TN3K dataset, they increase by 1.66 % and 3.05 %. Yuanhao Zou, Xiangjian He, Qing Xu 0014, Ming Liu 0021, Shengji Jin, Qian Zhang 0018, Maggie M. He, Jian Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2025 | NuSegDG: Integration of heterogeneous space and Gaussian kernel for domain-generalized nuclei segmentation
Zhenye Lou, Qing Xu 0014, Zekun Jiang, Xiangjian He, Chenxin Li, Zhen Chen 0013, Yi Wang 0037, Maggie M. He, Wenting Duan |
Knowl. Based Syst. | 4 |
| 2025 | FRFCNet: Feature Refinement and Flexible Concatenation for Object DetectionabstractThe state-of-the-art YOLO detection algorithms still suffer from the issue of redundant extraction of similar features during feature propagation, and the simplistic stacking approach of connecting different features limits the flexibility of feature fusion. We propose a new feature recombination mechanism involving refining feature extraction and flexible concatenation. It includes the HFConv (Hybrid Flexibility Convolution) module, the MFD (Multivariate Flexibility Downsampling) module, and the DFSPP (Deformable and Flexible Spatial Pyramid Pooling) module. Specifically, the HFConv module employs feature refinement and flexible connection strategies to optimize feature representation and reduce redundancy in a dynamic way, acquiring diverse feature information from local and surrounding regions. The MFD module leverages multiple downsampling methods to address the issue of feature redundancy that may arise from a single downsampling method, thereby enhancing feature diversity. The DFSPP module learns an offset corresponding to the pooling kernel size, allowing for the extraction of the most critical information in a dynamic manner. By incorporating these modules into the YOLO architecture, we develop a more robust network called FRFCNet, and the experimental results show a notable 4.1% and 2.8% improvement in AP values on the VOC2012 and COCO2017 datasets, respectively, compared to the baseline (YOLOV7-Tiny-SiLu), outperforming current one-stage detectors. Tao Zhang 0010, Xiangjian He, Qiang Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone ImageryabstractObject detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of objects in such images. Specifically, a set of patches potentially containing objects are first generated. A set of rewards measuring the localization accuracy, the accuracy of predicted labels, and the scale consistency among nearby patches are designed in the agent to guide the scale optimization. The proposed scale-consistency reward ensures similar scales for neighboring objects of the same category. Furthermore, a spatial-semantic attention mechanism is designed to exploit the spatial semantic relations between patches. The agent employs the proximal policy optimization strategy in conjunction with the evolutionary strategy, effectively utilizing both the current patch status and historical experience embedded in the agent. The proposed model is compared with state-of-the-art methods on two benchmark datasets for object detection on drone imagery. It significantly outperforms all the compared methods. Code is available at https://github.com/UNNC-CV/EvOD/. Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Yitian Zhao, Ruibin Bai, Xiangjian He, Jiang Liu 0001 |
AAAI | 8 |
| 2024 | Zero Trust for Intrusion Detection System: A Systematic Literature Review
Abeer Z. Alalmaie, Nazar Waheed, Mohrah Alalyan, Priyadarsi Nanda, Wenjing Jia, Xiangjian He |
ICAART (3) | 6 |
| 2024 | MambaVesselNet: A Hybrid CNN-Mamba Architecture for 3D Cerebrovascular SegmentationabstractSegmenting vessels in magnetic resonance imaging (MRI) stands as a mainstream approach for evaluating cerebrovascular conditions.Due to the complex semantics and topology of cerebrovascular structures, existing CNN-based segmentation methods often fail to correlate the topological structure and branch vessels, resulting in incomplete segmentation.To address the challenge of global dependencies modelling, transformer architectures have been employed due to their capability of capturing long-range dependencies, and they have shown promise in 3D medical image segmentation.However, the transformer architecture greatly increases the computational burden when processing high-dimensional 3D MRI images.In light of this, a selective state space model (SSM) Mamba has gained recognition for its adeptness in handling long-range dependencies in sequential data, particularly noted for its efficiency and speed in natural language processing applications.Mamba is now widely applied in various computer vision tasks.Based on these findings, in this study, we propose MambaVesselNet, a Hybrid CNN-Mamba network for 3D cerebrovascular segmentation.MambaVesselNet leverages CNNs to capture local features and incorporates the Mamba block at the bottleneck to model long-range dependencies within the whole-volume features.The effectiveness of MambaVesselNet is validated on a public cerebrovascular dataset, and our benchmark demonstrates new state-of-the-art performance. Xiangjian He |
MMAsia | 3 |
| 2024 | SS-FS CSA: Self-Supervised and Fully Supervised Integration for 3D Cerebrovascular SegmentationabstractThree-dimensional cerebrovascular segmentation is crucial for accurate diagnosis and treatment planning of cerebrovascular diseases.However, the lack of high-quality publicly labelled datasets can limit sufficient training, leading to inaccurate results.To address this issue, this study proposes a novel method that combines self-supervised and fully supervised learning, termed the SS-FS Cerebrovascular Segmentation Approach (SS-FS CSA).The method introduces publicly available unlabelled databases into the training process, alleviating the problem of insufficient high-quality labelled medical datasets.The SS-FS CSA method achieves a Dice Similarity Coefficient (DSC) of 82.82%, improving over 2% compared to the SOTA baseline, proving its validity and feasibility in 3D segmentation tasks. Chenxi Niu, Xiangjian He |
MMAsia | 3 |
| 2024 | Automatic quantitative stroke severity assessment based on Chinese clinical named entity recognition with domain-adaptive pre-trained large language modelabstractBACKGROUND: Stroke is a prevalent disease with a significant global impact. Effective assessment of stroke severity is vital for an accurate diagnosis, appropriate treatment, and optimal clinical outcomes. The National Institutes of Health Stroke Scale (NIHSS) is a widely used scale for quantitatively assessing stroke severity. However, the current manual scoring of NIHSS is labor-intensive, time-consuming, and sometimes unreliable. Applying artificial intelligence (AI) techniques to automate the quantitative assessment of stroke on vast amounts of electronic health records (EHRs) has attracted much interest. OBJECTIVE: This study aims to develop an automatic, quantitative stroke severity assessment framework through automating the entire NIHSS scoring process on Chinese clinical EHRs. METHODS: Our approach consists of two major parts: Chinese clinical named entity recognition (CNER) with a domain-adaptive pre-trained large language model (LLM) and automated NIHSS scoring. To build a high-performing CNER model, we first construct a stroke-specific, densely annotated dataset "Chinese Stroke Clinical Records" (CSCR) from EHRs provided by our partner hospital, based on a stroke ontology that defines semantically related entities for stroke assessment. We then pre-train a Chinese clinical LLM coined "CliRoberta" through domain-adaptive transfer learning and construct a deep learning-based CNER model that can accurately extract entities directly from Chinese EHRs. Finally, an automated, end-to-end NIHSS scoring pipeline is proposed by mapping the extracted entities to relevant NIHSS items and values, to quantitatively assess the stroke severity. RESULTS: Results obtained on a benchmark dataset CCKS2019 and our newly created CSCR dataset demonstrate the superior performance of our domain-adaptive pre-trained LLM and the CNER model, compared with the existing benchmark LLMs and CNER models. The high F1 score of 0.990 ensures the reliability of our model in accurately extracting the entities for the subsequent automatic NIHSS scoring. Subsequently, our automated, end-to-end NIHSS scoring approach achieved excellent inter-rater agreement (0.823) and intraclass consistency (0.986) with the ground truth and significantly reduced the processing time from minutes to a few seconds. CONCLUSION: Our proposed automatic and quantitative framework for assessing stroke severity demonstrates exceptional performance and reliability through directly scoring the NIHSS from diagnostic notes in Chinese clinical EHRs. Moreover, this study also contributes a new clinical dataset, a pre-trained clinical LLM, and an effective deep learning-based CNER model. The deployment of these advanced algorithms can improve the accuracy and efficiency of clinical assessment, and help improve the quality, affordability and productivity of healthcare services. Zhanzhong Gu, Xiangjian He, Ping Yu 0004, Wenjing Jia, Xiguang Yang, Penghui Hu, Shiyan Chen, Yiguang Lin |
Artif. Intell. Medicine | 2 |
| 2024 | MTDiff: Visual anomaly detection with multi-scale diffusion models
Wenju Li, Xiangjian He |
Knowl. Based Syst. | 3 |
| 2024 | WBNet: Weakly-supervised salient object detection via scribble and pseudo-background priorsabstractWeakly supervised salient object detection (WSOD) methods endeavor to boost sparse labels to get more salient cues in various ways. Among them, an effective approach is using pseudo labels from multiple unsupervised self-learning methods, but inaccurate and inconsistent pseudo labels could ultimately lead to detection performance degradation. To tackle this problem, we develop a new multi-source WSOD framework, WBNet, that can effectively utilize pseudo-background (non-salient region) labels combined with scribble labels to obtain more accurate salient features. We first design a comprehensive salient pseudo-mask generator from multiple self-learning features. Then, we pioneer the exploration of generating salient pseudo-labels via point-prompted and box-prompted Segment-Anything Models (SAM). Then, WBNet leverages a pixel-level Feature Aggregation Module (FAM), a mask-level Transformer-decoder (TFD), and an auxiliary Boundary Prediction Module (EPM) with a hybrid loss function to handle complex saliency detection tasks. Comprehensively evaluated with state-of-the-art methods on five widely used datasets, the proposed method significantly improves saliency detection performance. The code and results are publicly available at https://github.com/yiwangtz/WBNet. Yi Wang 0037, Ruili Wang 0001, Xiangjian He, Chi Lin 0001, Tianzhu Wang, Qi Jia 0001, Xin Fan 0001 |
Pattern Recognit. | 3 |
| 2024 | CARD: Semantic Segmentation With Efficient Class-Aware Regularized DecoderabstractSemantic segmentation has recently achieved notable advances by exploiting “class-level” contextual information during learning, e.g., the Object Contextual Representation (OCR) and Context Prior (CPNet) approaches. However, these approaches simply concatenate class-level information to pixel features to boost pixel representation learning, which cannot fully utilize intra-class and inter-class contextual information. Moreover, these approaches learn soft class centers based on coarse mask prediction, which is prone to error accumulation. To better exploit class-level information, we propose a universal Class-Aware Regularization (CAR) approach to optimize the intra-class variance and inter-class distance during feature learning, motivated by the fact that humans can recognize an object by itself no matter which other objects it appears with. Moreover, we design a dedicated decoder for CAR (named CARD), which consists of a novel spatial token mixer and an upsampling module, to maximize its gain for existing baselines while being highly efficient in terms of computational cost. Specifically, CAR consists of three novel loss functions. The first loss function encourages more compact class representations within each class, the second directly maximizes the distance between different class centers, and the third further pushes the distance between inter-class centers and pixels. Furthermore, the class center in our approach is directly generated from ground truth instead of from the error-prone coarse prediction. CAR can be directly applied to most existing segmentation models during training, including OCR and CPNet, and can largely improve their accuracy at no additional inference overhead. Extensive experiments and ablation studies conducted on multiple benchmark datasets demonstrate that the proposed CAR can boost the accuracy of all baseline models by up to 2.23% mIOU with superior generalization ability. CARD outperforms state-of-the-art approaches on multiple benchmarks with a highly efficient architecture. The code will be available at https://github.com/edwardyehuang/CAR. Liang Chen 0026, Wenjing Jia, Xiangjian He, Lixin Duan, Xuefei Zhe, Linchao Bao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Point Clouds are Specialized Images: A Knowledge Transfer Approach for 3D UnderstandingabstractSelf-supervised representation learning (SSRL) has gained increasing attention in point cloud understanding, in addressing the challenges posed by 3D data scarcity and high annotation costs. This paper presents PCExpert, a novel SSRL approach that reinterprets point clouds as “specialized images”. This conceptual shift allows PCExpert to leverage knowledge derived from large-scale image modality in a more direct and deeper manner, via extensively sharing the parameters with a pre-trained image encoder in a multi-way Transformer architecture. The parameter sharing strategy, combined with an additional pretext task for pre-training, i.e., transformation estimation, empowers PCExpert to outperform the state of the arts in a variety of tasks, with a remarkable reduction in the number of trainable parameters. Notably, PCExpert's performance underLINEARfine-tuning (e.g., yielding a 90.02% overall accuracy on ScanObjectNN) has already closely approximated the results obtained withFULLmodel fine-tuning (92.66%), demonstrating its effective representation capability. Jiachen Kang, Wenjing Jia, Xiangjian He, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Why Zero Trust Framework Adoption has Emerged During and After Covid-19 Pandemic
Abeer Z. Alalmaie, Priyadarsi Nanda, Xiangjian He, Mohrah Saad Alayan |
AINA (3) | 3 |
| 2023 | Pixels, Regions, and Objects: Multiple Enhancement for Salient Object DetectionabstractSalient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in complex scenes with multiple objects and background clutter. To address this issue, we propose a novel approach called Multiple Enhancement Network (MENet) that adopts the boundary sensibility, content integrity, iterative refinement, and frequency decomposition mechanisms of HVS. A multi-level hybrid loss is firstly designed to guide the network to learn pixel-level, region-level, and object-level features. A flexible multiscale feature enhancement module (ME-Module) is then designed to gradually aggregate and refine global or detailed features by changing the size order of the input feature sequence. An iterative training strategy is used to enhance boundary features and adaptive features in the dual-branch decoder of MENet. Comprehensive evaluations on six challenging benchmark datasets show that MENet achieves state-of-the-art results. Both the codes and results are publicly available at https://github.com/yiwangtz/MENet. Yi Wang 0037, Ruili Wang 0001, Xin Fan 0001, Tianzhu Wang, Xiangjian He |
CVPR | 5 |
| 2023 | ZT-NIDS: Zero Trust, Network Intrusion Detection System
Abeer Z. Alalmaie, Priyadarsi Nanda, Xiangjian He |
SECRYPT | 3 |
| 2023 | ABUSDet: A Novel 2.5D deep learning model for automated breast ultrasound tumor detection
Xudong Song, Xiaoyang Lu, Gengfa Fang, Xiangjian He, Xiaochen Fan, Le Cai, Wenjing Jia |
Appl. Intell. | 4 |
| 2023 | Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision
Jiachen Kang, Wenjing Jia, Xiangjian He |
Neurocomputing | 3 |
| 2023 | An Optimized IoT-Enabled Big Data Analytics Architecture for Edge-Cloud ComputingabstractThe awareness of edge computing is attaining eminence and is largely acknowledged with the rise of Internet of Things (IoT). Edge-enabled solutions offer efficient computing and control at the network edge to resolve the scalability and latency-related concerns. Though, it comes to be challenging for edge computing to tackle diverse applications of IoT as they produce massive heterogeneous data. The IoT-enabled frameworks for Big Data analytics face numerous challenges in their existing structural design, for instance, the high volume of data storage and processing, data heterogeneity, and processing time among others. Moreover, the existing proposals lack effective parallel data loading and robust mechanisms for handling communication overhead. To address these challenges, we propose an optimized IoT-enabled big data analytics architecture for edge-cloud computing using machine learning. In the proposed scheme, an edge intelligence module is introduced to process and store the big data efficiently at the edges of the network with the integration of cloud technology. The proposed scheme is composed of two layers: IoT-edge and Cloud-processing. The data injection and storage is carried out with an optimized MapReduce parallel algorithm. Optimized Yet Another Resource Negotiator (YARN) is used for efficiently managing the cluster. The proposed data design is experimentally simulated with an authentic dataset using Apache Spark. The comparative analysis is decorated with existing proposals and traditional mechanisms. The results justify the efficiency of our proposed work. Muhammad Babar 0001, Mian Ahmad Jan, Xiangjian He, Muhammad Usman Tariq, Spyridon Mastorakis, Ryan Alturki |
IEEE Internet Things J. | 3 |
| 2023 | FCH, an incentive framework for data-owner dominated federated learning
Priyadarsi Nanda, Christy Jie Liang, Xiangjian He |
J. Inf. Secur. Appl. | 4 |
| 2023 | Cross-domain learning for underwater image enhancement
Fei Li 0030, Jiangbin Zheng 0001, Yuan-fang Zhang, Wenjing Jia, Qianru Wei, Xiangjian He |
Signal Process. Image Commun. | 6 |
| 2023 | Self-learning and explainable deep learning network toward the security of artificial intelligence of things
Xiangjian He |
J. Supercomput. | 2 |
| 2023 | Arbitrary-Shape Scene Text Detection via Visual-Relational Rectification and Contour ApproximationabstractOne trend in the latest bottom-up approaches for arbitrary-shape scene text detection is to determine the links between text segments using Graph Convolutional Networks (GCNs). However, the performance of these bottom-up methods is still inferior to that of state-of-the-art top-down methods even with the help of GCNs. We argue that a cause of this is that bottom-up methods fail to make proper use of visual-relational features, which results in accumulated false detection, as well as the error-prone route-finding used for grouping text segments. In this paper, we improve classic bottom-up text detection frameworks by fusing the visual-relational features of text with two effective false positive/negative suppression (FPNS) mechanisms and developing a new shape-approximation strategy. First, dense overlapping text segments depicting the “characterness” and “streamline” properties of text are constructed and used in weakly supervised node classification to filter the falsely detected text segments. Then, relational features and visual features of text segments are fused with a novel Location-Aware Transfer (LAT) module and Fuse Decoding (FD) module to jointly rectify the detected text segments. Finally, a novel multiple-text-map-aware contour-approximation strategy is developed based on the rectified text segments, instead of the error-prone route-finding process, to generate the final contour of the detected text. Experiments conducted on five benchmark datasets demonstrate that our method outperforms the state-of-the-art performance when embedded in a classic text detection framework, which revitalizes the strengths of bottom-up methods. Chengpei Xu, Wenjing Jia, Tingcheng Cui, Ruomei Wang 0001, Yuan-fang Zhang, Xiangjian He |
IEEE Trans. Multim. | 6 |
| 2023 | MorphText: Deep Morphology Regularized Accurate Arbitrary-Shape Scene Text DetectionabstractBottom-up text detection methods play an important role in arbitrary-shape scene text detection but there are two restrictions preventing them from achieving their great potential, i.e., 1) the accumulation of false text segment detections, which affects subsequent processing, and 2) the difficulty of building reliable connections between text segments. Targeting these two problems, we propose a novel approach, named ``MorphText", to capture the regularity of texts by embedding deep morphology for arbitrary-shape text detection. Towards this end, two deep morphological modules are designed to regularize text segments and determine the linkage between them. First, a Deep Morphological Opening (DMOP) module is constructed to remove false text segment detections generated in the feature extraction process. Then, a Deep Morphological Closing (DMCL) module is proposed to allow text instances of various shapes to stretch their morphology along their most significant orientation while deriving their connections.Extensive experiments conducted on four challenging benchmark datasets (CTW1500, Total-Text, MSRA-TD500 and ICDAR2017) demonstrate that our proposed MorphText outperforms both top-down and bottom-up state-of-the-art arbitrary-shape scene text detection approaches. Chengpei Xu, Wenjing Jia, Ruomei Wang 0001, Xiangjian He |
IEEE Trans. Multim. | 5 |
| 2022 | Channelized Axial Attention - considering Channel Relation within Spatial Attention for Semantic SegmentationabstractSpatial and channel attentions, modelling the semantic interdependencies in spatial and channel dimensions respectively, have recently been widely used for semantic segmentation. However, computing spatial and channel attentions separately sometimes causes errors, especially for those difficult cases. In this paper, we propose Channelized Axial Attention (CAA) to seamlessly integrate channel attention and spatial attention into a single operation with negligible computation overhead. Specifically, we break down the dot-product operation of the spatial attention into two parts and insert channel relation in between, allowing for independently optimized channel attention on each spatial location. We further develop grouped vectorization, which allows our model to run with very little memory consumption without slowing down the running speed. Comparative experiments conducted on multiple benchmark datasets, including Cityscapes, PASCAL Context, and COCO-Stuff, demonstrate that our CAA outperforms many state-of-the-art segmentation models (including dual attention) on all tested datasets. Wenjing Jia, Liu Liu 0014, Xiangjian He |
AAAI | 5 |
| 2022 | CAR: Class-Aware Regularizations for Semantic Segmentation
Liang Chen 0026, Xuefei Zhe, Wenjing Jia, Linchao Bao, Xiangjian He |
ECCV (28) | 7 |
| 2022 | The Force of Compensation, a Multi-stage Incentive Mechanism Model for Federated Learning
Priyadarsi Nanda, Christy Jie Liang, Xiangjian He |
NSS | 4 |
| 2022 | Zero Trust-NIDS: Extended Multi-View Approach for Network Trace Anonymization and Auto-Encoder CNN for Network Intrusion DetectionabstractAs the enterprise networks are being constantly targeted by sophisticated cyber threats, Zero Trust Security has been suggested to address existing threats. Zero Trust Security models have been recently proposed for outsourcing network security monitoring to third-party analysts. Therefore, the current trends of security monitoring needs to shift to "Never Trust, Always Verify". There are no concerns about analysis accuracy, if a zero trust model is resistant against security attacks. In this paper, a modified multi-view approach is proposed to preserve privacy in network traces, emphasizing the challenges needed to be tackled. We then extend the multi-view approach for the features that are not in the known list of the analyzer and extend the partitioning methods to a more balanced approach. In addition, in order to send any data to the analyzer, we propose to use an Auto-Encoder Convolutional Neural Network, which has the ability to receive any type of input attributes for detecting intrusive behavior. Our proposed multi-view approach outperforms existing works and improves efficiency by improving indistinguishability and preserving privacy for any attributes. The proposed Intrusion Detection System also outperforms existing works by up to 1% higher accuracy without any need for feature engineering. Abeer Z. Alalmaie, Priyadarsi Nanda, Xiangjian He |
TrustCom | 3 |
| 2022 | An Empirical Assessment of Security and Privacy Risks of Web-Based Chatbots
Nazar Waheed, Muhammad Ikram 0001, Saad Sajid Hashmi, Xiangjian He, Priyadarsi Nanda |
WISE | 4 |
| 2022 | Secure and Reliable Indoor Localization Based on Multitask Collaborative Learning for Large-Scale BuildingsabstractAccurate and reliable indoor location estimate is crucial for many Internet-of-Things (IoT) applications in the era of smart buildings. However, the positioning accuracy and security of the existing positioning works cannot meet the demands in the large-scale smart buildings scenarios covering multiple multifloor buildings. Therefore, in this article, we focus on the reliable and accurate localization under multibuilding and multifloor environments. We propose two novel designs, including a two-step reliable feature selector and a multitask collaborative positioning model. First, we design a two-step reliable feature selector based on an access point (AP) confidence model and manifold learning, to help select the most representative and reliable fingerprint features. Second, we propose a multitask cooperative positioning model, which consists of a multiscale feature fusion module to adaptively fuse multiscale features and a multitask joint learning module to effectively constrain the cumulative error of multiscale position. Finally, based on the above two, we propose a reliable multibuilding and multifloor localization method (RMBMFL), which can achieve accurate and reliable location estimates with low computational complexity in a smart building complex. We did real-world experiments in a 20 000${m^{2}}$site that covers three multistory buildings to evaluate the performance of the proposed RMBMFL. The experimental results show that RMBMFL achieves a building identification accuracy and a floor identification accuracy of 99%, and a room-level indoor localization with an average positioning error within 2 m, and outperforms state-of-the-art solutions. Juan Luo, Xuan Liu 0001, Xiangjian He |
IEEE Internet Things J. | 4 |
| 2022 | EdgeLoc: A Robust and Real-Time Localization System Toward Heterogeneous IoT DevicesabstractIndoor localization has become an essential demand driven by indoor location-based services (ILBSs) for mobile users. With the rising of Internet of Things (IoT), heterogeneous smartphones and wearables have become ubiquitous. However, the ILBSs for heterogeneous IoT devices confront significant challenges, such as received signal strength (RSS) variances caused by hardware heterogeneity, multipath reflections from complex environments, and localization time restricted by computation resources. This article proposes EdgeLoc, a robust and real-time indoor localization system toward heterogeneous IoT devices to solve the above challenges. In particular, the RSS fingerprinting data of Wi-Fi is employed for localization and tackling the heterogeneity of IoT devices in twofold. First, feature-level and signal-level solutions are presented to address the random RSS variances. At the feature level, this work proposes a novel capsule neural network model to efficiently extract incremental features from RSS fingerprinting data. At the signal level, a multistep dataflow is further devised to process RSS fingerprints into image-like data, which utilizes the feature matrix to reduce absolute sensing errors introduced by hardware heterogeneity. Second, an edge-IoT framework is designed to utilize the edge server to train the deep learning model and further supports real-time localization for heterogeneous IoT devices. Extensive field experiments with over 33 600 data points are conducted to validate the effectiveness of EdgeLoc with a large-scale Wi-Fi fingerprint data set. The results show that EdgeLoc outperforms the state-of-the-art SAE-CNN method in localization accuracy by up to 14.4%, with an average error of 0.68 m and an average positioning time of 2.05 ms. Qianwen Ye, Hongxia Bie, Kuanching Li, Xiaochen Fan, Liangyi Gong, Xiangjian He, Gengfa Fang |
IEEE Internet Things J. | 6 |
| 2022 | Deep RGB-D Saliency Detection Without DepthabstractThe existing saliency detection models based on RGB colors only leverage appearance cues to detect salient objects. Depth information also plays a very important role in visual saliency detection and can supply complementary cues for saliency detection. Although many RGB-D saliency models have been proposed, they require to acquire depth data, which is expensive and not easy to get. In this paper, we propose to estimate depth information from monocular RGB images and leverage the intermediate depth features to enhance the saliency detection performance in a deep neural network framework. Specifically, we first use an encoder network to extract common features from each RGB image and then build two decoder networks for depth estimation and saliency detection, respectively. The depth decoder features can be fused with the RGB saliency features to enhance their capability. Furthermore, we also propose a novel dense multiscale fusion model to densely fuse multiscale depth and RGB features based on the dense ASPP model. A new global context branch is also added to boost the multiscale features. Experimental results demonstrate that the added depth cues and the proposed fusion model can both improve the saliency detection performance. Finally, our model not only outperforms state-of-the-art RGB saliency models, but also achieves comparable results compared with state-of-the-art RGB-D saliency models. Yuan-fang Zhang, Jiangbin Zheng 0001, Wenjing Jia, Wenfeng Huang, Long Li 0008, Nian Liu 0002, Fei Li 0030, Xiangjian He |
IEEE Trans. Multim. | 8 |
| 2021 | Synthetic CT images for semi-sequential detection and segmentation of lung nodules
Mohammad Hesam Hesamian, Wenjing Jia, Xiangjian He, Paul J. Kennedy |
Appl. Intell. | 3 |
| 2021 | Nighttime image dehazing based on Retinex and dark channel prior using Taylor series expansion
Qunfang Tang, Jie Yang 0022, Xiangjian He, Wenjing Jia, Qingnian Zhang |
Comput. Vis. Image Underst. | 3 |
| 2021 | Face hallucination based on cluster consistent dictionary learningabstractAbstract Face hallucination is a super‐resolution technique specially designed to reconstruct high‐resolution faces from low‐resolution faces. Most state‐of‐the‐art algorithms leverage position‐patch prior knowledge of human faces to better super‐resolve face images. However, most of them assume the training face dataset is sufficiently large, well cropped or aligned. This paper, proposes a novel example‐based face hallucination method, based on cluster consistent dictionary learning with the assumption that human faces have similar facial structures. In this method, the paired face image patches are firstly labelled as face areas including eyes, nose, mouth and other parts, as well as non‐face areas without requiring the training face images cropped and aligned. Then, the training patches are clustered according their labels and textures. The cluster consistent dictionary is learned to represent the low‐resolution patches and the high‐resolution patches. Finally, the high‐resolution patches of the input low‐resolution face image can be efficiently generated by using the adjusted anchored neighbourhood regression. As utilizing the labelled facial parts prior knowledge, the proposed method represents more details in the reconstruction. Experimental results demonstrate that the authors' algorithm outperforms many state‐of‐the‐art techniques for face hallucination under different datasets. Minqi Li, Xiangjian He, Kin-Man Lam 0001, Kaibing Zhang, Junfeng Jing |
IET Image Process. | 2 |
| 2021 | PDANet: Pyramid density-aware attention based network for accurate crowd counting
Saeed Amirgholipour Kasmani, Wenjing Jia, Lei Liu 0036, Xiaochen Fan, Dadong Wang, Xiangjian He |
Neurocomputing | 6 |
| 2021 | See more than once: Kernel-sharing atrous convolution for semantic segmentation
Wenjing Jia, Yue Lu 0001, Xiangjian He |
Neurocomputing | 6 |
| 2021 | Rethinking feature aggregation for deep RGB-D salient object detection
Yuanfang Zhang, Jiangbin Zheng 0001, Long Li 0008, Nian Liu 0002, Wenjing Jia, Xiaochen Fan, Chengpei Xu, Xiangjian He |
Neurocomputing | 8 |
| 2021 | Anomaly3D: Video anomaly detection based on 3D-normality clusters
Mujtaba Asad, Jie Yang 0002, Enmei Tu, Liming Chen 0001, Xiangjian He |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Editorial: Machine Learning and Big Data Analytics for IoT-Enabled Smart Cities
Mian Ahmad Jan, Xiangjian He, Houbing Song, Muhammad Babar 0001 |
Mob. Networks Appl. | 2 |
| 2021 | Novelty Detection and Online Learning for Chunk Data StreamsabstractDatastream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while incrementally updating the model efficiently and stably, especially for high-dimensional and/or large-scale data streams. This paper proposes an efficient framework for novelty detection and incremental learning for unlabeled chunk data streams. First, an accurate factorization-free kernel discriminative analysis (FKDA-X) is put forward through solving a linear system in the kernel space. FKDA-X produces a Reproducing Kernel Hilbert Space (RKHS), in which unlabeled chunk data can be detected and classified by multiple known-classes in a single decision model with a deterministic classification boundary. Moreover, based on FKDA-X, two optimal methods FKDA-CX and FKDA-C are proposed. FKDA-CX uses the micro-cluster centers of original data as the input to achieve excellent performance in novelty detection. FKDA-C and incremental FKDA-C (IFKDA-C) using the class centers of original data as their input have extremely fast speed in online learning. Theoretical analysis and experimental validation on under-sampled and large-scale real-world datasets demonstrate that the proposed algorithms make it possible to learn unlabeled chunk data streams with significantly lower computational costs and comparable accuracies than the state-of-the-art approaches. Yi Wang 0037, Xiangjian He, Xin Fan 0001, Chi Lin 0001, Fengqi Li, Tianzhu Wang, Zhongxuan Luo, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Enabling Privacy-Preserving Shortest Distance Queries on Encrypted Graph DataabstractWhen coming to perform shortest distance queries on encrypted graph data outsourced in external storage infrastructure such as cloud, a significant challenge is how to compute the shortest distance in an accurate, efficient and secure way. This issue is addressed by a recent work, which makes use of somewhat homomorphic encryption (SWHE) to encrypt distance values output by a 2-hop cover labeling (2HCL) scheme. However, it may import large errors and even yield negative results. Besides, SWHE would be too inefficient for normal clients. In this paper, we propose GENOA, a novel Graph ENcryption scheme for shOrtest distAnce queries. GENOA employs only efficient symmetric-key primitives while significantly enhances the accuracy compared to the prior work. As a reasonable trade-off, it additionally reveals the order information among queried distance values in the 2HCL index. We theoretically prove the accuracy and security of GENOA under rigorous cryptographic model. Detailed experiments on eight real-world graphs demonstrate that GENOA is efficient and can produce almost exact results. Chang Liu 0001, Liehuang Zhu, Xiangjian He, Jinjun Chen |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | BuildSenSys: Reusing Building Sensing Data for Traffic Prediction With Cross-Domain LearningabstractWith the rapid development of smart cities, smart buildings are generating a massive amount of building sensing data by the equipped sensors. Indeed, building sensing data provides a promising way to enrich a series of data-demanding and cost-expensive urban mobile applications. In this paper, as a preliminary exploration, we study how to reuse building sensing data to predict traffic volume on nearby roads. Compared with existing studies, reusing building sensing data has considerable merits of cost-efficiency and high-reliability. Nevertheless, it is non-trivial to achieve accurate prediction on such cross-domain data with two major challenges. First, relationships between building sensing data and traffic data are not unknown as prior, and the spatio-temporal complexities impose more difficulties to uncover the underlying reasons behind the above relationships. Second, it is even more daunting to accurately predict traffic volume with dynamic building-traffic correlations, which are cross-domain, non-linear, and time-varying. To address the above challenges, we design and implement BuildSenSys, a first-of-its-kind system for nearby traffic volume prediction by reusing building sensing data. Our work consists of two parts, i.e., Correlation Analysis and Cross-domain Learning. First, we conduct a comprehensive building-traffic analysis based on multi-source datasets, disclosing how and why building sensing data is correlated with nearby traffic volume. Second, we propose a novel recurrent neural network for traffic volume prediction based on cross-domain learning with two attention mechanisms. Specifically, a cross-domain attention mechanism captures the building-traffic correlations and adaptively extracts the most relevant building sensing data at each predicting step. Then, a temporal attention mechanism is employed to model the temporal dependencies of data across historical time intervals. The extensive experimental studies demonstrate that BuildSenSys outperforms all baseline methods with up to 65.3 percent accuracy improvement (e.g., 2.2 percent MAPE) in predicting nearby traffic volume. We believe that this work can open a new gate of reusing building sensing data for urban traffic sensing, thus establishing connections between smart buildings and intelligent transportation. Xiaochen Fan, Chaocan Xiang, Chao Chen 0004, Panlong Yang, Liangyi Gong, Xudong Song, Priyadarsi Nanda, Xiangjian He |
IEEE Trans. Mob. Comput. | 8 |
| 2021 | DENet: A Universal Network for Counting Crowd With Varying Densities and ScalesabstractCounting people or objects with significantly varying scales and densities has attracted much interest from the research community and yet it remains an open problem. In this paper, we propose a simple but efficient and effective network, named DENet, which is composed of two components,i.e., a detection network (DNet) and an encoder-decoder estimation network (ENet). We first run the DNet on the input image to detect and count individuals who can be segmented clearly. Then, the ENet is utilized to estimate the density maps of the remaining areas, typically with low resolution and high densities where individuals cannot be detected. For this purpose, we propose a modified Xception network as the encoder for feature extraction and a combination of dilated convolution and transposed convolution as the decoder. When evaluated on the ShanghaiTech Part A, UCF and WorldExpo’10 datasets, our DENet has achieved lower Mean Absolute Error (MAE) than those of the state-of-the-art methods. Lei Liu 0036, Jie Jiang 0005, Wenjing Jia, Saeed Amirgholipour Kasmani, Yi Wang 0037, Michelle Zeibots, Xiangjian He |
IEEE Trans. Multim. | 7 |
| 2021 | Multi-frame feature-fusion-based model for violence detection
Mujtaba Asad, Jie Yang 0002, Pourya Shamsolmoali, Xiangjian He |
Vis. Comput. | 5 |
| 2021 | A New Algorithm for Sketch-Based Fashion Image Retrieval Based on Cross-Domain TransformationabstractDue to the rise of e‐commerce platforms, online shopping has become a trend. However, the current mainstream retrieval methods are still limited to using text or exemplar images as input. For huge commodity databases, it remains a long‐standing unsolved problem for users to find the interested products quickly. Different from the traditional text‐based and exemplar‐based image retrieval techniques, sketch‐based image retrieval (SBIR) provides a more intuitive and natural way for users to specify their search need. Due to the large cross‐domain discrepancy between the free‐hand sketch and fashion images, retrieving fashion images by sketches is a significantly challenging task. In this work, we propose a new algorithm for sketch‐based fashion image retrieval based on cross‐domain transformation. In our approach, the sketch and photo are first transformed into the same domain. Then, the sketch domain similarity and the photo domain similarity are calculated, respectively, and fused to improve the retrieval accuracy of fashion images. Moreover, the existing fashion image datasets mostly contain photos only and rarely contain the sketch‐photo pairs. Thus, we contribute a fine‐grained sketch‐based fashion image retrieval dataset, which includes 36,074 sketch‐photo pairs. Specifically, when retrieving on our Fashion Image dataset, the accuracy of our model ranks the correct match at the top‐1 which is 96.6%, 92.1%, 91.0%, and 90.5% for clothes, pants, skirts, and shoes, respectively. Extensive experiments conducted on our dataset and two fine‐grained instance‐level datasets, i.e., QMUL‐shoes and QMUL‐chairs, show that our model has achieved a better performance than other existing methods. Hao-Peng Lei, Mingwen Wang 0001, Xiangjian He, Wenjing Jia |
Wirel. Commun. Mob. Comput. | 4 |
| 2021 | Binarized graph neural network
Hanchen Wang 0001, Defu Lian, Ying Zhang 0001, Lu Qin 0001, Xiangjian He, Yiguang Lin, Xuemin Lin 0001 |
World Wide Web | 5 |
| 2020 | Scale-Aware Rolling Fusion Network for Crowd CountingabstractDue to wide application prospects and various challenges such as large scale variation, inter-occlusion between crowd people and background noise, crowd counting is receiving increasing attention. In this paper, we propose a scale-aware rolling fusion network (SRF-Net) for crowd counting, which focuses on dealing with scale variation in highly congested noisy scenes. SRF-Net is a two-stage architecture that consists of a band-pass stage and a rolling guidance stage. Compared with the existing methods, SRF-Net achieves better results in retaining appropriate multi-level features and capturing multi-scale features, thus improving the quality of density estimation maps in crowded scenarios with large scale variation. We evaluate our method on three popular crowd counting datasets (ShanghaiTech, UCF_CC_50 and UCF-QNRF), and extensive experiments show its outperformance over the state-of-the-art approaches. Chengying Gao, Zhuo Su 0001, Xiangjian He |
ICME | 4 |
| 2020 | A Unified Host-based Intrusion Detection Framework using Spark in CloudabstractThe host-based intrusion detection system (HIDS) is an essential research domain of cybersecurity. HIDS examines log data of hosts to identify intrusive behaviors. The detection efficiency is a significant factor of HIDS. Traditionally, HIDS is often installed with a standalone mode. Training detection engines with a large amount of data on a single physical computer with limited computing resources may be time-consuming. Therefore, this paper offers a unified HIDS framework based on Spark and deployed in the Google cloud. The framework includes a unified machine learning pipeline to implement scalable and efficient HIDS. Ming Liu 0021, Zhi Xue, Xiangjian He |
TrustCom | 3 |
| 2020 | Security and Privacy Implementation in Smart Home: Attributes Based Access Control and Smart ContractsabstractThere has been wide range of applications involving smart home systems for user comfort and accessibility to essential commodities. Users enjoy featured home services supported by the IoT smart devices. These IoT devices are resource-constrained, incapable of securing themselves and can be easily hacked. Edge computing can provide localized computations and storage which can augment such capacity limitations for IoT devices. Furthermore, blockchain has emerged as technology with capabilities to provide secure access and authentication for IoT devices in decentralized manner. In this paper, we propose an authentication scheme which integrate attribute based access control using smart contracts with ERC-20 Token (Ethereum Request For Comments) and edge computing to construct a secure framework for IoT devices in Smart home system. The edge server provide scalability to the system by offloading heavier computation tasks to edge servers. We present system architecture and design and discuss various aspects related to testing and implementation of the smart contracts. We show that our proposed scheme is secure by thoroughly analysing its security goals with respect to confidentiality, integrity and availability. Finally, we conduct a performance evaluation to demonstrate the feasibility and efficiency of the proposed scheme. Amjad Qashlan, Priyadarsi Nanda, Xiangjian He |
TrustCom | 3 |
| 2020 | A PHP and JSP Web Shell Detection System With Text Processing Based On Machine LearningabstractWeb shell is one of the most common network attack methods, and traditional detection methods may not detect complex and flexible variants of web shell attacks. In this paper, we present a comprehensive detection system that can detect both PHP and JSP web shells. After file classification, we use different feature extraction methods, i.e. AST for PHP files and bytecode for JSP files. We present a detection model based on text processing methods including TF-IDF and Word2vec algorithms. We combine different kinds of machine learning algorithms and perform a comprehensively controlled experiment. After the experiment and evaluation, we choose the detection machine learning model of the best performance, which can achieve a high detection accuracy above 98%. Han Zhang 0027, Ming Liu 0021, Zihan Yue, Zhi Xue, Yong Shi 0009, Xiangjian He |
TrustCom | 6 |
| 2020 | Deep learning for intelligent traffic sensing and prediction: recent advances and future challenges
Xiaochen Fan, Chaocan Xiang, Liangyi Gong, Yuben Qu, Saeed Amirgholipour Kasmani, Priyadarsi Nanda, Xiangjian He |
CCF Trans. Pervasive Comput. Interact. | 9 |
| 2020 | FACLSTM: ConvLSTM with focused attention for scene text recognition
Wenjing Jia, Xiangjian He, Michael Blumenstein, Shujing Lyu, Yue Lu 0001 |
Sci. China Inf. Sci. | 4 |
| 2020 | Security, Trust and Privacy in Cyber (STPCyber): Future trends and challenges
Priyadarsi Nanda, Xiangjian He, Laurence T. Yang |
Future Gener. Comput. Syst. | 2 |
| 2020 | A hybrid encryption technique for Secure-GLOR: The adaptive secure routing protocol for dynamic wireless mesh networks
Ashish Nanda, Priyadarsi Nanda, Xiangjian He, Aruna Jamdagni, Deepak Puthal |
Future Gener. Comput. Syst. | 3 |
| 2020 | QASEC: A secured data communication scheme for mobile Ad-hoc networks
Muhammad Usman 0015, Mian Ahmad Jan, Xiangjian He, Priyadarsi Nanda |
Future Gener. Comput. Syst. | 3 |
| 2020 | Structural correlation filters combined with a Gaussian particle filter for hierarchical visual tracking
Manna Dai, Gao Xiao, Shuying Cheng, Dadong Wang, Xiangjian He |
Neurocomputing | 5 |
| 2020 | Beyond context: Exploring semantic similarity for small object detection in crowded scenes
Jiangbin Zheng 0001, Xiangjian He, Wenjing Jia, Yefan Xie, Mingchen Feng, Xiuxiu Li |
Pattern Recognit. Lett. | 3 |
| 2020 | Hybrid Tree-Rule Firewall for High Speed Data TransmissionabstractTraditional firewalls employ listed rules in both configuration and process phases to regulate network traffic. However, configuring a firewall with listed rules may create rule conflicts, and slows down the firewall. To overcome this problem, we have proposed a Tree-rule firewall in our previous study. Although the Tree-rule firewall guarantees no conflicts within its rule set and operates faster than traditional firewalls, keeping track of the state of network connections using hashing functions incurs extra computational overhead. In order to reduce this overhead, we propose a hybrid Tree-rule firewall in this paper. This hybrid scheme takes advantages of both Tree-rule firewalls and traditional listed-rule firewalls. The GUIs of our Tree-rule firewalls are utilized to provide a means for users to create conflict-free firewall rules, which are organized in a tree structure and called 'tree rules'. These tree rules are later converted into listed rules that share the merit of being conflict-free. Finally, in decision making, the listed rules are used to verify against packet header information. The rules which have matched with most packets are moved up to the top positions by the core firewall. The mechanism applied in this hybrid scheme can significantly improve the functional speed of a firewall. Thawatchai Chomsiri, Xiangjian He, Priyadarsi Nanda, Zhiyuan Tan 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2020 | A Distributed and Anonymous Data Collection Framework Based on Multilevel Edge Computing ArchitectureabstractIndustrial Internet of Things applications demand trustworthiness in terms of quality of service (QoS), security, and privacy, to support the smooth transmission of data. To address these challenges, in this article, we propose a distributed and anonymous data collection (DaaC) framework based on a multilevel edge computing architecture. This framework distributes captured data among multiple level-one edge devices (LOEDs) to improve the QoS and minimize packet drop and end-to-end delay. Mobile sinks are used to collect data from LOEDs and upload to cloud servers. Before data collection, the mobile sinks are registered with a level-two edge-device to protect the underlying network. The privacy of mobile sinks is preserved through group-based signed data collection requests. Experimental results show that our proposed framework improves QoS through distributed data transmission. It also helps in protecting the underlying network through a registration scheme and preserves the privacy of mobile sinks through group-based data collection requests. Muhammad Usman 0015, Mian Ahmad Jan, Alireza Jolfaei, Min Xu 0001, Xiangjian He, Jinjun Chen |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Vehicular networks with security and trust management solutions: proposed secured message exchange via blockchain technology
Nisha Malik, Priyadarsi Nanda, Xiangjian He, Ren Ping Liu 0001 |
Wirel. Networks | 3 |
| 2019 | A Novel Multi-Path Anonymous Randomized Key Distribution Scheme for Geo Distributed NetworksabstractA major concern in distributed networks is the ability to provide acceptable levels of security. This is achieved by using encryption and authentication mechanisms that depend on encryption keys. However, given the ever-expanding nature of the network, it is difficult to keep setting up authorities that can aid the key- exchange process. This paper presents a novel solution to the challenge of exchanging keys of a large, distributed network without the need to set up additional authorities. The key-exchange scheme presented takes advantage of features such as packet anonymity, random selection and a multi- path approach for the exchange process. The paper also discusses the effectiveness of the proposed scheme against various threat scenarios. Ashish Nanda, Priyadarsi Nanda, Mohammad S. Obaidat, Xiangjian He, Deepak Puthal |
GLOBECOM | 4 |
| 2019 | Atrous Convolution for Binary Semantic Segmentation of Lung NoduleabstractAccurately estimating the size of tumours and reproducing their boundaries from lung CT images provides crucial information for early diagnosis, staging and evaluating patients response to cancer therapy. This paper presents an advanced solution to segment lung nodules from CT images by employing a deep residual network structure with Atrous convolution. The Atrous convolution increases the field of view of the filters and helps to improve classification accuracy. Moreover, in order to address the significant class imbalance issue between the nodule pixels and background non-nodule pixels, a weighted loss function is proposed. We evaluate our proposed solution on the widely adopted benchmark dataset LIDC. A promising result of an average DCS of 81.24% is achieved, outperforming the state of the arts. This demonstrates the effectiveness and importance of applying the Atrous convolution and weighted loss for such problems. Mohammad Hesam Hesamian, Wenjing Jia, Xiangjian He, Paul J. Kennedy |
ICASSP | 3 |
| 2019 | DeepText: Detecting Text from the Wild with Multi-ASPP-Assembled DeepLababstractIn this paper, we address the issue of scene text detection in the way of direct regression and successfully adapt an effective semantic segmentation model, DeepLab v3+ [1], for this application. In order to handle texts with arbitrary orientations and sizes and improve the recall of small texts, we propose to extract features of multiple scales by inserting multiple Atrous Spatial Pyramid Pooling (ASPP) layers to the DeepLab after the feature maps with different resolutions. Then, we set multiple auxiliary IoU losses at the decoding stage and make auxiliary connections from the intermediate encoding layers to the decoder to assist network training and enhance the discrimination ability of lower encoding layers. Experiments conducted on the benchmark scene text dataset ICDAR2015 demonstrate the superior performance of our proposed network, named as DeepText, over the state-of-the-art approaches. Wenjing Jia, Xiangjian He, Yue Lu 0001, Michael Blumenstein, Shujing Lyu |
ICDAR | 3 |
| 2019 | Learning Transmission Filtering Network for Image-Based Pm2.5 EstimationabstractPM2.5 is an important indicator of the severity of air pollution and its level can be predicted through hazy photographs caused by its degradation. Image-based PM2.5 estimation is thus extensively employed in various multimedia applications but is challenging because of its ill-posed property. In this paper, we convert it to the problem of estimating the PM2.5-relevant haze transmission and propose a learning model called the transmission filtering network. Different from most methods that generate a transmission map directly from a hazy image, our model takes the coarse transmission map derived from the dark channel prior as the input. To obtain a transmission map that satisfies the local smoothness constraint without regional boundary degradation, our model performs the edge-preserving smoothing filtering as the refinement on the map. Moreover, we introduce the attention mechanism to the network architecture for more efficient feature extraction and smoothing effects in the transmission estimation. Experimental results prove that our model performs favorably against the state-of-the-art dehazing methods in a variety of hazy scenes. Yinghong Liao, Bin Qiu, Zhuo Su 0001, Ruomei Wang 0001, Xiangjian He |
ICME | 5 |
| 2019 | Residual Magnifier: A Dense Information Flow Network for Super ResolutionabstractRecently, deep learning methods have been successfully applied to single image super-resolution tasks. However, some networks with extreme depth failed to achieve better performance because of the insufficient utilization of the local residual information extracted at each stage. To solve the above question, we propose a Dense Information Flow Network (DIF-Net), which can fully extract and utilize the local residual information at each stage to accomplish a better reconstruction. Specifically, we present a Two-stage Residual Extraction Block (TREB) to extract the shallow and deep local residual information at each stage. The dense connection mechanism is introduced throughout the model and within TREBs to dramatically increase the information flow. Meanwhile this mechanism prevents the shallow features extracted earlier from being diluted. Finally, we propose a lightweight subnet (residual enhancer) to efficiently recycle the overflow residual information from the backbone net for detail enhancement of the residual image. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods with relatively-less parameters. Mengcheng Cheng, Zhuo Su 0001, Xiangjian He |
ICME | 5 |
| 2019 | SPFusionNet: Sketch Segmentation Using Multi-modal Data FusionabstractThe sketch segmentation problem remains largely unsolved because conventional methods are greatly challenged by the highly abstract appearances of freehand sketches and their numerous shape variations. In this work, we tackle such challenges by exploiting different modes of sketch data in a unified framework. Specifically, we propose a deep neural network SPFusionNet to capture the characteristic of sketch by fusing from its image and point set modes. The image modal component SketchNet learns hierarchically abstract ro-bust features and utilizes multi-level representations to produce pixel-wise feature maps, while the point set-modal component SPointNet captures local and global contexts of the sampled point set to produce point-wise feature maps. Then our framework aggregates these feature maps by a fusion network component to generate the sketch segmentation result. The extensive experimental evaluation and comparison with peer methods on our large SketchSeg dataset verify the effectiveness of the proposed framework. Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Xiangjian He |
ICME | 7 |
| 2019 | Feature Fusion Based Deep Spatiotemporal Model for Violence Detection in Videos
Mujtaba Asad, Zuopeng Yang, Zubair Khan, Jie Yang 0002, Xiangjian He |
ICONIP (1) | 5 |
| 2019 | High-Performance Light Field Reconstruction with Channel-wise and SAI-wise Attention
Zexi Hu, Vera Chung, Seid Miad Zandavi, Wanli Ouyang, Xiangjian He, Yuefang Gao |
ICONIP (5) | 5 |
| 2019 | Optimization of a Convolutional Neural Network Using a Hybrid AlgorithmabstractIn recent years, Convolutional Neural Networks (CNNs) have been widely used in image recognition due to their aptitude in large scale image processing. The CNN uses Back-propagation (BP) to train weights and biases, which in turn makes the error consistently smaller. The most common optimizers that uses a BP algorithm are Stochastic Gradient Decent (SGD), Adam, and Adadelta. These optimizers, however, have been proved to fall easily into the regional optimal solution. Little research has been conducted on the application of Soft Computing in CNN to fix the above problem, and most studies that have been conducted focus on Particle Swarm Optimization. Among them, the hybrid algorithm combined with SGD proposed by Albeahdili improves the image classification accuracy over that achieved by the original CNN. This study proposes the amalgamation of Improved Simplified Swarm Optimization (iSSO) with SGD, hence culminating in the iSSO-SGD which is intended train CNNs more efficiently to establish a better prediction model and improve the classification accuracy. The performance of the proposed iSSO-SGD can be affirmed through a comparison with the PSO-SGD, the Adam, Adadelta, rmsprop and momentum optimizers and their abilities in improving the accuracy of image classification. Chia-Ling Huang, Yan-Chih Shih, Chyh-Ming Lai, Vera Chung, Wenbo Zhu 0001, Wei-Chang Yeh 0001, Xiangjian He |
IJCNN | 7 |
| 2019 | Testbed evaluation of Lightweight Authentication Protocol (LAUP) for 6LoWPAN wireless sensor networksabstractSummary 6LoWPAN networks involving wireless sensors consist of resource starving miniature sensor nodes. Since secured authentication is one of the important considerations, the use of asymmetric key distribution scheme may not be a perfect choice. Recent research shows that Lucky Thirteen attack has compromised Datagram Transport Layer Security (DTLS) with Cipher Block Chaining (CBC) mode for key establishment. Even though EAKES6Lo and S3 K techniques for key establishment follow the symmetric key establishment method, they strongly rely on a remote server and trust anchor. Our proposed Lightweight Authentication Protocol (LAUP) used a symmetric key method with no preshared keys and comprised of four flights to establish authentication and session key distribution between sensors and Edge Router in a 6LoWPAN environment. Each flight uses freshly derived keys from existing information such as PAN ID (Personal Area Network IDentification) and device identities. We formally verified our scheme using the Scyther security protocol verification tool. We simulated and evaluated the proposed LAUP protocol using COOJA simulator and achieved less computational time and low power consumption compared to existing authentication protocols such as the EAKES6Lo and SAKES. LAUP is evaluated using real‐time testbed and achieved less computational time, which is supportive of our simulated results. Annie Gilda Roselin, Priyadarsi Nanda, Surya Nepal, Xiangjian He |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Performance-enhancing network pruning for crowd counting
Lei Liu 0036, Saeed Amirgholipour Kasmani, Jie Jiang 0005, Wenjing Jia, Michelle Zeibots, Xiangjian He |
Neurocomputing | 6 |
| 2019 | SAMS: A Seamless and Authorized Multimedia Streaming Framework for WMSN-Based IoMTabstractAn Internet of Multimedia Things (IoMT) architecture aims to provide a support for real-time multimedia applications by using wireless multimedia sensor nodes that are deployed for a long-term usage. These nodes are capable of capturing both multimedia and nonmultimedia data, and form a network known as Wireless Multimedia Sensor Network (WMSN). In a WMSN, underlying routing protocols need to provide an acceptable level of Quality of Service (QoS) support for multimedia traffic. In this paper, we propose a Seamless and Authorized Streaming (SAMS) framework for a cluster-based hierarchical WMSN. The SAMS uses authentication at different levels to form secured clusters. The formation of these clusters allows only legitimate nodes to transmit captured data to their Cluster Heads (CHs). Each node senses the environment, stores captured data in its buffer, and waits for its turn to transmit to its CH. This waiting may result in an excessive packet-loss and end-to-end delay for multimedia traffic. To address these issues, a channel allocation approach is proposed for an intercluster communication. In the case of a buffer overflow, a member node in one cluster switches to a neighboring CH provided that the latter has an available channel for allocation. The experimental results show that the SAMS provides an acceptable level of QoS and enhances security of an underlying network. Mian Ahmad Jan, Muhammad Usman 0015, Xiangjian He, Ateeq Ur Rehman 0001 |
IEEE Internet Things J. | 3 |
| 2019 | Exploiting the Remote Server Access Support of CoAP ProtocolabstractThe constrained application protocol (CoAP) is a specially designed Web transfer protocol for use with constrained nodes and low-power networks. The widely available CoAP implementations have failed to validate the remote CoAP clients. Each CoAP client generates a random source port number when communicating with the CoAP server. However, we observe that in such implementations it is difficult to distinguish the regular packet and the malicious packet, opening a door for a potential off-path attack. The off-path attack is considered a weak attack on a constrained network and has received a less attention from the research community. However, the consequences resulting from such an attack cannot be ignored in practice. In this article, we exploit the combination of IP spoofing vulnerability and the remote server access support of CoAP is to be launch an off-path attack. The attacker injects a fake request message to change the credentials of the 6LoWPAN smart door keypad lock system. This creates a request spoofing vulnerability in CoAP, and the attacker exploits this vulnerability to gain full access to the system. Through our implementation, we demonstrated the feasibility of the attack scenario on the 6LoWPAN-CoAP network using smart door keypad lock. We proposed a machine learning (ML)-based approach to mitigate such attacks. To the best of our knowledge, we believe that this is the first article to analyze the remote CoAP server access support and request spoofing vulnerability of CoAP to launch an off-path attack and demonstrate how an ML-based approach can be deployed to prevent such attacks. Annie Gilda Roselin, Priyadarsi Nanda, Surya Nepal, Xiangjian He, Jarod Wright |
IEEE Internet Things J. | 4 |
| 2019 | On exploiting priority relation graph for reliable multi-path communication in mobile social networks
Limei Lin, Li Xu 0002, Yanze Huang, Yang Xiang 0001, Xiangjian He |
Inf. Sci. | 5 |
| 2019 | P2DCA: A Privacy-Preserving-Based Data Collection and Analysis Framework for IoMT ApplicationsabstractThe concept of Internet of Multimedia Things (IoMT) is becoming popular nowadays and can be used in various smart city applications, e.g., traffic management, healthcare, and surveillance. In the IoMT, the devices, e.g., Multimedia Sensor Nodes (MSNs), are capable of generating both multimedia and non-multimedia data. The generated data are forwarded to a cloud server via a Base Station (BS). However, it is possible that the Internet connection between the BS and the cloud server may be temporarily down. The limited computational resources restrict the MSNs from holding the captured data for a longer time. In this situation, mobile sinks can be utilized to collect data from MSNs and upload to the cloud server. However, this data collection may create privacy issues, such as revealing identities and location information of MSNs. Therefore, there is a need to preserve the privacy of MSNs during mobile data collection. In this paper, we propose an efficient privacy-preserving-based data collection and analysis (P2DCA) framework for IoMT applications. The proposed framework partitions an underlying wireless multimedia sensor network into multiple clusters. Each cluster is represented by a Cluster Head (CH). The CHs are responsible to protect the privacy of member MSNs through data and location coordinates aggregation. Later, the aggregated multimedia data are analyzed on the cloud server using a counter-propagation artificial neural network to extract meaningful information through segmentation. Experimental results show that the proposed framework outperforms the existing privacy-preserving schemes, and can be used to collect multimedia data in various IoMT applications. Muhammad Usman 0015, Mian Ahmad Jan, Xiangjian He, Jinjun Chen |
IEEE J. Sel. Areas Commun. | 3 |
| 2019 | Object tracking in the presence of shaking motions
Manna Dai, Shuying Cheng, Xiangjian He, Dadong Wang |
Neural Comput. Appl. | 3 |
| 2019 | Error Concealment for Cloud-Based and Scalable Video Coding of HD VideosabstractThe encoding of HD videos faces two challenges: requirements for a strong processing power and a large storage space. One time-efficient solution addressing these challenges is to use a cloud platform and to use a scalable video coding technique to generate multiple video streams with varying bit-rates. Packet-loss is very common during the transmission of these video streams over the Internet and becomes another challenge. One solution to address this challenge is to retransmit lost video packets, but this will create end-to-end delay. Therefore, it would be good if the problem of packet-loss can be dealt with at the user's side. In this paper, we present a novel system that encodes and stores the videos using the Amazon cloud computing platform, and recover lost video frames on user side using a new Error Concealment (EC) technique. To efficiently utilize the computation power of a user's mobile device, the EC is performed based on a multiple-thread and parallel process. The simulation results clearly show that, on average, our proposed EC technique outperforms the traditional Block Matching Algorithm (BMA) and the Frame Copy (FC) techniques. Muhammad Usman 0015, Xiangjian He, Kin-Man Lam 0001, Min Xu 0001, Syed Mohsin Matloob Bokhari, Jinjun Chen, Mian Ahmad Jan |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | CTOM: Collaborative Task Offloading Mechanism for Mobile Cloudlet NetworksabstractMobile cloud computing has emerged as a pervasive paradigm to execute computing tasks for capacity- limited mobile devices. More specifically, at the network edge, the resource-rich and trusted cloudlet system is acting as a 'data center in a box' to support compute-intensive mobile applications. The mobile cloudlets can provide in-proximity services by executing the workloads for nearby devices. Nevertheless, load balancing in mobile cloudlet network is of great importance, as it has a huge impact on task response time. Existing methods for cloudlet load balancing basically rely on the strategic placement or user cooperation. However, the above solutions require the global task load information from the whole network, which is costly in both communication and computation. To achieve more efficient and low-cost load balancing, we propose 'CTOM', a Collaborative Task Offloading Mechanism for mobile cloudlet networks. Our solution is based on the balls-and-bins theory and can balance the task load only requiring limited information. Extensive simulations and evaluation based on mobility trace demonstrate that, our CTOM outperforms the conventional random and proportional allocation schemes by reducing the task gaps among mobile cloudlets by 65% and 55% respectively. Meanwhile, CTOM's performance is close to that of the greedy algorithm but with much lower computing complexity. Xiaochen Fan, Xiangjian He, Deepak Puthal, Shiping Chen 0001, Chaocan Xiang, Priyadarsi Nanda, Xunpeng Rao |
ICC | 2 |
| 2018 | A-CCNN: Adaptive CCNN for Density Estimation and Crowd CountingabstractCrowd counting, for estimating the number of people in a crowd using vision-based computer techniques, has attracted much interest in the research community. Although many attempts have been reported, real-world problems, such as huge variation in subjects' sizes in images and serious occlusion among people, make it still a challenging problem. In this paper, we propose an Adaptive Counting Convolutional Neural Network (A-CCNN) and consider the scale variation of objects in a frame adaptively so as to improve the accuracy of counting. Our method takes advantages of contextual information to provide more accurate and adaptive density maps and crowd counting in a scene. Extensively experimental evaluation is conducted using different benchmark datasets for object-counting and shows that the proposed approach is effective and outperforms state-of-the-art approaches. Saeed Amirgholipour Kasmani, Xiangjian He, Wenjing Jia, Dadong Wang, Michelle Zeibots |
ICIP | 2 |
| 2018 | Beyond Context: Exploring Semantic Similarity for Tiny Face DetectionabstractTiny face detection aims to find faces with high degrees of variability in scale, resolution and occlusion in cluttered scenes. Due to the very little information available on tiny faces, it is not sufficient to detect them merely based on the information presented inside the tiny bounding boxes or their context. In this paper, we propose to exploit the semantic similarity among all predicted targets in each image to boost current face detectors. To this end, we present a novel framework to model semantic similarity as pairwise constraints within the metric learning scheme, and then refine our predictions with the semantic similarity by utilizing the graph cut techniques. Experiments conducted on three widely-used benchmark datasets have demonstrated the improvement over the-state-of-the-arts gained by applying this idea. Jiangbin Zheng 0001, Xiangjian He, Wenjing Jia |
ICIP | 3 |
| 2018 | Trusted Guidance Pyramid Network for Human ParsingabstractHuman parsing, which segments a human-centric image into pixel-wise categorization, has a wide range of applications. However, none of the existing methods can productively solve the issue of label parsing fragmentation due to confused and complicated annotations. In this paper, we propose a novel Trusted Guidance Pyramid Network (TGPNet) to address this limitation. Based on a pyramid architecture, we design a Pyramid Residual Pooling (PRP) module setting at the end of a bottom-up approach to capture both global and local level context. In the top-down approach, we propose a Trusted Guidance Multi-scale Supervision (TGMS) that efficiently integrates and supervises multi-scale contextual information. Furthermore, we present a simple yet powerful Trusted Guidance Framework (TGF) which imposes global-level semantics into parsing results directly without extra ground truth labels in model training. Extensive experiments on two public human parsing benchmarks well demonstrate that our TGPNet has a strong ability in solving label parsing fragmentation problem and has an obtained improvement than other methods. Xianghui Luo, Zhuo Su 0001, Jiaming Guo, Gengwei Zhang, Xiangjian He |
ACM Multimedia | 5 |
| 2018 | User Relationship Classification of Facebook Messenger Mobile Data using WEKA
Amber Umair, Priyadarsi Nanda, Xiangjian He, Kim-Kwang Raymond Choo |
NSS | 3 |
| 2018 | Violence Detection Based on Spatio-Temporal Feature and Fisher Vector
HuangKai Cai, Xiaolin Huang, Jie Yang 0002, Xiangjian He |
PRCV (1) | 5 |
| 2018 | A Sybil attack detection scheme for a forest wildfire monitoring application
Mian Ahmad Jan, Priyadarsi Nanda, Xiangjian He, Ren Ping Liu 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | SUDMAD: Sequential and unsupervised decomposition of a multi-author document based on a hidden markov modelabstractDecomposing a document written by more than one author into sentences based on authorship is of great significance due to the increasing demand for plagiarism detection, forensic analysis, civil law (i.e., disputed copyright issues), and intelligence issues that involve disputed anonymous documents. Among existing studies for document decomposition, some were limited by specific languages, according to topics or restricted to a document of two authors, and their accuracies have big room for improvement. In this paper, we consider the contextual correlation hidden among sentences and propose an algorithm for Sequential and Unsupervised Decomposition of a Multi‐Author Document (SUDMAD) written in any language, disregarding topics, through the construction of a Hidden Markov Model (HMM) reflecting the authors' writing styles. To build and learn such a model, an unsupervised, statistical approach is first proposed to estimate the initial values of HMM parameters of a preliminary model, which does not require the availability of any information of author's or document's context other than how many authors contributed to writing the document. To further boost the performance of this approach, a boosted HMM learning procedure is proposed next, where the initial classification results are used to create labeled training data to learn a more accurate HMM. Moreover, the contextual relationship among sentences is further utilized to refine the classification results. Our proposed approach is empirically evaluated on three benchmark datasets that are widely used for authorship analysis of documents. Comparisons with recent state‐of‐the‐art approaches are also presented to demonstrate the significance of our new ideas and the superior performance of our approach. Khaled Aldebei, Xiangjian He, Wenjing Jia, Wei-Chang Yeh 0001 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2018 | Editorial: Current and Future Trends in Wireless Communications Protocols and Technologies
Muhammad Alam 0002, Mian Ahmad Jan, Lei Shu 0001, Xiangjian He, Yuanfang Chen |
Mob. Networks Appl. | 4 |
| 2018 | Hybrid generative-discriminative hash tracking with spatio-temporal contextual cues
Manna Dai, Shuying Cheng, Xiangjian He |
Neural Comput. Appl. | 3 |
| 2018 | Performance evaluation of High Definition video streaming over Mobile Ad Hoc Networks
Muhammad Usman 0015, Mian Ahmad Jan, Xiangjian He, Muhammad Alam 0002 |
Signal Process. | 3 |
| 2018 | Guest Editorial Introduction to the Special Issue on Dependable Wireless Vehicular Communications for Intelligent Transportation Systems (ITS)abstractOver the past couple of decades, transportation systems have begun to receive widespread attention from the scientific community and emerged toward Intelligent Transportation Systems (ITS). Effective vehicular connectivity techniques can significantly enhance efficiency of travel, reduce traffic incidents and improve safety, and alleviate the impact of congestion; devising the ITS experience. Furthermore, during the past decades, the volume and density of vehicles increased significantly, especially the road traffic; this lead to a dramatic increase in the number of accidents and congestion, with negative impacts on the economy, environment, and quality of people’s lives. In particular, according to the World Health Organization (WHO), road traffic injuries are estimated to be the leading cause of death for young people aged 15–29 and the ninth cause of death worldwide in 2015. The enabling communication technologies are intended to realize the frameworks that will spur an array of applications and use cases in the domain of road safety, traffic efficiency, and driver’s assistance. Although these applications will allow the dissemination and gathering of useful information among vehicles and between transportation infrastructure and vehicles in pursuance of assisting drivers to travel safely and comfortably, much effort is required to implement these practices for the success of these applications. Muhammad Alam 0002, Ammar Rayes, Xiangjian He, Mohammed Atiquzzaman, Jaime Lloret Mauri, Kim Fung Tsang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | A Joint Framework for QoS and QoE for Video Transmission over Wireless Multimedia Sensor NetworksabstractWith the emergence of Wireless Multimedia Sensor Networks (WMSNs), the distribution of multimedia contents have now become a reality. Without proper management, the transmission of multimedia data over WMSNs affects the performance of networks due to excessive packet-drop. The existing studies on Quality of Service (QoS) mostly deal with simple Wireless Sensor Networks (WSNs) and as such do not account for an increasing number of sensor nodes and an increasing volume of data. In this paper, we propose a novel framework to support QoS in WMSNs along with a light-weight Error Concealment (EC) scheme. The EC schemes play a vital role to enhance Quality of Experience (QoE) by maintaining an acceptable quality at the receiving ends. The main objectives of the proposed framework are to maximize the network throughput and to cover-up the effects produced by dropped video packets. To control the data-rate, Scalable High efficiency Video Coding (SHVC) is applied at multimedia sensor nodes with variable Quantization Parameters (QPs). Multi-path routing is exploited to support real-time video transmission. Experimental results show that the proposed framework can efficiently adjust large volumes of video data under certain network distortions and can effectively conceal lost video frames by producing better objective measurements. Muhammad Usman 0015, Ning Yang 0003, Mian Ahmad Jan, Xiangjian He, Min Xu 0001, Kin-Man Lam 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2017 | A comparison of modified evolutionary computation algorithms with applications to three-dimensional endoscopic camera motion trackingabstractEndoscope 3D motion tracking plays an irreplaceable role for computer-assisted endoscopy systems development. Without such tracking, it is impossible to synchronize pre- and intraoperative images in a reference coordinate frame. Currently available methods are comprised of video-based and electromagnetic tracking. These methods limit to either video image artifacts or inaccurate sensor measurements and dynamic errors. This paper proposes two modified evolutionary computation algorithms: (a) adaptive particle swarm optimization (APSO) and (b) observation-boosted differential evolution (OBDE), to augment current endoscopic camera motion tracking. The experimental results demonstrate that our modified algorithms, which combine endoscopic video images with sensor measurements to estimate endoscope movements, can improve tracking accuracy from 4.8 mm to 2.9 mm. OBDE outperforms APSO for endoscope tracking. Xióngbiao Luó, Xiangjian He |
ICIP | 3 |
| 2017 | Multi-view pairwise relationship learning for sketch based 3D shape retrievalabstractRecent progress in sketch-based 3D shape retrieval creates a novel and user-friendly way to explore massive 3D shapes on the Internet. However, current methods on this topic rely on designing invariant features for both sketches and 3D shapes, or complex matching strategies. Therefore, they suffer from problems like arbitrary drawings and inconsistent viewpoints. To tackle this problem, we propose a probabilistic framework based on Multi-View Pairwise Relationship (MVPR) learning. Our framework includes multiple views of 3D shapes as the intermediate layer between sketches and 3D shapes, and transforms the original retrieval problem into the form of inferring pairwise relationship between sketches and views. We accomplish pairwise relationship inference by a novel MVPR net, which can automatically predict and merge the pairwise relationships between a sketch and multiple views, thus freeing us from exhaustively selecting the best view of 3D shapes. We also propose to learn robust features for sketches and views via fine-tuning pre-trained networks. Extensive experiments on a large dataset demonstrate that the proposed method can outperform state-of-the-art methods significantly. Hefeng Wu, Xiangjian He, Shujin Lin, Ruomei Wang 0001 |
ICME | 3 |
| 2017 | A framework for data security in cloud using collaborative intrusion detection schemeabstractCloud computing offers an on demand, elastic, global network access to a shared pool of resources that can be configured on user demand. The advantages of cloud computing are lucrative for well-established organizations looking to reduce infrastructure cost overheads. However, the users are not quite confident in entrusting their data to the cloud due to security threats and risks perceived in the cloud domain. Issues involving privacy requirements for the cloud and best practices in the cloud are suggested in this paper. Although the cloud provider ensures security in the cloud yet the flow of data, storage location, data computing process and security breaches are not transparent to the cloud customer. This distrust and lack of control on data is a major hindrance for potential cloud customers in adopting the cloud models for their businesses. Intrusion Detection Systems (IDSs) are widely used to detect malicious activities. However existing solutions with IDSs involving DDoS and other non-detectable events may not be suitable in applying to the cloud due to distributed data storage and a major shift in Internet access mechanisms offered by cloud providers. Hence there is a strong need to analyze an appropriate IDS to counter DDoS attacks in the cloud. In this paper we propose a novel framework for data security in the cloud using Collaborative Intrusion Detection (CIDS) scheme. The benefits of CIDS scheme in cloud are enabling the end user to get comprehensive information in the event of a distributed attack on cloud. Upasana T. Nagar, Priyadarsi Nanda, Xiangjian He, Zhiyuan Tan 0001 |
SIN | 3 |
| 2017 | PAWN: a payload-based mutual authentication scheme for wireless sensor networksabstractSummary Wireless sensor networks (WSNs) consist of resource‐starving miniature sensor nodes deployed in a remote and hostile environment. These networks operate on small batteries for days, months, and even years depending on the requirements of monitored applications. The battery‐powered operation and inaccessible human terrains make it practically infeasible to recharge the nodes unless some energy‐scavenging techniques are used. These networks experience threats at various layers and, as such, are vulnerable to a wide range of attacks. The resource‐constrained nature of sensor nodes, inaccessible human terrains, and error‐prone communication links make it obligatory to design lightweight but robust and secured schemes for these networks. In view of these limitations, we aim to design an extremely lightweight payload‐based mutual authentication scheme for a cluster‐based hierarchical WSN. The proposed scheme, also known as payload‐based mutual authentication for WSNs, operates in 2 steps. First, an optimal percentage of cluster heads is elected, authenticated, and allowed to communicate with neighboring nodes. Second, each cluster head, in a role of server, authenticates the nearby nodes for cluster formation. We validate our proposed scheme using various simulation metrics that outperform the existing schemes. Mian Ahmad Jan, Priyadarsi Nanda, Muhammad Usman 0015, Xiangjian He |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | User relationship strength modeling for friend recommendation on Instagram
Dongyan Guo, Jingsong Xu, Jian Zhang 0002, Min Xu 0001, Xiangjian He |
Neurocomputing | 6 |
| 2017 | Cryptography-based secure data storage and sharing using HEVC and public clouds
Muhammad Usman 0015, Mian Ahmad Jan, Xiangjian He |
Inf. Sci. | 3 |
| 2017 | MoWLD: a robust motion image descriptor for violence detection
Tao Zhang 0010, Wenjing Jia, Baoqing Yang, Jie Yang 0002, Xiangjian He, Zhonglong Zheng |
Multim. Tools Appl. | 5 |
| 2017 | Discriminative Dictionary Learning With Motion Weber Local Descriptor for Violence DetectionabstractAutomatic violence detection from video is a hot topic for many video surveillance applications. However, there has been little success in developing an algorithm that can detect violence in surveillance videos with high performance. In this paper, following our recently proposed idea of motion Weber local descriptor (WLD), we make two major improvements and propose a more effective and efficient algorithm for detecting violence from motion images. First, we propose an improved WLD (IWLD) to better depict low-level image appearance information, and then extend the spatial descriptor IWLD by adding a temporal component to capture local motion information and hence form the motion IWLD. Second, we propose a modified sparse-representation-based classification model to both control the reconstruction error of coding coefficients and minimize the classification error. Based on the proposed sparse model, a class-specific dictionary containing dictionary atoms corresponding to the class labels is learned using class labels of training samples. With this learned dictionary, not only the representation residual but also the representation coefficients become discriminative. A classification scheme integrating the modified sparse model is developed to exploit such discriminative information. The experimental results on three benchmark data sets have demonstrated the superior performance of the proposed approach over the state of the arts. Tao Zhang 0010, Wenjing Jia, Xiangjian He, Jie Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Visual Tracking via Nonnegative Multiple CodingabstractIt has been extensively observed that an accurate appearance model is critical to achieving satisfactory performance for robust object tracking. Most existing top-ranked methods rely on linear representation over a single dictionary, which brings about improper understanding on the target appearance. To address this problem, in this paper, we propose a novel appearance model named as “nonnegative multiple coding” (NMC) to accurately represent a target. First, a series of local dictionaries are created with different predefined numbers of nearest neighbors, and then the contributions of these dictionaries are automatically learned. As a result, this ensemble of dictionaries can comprehensively exploit the appearance information carried by all the constituted dictionaries. Second, the existing methods explicitly impose the nonnegative constraint to coefficient vectors, but in the proposed model, we directly deploy an efficient 12 norm regularization to achieve the similar nonnegative purpose with theoretical guarantees. Moreover, an efficient occlusion detection scheme is designed to alleviate tracking drifts, which investigates whether negative templates are selected to represent the severely occluded target. Experimental results on two benchmarks demonstrate that our NMC tracker are able to achieve superior performance to state-of-the-art methods. Fanghui Liu 0001, Chen Gong 0002, Tao Zhou 0002, Keren Fu, Xiangjian He, Jie Yang 0002 |
IEEE Trans. Multim. | 5 |
| 2016 | Unsupervised Multi-Author Document Decomposition Based on Hidden Markov Modelabstract© 2016 Association tor Computational Linguistics. This paper proposes an unsupervised approach for segmenting a multiauthor document into authorial components. The key novelty is that we utilize the sequential patterns hidden among document elements when determining their authorships. For this purpose, we adopt Hidden Markov Model (HMM) and construct a sequential probabilistic model to capture the dependencies of sequential sentences and their authorships. An unsupervised learning method is developed to initialize the HMM parameters. Experimental results on benchmark datasets have demonstrated the significant benefit of our idea and our approach has outperformed the state-of-the-arts on all tests. As an example of its applications, the proposed approach is applied for attributing authorship of a document and has also shown promising results. Khaled Aldebei, Xiangjian He, Wenjing Jia, Jie Yang 0002 |
ACL (1) | 2 |
| 2016 | Unsupervised Video Hashing by Exploiting Spatio-Temporal Feature
Chao Ma 0005, Yun Gu, Wei Liu 0044, Jie Yang 0002, Xiangjian He |
ICONIP (3) | 5 |
| 2016 | Learning a Discriminative Dictionary with CNN for Image Classification
Tao Zhang 0010, Chao Ma 0005, Lei Zhou 0003, Jie Yang 0002, Xiangjian He |
ICONIP (2) | 6 |
| 2016 | A Generative Model for Recognizing Mixed Group Activities in Still Images
Kan Li 0001, Xiangjian He |
IJCAI | 3 |
| 2016 | A new method for violence detection in surveillance scenes
Tao Zhang 0010, Wenjing Jia, Baoqing Yang, Jie Yang 0002, Xiangjian He |
Multim. Tools Appl. | 6 |
| 2016 | Shape-appearance-correlated active appearance modelabstractAmong the challenges faced by current active shape or appearance models, facial-feature localization in the wild, with occlusion in a novel face image, i.e. in a generic environment, is regarded as one of the most difficult computer-vision tasks. In this paper, we propose an Active Appearance Model (AAM) to tackle the problem of generic environment. Firstly, a fast face-model initialization scheme is proposed, based on the idea that the local appearance of feature points can be accurately approximated with locality constraints. Nearest neighbors, which have similar poses and textures to a test face, are retrieved from a training set for constructing the initial face model . To further improve the fitting of the initial model to the test face, an orthogonal CCA (oCCA) is employed to increase the correlation between shape features and appearance features represented by Principal Component Analysis (PCA). With these two contributions, we propose a novel AAM, namely the shape-appearance-correlated AAM (SAC-AAM), and the optimization is solved by using the recently proposed fast simultaneous inverse compositional (Fast-SIC) algorithm. Experiment results demonstrate a 5–10% improvement on controlled and semi-controlled datasets, and with around 10% improvement on wild face datasets in terms of fitting accuracy compared to other state-of-the-art AAM models. Huiling Zhou, Kin-Man Lam 0001, Xiangjian He |
Pattern Recognit. | 3 |
| 2016 | A unified model sharing framework for moving object detection
Yingying Chen 0003, Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu |
Signal Process. | 4 |
| 2016 | Big data meets multimedia analytics
Tat-Seng Chua, Xiangjian He, Weifeng Liu 0001, Massimo Piccardi, Yonggang Wen 0001, Dacheng Tao |
Signal Process. | 2 |
| 2016 | Building an Intrusion Detection System Using a Filter-Based Feature Selection AlgorithmabstractRedundant and irrelevant features in data have caused a long-term problem in network traffic classification. These features not only slow down the process of classification but also prevent a classifier from making accurate decisions, especially when coping with big data. In this paper, we propose a mutual information based algorithm that analytically selects the optimal feature for classification. This mutual information based feature selection algorithm can handle linearly and nonlinearly dependent data features. Its effectiveness is evaluated in the cases of network intrusion detection. An Intrusion Detection System (IDS), named Least Square Support Vector Machine based IDS (LSSVM-IDS), is built using the features selected by our proposed feature selection algorithm. The performance of LSSVM-IDS is evaluated using three intrusion detection evaluation datasets, namely KDD Cup 99, NSL-KDD and Kyoto 2006+ dataset. The evaluation results show that our feature selection algorithm contributes more critical features for LSSVM-IDS to achieve better accuracy and lower computational cost compared with the state-of-the-art methods. Mohammed A. Ambusaidi, Xiangjian He, Priyadarsi Nanda, Zhiyuan Tan 0001 |
IEEE Trans. Computers | 2 |
| 2016 | Frame Interpolation for Cloud-Based Mobile Video StreamingabstractCloud-based High Definition (HD) video streaming is becoming popular day by day. On one hand, it is important for both end users and large storage servers to store their huge amount of data at different locations and servers. On the other hand, it is becoming a big challenge for network service providers to provide reliable connectivity to the network users. There have been many studies over cloud-based video streaming for Quality of Experience (QoE) for services like YouTube. Packet losses and bit errors are very common in transmission networks, which affect the user feedback over cloud-based media services. To cover up packet losses and bit errors, Error Concealment (EC) techniques are usually applied at the decoder/receiver side to estimate the lost information. This paper proposes a time-efficient and quality-oriented EC method. The proposed method considers H.265/HEVC based intra-encoded videos for the estimation of whole intra-frame loss. The main emphasis in the proposed approach is the recovery of Motion Vectors (MVs) of a lost frame in real-time. To boost-up the search process for the lost MVs, a bigger block size and searching in parallel are both considered. The simulation results clearly show that our proposed method outperforms the traditional Block Matching Algorithm (BMA) by approximately 2.5 dB and Frame Copy (FC) by up to 12 dB at a packet loss rate of 1%, 3%, and 5% with different Quantization Parameters (QPs). The computational time of the proposed approach outperforms the BMA by approximately 1788 seconds. Muhammad Usman 0015, Xiangjian He, Kin-Man Lam 0001, Min Xu 0001, Syed Mohsin Matloob Bokhari, Jinjun Chen |
IEEE Trans. Multim. | 2 |
| 2015 | Small target detection using an optimization-based filterabstractSmall target detection is a critical problem in the Infrared Search And Track (IRST) system. Although it has been studied for years, there are some challenges remained, e.g. cloud edges and horizontal lines are likely to cause false alarms. This paper proposes a novel method using an optimization-based filter to detect infrared small target in heavy clutter. First, we design a certain pixel area as active area. Second, a weighted quadratic cost function is performed in the active area. Finally, a filter based on statistics of active area is derived from the cost function. Our method could preserve heterogeneous area, meanwhile, remove target region. Experimental results show our method achieves satisfied performance in heavy clutter. Keren Fu, Tao Zhou 0002, Jie Yang 0002, Qiang Wu 0001, Xiangjian He |
ICASSP | 6 |
| 2015 | Face hallucination based on nonparametric Bayesian learningabstractIn this paper, we propose a novel example-based face hallucination method through nonparametric Bayesian learning based on the assumption that human faces have similar local pixel structure. We cluster the low resolution (LR) face image patches by nonparametric method distance dependent Chinese Restaurant process (ddCRP) and calculate the centres of the clusters (i.e., subspaces). Then, we learn the mapping coefficients from the LR patches to high resolution (HR) patches in each subspace. Finally, the HR patches of input low resolution face image can be efficiently generated by a simple linear regression. The spatial distance constraint is employed to aid the learning of subspace centers so that every subspace will better reflect the detailed information of image patches. Experimental results show our method is efficient and promising for face hallucination. Minqi Li, Xiangjian He |
ICIP | 3 |
| 2015 | Online learning of multi-feature weights for robust object trackingabstractSparse Representation based Classification (SRC) and its potential in object tracking have been explored in recent years. However, the trade-off between the discriminative ability of the overly emphasized sparse representation and the lack of insight on correlation of visual information has raised questions over the general applicability of such methods in object tracking. In addition, the need for the optimization of a series of l1-regularized least square norm, increases the computational complexity thereby limiting their usage in real-time applications. In this paper, a novel approach to robust object tracking is proposed. First, the variations in the appearance of the tracked target is modelled using PCA basis vectors, and further, a l2-regularized least square method is used to solve the proposed representation model. In order to improve the robustness of feature representation in object tracking applications, weights are associated with multiple trackers; each formulated using a different feature, and adapted via an online learning scheme. Finally, a decision fusion criterion is imposed to generate an optimized output through the weighted combination of different tracking results. Experiments on challenging video sequences have demonstrated the superior accuracy and robustness of the proposed method in comparison to thirteen other state-of-the-art baselines. Tao Zhou 0002, Harish Bhaskar, Jie Yang 0002, Xiangjian He |
ICIP | 5 |
| 2015 | A probability-dynamic Particle Swarm Optimization for object trackingabstractParticle Swarm Optimization has been used in many research and application domain popularly since its development and improvement. Due to its fast and accurate solution searching, PSO has become one of the high potential tools to provide better outcomes to solve many practical problems. In image processing and object tracking applications, PSO also indicates to have good performance in both linear and non-linear object moving pattern, many scientists conduct development and research to implement not only basic PSO but also improved methods in enhancing the efficiency of the algorithm to achieve precise object tracking orbit. This paper is aim to propose a new improved PSO by comparing the inertia weight and constriction factor of PSO. It provides faster and more accurate object tracking process since the proposed algorithm can inherit some useful information from the previous solution to perform the dynamic particle movement when other better solution exists. The testing experiments have been done for different types of video, results showed that the proposed algorithm can have better quality of tracking performance and faster object retrieval speed. The proposed approach has been developed in C++ environment and tested against videos and objects with multiple moving patterns to demonstrate the benefits with precise object similarity. Feng Sha, Changseok Bae, Guang Liu 0002, XiMeng Zhao, Vera Chung, Wei-Chang Yeh 0001, Xiangjian He |
IJCNN | 7 |
| 2015 | Solving reliability redundancy allocation problems with orthogonal simplified swarm optimizationabstractThis study applies a penalty guided strategy and the orthogonal array test (OA) based on the Simplified Swarm Optimization algorithm (SSO) to solve the reliability redundancy allocation problems (RRAP) in the series system, the series-parallel system, the complex (bridge) system, and the overspeed protection of gas turbine system. For several decades, the RRAP has been one of the most well known techniques. The maximization of system reliability, the number of redundant components, and the reliability of corresponding components in each subsystem have to be decided simultaneously with nonlinear constraints, acting as one difficulty for the use of the RRAP. In other words, the objective function of the RRAP is the mixed-integer programming problem with the nonlinear constraints. The RRAP is of the class of NP-hard. Hence, in this paper, the SSO algorithm is proposed to solve the RRAP and improve computation efficiency for these NP-hard problems. There are four RRAP problems used to illustrate the applicability and the effectiveness of the SSO. The experimental results are compared with previously developed algorithms in literature. Moreover, the maximum-possible-improvement (MPI) is used to measure the amount of improvement of the solution found by the SSO to the previous solutions. According to the results, the system reliabilities obtained by the proposed SSO for the four RRAP problems are as well as or better than the previously best-known solutions. Wei-Chang Yeh 0001, Vera Chung, Yunzhi Jiang, Xiangjian He |
IJCNN | 4 |
| 2015 | Recognizing Human Activity in Still Images by Integrating Group-Based Contextual CuesabstractImages with wider angles usually capture more persons in wider scenes, and recognizing individuals' activities in these images based on existing contextual cues usually meet difficulties. We instead construct a novel group-based cue to utilize the context carried by suitable surrounding persons. We propose a global-local cue integration model (GLCIM) to find a suitable group of local cues extracted from individuals and form a corresponding global cue. A fusion restricted Boltzmann machine, a focal subspace measurement and a cue integration algorithm based on entropy are proposed to enable the GLCIM to integrate most of the relevant local cues and least of the irrelevant ones into the group. Our experiments demonstrate how integrating group-based cues improves the activity recognition accuracies in detail and show that all of the key parts of GLCIM make positive contributions to the increases of the accuracies. Kan Li 0001, Xiangjian He |
ACM Multimedia | 3 |
| 2015 | Orderless and Blurred Visual Tracking via Spatio-temporal Context
Manna Dai, Peijie Lin, Lijun Wu 0002, Zhicong Chen, Songlin Lai, Shuying Cheng, Xiangjian He |
MMM (1) | 8 |
| 2015 | A New Image Decomposition and Reconstruction Approach - Adaptive Fourier Decomposition
Can He, Liming Zhang 0002, Xiangjian He, Wenjing Jia |
MMM (2) | 3 |
| 2015 | A Novel Fast Full Frame Video Stabilization via Three-Layer Model
Jie Yang 0002, Dacheng Song, Xiangjian He |
MMM (1) | 5 |
| 2015 | Community Detection Based on Links and Node Features in Social Networks
Fengli Zhang, Min Xu 0001, Xiangjian He |
MMM (1) | 6 |
| 2015 | Survey of Error Concealment techniques: Research directions and open issuesabstractError Concealment (EC) techniques use either spatial, temporal or a combination of both types of information to recover the data lost in transmitted video. In this paper, existing EC techniques are reviewed, which are divided into three categories, namely Intra-frame EC, Inter-frame EC, and Hybrid EC techniques. We first focus on the EC techniques developed for the H.264/AVC standard. The advantages and disadvantages of these EC techniques are summarized with respect to the features in H.264. Then, the EC algorithms are also analyzed. These EC algorithms have been recently adopted in the newly introduced H.265/HEVC standard. A performance comparison between the classic EC techniques developed for H.264 and H.265 is performed in terms of the average PSNR. Lastly, open issues in the EC domain are addressed for future research consideration. Muhammad Usman 0015, Xiangjian He, Min Xu 0001, Kin-Man Lam 0001 |
PCS | 2 |
| 2015 | Observation-driven adaptive differential evolution and its application to accurate and smooth bronchoscope three-dimensional motion tracking
Xióngbiao Luó, Xiangjian He, Kensaku Mori |
Medical Image Anal. | 3 |
| 2015 | A camera motion histogram descriptor for video shot classification
Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Yi Wang 0037 |
Multim. Tools Appl. | 3 |
| 2015 | Fast and robust head detection with arbitrary pose and occlusion
Tao Zhang 0010, Wenjing Jia, Qiang Wu 0001, Jie Yang 0002, Xiangjian He |
Multim. Tools Appl. | 6 |
| 2015 | Scalable Semi-Supervised Classification via Neumann Series
Chen Gong 0002, Keren Fu, Lei Zhou 0003, Jie Yang 0002, Xiangjian He |
Neural Process. Lett. | 5 |
| 2015 | Robust visual tracking via efficient manifold ranking with low-dimensional compressive features
Tao Zhou 0002, Xiangjian He, Keren Fu, Jie Yang 0002 |
Pattern Recognit. | 2 |
| 2015 | Online gesture-based interaction with visual oriental characters based on manifold learning
Yi Wang 0037, Xin Fan 0001, Xiangjian He, Qi Jia 0001, Renjie Gao |
Signal Process. | 4 |
| 2015 | Detection of Denial-of-Service Attacks Based on Computer Vision TechniquesabstractDetection of Denial-of-Service (DoS) attacks has attracted researchers since 1990s. A variety of detection systems has been proposed to achieve this task. Unlike the existing approaches based on machine learning and statistical analysis, the proposed system treats traffic records as images and detection of DoS attacks as a computer vision problem. A multivariate correlation analysis approach is introduced to accurately depict network traffic records and to convert the records into their respective images. The images of network traffic records are used as the observed objects of our proposed DoS attack detection system, which is developed based on a widely used dissimilarity measure, namely Earth Mover's Distance (EMD). EMD takes cross-bin matching into account and provides a more accurate evaluation on the dissimilarity between distributions than some other well-known dissimilarity measures, such as Minkowski-form distance Lpand X2statistics. These unique merits facilitate our proposed system with effective detection capabilities. To evaluate the proposed EMD-based detection system, ten-fold cross-validations are conducted using KDD Cup 99 dataset and ISCX 2012 IDS Evaluation dataset. The results presented in the system evaluation section illustrate that our detection system can detect unknown DoS attacks and achieves 99.95 percent detection accuracy on KDD Cup 99 dataset and 90.12 percent detection accuracy on ISCX 2012 IDS evaluation dataset with processing capability of approximately 59,000 traffic records per second. Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001, Jiankun Hu |
IEEE Trans. Computers | 3 |
| 2015 | Local N-Ary Pattern and Its Extension for Texture ClassificationabstractTexture image classification is important in computer vision research. To effectively capture texture patterns, a distinctive feature such as a local binary pattern (LBP) is needed. An LBP is robust against monotonic and gray-scale variations and it computes quickly. Its robustness and speed advantage have made it popular in various texture analysis applications. However, an LBP is sensitive to noise, particularly smooth weak illumination gradients in near-uniform regions. To mitigate the effect of noise and increase distinctiveness, a local ternary pattern (LTP) is proposed. Compared with a binary coding LBP, an LTP adopts ternary coding. As a result, an LTP can better tolerate noise and is significantly more distinctive. These advantages of an LTP effectively improve its classification accuracy. However, the potential of ternary coding is not fully explored in LTPs because a ternary pattern is split into a pair of binary patterns. In this paper, to fully explore the distinctiveness in the local pattern, the feature extraction process is formulated as an integer decomposition problem, which is a generalized version of the Bachet de Meziriac weight problem (BMWP). Following this generalization, a local n-ary pattern (LNP) is proposed, for which the LBP is a special case parametrized under n = 2. The LTP is not a special case of the LNP. Both LBP and LTP are used as benchmark methods to evaluate LNPs performance due to their well-recognized success. In addition, a rotation-invariant and uniform LNP is also proposed and compared with a rotation-invariant and uniform LBP. The proposed LNP achieves significantly improved texture classification accuracy compared with the LBP and also demonstrates considerable improvement over the LTP. Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Jie Yang 0002, Yi Wang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Macroscopic Indeterminacy Swarm Optimization (MISO) algorithm for real-parameter searchabstractSwarm Intelligence (SI) is a nature-inspired emergent artificial intelligence. They are often inspired by the phenomena in nature. Many proposed algorithms are focused on designing new update mechanisms with formulae and equations to emerge new solutions. Despite the techniques used in an algorithm being the key factor of the whole system, the evaluation of candidate solutions also plays an important role. In this paper, the proposed algorithm Macroscopic Indeterminacy Swarm Optimization (MISO) presents a new search scheme with indeterminate moment of evaluation. Here, we perform an experiment based on public benchmark functions. The results produced by MISO, Differential Evolution (DE) with various settings, Artificial Bee Colony (ABC), Simplified Swarm Optimization (SSO), and Particle Swarm Optimization (PSO) have been compared. The result shows MISO can achieve similar or even better performance than other algorithms. Po-Chun Chang, Xiangjian He |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Diversity-Enhanced Condensation Algorithm and Its Application for Robust and Accurate Endoscope Three-Dimensional Motion TrackingabstractThe paper proposes a diversity-enhanced condensation algorithm to address the particle impoverishment problem which stochastic filtering usually suffers from. The particle diversity plays an important role as it affects the performance of filtering. Although the condensation algorithm is widely used in computer vision, it easily gets trapped in local minima due to the particle degeneracy. We introduce a modified evolutionary computing method, adaptive differential evolution, to resolve the particle impoverishment under a proper size of particle population. We apply our proposed method to endoscope tracking for estimating three-dimensional motion of the endoscopic camera. The experimental results demonstrate that our proposed method offers more robust and accurate tracking than previous methods. The current tracking smoothness and error were significantly reduced from (3.7, 4.8) to (2.3 mm, 3.2 mm), which approximates the clinical requirement of 3.0 mm. Xióngbiao Luó, Xiangjian He, Jie Yang 0002, Kensaku Mori |
CVPR | 3 |
| 2014 | Visual tracking based on weighted subspace reconstruction errorabstractIt is a challenging task to develop an effective and robust visual tracking method due to factors such as pose variation, illumination change, occlusion, and motion blur. In this paper, a novel tracking algorithm based on weighted subspace reconstruction error is proposed. We first compute the discriminative weights by sparse construction error with template dictionary consisted of positive and negative samples, and then confidence map for candidates is computed through subspace reconstruction error. Finally, the location of the target object is estimated by maximizing the decision map which is combined discriminative weights and subspace reconstruction error. Furthermore, we use the new evaluation criterion to verify the robustness of the current tracking result, which can reduce the accumulated error effectively. Experimental results on some challenging video sequences show that the proposed algorithm performs favorably against seven state-of-the-art methods in terms of accuracy and robustness. Tao Zhou 0002, Jie Yang 0002, Xiangjian He |
ICIP | 5 |
| 2014 | Spectral salient object detectionabstractMany existing methods for salient object detection are performed by over-segmenting images into non-overlapping regions, which facilitate local/global color statistics for saliency computation. In this paper, we propose a new approach: spectral salient object detection, which is benefited from selected attributes of normalized cut, enabling better retaining of holistic salient objects as comparing to conventionally employed pre-segmentation techniques. The proposed saliency detection method recursively bi-partitions regions that render the lowest cut cost in each iteration, resulting in binary spanning tree structure. Each segmented region is then evaluated under criterion that fit Gestalt laws and statistical prior. Final result is obtained by integrating multiple intermediate saliency maps. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed method against 13 state-of-the-art approaches to salient object detection. Keren Fu, Chen Gong 0002, Irene Y. H. Gu, Jie Yang 0002, Xiangjian He |
ICME | 5 |
| 2014 | Visual tracking via graph-based efficient manifold ranking with low-dimensional compressive featuresabstractIn this paper, a novel and robust tracking method based on efficient manifold ranking is proposed. For tracking, tracked results are taken as labeled nodes while candidate samples are taken as unlabeled nodes, and the goal of tracking is to search the unlabeled sample that is the most relevant with existing labeled nodes by manifold ranking algorithm. Meanwhile, we adopt non-adaptive random projections to preserve the structure of original image space, and a very sparse measurement matrix is used to efficiently extract low-dimensional compressive features for object representation. Furthermore, spatial context is used to improve the robustness to appearance variations. Experimental results on some challenging video sequences show the proposed algorithm outperforms six state-of-the-art methods in terms of accuracy and robustness. Tao Zhou 0002, Xiangjian He, Keren Fu, Jie Yang 0002 |
ICME | 2 |
| 2014 | Feature Selection and Mass Classification Using Particle Swarm Optimization and Support Vector Machine
Man To Wong, Xiangjian He, Wei-Chang Yeh 0001, Zaidah Ibrahim, Vera Chung |
ICONIP (3) | 2 |
| 2014 | Estimate Gaze Density by Incorporating EmotionabstractGaze density estimation has attracted many research efforts in the past years. The factors considered in the existing methods include low level feature saliency, spatial position, and objects. Emotion, as an important factor driving attention, has not been taken into account. In this paper, we are the first to estimate gaze density through incorporating emotion. To estimate the emotion intensity of each position in an image, we consider three aspects, generic emotional content, facial expression intensity, and emotional objects. Generic emotional content is estimated by using Multiple instance learning, which is employed to train an emotion detector from weakly labeled images. Facial expression intensity is estimated by using a ranking method. Emotional objects are detected, by taking blood/injury and worm/snake as examples. Finally, emotion intensity, low level feature saliency, and spatial position, are fused, through a linear support vector machine, to estimate gaze density. The performance is tested on public eye tracking dataset. Experimental results indicate that incorporating emotion does improve the performance of gaze density estimation. Min Xu 0001, Xiangjian He, Jinqiao Wang |
ACM Multimedia | 3 |
| 2014 | A Novel Feature Selection Approach for Intrusion Detection Data ClassificationabstractIntrusion Detection Systems (IDSs) play a significant role in monitoring and analyzing daily activities occurring in computer systems to detect occurrences of security threats. However, the routinely produced analytical data from computer networks are usually of very huge in size. This creates a major challenge to IDSs, which need to examine all features in the data to identify intrusive patterns. The objective of this study is to analyze and select the more discriminate input features for building computationally efficient and effective schemes for an IDS. For this, a hybrid feature selection algorithm in combination with wrapper and filter selection processes is designed in this paper. Two main phases are involved in this algorithm. The upper phase conducts a preliminary search for an optimal subset of features, in which the mutual information between the input features and the output class serves as a determinant criterion. The selected set of features from the previous phase is further refined in the lower phase in a wrapper manner, in which the Least Square Support Vector Machine (LSSVM) is used to guide the selection process and retain optimized set of features. The efficiency and effectiveness of our approach is demonstrated through building an IDS and a fair comparison with other stateof-the-art detection approaches. The experimental results show that our hybrid model is promising in detection compared to the previously reported results. Mohammed A. Ambusaidi, Xiangjian He, Zhiyuan Tan 0001, Priyadarsi Nanda, Upasana T. Nagar |
TrustCom | 2 |
| 2014 | A Stateful Mechanism for the Tree-Rule FirewallabstractIn this paper, we propose a novel connection tracking mechanism for Tree-rule firewall which essentially organizes firewall rules in a designated Tree structure. A new firewall model based on the proposed connection tracking mechanism is then developed and extended from the basic model of Net filter's Conn Track module, which has been used by many early generation commercial and open source firewalls including IPTABLES, the most popular firewall. To reduce the consumption of memory space and processing time, our proposed model uses one node per connection instead of using two nodes as appeared in Net filter model. This can reduce memory space and processing time. In addition, we introduce an extended hash table with more hashing bits in our firewall model in order to accommodate more concurrent connections. Moreover, our model also applies sophisticated techniques (such as using static information nodes, and avoiding timer objects and memory management tasks) to improve its processing speed. Finally, we implement this model on Linux Cent OS 6.3 and evaluate its speed. The experimental results show that our model performs more efficiently in comparison with the Net filter/IPTABLES. Thawatchai Chomsiri, Xiangjian He, Priyadarsi Nanda, Zhiyuan Tan 0001 |
TrustCom | 2 |
| 2014 | A Robust Authentication Scheme for Observing Resources in the Internet of Things EnvironmentabstractThe Internet of Things is a vision that broadens the scope of the internet by incorporating physical objects to identify themselves to the participating entities. This innovative concept enables a physical device to represent itself in the digital world. There are a lot of speculations and future forecasts about the Internet of Things devices. However, most of them are vendor specific and lack a unified standard, which renders their seamless integration and interoperable operations. Another major concern is the lack of security features in these devices and their corresponding products. Most of them are resource-starved and unable to support computationally complex and resource consuming secure algorithms. In this paper, we have proposed a lightweight mutual authentication scheme which validates the identities of the participating devices before engaging them in communication for the resource observation. Our scheme incurs less connection overhead and provides a robust defence solution to combat various types of attacks. Mian Ahmad Jan, Priyadarsi Nanda, Xiangjian He, Zhiyuan Tan 0001, Ren Ping Liu 0001 |
TrustCom | 3 |
| 2014 | PASCCC: Priority-based application-specific congestion control clustering protocol
Mian Ahmad Jan, Priyadarsi Nanda, Xiangjian He, Ren Ping Liu 0001 |
Comput. Networks | 3 |
| 2014 | Improving cloud network security using the Tree-Rule firewall
Xiangjian He, Thawatchai Chomsiri, Priyadarsi Nanda, Zhiyuan Tan 0001 |
Future Gener. Comput. Syst. | 1 |
| 2014 | Interactive segmentation based on iterative learning for multiple-feature fusionabstractThis paper proposes a novel interactive segmentation method based on conditional random field (CRF) model to utilize the location and color information contained in user input. The CRF is configured with the optimal weights between two features, which are the color Gaussian Mixture Model (GMM) and probability model of location information. To construct the CRF model, we propose a method to collect samples for the cuttraining tasks of learning the optimal weights on a single image׳s basis and updating the parameters of features. To refine the segmentation results iteratively, our method applies the active learning strategy to guide the process of CRF model updating or guide users to input minimal training data for training the optimal weights and updating the parameters of features. Experimental results show that the proposed method demonstrates qualitative and quantitative improvement compared with the state-of-the-art interactive segmentation methods. The proposed method is also a convenient tool for interactive object segmentation. Lei Zhou 0003, Yu Qiao 0003, Yijun Li 0003, Xiangjian He, Jie Yang 0002 |
Neurocomputing | 4 |
| 2014 | A three-level framework for affective content analysis and its case studies
Min Xu 0001, Jinqiao Wang, Xiangjian He, Jesse S. Jin, Suhuai Luo, Hanqing Lu |
Multim. Tools Appl. | 3 |
| 2014 | A hybrid domain enhanced framework for video retargeting with spatial-temporal importance and 3D grid optimization
Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu, Doan B. Hoang |
Signal Process. | 3 |
| 2014 | Bayesian salient object detection based on saliency driven clustering
Lei Zhou 0003, Keren Fu, Yijun Li 0003, Yu Qiao 0001, Xiangjian He, Jie Yang 0002 |
Signal Process. Image Commun. | 5 |
| 2014 | CAMHID: Camera Motion Histogram Descriptor and Its Application to Cinematographic Shot ClassificationabstractIn this paper, we propose a nonparametric camera motion descriptor for video shot classification. In the proposed method, a motion vector field (MVF) is constructed for each consecutive video frame by computing the motion vector (MV) of each macroblock. Then, the MVFs are divided into a number of local region of equal size. Next, the inconsistent/noisy MVs of each local region are eliminated by a motion consistency analysis. The remaining MVs of each local region from a number of consecutive frames are further collected for a compact representation. Initially, a matrix is formed using the MVs. Then, the matrix is decomposed using a singular value decomposition technique to represent the dominant motion. Finally, the angle of the most variance retaining principal component is computed and quantized to represent the motion of a local region by using a histogram. In order to represent the global camera motion, the local histograms are combined. The effectiveness of the proposed motion descriptor for video shot classification is tested by using a support vector machine. First, the proposed camera motion descriptors for video shots classification are computed on a video data set consisting of regular camera motion patterns (e.g., pan, zoom, tilt, static). Then, we apply the camera motion descriptors with an extended set of features to the classification of cinematographic shots. The experimental results show that the proposed shot level camera motion descriptor has a strong discriminative capability to classify different camera motion patterns of different videos effectively. We also show that our approach outperforms state-of-the-art methods. Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | A System for Denial-of-Service Attack Detection Based on Multivariate Correlation AnalysisabstractInterconnected systems, such as Web servers, database servers, cloud computing servers and so on, are now under threads from network attackers. As one of most common and aggressive means, denial-of-service (DoS) attacks cause serious impact on these computing systems. In this paper, we present a DoS attack detection system that uses multivariate correlation analysis (MCA) for accurate network traffic characterization by extracting the geometrical correlations between network traffic features. Our MCA-based DoS attack detection system employs the principle of anomaly based detection in attack recognition. This makes our solution capable of detecting known and unknown DoS attacks effectively by learning the patterns of legitimate network traffic only. Furthermore, a triangle-area-based technique is proposed to enhance and to speed up the process of MCA. The effectiveness of our proposed detection system is evaluated using KDD Cup 99 data set, and the influences of both non-normalized data and normalized data on the performance of the proposed detection system are examined. The results show that our system outperforms two other previously developed state-of-the-art approaches in terms of detection accuracy. Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Generalized local N-ary patterns for texture classificationabstractLocal Binary Pattern (LBP) has been well recognised and widely used in various texture analysis applications of computer vision and image processing. It integrates properties of texture structural and statistical texture analysis. LBP is invariant to monotonic gray-scale variations and has also extensions to rotation invariant texture analysis. In recent years, various improvements have been achieved based on LBP. One of extensive developments was replacing binary representation with ternary representation and proposed Local Ternary Pattern (LTP). This paper further generalises the local pattern representation by formulating it as a generalised weight problem of Bachet de Meziriac and proposes Local N-ary Pattern (LNP). The encouraging performance is achieved based on three benchmark datasets when compared with its predecessors. Sheng Wang 0003, Xiangjian He, Qiang Wu 0001, Jie Yang 0002 |
AVSS | 2 |
| 2013 | Text detection in born-digital images using multiple layer imagesabstractIn this paper, a new framework for detecting text from webpage and email images is presented. The original image is split into multiple layer images based on the maximum gradient difference (MGD) values to detect text with both strong and weak contrasts. Connected component processing and text detection are performed in each layer image. A novel texture descriptor named T-LBP, is proposed to further filter out non-text candidates with a trained SVM classifier. The ICDAR 2011 born-digital image dataset is used to evaluate and demonstrate the performance of the proposed method. Following the same performance evaluation criteria, the proposed method outperforms the winner algorithm of the ICDAR 2011 Robust Reading Competition Challenge 1. Wenjing Jia, Xiangjian He |
ICASSP | 3 |
| 2013 | RePIDS: A multi tier Real-time Payload-based Intrusion Detection System
Aruna Jamdagni, Zhiyuan Tan 0001, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001 |
Comput. Networks | 3 |
| 2013 | A segmentation-free method for image classification based on pixel-wise matching
Jun Ma 0004, Long Zheng 0001, Mianxiong Dong, Xiangjian He, Minyi Guo, Yuichi Yaguchi |
J. Comput. Syst. Sci. | 4 |
| 2013 | An algorithm for accuracy enhancement of license plate recognition
Lihong Zheng, Xiangjian He, Bijan Samali, Laurence T. Yang |
J. Comput. Syst. Sci. | 2 |
| 2013 | MIL-SKDE: Multiple-instance learning with supervised kernel density estimation
Ruo Du, Qiang Wu 0001, Xiangjian He, Jie Yang 0002 |
Signal Process. | 3 |
| 2013 | Hierarchical affective content analysis in arousal and valence dimensions
Min Xu 0001, Changsheng Xu, Xiangjian He, Jesse S. Jin, Suhuai Luo, Yong Rui |
Signal Process. | 3 |
| 2012 | A radio frequency identification network design methodology for the decision problem in Mackay Memorial Hospital based on swarm optimizationabstractRadio-frequency identification (RFID) is an automatic identification system which has become a hot topic in the fields of manufacturing, logistics, and so on. The purpose of this research is to propose a methodology for designing the RFID network planning problem (RNP) for application in the Mackay Memorial Hospital in Hsinchu, Taiwan. In this study, the RFID network is first considered as a grid and divided into several small squares. A soft computing methodology called FKB-SSO is proposed to solve the RNP problem based on simplified swarm optimization (SSO) by integrating k-means, fuzzy adaptive resonance theory (fuzzy-ART), and binary search. The proposed FKB-SSO will provide the basis for strategic decisions in constructing the RFID network to reduce the number of RFID readers with a minimal budget under the constraint of 100% coverage rate. The proposed FKB-SSO is more efficient than PSO and experts' manual solution in both run time and solution quality. Wei-Chang Yeh 0001, Yuan-Ming Yeh, Chun-Hua Chou, Vera Chung, Xiangjian He |
IEEE Congress on Evolutionary Computation | 5 |
| 2012 | Shot Classification Using Domain Specific Features for Movie Management
Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Ling Chen 0006 |
DASFAA (2) | 3 |
| 2012 | On splitting dataset: Boosting Locally Adaptive Regression Kernels for car localizationabstractIn this paper, we study the impact of learning an Adaboost classifier with small sample set (i.e., with fewer training examples). In particular, we make use of car localization as an underlying application, because car localization can be widely used to various real world applications. In order to evaluate the performance of Adaboost learning with a few examples, we simply apply Adaboost learning to a recently proposed feature descriptor - Locally Adaptive Regression Kernel (LARK). As a type of state-of-the-art feature descriptor, LARK is robust against illumination changes and noises. More importantly, we use LARK because its spatial property is also favorable for our purpose (i.e., each patch in the LARK descriptor corresponds to one unique pixel in the original image). In addition to learning a detector from the entire training dataset, we also split the original training dataset into several sub-groups and then we train one detector for each sub-group. We compare those features associated using the detector of each sub-group with that of the detector learnt with the entire training dataset and propose improvements based on the comparison results. Our experimental results indicate that the Adaboost learning is only successful on a small dataset when those learnt features simultaneously satisfy two conditions that: 1. features are learnt from the Region of Interest (ROI), and 2. features are sufficiently far away from each other. Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Min Xu 0001 |
ICARCV | 3 |
| 2012 | Learning geodesic CRF model for image segmentationabstractGraph cut based on color model is sensitive to statistical information of images. Integrating priority information into graph cut approach, such as the geodesic distance information, may overcome the well-known drawback of bias towards shorter paths that occurred frequently with graph cut methods. In this paper, a conditional random field (CRF) model is formulated to combine color model and geodesic distance information into a graph cut optimization framework. A discriminative model is used to capture more comprehensive statistical information for geodesic distance. A simple and efficient parameter learning scheme based on feature fusion is proposed for CRF model construction. The method is evaluated by applying it to segmentation of natural images, medical images and low contrast images. The experimental results show that the geodesic information obtained by learning can provide more reliable object features. The dynamic parameter learning scheme is able to select best cues from geodesic map and color model for image segmentation. Lei Zhou 0003, Yu Qiao 0001, Jie Yang 0002, Xiangjian He |
ICIP | 4 |
| 2012 | The Extended Co-learning Framework for Robust Object TrackingabstractRecently, object tracking has been widely studied as a binary classification problem. Semi-supervised learning is particularly suitable for improving classification accuracy when large quantities of unlabeled samples are generated (just like tracking procedure). The purpose of this paper is to fulfill robust and stable tracking by using collaborative learning, which belongs to the scope of semi-supervised learning, among three classifiers. Different from [1], random fern classifier is incorporated to deal with 2bitBP feature newly added and certain constraints are specially implemented in our framework. Besides, the way for selecting positive samples is also altered by us in order to achieve more stable tracking. Algorithm proposed in this paper is validated by tracking pedestrian and cup under occlusion. Experiments and comparison show that our algorithm can avoid drifting problem to some degree and make tracking result more robust and adaptive. Chen Gong 0002, Yang Liu 0007, Tianyu Li 0003, Jie Yang 0002, Xiangjian He |
ICME | 5 |
| 2012 | Efficient Super-Resolution by Finer Sub-Pixel Motion Prediction and Bilateral FilteringabstractSuper-resolution reconstruction produces high-resolution images from a set of low-resolution images of the same scene. In the last two and a half decades, many super-resolution algorithms have been proposed. These algorithms are very sensitive to their assumed models of motion and noise, and computationally expensive for many practical applications. In this paper we adopt earlier reported fast prediction based sub-pixel motion estimation and a novel interpolation scheme based on the bilateral filter to produce a fast color super-resolution reconstruction that can accommodate arbitrary local motion patterns. The proposed algorithm exploits photometric proximity and available finer fractional motion information in the high resolution grid, to reconstruct enhanced super-resolved image frames. Experiments show a PSNR performance comparable to the state-of-the-art but at a fraction of their computational cost. Damith J. Mudugamuwa, Xiangjian He, Wenjing Jia |
ICME | 2 |
| 2012 | Mass Classification in Digitized Mammograms Using Texture Features and Artificial Neural Network
Man To Wong, Xiangjian He, Hung T. Nguyen 0001, Wei-Chang Yeh 0001 |
ICONIP (5) | 2 |
| 2012 | Battle-Lemarie wavelet pyramid for improved GSM image denoising
Damith J. Mudugamuwa, Xiangjian He, Wenjing Jia |
ICPR | 2 |
| 2012 | Border Gateway Protocol Anomaly Detection Using Failure Quality Control MethodabstractBorder Gateway Protocol (BGP) is the de-facto inter-domain routing protocol used across thousands of Autonomous Systems (AS) joined together in the Internet. Security has been a major issue for BGP. Nevertheless, BGP suffers from serious threats even today, like Denial of Service (DoS) attack and misconfiguration of routing information. BGP is one of the complex routing protocols and hard to configure against malicious attacks. However, it is important to detect such malicious activities in a network, which could otherwise cause problems for availability of services in the Internet. In this paper we use the Failure Quality Control (FQC), a technique to detect anomaly packets in the network for real time intrusion detection. Muhammad Mujtaba, Priyadarsi Nanda, Xiangjian He |
TrustCom | 3 |
| 2012 | Triangle-Area-Based Multivariate Correlation Analysis for Effective Denial-of-Service Attack DetectionabstractCloud computing plays an important role in current converged networks. It brings convenience of accessing services and information to users regardless of location and time. However, there are some critical security issues residing in cloud computing, such as availability of services. Denial of service occurring on cloud computing has even more serious impact on the Internet. Therefore, this paper studies the techniques for detecting Denial-of-Service (DoS) attacks to network services and proposes an effective system for DoS attack detection. The proposed system applies the idea of Multivariate Correlation Analysis (MCA) to network traffic characterization and employs the principal of anomaly-based detection in attack recognition. This makes our solution capable of detecting known and unknown DoS attacks effectively by learning the patterns of legitimate network traffic only. Furthermore, a triangle area technique is proposed to enhance and speed up the process of MCA. The effectiveness of our proposed detection system is evaluated on the KDD Cup 99 dataset, and the influence of both non-normalized and normalized data on the performance of the detection system is examined. The results presented in the system evaluation section illustrate that our DoS attack detection system outperforms two state-of-the-art approaches. Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001 |
TrustCom | 3 |
| 2012 | Directional high-pass filter for blurry image analysis
Jie Yang 0002, Qiang Wu 0001, Xiangjian He |
Signal Process. Image Commun. | 5 |
| 2011 | Image clustering using Particle Swarm OptimizationabstractThis paper proposes an image clustering algorithm using Particle Swarm Optimization (PSO) with two improved fitness functions. The PSO clustering algorithm can be used to find centroids of a user specified number of clusters. Two new fitness functions are proposed in this paper. The PSO-based image clustering algorithm with the proposed fitness functions is compared to the K-means clustering. Experimental results show that the PSO-based image clustering approach, using the improved fitness functions, can perform better than K-means by generating more compact clusters and larger inter-cluster separation. Man To Wong, Xiangjian He, Wei-Chang Yeh 0001 |
IEEE Congress on Evolutionary Computation | 2 |
| 2011 | An effective document image deblurring algorithmabstractDeblurring camera-based document image is an important task in digital document processing, since it can improve both the accuracy of optical character recognition systems and the visual quality of document images. Traditional deblurring algorithms have been proposed to work for natural-scene images. However the natural-scene images are not consistent with document images. In this paper, the distinct characteristics of document images are investigated. We propose a content-aware prior for document image deblurring. It is based on document image foreground segmentation. Besides, an upper-bound constraint combined with total variation based method is proposed to suppress the rings in the deblurred image. Comparing with the traditional general purpose deblurring methods, the proposed deblurring algorithm can produce more pleasing results on document images. Encouraging experimental results demonstrate the efficacy of the proposed method. Xiangjian He, Jie Yang 0002, Qiang Wu 0001 |
CVPR | 2 |
| 2011 | Multivariate Correlation Analysis Technique Based on Euclidean Distance Map for Network Traffic Characterization
Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001 |
ICICS | 3 |
| 2011 | Using context saliency for movie shot classificationabstractMovie shot classification is vital but challenging task due to various movie genres, different movie shooting techniques and much more shot types than other video domain. Variety of shot types are used in movies in order to attract audiences attention and enhance their watching experience. In this pa per, we introduce context saliency to measure visual attention distributed in keyframes for movie shot classification. Different from traditional saliency maps, context saliency map is generated by removing redundancy from contrast saliency and incorporating geometry constrains. Context saliency is later combined with color and texture features to generate feature vectors. Support Vector Machine (SVM) is used to classify keyframes into pre-defined shot classes. Different from the existing works of either performing in a certain movie genre or classifying movie shot into limited directing semantic classes, the proposed method has three unique features: 1) context saliency significantly improves movie shot classification; 2) our method works for all movie genres; 3) our method deals with the most common types of video shots in movies. The experimental results indicate that the proposed method is effective and efficient for movie shot classification. Min Xu 0001, Jinqiao Wang, Muhammad Abul Hasan, Xiangjian He, Changsheng Xu, Hanqing Lu, Jesse S. Jin |
ICIP | 4 |
| 2011 | An overcomplete pyramid representation for improved gsm image denoisingabstractRemoving noise from a digital image is a challenging problem. Application of Gaussian Scale Mixtures (GSM) in the wavelet domain has been reported to be one of the most effective denoising algorithms, published to date. In this paper we investigate the impact of overcomplete wavelet image representations on the GSM image denoising algorithm. We explore the desirable local characteristics of wavelet coefficients that can enhance the efficiency of GSM denoising and based on the findings, we devise an improved over-complete pyramid representation to enhance the GSM denoising performance. We present the experimental denoising results using the proposed pyramid representation, and they outperform state-of-the-art GSM denoising results reported in the literature. Damith J. Mudugamuwa, Wenjing Jia, Xiangjian He, Jie Yang 0002 |
ICME | 3 |
| 2011 | Denial-of-Service Attack Detection Based on Multivariate Correlation Analysis
Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001 |
ICONIP (3) | 3 |
| 2011 | Learning Global and Local Features for License Plate Detection
Sheng Wang 0003, Wenjing Jia, Qiang Wu 0001, Xiangjian He, Jie Yang 0002 |
ICONIP (3) | 4 |
| 2011 | Facial Expression Recognition on Hexagonal Structure Using LBP-Based Histogram Variances
Xiangjian He, Ruo Du, Wenjing Jia, Qiang Wu 0001, Wei-Chang Yeh 0001 |
MMM (2) | 2 |
| 2011 | More on Weak Feature: Self-correlate Histogram Distances
Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Wenjing Jia |
PSIVT (1) | 3 |
| 2010 | Canny Edge Detection Using Bilateral Filter on Real Hexagonal Structure
Xiangjian He, Daming Wei, Kin-Man Lam 0001, Wenjing Jia, Qiang Wu 0001 |
ACIVS (1) | 1 |
| 2010 | Efficient character segmentation on car license platesabstractIn this paper an improved hill climbing algorithm based method is presented to cut character out of the license plate images. Although there are many existing commercial LPR systems, with poor illumination conditions and moving vehicle the accuracy impaired. After examination and comparison of two different types of image segmentation approaches, the hill climbing algorithm based method gave a better image segmentation results. The hill climbing algorithm was modified by introducing automatic parameter determination and smart searching. After modification it efficiently detects the peaks (local maxima) that represent different clusters in the global histogram of an image. The process is successful by getting a clean license plate image removing all unwanted areas. While testing by the OCR software, the experimental results show a high accuracy of image segmentation and significantly higher recognition rate after non-character areas are removed. The recognition rate increased from about 30.6% before our proposed process to about 91.3% after all unwanted non-character areas are removed. Hence, the overall recognition accuracy of LPR was improved. Lihong Zheng, Junbin Gao, Xiangjian He |
ICARCV | 3 |
| 2010 | ECCH: A novel color coocurrence histogramabstractIn this paper, a novel color cooccurrence histogram method, named eCCH which stands for color cooccurrence histogram at edge points, is proposed to describe the spatial-color joint distribution of images. Unlike all existing ideas, we only investigate the color distribution of pixels located at the two sides of edge points on gradient direction lines. When measuring the similarity of two eCCHs, the Gaussian weighted histogram intersection method is adopted, where both identical and similar color pairs are considered to compensate color variations. Comparative experimental results demonstrate the performance of the proposed eCCH in terms of robustness to color variance and small computational complexity. Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
ICASSP | 2 |
| 2010 | A Two-Tier System for Web Attack Detection Using Linear Discriminant Method
Zhiyuan Tan 0001, Aruna Jamdagni, Xiangjian He, Priyadarsi Nanda, Ren Ping Liu 0001, Wenjing Jia, Wei-Chang Yeh 0001 |
ICICS | 3 |
| 2010 | Tensor error correction for corrupted values in visual dataabstractThe multi-channel image or the video clip has the natural form of tensor. The values of the tensor can be corrupted due to noise in the acquisition process. We consider the problem of recovering a tensor L of visual data from its corrupted observations X = L + S, where the corrupted entries S are unknown and unbounded, but are assumed to be sparse. Our work is built on the recent studies about the recovery of corrupted low-rank matrix via trace norm minimization. We extend the matrix case to the tensor case by the definition of tensor trace norm in [6]. Furthermore, the problem of tensor is formulated as a convex optimization, which is much harder than its matrix form. Thus, we develop a high quality algorithm to efficiently solve the problem. Our experiments show potential applications of our method and indicate a robust and reliable solution. Yin Li 0003, Yue Zhou 0005, Junchi Yan, Jie Yang 0002, Xiangjian He |
ICIP | 5 |
| 2010 | Action Recognition by Multiple Features and Hyper-Sphere Multi-class SVMabstractIn this paper we propose a novel framework for action recognition based on multiple features for improve action recognition in videos. The fusion of multiple features is important for recognizing actions as often a single feature based representation is not enough to capture the imaging variations (view-point, illumination etc.) and attributes of individuals (size, age, gender etc.). Hence, we use two kinds of features: i) a quantized vocabulary of local spatio-temporal (ST) volumes (cuboids and 2-D SIFT), and ii) the higher-order statistical models of interest points, which aims to capture the global information of the actor. We construct video representation in terms of local space-time features and global features and integrate such representations with hyper-sphere multi-class SVM. Experiments on publicly available datasets show that our proposed approach is effective. An additional experiment shows that using both local and global features provides a richer representation of human action when compared to the use of a single feature type. Jie Yang 0002, Yi Zhang 0001, Xiangjian He |
ICPR | 4 |
| 2010 | Intrusion detection using GSAD model for HTTP traffic on web servicesabstractIntrusion detection systems are widely used security tools to detect cyber-attacks and malicious activities in computer systems and networks. Hypertext Transport Protocol (HTTP) is used for new applications without much interference. In this paper, we focus on intrusion detection of HTTP traffic by applying pattern recognition techniques using our Geometrical Structure Anomaly Detection (GSAD) model. Experimental results reveal that features extracted from HTTP request using GSAD model can be used to distinguish anomalous traffic from normal traffic, and attacks carried out over HTTP traffic can be identified. We evaluate and compare our results with the results of PAYL intrusion detection systems for the test of DARPA 1999 IDS data set. The results show GSAD has high detection rates and low false positive rates. Aruna Jamdagni, Zhiyuan Tan 0001, Priyadarsi Nanda, Xiangjian He, Ren Ping Liu 0001 |
IWCMC | 4 |
| 2010 | A New Universal Generating Function Method for Estimating the Novel Multiresource Multistate Information Network ReliabilityabstractIn this article, we introduce a special novel multistate network that permits multiresource to be transmitted from the source node to multiple targets simultaneously without satisfying the flow conservation law. This network is called the multiresource multistate information network (MMIN). The one-to-many-targets (i.e. one-to-all-target-subset) reliability problem of the MMIN is considered next under limited cost and capacity constraints. A straightforward, exact algorithm derived from the universal generating function method (UGFM) is developed for this new problem. The correctness and computational complexity of the proposed UGFM will be analysed and proven. One example is given to illustrate how MMIN reliability is evaluated using the proposed UGFM. Wei-Chang Yeh 0001, Xiangjian He |
IEEE Trans. Reliab. | 2 |
| 2009 | Higher order prediction for sub-pixel motion estimationabstractEstimating motion between two frames of a video sequence, up to sub-pixel accuracy, is a critical task for many image processing applications. Efficient block matching algorithms were proposed for motion estimation up to pixel accuracy. Applying these fast block search algorithms to up-sampled and interpolated frames can produce good results but with significant increase in computations. To reduce the number of search points, and therefore the computational cost, quadratic prediction was proposed earlier to predict the location of minimum block matching error, and then to limit the search window to the vicinity of the predicted location. In this paper we investigate the typical behavior of block matching error surface and propose an improved higher order prediction that models the error surface more accurately, utilizing additional local image behavior. Initial experiments have proved promising results of about 50% more improvement in PSNR compared to quadratic prediction with only a marginal increase in the computational cost. Damith J. Mudugamuwa, Xiangjian He, Chung-Hyun Ahn, Jie Yang 0002 |
ICIP | 2 |
| 2009 | Web Service Locating Unit in RFID-Centric Anti-counterfeit SystemabstractThe problem of piracy has disturbed people’s daily life for hundreds of years and has not been relieved until now, though many existing anti-counterfeit solutions have been applied. However, due to the emergences of Radio Frequency IDentification (RFID) technologies, there is a more reliable alternative solution to construct authentication system. On the other hand, there arises another issue of how to simplify the deployment of RFID-centric anti-counterfeit system over the Internet. In this article, we propose an approach, Web Service Locating Unit (WSLU), to achieve this goal to manage numbers of RFID-centric authentication services (relied on web services). Zhiyuan Tan 0001, Xiangjian He, Priyadarsi Nanda |
ISPA | 2 |
| 2009 | Facial expression recognition using histogram variances facesabstractIn human's expression recognition, the representation of expression features is essential for the recognition accuracy. In this work we propose a novel approach for extracting expression dynamic features from facial expression videos. Rather than utilising statistical models e.g. Hidden Markov Model (HMM), our approach integrates expression dynamic features into a static image, the Histogram Variances Face (HVF), by fusing histogram variances among the frames in a video. The HVFs can be automatically obtained from videos with different frame rates and immune to illumination interference. In our experiments, for the videos picturing the same facial expression, e.g., surprise, happy and sadness etc., their corresponding HVFs are similar, even though the performers and frame rates are different. Therefore the static facial recognition approaches can be utilised for the dynamic expression recognition. We have applied this approach on the well-known Cohn-Kanade AU-Coded Facial Expression database then classified HVFs using PCA and Support Vector Machine (SVMs), and found the accuracy of HVFs classification is very encouraging. Ruo Du, Qiang Wu 0001, Xiangjian He, Wenjing Jia, Daming Wei |
WACV | 3 |
| 2008 | Using dynamic programming to match human behavior sequencesabstractThis paper proposed a new approach for recognition and matching the human behavior sequence. Each human behavior sequence is represented by its key postures to greatly reduce the computation time. Normalization is applied to all the behavior sequences key postures for matching. A dynamic time warping (DTW) algorithm is used to perform the alignment of two time series. Experiments are carried out on an open human behavior database and exciting results have been obtained. Yan Chen 0020, Qiang Wu 0001, Xiangjian He |
ICARCV | 3 |
| 2008 | An approach of canny edge detection with virtual hexagonal image structureabstractEdge detection plays an important role in the areas of image processing, multimedia and computer vision. Gradient-based edge detection is a straightforward method to identify the edge points in the original grey-level image. It is intuitive that, in the human vision system, the edge points always appear where the gradient magnitude assumes a maximum. Hexagonal structure is an image structure alternative to traditional square image structure. The geometrical arrangement of pixels on a hexagonal structure can be described as a collection of hexagonal pixels. Because all the existing hardware for capturing image and for displaying image are produced based on square structure, an approach that uses bilinear interpolation and tri-linear interpolation is applied for conversion between square and hexagonal structures. Based on this approach, an edge detection method is proposed. This method performs Gaussian filtering to suppress image noise and computes gradients on the hexagonal structure. The pixel edge strengths on the square structure are then estimated before Canny' edge detector is applied to determine the final edge map. The experimental results show that the proposed method improves the edge detection accuracy and efficiency. Xiangjian He, Wenjing Jia, Qiang Wu 0001 |
ICARCV | 1 |
| 2008 | Extracting key postures in a human action video sequenceabstractHuman key posture extraction from videos will benefit video storage, video retrieval, human action recognition, human behaviour understanding and so on. This paper presents an approach to select key postures from human action sequences using 2D information. There are two steps in the proposed method. Information measurement which is a kind of global feature of a frame is used to roughly find key posture candidates. Then, a body skeleton feature which is a kind of local feature is applied to select final key postures from the candidates obtained in the first step. The experiments show that the proposed method is efficient. Yan Chen 0020, Qiang Wu 0001, Xiangjian He, Chunhua Du, Jie Yang 0002 |
MMSP | 3 |
| 2008 | Segmentation of characters on car license platesabstractLicense plate recognition usually contains three steps, namely license plate detection/localization, character segmentation and character recognition. When reading characters on a license plate one by one after license plate detection step, it is crucial to accurately segment the characters. The segmentation step may be affected by many factors such as license plate boundaries (frames). The recognition accuracy will be significantly reduced if the characters are not properly segmented. This paper presents an efficient algorithm for character segmentation on a license plate. The algorithm follows the step that detects the license plates using an AdaBoost algorithm. It is based on an efficient and accurate skew and slant correction of license plates, and works together with boundary (frame) removal of license plates. The algorithm is efficient and can be applied in real-time applications. The experiments are performed to show the accuracy of segmentation. Xiangjian He, Lihong Zheng, Qiang Wu 0001, Wenjing Jia, Bijan Samali, Marimuthu Palaniswami |
MMSP | 1 |
| 2008 | Pedestrian detection using hybrid statistical featureabstractA novel approach for walking people detection is proposed in this paper, which is inspired by the idea of gait energy image (GEI). Unlike most of common human detection methods where usually a trained detector scans a single image and then generates a detection result, the proposed method detects people on a sequence of silhouettes which contain both appearance characteristics and motion characteristics. Thus, our method is more robust. Encouraging experimental results are obtained based on CASIA gait database and the additional non-human objects data. Qiang Wu 0001, Chunhua Du, Jie Yang 0002, Xiangjian He, Yan Chen 0020 |
MMSP | 4 |
| 2007 | Comparison of Image Conversions Between Square Structure and Hexagonal Structure
Xiangjian He, Tom Hintz |
ACIVS | 1 |
| 2007 | Parallel Edge Detection on a Virtual Hexagonal Structure
Xiangjian He, Wenjing Jia, Qiang Wu 0001, Tom Hintz |
GPC | 1 |
| 2007 | Local Binary Patterns for Human Detection on Hexagonal StructureabstractLocal binary pattern (LBP) was designed and has been widely used for efficient texture classification. LBP provides a simple and effective way to represent texture patterns. Uniform LBPs play an important role for LBP-based pattern/object recognition as they include majority of LBPs. On the other hand, Human detection based on Mahalanobis distance map (MDM) recognizes appearance of human based on geometrical structure. Each MDM shows a clear texture pattern that can be classified using LBPs. In this paper, we compute LBPs of MDMs on a hexagonal structure. The circular pixel arrangement in hexagonal structure results in higher accuracy for LBP representation than on square structure. Chi-square as a measure is used for human detection based on uniform LBPs obtained. We show that our method using LBPs built on MDMs has a higher human detection rate and a lower false positive rate compared to the method merely based on MDMs. We will also show using experimental results that LBPs on hexagonal structure lead to more robust human classification. Xiangjian He, Yan Chen 0020, Qiang Wu 0001, Wenjing Jia |
ISM | 1 |
| 2007 | Editorial
Xiangjian He |
J. Netw. Comput. Appl. | 1 |
| 2007 | Region-based license plate detection
Wenjing Jia, Huaifeng Zhang, Xiangjian He |
J. Netw. Comput. Appl. | 3 |
| 2006 | Symmetric Color Ratio in Spiral Architecture
Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001 |
ACCV (2) | 3 |
| 2006 | Estimation of Internal and External Parameters for Camera Calibration Using 1D PatternabstractCamera calibration is to estimate the intrinsic and extrinsic parameters of a camera. Most of object-based calibration methods used 3D or 2D pattern. A novel and more flexible 1D object-based calibration was introduced only a couple of years ago, but merely for estimation of intrinsic parameters. The estimation of extrinsic papers is essential when multiple cameras are involved for simultaneously taking images from different view angles and when the knowledge of relative locations between the cameras is required. Though it is relatively simple using 2D or 3D calibration pattern, the estimation of extrinsic parameters is not obvious using 1D pattern. In this paper, we will perform a 1D camera calibration involving both intrinsic and extrinsic parameters. Xiangjian He, Huaifeng Zhang, Namho Hur, Jinwoong Kim, Qiang Wu 0001, Taeone Kim |
AVSS | 1 |
| 2006 | A Comparison on Histogram Based Image Matching MethodsabstractUsing colour histogram as a stable representation over change in view has been widely used for object recognition. In this paper, three newly proposed histogram-based methods are compared with other three popular methods, including conventional histogram intersection (HI) method, Wong and Cheung's merged palette histogram matching (MPHM) method, and Gevers' colour ratio gradient (CRG) method. These methods are tested on vehicle number plate images for number plate classification. Experimental results disclose that, the CRG method is the best choice in terms of speed, and the GWHI method can give the best classification results. Overall, the CECH method produces the best performance when both speed and classification performance are concerned. Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001 |
AVSS | 3 |
| 2006 | Car Plate Detection Using Cascaded Tree-Style Learner Based on Hybrid Object FeaturesabstractCar plate detection is a key component in automatic license plate recognition system. This paper adopts an enhanced cascaded tree style learner framework for car plate detection using the hybrid object features including the simple statistical features and Harr-like features. The statistical features are useful for simplifying the process on cascade classifier. The cascaded tree-style detector design will further reduce the false alarm and the false dismissal while retaining a high detection ratio. The experimental results obtained by the proposed algorithm exhibit the encouraging performance. Qiang Wu 0001, Huaifeng Zhang, Wenjing Jia, Xiangjian He, Jie Yang 0002, Tom Hintz |
AVSS | 4 |
| 2006 | Number Plate Recognition Based on Support Vector MachinesabstractAutomatic number plate recognition method is required due to increasing traffic management. In this paper, we first briefly review some knowledge of Support Vector Machines (SVMs). Then a number plate recognition algorithm is proposed. This algorithm employs an SVM to recognize numbers. The algorithm starts from a collection of samples of numbers from number plates. Each character is recognized by an SVM, which is trained by some known samples in advance. In order to recognize a number plate correctly, all numbers are tested one by one using the trained model. The recognition results are achieved by finding the maximum value between the outputs of SVMs. In this paper, experimental results based on SVMs are given. From the experimental results, we can make the conclusion that SVM is bettr than others such as inductive learning-based number recognition Lihong Zheng, Xiangjian He |
AVSS | 2 |
| 2006 | Uniformly Partitioning Images on Virtual Hexagonal StructureabstractHexagonal structure is different from the traditional square structure for image representation. The geometrical arrangement of pixels on hexagonal structure can be described in terms of a hexagonal grid. Uniformly separating image into seven similar copies with a smaller scale has commonly been used for parallel and accurate image processing on hexagonal structure. However, all the existing hardware for capturing image and for displaying image are produced based on square architecture. It has become a serious problem affecting the advanced research based on hexagonal structure. Furthermore, the current techniques used for uniform separation of images on hexagonal structure do not coincide with the rectangular shape of images. This has been an obstacle in the use of hexagonal structure for image processing. In this paper, we briefly review a newly developed virtual hexagonal structure that is scalable. Based on this virtual structure, algorithms for uniform image separation are presented. The virtual hexagonal structure retains image resolution during the process of image separation, and does not introduce distortion. Furthermore, images can be smoothly and easily transferred between the traditional square structure and the hexagonal structure while the image shape is kept in rectangle Xiangjian He, Huaqing Wang, Namho Hur, Wenjing Jia, Qiang Wu 0001, Jinwoong Kim, Tom Hintz |
ICARCV | 1 |
| 2006 | Learning-Based Number Recognition on Spiral ArchitectureabstractIn this paper, a number recognition algorithm is proposed on spiral architecture, a hexagonal image structure. This algorithm employs RULES-3 inductive learning method to recognize numbers. The algorithm starts from a collection of samples of numbers from number plates. Edge maps of the samples are then detected based on spiral architecture. A set of rules are extracted using these samples by RULES-3. The rules describe the frequencies of 9 different edge masks appearing in the samples. Each mask is a cluster of 7 hexagonal pixels. In order to recognize a number plate, all numbers are tested one by one using the extracted rules. The number recognition is achieved by counting the frequencies of the 9 masks. In this paper, a comparison between results based on rectangular structure and the results based on spiral architecture is given. From the experimental results, we can make the conclusion that Spiral Architecture is better than rectangular structure for inductive learning-based number recognition Lihong Zheng, Xiangjian He, Qiang Wu 0001, Tom Hintz |
ICARCV | 2 |
| 2006 | A New Approach for SA-Based Fractal Image CompressionabstractSpiral Architecture based fractal image compression is proposed in this paper. Perceptually, a new definition of range block and domain block is presented on such enhanced image structure. Compared with the common square image architecture, spiral architecture provides higher fidelity to fractal image compression, which is demonstrated by the experimental results. Huaqing Wang, Qiang Wu 0001, Xiangjian He, Tom Hintz |
ICIP | 3 |
| 2006 | Image Matching Using Colour Edge Cooccurrence HistogramsabstractIn this paper, a novel colour edge cooccurrence histogram (CECH) method is proposed to match images by measuring similarities between their CECH histograms. Unlike the previous colour edge cooccurrence histogram proposed by Crandall and Luo (2004 ) we only investigate those pixels which are located at the two sides of edge points in their gradient direction lines and at a distance away from the edge points. When measuring similarities between two CECH histograms, a newly proposed Gaussian weighted histogram intersection (GWHI) method is extended for this purpose. Both identical colour pairs and similar colour pairs are taken into account in our algorithm, and the weights are decided by the larger distance between two colour pairs involved in matching. The proposed algorithm is tested for matching vehicle number plate images captured under various illumination conditions. Experimental results demonstrate that the proposed algorithm can be used to compare images in real-time, and is robust to illumination variations and insensitive to the model images selected. Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001 |
SMC | 3 |
| 2006 | A Fast Algorithm for License Plate Detection in Various ConditionsabstractThis paper proposes a fast algorithm detecting license plates in various conditions. There are three main contributions in this paper. The first contribution is that we define a new vertical edge map, with which the license plate detection algorithm is extremely fast. The second contribution is that we construct a cascade classifier which is composed of two kinds of classifiers. The classifiers based on statistical features decrease the complexity of the system. They are followed by the classifiers based on Haar-features, which make it possible to detect license plate in various conditions. Our algorithm is robust to the variance of the illumination, view angle, the position, size and color of the license plates when working in complex environment. The third contribution is that we experimentally analyze the relations of the scaling factor with detection rate and processing time. On the basis of the analysis, we select the optimal scaling factor in our algorithm. In the experiments, both high detection rate (with low false positive rate) and high speed are achieved when the algorithm is used to detect license plates in various complex conditions. Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
SMC | 3 |
| 2006 | Real-Time License Plate Detection Under Various Conditions
Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
UIC | 3 |
| 2005 | Bi-Lateral Filtering Based Edge Detection on Hexagonal ArchitectureabstractEdge detection plays an important role in image processing but is still an open problem. This paper presents a novel edge detection method based on bi-lateral filtering which achieves better performance than single Gaussian filtering. In this form of filtering, both spatial closeness and intensity similarity of pixels are considered in order to preserve important visual cues provided by edges and reduce the sharpness of transitions in intensity values as well. In addition, the edge detection method proposed in this paper is achieved on hexagonally sampled images. Due to the compact and circular nature of the hexagonal lattice, a better quality edge map is obtained on hexagonal architecture than common edge detection on square architecture. Experimental results using our proposed method in this paper exhibit encouraging performance. Qiang Wu 0001, Xiangjian He, Tom Hintz |
ICASSP (2) | 2 |
| 2005 | Modified Color Ratio GradientabstractColor ratio gradient is an efficient method used for color image retrieval and object recognition, which is shown to be illumination-independent and geometry-insensitive when tested on scenery images. However, color ratio gradient produces unsatisfied matching result while dealing with relatively uniform objects without rich color texture. In addition, performance of color ratio gradient degenerates while processing unsaturated color image objects. In this paper, a scheme with modified color ratio gradient is presented, which addresses the two problems above. Experimental results using the proposed method in this paper exhibit more robust performance Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
MMSP | 3 |
| 2003 | Complete Image Partitioning on Spiral Architecture
Qiang Wu 0001, Xiangjian He, Tom Hintz, Yuhuang Ye |
ISPA | 2 |
| 2001 | A Skeleton Algorithm on Clusters for Image Edge DetectionabstractImage edge detection in computer vision and image processing is a process which detects one kind of significant feature in an image that appears as large delta values in intensities. In this paper, a parallel algorithmic skeleton for edge detection is proposed based on the Spiral Architecture and the Gaussian multi-scale theory. UNIX-based network programming mechanisms in C are used for the implementation on a cluster of Sun-workstations. Our work provides an efficient algorithm for edge detection and is robust to noise. Xiangjian He, Tom Hintz, Qiang Wu 0001 |
IPDPS | 1 |
| 1998 | Replicated shared object model for parallel edge detection algorithm based on spiral architecture
Xiangjian He, Tom Hintz, Ury Szewcow |
Future Gener. Comput. Syst. | 1 |