Zhijun Fang 0001

dblp:21/4944-1 · also Zhi-Jun Fang 0001, ZhiJun Fang 0001 · DBLP profile ↗
← Back
111ranked-venue papers
2as first author
80since 2021 · last 2026
0000-0001-8563-5678ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 1 first-author · 31 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Security and privacy · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Generating Sketches in a Hierarchical Auto-Regressive Process for Flexible Sketch Drawing Manipulation at Stroke-Level
abstract
Generating sketches with specific patterns as expected, i.e., manipulating sketches in a controllable way, is a popular task. Recent studies control sketch features at stroke-level by editing values of stroke embeddings as conditions. However, in order to provide generator a global view about what a sketch is going to be drawn, all these edited conditions should be collected and fed into generator simultaneously before generation starts, i.e., no further manipulation is allowed during sketch generating process. In order to realize sketch drawing manipulation more flexibly, we propose a hierarchical auto-regressive sketch generating process. Instead of generating an entire sketch at once, each stroke in a sketch is generated in a three-staged hierarchy: 1) predicting a stroke embedding to represent which stroke is going to be drawn, and 2) anchoring the predicted stroke on the canvas, and 3) translating the embedding to a sequence of drawing actions to form the full sketch. Moreover, the stroke prediction, anchoring and translation are proceeded auto-regressively, i.e., both the recently generated strokes and their positions are considered to predict the current one, guiding model to produce an appropriate stroke at a suitable position to benefit the full sketch generation. It is flexible to manipulate stroke-level sketch drawing at any time during generation by adjusting the exposed editable stroke embeddings.
Sicong Zang, Shuhui Gao, Zhijun Fang 0001
AAAI3
2026 KGFR: Knowledge-infused GraphFuseRec - a dual-channel graph fusion recommender for industrial expert systems
Shize Yan, Zhijun Fang 0001
Appl. Intell.2
2026 Entity completion for industrial knowledge graph based on zero-shot learning
Yin Cai, Zhijun Fang 0001, Anjie Wang, Zheyi Cheng
Data Min. Knowl. Discov.2
2026 Multigranularity-3DQA: Dynamic multi-granularity perception and gated dual-stage encoder-decoder for 3D question answering
Letao Zhang, Weibing Wan, Zhijun Fang 0001
Expert Syst. Appl.4
2026 HgCA: Hypergraph neural network with cross-attention for point cloud analysis
Xinxin Hou, Hui Feng 0001, Zhengpin Li, Shubo Zhou, Jian Wang 0016, Zhijun Fang 0001, Xueqin Jiang 0001
Neurocomputing6
2026 Continual few-shot relation extraction via multi-task balanced dual-branch network
Chenyang Shan, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014
Neurocomputing3
2026 Optimizing knowledge reasoning with combined graph neural networks and large language models
Weibing Wan, Zhijun Fang 0001
Neurocomputing4
2026 Large-scale generative dataset, benchmark, and edge-empowered transformer for crack detection
Haoran He, Zhijun Fang 0001
Inf. Sci.5
2026 Ibcl: instance-aware bias calibration with contrastive learning for visual question answering
Chuanfeng Liu, Benxue Sun, Zhijun Fang 0001
Multim. Syst.4
2026 MAFIFusion: a multi-attention and feature interaction network for infrared and visible image fusion
Haochen Yu, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014, Yadong Zhu
Multim. Syst.3
2026 MIMAR-OSA: Enhancing obstructive sleep apnea diagnosis through multimodal data integration and missing modality reconstruction
Xihe Qiu, Yingchen Wei, Xiaoyu Tan, Weidi Xu, Jingru Ma, Zhijun Fang 0001
Pattern Recognit.8
2026 Enhancing lightweight image super-resolution with hybrid convolution and attention
Hanwen Shi, Shubo Zhou, Yinghua Xie, Zhijun Fang 0001, Xueqin Jiang 0001
Pattern Recognit. Lett.5
2026 ReMALIS: Inference-Guided Intention Propagation for Multiagent Stochastic Task Coordination With Large Language Models in Complex Networks Domains
Xihe Qiu, Haoyu Wang 0011, Xiaoyu Tan, Yujie Xiong, Zhijun Fang 0001
IEEE Trans. Comput. Soc. Syst.6
2026 Learning Monocular Depth via Cascaded Iterative Refinement in Visual-Echo Scenes
abstract
In recent years, integrating multimodal information, particularly visual and echo data, has shown great promise for improving depth estimation performance. While existing works demonstrate that combining binaural echo features with image attributes can enhance depth estimation, they often use rudimentary feature alignment and fusion methods, failing to fully exploit the complementary nature of cross-modal information and limiting integration effectiveness. To address these challenges, this paper introduces an innovative multimodal fusion framework. First, the framework incorporates a combination of multi-scale self-attention and cross-attention mechanisms, establishing correlations between features and facilitating cohesive interactions between the visual and echo domains. Furthermore, we propose an incremental feature updating mechanism based on Convolutional Gated Recurrent Units (ConvGRU), which implements cascaded iterative optimization, integrating contextual features with the multi-scale fused features from both echo and image modalities. In each iteration, the framework preserves contextual information from previous steps while employing a multi-level loss function to guide result updates. This approach effectively captures spatial structural information and progressively enhances depth estimation accuracy. Comprehensive experimental evaluations on the Replica, Matterport3D and BatVision (BV1) datasets validate the effectiveness of the proposed method. Comparative analyses with state-of-the-art monocular plus echo methods underscore the superior performance achievable through this novel framework.
Anjie Wang, Zhijun Fang 0001, Leidong Fan, Guibiao Liao, Siwei Ma 0001, Jenq-Neng Hwang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Learning Adaptive Basis Fonts to Fuse Content Features for Few-Shot Font Generation
Keyang Lin, Zhijun Fang 0001, Sicong Zang
CVM (3)2
2025 Sparse-view 3D Open-vocabulary Gaussian Splatting via Collaborative Contrastive Learning
abstract
3D Gaussian Splatting-based Open-vocabulary 3D segmentation has shown impressive performance with dense input images. However, existing methods exhibit poor results when confronted with sparse inputs, primarily due to limited overlap among input views and insufficient view supervision provided. To tackle these challenges, we propose SpContrast, a novel framework that creates additional semantic constraints to enhance sparse-view 3D open-vocabulary segmentation. First, we introduce Collaborative Contrastive Learning (CCL), which creates instructive multi-view semantic constraints by collaboratively mining semantic interactions between training and online-rendered novel views. Motivated by the principle that semantically consistent features should converge and divergent ones separate, CCL establishes cross-view contrastive constraints to enhance semantic coherence. Second, to alleviate the adverse impact of false negative samples caused by semantic inconsistencies within the same object, we present Region-aware Negative Sampling (RNS). RNS rectifies these false negative samples, and treats them as hard samples during our contrastive optimization, leading to improved object completeness and more accurate segmentation. Extensive experiments on challenging sparse-input datasets, including Replica and ScanNet, demonstrate the superiority of SpContrast, achieving 7.6% and 8.3% mIoU improvements for 3D open-vocabulary segmentation.
Guibiao Liao, Anjie Wang, Mingxuan Chen, Zhijun Fang 0001
ICME4
2025 A Dual-Agent Framework for Condition-Based Maintenance of Production Systems
Linsheng Guo, Bo Huang 0014, Zhijun Fang 0001
IEA/AIE (2)5
2025 MST-SGAN-KGQA: An Approach for Industrial Knowledge Graph Quality Assessment
Linsheng Guo, Bo Huang 0014, Zhijun Fang 0001
IEA/AIE (2)5
2025 Adversarial Learning Based Error Detection for Industrial Knowledge Graphs
Linsheng Guo, Bo Huang 0014, Zhijun Fang 0001
IEA/AIE (2)5
2025 Inverse Farthest Point Sampling (IFPS): A Universal and Hierarchical Shell Representation for Discrete Data
Nayu Ding, Long Wan, Zhijun Fang 0001, Shen Cai, Lin Gao 0004
ICMR6
2025 Learning Basis Fonts on a Hypersphere to Guide Gated Content Features Fusion for Few-Shot Font Generation
Keyang Lin, Zhijun Fang 0001, Sicong Zang
PRCV (4)3
2025 An end-to-end audio classification framework with diverse features for obstructive sleep apnea-hypopnea syndrome diagnosis
Bin Li 0091, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001
Appl. Intell.6
2025 Human-object interaction detection based on adaptive contrastive learning and class-specific feature enhancement
Huanchun Peng, Kejun Xue, Xincheng Wang 0001, Yongbin Gao, Zhijun Fang 0001, Chenmou Wu
Appl. Intell.5
2025 Equipping sketch patches with context-aware positional encoding for graphic sketch representation
Sicong Zang, Zhijun Fang 0001
Comput. Vis. Image Underst.2
2025 Learning prototypes from background and latent objects for few-shot semantic segmentation
Yicong Wang, Rong Huang 0003, Shubo Zhou, Xueqin Jiang 0001, Zhijun Fang 0001
Knowl. Based Syst.5
2025 CLIP guided image caption decoding based on monte carlo tree search
Guangsheng Luo, Zhijun Fang 0001, JianLing Liu, YiFanBai Bai
Multim. Syst.2
2025 MSCC-RetNet: a multi-scale color corrected retinex network for underwater image enhancement
Benxue Sun, Mingxuan Chen, Liming Hu, Anjie Wang, Zhijun Fang 0001
Multim. Syst.5
2025 Fully exploring object relation interaction and hidden state attention for video captioning
Feiniu Yuan, Sipei Gu, Xiangfen Zhang, Zhijun Fang 0001
Pattern Recognit.4
2025 Semantic and Saliency-Aware Scalable Image Coding Toward Human-Machine Collaboration
abstract
With the widespread deployment of intelligent vision applications, a substantial amount of visual data is being transmitted to machines for automated analysis to alleviate the human burden. However, research on the interaction between human and machine vision remains limited, significantly constraining the collaborative potential of both systems. In this paper, we investigate task-oriented high- and low-level representations, and how they can be used to construct a scalable coding model for human-machine collaboration. First, we propose a semantic-aware base layer combined with an implicit semantic module, designed to encourage the network to learn compact representations for machine tasks under semantic consistency constraints. Second, we propose a saliency-aware enhancement layer, which navigates the compression of visual signals with a saliency prior derived from the base layer so as to construct high-quality visual perceptions that fit the collaborative scene. By interacting and recombining the decoupled features, our model further bridges the gap between high- and low-level representations, so that the learned representations enjoy both machine analysis and visual reconstruction. Experimental results demonstrate that the proposed method outperforms state-of-the-art machine vision codecs on several machine vision tasks, and can also achieve comparable or even better reconstruction quality, while maintaining a modest bit rate cost.
Tengyao Cui, Yihan Wang 0008, Zhijun Fang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 ECKT: enhancing cross-task knowledge transfer in continual few-shot relation extraction
Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao
J. Supercomput.3
2025 IGES-RCI: Improved Greedy Equivalence Search and Recursive Causal Inference for Industrial Equipment Failure Prediction
abstract
Predicting equipment failures plays a pivotal role in minimizing maintenance costs and boosting production efficiency within the industrial sector. This paper introduces a novel approach that integrates Causal Inference with predictive modeling to enhance prediction accuracy, tackling key challenges such as noise interference, insufficient causal validation, and missing data. We first validate the causal connections identified by the Greedy Equivalence Search algorithm using conditional mutual information to strengthen the reliability of the causal graph. An information bottleneck strategy is then employed to isolate essential causal features, effectively filtering out irrelevant noise and refining the causal structure. Crucially, in the actual prediction phase, we propose a recursive causal inference-based imputation method to handle missing data, leveraging the causal graph to iteratively infer and fill gaps, thereby improving data completeness and prediction accuracy. Experimental results demonstrate that the proposed method significantly outperforms existing approaches, exhibiting superior accuracy and robustness in managing complex industrial datasets.
Weibing Wan, Zhijun Fang 0001
IEEE Trans. Knowl. Data Eng.3
2025 Depth-guided color correction and multi-scale Retinex network for underwater image enhancement
Zhan Hu, Juan Zhang 0001, Yongbin Gao, Bo Huang 0014, Zhijun Fang 0001
Vis. Comput.5
2024 Multi-modal Scene Global Fusion Framework for Enhanced Depth Estimation
Anjie Wang, Xujun Wei, Mingxuan Chen, Yongbin Gao, Zhijun Fang 0001, Siwei Ma 0001
ICONIP (9)6
2024 LiDUT-Depth: A Lightweight Self-supervised Depth Estimation Model Featuring Dynamic Upsampling and Triplet Loss Optimization
Hao Jiang 0014, Zhijun Fang 0001, Xuan Shao, Jenq-Neng Hwang
ICPR (16)2
2024 Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor Distillation
abstract
Correspondence matching plays a crucial role in numerous robotics applications. In comparison to conventional hand-crafted methods and recent data-driven approaches, there is significant interest in plug-and-play algorithms that make full use of pre-trained backbone networks for multi-scale feature extraction and leverage hierarchical refinement strategies to generate matched correspondences. The primary focus of this paper is to address the limitations of deep feature matching (DFM), a state-of-the-art (SoTA) plug-and-play correspondence matching approach. First, we eliminate the pre-defined threshold employed in the hierarchical refinement process of DFM by leveraging a more flexible nearest neighbor search strategy, thereby preventing the exclusion of repetitive yet valid matches during the early stages. Our second technical contribution is the integration of a patch descriptor, which extends the applicability of DFM to accommodate a wide range of backbone networks pre-trained across diverse computer vision tasks, including image classification, semantic segmentation, and stereo matching. Taking into account the practical applicability of our method in real-world robotics applications, we also propose a novel patch descriptor distillation strategy to further reduce the computational complexity of correspondence matching. Extensive experiments conducted on three public datasets demonstrate the superior performance of our proposed method. Specifically, it achieves an overall performance in terms of mean matching accuracy of 0.68, 0.92, and 0.95 with respect to the tolerances of 1, 3, and 5 pixels, respectively, on the HPatches dataset, outperforming all other SoTA algorithms. Our source code, demo video, and supplement are publicly available at mias.group/GCM.
Ziwei Long, Yanting Zhang 0001, Jin Wu 0002, Zhijun Fang 0001, Rui Fan 0001
ICRA5
2024 Taming Diffusion for Fashion Clothing Generation with Versatile Condition
Yanting Zhang 0001, Jingyi Guo, Cairong Yan, Zhijun Fang 0001
PRCV (5)4
2024 A Medical Image Segmentation Method based on Multi-scale Features and Contour Loss Constrain (S)
abstract
In recent years, deep learning has made breakthroughs in medical image segmentation, especially the U-Net architecture, which is becoming a benchmark for various medical image segmentation tasks due to the accuracy of its segmentation results.Although U-Net has achieved great success in many medical image segmentation tasks, it is still unsatisfactory in segmenting object boundaries and small objects.This is due to the fact that segmentation networks gradually lose information, especially edge information and small object information, during the process of convolution and downsampling of features.In order to solve the above problems, we design a novel method that uses a multi-scale module as a feature extractor in the contraction path of the Ushaped structure, which better captures the scale changes of the target object by acquiring image features at different scales; in the prediction stage, a contour prediction branch is constructed to constrain the loss of the target's contour, so that the segmentation network pays more attention to the boundaries of the target.We have validated the performance of our method on the Automated Cardiac Diagnosis Challenge (ACDC) and the spleen segmentation tasks of the Medical Segmentation Decathlon (MSD).The results show that our method obtained the best 95% Hausdorff Distance (HD) metrics on both the ACDC dataset and the Spleen dataset, as well as being quite competitive with other state-of-the-art methods in terms of Dice scores.
Jian Niu, Zhijun Fang 0001, Xihe Qiu
SEKE2
2024 Dynamic Convolution Based Intelligent Algorithm for YOLOv5 Underwater Target Detection
abstract
With the growing importance of marine resources and the increasing demand for exploration of underwater environments, underwater target detection technology has become one of the key technologies in the fields of ocean engineering, underwater archaeology, and intelligent agriculture. However, due to the complexity and uncertainty of underwater environments, such as light attenuation, water turbidity, and dynamic changes of water currents, current target detection methods often perform poorly in underwater scenes. To solve the corresponding problems, this paper proposes the YOLOv5_OD_Conv model, which aims to improve the accuracy and generalisation ability of the model by introducing the OD_Conv full-dimensional dynamic convolution in the YOLOv5 Neck part. Simulation and experimental results show that the proposed method increases the detection accuracy P by 1.05%, the precision mAP0.5 by 1.5%, and the recall R by 0.43% compared to YOLOv5s. The improvement of detection effect is obvious, which proves the effectiveness of the method.
Jialing Jiang, Bo Huang 0014, Zhijun Fang 0001, Yongbin Gao
SoMeT3
2024 Adaptive multimodal prompt for human-object interaction with local feature enhanced transformer
Kejun Xue, Yongbin Gao, Zhijun Fang 0001, Mingxuan Chen, Chenmou Wu
Appl. Intell.3
2024 A lightweight RGB superposition effect adjustment network for low-light image enhancement and denoising
Pei-Dong Chen, Juan Zhang 0001, Yongbin Gao, Zhijun Fang 0001, Jenq-Neng Hwang
Eng. Appl. Artif. Intell.4
2024 A pyramid Gaussian pooling based CNN and transformer hybrid network for smoke segmentation
abstract
Abstract Visual smoke semantic segmentation is a challenging task due to semi‐transparency, variable shapes, and complex textures of smoke. To improve segmentation performance, a convolutional neural network and transformer hybrid network are proposed based on pyramid Gaussian pooling (PGP) for smoke segmentation. In order to utilize low‐pass filtering to suppress noise, a PGP method is designed. Then, the output of PGP is reshaped to construct a set of visual tokens for transformers, thus a PGP‐transformer module is presented to make full use of the self‐attention mechanism. Finally, the PGP‐transformer module is inserted into the U‐shaped architecture with skip connections. A large number of experiments have proved that the method is significantly superior to existing state‐of‐the‐art algorithms on virtual and real smoke datasets, and ablation experiments have also verified the effectiveness of the proposed modules.
Guiqian Wang, Feiniu Yuan, Hongdi Li, Zhijun Fang 0001
IET Image Process.4
2024 Occupancy map-based low complexity motion prediction for video-based point cloud compression
Yihan Wang 0008, Tengyao Cui, Zhijun Fang 0001
J. Vis. Commun. Image Represent.4
2024 Dual-stream multi-label image classification model enhanced by feature reconstruction
Liming Hu, Mingxuan Chen, Anjie Wang, Zhijun Fang 0001
Multim. Syst.4
2024 Entity alignment based on informative neighbor sampling and multi-embedding graph matching
Yongbin Gao, Zhijun Fang 0001
Multim. Tools Appl.3
2024 Self-Enhanced Attention for Image Captioning
abstract
Abstract Image captioning, which involves automatically generating textual descriptions based on the content of images, has garnered increasing attention from researchers. Recently, Transformers have emerged as the preferred choice for the language model in image captioning models. Transformers leverage self-attention mechanisms to address gradient accumulation issues and eliminate the risk of gradient explosion commonly associated with RNN networks. However, a challenge arises when the input features of the self-attention mechanism belong to different categories, as it may result in ineffective highlighting of important features. To address this issue, our paper proposes a novel attention mechanism called Self-Enhanced Attention (SEA), which replaces the self-attention mechanism in the decoder part of the Transformer model. In our proposed SEA, after generating the attention weight matrix, it further adjusts the matrix based on its own distribution to effectively highlight important features. To evaluate the effectiveness of SEA, we conducted experiments on the COCO dataset, comparing the results with different visual models and training strategies. The experimental results demonstrate that when using SEA, the CIDEr score is significantly higher compared to the scores obtained without using SEA. This indicates the successful addressing of the challenge of effectively highlighting important features with our proposed mechanism.
Qingyu Sun, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao
Neural Process. Lett.3
2024 F-SCP: An automatic prompt generation method for specific classes based on visual language pre-training models
Baihong Han, Zhijun Fang 0001, Hamido Fujita, Yongbin Gao
Pattern Recognit.3
2024 Enhancing Few-Shot Out-of-Distribution Detection With Pre-Trained Model Features
abstract
Ensuring the reliability of open-world intelligent systems heavily relies on effective out-of-distribution (OOD) detection. Despite notable successes in existing OOD detection methods, their performance in scenarios with limited training samples is still suboptimal. Therefore, we first construct a comprehensive few-shot OOD detection benchmark in this paper. Remarkably, our investigation reveals that Parameter-Efficient Fine-Tuning (PEFT) techniques, such as visual prompt tuning and visual adapter tuning, outperform traditional methods like fully fine-tuning and linear probing tuning in few-shot OOD detection. Considering that some valuable information from the pre-trained model, which is conducive to OOD detection, may be lost during the fine-tuning process, we reutilize features from the pre-trained models to mitigate this issue. Specifically, we first propose a training-free approach, termed uncertainty score ensemble (USE). This method integrates feature-matching scores to enhance existing OOD detection methods, significantly narrowing the gap between traditional fine-tuning and PEFT techniques. However, due to its training-free property, this method is unable to improve in-distribution accuracy. To this end, we further propose a method called Domain-Specific and General Knowledge Fusion (DSGF) to improve few-shot OOD detection performance and ID accuracy under different fine-tuning paradigms. Experiment results demonstrate that DSGF enhances few-shot OOD detection across different fine-tuning strategies, shot settings, and OOD detection methods. We believe our work can provide the research community with a novel path to leveraging large-scale visual pre-trained models for addressing FS-OOD detection. The code will be released.
Jiuqing Dong, Yongbin Gao, Zhijun Fang 0001
IEEE Trans. Image Process.6
2024 TV-Net: A Structure-Level Feature Fusion Network Based on Tensor Voting for Road Crack Segmentation
abstract
Pavement cracks are a common and significant problem for intelligent pavement maintainment. However, the features extracted in pavement images are often texture-less, and noise interference can be high. Segmentation using traditional convolutional neural network training can lose feature information when the network depth goes larger, which makes accurate prediction a challenging topic. To address these issues, we propose a new approach that features an enhanced tensor voting module and a customized pixel-level pavement crack segmentation network structure, called TV-Net. We optimize the tensor voting framework and find the relationship between tensor scale factors and crack distributions. A tensor voting fusion module is introduced to enhance feature maps by incorporating significant domain maps generated by tensor voting. Additionally, we propose a structural consistency loss function to improve segmentation accuracy and ensure consistency with the structural characteristics of the cracks obtained through tensor voting. The sufficient experimental analysis demonstrates that our method outperforms existing mainstream pixel-level segmentation networks on the same road crack dataset. Our proposed TV-Net has an excellent performance in avoiding noise interference and strengthening the structure of the fracture site of pavement cracks. Code is available at https://github.com/sues-vision/ TV-Net.git.
Wenwen Zheng, Zhijun Fang 0001, Yongbin Gao
IEEE Trans. Intell. Transp. Syst.3
2024 Cluster knowledge-driven vertical federated learning
Zilong Yin, Xiaoli Zhao 0003, Haoyu Wang 0011, Zhijun Fang 0001
J. Supercomput.6
2024 Monocular Depth and Ego-motion Estimation with Scale Based on Superpixel and Normal Constraints
abstract
Three-dimensional perception in intelligent virtual and augmented reality (VR/AR) and autonomous vehicles (AV) applications is critical and attracting significant attention. The self-supervised monocular depth and ego-motion estimation serves as a more intelligent learning approach that provides the required scene depth and location for 3D perception. However, the existing self-supervised learning methods suffer from scale ambiguity, boundary blur, and imbalanced depth distribution, limiting the practical applications of VR/AR and AV. In this article, we propose a new self-supervised learning framework based on superpixel and normal constraints to address these problems. Specifically, we formulate a novel 3D edge structure consistency loss to alleviate the boundary blur of depth estimation. To address the scale ambiguity of estimated depth and ego-motion, we propose a novel surface normal network for efficient camera height estimation. The surface normal network is composed of a deep fusion module and a full-scale hierarchical feature aggregation module. Meanwhile, to realize the global smoothing and boundary discriminability of the predicted normal map, we introduce a novel fusion loss which is based on the consistency constraints of the normal in edge domains and superpixel regions. Experiments are conducted on several benchmarks, and the results illustrate that the proposed approach outperforms the state-of-the-art methods in depth, ego-motion, and surface normal estimation.
Junxin Lu, Yongbin Gao, Jieyu Chen, Jenq-Neng Hwang, Hamido Fujita, Zhijun Fang 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Self-Supervised Learning of Depth and Ego-Motion for 3D Perception in Human Computer Interaction
abstract
3D perception of depth and ego-motion is of vital importance in intelligent agent and Human Computer Interaction (HCI) tasks, such as robotics and autonomous driving. There are different kinds of sensors that can directly obtain 3D depth information. However, the commonly used Lidar sensor is expensive, and the effective range of RGB-D cameras is limited. In the field of computer vision, researchers have done a lot of work on 3D perception. While traditional geometric algorithms require a lot of manual features for depth estimation, Deep Learning methods have achieved great success in this field. In this work, we proposed a novel self-supervised method based on Vision Transformer (ViT) with Convolutional Neural Network (CNN) architecture, which is referred to as ViT-Depth . The image reconstruction losses computed by the estimated depth and motion between adjacent frames are treated as supervision signal to establish a self-supervised learning pipeline. This is an effective solution for tasks that need accurate and low-cost 3D perception, such as autonomous driving, robotic navigation, 3D reconstruction, and so on. Our method could leverage both the ability of CNN and Transformer to extract deep features and capture global contextual information. In addition, we propose a cross-frame loss that could constrain photometric error and scale consistency among multi-frames, which lead the training process to be more stable and improve the performance. Extensive experimental results on autonomous driving dataset demonstrate the proposed approach is competitive with the state-of-the-art depth and motion estimation methods.
Shanbao Qiao, Naixue Xiong, Yongbin Gao, Zhijun Fang 0001, Juan Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 A Multiscale Coarse-to-Fine Human Pose Estimation Network With Hard Keypoint Mining
abstract
Current convolution neural network (CNN)-based multiperson pose estimators have achieved great progress, however, they pay no or less attention to “hard” samples, such as occluded keypoints, small and nearly invisible keypoints, and ambiguous keypoints. In this article, we explicitly deal with these “hard” samples by proposing a novel multiscale coarse-to-fine human pose estimation network (HM2PN), which includes two sequential subnetworks: CoarseNet and FineNet. CoarseNet conducts a coarse prediction to locate “simple” keypoints like hands and ankles with a multiscale fusion module, which is integrated with bottleneck, resulting in a novel module called multiscale bottleneck. The new module improves the multiscale representation ability of the network in a fine-grained level, while marginally reducing the computation cost because of group convolution. FineNet further infers “hard” keypoints and refines “simple” keypoints simultaneously with a hard keypoint mining loss. Distinct from the previous works, the proposed loss deals with “hard” keypoints differentially and prevents “simple” keypoints from dominating the computed gradients during training. Experiments on the COCO keypoint benchmark show that our approach achieves superior pose estimation performance compared with other state-of-the-art methods. Source code is available for further research:https://github.com/sues-vision/C2F-HumanPoseEstimation.
Hangyu Tao, Jenq-Neng Hwang, Zhijun Fang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Fast-Fading Channel and Power Optimization of the Magnetic Inductive Cellular Network
abstract
The cellular network of magnetic Induction (MI) communication holds promise in long-distance underground environments. In the traditional MI communication, there is no fast-fading channel since the MI channel is treated as a quasi-static channel. However, for the vehicle (mobile) MI (VMI) communication, the unpredictable antenna vibration brings the remarkable fast-fading. As such fast-fading cannot be modeled by the central limit theorem, it differs radically from other wireless fast-fading channels. Unfortunately, few studies focus on this phenomenon. In this paper, using a novel space modeling based on the electromagnetic field theorem, we propose a 3-dimension model of the VMI antenna vibration. By proposing “conjugate pseudo-piecewise functions” and boundary$p(x)$distribution, we derive the cumulative distribution function (CDF), probability density function (PDF) and the expectation of the VMI fast-fading channel. We also theoretically analyze the effects of the VMI fast-fading on the network throughput, including the VMI outage probability which can be ignored in the traditional MI channel study. We draw several intriguing conclusions different from those in wireless fast-fading studies. For instance, the fast-fading brings more uniformly distributed channel coefficients. Finally, we propose the power control algorithm using the non-cooperative game and multiagent Q-learning methods to optimize the throughput of the cellular VMI network. Simulations validate the derivation and the proposed algorithm.
Honglei Ma, Erwu Liu, Zhijun Fang 0001, Rui Wang 0001, Yongbin Gao, Dongming Zhang 0001
IEEE Trans. Wirel. Commun.3
2023 Gram-based Attentive Neural Ordinary Differential Equations Network for Video Nystagmography Classification
abstract
Video nystagmography (VNG) is the diagnostic gold standard of benign paroxysmal positional vertigo (BPPV), which requires medical professionals to examine the direction, frequency, intensity, duration, and variation in the strength of nystagmus on a VNG video. This is a tedious process heavily influenced by the doctor’s experience, which is error-prone. Recent automatic VNG classification methods approach this problem from the perspective of video analysis without considering medical prior knowledge, resulting in unsatisfactory accuracy and limited diagnostic capability for nystagmographic types, thereby preventing their clinical application. In this paper, we propose an end-to-end data-driven novel BPPV diagnosis framework (TC-BPPV) by considering this problem as an eye trajectory classification problem due to the disease’s symptoms and experts’ prior knowledge. In this framework, we utilize an eye movement tracking system to capture the eye trajectory and propose the Gram-based attentive neural ordinary differential equations network (Gram-AODE) to perform classification. We validate our framework using the VNG dataset provided by the collaborative university hospital and achieve state-of-the-art performance. We also evaluate Gram-AODE on multiple open-source benchmarks to demonstrate its effectiveness in trajectory classification. Code is available at https://github.com/XiheQiu/Gram-AODE.
Xihe Qiu, Shaojie Shi, Xiaoyu Tan, Chao Qu, Zhijun Fang 0001, Yongbin Gao, Peixia Wu
ICCV5
2023 Depth Estimation of Multi-Modal Scene Based on Multi-Scale Modulation
abstract
As multimodal information is complementary, effectively utilizing scene multimodal information has become an increasingly important research topic for many scholars. This paper proposes a novel multi-scale global learning strategy that utilizes both echo and visual modal data as inputs to estimate scene depth. The framework involves constructing a multi-scale feature extraction method using pyramid pooling modules to aggregate contextual information from different regions and improve global information acquisition ability. Furthermore, a recurrent multi-scale feature modulation module is introduced to generate more semantic and accurate spatial representations in each iteration update process. Additionally, a multi-scale fusion method is constructed for the fusion of echo and visual modalities. The proposed method's superior performance is demonstrated through sufficient experiments conducted on the Replica dataset.
Anjie Wang, Zhijun Fang 0001, Yongbin Gao, Gaofeng Cao, Siwei Ma 0001
ICIP2
2023 Online object-level SLAM with dual bundle adjustment
Yongbin Gao, Zhijun Fang 0001
Appl. Intell.4
2023 GsNeRF: Fast novel view synthesis of dynamic radiance fields
Dezhi Liu, Weibing Wan, Zhijun Fang 0001, Xiuyuan Zheng
Comput. Graph.3
2023 Corrigendum to "Single-image deraining via a Recurrent Memory Unit Network" [Knowl.-Based Syst. 218 (2021) 106832]
Yan Zhang 0116, Juan Zhang 0001, Bo Huang 0014, Zhijun Fang 0001
Knowl. Based Syst.4
2023 A lightweight network for smoke semantic segmentation
Feiniu Yuan, Zhijun Fang 0001
Pattern Recognit.4
2023 An effective CNN and Transformer complementary network for medical image segmentation
Feiniu Yuan, Zhengxiao Zhang, Zhijun Fang 0001
Pattern Recognit.3
2023 An angular shrinkage BERT model for few-shot relation extraction with none-of-the-above detection
Junwen Wang, Yongbin Gao, Zhijun Fang 0001
Pattern Recognit. Lett.3
2023 LFT-Net: Local Feature Transformer Network for Point Clouds Analysis
abstract
6G network enables the rapid connection of autonomous vehicles, the generated internet of vehicles establishes a large-scale point cloud, which requires automatic point cloud analysis to build an intelligent transportation system in terms of the 3D object detection and segmentation. Recently, a great variety of deep convolution networks have been proposed for 3D data analysis, making significant progress in the application of deep learning in 3D computer vision. Inspired by the application of transformer network in 2D computer visual tasks, and in order to increase the expression ability of local fine-grained features, we propose an effective local feature transformer network to learn local feature information and correlations between point clouds. Our network is adaptive to the arrangement of set elements through transformer module, so it is suitable for the feature extraction of local point clouds. In addition, experimental results demonstrate that our LFT-network outperforms the state-of-the-art in 3D model classification tasks on ModelNet40 dataset and segmentation tasks on S3DIS dataset.
Yongbin Gao, Xuebing Liu, Jun Li 0036, Zhijun Fang 0001, Kazi Mohammed Saidul Huq
IEEE Trans. Intell. Transp. Syst.4
2023 Joint Optimization of Depth and Ego-Motion for Intelligent Autonomous Vehicles
abstract
The three-dimensional (3D) perception of autonomous vehicles is crucial for localization and analysis of the driving environment, while it involves massive computing resources for deep learning, which can’t be provided by vehicle-mounted devices. This requires the use of seamless, reliable, and efficient massive connections provided by the 6G network for computing in the cloud. In this paper, we propose a novel deep learning framework with 6G enabled transport system for joint optimization of depth and ego-motion estimation, which is an important task in 3D perception for autonomous driving. A novel loss based on feature map and quadtree is proposed, which uses feature value loss with quadtree coding instead of photometric loss to merge the feature information at the texture-less region. Besides, we also propose a novel multi-level V-shaped residual network to estimate the depths of the image, which combines the advantages of V-shaped network and residual network, and solves the problem of poor feature extraction results that may be caused by the simple fusion of low-level and high-level features. Lastly, to alleviate the influence of image noise on pose estimation, we propose a number of parallel sub-networks that use RGB image and its feature map as the input of the network. Experimental results show that our method significantly improves the quality of the depth map and the localization accuracy and achieves the state-of-the-art performance.
Yongbin Gao, Jun Li 0036, Zhijun Fang 0001, Saba Al-Rubaye, Yier Yan
IEEE Trans. Intell. Transp. Syst.4
2022 An end-to-end deep learning model for robust smooth filtering identification
Luo Yu, Zhijun Fang 0001, Naixue Xiong, Haiyue Tian
Future Gener. Comput. Syst.3
2022 A model-based hybrid soft actor-critic deep reinforcement learning algorithm for optimal ventilator settings
Shaotao Chen, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001, Yaochu Jin
Inf. Sci.4
2022 Aspect-level sentiment analysis with aspect-specific context position information
Bo Huang 0014, Ruyan Guo, Zhijun Fang 0001, Guohui Zeng, Jin Liu 0010, Yini Wang, Hamido Fujita, Zhicai Shi
Knowl. Based Syst.4
2022 An anisotropic non-local attention network for image segmentation
Feiniu Yuan, Yaowen Zhu, Zhijun Fang 0001, Jinting Shi
Mach. Vis. Appl.4
2022 A sparse graph wavelet convolution neural network for video-based person re-identification
Yingmao Yao, Hamido Fujita, Zhijun Fang 0001
Pattern Recognit.4
2022 Depth Estimation Using a Self-Supervised Network Based on Cross-Layer Feature Fusion and the Quadtree Constraint
abstract
Depth estimation from a camera is an important task for 3D perception. Recently, without using the labeled ground truth of depth map, a self-supervised deep learning network can use relative pose to synthesize the target image from the reference image, and the photometric error between synthesized reference image and real one is used as self-supervisory signal. In this paper, we propose a novel self-supervised depth estimation network, which takes advantage of the quadtree constraint to optimize the depth estimation network. Based on the quadtree constraint, the photometric loss and depth loss of quadtree are proposed. In order to solve the problem that multiple depth values in repeated structures and uniform texture regions can cause relatively low photometric loss, we use quadtree-based photometric loss, which calculates the averaged photometric loss in quadtree blocks instead of the pixel-wise loss. For the problem of imbalanced depth distribution, we use quadtree depth loss, which constrains the depth inconsistency within quadtree blocks. The depth estimation network is composed of deep fusion module and cross-layer feature fusion module, which can better extract the feature information of RGB image and sparse keypoints depths, and makes full use of the detail information of the shallow feature map and the semantic information of the deep feature map to enrich the feature information extraction. Experimental results demonstrate that our method outperforms the state-of-the-art approaches of depth estimation.
Yongbin Gao, Zhijun Fang 0001, Yuming Fang 0001, Hamido Fujita, Jenq-Neng Hwang
IEEE Trans. Circuits Syst. Video Technol.3
2021 Point AE-DCGAN: A deep learning model for 3D point cloud lossy geometry compression
abstract
3D point cloud has been widely applied in virtual reality and augmented reality. A complex 3D scene always needs a large number of the point cloud to represent and demands a lot of space to store. Thus, point cloud compression becomes a crucial issue to research. In this paper, we propose a novel lossy geometric compression method of autoencoder based on DCGAN optimization. This method can reconstruct a high-quality point cloud and solves a large area of missing points in the process of compression and decompression. To improve the point cloud codec performance, we propose a multi-scale 3D deconvolution hopping connection structure to obtain a better-quality reconstructed point cloud under low bit rates. Our approach is the first GAN-based point cloud compression algorithm to our knowledge. Compared with state-of-the-art methods on the MVUB dataset, our approach achieves a better rate-distortion performance and visual quality.
Zhijun Fang 0001, Yongbin Gao, Siwei Ma 0001, Yaochu Jin, Anjie Wang
DCC2
2021 An efficient attention module for 3d convolutional neural networks in action recognition
Guanghao Jiang, Zhijun Fang 0001
Appl. Intell.3
2021 PointFusionNet: Point feature fusion network for 3D point clouds analysis
Pan Liang, Zhijun Fang 0001, Bo Huang 0014, Xianhua Tang, Cengsi Zhong
Appl. Intell.2
2021 Automatic coronary artery segmentation algorithm based on deep learning and digital image processing
Yongbin Gao, Zhijun Fang 0001
Appl. Intell.3
2021 SAT-Net: a side attention network for retinal image segmentation
Huilin Tong, Zhijun Fang 0001, Ziran Wei, Qingping Cai, Yongbin Gao
Appl. Intell.2
2021 A novel IoT network intrusion detection approach based on Adaptive Particle Swarm Optimization Convolutional Neural Network
Xiu Kan, Yixuan Fan, Zhijun Fang 0001, Le Cao, Naixue Xiong
Inf. Sci.3
2021 Line-based visual odometry using local gradient fitting
Junxin Lu, Zhijun Fang 0001, Yongbin Gao, Jieyu Chen
J. Vis. Commun. Image Represent.2
2021 3D reconstruction with auto-selected keyframes based on depth completion correction and pose fusion
Yongbin Gao, Zhijun Fang 0001, Shuqun Yang
J. Vis. Commun. Image Represent.3
2021 Single-image deraining via a Recurrent Memory Unit Network
Yan Zhang 0116, Juan Zhang 0001, Bo Huang 0014, Zhijun Fang 0001
Knowl. Based Syst.4
2021 Photometric transfer for direct visual odometry
Kaiying Zhu, Zhijun Fang 0001, Yongbin Gao, Hamido Fujita, Jenq-Neng Hwang
Knowl. Based Syst.3
2021 Joint Design of Beamforming and Edge Caching in Fog Radio Access Networks
abstract
In this paper, we study a novel transmission framework based on statistical channel state information (SCSI) by incorporating edge caching and beamforming in a fog radio access network (F-RAN) architecture. By optimizing the statistical beamforming and edge caching, we formulate a comprehensive nonconvex optimization problem to minimize the backhaul cost subject to the BS transmission power, limited caching capacity, and quality-of-service (QoS) constraints. By approximating the problem using the l 0 -norm, Taylor series expansion, and other processing techniques, we provide a tailored second-order cone programming (SOCP) algorithm for the unicast transmission scenario and a successive linear approximation (SLA) algorithm for the joint unicast and multicast transmission scenario. This is the first attempt at the joint design of statistical beamforming and edge caching based on SCSI under the F-RAN architecture.
Wenjing Lv, Rui Wang 0001, Jun Wu 0006, Zhijun Fang 0001, Songlin Cheng
Secur. Commun. Networks4
2020 Nested rings: a simple scalable ring-based ROADM structure for neural application computing in mega datacenters
Zhijun Fang 0001
Neural Comput. Appl.2
2020 Feature fusion network based on attention mechanism for 3D semantic segmentation of point clouds
Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014, Cengsi Zhong, Ruoxi Shang
Pattern Recognit. Lett.2
2020 Adversarial Learning for Joint Optimization of Depth and Ego-Motion
abstract
In recent years, supervised deep learning methods have shown a great promise in dense depth estimation. However, massive high-quality training data are expensive and impractical to acquire. Alternatively, self-supervised learning-based depth estimators can learn the latent transformation from monocular or binocular video sequences by minimizing the photometric warp error between consecutive frames, but they suffer from the scale ambiguity problem or have difficulty in estimating precise pose changes between frames. In this paper, we propose a joint self-supervised deep learning pipeline for depth and ego-motion estimation by employing the advantages of adversarial learning and joint optimization with spatial-temporal geometrical constraints. The stereo reconstruction error provides the spatial geometric constraint to estimate the absolute scale depth. Meanwhile, the depth map with an absolute scale and a pre-trained pose network serves as a good starting point for direct visual odometry (DVO). DVO optimization based on spatial geometric constraints can result in a fine-grained ego-motion estimation with the additional backpropagation signals provided to the depth estimation network. Finally, the spatial and temporal domain-based reconstructed views are concatenated, and the iterative coupling optimization process is implemented in combination with the adversarial learning for accurate depth and precise ego-motion estimation. The experimental results show superior performance compared with state-of-the-art methods for monocular depth and ego-motion estimation on the KITTI dataset and a great generalization ability of the proposed approach.
Anjie Wang, Zhijun Fang 0001, Yongbin Gao, Songchao Tan, Shanshe Wang, Siwei Ma 0001, Jenq-Neng Hwang
IEEE Trans. Image Process.2
2019 Unsupervised Learning of Depth and Ego-Motion with Spatial-Temporal Geometric Constraints
abstract
In this paper, we propose an unsupervised joint deep learning pipeline for depth and ego-motion estimation that explicitly incorporated with traditional spatial-temporal geometric constraints. The stereo reconstruction error provides the spatial geometric constraint to estimate the absolute scale depth. Meanwhile, the depth map with absolute scale and a pre-trained pose network serve as a good starting point for direct visual odometry (DVO), resulting in a fine-grained ego-motion estimation with the additional back-propagation signals provided to the depth estimation network. The proposed joint training pipeline enables an iterative coupling optimization process for accurate depth and precise ego-motion estimation. The experimental results show the state-of-the-art performance for monocular depth and ego-motion estimation on the KITTI dataset and a great generalization ability of the proposed approach.
Anjie Wang, Yongbin Gao, Zhijun Fang 0001, Shanshe Wang, Siwei Ma 0001, Jenq-Neng Hwang
ICME3
2019 DD-CycleGAN: Unpaired image dehazing via Double-Discriminator Cycle-Consistent Generative Adversarial Network
Jingming Zhao, Juan Zhang 0001, Zhi Li 0049, Jenq-Neng Hwang, Yongbin Gao, Zhijun Fang 0001, Bo Huang 0014
Eng. Appl. Artif. Intell.6
2019 Unsupervised learning of depth estimation based on attention model and global pose optimization
Renyue Dai, Yongbin Gao, Zhijun Fang 0001, Anjie Wang, Juan Zhang 0001, Cengsi Zhong
Signal Process. Image Commun.3
2019 Deep3DSaliency: Deep Stereoscopic Video Saliency Detection Model by 3D Convolutional Networks
abstract
Stereoscopic saliency detection plays an important role in various stereoscopic video processing applications. However, conventional stereoscopic video saliency detection methods mainly use independent low-level features instead of extracting them automatically, and thus, they ignore the intrinsic relationship between the spatial and temporal information. In this paper, we propose a novel stereoscopic video saliency detection method based on 3D convolutional neural networks, namely Deep 3D Video Saliency (Deep3DSaliency). The proposed network consists of two sub-models: Spatiotemporal Saliency Model (STSM), and Stereoscopic Saliency Aware Model (SSAM). STSM directly takes three consecutive video frames as the input to extract visual spatiotemporal features, while SSAM attempts to further infer the depth and semantic features from the left and right video frames by shared parameters from STSM. The visual spatiotemporal features from STSM, and the depth and semantic features from SSAM are learned by an alternating optimization scheme. Finally, all these saliency-related features are combined together for the final stereoscopic saliency detection via 3D deconvolution. Experimental results show the superior performance of the proposed model over other existing ones in saliency estimation for 3D video sequences.
Yuming Fang 0001, Guanqun Ding, Jia Li 0003, Zhijun Fang 0001
IEEE Trans. Image Process.4
2018 Gradient-based adaptive particle swarm optimizer with improved extremal optimization
Xiaoli Zhao 0003, Jenq-Neng Hwang, Zhijun Fang 0001
Appl. Intell.3
2018 Image salient regions encryption for generating visually meaningful ciphertext image
Wenying Wen, Yushu Zhang 0001, Yuming Fang 0001, Zhijun Fang 0001
Neural Comput. Appl.4
2017 Inter-camera tracking based on fully unsupervised online learning
abstract
In this paper, we present a novel fully automatic approach to track the same human across multiple disjoint cameras. Our framework includes a two-phase feature extractor and an online-learning-based camera link model estimation. We introduce an effective and robust integration of appearance and context features. Couples are detected automatically, and the couple feature is also integrated with appearance features effectively. The proposed algorithm is scalable with the use of a fully unsupervised online learning framework. In the experiments, it outperforms all the state-of-the-art methods on the benchmark NLPR_MCT dataset.
Young-Gun Lee, Jenq-Neng Hwang, Zhijun Fang 0001
ICIP4
2017 A New Code Generation Method for Software Engineering: From Requirements Model to Source Code
abstract
the existing software engineering techniques for software synthesis from requirements model to source code have many limitations. The synthesis approach shows that these limitations focused on the refinement relationship between the requirements specification and the desired system. We have to propose a new approach for code generation to overcome such limitations, i.e. refining the software behaviors in requirements model to code, distinguishing function information and architecture information from requirements model, among others. Hence, in this thesis we aim at the problems that how to modeling based on software behaviors, how to delimitate the system architecture and so on. And we also will show a sample, ultimately, to demonstrate our approach. Meanwhile, some additional techniques for synthesis to ensure the correctness of the source code will be recommended.
Bo Huang 0014, Zhijun Fang 0001, Yongbin Gao
SoMeT2
2017 Exploring finger vein based personal authentication for secure IoT
Yu Lu 0006, Shiqian Wu, Zhijun Fang 0001, Naixue Xiong, Sook Yoon, Dong Sun Park
Future Gener. Comput. Syst.3
2017 Visual topic discovering, tracking and summarization from social media streams
Yu-Ru Lin, Naixue Xiong, Zhijun Fang 0001
Multim. Tools Appl.5
2017 Optimized Multioperator Image Retargeting Based on Perceptual Similarity Measure
abstract
With various emerging mobile devices, the visual content have be to resized into different sizes or aspect ratios for good viewing experiences. In this paper, we propose a new multioperator retargeting algorithm by using four retargeting operators of seam carving, cropping, warping, and scaling iteratively. To determine which retargeting operator should be used at each iteration, we adopt structural similarity (SSIM) to evaluate the similarity between the original and retargeted images. The retargeting operator sequence is constructed based on the four types of retargeting operators by an optimization process. Since the sizes of original and retargeted images are different, scale-invariant feature transform flow is used for dense correspondence between the original and retargeted images for similarity evaluation. Additionally, visual saliency is used to weight SSIM results based on the characteristics of the human visual system. Experimental results on a public image retargeting database have shown the promising performance of the proposed multioperator retargeting algorithm.
Yuming Fang 0001, Zhijun Fang 0001, Feiniu Yuan, Yong Yang 0001, Shouyuan Yang, Naixue Xiong
IEEE Trans. Syst. Man Cybern. Syst.2
2016 Camera self-calibration from tracking of moving persons
abstract
In a video surveillance system with a single static camera, tracking results of moving persons can be effectively used for camera self-calibration. However, the current methods need to depend on robustness of both tracking and segmentation procedures. RANSAC has been widely used to remove outliers in finding the vertical vanishing point and the horizon line, but the performance is degraded when the proportion of outliers is high. Last but not least, all of them require excessive simplifications in the algorithmic procedures resulting in increasing reprojection error. In this paper, a robust segmentation and tracking system is applied to provide accurate estimation of head and foot locations of moving persons. The noise in the computation of vanishing points is handled by mean shift clustering and Laplace linear regression through convex optimization. We also propose to use the estimation of distribution algorithm (EDA) to search for the local optimal solution for camera calibration that minimizes average reprojection error on the ground plane, while relaxing the assumptions on camera parameters. Promising evaluations of the performance of our proposed method on real scenes are presented.
Yen-Shuo Lin, Kuan-Hui Lee, Jenq-Neng Hwang, Jen-Hui Chuang, Zhijun Fang 0001
ICPR6
2016 A novel selective image encryption method based on saliency detection
abstract
Salient regions usually carry important information in images. Existing feature encryption algorithms aim at extracting edge features as significant information rather than salient regions for encryption purpose. Moreover, most of them protect significant information by transforming the input image into texture-like or noise-like encrypted image which is obviously a visual sign of encrypted image, and thus can be easily attacked. In this paper, we propose a salient regions encryption scheme to generate visually meaningful ciphertext. First, salient regions are efficiently extracted by a saliency detection model in the compressed domain. Then we pre-encrypt these salient regions by a chaos-based encryption algorithm. With optical encryption theory, the pre-encrypted salient regions are finally transformed into a visually meaningful ciphertext. To the best of our knowledge, it is the first time to use salient regions as important visual information for encryption to obtain cipertext in images. The experimental results demonstrate that the salient regions can be largely hidden with the proposed method.
Wenying Wen, Yushu Zhang 0001, Yuming Fang 0001, Zhijun Fang 0001
VCIP4
2016 A mutual local-ternary-pattern based method for aligning differently exposed images
Shiqian Wu, Lingxian Yang, Wangming Xu, Jinghong Zheng 0001, Zhengguo Li, Zhijun Fang 0001
Comput. Vis. Image Underst.6
2016 A general effective rate control system based on matching measurement and inter-quantizer
Zhijun Fang 0001, Yongbin Gao, Naixue Xiong, Athanasios V. Vasilakos, Yuming Fang 0001
Inf. Sci.1
2016 High-order local ternary patterns with locality preserving projection for smoke detection and image classification
Feiniu Yuan, Jinting Shi, Xue Xia 0005, Yuming Fang 0001, Zhijun Fang 0001, Tao Mei 0001
Inf. Sci.5
2016 Abnormal event detection in crowded scenes based on deep learning
Zhijun Fang 0001, Fengchang Fei, Yuming Fang 0001, Changhoon Lee, Naixue Xiong, Lei Shu 0001
Multim. Tools Appl.1
2016 Feedback Control Scheduling in Energy-Efficient and Thermal-Aware Data Centers
abstract
This paper presents a model-predictive control-based scheduling strategy called ThermoRing to reduce cooling costs in data centers. ThermoRing makes use of an online feedback control mechanism to improve thermal management of energy-efficient clusters in a data center. ThermoRing aims at keeping the maximum inlet temperatures of the nodes under a redline temperature limit with little stability errors. Importantly, the ThermoRing approach is capable of dealing with emergency conditions (e.g., node fan shutdown and unexpected rising task arrival rates) by dynamically balancing load among the nodes. ThermoRing incorporates a heat distribution matrix to model the thermal characteristics of a data center housing cluster. ThermoRing is conducive to thermal management in data centers with high-scheduling performance and stability. Using a real-world online bookstore trace, we conduct extensive experiments to compare the performance of ThermoRing with three existing solutions (i.e., C-Oracle, Ad-hoc, and MinHR). The experimental results show that ThermoRing improved the system throughput by more than 10% under regular load conditions and by 40% in emergency cases. ThermoRing also significantly improves the energy efficiency of MinHR, which is a thermal-aware scheduler.
Tao Peng 0006, Xiao Qin 0001, Qiping Hu, Zhijun Fang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2015 Combined estimation of camera link models for human tracking across nonoverlapping cameras
abstract
Human tracking across multiple cameras is highly demanded for large scale video surveillance. To successfully track human across multiple uncalibrated cameras that have no overlapping field of views, a system to train more reliable camera link models is proposed in this paper. We employ a novel approach of combining multiple camera links and building bidirectional transition time distribution in the process of estimation. Through the unsupervised scheme, the system builds several camera link models simultaneously for the camera network that has multi-path in presence of the outliers. Our proposed method decreases incorrect correspondences and results in more accurate camera link model for higher tracking accuracy. The proposed algorithm shows the effectiveness by evaluating in the real-world camera network scenarios.
Young-Gun Lee, Jenq-Neng Hwang, Zhijun Fang 0001
ICASSP3
2015 Real-time image smoke detection using staircase searching-based dual threshold AdaBoost and dynamic analysis
abstract
It is very challenging to accurately detect smoke from images because of large variances of smoke colour, textures, shapes and occlusions. To improve performance, the authors combine dual threshold AdaBoost with staircase searching technique to propose and implement an image smoke detection method. First, extended Haar‐like features and statistical features are efficiently extracted from integral images from both intensity and saturation components of RGB images. Then, a dual threshold AdaBoost algorithm with a staircase searching technique is proposed to classify the features of smoke for smoke detection. The staircase searching technique aims at keeping consistency of training and classifying as far as possible. Finally, dynamic analysis is proposed to further validate the existence of smoke. Experimental results demonstrate that the proposed system has a good robustness in terms of early smoke detection and low false alarm rate, and it can detect smoke from videos with size of 320 × 240 in real time.
Feiniu Yuan, Zhijun Fang 0001, Shiqian Wu, Yong Yang 0001, Yuming Fang 0001
IET Image Process.2
2015 Visual acuity inspired saliency detection by using sparse features
Yuming Fang 0001, Weisi Lin, Zhijun Fang 0001, Zhenzhong Chen 0001, Chia-Wen Lin, Chenwei Deng
Inf. Sci.3
2015 No-Reference Quality Assessment of Contrast-Distorted Images Based on Natural Scene Statistics
abstract
Contrast distortion is often a determining factor in human perception of image quality, but little investigation has been dedicated to quality assessment of contrast-distorted images without assuming the availability of a perfect-quality reference image. In this letter, we propose a simple but effective method for no-reference quality assessment of contrast distorted images based on the principle of natural scene statistics (NSS). A large scale image database is employed to build NSS models based on moment and entropy features. The quality of a contrast-distorted image is then evaluated based on its unnaturalness characterized by the degree of deviation from the NSS models. Support vector regression (SVR) is employed to predict human mean opinion score (MOS) from multiple NSS features as the input. Experiments based on three publicly available databases demonstrate the promising performance of the proposed method.
Yuming Fang 0001, Kede Ma, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001, Guangtao Zhai
IEEE Signal Process. Lett.5
2014 Video Saliency Incorporating Spatiotemporal Cues and Uncertainty Weighting
abstract
We propose a novel algorithm to detect visual saliency from video signals by combining both spatial and temporal information and statistical uncertainty measures. The main novelty of the proposed method is twofold. First, separate spatial and temporal saliency maps are generated, where the computation of temporal saliency incorporates a recent psychological study of human visual speed perception. Second, the spatial and temporal saliency maps are merged into one using a spatiotemporally adaptive entropy-based uncertainty weighting approach. The spatial uncertainty weighing incorporates the characteristics of proximity and continuity of spatial saliency, while the temporal uncertainty weighting takes into account the variations of background motion and local contrast. Experimental results show that the proposed spatiotemporal uncertainty weighting algorithm significantly outperforms state-of-the-art video saliency detection models.
Yuming Fang 0001, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001
IEEE Trans. Image Process.4
2013 A saliency detection model based on sparse features and visual acuity
abstract
In this paper, we propose a novel computational model of visual attention based on the relevant characteristics of the Human Visual System (HVS). The input image is firstly divided into small image patches. Then the sparse features for each image patch are extracted based on the learned sparse coding basis. The human visual acuity is adopted in the calculation of the center-surround feature differences for saliency detection. In addition, the neighboring image patches for computing the saliency value of each center image patch are selected based on the characteristics of HVS. Experimental results show that the proposed saliency detection algorithm outperforms other existing schemes tested with a large public image database.
Yuming Fang 0001, Weisi Lin, Zhenzhong Chen 0001, Chia-Wen Lin, Zhijun Fang 0001, Chenwei Deng
ISCAS5
2013 Fingerprint matching based on extreme learning machine
Shan Juan Xie, Sook Yoon, Dong Sun Park, Zhijun Fang 0001, Shouyuan Yang
Neural Comput. Appl.5
2009 Infrared Face Recognition Based on Radiant Energy and Curvelet Transformation
abstract
In this paper, a infrared face recognition method using radiant energy conversion and curvelet transformation is proposed. Firstly, to get the stable feature of thermal face, thermal images are converted into radiant energy images according to Stefan-Boltzmann's law. Secondly, curvelet transform has better directional and edge representation abilities than widely used wavelet transformation and other classic transformations. Inspired by these attractive attributes of curvelets in sparse representation of the images, we introduce the idea of decomposing images into their curvelet subbands to extract the principal representative feature, which saves the computational complexity and storage units. Finally, the nearest neighbor classifier is chosen to get the system recognition result. The experiments illustrate that compared with traditional PCA based systems, the proposed system has better performance and requires fewer computations and memory units.
Zhihua Xie 0002, Shiqian Wu, Zhijun Fang 0001
IAS4
2006 An Efficient Mobility Management Scheme for Hierarchical Mobile IPv6 Networks
Zhengyou Wang, Zhijun Fang 0001, Weiming Zeng, Shiqian Wu
ICCSA (2)3
2006 A New Color Image Enhancement Algorithm for Camera-Equipped Mobile Telephone
Zhengyou Wang, Quan Xue, Guobin Chen, Weiming Zeng, Zhijun Fang 0001, Shiqian Wu
KES (1)5