Xianping Fu

dblp:11/7460 · DBLP profile ↗
← Back
118ranked-venue papers
5as first author
100since 2021 · last 2026
0000-0001-9888-9327ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 59 · 2 first-author · 53 since 2021Artificial intelligence and machine learning · 49 · 2 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Instance-Guided Scene Adaptation for Unsupervised Person Search
abstract
Unsupervised Domain Adaptation (UDA) is a challenging task in person search. It adapts a well-trained model from a labeled source domain to an unlabeled target domain for privacy and efficiency. Currently, most of the state-of-the-art UDA person search methods adopt multi-scale feature alignment techniques to learn domain-invariant representations. However, person search is a multi-granularity task, and such an indiscriminate method of bridging the differences between domains misleads the identity learning process, which significantly limits the model's performance. In this paper, we propose an Instance-Guided Scene Adaptation (IGSA) framework by eradicating scene disparities and focusing the tasks on instances, effectively eliminating the contradiction between person search and domain adaptation. In IGSA, a Scene-Aware Bidirectional Filter (SABF) is designed to divide the image features into background and foreground to perform bidirectional modulations, thereby achieving simultaneous scene elimination and instance enhancement. To further improve the reliability of identity learning, we also propose an Instance Consistency Contrastive Learning (ICCL) method. By performing cross-epoch updates on the instance-level memory bank and re-initializing the cluster-level memory bank, the problem of inconsistent training across epochs caused by instance identity drift can be alleviated. Through the above designs, our method can achieve state-of-the-art performance on two benchmark datasets, with 82.1% mAP and 83.8% top-1 on the CUHK-SYSU dataset and 41.1% mAP and 82.3% top-1 on the PRW dataset, which is even better than some supervised methods.
Huibing Wang, Jinjia Peng, Xianping Fu, Jiqing Zhang
AAAI4
2026 Conditional Prompt Learning via Degradation Perception for Underwater Image Enhancement
abstract
Underwater Image Enhancement (UIE) focuses on improving visual quality from various underwater scenes. Existing methods simplistically treat various degradations as homogeneous, disregarding their intrinsic connections and causing models to blindly learn, resulting in conflicting optimization goals and visual distortions. To address above limitations, we propose a Conditional Prompt Learning via Degradation Perception (CPLDP) model, which employs conditional prompt as degradation perception priors and guides underwater image enhancement. Specifically, we show that the natural language prompts not only promote distinguishing different degraded images, but also aid in exploring more details with semantic information. Therefore, our method generates five key degradation prompts (green/blue/green-blue color casts, uneven illumination and haze) with conditional prompt learning. Subsequently, considering the intrinsic relationships among different degradations, we employ degradation perceptions as priors and fine-tune the learning strategy to enhance underwater images. During training, an adaptive loss function with multi-degradations is designed, allowing it to effectively handle the task conflicts among multiple underwater degradations. Additionally, we conduct a human visual-based underwater dataset with various degradation types by subjective statistics. Extensive experiments on both full-reference and non-reference datasets demonstrate that our CPLDP can achieve better visual results and outperforms state-of-the-art UIE methods across various degradation scenarios.
Mingze Yao, Zhiying Jiang, Xianping Fu, Huibing Wang
AAAI3
2026 Event-based Gaze Estimation via Spatial-Temporal Interaction
abstract
Event cameras respond in microseconds, enabling them to capture rapid changes in eye movements. However, the inherent variability in eye movements usually introduces uncertainty, complicating accurate predictions based solely on temporal trends. To address this, we propose an event-based gaze estimation method that leverages spatial-temporal interactions to integrate both spatial and temporal distributions for enhanced eye movement prediction. The temporal features are generated by the gated recurrent unit (GRU), while the multi-level spatial features are extracted by convolutions. To preserve fine-grained spatial details, cross-channel information fusion is employed to merge spatiotemporal features, with skip connections ensuring retention of the original input’s spatial information. Experimental results on the 3ET+ benchmark dataset demonstrate that our method achieves a competitive accuracy of 1.52 pixels.
Yafei Wang 0004, Runze Yan, Xianping Fu
ETRA4
2026 GLGaze: In-vehicle gaze estimation via joint global-local enhancement in the spatial and scale domains
Yafei Wang 0004, Fushuo Huo, Runze Yan, Niuniu Zhang, Xianping Fu
Expert Syst. Appl.6
2026 Efficient Vision Transformer with Token Sparsification for Event-Based Object Tracking
Jiqing Zhang, Xin Yang 0011, Haoming Tang, Yuanchen Wang, Huibing Wang, Xianping Fu
Int. J. Comput. Vis.7
2026 DBWaterNet: Dual-branch joint refinement for underwater image enhancement
Muazzamu Ibrahim, Zayyanu Shuaibu, Zhexiang Zhang, Guipeng Zhu, Jianfeng Zhong, Xinbo Zhang, Yafei Wang 0004, Xianping Fu
J. Vis. Commun. Image Represent.9
2026 PIGaze: Personalized in-vehicle gaze estimation with plug-and-play adaptation
Yafei Wang 0004, Haiheng Nan, Runze Yan, Xueyan Ding, Xianping Fu
Knowl. Based Syst.6
2026 Semantic Contrast for Domain-Robust Underwater Image Quality Assessment
abstract
Underwater image quality assessment (UIQA) is hindered by complex degradation and domain shifts across aquatic environments. Existing no-reference IQA methods rely on costly and subjective mean opinion scores (MOS), which limit their generalization to unseen domains. To overcome these challenges, we propose SCUIA, an unsupervised UIQA framework leveraging semantic contrastive learning for quality prediction without human annotations. Specifically, we introduce a vision-language contrastive learning strategy that aligns image features with textual embeddings in a unified semantic space, capturing implicit degradation-quality correlations. We further enhance quality discrimination with a hierarchical contrastive learning mechanism that combines image-specific statistical priors and semantic prompts. A triplet-based inter-group contrastive loss explicitly models relative quality relationships. To tackle cross-domain variations, we develop an unsupervised domain adaptation module that uses local statistical features to guide CLIP fine-tuning to disentangle domain-invariant quality representations from domain-specific noise. This enables zero-shot cross-domain quality prediction without labeled data. Extensive experiments on public UIQA benchmarks demonstrate significant improvements over existing methods, highlighting superior generalization and domain adaptability.
Jingchun Zhou, Chunjiang Liu, Qiuping Jiang, Xianping Fu, Junhui Hou, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 BCDnet: Balanced coupling and decoupling network for person search
Zhengjie Lu, Jinjia Peng, Huibing Wang, Xianping Fu
Pattern Recognit.4
2026 Bridging the gap : Learning adaptive knowledge transition for lifelong person re-identification
Jinjia Peng, Jican Tan, Huibing Wang, Xianping Fu
Pattern Recognit.5
2026 MSLiR-Net: Multi-scale lightweight real-time underwater image enhancement with Spatial-Frequency Features Interaction
Muazzamu Ibrahim, Zayyanu Shuaibu, Zhexiang Zhang, Guipeng Zhu, Jianfeng Zhong, Xinbo Zhang, Yafei Wang 0004, Xianping Fu
Signal Process. Image Commun.9
2026 Boosting Underwater Object Detection via Differential Attention
GuoLiang Yuan 0001, Junchi Li, Hongming Chen 0004, Xianping Fu, Yafei Wang 0004
IEEE Signal Process. Lett.5
2026 A Lightweight Polarization-Guided Plug-In for Underwater Image Enhancement
abstract
Underwater images play a vital role in marine exploration, but are often severely degraded due to complex imaging conditions, including color distortion, haze effects, and non-uniform illumination. Existing deep learning-based enhancement methods predominantly rely on conventional RGB sensors, which struggle to distinguish between scattered and reflected light, thereby limiting enhancement performance. Polarization imaging, with its capability to capture directional light information, offers promising potential for underwater image enhancement. In this paper, we propose a lightweight yet effective polarization feature extractor that captures global spatial cues from polarization images. Additionally, we design a polarization-guided feature integration module that adaptively enhances the representational capacity of RGB features. Notably, the proposed module is plug-in and can be seamlessly integrated into existing RGB-based enhancement networks. Extensive experiments across multiple datasets demonstrate that incorporating polarization information significantly improves enhancement performance, highlighting its effectiveness as a valuable cue for underwater image enhancement. The code and pretrained models are at https://github.com/jgy0/UPGD.
Guangyao Ju, Jiqing Zhang, Jingqi Zang, Zetian Mi, Xin Yang 0011, Huibing Wang, Jiarui Fan, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.10
2026 UVE-LUT: Learnable Attenuation-Aware Lookup Table for Underwater Video Enhancement
abstract
Deep learning techniques are increasingly being employed for underwater video enhancement (UVE), with the key objectives not only encompassing color cast correction and image quality improvement but also maintaining temporal consistency. Most existing underwater image enhancement (UIE) methods, when directly applied to video sequences, often introduce inter-frame flickering artifacts, thereby impairing the comprehension of video content. Although 3D convolutional neural network-based approaches can preserve temporal coherence, they are characterized by substantial parameter counts and computational complexity, leading to inefficient model performance. To address these limitations, we propose a learnable Look-Up-Table (LUT) based on attenuation-aware and wavelet prior for UVE, abbreviated UVE-LUT, which effectively maintains inter-frame consistency in color and brightness distribution by leveraging look-up table (LUT) technology. The proposed method incorporates a learnable Attenuation-Aware LUT (AA-LUT) module that performs low-latency, low-complexity adaptive enhancement on the low-frequency components of frame sequences in the wavelet domain. Additionally, we design a high-frequency detail enhancement module to conduct three-dimensional detail enhancement on the high-frequency components. Extensive subjective and objective evaluations on underwater video and image datasets demonstrate that the proposed UVE-LUT achieves a favorable balance between image quality and temporal consistency. Our source code will be available at Github.
Zetian Mi, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.4
2026 Interaction-Driven Edge Crisping for Underwater Salient Object Detection
abstract
Underwater salient object detection (USOD) faces greater challenges than general scenes due to the edge blurring which is caused by light absorption and scattering in water. Existing methods employ unrefined edge feature to perform unidirectional guidance on saliency feature, resulting in the coarse edge of saliency map. To address this issue, we propose a novel interaction-driven edge crisping network (IDENet) for underwater salient object detection. IDENet facilitates the bidi-rectional modulation of inter-features and the self-refinement of intra-feature, generates crisp saliency map and edge map. In IDENet, the interaction-driven edge guidance module (IDEGM) is designed to utilize cross-feature interaction by leveraging their correlations, facilitating saliency feature’s awareness of edge information, mitigating the interference of non-salient objects in edge feature. To learn more accurate edge region of the salient object, the edge intersection-and-union loss function (EIUL) is introduced to restrict the intersection and union of predicted saliency maps and edge maps to prevent over-expansion or under-contraction. Experimental results on two latest underwater datasets demonstrate the superiority of the proposed method over the state-of-the-art models. The source code of our method will be made available at https://github.com/ UnderwaterVisionMZTdlmu/IDENet.
Zetian Mi, Shuaiyong Jiang, Guanxi Li, Jiqing Zhang, Huibing Wang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.7
2026 Towards Ultra-High-Definition Image Deraining: A Benchmark and an Efficient Method
abstract
Despite significant advancements in image deraining, most existing methods are carried out on low-resolution images, leaving their effectiveness on high-resolution images uncertain. This limitation becomes even more pronounced with the rise of ultra-high-definition (UHD) imaging. In this paper, we tackle the challenge of UHD image deraining and introduce 4K-Rain13 k, the first large-scale UHD image deraining dataset, featuring 13,000 paired images at 4 K resolution. Leveraging this dataset, we conduct a benchmark study on existing methods for processing UHD images. To better address this task, we propose UDR-Mixer, an efficient and effective architecture tailored for UHD image deraining. Our model comprises two key components: a spatial feature rearrangement layer, which captures long-range dependencies in UHD images, and a frequency feature modulation layer, which enhances high-fidelity image reconstruction. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while maintaining lower model complexity. The source code and proposed dataset are available athttps://github.com/cschenxiang/UDR-Mixer.
Hongming Chen 0004, Xiang Chen 0015, Chen Wu 0006, Zhuoran Zheng, Jinshan Pan, Xianping Fu
IEEE Trans. Multim.6
2026 Tensorized Fine-Grained Incomplete Multiview Clustering via Intrinsic Structure Recovery
abstract
Incomplete multiview clustering (IMVC) focuses on exploiting the complementary and consistent information from multiple incomplete views for dividing unlabeled multiview data into corresponding clusters. Most existing methods seek to recover the missing samples of the views while inevitably having an influence on the intrinsic structure of the original space. Moreover, previous IMVC algorithms treat the samples of each view equally, which learns the consensus representation in a view-level manner and thus neglects that different views contribute to each individual sample diversely. These situations are a limitation to effective recovery of absent samples and fuse the heterogeneous information, which leads to suboptimal clustering performance. To tackle the above issues, this article proposes a novel approach called tensorized fine-grained IMVC via intrinsic structure recovery (TFIR), which flexibly recovers the intrinsic structure of the original data and attains the unified representation based on the fine-grained fusion strategy. Specifically, TFIR infers the incomplete data via available instances’ relations to improve the accuracy of learned intersample-specific representations. Afterward, TFIR stacks the specific representations of multiple views into a tensor to preserve the intrasample consistency. Finally, TFIR explores the complementarity among various views from the fine-grained sample perspective to obtain a rich underlying structure for clustering. Experimental results on the eight benchmark datasets clearly verify the remarkable superiority of TFIR compared with the most state-of-the-art methods.
Huibing Wang, Luyan Cui, Mingze Yao, Xianping Fu
IEEE Trans. Syst. Man Cybern. Syst.4
2025 Towards Implicit Personal Eye Gaze Calibration in Real-world Driving Scenarios
Yafei Wang 0004, GuoLiang Yuan 0001, Xianping Fu
ETRA3
2025 Few-shot Personalized Gaze Estimation in Natural Driving Environment
Yafei Wang 0004, Haiheng Nan, Xianping Fu
ETRA3
2025 PTGaze: Cross-Domain Gaze Estimation via Proxy Tuning
Yafei Wang 0004, Runze Yan, Yaxiong Lei, Xianping Fu
ETRA4
2025 Efficient Cross-Boundary Grasping in Stacked Clutter with Single-Visual Mapping Multi-Step
abstract
In logistics applications, the vision-based technology for grasping target objects in the air is relatively mature. However, when operating across the air and water such as grasping marine products from the water, the visual information collected by the camera will be disturbed by ripples and bubbles on the water surface, resulting in low grasping efficiency. Therefore, we introduce a grasping strategy based on single-visual mapping for multi-step (SVMMS) strategy to achieve cross-medium operations involving stacked objects. Specifically, we design a multifunctional integrated Deep Q-learning-based network model to extract visual features from the scene to effectively detect stacked objects and outputs their hierarchical relationships. Moreover, we quantify the underlying relationship between motion logic during action execution and changes in RGB-D during action execution to help the robot achieve efficient and collision-free operations. Our approach also incorporates a time-series design with prioritized experience replay to globally optimize the action sequence. Additionally, we propose a novel sim2real method by combining domain randomization to address the difference in object sizes between the simulation and the real world. Extensive experiments in both simulation and physical environments show that SVMMS-Grasp significantly outperforms existing methods in terms of task success rate, stability, and operational efficiency.
Yudong Luo, Feiyu Xie, Na Zhao 0008, Xianping Fu, Yantao Shen 0001
ICRA5
2025 Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning
abstract
Incomplete multi-view clustering (IMC) has garnered substantial attention due to its capacity to handle unlabeled data. Existing methods predominantly explore pairwise consistency between every two views. However, such consistency is highly susceptible to missing samples and outliers within a certain view and thus deviates from the true clustering distribution. Moreover, dual-view interaction neglects the collaboration effects of multiple views, making it challenging to capture the holistic characteristics across views. In response to these issues, we propose a novel Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning (CAL). Specifically, CAL reconstructs views with available instances to mine sample-wise affinities and harness comprehensive content information within views. Subsequently, to extract clean structural information, CAL imposes a structured sparse constraint on the representation tensor to eliminate biased errors. Furthermore, by integrating the consensus representation into a representation tensor, CAL can employ high-order interaction of multiple views to depict the semantic correlation between views while acquiring a unified structural graph across multiple views. Extensive experiments on seven benchmark datasets demonstrate that CAL outperforms some state-of-the-art methods in clustering performance. The code is available at https://github.com/whbdmu/CAL.
Huibing Wang, Jinjia Peng, Yawei Chen, Mingze Yao, Xianping Fu, Yang Wang 0023
IJCAI6
2025 AIS Data-Driven Maritime Monitoring Based on Transformer: A Comprehensive Review
abstract
With the increasing demands for safety, efficiency, and sustainability in global shipping, Automatic Identification System (AIS) data plays an increasingly important role in maritime monitoring. AIS data contains spatial-temporal variation patterns of vessels that hold significant research value in the marine domain. However, due to its massive scale, the full potential of AIS data has long remained untapped. With its powerful sequence modeling capabilities—particularly its ability to capture long-range dependencies and complex temporal dynamics—the Transformer model has emerged as an effective tool for processing AIS data. Therefore, this paper reviews the research on Transformer-based AIS data-driven maritime monitoring, providing a comprehensive overview of the current applications of Transformer models in the marine field. The focus is on Transformer-based trajectory prediction methods, behavior detection, and prediction techniques. Additionally, this paper collects and organizes publicly available AIS datasets from the reviewed papers, performing data filtering, cleaning, and statistical analysis. The statistical results reveal the operational characteristics of different vessel types, providing data support for further research on maritime monitoring tasks. Finally, we offer valuable suggestions for future research, identifying two promising research directions. Datasets are available at https://github.com/eyesofworld/Maritime-Monitoring.
Zhiye Xie, Enmei Tu, Xianping Fu, GuoLiang Yuan 0001
IJCNN3
2025 Data-Driven MPC for Attitude Control of Autonomous Underwater Robot
abstract
High maneuverability is essential to the autonomous operation of underwater robots. To achieve real-time maneuvering motion, the control strategy must take into account nonlinear hydrodynamic effects, which are extremely difficult to accurately capture during motion and therefore a balance must be struck between accuracy and real-time computational efficiency. Therefore, this paper proposes a data-driven approach to model the dynamics of the underwater robot using Sparse Identification of Nonlinear Dynamics (SINDy). Compared with existing works, our method does not require any physical prior knowledge and only uses a short period of onboard sensor data. Subsequently, the learned dynamic model is incorporated into a model predictive controller (MPC) to enable precise attitude control. Finally, the proposed method is implemented on our developed fully vectored propulsion underwater robot, and a series of attitude tracking experiments are conducted in an indoor water tank. Experimental results reveal that our approach significantly improves the model accuracy and reduces the attitude tracking errors by over 79% at a control frequency of 20 Hz, which proves the effectiveness and real-time performance of the method.
Tianzhu Gao, Yudong Luo, Na Zhao 0008, Yuanchu Yan, Xianping Fu, Yantao Shen 0001
IROS6
2025 Dual-Constraint Multi-view Fuzzy Clustering with Scalable Anchor Graph Learning
Luyan Cui, Huibing Wang, Yawei Chen, Mingze Yao, Xianping Fu, Jiqing Zhang
ACM Multimedia5
2025 Eye-based Emotion Recognition via Event-Driven Sparse Transformers
abstract
Event-driven eye-based emotion recognition has attracted increasing attention due to the high temporal resolution and dynamic range inherent to event cameras. The intrinsic spatial sparsity of event data, combined with the eye-based emotion recognition task's reliance on localized features such as eyebrows and eyelids, makes it intuitive and efficient to discard less informative regions. However, integrating such sparsification into CNNs remains challenging due to their reliance on dense grid-based operations. In this paper, we propose an efficient vision transformer framework for eye-based emotion recognition with event cameras. Specifically, we present window selection and token selection schemes tailored for event data and eye-based emotion recognition, which can diminish computing demands while enhancing performance. Firstly, we estimate the importance of all local windows and discard those with limited information, reducing computational cost while emphasizing attention on the periocular region. Secondly, we further introduce an adaptive token pruning mechanism that jointly evaluates the input event data and tokens to predict a binary decision mask, identifying and discarding uninformative tokens. Extensive experiments validate that the proposed approach outperforms existing state-of-the-art methods in accuracy by a significant margin.
Zixuan Wan, Jiqing Zhang, Yafei Wang 0004, Zetian Mi, Xin Yang 0011, Xianping Fu, Huibing Wang
ACM Multimedia8
2025 Consensus guided incomplete multi-view clustering via geometric consistency learning
Huibing Wang, Mingze Yao, Yawei Chen, Jinjia Peng, Guangqi Jiang, Xianping Fu
Appl. Intell.7
2025 Enhancing cyber safety in e-learning environment through cybersecurity awareness and information security compliance: PLS-SEM and FsQCA analysis
Chrispus Zacharia Oroni, Xianping Fu, Daniela Daniel Ndunguru, Arsenyan Ani
Comput. Secur.2
2025 Detail-focused and polarization-guided multi-modality fusion for underwater image clarity enhancing
Mingze Yao, Huibing Wang, Xianping Fu
Eng. Appl. Artif. Intell.5
2025 Enhancing Image Security With a Novel Chaotic System: A Focus on Multiface Image Encryption in Smart Applications
abstract
To ensure stringent security strategies for image information involving personal privacy during conveyance and storage, we propose an innovative multiface privacy protection scheme based on chaos theory. Compared to single-face encryption algorithms, the proposed scheme has broader potential applications in fields, such as smart cities and smart transportation. Specifically, a new spatiotemporal chaotic system named the sine-cosine coupled mapping lattice system (SCCML) is designed. It features a larger parameter domain, complexity, and profound unpredictability, yet maintains a simpler construction aimed at providing potential benefits and implementations in the field of information security. In the proposed multiface privacy protection scheme, multiple faces within an image are rapidly and accurately identified and then encrypted using the proposed SCCML-based digital separation loop encryption algorithm. The encryption algorithm exhibits a synchronous scrambling diffusion mechanism. Additionally, the introduction of mixed multibase cascade diffusion offers multiple layers of security for facial data, prevents diffusion singularity, and enhances diversity, making it significantly more challenging to crack. Experimental verification on a real multiface image dataset shows that the algorithm is superior, practical, safe, and efficient.
Pengbo Liu 0001, Huipeng Liu, Herbert H. C. Iu, Xiaopeng Yan, Xianping Fu
IEEE Internet Things J.7
2025 Depth-Consistent Monocular Visual Trajectory Estimation for AUVs
abstract
Visual trajectory estimation can endow autonomous underwater vehicles (AUVs) with environmental perception capabilities and has broad application prospects in the fields, such as oceanographic surveys, underwater construction, and marine ranching. However, due to featureless images and depth ambiguity issues caused by underwater multiple mediums environments, monocular visual trajectory estimation in complex underwater environments remains a challenging problem. In this article, we propose a monocular visual trajectory estimation method for AUVs, which can address the challenges of featureless and depth ambiguity by leveraging deep image representations and multiview geometry. Specifically, we design a bidirectional optical flow consistency scheme that selects sparse correspondences from monocular dense predictions to deal with the featureless images, and then achieve AUV trajectory estimation through epipolar constraints. Furthermore, we propose an iterative depth-consistent method, which solves the problem of depth ambiguity by aligning geometrically triangulated depths to the scale-consistent deep depths. We also develop a low-cost, agile, and portable AUV picking system with real-time trajectory estimation capabilities, and carry out extensive experiments in the Yellow Sea to test its performance. The experimental results demonstrate the effectiveness of the proposed method.
Yangyang Wang 0005, Dongbing Gu, Jie Wang 0003, Xianping Fu
IEEE Internet Things J.5
2025 High sensitivity image encryption algorithm based on cascaded chaotic system
Pengbo Liu 0001, Herbert H. C. Iu, Qi Li 0029, Xianping Fu
J. Inf. Secur. Appl.6
2025 A real-world underwater turbid image enhancement benchmark and beyond
Yafei Wang 0004, Yuán-Ruì Yáng, Xianping Fu
Knowl. Based Syst.5
2025 A novel hybrid YOLO-O SegNet for object detection and optimization with DCNN-based object recognition in federated learning
Soomro Pir Dino, Xianping Fu, Santosh Kumar Banbhrani, Muhammad Asad 0002, Zayyanu Shuaibu
Multim. Tools Appl.2
2025 Traffic sign recognition model based on scale sequence features and high-order spatial interactions
Yafei Wang 0004, Wenju Li, Xianping Fu
Neural Comput. Appl.4
2025 Spatial Residual for Underwater Object Detection
abstract
Feature drift is caused by the dynamic coupling of target features and degradation factors, which reduce underwater detector performance. We redefine feature drift as the instability of target features within boundary constraints while solving partial differential equations (PDEs). From this insight, we propose the Spatial Residual (SR) block, which uses SkipCut to establish effective constraints across the network width for solving PDEs and optimizes the solution space. It is implemented as a general-purpose backbone with 5 Spatial Residuals (BSR5) for complex feature scenarios. Specifically, BSR5 extracts discrete channel slices through SkipCut, where each sliced feature is parsed within the appropriate data capacity. In gradient backpropagation, SkipCut functions as a ShortCut, optimizing information flow and gradient allocation to enhance performance and accelerate training. Experiments on the RUOD dataset show that BSR5-integrated DETRs and YOLOs achieve state-of-the-art results for conventional and end-to-end detectors. Specifically, our BSR5-DETR improves 1.3% and 2.7% AP than RT-DETR with ResNet-101, while reducing parameters by 41.6% and 6.6%, respectively. Further validation highlights BSR5's strong convergence and robustness, especially in training from scratch scenarios, making it well suited for data-scarce, resource-constrained, and real-time tasks.
Jingchun Zhou, Zongxin He, Dehuan Zhang, Siyuan Liu 0004, Xianping Fu, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Underwater image enhancement via brightness mask-guided multi-attention embedding
Zetian Mi, Xianping Fu
Signal Process. Image Commun.4
2025 Multi-Agent Q-Net Enhanced Coevolutionary Algorithm for Resource Allocation in Emergency Human-Machine Fusion UAV-MEC System
abstract
Unmanned aerial vehicle (UAV) assisted communication has emerged as a powerful technology for reliable and flexible emergency communications (e.g., earthquakes, hurricanes and floods), especially when the mobile infrastructure is seriously damaged. UAV assisted mobile edge computing (UAV-MEC) system can be deployed in the natural disaster area as communication relay or air mobile base stations to resume communication and provide computing resources for the users in disaster areas. However, the optimized resource allocation performance of UAV-MEC system can be further guaranteed with the human-fusion decision making. In this paper, we construct an emergency human-machine fusion UAV-MEC system consisting of multiple UAVs equipped with computing resources, and the human-machine decision makings are fused for UAV deployment. In order to solve the resource allocation problem of human-machine fusion UAV-MEC system, we establish an human-machine deep integration model for UAV-MEC system, and the UAVs are dispatched reasonably through human-machine fusion decision makings to maintain efficient communication in emergency communication areas. To minimize task latency and improve the computation efficiency in emergency human-machine fusion UAV-MEC system, we consider the number of dispatched UAVs, deployment plans, flight plans, and simultaneously optimize the task allocation scheme, priority order, and task offloading ratio. We propose a reinforcement learning framework combined with evolutionary algorithms, which is named as multi-agent Q-net enhanced cooperative genetic algorithm (MQCGA), for resource allocation of UAV. Based on neural network forecasts, the greedy rate during training processing can be dynamically controlled, and the learning ability of different agents can be strengthened. Simulation experiments are conducted to evaluate the proposed framework, and the results show that our proposed MQCGA algorithm is significantly superior to other algorithms in terms of latency and energy consumption. Note to Practitioners—With the development of MEC and the popularity of UAVs, the potential of UAV-assisted MEC draw much attention from industrial field. Considering the lack of communication capabilities in a certain area in an unexpected situation, UAVs can be quickly deployed to corresponding locations and provide computing services. In this paper, an UAV-MEC system that integrates human-machine decision-making for emergency communication situations, named human-machine fusion UAV-MEC system, is considered. The system divides the scene into regions, models users and UAVs, and provides detailed deployment schemes, maximizing the practicality and applicability of the scene. In order to improve the communication efficiency in the case of emergency communication, this paper proposes a new resource scheduling algorithm and adds human-machine decision-making to enable UAVs to continuously provide efficient services for a certain area. The experimental results provide practitioners with a theoretical basis, such as the task completion time, UAV energy consumption and computing resource scheduling. Applying the system to actual scenarios also requires two preconditions of the system, one is the information collected by the large UAV, and the other is the communication among the UAVs. These two preconditions facilitate the human-machine fusion UAV-MEC system deployment in practical applications.
Lu Sun 0004, Zhaolong Ning, Jie Wang 0003, Xianping Fu
IEEE Trans Autom. Sci. Eng.5
2025 Focus More on What? Guiding Multi-Task Training for End-to-End Person Search
abstract
End-to-end person search is a research domain that executes the task of locating and identifying a target individual from a large number of scene images via a multi-task framework. However, a major challenge for the learning of end-to-end methods is the inherent conflicts between the two sub-tasks: person detection focuses on identifying generic features of persons, while person re-identification (Re-ID) strives to find unique, distinguishing features for matching the target person. Unlike previous research focusing on model architectures, this paper delves into the end-to-end person search training process. We find that the unbalanced and conflicting training issues significantly impair the learning efficiency of the Re-ID sub-task, which directly influences person search accuracy. To address this, we propose a novel Guiding Multi-Task Training (GMT) framework that facilitates end-to-end balanced learning for person search. We introduce a Guiding Multi-Task Harmonious Learning (GMHL) module, which decouple the features and then performs intra- and cross-task feature interaction to enhance the learning of each sub-task. Moreover, GMT employs a Balancing Multi-Task Oriented Fusing (BMOF) method to explicitly enhance Re-ID sub-task learning through additional Re-ID training and target-guided multi-model parameters fusion. Extensive experiments on 2 benchmark datasets, CUHK-SYSU and PRW, show that GMT achieves leading performance with 96.0% mAP and 61.3% mAP, respectively.
Boyu Cai, Huibing Wang, Mingze Yao, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Fusion-Based Channel-Wise Isotropic Convergent Real-Time Underwater Image Enhancement
abstract
Existing underwater image enhancement (UIE) methods typically prioritize improving image quality at the expense of algorithmic efficiency. In this paper, we propose a fusion-based, channel-wise isotropic convergent UIE method designed for real-time performance. The proposed approach comprises three key modules: (i) a non-linear transformation module that corrects color casts and aligns the pixel distribution with the gray-world assumption (GWA); (ii) a channel-wise isotropic convergence scheme that reduces intensity distribution disparities across channels, promoting balanced convergence; and (iii) a patch-based enhancement strategy that divides the image into smaller patches to better capture local features and improve adaptability to non-uniform degradation. Moreover, certain critical steps in our method are optimized to achieve O(1) time complexity, allowing it to meet real-time requirements. Extensive experiments validate the effectiveness of each module in the proposed method, showcasing its superiority when compared to the existing state-of-the-art (SOTA) approaches. Code has been released at https://github.com/JohnChenS/FCICE_UnderwaterImageEnhancement.
Yuehan Chen, Jiqing Zhang, Haoming Tang, Huibing Wang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.7
2025 An Underwater Image Restoration Method With Polarization Imaging Optimization Model for Poor Visible Conditions
abstract
Polarization imaging is extensively employed in underwater image restoration due to its effectiveness in removing backscattered light. However, existing polarization imaging methods generally assume the degree of polarization (DoP) of the backscattering is spatially constant and estimate it from the background region, limiting their practical applications. To address these challenges, we propose an underwater image restoration method based on a polarization imaging optimization model (PIOM). First, we develop a novel polarization image formation model by fusing the DoP and angle of polarization (AoP) of backscattered light. Second, we introduce an adaptive particle swarm local optimization (APSLO) method based on the PIOM. This method decomposes the image into small blocks and employs an objective optimization function to estimate the local optimal fusion parameters. Additionally, we propose a robust polynomial spatial fitting method to reduce block artifacts and noise disturbances, achieving globally optimal fusion parameters. Finally, we fully consider the advantages of gamma correction, and propose an adaptive contrast enhancement method to balance brightness and contrast. Experimental results show that our PIOM effectively removes backscattering while preserving finer details, colors, and contours. The code and datasets will be available athttps://github.com/liyafengLYF/UIRPIOM.
Yuehan Chen, Jiqing Zhang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Robust Image Steganography via Color Conversion
abstract
In this paper, we propose a robust image steganography method utilizing color conversion, leveraging de-colorization and colorization models to achieve covert transmission of secret information. The motivation is to use color conversion of the stego image to conceal steganographic behavior. For the sender, secret information is embedded into the color cover image using a robust embedding algorithm based on quaternion exponent moments. The stego images are then de-colorized to obtain grayscale images, which can be transmitted over public channels. For the receiver, a corresponding colorization network is designed to reconstruct the stego image and extract the secret information. Additionally, an attack module using Gaussian noise is implemented to enhance the robustness of the proposed steganography. Given a color image, its grayscale version can be chosen from various options, making it difficult for attackers to detect steganographic activity as long as the generated grayscale image appears normal and meaningful. Extensive simulation results demonstrate the feasibility and scalability of the proposed steganography method.
Qi Li 0029, Bin Ma 0003, Xianping Fu, Xiaoyu Wang 0011, Chunpeng Wang 0001, Xiaolong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 TAFormer: A Transmission-Aware Transformer for Underwater Image Enhancement
abstract
The attenuation and scattering of different colors of light underwater are wavelength- and distance-dependent, leading to various degradation problems in underwater images. When enhancing underwater images, many deep learning-based methods rely solely on convolutional neural networks to learn a mapping from degraded images to clear images to achieve enhanced effects. However, such methods have limitations in capturing long-term dependencies, preventing them from accurately capturing the global information of images. Although Transformers can solve this problem, there is a lack of inductive bias in training due to the limited number of training datasets with certain degradation phenomena. To address this issue, a novel Swin Transformer based on physical perception is proposed for the first time. Swin Transformer is used to solve the long- and short-distance dependency problem. Additionally, the underwater image degradation process is considered in network design to solve the problem of poor inductive bias. Combining the advantages of physical imaging, convolutional neural networks and Transformer can effectively improve the visual quality of underwater images. Rich qualitative and quantitative experimental results show that our Transformer achieves competitive performance on 5 benchmark datasets.
Zetian Mi, Yulin Wang 0003, Shuaiyong Jiang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Underwater Vignetting Image Correction Based on Binary Polynomial Regularization and Latent Low-Rank Representation
abstract
Due to light attenuation and complex environments, underwater robots need to carry artificial light to improve visibility, which leads to issues such as brightness vignetting, low contrast, and color distortion in the captured underwater images. However, existing methods for enhancing underwater images often overlook the challenges caused by artificial light. To address these challenges, we construct a novel underwater vignetting image formation model and propose a correction method called UVIC. The method consists of three main modules: separating the vignetting component, separating the backscattering component, and adaptive brightness and color correction. In our model, based on the linear relationship between the image gradient and the coefficients of the binary polynomial, we introduce a binary polynomial regularization to separate the vignetting component without estimating the center of the vignetting. Additionally, the backscattering can be effectively separated by introducing a latent low-rank representation based on local consistency, without estimating atmospheric light and transmission parameters. Furthermore, we design an adaptive brightness and color correction module using the global brightness of the image L layer and the histogram distribution characteristics of the a and b layers to adjust the brightness and color bias of the image. Particularly, there are both additive and multiplicative operations, and we decompose the objective function into two submodels and solve them by the iterative reweighted least squares and alternating direction multiplier methods, respectively. Numerous experiments demonstrate that UVIC not only effectively corrects image brightness vignetting, but also improves color bias, contrast, and sharpness.
Yulin Wang 0003, Yueming Ma, Jiqing Zhang, Zetian Mi, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.6
2025 Tensor Completion Framework by Graph Refinement for Incomplete Multi-View Clustering
abstract
Incomplete Multi-view Clustering (IMVC) endeavors to harness information from multiple incomplete views to partition multi-view data into their respective clusters. How to recover missing information with lossless fidelity is the core of IMVC, which is of vital importance but challenging. Most of the existing methods include a feature recovery step to mitigate the negative impact of missing samples on the feature graph, however, these IMVC algorithms simply utilize the correlation between samples to recover the relationship between the unmissing instances and the missing instances while ignoring the consistency between views, which leads to often unsatisfactory recovery results. In addition, previous IMVC algorithms focus more on the recovery of incomplete data, ignoring the effect of the error term on incomplete graphs. This can mislead the recovery process of IMVC algorithm and the feature graph can be affected by anomalous information, which leads to degradation of clustering performance. To address this gap, this paper introduces the Tensor Completion Framework by Graph Refinement for Incomplete Multi-view Clustering (IMVC-TGR). IMVC-TGR separates the redundant information in each affine graph by graph refinement operation, aiming to mitigate the negative impact of error terms and redundant information on the feature graph during the recovery process. Meanwhile, IMVC-TGR stacks the feature graphs into tensors to explore intra-view correlation and inter-view consistency, so as to recover the relationship between missing samples and non-missing samples, and improve the quality of the feature graphs. Finally, IMVC-TGR introduces semantic consistency constraints and self-weighted fusion strategies into the high-quality feature graphs, aiming at preserving the complementary information between different views while balancing the contributions of the refined representation matrices of different views. The experimental results on multiple different datasets indicate that IMVC-TGR can achieve state-of-the-art performance.
Huibing Wang, Yawei Chen, Mingze Yao, Jinjia Peng, Xianping Fu
IEEE Trans. Multim.6
2025 Between/Within View Information Completing for Tensorial Incomplete Multi-View Clustering
abstract
Incomplete Multi-view Clustering (IMvC) receives increasing attention due to its effectiveness in solving data-missing problems. With the information loss in incomplete situations, the core of IMvC needs to consider effectively overcoming the challenge of missing views, that is, exploring the underlying correlations from available data and recovering the missing information. However, most existing IMvC methods overemphasize the recovery-first principle with integrating the existing data from different views while neglecting the influence of view consistency in IMvC task together with valuable within view information. In this paper, a novel Between/Within View Information Completing for Tensorial Incomplete Multi-view Clustering (BWIC-TIMC) has been proposed, in which between/within view information is jointly exploited for effectively completing the missing views. Specifically, the proposed method designs a dual tensor constraint module, which focuses on simultaneously exploring the view-specific correlations of incomplete views and enforcing the between view consistency across different views. With the dual tensor constraint, between/within view information can be effectively integrated for completing missing views for IMvC task. Furthermore, in order to balance different contributions of multiple views and alleviate the problem of feature degeneration, BWIC-TIMC implements an adaptive fusion graph learning strategy for consensus representation learning. Extensive comparative experiments with the-state-of-art baselines can demonstrate the effectiveness of BWIC-TIMC.
Mingze Yao, Huibing Wang, Yawei Chen, Xianping Fu
IEEE Trans. Multim.4
2025 Dark Channel Low-Rank Prior for Enhanced Single Underwater Image Restoration
Yulin Wang 0003, Zheng Liang 0001, Zetian Mi, Jiqing Zhang, Xianping Fu
Vis. Comput.5
2024 Model Predictive Control for an Autonomous Underwater Robot with Fully Vectored Propulsion
abstract
Due to the low motion efficiency and maneuver-ability of underwater robots with six degrees of freedom, it is challenging for them to respond quickly to the attitude requirements during underwater autonomous manipulation. This paper presents a novel autonomous underwater robot with fully vectored propulsion and a model predictive control method to achieve more agile and efficient movements autonomously. In detail, we first design a robot with eight vector-distributed thruster layouts for fully vectored propulsion and construct the software architecture based on the robot operating system (ROS). Then, we establish the hydrodynamic model by adopting the Fossen approach and construct a 13-dimensional system state-space equation, which is discretized using the explicit fourth-order Runge-Kutta method. To achieve autonomous manipulation, model predictive control is employed along with physical constraints of the custom-built robot to enable real-time prediction and optimization of the robot’s states for control purposes. Finally, numerical simulations and experiments of the Point-to-Point Motion are conducted to test the robot’s performance. Experimental results reveal that the average error of each direction is 0.0027 m, 0.0031 m, and 0.0368 m in the x-axis, y-axis, and z-axis, respectively, and 0.8502°, 2.1941°, 0.2408° corresponding to three attitude angles, which verify the performance of employing MPC to control an autonomous underwater robot with fully vectored propulsion.
Tianzhu Gao, Yudong Luo, Weirong Luo, Xianping Fu, Na Zhao 0008, Yantao Shen 0001
ICRA5
2024 Fast One-Stage Unsupervised Domain Adaptive Person Search
Tianxiang Cui, Huibing Wang, Jinjia Peng, Ruoxi Deng, Xianping Fu, Yang Wang 0023
IJCAI5
2024 Scene-Adaptive Person Search via Bilateral Modulations
Huibing Wang, Jinjia Peng, Xianping Fu, Yang Wang 0023
IJCAI4
2024 A depth map stitching framework based on salient region matching
Zetian Mi, Haixia Qi, Huibing Wang, Xianping Fu
Eng. Appl. Artif. Intell.7
2024 HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image Enhancement
abstract
Underwater image enhancement presents a significant challenge due to the complex and diverse underwater environments that result in severe degradation phenomena such as light absorption, scattering, and color distortion. More importantly, obtaining paired training data for these scenarios is a challenging task, which further hinders the generalization performance of enhancement models. To address these issues, we propose a novel approach, the Hybrid Contrastive Learning Regularization (HCLR-Net). Our method is built upon a distinctive hybrid contrastive learning regularization strategy that incorporates a unique methodology for constructing negative samples. This approach enables the network to develop a more robust sample distribution. Notably, we utilize non-paired data for both positive and negative samples, with negative samples are innovatively reconstructed using local patch perturbations. This strategy overcomes the constraints of relying solely on paired data, boosting the model’s potential for generalization. The HCLR-Net also incorporates an Adaptive Hybrid Attention module and a Detail Repair Branch for effective feature extraction and texture detail restoration, respectively. Comprehensive experiments demonstrate the superiority of our method, which shows substantial improvements over several state-of-the-art methods in terms of quantitative metrics, significantly enhances the visual quality of underwater images, establishing its innovative and practical applicability. Our code is available at: https://github.com/zhoujingchun03/HCLR-Net .
Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu
Int. J. Comput. Vis.8
2024 Correction: HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image Enhancement
Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu
Int. J. Comput. Vis.8
2024 Mobility-Enhancement Simultaneous Optical Wireless Communication and Energy Harvesting System for IoUT
abstract
For the Internet of Underwater Things (IoUT), a large-capacity communication mechanism with high mobility is crucial to achieving information interaction at both short and medium ranges. In addition, to alleviate the energy anxiety of IoUT devices, a mobility enhanced simultaneous underwater optical wireless communication (UOWC) and energy harvesting system is proposed and experimentally demonstrated in this paper. Compared with traditional UOWC systems, we first employ light-emitting diode arrays and solar panels as system transceivers, which can not only realize high-speed data transmission without complex alignment operations, but also achieves efficient energy harvesting. An experimental demonstration system with a channel length of 0.5 m is established and the energy harvesting and communication performance of the proposed system are analyzed under four typical water types. For the harvesting part, the optimal load resistance is measured for better charging efficiency. The maximum output power measured in pure water is 1.6 times more than that in turbid seawater. For the communication part, with a hardware pre-equalization circuit, the -3dB bandwidth of the system is increased from 2.3 MHz to 14 MHz. The on-off-keying (OOK) modulation signals with a maximum data rate of 19 Mbps are successfully transmitted through different water types. Furthermore, the proposed system can achieve reliable communication at data rates greater than 16 Mbps with a wide 90-degree field-of-view, which greatly improves the system response range. The proposed simultaneous UOWC and energy harvesting system shows great potential for the construction of high-dynamic and energy-efficient IoUT networks.
Anliang Liu, Xinping Liu, Xianping Fu
IEEE Internet Things J.3
2024 Underwater image restoration via spatially adaptive polarization imaging and color correction
Jiqing Zhang, Yuehan Chen, Haoming Tang, Xianping Fu
Knowl. Based Syst.6
2024 High-strength synergic-calibration attention system in YOLO for underwater object detection application
GuoLiang Yuan 0001, Huibing Wang, Xianping Fu
Multim. Syst.4
2024 Criss-cross global interaction-based selective attention in YOLO for underwater object detection
Huibing Wang, Tianzhu Gao, Xianping Fu
Multim. Tools Appl.5
2024 Towards applying image retrieval approach for finding semantic locations in autonomous vehicles
Salahuddin Unar, Yining Su, Xiu Zhao, Pengbo Liu 0001, Yafei Wang 0004, Xianping Fu
Multim. Tools Appl.6
2024 Low-rank matrix recovery via novel double nonconvex nonsmooth rank minimization with ADMM
Yulin Wang 0003, Xianping Fu
Multim. Tools Appl.3
2024 Underwater image dehazing using a novel color channel based dual transmission map estimation
Xiaohong Yan, Guangyuan Wang, Yafei Wang 0004, Xianping Fu
Multim. Tools Appl.6
2024 Adapt only once: Fast unsupervised person re-identification via relevance-aware guidance
Jinjia Peng, Jiazuo Yu 0002, Huibing Wang, Xianping Fu
Pattern Recognit.5
2024 IACC: Cross-Illumination Awareness and Color Correction for Underwater Images Under Mixed Natural and Artificial Lighting
abstract
Enhancing underwater images captured under mixed artificial and natural lighting conditions presents two critical challenges. Existing methods lack a unified luminance feature extraction paradigm for mixed lighting scenes, leading to imbalance in luminance features, and consequent local overexposure or underexposure. Additionally, some color correction methods, through the fusion of features across multiple color spaces neglect the information loss due to the absence of feature alignment in cross-space fusion. To address these challenges, we propose a specialized method, namely IACC, which unifies the luminance features of underwater images under mixed lighting and guides consistent enhancement across similar luminance regions. Furthermore, complementary colors are introduced to globally guide the correction of color discrepancies, preserving the structural consistency and mitigating potential structural information loss during the original image feature extraction. Extensive experiments on various underwater datasets demonstrate the superiority of our method, which outperforms state-of-the-art methods in both machine and human visual perception. Our code is available athttps://github.com/zhoujingchun03/IACC.
Jingchun Zhou, Qilin Gai, Dehuan Zhang, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu
IEEE Trans. Geosci. Remote. Sens.6
2024 Underwater Color Correction Network With Knowledge Transfer
abstract
Underwater images suffer from severe color distortion, due to the wavelength-dependent light attenuation and scattering. Various underwater image enhancement methods have been developed to improve the quality of degraded underwater images. However, contemporary approaches often overlook the impact of different scene colors on the overall process, potentially leading to undesired outcomes, such as enhanced images exhibiting excessive redness. In this paper, we observe that the color tones of degraded underwater images exhibit variability under the influence of different underwater targets and scenes. Each degraded color channel can be utilized to guide the color correction of other channels. Given this, a light-weight underwater color correction network, dubbed UCCNet, is presented to alleviate the issue of color corruption. In UCCNet, three parallel branches are designed to excavate the residual information within each color channel, subsequently leveraging these features to improve the quality of underwater images. Moreover, facing the challenge of effectively enhancing underwater images in diverse and complex scenes, the model UCCNet-KT is established based on UCCNet. In UCCNet-KT, the technology of knowledge transfer is designed to improve the generalization ability by enriching the dataset and constructing the loss function. Extensive experiments on various underwater datasets indicate the impressive performance of the UCCNet and UCCNet-KT qualitatively and quantitatively.
Yafei Wang 0004, Xianping Fu
IEEE Trans. Multim.5
2024 Manifold-Based Incomplete Multi-View Clustering via Bi-Consistency Guidance
abstract
Incomplete multi-view clustering primarily focuses on dividing unlabeled data into corresponding categories with missing instances, and has received intensive attention due to its superiority in real applications. Considering the influence of incomplete data, the existing methods mostly attempt to recover data by adding extra terms. However, for the unsupervised methods, a simple recovery strategy will cause errors and outlying value accumulations, which will affect the performance of the methods. Broadly, the previous methods have not taken the effectiveness of recovered instances into consideration, or cannot flexibly balance the discrepancies between recovered data and original data. To address these problems, we propose a novel method termed Manifold-based Incomplete Multi-view clustering via Bi-consistency guidance (MIMB), which flexibly recovers incomplete data among various views, and attempts to achieve biconsistency guidance via reverse regularization. In particular, MIMB adds reconstruction terms to representation learning by recovering missing instances, which dynamically examines the latent consensus representation. Moreover, to preserve the consistency information among multiple views, MIMB implements a biconsistency guidance strategy with reverse regularization of the consensus representation and proposes a manifold embedding measure for exploring the hidden structure of the recovered data. Notably, MIMB aims to balance the importance of different views, and introduces an adaptive weight term for each view. Finally, an optimization algorithm with an alternating iteration optimization strategy is designed for final clustering. Extensive experimental results on 6 benchmark datasets are provided to confirm that MIMB can significantly obtain superior results as compared with several state-of-the-art baselines.
Huibing Wang, Mingze Yao, Yawei Chen, Yunqiu Xu, Haipeng Liu 0004, Wei Jia 0001, Xianping Fu, Yang Wang 0023
IEEE Trans. Multim.7
2024 Graph-Collaborated Auto-Encoder Hashing for Multiview Binary Clustering
abstract
Unsupervised hashing methods have attracted widespread attention with the explosive growth of large-scale data, which can greatly reduce storage and computation by learning compact binary codes. Existing unsupervised hashing methods attempt to exploit the valuable information from samples, which fails to take the local geometric structure of unlabeled samples into consideration. Moreover, hashing based on auto-encoders aims to minimize the reconstruction loss between the input data and binary codes, which ignores the potential consistency and complementarity of multiple sources data. To address the above issues, we propose a hashing algorithm based on auto-encoders for multiview binary clustering, which dynamically learns affinity graphs with low-rank constraints and adopts collaboratively learning between auto-encoders and affinity graphs to learn a unified binary code, called graph-collaborated auto-encoder (GCAE) hashing for multiview binary clustering. Specifically, we propose a multiview affinity graphs' learning model with low-rank constraint, which can mine the underlying geometric information from multiview data. Then, we design an encoder-decoder paradigm to collaborate the multiple affinity graphs, which can learn a unified binary code effectively. Notably, we impose the decorrelation and code balance constraints on binary codes to reduce the quantization errors. Finally, we use an alternating iterative optimization scheme to obtain the multiview clustering results. Extensive experimental results on five public datasets are provided to reveal the effectiveness of the algorithm and its superior performance over other state-of-the-art alternatives.
Huibing Wang, Mingze Yao, Guangqi Jiang, Zetian Mi, Xianping Fu
IEEE Trans. Neural Networks Learn. Syst.5
2024 Multiple information perception-based attention in YOLO for underwater object detection
Huibing Wang, Tianxiang Cui, Zhicheng Guo, Xianping Fu
Vis. Comput.5
2024 Publisher Correction: Multiple information perception-based attention in YOLO for underwater object detection
Huibing Wang, Tianxiang Cui, Zhicheng Guo, Xianping Fu
Vis. Comput.5
2023 Target-Oriented Multi-criteria Band Selection for Hyperspectral Image
Huijuan Pang, Xudong Sun 0009, Xianping Fu, Huibing Wang
PRCV (7)3
2023 Learning interpretable shared space via rank constraint for multi-view clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu
Appl. Intell.5
2023 Unsupervised underwater image enhancement via content-style representation disentanglement
Pengli Zhu, Yancheng Liu, Yuanquan Wen, Minyi Xu, Xianping Fu, Siyuan Liu 0004
Eng. Appl. Artif. Intell.5
2023 Jointly adversarial networks for wavelength compensation and dehazing of underwater images
Xianping Fu, Xueyan Ding, Zheng Liang 0001, Yafei Wang 0004
Multim. Tools Appl.1
2023 Multi-dimensional, multi-functional and multi-level attention in YOLO for underwater object detection
Xudong Sun 0009, Huibing Wang, Xianping Fu
Neural Comput. Appl.4
2023 Visualized Multiple Image Selection Encryption Based on Log Chaos System and Multilayer Cellular Automata Saliency Detection
abstract
For the security of multiple image regions of interest, this paper designs a multi-image visual encryption algorithm based on region of interest (ROI). Firstly, a new one-dimensional chaotic system (1-DLCP) is proposed. The performance analysis shows that the system has better dynamic behavior. Then, the salient parts of multiple images are extracted using multilayer cellular automata saliency detection. The measured matrix generated by the proposed chaotic system compresses the extracted plaintext images respectively. The compressed image and the extracted salient parts are recombined into one image, and a scrambling-diffusion synchronization operation is performed on the recombined image to obtain a secret image. The final embedding stage uses integer wavelet transform (IWT) to embed the secret image into a color carrier image. It is worth mentioning that this design scheme not only reconstructs the significant part without damage, but also reduces the storage space.
Yining Su, Pengbo Liu 0001, Salahuddin Unar, Xingyuan Wang 0001, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.6
2023 Towards Adaptive Consensus Graph: Multi-View Clustering via Graph Collaboration
abstract
Multi-view clustering is a long-standing important task, however, it remains challenging to exploit valuable information from the complex multi-view data located in diverse high-dimensional spaces. The core issue is the effective collaboration of multiple views to holistically uncover the essential correlations between multi-view data through graph learning. Furthermore, it is indispensable for most existing methods to introduce an additional clustering step to produce the final clusters, which evidently reduces the uniform relationship between graph learning and clustering. Based on the above considerations, in this paper, we present a novel method named multi-view clustering via graph collaboration (MCGC). Based on the low-dimensional representation space developed by MCGC, it first perceives the correlations between samples in each individual view under the supervision of the Hilbert-Schmidt independence criterion (HSIC). Then, MCGC proposes learning a consensus graph by adaptively collaborating between all the views, which is able to uncover the essential structure of the multi-view data. Meanwhile, by imposing the rank constraint on the Laplacian matrix of the consensus graph to partition the multi-view data naturally into the required number of clusters, the optimal clustering results can be obtained directly without any postprocessing steps. Finally, the resulting optimization problem is solved by an alternating optimization scheme with guaranteed fast convergence. Extensive experiments on 5 benchmark multi-view datasets demonstrate that MCGC markedly outperforms the state-of-the-art baselines.
Huibing Wang, Guangqi Jiang, Jinjia Peng, Ruoxi Deng, Xianping Fu
IEEE Trans. Multim.5
2022 Parallelism Network with Partial-aware and Cross-correlated Transformer for Vehicle Re-identification
abstract
Vehicle re-identification (ReID) aims to identify a specific vehicle in the dataset captured by non-overlapping cameras, which plays a great significant role in the development of intelligent transportation systems. Even though CNN-based model achieves impressive performance for the ReID task, its Gaussian distribution of effective receptive fields has limitations in capturing the long-term dependence between features. Moreover, it is crucial to capture fine-grained features and the relationship between features as much as possible from vehicle images.
Guangqi Jiang, Huibing Wang, Jinjia Peng, Xianping Fu
ICMR4
2022 Learning latent features with local channel drop network for vehicle re-identification
Xianping Fu, Jinjia Peng, Guangqi Jiang, Huibing Wang
Eng. Appl. Artif. Intell.1
2022 A unified total variation method for underwater image enhancement
abstract
Underwater images usually suffer from color casts and low contrast due to the absorption and scattering of light by the water medium. The degradation is caused not only by the light attenuation on the scene-sensor path but also by the light attenuation on the water surface-scene path. To eliminate the dual-path light attenuation, we propose a novel unified total variation method based on an extended underwater imaging model. Unlike previous variation-based methods that only consider light propagation along the scene-sensor path, we additionally include light propagation along the water surface-scene path in the underwater imaging model. In the proposed variational framework, we transform underwater image enhancement into two subproblems and construct different prior knowledge-guided optimization functions for them. The two subproblems aim to remove the light attenuation along the scene-sensor and surface-scene paths. Moreover, we present an alternating direction minimization algorithm based on an augmented Lagrange multiplier to address the optimization problems. The subjective and objective experimental results on underwater images with different attenuation characteristics demonstrate that the proposed method achieves good performance in underwater image enhancement.
Xueyan Ding, Yafei Wang 0004, Zheng Liang 0001, Xianping Fu
Knowl. Based Syst.4
2022 Attention-guided dynamic multi-branch neural network for underwater image enhancement
Xiaohong Yan, Wenqiang Qin, Yafei Wang 0004, Guangyuan Wang, Xianping Fu
Knowl. Based Syst.5
2022 Self-calibrated driver gaze estimation via gaze pattern learning
GuoLiang Yuan 0001, Yafei Wang 0004, Huizhu Yan, Xianping Fu
Knowl. Based Syst.4
2022 An Image Dehazing Approach With Adaptive Color Constancy for Poor Visible Conditions
abstract
The presence of suspended particles in turbid media not only superimposes a veil on the scene but also changes the color of the scene, resulting in poor visibility and low contrast of images. It is essential to overcome these effects for further applications of the images. However, these effects are difficult to eliminate simultaneously and thus the restoration for the images becomes multitasked. In this letter, we propose an extended image formation model that introduces color constancy theory to discuss the change of light color when natural light penetrates the medium to the object. Based on the model, the image restoration task is divided into two parts: dehazing and color correction, and for these two parts, scene depth fusion-based dehazing and adaptive color constancy are proposed. The dehazing approach introduces a weighted fusion strategy to estimate scene depth to achieve a better dehazing effect. The color constancy approach improves the Gray World hypothesis with gain factors to adaptively compensate for the chromatic loss and extends the color deviation solved by color constancy from color temperature to medium. Extensive experiments on images of different scenes prove the effectiveness of the proposed method in image restoration.
Xueyan Ding, Yafei Wang 0004, Xianping Fu
IEEE Geosci. Remote. Sens. Lett.3
2022 A Color Cast Image Enhancement Method Based on Affine Transform in Poor Visible Conditions
abstract
In this letter, a simple yet effective dehazing framework is proposed, which consists of a novel color correction and a contrast enhancement. Most of the existing dehazing works focus on enhancing the contrast of the degraded images, but rarely of them concern about the color cast, which is ubiquitous in the scattering medium. To address the color distortion, an affine transform model-based color correction method is first proposed to improve the appearance of the image while preserving the details, which is inspired by the traditional color transfer. The color transfer alters the color values of a source image by sharing the global color statistics of a reference image, which makes it unsuitable to address the locally variable color deviations encountered in highly color distorted images as in poor visibility conditions (sandstorms and underwater). To alter color correction locally, we add local color fidelity and gradient constraint to the proposed technique, which overcomes the limitation that the traditional method depends too much on the global color statistics of the reference image and encourages it to handle the degraded image with various color casts and light conditions. In addition, a multiscale gradient-domain processing is applied to enhance the contrast. In this procedure, by extracting the information of different layers, we can easily restore the contrast while limiting the significant amplification of noise. The extensive qualitative and quantitative experiments reveal that the color and the contrast can be significantly improved by the proposed technique.
Zheng Liang 0001, Xueyan Ding, Yafei Wang 0004, Yulin Wang 0003, Xianping Fu
IEEE Geosci. Remote. Sens. Lett.6
2022 Effective Polarization-Based Image Dehazing With Regularization Constraint
abstract
Image taken in turbid media generally exists poor visibility and low contrast, which results from attenuation of the propagated light. In this letter, an effective polarization-based image dehazing method is proposed, which relies on the relationship between the angle of polarization (AoP) from the Stokes vector and the scattered light. To avoid the influence of noise, AoP is optimized based on regularization constraints. The regularization function is made using an assumption that adjacent pixels with similar colors have similar values of AoP. Moreover, according to the revised AoP information, all the key parameters can be effectively and automatically estimated without considering the no-object region (or the sky region) exists or not, which relies on a frequency prior strategy. Extensive experiments on real-world images demonstrate that the proposed method is more effective than several previous image restoration or enhancement works.
Zheng Liang 0001, Xueyan Ding, Zetian Mi, Yafei Wang 0004, Xianping Fu
IEEE Geosci. Remote. Sens. Lett.5
2022 A Generalized Enhancement Framework for Hazy Images With Complex Illumination
abstract
Images captured under low-light conditions are generally characterized by poor illumination, low contrast, and nonignorable large amount of noise. In order to improve the visibility in weak illumination scenes, multiple artificial light sources are used, which leads to severe uneven illumination of the scene. The main challenges of dehazing images with complex illumination are to suppress the boosting of unsightly noise when enhancing contrast and avoid overenhancement in bright glow regions. To circumvent problems above, this letter proposes a generalized enhancement framework, which works well not only in uniform light conditions but also in strongly nonuniform illumination low-light scenes. To achieve this, we first decompose the input hazy image into a structure layer containing low-frequency illumination variance and a texture layer containing large amount of high-frequency details. Sequentially, benefit from two derived masks that are intrinsically similar to weight maps, the proposed framework can perform regional adaptive brightness adjustment on the structure layer according to the distribution of light in the input image. Meanwhile, regions of effective details in the texture layer are assigned higher weights, while regions that belong to noise are suppressed. Finally, adding the enhanced texture layer back to the brightened structure layer, visually appealing results are generated. Experimental results on various scenarios demonstrate the superiority of the proposed framework over state-of-the-art methods in terms of both qualitative and quantitative.
Zetian Mi, Zheng Liang 0001, Xianping Fu
IEEE Geosci. Remote. Sens. Lett.5
2022 A Global-Local Spectral Weight Network Based on Attention for Hyperspectral Band Selection
abstract
Band selection (BS) methods based on deep learning have achieved significant development. However, most existing band selection methods commonly utilize a fully connected neural network (FCN) or convolutional neural network (CNN) to explore the correlation among bands and rarely combine the two styles of the network to select bands. Moreover, almost all the methods employ the form of the combination of$L_{1}$norm and Sigmoid to constitute attention model, which may lead to losing some informative band feature. To tackle these troubles, this letter proposes a novel band selection network using FCN and CNN, termed as global-local spectral weight network based on attention (GLSWA), in which the band features of each pixel is mined using the network of two types, and designing an attention-based scoring module (ASM) and a convolutional reconstruction module (CRM), respectively, so that each attention of band is adjusted by simultaneous considering the entire band features and successive one. Experimental results on three real hyperspectral image (HSI) datasets show that the proposed method achieves satisfactory accuracy than some state-of-the-art algorithms.
Xudong Sun 0009, Yuan Zhu 0004, Fengqiang Xu, Xianping Fu
IEEE Geosci. Remote. Sens. Lett.5
2022 Eliminating cross-camera bias for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Dongyan Chen, Huibing Wang, Xianping Fu
Multim. Tools Appl.6
2022 A natural-based fusion strategy for underwater image enhancement
Xiaohong Yan, Guangxin Wang, Guangqi Jiang, Yafei Wang 0004, Zetian Mi, Xianping Fu
Multim. Tools Appl.6
2022 Refined marine object detector with attention-based spatial pyramid pooling networks and bidirectional feature fusion strategy
Fengqiang Xu, Huibing Wang, Xudong Sun 0009, Xianping Fu
Neural Comput. Appl.4
2022 Conditional generative adversarial network with dual-branch progressive generator for underwater image enhancement
Yafei Wang 0004, Guangyuan Wang, Xiaohong Yan, Guangqi Jiang, Xianping Fu
Signal Process. Image Commun.6
2022 A novel biologically-inspired method for underwater image enhancement
Xiaohong Yan, Guangxin Wang, Guangyuan Wang, Yafei Wang 0004, Xianping Fu
Signal Process. Image Commun.5
2022 Tensorial Multi-View Clustering via Low-Rank Constrained High-Order Graph Learning
abstract
Multi-view clustering aims to partition multi-view data into different categories by optimally exploring the consistency and complementary information from multiple sources. However, most existing multi-view clustering algorithms heavily rely on the similarity graphs from respective views and fail to comprehend multiple views holistically. Moreover, due to the noise and redundancy maintained in the original data, the original errors of multiple similarity graphs will continue to accumulate in the process of constructing consistent graphs. These situations always lead to the limitation to effective fuse the essential information from multiple views, which always influences the clustering performance and cries out for reliable solutions. Based on the above considerations, we propose a novel method termed Tensorial Multi-view Clustering (TMvC), which learns high-order graph by low-rank tensor constraint to uncover the essential information stored in multiple views. TMvC first learns the Laplacian graphs of all views and stacks them into a tensor which can be viewed as a high-order graph. With the high-order graph, consistency and complementary information from different views can be propagated smoothly across all views. Then, based on low-rank constraint, high-order graph is constrained in the horizontal and vertical directions to better uncover the inter-view and inter-class correlations between multi-view data, which is of vital importance for multi-view clustering. Extensive experiments on document and image datasets demonstrate that TMvC can achieve the state-of-the-art performance for multi-view clustering.
Guangqi Jiang, Jinjia Peng, Huibing Wang, Zetian Mi, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.5
2022 GUDCP: Generalization of Underwater Dark Channel Prior for Underwater Image Restoration
abstract
This letter introduces an underwater image enhancement method to handle low contrast and color cast of underwater images. Firstly, with the help of hierarchical searching technique, we propose a novel backscattered light estimation method. And in this procedure, a novel scoring formula is considered into our method, which comprehensively considers multiple prior knowledge. Then, we generalize underwater dark channel prior (UDCP) approach to obtain more robust transmission estimation. In addition, we also develop a white balance method to further modify the appearance of the resultant image. Extensive experiments on real-world images demonstrate that the proposed method outperforms several previous image restoration or enhancement works.
Zheng Liang 0001, Xueyan Ding, Yafei Wang 0004, Xiaohong Yan, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.5
2021 Learning Multiple Semantic Knowledge For Cross-Domain Unsupervised Vehicle Re-Identification
abstract
Unsupervised Vehicle re-identification (reID) aims at searching the similar vehicles’ images from large unlabelled datasets captured in a multiple camera network, which is still a challenging task. In this paper, a multiple semantic knowledge learning approach is proposed to exploit the potential similarity of unlabeled samples, which builds multiple clusters from different views automatically with different cues. Specially, different from some existing works focus on the knowledge of one view, for each vehicle in the target domain, different semantic knowledge could be learned with the proposed focal drop network and several different labels can be assigned according these knowledge, which would be employed to train the vehicle reID model jointly. In addition, due to the unreliability of pseudo labels assigned by the clustering, the hard triplet center loss is proposed to take the difference of intra-cluster and inter-cluster into consideration for better training the unsupervised framework to adapt the unknown domain. Comprehensive experimental results clearly demonstrate that our method achieves excellent performance on both VehicleID dataset and VeRi-776 dataset.
Huibing Wang, Jinjia Peng, Guangqi Jiang, Xianping Fu
ICME4
2021 Band Selection for Specific Target Detection of Hyperspectral Imagery
abstract
Band selection (BS) is considered as an effective method for dimensionality reduction of hyperspectral data. As an important application for hyperspectral remote sensing, target detection is widely concerned. Therefore, how to select more representational band subset for specific target to improve performance of detection is worth discussing. This letter proposed a BS method for specific target detection, called constrained target band selection under adaptive subspace partitioning (CTASPBS). Firstly, all bands are partitioned into multiple weakly correlated subsets via adaptive subspace partition strategy (ASPS). Then, according to a target-constrained band prioritization (BP) criterion, the band with the highest priority in each subset is selected to form the optimal band subset. Due to the application of different BP criterion, two BS methods ASPS_MinV and ASPS_MaxV are proposed. Finally, experimental results on real hyperspectral data show that CTASPBS is an effective BS method for specific target detection.
Xudong Sun 0009, Site Li, Fengqiang Xu, Xianping Fu
IGARSS5
2021 MSAV: An Unified Framework for Multi-view Subspace Analysis with View Consistence
abstract
With the development of multimedia period, information is always caputred with multiple views, which causes a research upsurge on multi-view learning. It is obvious that multi-view data contains more information than those single view ones. Therefore, it is crucial to develop the multi-view algorithms to adapt the demand of many applications. Even though some excellent multi-view algorithms were proposed, most of them can only deal with the specific problems. To tacle this problem, this paper proposes an unified framework named Multi-view Subspace Analysis with View Consistence (MSAV), which provides an unified means to extend those single-view dimension reduciton algorithms into multi-view versions. MSAV first extends multi-view data into kernel space to avoid the problem caused by different dimensions of the data from multiple views. Then, we introduced a self-weighted learning strategy to automatically assign weights for all views according to their importance. Finally, in order to promote the consistence of all views, Hilbert-Schmidt Independence Criterion is adopted by MSAV. Furthermore, We conducted experiments on several benchmark datasets to verify the performance of MSAV.
Huibing Wang, Guangqi Jiang, Jinjia Peng, Xianping Fu
ICMR4
2021 Graph-based Multi-view Binary Learning for image clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu
Neurocomputing5
2021 Single underwater image enhancement by attenuation map guided color correction and detail preserved dehazing
Zheng Liang 0001, Yafei Wang 0004, Xueyan Ding, Zetian Mi, Xianping Fu
Neurocomputing5
2021 Discriminative feature and dictionary learning with part-aware model for vehicle re-identification
Huibing Wang, Jinjia Peng, Guangqi Jiang, Fengqiang Xu, Xianping Fu
Neurocomputing5
2021 Scale-aware feature pyramid architecture for marine object detection
Fengqiang Xu, Huibing Wang, Jinjia Peng, Xianping Fu
Neural Comput. Appl.4
2021 Depth-aware total variation regularization for underwater image dehazing
Xueyan Ding, Zheng Liang 0001, Yafei Wang 0004, Xianping Fu
Signal Process. Image Commun.4
2021 Kernelized Multiview Subspace Analysis By Self-Weighted Learning
abstract
With the popularity of multimedia technology, information is always represented from multiple views. Even though multiview data can reflect the same sample from different perspectives, multiple views are consistent to some extent because they are representations of the same sample. Most of the existing algorithms are graph-based ones to learn the complex structures within multiview data but overlook the information within data representations. Furthermore, many existing works treat multiple views discriminatively by introducing some hyperparameters, which is undesirable in practice. To this end, abundant multiview-based methods have been proposed for dimension reduction. However, there is still no research that leverages the existing work into a unified framework. In this paper, we propose a general framework for multiview data dimension reduction, named kernelized multiview subspace analysis (KMSA) to handle multiview feature representation in the kernel space, providing a feasible channel for multiview data with different dimensions. Compared with the graph-based methods, KMSA can fully exploit information from multiview data with nothing to lose. Since different views have different influences on KMSA, we propose a self-weighted strategy to treat different views discriminatively. A co-regularized term is proposed to promote the mutual learning from multiviews. KMSA combines self-weighted learning with the co-regularized term to learn the appropriate weights for all views. We evaluate our proposed framework on 6 multiview datasets for classification and image retrieval. The experimental results validate the advantages of our proposed method.
Huibing Wang, Yang Wang 0023, Zhao Zhang 0001, Xianping Fu, Li Zhuo 0001, Mingliang Xu 0001, Meng Wang 0001
IEEE Trans. Multim.4
2020 Unsupervised Vehicle Re-identification with Progressive Adaptation
abstract
Vehicle re-identification (reID) aims at identifying vehicles across different non-overlapping cameras views. The existing methods heavily relied on well-labeled datasets for ideal performance, which inevitably causes fateful drop due to the severe domain bias between the training domain and the real-world scenes; worse still, these approaches required full annotations, which is labor-consuming. To tackle these challenges, we propose a novel Progressive Adaptation Learning method for vehicle reID, named PAL, which infers from the abundant data without annotations. For PAL, a data adaptation module is employed for source domain, which generates the images with similar data distribution to unlabeled target domain as “pseudo target samples”. These pseudo samples are combined with the unlabeled samples that are selected by a dynamic sampling strategy to make training faster. We further proposed a weighted label smoothing (WLS) loss, which considers the similarity between samples with different clusters to balance the confidence of pseudo labels. Comprehensive experimental results validate the advantages of PAL on both VehicleID and VeRi-776 dataset.
Jinjia Peng, Yang Wang 0023, Huibing Wang, Zhao Zhang 0001, Xianping Fu, Meng Wang 0001
IJCAI5
2020 Purifying real images with an attention-guided style transfer network for gaze estimation
Xianping Fu, Yuxiao Yan, Jinjia Peng, Huibing Wang
Eng. Appl. Artif. Intell.1
2020 Joint rain and atmospheric veil removal from single image
abstract
In natural rainy scenes, visibility is significantly degraded by two types of phenomena: specular highlights of nearby individual rain streaks and atmospheric veiling effect caused by distant accumulated rain. However, most existing deraining methods only take the first kind of degradation into consideration, which limits their potential application in heavy rain. In this study, a joint rain and atmospheric veil removal framework is proposed to address this problem. Since rain streaks and rain accumulation are entangled with each other, which is intractable to simulate, causing clean/rainy image pairs of real‐world are hard to generate. Hence, after introducing a generalised rain model, which can represent both rain streaks and atmospheric veil physically, the authors do not learn the mapping function between image pairs using deep‐learning architecture, but estimate the rain streaks, transmission, and atmospheric light via Gaussian mixture model patch prior and dark channel prior to solve the rain model instead. According to the comprehensive experimental evaluations, the proposed method outperforms other state‐of‐the‐art methods in terms of both high visibility and vivid colour, especially in natural heavy rain scenario.
Zetian Mi, Yafei Wang 0004, Congcong Zhao, Fengming Du, Xianping Fu
IET Image Process.5
2020 Cross domain knowledge learning with dual-branch adversarial network for vehicle re-identification
Jinjia Peng, Huibing Wang, Fengqiang Xu, Xianping Fu
Neurocomputing4
2020 Vehicle re-identification using multi-task deep learning network and spatio-temporal model
Jinjia Peng, Fengqiang Xu, Xianping Fu
Multim. Tools Appl.4
2019 Purifying naturalistic images through a real-time style transfer semantics network
Yuxiao Yan, Ibrahim Shehi Shehu, Xianping Fu, Huibing Wang
Eng. Appl. Artif. Intell.4
2019 Learning multi-region features for vehicle re-identification with context-based ranking method
Jinjia Peng, Huibing Wang, Xianping Fu
Neurocomputing4
2019 Co-regularized multi-view sparse reconstruction embedding for dimension reduction
Huibing Wang, Jinjia Peng, Xianping Fu
Neurocomputing3
2019 Auto-weighted Mutli-view Sparse Reconstructive Embedding
Huibing Wang, Haohao Li, Xianping Fu
Multim. Tools Appl.3
2019 Multi-feature distance metric learning for non-rigid 3D shape retrieval
Huibing Wang, Haohao Li, Jinjia Peng, Xianping Fu
Multim. Tools Appl.4
2019 Guiding intelligent surveillance system by learning-by-synthesis gaze estimation
Yuxiao Yan, Jinjia Peng, Zetian Mi, Xianping Fu
Pattern Recognit. Lett.5
2018 Refining Synthetic Images with Semantic Layouts by Adversarial Training
abstract
Recently, progress in learning-by-synthesis has proposed training models on synthetic images, which can effectively reduce the cost of manpower and material resources. However, learning from synthetic images still fails to achieve the desired performance compared to naturalistic images due to the different distribution of synthetic images. In an attempt to address this issue, previous methods were to improve the realism of synthetic images by learning a model. However, the disadvantage of the method is that the distortion has not been improved and the authenticity level is unstable. To solve this problem, we put forward a new structure to improve synthetic images, via the reference to the idea of style transformation, through which we can efficiently reduce the distortion of pictures and minimize the need of real data annotation. We estimate that this enables generation of highly realistic images, which we demonstrate both qualitatively and with a user study. We quantitatively evaluate the generated images by training models for gaze estimation. We show a significant improvement over using synthetic images, and achieve state-of-the-art results on various datasets including MPIIGaze dataset.
Yuxiao Yan, Jinjia Peng, HaoHui Wei, Xianping Fu
ACML5
2018 Image Purification Networks: Real-time Style Transfer with Semantics through Feed-forward Synthesis
abstract
Image synthesis has been widely accepted as a cost effective way to learn models because it provides training sets that are large, diverse and accurately labeled. However, the realism of the synthetic image is not enough, this affects generalization on naturalistic test image. In an attempt to address this issue, previous methods learn a model to improve the realism of synthetic image. Differently, from previous methods, we take the first step towards purifying the naturalistic image to weaken the influence of light and convert the distribution of an outdoor naturalistic image through a real-time style transfer task to that of indoor synthetic image. This paper proposes, therefore a real-time image purification networks that transfer style information with semantics through a feed-forward synthesis. Results from our experiments demonstrate that images purified through the proposed networks architecture trained models for gaze estimation more accurately on cross-datasets over using raw naturalistic images and when compared to baseline methods.
Yuxiao Yan, Ibrahim Shehi Shehu, Xianping Fu
IJCNN4
2018 Learning a gaze estimator with neighbor selection from large-scale synthetic eye images
Yafei Wang 0004, Xueyan Ding, Jinjia Peng, Jiming Bian, Xianping Fu
Knowl. Based Syst.6
2016 Appearance-based gaze estimation using deep features and random forest regression
Yafei Wang 0004, Tianyi Shen, GuoLiang Yuan 0001, Jiming Bian, Xianping Fu
Knowl. Based Syst.5
2015 Experience of experimental teaching and management based on cloud computing
abstract
Many laboratories were set up in Chinese higher education institutions in recent years. Some of them in the same university are equipped with similar rigs, which may lead to waste of money and human power in their setup and maintenance, and inconvenience and inefficiency for both faculties and students. An experimental teaching and management system, Open Laboratory Access (OLA), based on cloud computing technology is presented to enable full sharing and deep integration of experimental resources among laboratories and provide improved user experience. OLA provides a comprehensive and multi-angle virtual experimental teaching and management functions, including online experiment schedule and laboratory rig reservation, intelligent experiment guidance, teaching effectiveness evaluation, automatical experiment results statistics, etc. Thin-clients are adopted to reduce cost and ensure consistent user experience. After a one-year trial, the results of a user experience survey indicate that OLA provides good experience of experiment for both faculties and students.
Xueyu Geng, Xianping Fu
FIE4
2013 Automatic Calibration Method for Driver's Head Orientation in Natural Driving Environment
abstract
Gaze tracking is crucial for studying driver's attention, detecting fatigue, and improving driver assistance systems, but it is difficult in natural driving environments due to nonuniform and highly variable illumination and large head movements. Traditional calibrations that require subjects to follow calibrators are very cumbersome to be implemented in daily driving situations. A new automatic calibration method, based on a single camera for determining the head orientation and which utilizes the side mirrors, the rear-view mirror, the instrument board, and different zones in the windshield as calibration points, is presented in this paper. Supported by a self-learning algorithm, the system tracks the head and categorizes the head pose in 12 gaze zones based on facial features. The particle filter is used to estimate the head pose to obtain an accurate gaze zone by updating the calibration parameters. Experimental results show that, after several hours of driving, the automatic calibration method without driver's corporation can achieve the same accuracy as a manual calibration method. The mean error of estimated eye gazes was less than 5°in day and night driving.
Xianping Fu, Xiao Guan, Eli Peli, Hongbo Liu 0001, Gang Luo 0003
IEEE Trans. Intell. Transp. Syst.1
2003 Video coding of model based at very low bit rates
Xianping Fu
VCIP1