Lei Zhang 0038

dblp:64/5666-38 · DBLP profile ↗
← Back
121ranked-venue papers
25as first author
76since 2021 · last 2026
0000-0002-5305-8543ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 12 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 9 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Computer networks · 11 · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 ADPS-Sat: Adaptive Distributed Patch-Sequence Scheduling for Satellite-Edge Vision Transformers
Haochun Lei, Yuben Qu, Zhen Qin 0005, Lei Zhang 0038, Kefeng Guo, Chao Dong 0001, Qihui Wu 0001, Kapal Dev
ICC4
2026 Exploiting foundational spatial semantic prior for accurate few-shot hyperspectral image classification
Xingbing Zhao, Lei Zhang 0038, Weixin Ren, Pengfei Bai
Expert Syst. Appl.3
2026 Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm
Fuxiang Huang, Xiaowei Fu, Shiyu Ye, Wen Li 0001, Xinbo Gao 0001, David Zhang 0001, Lei Zhang 0038
Int. J. Comput. Vis.8
2026 Implicit Non-Causal Factors are Out via Dataset Splitting for Domain Generalization Object Detection
Lei Zhang 0038, Shuyin Xia, Guoyin Wang 0001, Fuxiang Huang
Int. J. Comput. Vis.2
2026 Reliable Covert Communication in NOMA-Aided Cognitive Satellite Aerial Terrestrial Integrated Networks
abstract
NOMA-aided cognitive satellite aerial terrestrial integrated networks (CSATINs) are considered revolutionary and key technologies for 6G Internet of Things (6G-IoT), offering enhanced connectivity, high spectral efficiency, and broad coverage. In this article, we first establish trustworthy CSATINs with multiple aerial relays, aiming to achieve reliable communication in the presence of an eavesdropper. Then, to enhance the system’s covert performance, we propose an unmanned aerial vehicle scheduling scheme. Moreover, based on the established covert system model, we derive the closed-form expressions of detection error probability (DEP), covert outage probability (COP), and effective covert rate (ECR). Particularly, an optimization is proposed to enhance the covert performance of the considered system. Finally, Monte Carlo simulations are given to validate the correctness of the theoretical analysis, demonstrating that the reliability and covertness of the proposed system can be simultaneously enhanced by appropriately adjusting the power allocation coefficients, jamming power, and the transmission power of the satellite and UAVs.
Peilin Qi, Kefeng Guo, Ali Nauman, Qihui Wu 0001, Lei Zhang 0038, Zeke Wu, Keshav Singh 0001
IEEE Internet Things J.5
2026 M3C: Resist Agnostic Attacks by Mitigating Consistent Class Confusion Prior
abstract
Adversarial attack is a major obstacle to the deployment of deep neural networks (DNNs) for security-sensitive applications. To address these adversarial perturbations, various adversarial defense strategies have been developed, with Adversarial Training (AT) being one of the most effective methods to protect neural networks from adversarial attacks. However, existing AT methods struggle against training-agnostic attacks due to their limited generalizability. This suggests that the AT models lack a unified perspective for various attacks to conduct universal defense. This paper sheds light on a generalizable prior under various attacks: consistent class confusion (3C), i.e., an AT classifier often confuses the predictions between correct and ambiguous classes in a highly similar pattern among diverse attacks. Relying on this latent prior as a bridge between seen and agnostic attacks, we propose a more generalized AT model by mitigating consistent class confusion (M3C) to resist training-agnostic attacks. Specifically, we optimize an Adversarial Confusion Loss (ACL), which is weighted by uncertainty, to distinguish the most confused classes and encourage the AT model to focus on these confused samples. To suppress malignant features affecting correct predictions and producing significant class confusion, we propose a Gradient-Aware Attention (GAA) mechanism to enhance the classification confidence of correct classes and eliminate class confusion. Experiments on multiple benchmarks and network frameworks demonstrate that our M3C model significantly improves the generalization of AT robustness against agnostic attacks. The finding of the 3C prior reveals the potential and possibility for defending against a wide range of attacks, and provides a new perspective to overcome such challenge in this field.
Xiaowei Fu, Fuxiang Huang, Guoyin Wang 0001, Xinbo Gao 0001, Lei Zhang 0038
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 AMG-Net: A multitask network with adaptive mutual guidance for Semantic Change Detection
Yuduo Bian, Wei Wei 0008, Chen Ding 0002, Lei Zhang 0038, Jiangbin Zheng 0001, Yanning Zhang 0001
Pattern Recognit.4
2026 Generative model-based mixed-semantic enhancement for transductive zero-shot learning
Huaizhou Qi, Yang Liu 0069, Jungong Han, Lei Zhang 0038
Pattern Recognit.4
2026 ARKFNet: A neural network-enhanced anomaly-robust Kalman filter
Shanli Chen, Dongyuan Lin, Peng Cai 0002, Lei Zhang 0038
Signal Process.5
2026 Reliable Covert Communication for Integrated Cognitive Satellite-Aerial-Terrestrial Networks With NOMA and Poisson-Distributed Jammers
Kefeng Guo, Peilin Qi, Shahid Mumtaz, Yuzhen Huang 0001, Ali Nauman, Lei Zhang 0038, Qihui Wu 0001
IEEE Trans. Commun.6
2026 Token Calibration for Transformer-Based Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain by learning domain-invariant representations. Motivated by the recent success of Vision Transformers (ViTs), several UDA approaches have adopted ViT architectures to exploit fine-grained patch-level representations, which are unified as Transformer-based $D$ omain $A$ daptation (TransDA) independent of CNN-based. However, we have a key observation in TransDA: due to inherent domain shifts, patches (tokens) from different semantic categories across domains may exhibit abnormally high similarities, which can mislead the self-attention mechanism and degrade adaptation performance. To solve that, we propose a novel $P$ atch- $A$ daptation Transformer (PATrans), which first identifies similarity-anomalous patches and then adaptively suppresses their negative impact to domain alignment, i.e. token calibration. Specifically, we introduce a $P$ atch- $A$ daptation $A$ ttention (PAA) mechanism to replace the standard self-attention mechanism, which consists of a weight-shared triple-branch mixed attention mechanism and a patch-level domain discriminator. The mixed attention integrates self-attention and cross-attention to enhance intra-domain feature modeling and inter-domain similarity estimation. Meanwhile, the patch-level domain discriminator quantifies the anomaly probability of each patch, enabling dynamic reweighting to mitigate the impact of unreliable patch correspondences. Furthermore, we introduce a contrastive attention regularization strategy, which leverages category-level information in a contrastive learning framework to promote class-consistent attention distributions. Extensive experiments on four benchmark datasets demonstrate that PATrans attains significant improvements over existing state-of-the-art UDA methods (e.g., 89.2% on the VisDA-2017). Code is available at: https://github.com/YSY145/PATrans.
Xiaowei Fu, Shiyu Ye, Chenxu Zhang 0001, Fuxiang Huang, Xin Xu 0001, Lei Zhang 0038
IEEE Trans. Image Process.6
2026 Rectifying Adversarial Sample With Low Entropy Prior for Test-Time Defense
abstract
Existing defense methods fail to defend against un known attacks and thus raise generalization issue of adversarial robustness. To remedy this problem, we attempt to delve into some underlying common characteristics among various attacks for generality. In this work, we reveal the commonly overlooked low entropy prior (LE) implied in various adversarial samples, and shed light on the universal robustness against unseen attacks in inference phase. LE prior is elaborated as two properties across various attacks as shown in Fig. 1 and 2: 1) low entropy misclassification for adversarial samples and 2) lower entropy prediction for higher attack intensity. This phenomenon stands in stark contrast to the naturally distributed samples. The LE prior can instruct existing test-time defense methods, thus we propose a two-stage REAL approach: Rectify Adversarial sample based on LE prior for test-time adversarial rectification. Specifically, to align adversarial samples more closely with clean samples, we propose to first rectify adversarial samples misclassified with low entropy by reverse maximizing prediction entropy, thereby eliminating their adversarial nature. To ensure the rectified samples can be correctly classified with low entropy, we carry out secondary rectification by forward minimizing prediction entropy, thus creating a Max-Min entropy optimization scheme. Further, based on the second property, we propose an attack aware weighting mechanism to adaptively adjust the strengths of Max-Min entropy objectives. Experiments on several datasets show that REAL can greatly improve the performance of existing sample rectification models.
Xiaowei Fu, Fuxiang Huang, Xinbo Gao 0001, Lei Zhang 0038
IEEE Trans. Multim.5
2026 Vision Mamba: A Comprehensive Survey and Taxonomy
abstract
State space model (SSM) is a mathematical model used to describe and analyze the behavior of dynamic systems. This model has witnessed numerous applications in several fields, including control theory, signal processing, economics, and machine learning. In the field of deep learning, SSMs are used to process sequence data, such as time series analysis, natural language processing (NLP), and video understanding. By mapping sequence data to state space, long-term dependencies in the data can be better captured. In particular, modern SSMs have shown strong representational capabilities in NLP, especially in long sequence modeling, while maintaining linear time complexity. In particular, based on the latest SSMs, Mamba merges time-varying parameters into SSMs toward efficient training and inference. Given its impressive efficiency and strong long-range dependency modeling capability, Mamba is expected to become a new AI architecture that may be capable of surpassing Transformer. Recently, a number of works attempt to study the potential of Mamba in various fields, such as general vision, multimodal learning, medical image analysis, and remote sensing image analysis, by extending Mamba from natural language domain to visual domain. To fully understand Mamba in the visual domain, we conduct a comprehensive survey and present a taxonomy study. This survey focuses on Mamba's application to a variety of visual tasks and data types, and discusses its predecessors, recent advances, and far-reaching impact on a wide range of domains.
Chenxu Zhang 0001, Fuxiang Huang, Shuyin Xia, Guoyin Wang 0001, Lei Zhang 0038
IEEE Trans. Neural Networks Learn. Syst.6
2025 Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot Learning
Fei Zhou 0008, Lei Zhang 0038, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001
ICCV3
2025 Delay Optimization in Remote ID-Based UAV Communication via BLE and Wi-Fi Switching
abstract
The remote identification (Remote ID) broadcast capability allows unmanned aerial vehicles (UAVs) to exchange messages, which is a pivotal technology for inter-UAV communications. Although this capability enhances the operational visibility, low delay in Remote ID-based communications is critical for ensuring the efficiency and timeliness of multi-UAV operations in dynamic environments. To address this challenge, we first establish delay models for Remote ID communications by considering packet reception and collisions across both BLE 4 and Wi-Fi protocols. Building upon these models, we formulate an optimization problem to minimize the long-term communication delay through adaptive protocol selection. Since the delay performance varies with the UAV density, we propose an adaptive BLE/Wi-Fi switching algorithm based on the multi-agent deep Q-network approach. Experimental results demonstrate that in dynamic-density scenarios, our strategy achieves 32.1% and 37.7% lower latency compared to static BLE 4 and Wi-Fi modes respectively.
Ziye Jia, Lei Zhang 0038, Qiuming Zhu, Qihui Wu 0001
PIMRC3
2025 UAV-Assisted MEC for Disaster Response: Stackelberg Game-Based Resource Optimization
abstract
The unmanned aerial vehicle assisted multi-access edge computing (UAV-MEC) technology has been widely applied in the sixth-generation era. However, due to the limitations of energy and computing resources in disaster areas, how to efficiently offload the tasks of damaged user equipments (UEs) to UAVs is a key issue. In this work, we consider a multiple UAVMECs assisted task offloading scenario, which is deployed inside the three-dimensional corridors and provide computation services for UEs. In detail, a ground UAV controller acts as the central decision-making unit for deploying the UAV-MECs and allocates the computational resources. Then, we model the relationship between the UAV controller and UEs based on the Stackelberg game. The problem is formulated to maximize the utility of both the UAV controller and UEs. To tackle the problem, we design a K-means based UAV localization and availability response mechanism to pre-deploy the UAV-MECs. Then, a chess-like particle swarm optimization probability based strategy selection learning optimization algorithm is proposed to deal with the resource allocation. Finally, extensive simulation results verify that the proposed scheme can significantly improve the utility of the UAV controller and UEs in various scenarios compared with baseline schemes.
Yafei Guo, Ziye Jia, Lei Zhang 0038, Yu Zhang 0047, Qihui Wu 0001
VTC2025-Spring3
2025 Joint ADS-B in B5G for Hierarchical AAV Networks: Performance Analysis and MEC-Based Optimization
abstract
Autonomous aerial vehicles (AAVs) play significant roles in multiple fields, which brings great challenges for the airspace safety. In order to achieve efficient surveillance and break the limitation of application scenarios caused by single communication, we propose the collaborative surveillance model for hierarchical AAVs based on the cooperation of automatic dependent surveillance-broadcast (ADS-B) and 5G. Specifically, AAVs are hierarchical deployed, with the low-altitude central AAV equipped with the 5G module, and the high-altitude central AAV with ADS-B, which helps automatically broadcast the flight information to surrounding aircraft and ground stations. First, we build the framework, derive the analytic expression, and analyze the channel performance of both air-to-ground (A2G) and air-to-air (A2A). Then, since the redundancy or information loss during transmission aggravates the monitoring performance, the mobile edge computing (MEC) based on-board processing algorithm is proposed. Finally, the performances of the proposed model and algorithm are verified through both simulations and experiments. In detail, the redundant data filtered out by the proposed algorithm accounts for 53.48%, and the supplementary data accounts for 16.42% of the optimized data. This work designs a AAV monitoring framework and proposes an algorithm to enhance the observability of trajectory surveillance, which helps improve the airspace safety and enhance the air traffic flow management.
Chao Dong 0001, Yiyang Liao, Ziye Jia, Qihui Wu 0001, Lei Zhang 0038
IEEE Internet Things J.5
2025 Zero-Shot Sketch-Based Image Retrieval with teacher-guided and student-centered cross-modal bidirectional knowledge distillation
Jiale Du, Yang Liu 0069, Xinbo Gao 0001, Jungong Han, Lei Zhang 0038
Pattern Recognit.5
2025 Remove to Regenerate: Boosting Adversarial Generalization With Attack Invariance
abstract
Adversarial attacks pose a huge challenge to the deployment of deep neural networks (DNNs) in security-sensitive applications. Adversarial defense methods are developed to resist adversarial perturbation. However, most defenses overlook the generalization to various attacks. In medical field, it is known that targeted therapy is a treatment approach at the cellular and molecular levels that targets already identified carcinogenic sites. Inspired by the popular targeted therapies for cancer, we view adversarial attacks as local lesions of natural benign samples. The mechanism behind this assumption implies our key finding that the salient attack components in an adversarial sample dominate the attacking process, while trivial attack components unexpectedly provide trustworthy evidence for obtaining generalizable robustness. Based on this finding, an explainable but efficient Adversarial Surgery and Regeneration (ASR) model following the targeted therapy mechanism is developed to improve the adversarial generalization of DNNs, which has three merits: 1) A score-based Pixel Surgery (PS) module is proposed to remove the salient attack components while retaining the trivial attack components as a kind of attack-invariant information. 2) A Semantic Regeneration module (SR) based on a conditional alignment extrapolator is proposed to restore the discriminative content from the attack-free trivial components, which achieves pixel and semantic consistency for adversarial samples. 3) To further harmonize robustness and accuracy and address such an intractable problem in adversarial defense, a self-augmentation regularizer with adversarial R-drop (ARD) is designed. Experiments on numerous benchmarks show the superiority of the proposed ASR approach. The code can be found inhttps://github.com/fxw13/ASR.
Xiaowei Fu, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multi-Model Synergy Perception for Open-World Person Re-Identification
abstract
Open-world person re-identification aims to train a model on source doamins and generalize well on unseen domains. Existing domain generalizable person re-identification methods primarily employ the equality training paradigm to train the model on multi-source domains. However, in open-world scenarios, domain imbalance often causes domain bias issue that leads to sub-optimal generalization ability, which is seriously overlooked. In this paper, we propose a Multi-model Synergy Perception (MSP) framework equipped with an Asynchronous Training Paradigm (ATP) on biased domains to maintain the domain balance for exploring the domain-invariant features. With the philosophy of divide and conquer, we divide the biased source domains into multiple debiased sub-source domains and employ a multi-network architecture to learn these sub-source domains in parallel. Additionally, to better generalize knowledge across these sub-source domains, we propose a Structure Synergy Perception (SSP) module that constructs the feature relationship distribution for each sub-domain and aligns them to map the unique knowledge to each other. Furthermore, considering the consistency of sub-source domains, we further propose a Synergy Distillation Perception (SDP) to improve the model both semantic and domain generalization ability. The main idea of SDP is to use the center guided soft label and the part based triplet graph to distill each submodel, which can facilitate the network to explore domain-invariant representations of images. Extensive experiments demonstrate that our method outperforms state-of-the-arts for open-domain person ReID.
Zhipu Liu, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.2
2025 D3Former: Dual-Decoder Dual-Transformer Reconstruction Network for Hyperspectral Anomaly Detection
Tan Guo, Yukun Yang 0006, Fulin Luo, Chuan Fu, Lei Zhang 0038
IEEE Trans. Geosci. Remote. Sens.5
2025 Infrared Small Target Detection via Diverse Feature Harmonization
abstract
Infrared small target detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, poor contrast in infrared images, and the tendency for small and dim targets to be obscured by cluttered backgrounds. These factors complicate the extraction of diverse and effective information for target detection, with existing encoder-decoder architectures often causing irreversible loss of small target features through continuous down-sampling, resulting in missed and false detections. To address these challenges, we propose the Diverse Feature Capture and Harmonization Network (DFCHNet), which learns and harmonizes diverse features through multiple encoding paths. DFCHNet includes an infrared image reconstruction branch running in parallel with the detection branch, preserving small target information and reducing feature loss via complementary contextual encoding. Additionally, we introduce FTConv to capture target edges while suppressing background noise. A Cross-layer Feature Autonomous Selection (CFAS) method adaptively harmonizes cross-layer features. We also propose a Coordinate Calibration (CC) loss function and a two-stage training strategy to refine predicted target positions. Experimental results on three IRSTD datasets demonstrate that DFCHNet outperforms current state-of-the-art methods.
Tan Guo, Baojiang Zhou, Fulin Luo, Lei Zhang 0038
IEEE Trans. Geosci. Remote. Sens.4
2025 DENet: Direction and Edge Co-Awareness Network for Road Extraction From High-Resolution Remote Sensing Imagery
abstract
Automatic road extraction from high-resolution remote sensing images has greatly facilitated the applications of high-precision road mapping in autonomous driving and intelligent transportation. However, challenges such as occlusions from buildings, trees, and complex road shapes bring great difficulties to precise road extraction. Also, existing methods often overlook the integrity of road direction and edge, leading to unsatisfactory extraction results. To alleviate the issue, this paper has presented a direction and edge co-awareness network (DENet). Firstly, the road edge detector (RED) is introduced to extract coarse road edges with abundant directional information. By leveraging the edge enhancement blocks, the edge structures of road can be efficiently refined, achieving the extraction of intricate narrow and elongated road shapes. Secondly, we incorporate the directional spatial attention (DSA) mechanism within the dual encoders and decoders to promote the extraction and fusion of road directional information and elongated features from different orientations, thus greatly mitigating the road occlusion issue. Finally, to fully interlace potential road information, a grouped local-global feature fusion (GLFF) is specifically designed to exchange multi-scale semantic information across different channels, simultaneously emphasizing road features and suppressing irrelevant background features. Numerous experimental results on three public datasets demonstrate the effectiveness and efficiency of the proposed DENet for road extraction, achieving F1 scores of 78.51% on the CHN6-CUG dataset, 79.35% on the Massachusetts road dataset, and 77.90% on the GF2-FC dataset, outperforming several existing state-of-the-art methods. The code is available at:https://github.com/gwy103/DENet.
Tan Guo, Fulin Luo, Lei Zhang 0038, Bo Du 0001, Xinbo Gao 0001
IEEE Trans. Intell. Transp. Syst.4
2024 Dynamic Weighted Combiner for Mixed-Modal Image Retrieval
abstract
Mixed-Modal Image Retrieval (MMIR) as a flexible search paradigm has attracted wide attention. However, previous approaches always achieve limited performance, due to two critical factors are seriously overlooked. 1) The contribution of image and text modalities is different, but incorrectly treated equally. 2) There exist inherent labeling noises in describing users' intentions with text in web datasets from diverse real-world scenarios, giving rise to overfitting. We propose a Dynamic Weighted Combiner (DWC) to tackle the above challenges, which includes three merits. First, we propose an Editable Modality De-equalizer (EMD) by taking into account the contribution disparity between modalities, containing two modality feature editors and an adaptive weighted combiner. Second, to alleviate labeling noises and data bias, we propose a dynamic soft-similarity label generator (SSG) to implicitly improve noisy supervision. Finally, to bridge modality gaps and facilitate similarity learning, we propose a CLIP-based mutual enhancement module alternately trained by a mixed-modality contrastive loss. Extensive experiments verify that our proposed model significantly outperforms state-of-the-art methods on real-world datasets. The source code is available at https://github.com/fuxianghuang1/DWC.
Fuxiang Huang, Lei Zhang 0038, Xiaowei Fu, Suqi Song
AAAI2
2024 Implicit Discriminative Knowledge Learning for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-identification (VI-ReID) is a challenging cross-modal pedestrian retrieval task, due to significant intra-class variations and cross-modal discrepancies among different cameras. Existing works mainly focus on embedding images of different modalities into a unified space to mine modality-shared features. They only seek distinctive information within these shared features, while ignoring the identity-aware useful information that is implicit in the modality-specific features. To address this issue, we propose a novel Implicit Discriminative Knowledge Learning (IDKL) network to uncover and leverage the implicit discriminative information contained within the modality-specific. First, we extract modality-specific and modality-shared features using a novel dual-stream network. Then, the modality-specific features undergo purification to reduce their modality style discrepancies while preserving identity-aware discriminative knowledge. Subsequently, this kind of implicit knowledge is distilled into the modality-shared feature to enhance its distinctiveness. Finally, an alignment loss is proposed to minimize modality discrepancy on enhanced modality-shared features. Extensive experiments on multiple public datasets demonstrate the superiority of IDKL network over the state-of-the-art methods. Code is available at https://github.com/1KK077/IDKL.
Kaijie Ren, Lei Zhang 0038
CVPR2
2024 Robust Overfitting Does Matter: Test-Time Adversarial Purification with FGSM
abstract
Numerous studies have demonstrated the susce, of deep neural networks (DNNs) to subtle adversar turbations, prompting the development of many at adversarial defense methods aimed at mitigating ac ial attacks. Current defense strategies usually train for a specific adversarial attack method and can good robustness in defense against this type of ac ial attack. Nevertheless, when subjected to evan involving unfamiliar attack modalities, empirical e reveals a pronounced deterioration in the robust DNNs. Meanwhile, there is a trade-off between the classification accuracy of clean examples and adversarial examples. Most defense methods often sacrifice the accuracy of clean examples in order to improve the adversarial robustness of DNNs. To alleviate these problems and enhance the overall robust generalization of DNNs, we propose the Test-Time Pixel-Level Adversarial Purification (TPAP) method. This approach is based on the robust overfitting character-istic of DNNs to the fast gradient sign method (FGSM) on training and test datasets. It utilizes FGSM for adversarial purification, to process images for purifying unknown adversarial perturbations from pixels at testing time in a “counter changes with changelessness” manner, thereby enhancing the defense capability of DNNs against various unknown adversarial attacks. Extensive experimental results show that our method can effectively improve both overall robust generalization of DNNs, notably over pre-vious methods. Code is available https://github.com/tly18/TPAP.
Linyu Tang, Lei Zhang 0038
CVPR2
2024 Urban Waterlogging Detection: A Challenging Benchmark and Large-Small Model Co-adapter
Suqi Song, Chenxu Zhang 0001, Pengkun Li, Fenglong Song, Lei Zhang 0038
ECCV (36)6
2024 DNN Tasks Offloading and Bandwidth Optimization for Satellite-Terrestrial Collaborative Intelligence
abstract
Deep Neural Networks (DNNs) are now widely used in Low Earth Orbit (LEO) satellites, such as in remote sensing and environmental monitoring. DNN tasks are generally resource-intensive, while the resources of LEO satellites including computation and storage resources are usually limited, which implies directly running high-precision and complex DNNs on them is extremely challenging. A promising way is leveraging the layered structure of DNNs and executing DNN tasks collaboratively between satellites and ground, i.e., satellite-terrestrial collaborative inference. However, most existing works about satellite- terrestrial collaborative inference mainly focus on the optimization of DNN offloading strategy in terms of latency and energy minimization, without considering how to minimize the highly precious satellite communication resources in the collaboration. In this paper, we study how to jointly optimize the offloading decision and satellites' communication bandwidth, to achieve the minimization of weighted sum of latency, energy consumption, and communication bandwidth consumption. The aforementioned problem is a Mixed Integer Nonlinear Programming (MINLP) problem and hard to resolve. We design an alternating optimization algorithm combining branch-and-bound and gradient descent methods (AO-SA) to obtain an efficient solution. Extensive simulations validate the efficiency of the proposed algorithm: compared to existing satellite-terrestrial offloading algorithms, it improves the performance in terms of latency and energy consumption by up to 31 %, while saving the bandwidth resource of satellites by 28 % on average.
Haochun Lei, Yuben Qu, Lei Zhang 0038, Lingyuan Zhao, Guangxia Li, Qihui Wu 0001
MSN3
2024 Joint ADS-B in 5G for Hierarchical Aerial Networks: Performance Analysis and Optimization
abstract
Unmanned aerial vehicles (UAVs) are widely applied in multiple fields, which emphasizes the challenge of obtaining UAV flight information to ensure the airspace safety. UAVs equipped with automatic dependent surveillance-broadcast (ADSB) devices are capable of sending flight information to nearby aircrafts and ground stations (GSs). However, the saturation of limited frequency bands of ADS-B leads to interferences among UAVs and impairs the monitoring performance of GS to civil planes. To address this issue, the integration of the 5th generation mobile communication technology (5G) with ADS-B is proposed for UAV operations in this paper. Specifically, a hierarchical structure is proposed, in which the high-altitude central UAV is equipped with ADS-B and the low-altitude central UAV utilizes 5G modules to transmit flight information. Meanwhile, based on the mobile edge computing technique, the flight information of sub-UAVs is offloaded to the central UAV for further processing, and then transmitted to GS. We present the deterministic model and stochastic geometry based model to build the air-to-ground channel and air-to-air channel, respectively. The effectiveness of the proposed monitoring system is verified via simulations and experiments. This research contributes to improving the airspace safety and advancing the air traffic flow management.
Ziye Jia, Yiyang Liao, Chao Dong 0001, Lijun He 0005, Qihui Wu 0001, Lei Zhang 0038
PIMRC6
2024 Efficient Pipeline Collaborative DNN Inference in Resource-Constrained UAV Swarm
abstract
Recent advancements in unmanned aerial vehicle (UAV) technology have propelled the popularity of edge intelligence (EI) applications with deep learning in UAV swarm. Nevertheless, the high computational demands of deep neural networks (DNNs) conflict with the limited computing power and battery capacity of UAV. Furthermore, many UAV applications require real-time performance such as object detection and recognition. In this paper, we study how to achieve fast DNN inference in UAV swarm by the collaboration of multiple UAVs, and formulate the problem of minimizing the completion time of a series of arriving DNN inference tasks, under memory and energy constraints. To solve the aforementioned challenging problem with combinatorial explosion, we propose an efficient solution exploiting deep reinforcement learning (DRL) with action space simplification to find the allocation strategy of each DNN inference task within a resource-constrained UAV swarm. Simulation results validate the effectiveness of the proposed solution compared to five benchmark algorithms.
Weiqing Ren, Yuben Qu, Zhen Qin 0005, Chao Dong 0001, Fuhui Zhou, Lei Zhang 0038, Qihui Wu 0001
WCNC6
2024 An Open-World, Diverse, Cross-Spatial-Temporal Benchmark for Dynamic Wild Person Re-Identification
Lei Zhang 0038, Xiaowei Fu, Fuxiang Huang, Yi Yang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.1
2024 Gradient Harmonization in Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) intends to transfer knowledge from a labeled source domain to an unlabeled target domain. Many current methods focus on learning feature representations that are both discriminative for classification and invariant across domains by simultaneously optimizing domain alignment and classification tasks. However, these methods often overlook a crucial challenge: the inherent conflict between these two tasks during gradient-based optimization. In this paper, we delve into this issue and introduce two effective solutions known as Gradient Harmonization, including GH and GH++, to mitigate the conflict between domain alignment and classification tasks. GH operates by altering the gradient angle between different tasks from an obtuse angle to an acute angle, thus resolving the conflict and trade-offing the two tasks in a coordinated manner. Yet, this would cause both tasks to deviate from their original optimization directions. We thus further propose an improved version, GH++, which adjusts the gradient angle between tasks from an obtuse angle to a vertical angle. This not only eliminates the conflict but also minimizes deviation from the original gradient directions. Finally, for optimization convenience and efficiency, we evolve the gradient harmonization strategies into a dynamically weighted loss function using an integral operator on the harmonized gradient. Notably, GH/GH++ are orthogonal to UDA and can be seamlessly integrated into most existing UDA models. Theoretical insights and experimental analyses demonstrate that the proposed approaches not only enhance popular UDA baselines but also improve recent state-of-the-art models.
Fuxiang Huang, Suqi Song, Lei Zhang 0038
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Meta Invariance Defense Towards Generalizable Robustness to Unknown Adversarial Attacks
abstract
Despite providing high-performance solutions for computer vision tasks, the deep neural network (DNN) model has been proved to be extremely vulnerable to adversarial attacks. Current defense mainly focuses on the known attacks, but the adversarial robustness to the unknown attacks is seriously overlooked. Besides, commonly used adaptive learning and fine-tuning technique is unsuitable for adversarial defense since it is essentially a zero-shot problem when deployed. Thus, to tackle this challenge, we propose an attack-agnostic defense method named Meta Invariance Defense (MID). Specifically, various combinations of adversarial attacks are randomly sampled from a manually constructed Attacker Pool to constitute different defense tasks against unknown attacks, in which a student encoder is supervised by multi-consistency distillation to learn the attack-invariant features via a meta principle. The proposed MID has two merits: 1) Full distillation from pixel-, feature- and prediction-level between benign and adversarial samples facilitates the discovery of attack-invariance. 2) The model simultaneously achieves robustness to the imperceptible adversarial perturbations in high-level image classification and attack-suppression in low-level robust image regeneration. Theoretical and empirical studies on numerous benchmarks such as ImageNet verify the generalizable robustness and superiority of MID under various attacks.
Lei Zhang 0038, Yi Yang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Semi-supervised domain generalization with evolving intermediate domain
Luojun Lin, Zhishu Sun, Weijie Chen 0006, Wenxi Liu, Yuanlong Yu 0001, Lei Zhang 0038
Pattern Recognit.7
2024 Learnable Background Endmember With Subspace Representation for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to label each hyperspectral image (HSI) pixel as background or anomaly, in a totally unsupervised manner. Thus, a fine background representation is vital to obtain good HAD performance. This article introduces background endmember representation and proposes a novel HAD method termed learnable background endmember with subspace representation (LEBSR). First, the HSI is unmixed to simultaneously obtain the background endmembers and their abundances. The three constraints of$p$-norm, sum-to-one, and nonnegativity work together to promote a more meaningful and accurate background endmember representation. In addition, a mapping matrix with orthogonality is jointly optimized to transform the priori backgrounds into the background endmember subspace, and then the mapped priori backgrounds are approximated to the low-rank representation (LRR) with the background endmembers. With the methodology, the backgrounds can be well reconstructed under the guidance of the priori background information to accurately detect anomalous pixels with the reconstructed residuals. The experimental results on several HSI datasets verify the superior performance of LEBSR than the state-of-the-art methods.https://github.com/HalongL/HAD-LEBSR
Tan Guo, Fulin Luo, Xiuwen Gong, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 DMFNet: Dual-Encoder Multistage Feature Fusion Network for Infrared Small Target Detection
abstract
Infrared Small Target Detection (IRSTD) is a challenging task of identifying small targets with low signal-to-noise ratios in complex backgrounds. Traditional methods in the complex background of IRSTD lead to a large number of false alarms and missdetections. Although CNN-based methods have made progress in IRSTD, how to extract more effective information and fully utilize inter-layer information remains an unresolved issue. Therefore, this paper proposed a Dual-Encoder Multi-stage Feature Fusion Network (DMFNet). Specifically, we designed a dual-encoder with different inputs to capture more effective small target feature information. We then designed a Receptive Field Expansion Attention Module (REAM) to incorporate non-local contextual information. In the decoding phase, the Triple Cross-layer Fusion Module (TCFM) was developed to exchange the low-level spatial details and the high-level semantic information for preserving more small target information in deeper layers. Finally, by concatenating multi-scale features from various layers of the decoder, more discriminative feature maps were generated to clearly describe the infrared small targets. Experimental results on the NUDT-SIRST, NUAA-SIRST, and IRSTD-1k datasets demonstrated that DMFNet outperforms some other state-of-the-art methods, achieving superior detection performance. Codes: https://github.com/BJZHOU2000/DMFNet.
Tan Guo, Baojiang Zhou, Fulin Luo, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 All-Sky Autonomous Computing in UAV Swarm
abstract
Unmanned aerial vehicles (UAVs) play an essential role in emergency cases and adverse environments for applications like disaster detection and mine exploration. To process the massive volume of sensing data generated by various sensory payloads in these applications, existing works either compress deep learning (DL) models to conduct onboard computing, or offload raw data back to the resourceful ground station with the help of relay UAVs due to base station damage. However, the former sacrifices the inference accuracy of DL models (up to 10% accuracy loss), while the latter achieves high accuracy at the cost of significant latency, due to limited wireless communication resources in the multi-hop transmission. To address the problem, exploiting the resources of the UAV swarm including both task UAVs and relay UAVs, we build up anall-skyautonomous computing (ASAP) system to autonomously conduct collaborative computing in the swarm, to achieve both high accuracy and low latency of sensing data processing. In detail, we first propose a novel UAV swarm-native collaborative computing architecture, considering the general hierarchy and clustering structure of UAV swarms, as well as the characteristic of DL model execution. We then design an elastic efficient task scheduler to allocate computing tasks for UAVs, and update the scheduling scheme online when some UAVs are unavailable, with the aid of a lightweight and accurate DL inference performance predictor. Finally, we design an adaptive inter-UAV data compressor, to adapt to the limited and dynamic communication resources between UAVs. Experiment results on 24 airborne computers and five real-world UAVs show that, the proposed system can perform collaborative computing in a timely manner and effectively deal with situations when some UAVs become unavailable.
Yuben Qu, Chao Dong 0001, Haipeng Dai 0001, Zhenhua Li 0001, Lei Zhang 0038, Qihui Wu 0001, Song Guo 0001
IEEE Trans. Mob. Comput.6
2024 Balancing Transferability and Discriminability for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to leverage a sufficiently labeled source domain to classify or represent the fully unlabeled target domain with a different distribution. Generally, the existing approaches try to learn a domain-invariant representation for feature transferability and add class discriminability constraints for feature discriminability. However, the feature transferability and discriminability are usually not synchronized, and there are even some contradictions between them, which is often ignored and, thus, reduces the accuracy of recognition. In this brief, we propose a deep multirepresentations adversarial learning (DMAL) method to explore and mitigate the inconsistency between feature transferability and discriminability in UDA task. Specifically, we consider feature representation learning at both the domain level and class level and explore four types of feature representations: domain-invariant, domain-specific, class-invariant, and class-specific. The first two types indicate the transferability of features, and the last two indicate the discriminability. We develop an adversarial learning strategy between the four representations to make the feature transferability and discriminability to be gradually synchronized. A series of experimental results verify that the proposed DMAL achieves comparable and promising results on six UDA datasets.
Jingke Huang, Ni Xiao, Lei Zhang 0038
IEEE Trans. Neural Networks Learn. Syst.3
2024 Transfer Adaptation Learning: A Decade Survey
abstract
The world we see is ever-changing and it always changes with people, things, and the environment. Domain is referred to as the state of the world at a certain moment. A research problem is characterized as transfer adaptation learning (TAL) when it needs knowledge correspondence between different moments/domains. TAL aims to build models that can perform tasks of target domain by learning knowledge from a semantic-related but distribution different source domain. It is an energetic research field of increasing influence and importance, which is presenting a blowout publication trend. This article surveys the advances of TAL methodologies in the past decade, and the technical challenges and essential problems of TAL have been observed and discussed with deep insights and new perspectives. Broader solutions of TAL being created by researchers are identified, i.e., instance reweighting adaptation, feature adaptation, classifier adaptation, deep network adaptation, and adversarial adaptation, which are beyond the early semisupervised and unsupervised split. The survey helps researchers rapidly but comprehensively understand and identify the research foundation, research status, theoretical limitations, future challenges, and understudied issues (universality, interpretability, and credibility) to be broken in the field toward generalizable representation in open-world scenarios.
Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Multi-view Adversarial Discriminator: Mine the Non-causal Factors for Object Detection in Unseen Domains
abstract
Domain shift degrades the performance of object detection models in practical applications. To alleviate the influence of domain shift, plenty of previous work try to decouple and learn the domain-invariant (common) features from source domains via domain adversarial learning (DAL). However, inspired by causal mechanisms, we find that previous methods ignore the implicit insignificant non-causal factors hidden in the common features. This is mainly due to the single-view nature of DAL. In this work, we present an idea to remove non-causal factors from common features by multi-view adversarial training on source domains, because we observe that such insignificant non-causal factors may still be significant in other latent spaces (views) due to the multi-mode structure of data. To summarize, we propose a Multi-view Adversarial Discriminator (MAD) based domain generalization model, consisting of a Spurious Correlations Generator (SCG) that increases the diversity of source domain by random augmentation and a Multi-View Domain Classifier (MVDC) that maps features to multiple latent spaces, such that the non-causal factors are removed and the domain-invariant features are purified. Extensive experiments on six benchmarks show our MAD obtains state-of-the-art performance.
Mingjun Xu, Lingyun Qin, Weijie Chen 0006, Shiliang Pu, Lei Zhang 0038
CVPR5
2023 Parameter Exchange for Robust Dynamic Domain Generalization
abstract
Agnostic domain shift is the main reason of model degradation on the unknown target domains, which brings an urgent need to develop Domain generalization (DG). Recent advances at DG use dynamic networks to achieve training-free adaptation on the unknown target domains, termed Dynamic Domain Generalization (DDG), which compensates for the lack of self-adaptability in static models with fixed weights. The parameters of dynamic networks can be decoupled into a static and a dynamic component, which are designed to learn domain-invariant and domain-specific features, respectively. Based on the existing arts, in this work, we try to push the limits of DDG by disentangling the static and dynamic components more thoroughly from an optimization perspective. Our main consideration is that we can enable the static component to learn domain-invariant features more comprehensively by augmenting the domain-specific information. As a result, the more comprehensive domain-invariant features learned by the static component can then enforce the dynamic component to focus more on learning adaptive domain-specific features. To this end, we propose a simple yet effective Parameter Exchange (PE) method to perturb the combination between the static and dynamic components. We optimize the model using the gradients from both the perturbed and non-perturbed feed-forward jointly to implicitly achieve the aforementioned disentanglement. In this way, the two components can be optimized in a mutually-beneficial manner, which can resist the agnostic domain shifts and improve the self-adaptability on the unknown target domain. Extensive experiments show that PE can be easily plugged into existing dynamic networks to improve their generalization ability without bells and whistles.
Luojun Lin, Zhifeng Shen, Zhishu Sun, Yuanlong Yu 0001, Lei Zhang 0038, Weijie Chen 0006
ACM Multimedia5
2023 Multi-adversarial Faster-RCNN with Paradigm Teacher for Unrestricted Object Detection
Zhenwei He, Lei Zhang 0038, Xinbo Gao 0001, David Zhang 0001
Int. J. Comput. Vis.2
2023 PSGAN: Revisit the binary discriminator and an alternative for face frontalization
Qingyan Duan, Lei Zhang 0038, Yan Zhang 0108, Xinbo Gao 0001
Neurocomputing2
2023 Coarse-to-fine sparse self-attention for vehicle re-identification
Fuxiang Huang, Xuefeng Lv, Lei Zhang 0038
Knowl. Based Syst.3
2023 BP-triplet net for unsupervised domain adaptation: A Bayesian perspective
Shanshan Wang 0008, Lei Zhang 0038, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001
Pattern Recognit.2
2023 Exploring Implicit Domain-Invariant Features for Domain Adaptive Object Detection
abstract
Recent researches have made a great progress in domain adaptive object detectors. These detectors aim to learn explicit domain-invariant features by adversarially mitigating domain divergence and simultaneously optimizing source risks. However, an inherent problem is that they ignore the informative knowledge implied in domain-specific features, which is recognized as implicit domain-invariant feature. This is mainly caused by the multimode structure underlying target distribution, characterized by various scales and categories of objects in target images. To solve that, we propose the Implicit Domain-invariant Faster R-CNN (IDF) by using non-adversarial domain discriminator, dual attention mechanism and selective feature perception. This idea is implemented on the Faster R-CNN backbone, but with an improved architecture of two branches, i.e. domain-invariant branch and domain-specific branch. The former can clearly learn explicit domain adaptive features w.r.t. easy samples, while the latter aims to learn implicit domain-invariant features w.r.t. hard samples. Experiments on numerous benchmark datasets, including the Cityscapes, Foggy Cityscapes, KITTI and SIM10K, show the superiority of our IDF over other state-of-the-art domain adaptive object detectors. The demo code is released inhttps://github.com/sea123321/IDF.
Qinghai Lang, Lei Zhang 0038, Wenxu Shi, Weijie Chen 0006, Shiliang Pu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Neural Image Parts Group Search for Person Re-Identification
abstract
Employing partition strategy to explore fine-grained features has been verified to be beneficial for person re-identification in recent literature. However, existing methods primarily rely on expert experience to manually design various partition strategies, which may lead to a sub-optimal solution for fine-grained features exploration. In this paper, we propose a Neural Parts Group Search (NPGS) strategy that auto-searches the optimal parts group via evolutionary algorithm (EA) to facilitate the network to exploit the local details. And during search process, designing a high-quality search space is especially crucial for an efficient optimization. Considering the human top-down structure and the semantic coherence of parts, we design a coarse-to-fine parts search space (C2F-PSP) in NPGS, which effectively reduce the search complexity without the loss of parts expressivity. Additionally, since only employing the high-level semantic features is insufficient for the NPGS to search effective parts, we further develop an efficient feature aggregation strategy named hierarchical low-rank bilinear pooling that progressively integrates the high-level semantic property and the low-level fine-grained details to facilitate the NPGS to explore the fine-grained features. Furthermore, to relieve the interference of background during parts search process, we propose a novel Relational Attention Module (RAM) by exploiting the channel and spatial structural interdependence of pixels to strengthen the discriminative regions. Extensive experiments on the mainstream evaluation datasets demonstrate that our method outperforms the recent state-of-the-art Re-ID models.
Zhipu Liu, Lei Zhang 0038, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Stochastic Gradient Perturbation: An Implicit Regularizer for Person Re-Identification
abstract
Generalization of the person re-identification (ReID) model plays an important role in practical application, and we discuss a simple yet effective regularizer to improve it inspired by Adversarial Training (AT). AT has been indicated as an advanced regularizer due to its adversarial mechanism, ability to mine hard samples, and nature of data augmentation. However, serving as an augmentation-based regularizer, AT shows low diversity of the perturbation, excessive computational cost, and the optimization dilemma between adversarial robustness and accuracy for ReID task, and is thus suboptimal. To tackle these limitations and get a more effective regularizer for ReID, we rethink the nature of AT and unveil that the adversarial data augmentation is essentially reflected by gradients. Based on this, a novel implicit regularizer, named Stochastic Gradient Perturbation (SGP), is proposed, which naturally brings three merits: 1) Better diversity of the perturbation due to the proposed non-directional stochastic perturbations rather than directional adversarial perturbations. 2) Lower computational cost due to the proposed implicit gradient augmentation rather than explicitly additional data. 3) The optimization dilemma of the adversarial robustness and generalization is naturally overcome since SGP contains the adversarial gradient perturbation. Further, we put forward a perspective that the generalization and adversarial robustness may have an inter unity. Experiments on the baseline and SOTA models demonstrate powerful performances of the plugged-played SGP, and both generalization and adversarial robustness can be guaranteed.
Fuxiang Huang, Weijie Chen 0006, Shiliang Pu, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.5
2023 Anomaly Detection of Hyperspectral Image With Hierarchical Antinoise Mutual-Incoherence- Induced Low-Rank Representation
abstract
Hyperspectral image (HSI) anomaly detection (AD) generally considers background pixels as low-rank distribution and anomaly pixels as sparse distribution. However, it is usually difficult to construct an accurate background dictionary for the background pixels composed of different land-covers, and completely separate sparse anomaly targets from various complicated background pixels with complex mixed noise interference. To address these challenges, we propose an anti-noise hierarchical mutual-incoherence-induced discriminative learning (AHMID) method for AD of HSI. A structural incoherence constraint is designed to constrain the inherent dissimilarity and incoherence between background and anomalies for improving their separability. Then, a first-order statistic constraint is conducted on targets to enhance the anomaly representation, and a decentralization constraint is used on background to suppress the background representation. Meanwhile, a mixed noise model is constructed by ℓ1,1-norm and Frobenius norm to improve the anti-noise performance. Finally, a hierarchical alternating strategy is developed to gradually optimize the background and anomalies. Experiments on six HSI AD datasets show that the proposed method outperforms a few state-of-the-art AD algorithms. Code: https://github.com/HalongL/HAD-AHMID.
Tan Guo, Fulin Luo, Xiuwen Gong, Lei Zhang 0038
IEEE Trans. Geosci. Remote. Sens.6
2023 Dual-View Spectral and Global Spatial Feature Fusion Network for Hyperspectral Image Classification
abstract
For hyperspectral image (HSI) classification, two branch networks generally use the convolution neural networks (CNNs) to extract the spatial features and the long short-term memory (LSTM) to learn the spectral features. However, CNN with a local kernel neglects the global properties of the whole HSI. LSTM doesn’t consider the macroscopic and detailed information of spectra. In this paper, we propose a dual-view spectral and global spatial feature fusion network (DSGSF) to extract the spatial-spectral features for HSI classification, including a spatial subnetwork and a spectral subnetwork. In the spatial subnetwork, we propose a global spatial feature representation model based on the encoder-decoder structure with channel attention and spatial attention to learn the global spatial features. In the spectral subnetwork, we design a dual-view spectral feature aggregation model with view attention to learn the diversity of spectral features. By fusing the two subnetworks, we construct DSGSF to extract the spatial-spectral features of HSI with strong discriminating performance. Experimental results on three public datasets illustrate that the proposed method can achieve competitive results compared with the state-of-the-art methods. Code: https://github.com/RZWang-WH/DSGSF.
Tan Guo, Fulin Luo, Xiuwen Gong, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Style Uncertainty Based Self-Paced Meta Learning for Generalizable Person Re-Identification
abstract
Domain generalizable person re-identification (DG ReID) is a challenging problem, because the trained model is often not generalizable to unseen target domains with different distribution from the source training domains. Data augmentation has been verified to be beneficial for better exploiting the source data to improve the model generalization. However, existing approaches primarily rely on pixel-level image generation that requires designing and training an extra generation network, which is extremely complex and provides limited diversity of augmented data. In this paper, we propose a simple yet effective feature based augmentation technique, named Style-uncertainty Augmentation (SuA). The main idea of SuA is to randomize the style of training data by perturbing the instance style with Gaussian noise during training process to increase the training domain diversity. And to better generalize knowledge across these augmented domains, we propose a progressive learning to learn strategy named Self-paced Meta Learning (SpML) that extends the conventional one-stage meta learning to multi-stage training process. The rationality is to gradually improve the model generalization ability to unseen target domains by simulating the mechanism of human learning. Furthermore, conventional person Re-ID loss functions are unable to leverage the valuable domain information to improve the model generalization. So we further propose a distance-graph alignment loss that aligns the feature relationship distribution among domains to facilitate the network to explore domain-invariant representations of images. Extensive experiments on four large-scale benchmarks demonstrate that our SuA-SpML achieves state-of-the-art generalization to unseen domains for person ReID.
Lei Zhang 0038, Zhipu Liu, Wensheng Zhang 0002, David Zhang 0001
IEEE Trans. Image Process.1
2023 Randomized Spectrum Transformations for Adapting Object Detector in Unseen Domains
abstract
We propose a Meta Learning on Randomized Transformations (MLRT) to learn domain invariant object detectors. Domain generalization is a problem about learning an invariant model from multiple source domains which can generalize well on unseen target domains. This problem is overlooked in object detection field, which is formally named as domain generalizable object detection (DGOD). Moreover, existing domain generalization methods have the problem of domain bias so that they can easily overfit to some specific domain (e.g., source domain). In order to alleviate the domain bias, in MLRT model, a novel randomized spectrum transformation (RST) module is proposed to increase the diversity of source domains. Specifically, RST randomizes the domain specific information of images in frequency-space, which can transform single or multiple source domains into various new domains. Besides, we observe a prior that the gradient imbalance degree among domains can also reflect the domain bias. Therefore, we further propose to alleviate the domain bias from the perspective of gradient balancing, and a novel gradient weighting (GW) module is proposed to balance the gradients over all domains via a hand-crafted weight. Finally we embed our RST and GW into a general meta learning framework and the proposed MLRT model is formalized for DGOD task. Extensive experiments are conducted on six benchmarks, and our method achieves the SOTA performance.
Lei Zhang 0038, Lingyun Qin, Mingjun Xu, Weijie Chen 0006, Shiliang Pu, Wensheng Zhang 0002
IEEE Trans. Image Process.1
2023 Adversarial and Isotropic Gradient Augmentation for Image Retrieval With Text Feedback
abstract
Image Retrieval with Text Feedback (IRTF) is an emerging research topic where the query consists of an image and a text expressing a requested attribute modification. The goal is to retrieve the target images similar to the query text modified query image. The existing methods usually adopt feature fusion of the query image and text to match the target image. However, they ignore two crucial issues: overfitting and low diversity of training data, which make the feature fusion based IRTF task not generalizable. Conventional generation based data augmentation is an effective way to alleviate overfitting and improve diversity, but increases the volume of training data and generation model parameters, which is bound to bring huge computation costs. By rethinking the conventional data augmentation mechanism, we propose a plug-and-play Gradient Augmentation (GA) based regularization approach. Specifically, GA contains two items: 1) To alleviate model overfitting on the training set, we deduce anexplicit adversarial gradient augmentationfrom the perspective of adversarial training, which challenges the“no free lunch”philosophy. 2) To improve the diversity of training set, we propose animplicit isotropic gradient augmentationfrom the perspective of gradient descent-based optimization, which achieves the goal ofbig gain but no pain. Besides, we introduce deep metric learning to train the model and provide theoretical insights of GA on generalisation. Finally, we propose a new evaluation protocol called Weighted Harmonic Mean (WHM) to assess the model generalisation. Experiments show that our GA outperforms the state-of-the-art methods by 6.2 and 4.7% on CSS and Fashion200 k datasets, respectively, without bells and whistles.
Fuxiang Huang, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Multim.2
2023 Compound Projection Learning for Bridging Seen and Unseen Objects
abstract
Zero-shot Learning (ZSL) aims to transfer knowledge from seen image categories to unseen ones by leveraging semantic information. It is generally assumed that the seen and unseen objects share a common semantic space. Most of existing ZSL methods focus on how to connect the visual space and the semantic space. However, since there are some visual distribution differences between seen and unseen objects, the projection function learned by those seen classes is biased when transferring knowledge to unseen classes. We argue that, although the unseen objects are class-agnostic, the visual distribution information of unseen samples can be generated by exploiting semantic features. In this paper, we propose aCompoundProjectionLearning (CPL) model to transfer knowledge from seen to unseen objects by exploiting the information of both seen and class-agnostic samples. With the projected semantic representation by CPL, effective constraints such as projection loss and semantic reconstruction loss can be explored for seen and unseen objects, respectively, such that the semantic ambiguity across seen and unseen objects is reduced. Additionally, we utilize a similarity network to further explore the inter-class relationship by employing the labels and the similarities between seen and unseen classes. Extensive experiments on ZSL benchmark datasets show the effectiveness of our proposed approach.
Wenli Song, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Multim.2
2022 Attention Diversification for Domain Generalization
Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Mingli Song, Di Xie, Shiliang Pu
ECCV (34)7
2022 Recurrent LSTM-based UAV Trajectory Prediction with ADS-B Information
abstract
Recently, unmanned aerial vehicles (UAVs) are gathering increasing attentions from both the academia and industry. The ever-growing number of UAV brings challenges for air traffic control (ATC), and thus trajectory prediction plays a vital role in ATC, especially for avoiding collisions among UAVs. However, the dynamic flight of UAV aggravates the complexity of trajectory prediction. Different with civil aviation aircrafts, the most intractable difficulty for UAV trajectory prediction depends on acquiring effective location information. Fortunately, the automatic dependent surveillance-broadcast (ADS-B) is an effective technique to help obtain positioning information. It is widely used in the civil aviation aircraft, due to its high data update frequency and low cost of corresponding ground stations construction. Hence, in this work, we consider leveraging ADS-B to help UAV trajectory prediction. However, with the ADS-B information for a UAV, it still lacks efficient mechanism to predict the UAV trajectory. It is noted that the recurrent neural network (RNN) is available for the UAV trajectory prediction, in which the long short-term memory (LSTM) is specialized in dealing with the time-series data. As above, in this work, we design a system of UAV trajectory prediction with the ADS-B information, and propose the recurrent LSTM (RLSTM) based algorithm to achieve the accurate prediction. Finally, extensive simulations are conducted by Python to evaluate the proposed algorithms, and the results show that the average trajectory prediction error is satisfied, which is in line with expectations.
Ziye Jia, Chao Dong 0001, Yuntian Liu, Lei Zhang 0038, Qihui Wu 0001
GLOBECOM5
2022 Learning Domain Adaptive Object Detection with Probabilistic Teacher
abstract
Self-training for unsupervised domain adaptive object detection is a challenging task, of which the performance depends heavily on the quality of pseudo boxes. Despite the promising results, prior works have largely overlooked the uncertainty of pseudo boxes during self-training. In this paper, we present a simple yet effective framework, termed as Probabilistic Teacher (PT), which aims to capture the uncertainty of unlabeled target data from a gradually evolving teacher and guides the learning of a student in a mutually beneficial manner. Specifically, we propose to leverage the uncertainty-guided consistency training to promote classification adaptation and localization adaptation, rather than filtering pseudo boxes via an elaborate confidence threshold. In addition, we conduct anchor adaptation in parallel with localization adaptation, since anchor can be regarded as a learnable parameter. Together with this framework, we also present a novel Entropy Focal Loss (EFL) to further facilitate the uncertainty-guided self-training. Equipped with EFL, PT outperforms all previous baselines by a large margin and achieve new state-of-the-arts.
Meilin Chen, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Yunfeng Yan, Donglian Qi, Yueting Zhuang, Di Xie, Shiliang Pu
ICML6
2022 Universal Domain Adaptive Object Detector
abstract
Universal domain adaptive object detection (UniDAOD) is more challenging than domain adaptive object detection (DAOD) since the label space of the source domain may not be the same as that of the target and the scale of objects in the universal scenarios can vary dramatically (i.e, category shift and scale shift). To this end, we propose US-DAF, namely Universal Scale-Aware Domain Adaptive Faster RCNN with Multi-Label Learning, to reduce the negative transfer effect during training while maximizing transferability as well as discriminability in both domains under a variety of scales. Specifically, our method is implemented by two modules: 1) We facilitate the feature alignment of common classes and suppress the interference of private classes by designing a Filter Mechanism module to overcome the negative transfer caused by category shift. 2) We fill the blank of scale-aware adaptation in object detection by introducing a new Multi-Label Scale-Aware Adapter to perform individual alignment between corresponding scale for two domains. Experiments show that US-DAF achieves state-of-the-art results on three scenarios (\emphi.e, Open-Set, Partial-Set, and Closed-Set) and yields 7.1% and 5.9% relative improvement on benchmark datasets Clipart1k and Watercolor in particular.
Wenxu Shi, Lei Zhang 0038, Weijie Chen 0006, Shiliang Pu
ACM Multimedia2
2022 IDEA: intelligent divine eye on air through multi-UAV collaborative inference
abstract
This demonstration shows a working prototype of IDEA, Intelligent Divine Eye on Air through multi-UAV collaborative inference, to improve the throughput and accuracy of onboard object detection. Different from most existing UAV airborne object detection systems relying single UAV to run the convolutional neural networks (CNN)-based inference independently, IDEA leverages the abundant resources of multiple UAVs in a swarm, and collaboratively executes the inference task. Specifically, IDEA divides the CNN model into multiple submodels (each consisting of several successive layers), and distributes each submodel to a UAV, where the execution sequence of the submodels is coordinated to output the final prediction. The prominent advantage of IDEA lies in that, it can not only run highly accurate complex CNN models, but also perform the object detection tasks in a pipeline manner, which thus boosts high detection accuracy and frame rate. IDEA prototype with three self-constructed real-world UAVs shows ~2.6× frame rate improvement over that with one single UAV, while achieving higher detection accuracy.
Chao Dong 0001, Yuben Qu, Feiyu Wu, Lei Zhang 0038, Qihui Wu 0001
MobiSys5
2022 Domain adaptive subspace transfer model for sensor drift compensation in biologically inspired electronic nose
Tan Guo, Xiaoheng Tan, Liu Yang 0003, Zhifang Liang, Bob Zhang 0001, Lei Zhang 0038
Expert Syst. Appl.6
2022 Refining pseudo labels for unsupervised Domain Adaptive Re-Identification
Shanshan Wang 0008, Lei Zhang 0038, Fan Wang 0019, Hao Li 0030
Knowl. Based Syst.2
2022 Comments on "Hierarchical Suppression Method for Hyperspectral Target Detection"
abstract
The hierarchical constrained energy minimization (hCEM) algorithm, published in TGRS, has received more attentions in the field of hyperspectral target detection since publication. Using the classical constrained energy minimization (CEM) detector as the basic unit, it designs a hierarchical structure to gradually suppress the background and to enhance the target detection performance. The authors claimed that the convergence of the hCEM algorithm can be theoretically guaranteed by analyzing the realtionship between different layers of the hierarchical output. However, after some investigations, we found that the key formula presented in the paper is theoretically defective. This implies that the theoretical results do not hold and the convergence of the algorithm cannot be ensured.
Lei Wang 0112, Luyan Ji, Xiurui Geng, Lei Zhang 0038
IEEE Geosci. Remote. Sens. Lett.4
2022 Evaluation Model of Click Rate of Electronic Commerce Advertising Based on Fuzzy Genetic Algorithm
Peisen Song, Lei Zhang 0038
Mob. Networks Appl.3
2022 Tasks Integrated Networks: Joint Detection and Retrieval for Image Search
abstract
The traditional object (person) retrieval (re-identification) task aims to learn a discriminative feature representation with intra-similarity and inter-dissimilarity, which supposes that the objects in an image are manually or automatically pre-cropped exactly. However, in many real-world searching scenarios (e.g., video surveillance), the objects (e.g., persons, vehicles, etc.) are seldom accurately detected or annotated. Therefore, object-level retrieval becomes intractable without bounding-box annotation, which leads to a new but challenging topic, i.e., image-level search with multi-task integration of joint detection and retrieval. In this paper, to address the image search issue, we first introduce an end-to-end Integrated Net (I-Net), which has three merits: 1) A Siamese architecture and an on-line pairing strategy for similar and dissimilar objects in the given images are designed. Benefited by the Siamese structure, I-Net learns the shared feature representation, because, on which, both object detection and classification tasks are handled. 2) A novelon-linepairing (OLP) loss is introduced with a dynamic feature dictionary, which alleviates the multi-task training stagnation problem, by automatically generating a number of negative pairs to restrict the positives. 3) Ahardexamplepriority (HEP) based softmax loss is proposed to improve the robustness of classification task by selecting hard categories. The shared feature representation of I-Net may restrict the task-specific flexibility and learning capability between detection and retrieval tasks. Therefore, with the philosophy ofdivide andconquer, we further propose an improved I-Net, called DC-I-Net, which makes two new contributions: 1) two modules are tailored to handle different tasks separately in the integrated framework, such that the task specification is guaranteed. 2) A class-center guided HEP loss (C$^2$HEP) by exploiting the stored class centers is proposed, such that the intra-similarity and inter-dissimilarity can be captured for ultimate retrieval. Extensive experiments on famous image-level search oriented benchmark datasets, such as CUHK-SYSU dataset and PRW dataset for person search and the large-scale WebTattoo dataset for tattoo search, demonstrate that the proposed DC-I-Net outperforms the state-of-the-art tasks-integrated and tasks-separated image search models.
Lei Zhang 0038, Zhenwei He, Yi Yang 0001, Liang Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Simultaneous Face Completion and Frontalization via Mask Guided Two-Stage GAN
abstract
Pose variation and occlusion are two key factors that affect the accuracy of face recognition. Most of the previous work alleviate the impacts of pose and occlusion by performing the tasks of face frontalization and face completion, respectively. Specially, generative adversarial networks (GANs) based methods have made a great progress on both of these two tasks. However, the two tasks are rarely paid attention simultaneously. Hence, the synthesis and recognition from the profile but occluded facial image is still an understudied and challenging problem in two aspects. 1) Occlusion mask, as a kind of noise, can be some very important prior information in the corrupted image. In particular, the occlusion mask is often used to fit this noise and help face restoration of the occluded region. However, a prior work, such as BoostGAN, failed to utilize the mask guided noise prior information. 2) The two tasks, de-occlusion and frontalization, are collaborative, so the identity discriminative information is easily lost if the two tasks are not organically unified in the training phase. In order to overcome these challenges, we propose a novel mask guided two-stage generative adversarial network (TSGAN). There are two major contributions in this work: 1) In order to utilize the prior information of noise, an occluded mask based attention model is proposed, which is integrated into the two stages via U-connection. This mask works as a guidance to simultaneously repair and frontalize the profile but occluded faces. 2) An end-to-end paradigm with a two-stage architecture (i.e.,face deocclusion network and face frontalization network) is proposed to complete the two tasks jointly. Furthermore, for preserving the discriminative identity information in both stages more effectively, we propose a novel dual triplet loss, consisting of a deocclusion triplet loss and a frontalization triplet loss. Qualitative and quantitative experiments on both constrained and unconstrained face datasets with regular and irregular (natural) occlusions demonstrate the superiority of our approach.
Qingyan Duan, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Cross-Modal Cross-Domain Dual Alignment Network for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared cross-modal person re-identification (Re-ID) has drawn increasing attention due to its application value in practice. Most of the current works rely on a supervised training manner. However, in real-world applications, manual collection of pair-wise RGB-Infrared (IR) person data is labor-intensive and time-consuming. Moreover, when a trained model is directly used in another domain, there is usually a significant performance drop. To overcome the above problems, we make the first attempt to transfer the learned model to a new RGB-IR domain which is unlabeled. The practical problem covers two kinds of challenges, i.e., cross-modal (RGB-Infrared) and cross-domain (different dataset) person Re-ID. Previous works have often considered only one of them either cross-modal or cross-domain. In this work, we propose a dual alignment network (DAN) to solve the RGB-Infrared cross-modal cross-domain person Re-ID problem. This network consists of three parts: Domain Adversarial Alignment component (DAA), Pseudo Label Generation module for target domain (PLG), and Cross-Modal Alignment component (CMA). These three modules complement and promote the model to learn domain-invariant and modality-invariant person representations. Further, we propose a protocol of cross-modal cross-domain person Re-ID by synthesizing target domains by adding random noise, adjusting the lighting intensity, and changing the background color, respectively. Experiments on real and synthetic datasets under the same cross-modalities across domains demonstrate the effectiveness of our method.
Xiaowei Fu, Fuxiang Huang, Huimin Ma 0001, Xin Xu 0001, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.6
2022 Partial Alignment for Object Detection in the Wild
abstract
Conventional object detectors often encounter remarkable performance drops due to the domain shift caused by environmental changes. However, labeling sufficient training data drawn from various domains is cost-ineffective and labor-intensive. To this end, unsupervised domain adaptive object detection (DAOD) has attracted much attention, in which the detector is transferred from the label-rich source domain to a label-agnostic target domain. Most of the existing cross-domain object detectors are designed with a parameter-shared network architecture, which, however, has an inherent flaw to accumulate the source errors caused by the inaccurate target distribution. As a result, the target data instead of decisive source data may dominate the learning towards model collapse. To overcome the risk, we propose an Asymmetric Tri-way Faster-RCNN (ATF) for DAOD tasks, in which a novel ancillary net is deployed with the chief net and it brings two merits. 1) The ancillary net enables the source distribution to be preserved without accumulating source error/distortion, such that the target distribution is dominated by source data and the model collapse is alleviated. 2) Ancillary target features can be generated by the ancillary net, which further enhances the discrimination of the target domain object detector. Furthermore, in order to remove the useless source-specific knowledge and exploit the informative domain-invariant knowledge, we propose a Partial Alignment based ATF (PA-ATF) model, in which an adversarial shuffling based mutual-information minimization strategy for domain disentanglement is provided. Intuitively, transferring only the source-domain invariant information to the target domain is more conscionable. Extensive experiments on benchmark datasets, including the Cityscapes, Foggy Cityscapes, Pascal VOC, Clipart, Watercolor, SIM10K, and KITTI demonstrate the remarkable performance of our models over other state-of-the-art approaches.
Zhenwei He, Lei Zhang 0038, Yi Yang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Efficient Context-Guided Stacked Refinement Network for RGB-T Salient Object Detection
abstract
RGB-T salient object detection (SOD) aims at utilizing the complementary cues of RGB and Thermal (T) modalities to detect and segment the common objects. However, on one hand, existing methods simply fuse the features of two modalities without fully considering the characters of RGB and T. On the other hand, the high computational cost of existing methods prevents them from real-world applications (e.g., automatic driving, abnormal detection, person re-ID). To this end, we proposed an efficient encoder-decoder network named Context-guided Stacked Refinement Network (CSRNet). Specifically, we utilize a lightweight backbone and design efficient decoder parts, which greatly reduce the computational cost. To fuse RGB and T modalities, we proposed an efficient Context-guided Cross Modality Fusion (CCMF) module to filter the noise and explore the complementation of two modalities. Besides, Stacked Refinement Network (SRN) progressively refines the features from top to down via the interaction of semantic and spatial information. Extensive experiments show that our method performs favorably against state-of-the-art algorithms on RGB-T SOD task while with small model size (4.6M), few FLOPs (4.2G), and real-time speed (38fps). Our codes is available at:https://github.com/huofushuo/CSRNet.
Fushuo Huo, Xuegui Zhu, Lei Zhang 0038, Yu Shu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Domain Adaptation Preconceived Hashing for Unconstrained Visual Retrieval
abstract
Learning to hash has been widely applied for image retrieval due to the low storage and high retrieval efficiency. Existing hashing methods assume that the distributions of the retrieval pool (i.e., the data sets being retrieved) and the query data are similar, which, however, cannot truly reflect the real-world condition due to the unconstrained visual cues, such as illumination, pose, background, and so on. Due to the large distribution gap between the retrieval pool and the query set, the performances of traditional hashing methods are seriously degraded. Therefore, we first propose a new efficient but transferable hashing model for unconstrained cross-domain visual retrieval, in which the retrieval pool and the query sample are drawn from different but semantic relevant domains. Specifically, we propose a simple yet effective unsupervised hashing method, domain adaptation preconceived hashing (DAPH), toward learning domain-invariant hashing representation. Three merits of DAPH are observed: 1) to the best of our knowledge, we first propose unconstrained visual retrieval by introducing DA into hashing for learning transferable hashing codes; 2) a domain-invariant feature transformation with marginal discrepancy distance minimization and feature reconstruction constraint is learned, such that the hashing code is not only domain adaptive but content preserved; and 3) a DA preconceived quantization loss is proposed, which further guarantees the discrimination of the learned hashing code for sample retrieval. Extensive experiments on various benchmark data sets verify that our DAPH outperforms many state-of-the-art hashing methods toward unconstrained (unrestricted) instance retrieval in both single- and cross-domain scenarios.
Fuxiang Huang, Lei Zhang 0038, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Dynamic Weighted Learning for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to improve the classification performance on an unlabeled target domain by leveraging information from a fully labeled source domain. Recent approaches explore domain-invariant and class-discriminant representations to tackle this task. These methods, however, ignore the interaction between domain alignment learning and class discrimination learning. As a result, the missing or inadequate tradeoff between domain alignment and class discrimination are prone to the problem of negative transfer. In this paper, we propose Dynamic Weighted Learning (DWL) to avoid the discriminability vanishing problem caused by excessive alignment learning and domain misalignment problem caused by excessive discriminant learning. Technically, DWL dynamically weights the learning losses of alignment and discriminability by introducing the degree of alignment and discriminability. Besides, the problem of sample imbalance across domains is first considered in our work, and we solve the problem by weighing the samples to guarantee information balance across domains. Extensive experiments demonstrate that DWL has an excellent performance in several benchmark datasets. Code is available at https://github.com/NiXiao-cqu/TransferLearning-dwl-cvpr2021.
Ni Xiao, Lei Zhang 0038
CVPR2
2021 Reconcile Prediction Consistency for Balanced Object Detection
abstract
Classification and regression are two pillars of object detectors. In most CNN-based detectors, these two pillars are optimized independently. Without direct interactions be-tween them, the classification loss and the regression loss can not be optimized synchronously toward the optimal direction in the training phase. This clearly leads to lots of inconsistent predictions with high classification score but low localization accuracy or low classification score but high localization accuracy in the inference phase, especially for the objects of irregular shape and occlusion, which severely hurts the detection performance of existing detectors after N-MS. To reconcile prediction consistency for balanced object detection, we propose a Harmonic loss to harmonize the optimization of classification branch and localization branch. The Harmonic loss enables these two branches to supervise and promote each other during training, thereby producing consistent predictions with high co-occurrence of top classification and localization in the inference phase. Furthermore, in order to prevent the localization loss from being dominated by outliers during training phase, a Harmonic IoU loss is proposed to harmonize the weight of the localization loss of different IoU-level samples. Comprehensive experiments on benchmarks PASCAL VOC and MS COCO demonstrate the generality and effectiveness of our model for facilitating existing object detectors to state-of-the-art accuracy.
Keyang Wang, Lei Zhang 0038
ICCV2
2021 Layer-Wise Customized Weak Segmentation Block And Aiou Loss For Accurate Object Detection
abstract
The anchor-based detectors handle the problem of scale variation by building the feature pyramid and directly setting different scales of anchors on each cell in different layers. However, it is difficult for box-wise anchors to guide the adaptive learning of scale-specific features in each layer because there is no one-to-one correspondence between box-wise anchors and pixel-level features. In order to alleviate the problem, in this paper, we propose a scale-customized weak segmentation (SCWS) block at the pixel level for scale customized object feature learning in each layer. By integrating the SCWS blocks into the single-shot detector, a scale-aware object detector (SCOD) is constructed to detect objects of different sizes naturally and accurately. Furthermore, the standard location loss neglects the fact that the hard and easy samples may be seriously imbalanced. A forthcoming problem is that it is unable to get more accurate bounding boxes due to the imbalance. To address this problem, an adaptive IoU (AIoU) loss via a simple yet effective squeeze operation is specified in our SCOD. Extensive experiments on PASCAL VOC and MS COCO demonstrate the superiority of our SCOD.
Keyang Wang, Lei Zhang 0038, Wenli Song, Qinghai Lang, Lingyun Qin
ICIP2
2021 Label Disentangled Analysis for unsupervised visual domain adaptation
Ni Xiao, Lei Zhang 0038, Xin Xu 0001, Tan Guo, Huimin Ma 0001
Knowl. Based Syst.2
2021 Adversarial View Confusion Feature Learning for Person Re-Identification
abstract
The performances of person re-identification tasks can be seriously degraded because of variations caused by view changes. In recent years, there are many methods focusing on how to solve cross view challenges which can be roughly divided into two categories: 1) learning view-invariant features without the help of view information. 2) combining view-wise features with the guide of view information. However, these methods are neither perfect enough. Methods of the first category are not roust enough for different kinds of view-invariants while methods of the other category can not generalize well in real-world applications. In this paper, we aim to learn view-invariant features with the help of view information. We proposed an end-to-end trainable framework, called View Confusion Feature Learning (VCFL), to learn view-invariant features by getting rid of view specific information. To the best of our knowledge, VCFL is originally proposed to learn view-invariant identity-wise features, and it is a kind of combination of view-generic and view-specific methods. The whole view confusion learning mechanism consists of three parts: 1) adversarial learning between feature extractor and the view classifier; 2) drawing the features with the same ID close to centers; 3) the guidance of SIFT, for seamlessly integration of hand-crafted features and deep features. In order to make the whole confusion mechanism work better, we further propose a VCFL+ model, which improves the fusion process in the feature map level through the thoughts of attention mechanism. Experiments on three benchmark datasets including Market1501, CUHK03, and DukeMTMC prove the superiority of our method over state-of-the-art approaches.
Lei Zhang 0038, Fangyi Liu, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 AdvKin: Adversarial Convolutional Network for Kinship Verification
abstract
Kinship verification in the wild is an interesting and challenging problem. The goal of kinship verification is to determine whether a pair of faces are blood relatives or not. Most previous methods for kinship verification can be divided as handcrafted features-based shallow learning methods and convolutional neural network (CNN)-based deep-learning methods. Nevertheless, these methods are still facing the challenging task of recognizing kinship cues from facial images. The reason is that the family ID information and the distribution difference of pairwise kin-faces are rarely considered in kinship verification tasks. To this end, a family ID-based adversarial convolutional network (AdvKin) method focused on discriminative Kin features is proposed for both small-scale and large-scale kinship verification in this article. The merits of this article are four-fold: 1) for kin-relation discovery, a simple yet effective self-adversarial mechanism based on a negative maximum mean discrepancy (NMMD) loss is formulated as attacks in the first fully connected layer; 2) a pairwise contrastive loss and family ID-based softmax loss are jointly formulated in the second and third fully connected layer, respectively, for supervised training; 3) a two-stream network architecture with residual connections is proposed in AdvKin; and 4) for more fine-grained deep kin-feature augmentation, an ensemble of patch-wise AdvKin networks is proposed (E-AdvKin). Extensive experiments on 4 small-scale benchmark KinFace datasets and 1 large-scale families in the wild (FIW) dataset from the first Large-Scale Kinship Recognition Data Challenge, show the superiority of our proposed AdvKin model over other state-of-the-art approaches.
Lei Zhang 0038, Qingyan Duan, David Zhang 0001, Wei Jia 0001, Xizhao Wang
IEEE Trans. Cybern.1
2021 Look More Into Occlusion: Realistic Face Frontalization and Recognition With BoostGAN
abstract
Many factors can affect face recognition, such as occlusion, pose, aging, and illumination. First and foremost are occlusion and large-pose problems, which may even lead to more than 10% accuracy degradation. Recently, generative adversarial net (GAN) and its variants have been proved to be effective in processing pose and occlusion. For the former, pose-invariant feature representation and face frontalization based on GAN models have been studied to solve the pose variation problem. For the latter, frontal face completion on occlusions based on GAN models have also been presented, which is much concerned with facial structure and realistic pixel details rather than identity preservation. However, synthesizing and recognizing the occluded but profile faces is still an understudied problem. Therefore, in this article, to address this problem, we contribute an efficient but effective solution on how to synthesize and recognize faces with large-pose variations and simultaneously corrupted regions (e.g., nose and eyes). Specifically, we propose a boosting GAN (BoostGAN) for occluded but profile face frontalization, deocclusion, and recognition, which has two aspects: 1) with the assumption that face occlusion is incomplete and partial, multiple images with patch occlusion are fed into our model for knowledge boosting, i.e., identity and texture information and 2) a new aggregation structure integrated with a deep encoder-decoder network for coarse face synthesis and a boosting network for fine face generation is carefully designed. Exhaustive experiments on benchmark data sets with regular and irregular occlusions demonstrate that the proposed model not only shows clear photorealistic images but also presents powerful recognition performance over state-of-the-art GAN models for occlusive but profile face recognition in both the controlled and uncontrolled environments. To the best of our knowledge, this article proposes to solve face synthesis and recognition under poses and occlusions for the first time.
Qingyan Duan, Lei Zhang 0038
IEEE Trans. Neural Networks Learn. Syst.2
2020 Probability Weighted Compact Feature for Domain Adaptive Retrieval
abstract
Domain adaptive image retrieval includes single-domain retrieval and cross-domain retrieval. Most of the existing image retrieval methods only focus on single-domain retrieval, which assumes that the distributions of retrieval databases and queries are similar. However, in practical application, the discrepancies between retrieval databases often taken in ideal illumination/pose/background/camera conditions and queries usually obtained in uncontrolled conditions are very large. In this paper, considering the practical application, we focus on challenging cross-domain retrieval. To address the problem, we propose an effective method named Probability Weighted Compact Feature Learning (PWCF), which provides inter-domain correlation guidance to promote cross-domain retrieval accuracy and learns a series of compact binary codes to improve the retrieval speed. First, we derive our loss function through the Maximum A Posteriori Estimation (MAP): Bayesian Perspective (BP) induced focal-triplet loss, BP induced quantization loss and BP induced classification loss. Second, we propose a common manifold structure between domains to explore the potential correlation across domains. Considering the original feature representation is biased due to the inter-domain discrepancy, the manifold structure is difficult to be constructed. Therefore, we propose a new feature named Histogram Feature of Neighbors (HFON) from the sample statistics perspective. Extensive experiments on various benchmark databases validate that our method outperforms many state-of-the-art image retrieval methods for domain adaptive image retrieval. The source code is available at {https://github.com/fuxianghuang1/PWCF}.
Fuxiang Huang, Lei Zhang 0038, Yang Yang 0002, Xichuan Zhou
CVPR2
2020 Domain Adaptive Object Detection via Asymmetric Tri-Way Faster-RCNN
Zhenwei He, Lei Zhang 0038
ECCV (24)2
2020 Domain Private and Agnostic Feature for Modality Adaptive Face Recognition
abstract
Heterogeneous face recognition is a challenging task due to the large modality discrepancy and insufficient cross-modal samples. Most existing works focus on discriminative feature transformation, metric learning and cross-modal face synthesis. However, the fact that cross-modal faces are always coupled by domain (modality) and identity information has received little attention. Therefore, how to learn and utilize the domain-private feature and domain-agnostic feature for modality adaptive face recognition is the focus of this work. Specifically, this paper proposes a Feature Aggregation Network (FAN), which includes disentangled representation module (DRM), feature fusion module (FFM) and adaptive penalty metric (APM) learning session. First, in DRM, two subnetworks, i.e. domain-private network and domain-agnostic network are specially designed for learning modality features and identity features, respectively. Second, in FFM, the identity features are fused with domain features to achieve cross-modal bidirectional identity feature transformation, which, to a large extent, further disentangles the modality information and identity information. Third, considering that the distribution imbalance between easy and hard pairs exists in cross-modal datasets, which increases the risk of model bias, the identity preserving guided metric learning with adaptive hard pairs penalization is proposed in our FAN. The proposed APM also guarantees the cross-modality intra-class compactness and inter-class separation. Extensive experiments on benchmark cross-modal face datasets show that our FAN outperforms SOTA methods.
Yingguo Xu, Lei Zhang 0038, Qingyan Duan
IJCB2
2020 Self-adaptive Re-weighted Adversarial Domain Adaptation
abstract
Existing adversarial domain adaptation methods mainly consider the marginal distribution and these methods may lead to either under transfer or negative transfer. To address this problem, we present a self-adaptive re-weighted adversarial domain adaptation approach, which tries to enhance domain alignment from the perspective of conditional distribution. In order to promote positive transfer and combat negative transfer, we reduce the weight of the adversarial loss for aligned features while increasing the adversarial force for those poorly aligned measured by the conditional entropy. Additionally, triplet loss leveraging source samples and pseudo-labeled target samples is employed on the confusing domain. Such metric loss ensures the distance of the intra-class sample pairs closer than the inter-class pairs to achieve the class-level alignment. In this way, the high accurate pseudolabeled target samples and semantic alignment can be captured simultaneously in the co-training process. Our method achieved low joint error of the ideal source and target hypothesis. The expected target error can then be upper bounded following Ben-David’s theorem. Empirical evidence demonstrates that the proposed model outperforms state of the arts on standard domain adaptation datasets.
Shanshan Wang 0008, Lei Zhang 0038
IJCAI2
2020 Hierarchical Bi-Directional Feature Perception Network for Person Re-Identification
abstract
Previous Person Re-Identification (Re-ID) models aim to focus on the most discriminative region of an image, while its performance may be compromised when that region is missing caused by camera viewpoint changes or occlusion. To solve this issue, we propose a novel model named Hierarchical Bi-directional Feature Perception Network (HBFP-Net) to correlate multi-level information and reinforce each other. First, the correlation maps of cross-level feature-pairs are modeled via low-rank bilinear pooling. Then, based on the correlation maps, Bi-directional Feature Perception (BFP) module is employed to enrich the attention regions of high-level feature, and to learn abstract and specific information in low-level feature. And then, we propose a novel end-to-end hierarchical network which integrates multi-level augmented features and inputs the augmented low- and middle-level features to following layers to retrain a new powerful network. What's more, we propose a novel trainable generalized pooling, which can dynamically select any value of all locations in feature maps to be activated. Extensive experiments implemented on the mainstream evaluation datasets including Market-1501, CUHK03 and DukeMTMC-ReID show that our method outperforms the recent SOTA Re-ID models.
Zhipu Liu, Lei Zhang 0038, Yang Yang 0002
ACM Multimedia2
2020 Text-Embedded Bilinear Model for Fine-Grained Visual Recognition
abstract
Fine-grained visual recognition, which aims to identify subcategories of the same base-level category, is a challenging task because of its large intra-class variances and small inter-class variances. Human beings can perform object recognition task based on not only the visual appearance but also the knowledge from texts, as texts can point out the discriminative parts or characteristics which are always the key to distinguishing different subcategories. This is an involuntary transfer from human textual attention to visual attention, suggesting that texts are able to assist fine-grained recognition. In this paper, we propose a Text-Embedded Bilinear (TEB) model which incorporates texts as extra guidance for fine-grained recognition. Specially, we first conduct a text-embedded network to embed text feature into the discriminative image feature learning to get a embedded feature. In addition, since the cross-layer part feature interaction and fine-grained feature learning are mutually correlated and can reinforce each other, we also extract a candidate feature from the text encoder and embed it into the inter-layer feature of the image encoder to get an embedded candidate feature. At last we utilize a cross-layer bilinear network to fuse the two embedded features. Comparing with state-of-the-art methods on the widely used CUB-200-2011 dataset and Oxford Flowers-102 dataset for fine-grained image recognition, the experimental results demonstrate our TEB model achieves the best performance.
Xiang Guan, Yang Yang 0002, Lei Zhang 0038
ACM Multimedia4
2020 Single-Shot Two-Pronged Detector with Rectified IoU Loss
abstract
In the CNN based object detectors, feature pyramids are widely exploited to alleviate the problem of scale variation across object instances. These object detectors, which strengthen features via a top-down pathway and lateral connections, are mainly to enrich the semantic information of low-level features, but ignore the enhancement of high-level features. This can lead to an imbalance between different levels of features, in particular a serious lack of detailed information in the high-level features, which makes it difficult to get accurate bounding boxes. In this paper, we introduce a novel two-pronged transductive idea to explore the relationship among different layers in both backward and forward directions, which can enrich the semantic information of low-level features and detailed information of high-level features at the same time. Under the guidance of the two-pronged idea, we propose a Two-Pronged Network (TPNet) to achieve bidirectional transfer between high-level features and low-level features, which is useful for accurately detecting object at different scales. Furthermore, due to the distribution imbalance between the hard and easy samples in single-stage detectors, the gradient of localization loss is always dominated by the hard examples that have poor localization accuracy. This will enable the model to be biased toward the hard samples. So in our TPNet, an adaptive IoU based localization loss, named Rectified IoU (RIoU) loss, is proposed to rectify the gradients of each kind of samples. The Rectified IoU loss increases the gradients of examples with high IoU while suppressing the gradients of examples with low IoU, which can improve the overall localization accuracy of model. Extensive experiments demonstrate the superiority of our TPNet and RIoU loss.
Keyang Wang, Lei Zhang 0038
ACM Multimedia2
2020 Hard Negative Samples Emphasis Tracker without Anchors
abstract
Trackers based on Siamese network have shown tremendous success, because of their balance between accuracy and speed. Nevertheless, with tracking scenarios becoming more and more sophisticated, most existing Siamese-based approaches ignore the addressing of the problem that distinguishes the tracking target from hard negative samples in the tracking phase. The features learned by these networks lack of discrimination, which significantly weakens the robustness of Siamese-based trackers and leads to suboptimal performance. To address this issue, we propose a simple yet efficient hard negative samples emphasis method, which constrains Siamese network to learn features that are aware of hard negative samples and enhance the discrimination of embedding features. Through a distance constraint, we force to shorten the distance between exemplar vector and positive vectors, meanwhile, enlarge the distance between exemplar vector and hard negative vectors. Furthermore, we explore a novel anchor-free tracking framework in a per-pixel prediction fashion, which can significantly reduce the number of hyper-parameters and simplify the tracking process by taking full advantage of the representation of convolutional neural network. Extensive experiments on six standard benchmark datasets demonstrate that the proposed method can perform favorable results against state-of-the-art approaches.
Zhongzhou Zhang, Lei Zhang 0038
ACM Multimedia2
2020 Adversarial transfer learning for cross-domain visual recognition
Shanshan Wang 0008, Lei Zhang 0038, Jingru Fu
Knowl. Based Syst.2
2020 Target Detection in Hyperspectral Imagery via Sparse and Dense Hybrid Representation
abstract
Representation-based target detectors for hyperspectral imagery (HSI) have recently aroused a lot of interests. However, existing methods ignore the dictionary structure and cannot guarantee an informative and discriminative representation of test pixels for target detection. To alleviate the problem, this letter proposes a novel sparse and dense hybrid representation-based target detector (SDRD). The proposed detector adopts the idea that the relationship between the background and the target sub-dictionaries is a collaborative competition. The structure of the dictionary is discovered and preserved by learning a sparse and dense hybrid representation for test pixel. Benefitting from this, a compact and discriminative representation can be obtained to better represent the test pixel for an improved detection performance. Experimental results on several HSI data sets verify the effectiveness of SDRD in comparison with several state-of-the-art methods.
Tan Guo, Fulin Luo, Lei Zhang 0038, Xiaoheng Tan, Juhua Liu, Xiaocheng Zhou
IEEE Geosci. Remote. Sens. Lett.3
2020 Optimal Projection Guided Transfer Hashing for Image Retrieval
abstract
Recently, learning to hash has been widely studied for image retrieval thanks to the computation and storage efficiency of binary codes. Most existing learning to hash methods have yielded significant performance. However, for most existing learning to hash methods, sufficient training images are required and used to learn precise hashing codes. In some real-world applications, there are not always sufficient training images in the domain of interest. In addition, some existing supervised approaches need a amount of labeled data, which is an expensive process in terms of time, labor and human expertise. To handle such problems, inspired by transfer learning, we propose a simple yet effective unsupervised hashing method named Optimal Projection Guided Transfer Hashing (GTH) where we borrow the images of other different but related domain i.e., source domain to help learn precise hashing codes for the domain of interest i.e., target domain. In GTH, we aim to learn domain-invariant hashing functions. To achieve that, we propose to minimize the error matrix between two hashing projections of target and source domains. We seek for the maximum likelihood estimation (MLE) solution of the error matrix between the two hashing projections due to the domain gap. Furthermore, an alternating optimization method is adopted to obtain the two projections of target and source domains. By doing so, two projections can be progressively aligned. Extensive experiments on various benchmark databases for cross-domain visual recognition verify that our method outperforms many state-of-the-art learning to hash methods. The source code is available at https://github.com/liuji93/GTH.
Lei Zhang 0038, Ji Liu 0002, Yang Yang 0002, Fuxiang Huang, Feiping Nie 0001, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Class-Specific Reconstruction Transfer Learning for Visual Recognition Across Domains
abstract
Subspace learning and reconstruction have been widely explored in recent transfer learning work. Generally, a specially designed projection and reconstruction transfer functions bridging multiple domains for heterogeneous knowledge sharing are wanted. However, we argue that the existing subspace reconstruction based domain adaptation algorithms neglect the class prior, such that the learned transfer function is biased, especially when data scarcity of some class is encountered. Different from those previous methods, in this paper, we propose a novel class-wise reconstruction-based adaptation method called Class-specific Reconstruction Transfer Learning (CRTL), which optimizes a well modeled transfer loss function by fully exploiting intra-class dependency and inter-class independency. The merits of the CRTL are three-fold. 1) Using a class-specific reconstruction matrix to align the source domain with the target domain fully exploits the class prior in modeling the domain distribution consistency, which benefits the cross-domain classification. 2) Furthermore, to keep the intrinsic relationship between data and labels after feature augmentation, a projected Hilbert-Schmidt Independence Criterion (pHSIC), that measures the dependency between data and label, is first proposed in transfer learning community by mapping the data from raw space to RKHS. 3) In addition, by imposing low-rank and sparse constraints on the class-specific reconstruction coefficient matrix, the global and local data structure that contributes to domain correlation can be effectively preserved. Extensive experiments on challenging benchmark datasets demonstrate the superiority of the proposed method over state-of-the-art representation-based domain adaptation methods. The demo code is available in https://github.com/wangshanshanCQU/CRTL.
Shanshan Wang 0008, Lei Zhang 0038, Wangmeng Zuo, Bob Zhang 0001
IEEE Trans. Image Process.2
2020 Deep-Like Hashing-in-Hash for Visual Retrieval: An Embarrassingly Simple Method
abstract
Existing hashing methods have yielded significant performance in image and multimedia retrieval, which can be categorized into two groups: shallow hashing and deep hashing. However, there still exist some intrinsic limitations among them. The former generally adopts a one-step strategy to learn the hashing codes for discovering the discriminative binary feature, but the latent discriminative information in the learned hashing codes is not well exploited. The latter, as deep neural network based hashing models, can learn highly discriminative and compact features, but relies on large-scale data and computation resources for numerous network parameters tuning with back-propagation optimization. Straightforward training of deep hashing models from scratch on small-scale data is almost impossible. Therefore, in order to develop efficient but effective learning to hash algorithm that depends only on small-scale data, we propose a novel non-neural network based deep-like learning framework, i.e. multi-level cascaded hashing (MCH) approach with hierarchical learning strategy, for image retrieval. The contributions are threefold. First, a hashing-in-hash architecture is designed in MCH, which inherits the excellent traits of traditional neural networks based deep learning, such that discriminative binary features that are beneficial to image retrieval can be effectively captured. Second, in each level the binary features of all preceding levels and the visual appearance feature are simultaneously cascaded as inputs of all subsequent levels to retrain, which fully exploits the implicated discriminative information. Third, a basic learning to hash (BLH) model with label constraint is proposed for hierarchical learning. Without loss of generality, the existing hashing models can be easily integrated into our MCH framework. We show experimentally on small- and large-scale visual retrieval tasks that our method outperforms several state-of-the-arts.
Lei Zhang 0038, Ji Liu 0002, Fuxiang Huang, Yang Yang 0002, David Zhang 0001
IEEE Trans. Image Process.1
2020 Deep Cascade Model-Based Face Recognition: When Deep-Layered Learning Meets Small Data
abstract
Sparse representation based classification (SRC), nuclear-norm matrix regression (NMR), and deep learning (DL) have achieved a great success in face recognition (FR). However, there still exist some intrinsic limitations among them. SRC and NMR based coding methods belong to one-step model, such that the latent discriminative information of the coding error vector cannot be fully exploited. DL, as a multi-step model, can learn powerful representation, but relies on large-scale data and computation resources for numerous parameters training with complicated back-propagation. Straightforward training of deep neural networks from scratch on small-scale data is almost infeasible. Therefore, in order to develop efficient algorithms that are specifically adapted for small-scale data, we propose to derive the deep models of SRC and NMR. Specifically, in this paper, we propose an end-to-end deep cascade model (DCM) based on SRC and NMR with hierarchical learning, nonlinear transformation and multi-layer structure for corrupted face recognition. The contributions include four aspects. First, an end-to-end deep cascade model for small-scale data without back-propagation is proposed. Second, a multi-level pyramid structure is integrated for local feature representation. Third, for introducing nonlinear transformation in layer-wise learning, softmax vector coding of the errors with class discrimination is proposed. Fourth, the existing representation methods can be easily integrated into our DCM framework. Experiments on a number of small-scale benchmark FR datasets demonstrate the superiority of the proposed model over state-of-the-art counterparts. Additionally, a perspective that deep-layered learning does not have to be convolutional neural network with back-propagation optimization is consolidated. The demo code is available in https://github.com/liuji93/DCM.
Lei Zhang 0038, Ji Liu 0002, Bob Zhang 0001, David Zhang 0001, Ce Zhu
IEEE Trans. Image Process.1
2020 Guide Subspace Learning for Unsupervised Domain Adaptation
abstract
A prevailing problem in many machine learning tasks is that the training (i.e., source domain) and test data (i.e., target domain) have different distribution [i.e., non-independent identical distribution (i.i.d.)]. Unsupervised domain adaptation (UDA) was proposed to learn the unlabeled target data by leveraging the labeled source data. In this article, we propose a guide subspace learning (GSL) method for UDA, in which an invariant, discriminative, and domain-agnostic subspace is learned by three guidance terms through a two-stage progressive training strategy. First, the subspace-guided term reduces the discrepancy between the domains by moving the source closer to the target subspace. Second, the data-guided term uses the coupled projections to map both domains to a unified subspace, where each target sample can be represented by the source samples with a low-rank coefficient matrix that can preserve the global structure of data. In this way, the data from both domains can be well interlaced and the domain-invariant features can be obtained. Third, for improving the discrimination of the subspaces, the label-guided term is constructed for prediction based on source labels and pseudo-target labels. To further improve the model tolerance to label noise, a label relaxation matrix is introduced. For the solver, a two-stage learning strategy with teacher teaches and student feedbacks mode is proposed to obtain the discriminative domain-agnostic subspace. In addition, for handling nonlinear domain shift, a nonlinear GSL (NGSL) framework is formulated with kernel embedding, such that the unified subspace is imposed with nonlinearity. Experiments on various cross-domain visual benchmark databases show that our methods outperform many state-of-the-art UDA methods. The source code is available at https://github.com/Fjr9516/GSL.
Lei Zhang 0038, Jingru Fu, Shanshan Wang 0008, David Zhang 0001, Zhao Yang Dong, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.1
2020 Exploring nonnegative and low-rank correlation for noise-resistant spectral clustering
Zheng Wang 0044, Lin Zuo, Jing Ma 0004, Jingjing Li 0001, Zhao Kang 0001, Lei Zhang 0038
World Wide Web7
2019 Optimal Projection Guided Transfer Hashing for Image Retrieval
abstract
Recently, learning to hash has been widely studied for image retrieval thanks to the computation and storage efficiency of binary codes. For most existing learning to hash methods, sufficient training images are required and used to learn precise hashing codes. However, in some real-world applications, there are not always sufficient training images in the domain of interest. In addition, some existing supervised approaches need a amount of labeled data, which is an expensive process in terms of time, labor and human expertise. To handle such problems, inspired by transfer learning, we propose a simple yet effective unsupervised hashing method named Optimal Projection Guided Transfer Hashing (GTH) where we borrow the images of other different but related domain i.e., source domain to help learn precise hashing codes for the domain of interest i.e., target domain. Besides, we propose to seek for the maximum likelihood estimation (MLE) solution of the hashing functions of target and source domains due to the domain gap. Furthermore, an alternating optimization method is adopted to obtain the two projections of target and source domains such that the domain hashing disparity is reduced gradually. Extensive experiments on various benchmark databases verify that our method outperforms many state-of-the-art learning to hash methods. The implementation details are available at https://github.com/liuji93/GTH.
Ji Liu 0002, Lei Zhang 0038
AAAI2
2019 An Efficient Compressive Convolutional Network for Unified Object Detection and Image Compression
abstract
This paper addresses the challenge of designing efficient framework for real-time object detection and image compression. The proposed Compressive Convolutional Network (CCN) is basically a compressive-sensing-enabled convolutional neural network. Instead of designing different components for compressive sensing and object detection, the CCN optimizes and reuses the convolution operation for recoverable data embedding and image compression. Technically, the incoherence condition, which is the sufficient condition for recoverable data embedding, is incorporated in the first convolutional layer of the CCN model as regularization; Therefore, the CCN convolution kernels learned by training over the VOC and COCO image set can be used for data embedding and image compression. By reusing the convolution operation, no extra computational overhead is required for image compression. As a result, the CCN is 3.1 to 5.0 fold more efficient than the conventional approaches. In our experiments, the CCN achieved 78.1 mAP for object detection and 3.0 dB to 5.2 dB higher PSNR for image compression than the examined compressive sensing approaches.
Xichuan Zhou, Shujun Liu, Yingcheng Lin, Lei Zhang 0038, Cheng Zhuo
AAAI5
2019 Multi-Adversarial Faster-RCNN for Unrestricted Object Detection
abstract
Conventional object detection methods essentially suppose that the training and testing data are collected from a restricted target domain with expensive labeling cost. For alleviating the problem of domain dependency and cumbersome labeling, this paper proposes to detect objects in unrestricted environment by leveraging domain knowledge trained from an auxiliary source domain with sufficient labels. Specifically, we propose a multi-adversarial Faster-RCNN (MAF) framework for unrestricted object detection, which inherently addresses domain disparity minimization for domain adaptation in feature representation. The paper merits are in three-fold: 1) With the idea that object detectors often becomes domain incompatible when image distribution resulted domain disparity appears, we propose a hierarchical domain feature alignment module, in which multiple adversarial domain classifier submodules for layer-wise domain feature confusion are designed; 2) An information invariant scale reduction module (SRM) for hierarchical feature map resizing is proposed for promoting the training efficiency of adversarial domain adaptation; 3) In order to improve the domain adaptability, the aggregated proposal features with detection results are feed into a proposed weighted gradient reversal layer (WGRL) for characterizing hard confused domain samples. We evaluate our MAF on unrestricted tasks including Cityscapes, KITTI, Sim10k, etc. and the experiments show the state-of-the-art performance over the existing detectors.
Zhenwei He, Lei Zhang 0038
ICCV2
2019 View Confusion Feature Learning for Person Re-Identification
abstract
Person re-identification is an important task in video surveillance that aims to associate people across camera views at different locations and time. View variability is always a challenging problem seriously degrading person re-identification performance. Most of the existing methods either focus on how to learn view invariant feature or how to combine viewwise features. In this paper, we mainly focus on how to learn view-independent features by getting rid of view specific information through a view confusion learning mechanism. Specifically, we propose an end-to-end trainable framework, called View Confusion Feature Learning (VCFL), for person Re-ID across cameras. To the best of our knowledge, VCFL is originally proposed to learn view-independent identity-wise features, and it's a kind of combination of view-generic and view-specific methods. Furthermore, we extract sift-guided features by using bag-of-words model to help supervise the training of deep networks and enhance the view invariance of features. In experiments, our approach is validated on three benchmark datasets including CUHK01, CUHK03, and MARKET1501, which show the superiority of the proposed method over several state-of-the-art approaches.
Fangyi Liu, Lei Zhang 0038
ICCV2
2019 A deep manifold learning approach for spatial-spectral classification with limited labeled training samples
Xichuan Zhou, Fang Tang, Yingjun Zhao, Lei Zhang 0038, Dong Li 0007
Neurocomputing6
2019 Data induced masking representation learning for face data analysis
Tan Guo, Lei Zhang 0038, Xiaoheng Tan, Liu Yang 0003, Zhifang Liang
Knowl. Based Syst.2
2019 Learning Robust Weighted Group Sparse Graph for Discriminant Visual Analysis
Tan Guo, Xiaoheng Tan, Lei Zhang 0038, Chaochen Xie
Neural Process. Lett.3
2019 Taste Recognition in E-Tongue Using Local Discriminant Preservation Projection
abstract
Electronic tongue (E-Tongue), as a novel taste analysis tool, shows a promising perspective for taste recognition. In this paper, we constructed a voltammetric E-Tongue system and measured 13 different kinds of liquid samples, such as tea, wine, beverage, functional materials, etc. Owing to the noise of system and a variety of environmental conditions, the acquired E-Tongue data shows inseparable patterns. To this end, from the viewpoint of algorithm, we propose a local discriminant preservation projection (LDPP) model, an under-studied subspace learning algorithm, that concerns the local discrimination and neighborhood structure preservation. In contrast with other conventional subspace projection methods, LDPP has two merits. On one hand, with local discrimination it has a higher tolerance to abnormal data or outliers. On the other hand, it can project the data to a more separable space with local structure preservation. Further, support vector machine, extreme learning machine (ELM), and kernelized ELM (KELM) have been used as classifiers for taste recognition in E-Tongue. Experimental results demonstrate that the proposed E-Tongue is effective for multiple tastes recognition in both efficiency and effectiveness. Particularly, the proposed LDPP-based KELM classifier model achieves the best taste recognition performance of 98%. The developed benchmark data sets and codes will be released and downloaded in http://www.leizhang.tk/ tempcode.html.
Lei Zhang 0038, Xuehan Wang, Guang-Bin Huang, Tao Liu 0014, Xiaoheng Tan
IEEE Trans. Cybern.1
2019 Manifold Criterion Guided Transfer Learning via Intermediate Domain Generation
abstract
In many practical transfer learning scenarios, the feature distribution is different across the source and target domains (i.e., nonindependent identical distribution). Maximum mean discrepancy (MMD), as a domain discrepancy metric, has achieved promising performance in unsupervised domain adaptation (DA). We argue that the MMD-based DA methods ignore the data locality structure, which, up to some extent, would cause the negative transfer effect. The locality plays an important role in minimizing the nonlinear local domain discrepancy underlying the marginal distributions. For better exploiting the domain locality, a novel local generative discrepancy metric-based intermediate domain generation learning called Manifold Criterion guided Transfer Learning (MCTL) is proposed in this paper. The merits of the proposed MCTL are fourfold: 1) the concept of manifold criterion (MC) is first proposed as a measure validating the distribution matching across domains, and DA is achieved if the MC is satisfied; 2) the proposed MC can well guide the generation of the intermediate domain sharing similar distribution with the target domain, by minimizing the local domain discrepancy; 3) a global generative discrepancy metric is presented, such that both the global and local discrepancies can be effectively and positively reduced; and 4) a simplified version of MCTL called MCTL-S is presented under a perfect domain generation assumption for more generic learning scenario. Experiments on a number of benchmark visual transfer tasks demonstrate the superiority of the proposed MC guided generative transfer method, by comparing with the other state-of-the-art methods. The source code is available in https://github.com/wangshanshanCQU/MCTL.
Lei Zhang 0038, Shanshan Wang 0008, Guang-Bin Huang, Wangmeng Zuo, Jian Yang 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2019 Abnormal Odor Detection in Electronic Nose via Self-Expression Inspired Extreme Learning Machine
abstract
The electronic nose (E-nose), as a metal oxide semiconductor gas sensor system coupled with pattern recognition algorithms, is developed for approximating artificial olfaction functions. Ideal gas sensors should be with selectivity, reliability, and cross-sensitivity to different odors. However, a new problem is that abnormal odors (e.g., perfume, alcohol, etc.) would show strong sensor response, such that they deteriorate the usual usage of E-nose for target odor analysis. An intuitive idea is to recognize abnormal odors and remove them online. A known truth is that the kinds of abnormal odors are countless in real-world scenarios. Therefore, general pattern classification algorithms lose effect because it is expensive and unrealistic to obtain all kinds of abnormal odors data. In this paper, we propose two simple yet effective methods for abnormal odor (outlier) detection: 1) a self-expression model (SEM) with l1/l2-norm regularizer is proposed, which is trained on target odor data for coding and then a very few abnormal odor data is used as prior knowledge for threshold learning and 2) inspired by self-expression mechanism, an extreme learning machine (ELM) based self-expression (SE2LM) is proposed, which inherits the advantages of ELM in solving a single hidden layer feed-forward neural network. Experiments on several datasets by an E-nose system fabricated in our laboratory prove that the proposed SEM and SE2LM methods are significantly effective for real-time abnormal odor detection.
Lei Zhang 0038, Pingling Deng
IEEE Trans. Syst. Man Cybern. Syst.1
2018 Learning a Wavelet-Like Auto-Encoder to Accelerate Deep Neural Networks
abstract
Accelerating deep neural networks (DNNs) has been attracting increasing attention as it can benefit a wide range of applications, e.g., enabling mobile systems with limited computing resources to own powerful visual recognition ability. A practical strategy to this goal usually relies on a two-stage process: operating on the trained DNNs (e.g., approximating the convolutional filters with tensor decomposition) and fine-tuning the amended network, leading to difficulty in balancing the trade-off between acceleration and maintaining recognition performance. In this work, aiming at a general and comprehensive way for neural network acceleration, we develop a Wavelet-like Auto-Encoder (WAE) that decomposes the original input image into two low-resolution channels (sub-images) and incorporate the WAE into the classification neural networks for joint training. The two decomposed channels, in particular, are encoded to carry the low-frequency information (e.g., image profiles) and high-frequency (e.g., image details or noises), respectively, and enable reconstructing the original input image through the decoding process. Then, we feed the low-frequency channel into a standard classification network such as VGG or ResNet and employ a very lightweight network to fuse with the high-frequency channel to obtain the classification result. Compared to existing DNN acceleration solutions, our framework has the following advantages: i) it is tolerant to any existing convolutional neural networks for classification without amending their structures; ii) the WAE provides an interpretable way to preserve the main components of the input image for classification.
Tianshui Chen, Liang Lin 0004, Wangmeng Zuo, Lei Zhang 0038
AAAI5
2018 End-to-End Detection and Re-identification Integrated Net for Person Search
Zhenwei He, Lei Zhang 0038
ACCV (2)2
2018 FM-MAC: A Multi-Channel MAC Protocol for FANETs with Directional Antenna
abstract
Nowadays, Flying Ad hoc NETworks (FANETs) which consists of multiple Unmanned Aerial Vehicles (UAVs) are widely used in various military and civilian applications which need different Quality of Service (QoS) guarantees, for example, low delay for safety packet and high throughput for service packet. Meanwhile, to provide high bandwidth and spatial reuse, directional antennas are equipped on the UAVs more and more. However, due to the high mobility of UAVs, how to support different QoS with directional antenna is challenging for Media Access Control (MAC) protocol of FANETs. Recently, multi-channel MAC protocol has been proved to be effective to support different QoS. In this paper, we propose a FANETs multi-channel MAC protocol called FM-MAC, which combines the advantages of multi-channel and directional antenna to provide different QoS guarantees. Firstly, a reservation scheme based on mobile prediction is proposed to address the link interruption issue brought by high mobility of UAVs. Secondly, we propose a preemption mechanism to provide priority for service packets. Simulation results show that compared with other two representative protocols, FM-MAC not only improves the throughput of service packets, but also achieves lower delay and higher reliability for safety packets.
Guodong Wu, Chao Dong 0001, Aijing Li, Lei Zhang 0038, Qihui Wu 0001
GLOBECOM4
2018 LSTN: Latent Subspace Transfer Network for Unsupervised Domain Adaptation
Shanshan Wang 0008, Lei Zhang 0038
PRCV (2)2
2018 Pixel Saliency Based Encoding for Fine-Grained Image Classification
Lei Zhang 0038, Ji Liu 0002
PRCV (1)2
2018 A Spatial-Temporal Method to Detect Global Influenza Epidemics Using Heterogeneous Data Collected from the Internet
abstract
The 2009 influenza pandemic teaches us how fast the influenza virus could spread globally within a short period of time. To address the challenge of timely global influenza surveillance, this paper presents a spatial-temporal method that incorporates heterogeneous data collected from the Internet to detect influenza epidemics in real time. Specifically, the influenza morbidity data, the influenza-related Google query data and news data, and the international air transportation data are integrated in a multivariate hidden Markov model, which is designed to describe the intrinsic temporal-geographical correlation of influenza transmission for surveillance purpose. Respective models are built for 106 countries and regions in the world. Despite that the WHO morbidity data are not always available for most countries, the proposed method achieves 90.26 to 97.10 percent accuracy on average for real-time detection of global influenza epidemics during the period from January 2005 to December 2015. Moreover, experiment shows that, the proposed method could even predict an influenza epidemic before it occurs with 89.20 percent accuracy on average. Timely international surveillance results may help the authorities to prevent and control the influenza disease at the early stage of a global influenza pandemic.
Xichuan Zhou, Fan Yang 0021, Qin Li 0006, Fang Tang, Shengdong Hu, Zhi Lin 0002, Lei Zhang 0038
IEEE ACM Trans. Comput. Biol. Bioinform.8
2018 DANoC: An Efficient Algorithm and Hardware Codesign of Deep Neural Networks on Chip
abstract
Deep neural networks (NNs) are the state-of-the-art models for understanding the content of images and videos. However, implementing deep NNs in embedded systems is a challenging task, e.g., a typical deep belief network could exhaust gigabytes of memory and result in bandwidth and computational bottlenecks. To address this challenge, this paper presents an algorithm and hardware codesign for efficient deep neural computation. A hardware-oriented deep learning algorithm, named the deep adaptive network, is proposed to explore the sparsity of neural connections. By adaptively removing the majority of neural connections and robustly representing the reserved connections using binary integers, the proposed algorithm could save up to 99.9% memory utility and computational resources without undermining classification accuracy. An efficient sparse-mapping-memory-based hardware architecture is proposed to fully take advantage of the algorithmic optimization. Different from traditional Von Neumann architecture, the deep-adaptive network on chip (DANoC) brings communication and computation in close proximity to avoid power-hungry parameter transfers between on-board memory and on-chip computational units. Experiments over different image classification benchmarks show that the DANoC system achieves competitively high accuracy and efficiency comparing with the state-of-the-art approaches.
Xichuan Zhou, Shengli Li 0007, Fang Tang, Shengdong Hu, Zhi Lin 0002, Lei Zhang 0038
IEEE Trans. Neural Networks Learn. Syst.6
2018 Efficient Solutions for Discreteness, Drift, and Disturbance (3D) in Electronic Olfaction
abstract
In this paper, we aim at presenting the new challenges of electronic noses (E-noses) and proposing effective methods for handling the new challenging scientific issues to be solved, such as signal discreteness (reproducibility), systematical drift and nontarget disturbances. We first review the progress of E-noses in applications, systems, and algorithms during the past two decades. Recall a number of significant achievements and motivated by the current issues that hinder large-scale application pace of E-nose technology, we propose to address three key issues: 1) discreteness; 2) drift; and 3) disturbance (simplified as 3D issues), which are sensor induced and sensor specific. For each issue, a highly effective and efficient method is proposed. Specifically, for discreteness issue, a global affine transformation method is introduced for E-nose instruments batch calibration; for drift issue, an unsupervised feature adaptation model is proposed to achieve effective drift adaptation; additionally, for disturbance issue, we proposed a simple targets-to-targets self-representation classifier method for fast nontargets detection, without knowing any prior knowledge of thousands of nontarget disturbances in real world. For each method, a closed form solution can be analytically determined and the simplicity is guaranteed. Experiments demonstrate the effectiveness and efficiency of the proposed methods for addressing the proposed 3D issues in real applications of E-noses.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Facial beauty analysis based on geometric feature: Toward attractiveness assessment application
Lei Zhang 0038, David Zhang 0001, Mingming Sun 0006, Fangmei Chen
Expert Syst. Appl.1
2017 Deep object recognition across domains based on adaptive extreme learning machine
Lei Zhang 0038, Zhenwei He
Neurocomputing1
2017 Domain class consistency based transfer learning for image classification across domains
Lei Zhang 0038, Jian Yang 0003, David Zhang 0001
Inf. Sci.1
2017 Evolutionary Cost-Sensitive Extreme Learning Machine
abstract
Conventional extreme learning machines (ELMs) solve a Moore-Penrose generalized inverse of hidden layer activated matrix and analytically determine the output weights to achieve generalized performance, by assuming the same loss from different types of misclassification. The assumption may not hold in cost-sensitive recognition tasks, such as face recognition-based access control system, where misclassifying a stranger as a family member may result in more serious disaster than misclassifying a family member as a stranger. Though recent cost-sensitive learning can reduce the total loss with a given cost matrix that quantifies how severe one type of mistake against another, in many realistic cases, the cost matrix is unknown to users. Motivated by these concerns, this paper proposes an evolutionary cost-sensitive ELM, with the following merits: 1) to the best of our knowledge, it is the first proposal of ELM in evolutionary cost-sensitive classification scenario; 2) it well addresses the open issue of how to define the cost matrix in cost-sensitive learning tasks; and 3) an evolutionary backtracking search algorithm is induced for adaptive cost matrix optimization. Experiments in a variety of cost-sensitive tasks well demonstrate the effectiveness of the proposed approaches, with about 5%-10% improvements.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2016 Discriminative kernel transfer learning via l2, 1-norm minimization
abstract
In this paper, we propose a l2,1-norm based discriminative robust transfer learning (DKTL) method for domain adaptation tasks. The key idea is to simultaneously learn discriminative subspaces by using the proposed domain-class-consistency (DCC) metric, and the representation based robust transfer model between source domain and target domain via l21-norm minimization. The DCC metric includes two parts: domain-consistency used to measure the between-domain distribution discrepancy and class-consistency used to measure the within-domain class separability. The objective of transfer learning is to maximize the proposed metric, while for easily formulating this metric in model, we propose to minimize the domain-class-inconsistency, such that both domain distribution mismatch and class inseparability are well addressed. Two advantages of the proposed method are that on one hand the robust sparse coding selects a few valuable source data with noises (outliers) removed during knowledge transfer, and the proposed DCC metric can help to pursue discriminative subspaces of different domains for classification based transfer learning tasks. Extensive experiments demonstrate the superiority of the proposed method over other state-of-the-art domain adaptation methods.
Lei Zhang 0038, Sunil Kr. Jha, Tao Liu 0014, Guangshu Pei
IJCNN1
2016 Measurement and Analysis of the Broadband Radio Propagation in a High-Speed Railway Station
abstract
To ensure the stability of radio signal transmission in high-speed railway station, a detailed channel modeling for this special environment is essential and valuable. In this work, an evaluation of the performance of radio propagation in a real high speed railway station scenario is conducted. The corresponding measurements in the actual environment are carried out with the assistance of hardware testbed at 950 MHz and 2150 MHz, respectively. The propagation in both Line of Sight (LOS) and None-Line of Sight (NLOS) conditions are considered and tested. The measurement scenarios and configurations are presented in detail. Furthermore, the key parameters, such as power delay profile, reverberation time are extracted from the broadband measurement results and compared at both frequencies to achieve a rational solution for broadband communication system applying in the development of next generation high-speed railway.
Lei Zhang 0038, Jianwen Ding, Bei Zhang 0003, Cesar Briso-Rodríguez, Ke Guan
VTC Spring1
2016 Robust Visual Knowledge Transfer via Extreme Learning Machine-Based Domain Adaptation
abstract
We address the problem of visual knowledge adaptation by leveraging labeled patterns from source domain and a very limited number of labeled instances in target domain to learn a robust classifier for visual categorization. This paper proposes a new extreme learning machine based cross-domain network learning framework, that is called Extreme Learning Machine (ELM) based Domain Adaptation (EDA). It allows us to learn a category transformation and an ELM classifier with random projection by minimizing the -norm of the network output weights and the learning error simultaneously. The unlabeled target data, as useful knowledge, is also integrated as a fidelity term to guarantee the stability during cross domain learning. It minimizes the matching error between the learned classifier and a base classifier, such that many existing classifiers can be readily incorporated as base classifiers. The network output weights cannot only be analytically determined, but also transferrable. Additionally, a manifold regularization with Laplacian graph is incorporated, such that it is beneficial to semi-supervised learning. Extensively, we also propose a model of multiple views, referred as MvEDA. Experiments on benchmark visual datasets for video event recognition and object recognition, demonstrate that our EDA methods outperform existing cross-domain learning methods.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Image Process.1
2016 LSDT: Latent Sparse Domain Transfer Learning for Visual Adaptation
abstract
We propose a novel reconstruction-based transfer learning method called latent sparse domain transfer (LSDT) for domain adaptation and visual categorization of heterogeneous data. For handling cross-domain distribution mismatch, we advocate reconstructing the target domain data with the combined source and target domain data points based on ℓ1-norm sparse coding. Furthermore, we propose a joint learning model for simultaneous optimization of the sparse coding and the optimal subspace representation. In addition, we generalize the proposed LSDT model into a kernel-based linear/nonlinear basis transformation learning framework for tackling nonlinear subspace shifts in reproduced kernel Hilbert space. The proposed methods have three advantages: 1) the latent space and the reconstruction are jointly learned for pursuit of an optimal subspace transfer; 2) with the theory of sparse subspace clustering, a few valuable source and target data points are formulated to reconstruct the target data with noise (outliers) from source domain removed during domain adaptation, such that the robustness is guaranteed; and 3) a nonlinear projection of some latent space with kernel is easily generalized for dealing with highly nonlinear domain shift (e.g., face poses). Extensive experiments on several benchmark vision data sets demonstrate that the proposed approaches outperform other state-of-the-art representation-based domain adaptation methods.
Lei Zhang 0038, Wangmeng Zuo, David Zhang 0001
IEEE Trans. Image Process.1
2016 Visual Understanding via Multi-Feature Shared Learning With Global Consistency
abstract
Image/video data is usually represented with multiple visual features. Fusion of multi-source information for establishing attributes has been widely recognized. Multi- feature visual recognition has recently received much attention in multimedia applications. This paper studies visual understanding via a newly proposed l2-norm-based multi-feature shared learning framework, which can simultaneously learn a global label matrix and multiple sub-classifiers with the labeled multi-feature data. Additionally, a group graph manifold regularizer composed of the Laplacian and Hessian graph is proposed. It can better preserve the manifold structure of each feature, such that the label prediction power is much improved through semi-supervised learning with global label consistency. For convenience, we call the proposed approach global-label- consistent classifier (GLCC). The merits of the proposed method include the following: 1) the manifold structure information of each feature is exploited in learning, resulting in a more faithful classification owing to the global label consistency; 2) a group graph manifold regularizer based on the Laplacian and Hessian regularization is constructed ; and 3) an efficient alternative optimization method is introduced as a fast solver owing its speed to convex sub-problems. Experiments on several benchmark visual datasets-the 17-category Oxford Flower dataset, the challenging 101- category Caltech dataset, the YouTube and Consumer Videos dataset, and the large-scale NUS-WIDE dataset-have been used for multimedia understanding . The results demonstrate that the proposed approach compares favorably with state-of-the-art algorithms. An extensive experiment using the deep convolutional activation features also shows the effectiveness of the proposed approach. The code will be available on http://www.escience.cn/people/lei/index.html.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Multim.1
2015 Measurements and Analysis of Large-Scale Fading Characteristics in Curved Subway Tunnels at 920 MHz, 2400 MHz, and 5705 MHz
abstract
Wave propagation characteristics in curved tunnels are of importance for designing reliable communications in subway systems. This paper presents the extensive propagation measurements conducted in two typical types of subway tunnels—traditional arched “Type I” tunnel and modern arched “Type II” tunnel—with 300- and 500-m radii of curvature with different configurations—horizontal and vertical polarizations at 920, 2400, and 5705 MHz, respectively. Based on the measurements, statistical metrics of propagation loss and shadow fading (path-loss exponent, shadow fading distribution, autocorrelation, and crosscorrelation) in all the measurement cases are extracted. Then, the large-scale fading characteristics in the curved subway tunnels are compared with the cases of road and railway tunnels, the other main rail traffic scenarios, and some “typical” scenarios to give a comprehensive insight into the propagation in various scenarios where the intelligent transportation systems are deployed. Moreover, for each of the large-scale fading parameters, extensive analysis and discussions are made to reflect the physical laws behind the observations. The quantitative results and findings are useful to realize intelligent transportation systems in the subway system.
Ke Guan, Bo Ai 0001, Zhangdui Zhong, Carlos F. López, Lei Zhang 0038, Cesar Briso-Rodríguez, Andrej Hrovat, Bei Zhang 0003, Ruisi He
IEEE Trans. Intell. Transp. Syst.5
2009 An orthogonal multi-objective evolutionary algorithm with lower-dimensional crossover
abstract
This paper proposes an multi-objective evolutionary algorithm. The algorithm is based on OMOEA-II. A new linear breeding operator with lower-dimensional crossover and copy operation is used. By using the lower-dimensional crossover, the complexity of searching is decreased so the algorithm converges faster. The orthogonal crossover increase probability of producing potential superior solutions, which helps the algorithm get better results. Ten unconstrained problems are used to test the algorithm. For three problems, the obtained solutions are very close to the true Pareto front, and for one problem, the obtained solutions distribute on part of the true Pareto front.
Sanyou Zeng, Lei Zhang 0038, Yulong Shi, Xin Tian 0001, Yang Yang 0002, Haoqiu Long, Xianqiang Yang 0002, Danpin Yu, Zu Yan
IEEE Congress on Evolutionary Computation4