Yunpeng Bai

dblp:256/5589 · DBLP profile ↗
← Back
57ranked-venue papers
11as first author
56since 2021 · last 2026
0000-0002-6923-672XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 7 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Security and privacy · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PacketPatch: Practical generation and deployment of adversarial packets for byte-feature-based encrypted traffic classification
Yuwei Xu 0001, Yunpeng Bai, Kehui Song, Jie Cao 0009, Qiao Xiang, Guang Cheng 0001
Comput. Secur.3
2026 Cross-Modal Bayesian Inference for training-free open-vocabulary object detection in remote sensing images
Yan Li 0171, Yunpeng Bai, Xingguo Zhang, Ying Li 0017, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.2
2025 Chartist: Task-driven Eye Movement Control for Chart Reading
abstract
| openaire: EC/HE/101141916/EU//Artificial User
Danqing Shi, Yao Wang 0018, Yunpeng Bai, Andreas Bulling, Antti Oulasvirta
CHI3
2025 BrepGiff: Lightweight Generation of Complex B-rep with 3D GAT Diffusion
abstract
Despite advancements in Computer-Aided-Design (CAD) generation, direct generation of complex Boundary Representation (B-rep) CAD models remains challenging. The difficulty arises from the parametric nature of B-rep data, complicating the encoding and generation of its geometric and topological information. In this paper, we introduce BrepGiff, a lightweight generation approach for high-quality and complex B-rep based on 3D Graph Diffusion. First, we transfer B-rep models into 3D graphs representation. Specifically, BrepGiff extracts and integrates topological and geometric features to construct a 3D graph where nodes correspond to face centroids in 3D space, preserving adjacency and degree information. Geometric features are derived by sampling points in the UV domain and extracting face and edge features. BrepGiff then applies Graph Attention Network (GAT) to enforce topological constraints from local to global during the degree-guided diffusion process. With the 3D graph representation and diffusion process, BrepGiff significantly reduces the computational cost and improves the quality, thus achieving lightweight generation of complex models. Experiments show that BrepGiff can generate complex B-rep models (>100 faces) using only 2 RTX4090 GPUs, achieving state-of-the-art performance in B-rep generation.
Xiaoshui Huang, Jiacheng Hao, Yunpeng Bai, Hongping Gan, Yilei Shi
CVPR4
2025 Decoding Cognitive Load: Eye-Tracking Insights into Working Memory and Visual Attention
abstract
Publisher Copyright: © 2025 Copyright held by the owner/author(s).
Xiaofu Jin, Yunpeng Bai, Shuai Ma 0005, Danqing Shi, Luwen Yu, Mingming Fan 0001
ETRA2
2025 PacketMorph: Generation of Recoverable Adversarial Packets Against Encrypted Traffic Classification via Class-Wise Universal Perturbation
Yuwei Xu 0001, Yunpeng Bai, Jie Cao 0009, Kehui Song, Guang Cheng 0001
ICA3PP (4)3
2025 Appearance- and Orientation-aware Fine-grained Rotated Ship Detection in High-Resolution Satellite Imagery
abstract
Ship detection using remote sensing imagery is a crucial research area with both military and civilian applications. However, it remains challenging due to limitations in current ship datasets, such as insufficient volume, incomplete annotations, and inaccuracies. Additionally, ships often exhibit arbitrary orientations, dense clustering, varying aspect ratios, and significant dimensional changes. To address these issues, this paper advances ship detection from both data and methodological perspectives. First, a new dataset, ORSISOD, is introduced. This dataset includes seven finely categorized ship types, annotated with rotated bounding boxes, which are more appropriate for ship detection than traditional horizontal boxes. Second, a novel rotated ship detection method is proposed, incorporating a Dynamic IOU Threshold Selection (DITS) module and a Positive Sample Quality Assessment (PSQA) module. DITS adjusts the IOU threshold based on ship size and shape, while PSQA assesses sample quality using ship aspect ratio and angle information. The ORSISOD dataset was tested on 12 object detection algorithms, providing benchmarks for ship detection. Furthermore, the proposed method was evaluated on both ORSISOD and DOTA datasets, demonstrating superior performance.
Yan Li 0171, Lingyi Liu, Yunpeng Bai, Ying Li 0017, Qiang Shen 0001
ICASSP3
2025 FiffDepth: Feed-Forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation
abstract
Monocular Depth Estimation (MDE) is a fundamental 3D vision problem with numerous applications such as 3D scene reconstruction, autonomous navigation, and AI content creation. However, robust and generalizable MDE remains challenging due to limited real-world labeled data and distribution gaps between synthetic datasets and real data. Existing methods often struggle with real-world test data with low efficiency, reduced accuracy, and lack of detail. To address these issues, we propose an efficient MDE approach named FiffDepth. The key feature of FiffDepth is its use of diffusion priors. It transforms diffusion-based image generators into a feed-forward architecture for detailed depth estimation. FiffDepth preserves key generative features and integrates the strong generalization capabilities of models like DINOv2. Through benchmark evaluations, we demonstrate that FiffDepth achieves exceptional accuracy, stability, and fine-grained detail, offering significant improvements in MDE performance against state-of-the-art MDE approaches. The paper's source code is available here: https://yunpeng1998.github.io/FiffDepth/
Yunpeng Bai, Qixing Huang
ICCV1
2025 MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
Liyuan Deng, Yunpeng Bai, Yongkang Dai, Xiaoshui Huang, Hongping Gan, Dongshuo Huang, Jiacheng Hao, Yilei Shi
ICCV2
2025 BRepFormer: Transformer-Based B-rep Geometric Feature Recognition
Yongkang Dai, Xiaoshui Huang, Yunpeng Bai, Hongping Gan, Yilei Shi
ICMR3
2025 SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have greatly improved 3D reconstruction. However, its substantial data size poses a significant challenge for transmission and storage. While many compression techniques have been proposed, they fail to efficiently adapt to fluctuating network bandwidth, leading to resource wastage. We address this issue from the perspective of size-aware compression, where we aim to compress 3DGS to a desired size by quickly searching for suitable hyperparameters. Through a measurement study, we identify key hyperparameters that affect the size - namely, the reserve ratio of Gaussians and bit-width settings for Gaussian attributes. Then, we formulate this hyperparameter optimization problem as a mixed-integer nonlinear programming (MINLP) problem, with the goal of maximizing visual quality while respecting the size budget constraint. To solve the MINLP, we decouple this problem into two parts: discretely sampling the reserve ratio and determining the bit-width settings using integer linear programming (ILP). To solve the ILP more quickly and accurately, we design a quality loss estimator and a calibrated size estimator, as well as implement a CUDA kernel. Extensive experiments on multiple 3DGS variants demonstrate that our method achieves state-of-the-art performance in post-training compression. Furthermore, our method can achieve comparable quality to leading training-required methods after fine-tuning.
Shuzhao Xie, Weixiang Zhang, Shijia Ge, Sicheng Pan, Yunpeng Bai, Cong Zhang 0002, Xiaoyi Fan 0001, Zhi Wang 0001
ACM Multimedia7
2025 GeoVideo: Introducing Geometric Regularization into Video Generation Model
abstract
Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches operate purely in the 2D pixel space and lack explicit mechanisms for modeling 3D structures, often resulting in temporally inconsistent geometries, implausible motions, and structural artifacts. In this work, we introduce geometric regularization losses into video generation by augmenting latent diffusion models with per-frame depth prediction. We adopted depth as the geometric representation because of the great progress in depth prediction and its compatibility with image-based latent encoders. Specifically, to enforce structural consistency over time, we propose a multi-view geometric loss that aligns the predicted depth maps across frames within a shared 3D coordinate system. Our method bridges the gap between appearance generation and 3D structure modeling, leading to improved spatio-temporal coherence, shape consistency, and physical plausibility. Experiments across multiple datasets show that our approach produces significantly more stable and geometrically consistent results than existing baselines.
Yunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang 0019, Qixing Huang
NeurIPS1
2025 Causal inference model for accurate medical diagnosis in Coronary Artery Bypass Graft operation
Qiyi Zhang, Wei Zhang 0390, Qiang Li 0048, Yunpeng Bai, Weizhi Nie, Keliang Xie
Artif. Intell. Medicine4
2025 Language-Guided Change Detection for high-resolution remote sensing imagery with limited labelled data
Yunpeng Bai, Yefan Xie, Ying Li 0017, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.2
2025 Self-supervised multimodal change detection based on difference contrast learning for remote sensing imagery
Yunpeng Bai, Yefan Xie, Ying Li 0017, Changjing Shang, Qiang Shen 0001
Pattern Recognit.2
2025 TextIR: A Simple Framework for Text-Based Editable Image Restoration
abstract
Many current image restoration approaches utilize neural networks to acquire robust image-level priors from extensive datasets, aiming to reconstruct missing details. Nevertheless, these methods often falter with images that exhibit significant information gaps. While incorporating external priors or leveraging reference images can provide supplemental information, these strategies are limited in their practical scope. Alternatively, textual inputs offer greater accessibility and adaptability. In this study, we develop a sophisticated framework enabling users to guide the restoration of deteriorated images via textual descriptions. Utilizing the text-image compatibility feature of CLIP enhances the integration of textual and visual data. Our versatile framework supports multiple restoration activities such as image inpainting, super-resolution, and colorization. Comprehensive testing validates our technique's efficacy.
Yunpeng Bai, Cairong Wang, Shuzhao Xie, Chao Dong 0005, Chun Yuan 0003, Zhi Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Heads-Up Multitasker: Simulating Attention Switching On Optical Head-Mounted Displays
abstract
Optical Head-Mounted Displays (OHMDs) allow users to read digital content while walking. A better understanding of how users allocate attention between these two tasks is crucial for improving OHMD interfaces. This paper introduces a computational model for simulating users’ attention switches between reading and walking. We model users’ decision to deploy visual attention as a hierarchical reinforcement learning problem, wherein a supervisory controller optimizes attention allocation while considering both reading activity and walking safety. Our model simulates the control of eye movements and locomotion as an adaptation to the given task priority, design of digital content, and walking speed. The model replicates key multitasking behaviors during OHMD reading while walking, including attention switches, changes in reading and walking speeds, and reading resumptions.
Yunpeng Bai, Aleksi Ikkala, Antti Oulasvirta, Shengdong Zhao 0001, Lucia J. Wang, Pengzhi Yang, Peisen Xu
CHI1
2024 DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
Yunpeng Bai, Xintao Wang 0002, Yan-Pei Cao 0001, Yixiao Ge, Chun Yuan 0003, Ying Shan
ECCV (31)1
2024 MesonGS: Post-training Compression of 3D Gaussians via Efficient Attribute Transformation
Shuzhao Xie, Weixiang Zhang, Yunpeng Bai, Rongwei Lu, Shijia Ge, Zhi Wang 0001
ECCV (33)4
2024 Dual-view Traffic Identification for Open Source Proxy Software through Early Flows
abstract
Open Source Proxy Software (OSPS) provides privacy protection for users accessing the Internet by constructing a private anonymizing network. However, there is a growing concern about whether OSPS can actually prevent privacy leaks as it claims. Researchers have attempted to use AI-based techniques to identify OSPS, but there are two shortcomings in the current studies. First, there is no complete public dataset to support the identification tasks for different requirements. The existing datasets do not cover the most commonly used OSPS tools and their typical configurations. Second, with the introduction of deep learning techniques, the models continue to become complex, resulting in significant computational overhead. Using early flows for identification may make the model lighter, but result in weaker representations and lower classification performance. To address the above shortcomings, we have carried out pioneering work on OSPS traffic identification through early flows. First, we collect the access traffic of three OSPS tools and create a dataset with 8 protocol configurations. Second, we present an innovative Dual-View Identification (DVI) method for OSPS traffic. By considering both static and dynamic views, DVI effectively characterizes early flows and achieves accurate classification through feature fusion. In the static view, spatial distribution features are extracted by representing the early flows as a grayscale picture. In the dynamic view, spatial features and temporal correlations are represented using a flow with multiple packets, similar to a video with multiple frames. Comparative experiments show that DVI achieves over 90% accuracy and F1 scores in all three tasks, which greatly improves its ability to identify different protocol configurations and access sites. Besides, DVI outperforms 5 state-of-the-art methods and achieves low parameters and FLOPS through early flows.
Yuwei Xu 0001, Yunpeng Bai, Yuquan Zhang, Yige Song, Qiao Xiang, Guang Cheng 0001
ISPA2
2024 Tarnhelm: Using Adversarial Samples to Protect User Privacy Against Traffic Identification
Yuwei Xu 0001, Yunpeng Bai, Jie Cao 0009, Liang He 0002, Guang Cheng 0001
SecureComm (3)2
2024 Baffle: Hiding Backdoors in Offline Reinforcement Learning Datasets
abstract
Reinforcement learning (RL) makes an agent learn from trial-and-error experiences gathered during the interaction with the environment. Recently, offline RL has become a popular RL paradigm because it saves the interactions with environments. In offline RL, data providers share large pre-collected datasets, and others can train high-quality agents without interacting with the environments. This paradigm has demonstrated effectiveness in critical tasks like robot control, autonomous driving, etc. However, less attention is paid to investigating the security threats to the offline RL system. This paper focuses on backdoor attacks, where some perturbations are added to the data (observations) such that given normal observations, the agent takes high-rewards actions, and low-reward actions on observations injected with triggers. In this paper, we propose Baffle (Backdoor Attack for Offline Reinforcement Learning), an approach that automatically implants backdoors to RL agents by poisoning the offline RL dataset, and evaluate how different offline RL algorithms react to this attack. Our experiments conducted on four tasks and nine offline RL algorithms expose a disquieting fact: none of the existing offline RL algorithms has been immune to such a backdoor attack. More specifically, Baffle modifies 10% of the datasets for four tasks (3 robotic controls and 1 autonomous driving). Agents trained on the poisoned datasets perform well in normal settings. However, when triggers are presented, the agents’ performance decreases drastically by 63.2%, 53.9%, 64.7%, and 47.4% in the four tasks on average. The backdoor still persists after fine-tuning poisoned agents on clean datasets. We further show that the inserted backdoor is also hard to be detected by a popular defensive method. This paper calls attention to developing more effective protection for the open-source offline RL dataset.
Chen Gong 0005, Zhou Yang 0003, Yunpeng Bai, Junda He, Jieke Shi, Kecen Li, Arunesh Sinha, Xinwen Hou, David Lo 0001, Tianhao Wang 0001
SP3
2024 Perturbing Vulnerable Bytes in Packets to Generate Adversarial Samples Resisting DNN-Based Traffic Monitoring
abstract
Leveraging the advanced capabilities of Deep Neural Networks (DNNs), attackers can precisely detect users' online activities through traffic monitoring, nullifying the efficacy of current encrypted communication tools/protocols and progressively resulting in privacy leakage. Several defensive methods against DNN-based traffic monitoring (DTM) have been proposed; however, these methods often rely excessively on prior knowledge and incur inevitable additional bandwidth overhead (BWO). Moreover, they frequently generate invalid packets that violate network transmission constraints. To address these drawbacks, in this paper, we propose BYTEFLIPPING, a byte-space grey-box defensive method, which perturbs vulnerable bytes in the transport layer payload to generate adversarial sample packets. We design a Payload Byte Vulnerability Ranking algorithm to pinpoint the most vulnerable bytes and based on this generate adversarial packets to defend DTM. Extensive experiments reveal that ByteFLIPPING performs well in protecting against three DTM methods across two benchmark datasets, significantly decreasing the accuracy of the state-of-the-art ET-BERT by 94%. Compared to baseline defensive methods, BYTEFLIPPING incurs no extra BWO, offers more dependable packet validity, and boasts greater feasibility.
Jie Cao 0009, Zhengxin Xu, Yunpeng Bai, Yuwei Xu 0001, Qiao Xiang, Guang Cheng 0001
TrustCom3
2024 Early prediction of sepsis using chatGPT-generated summaries and structured data
Qiang Li 0048, Hanbo Ma, Dan Song 0006, Yunpeng Bai, Keliang Xie
Multim. Tools Appl.4
2024 A Multitask Network for Joint Multispectral Pansharpening on Diverse Satellite Data
abstract
Despite the rapid advance in multispectral (MS) pansharpening, existing convolutional neural network (CNN)-based methods require training on separate CNNs for different satellite datasets. However, such a single-task learning (STL) paradigm often leads to overlooking any underlying correlations between datasets. Aiming at this challenging problem, a multitask network (MTNet) is presented to accomplish joint MS pansharpening in a unified framework for images acquired by different satellites. Particularly, the pansharpening process of each satellite is treated as a specific task, while MTNet simultaneously learns from all data obtained from these satellites following the multitask learning (MTL) paradigm. MTNet shares the generic knowledge between datasets via task-agnostic subnetwork (TASNet), utilizing task-specific subnetworks (TSSNets) to facilitate the adaptation of such knowledge to a certain satellite. To tackle the limitation of the local connectivity property of the CNN, TASNet incorporates Transformer modules to derive global information. In addition, band-aware dynamic convolutions (BDConvs) are proposed that can accommodate various ground scenes and bands by adjusting their respective receptive field (RF) size. Systematic experimental results over different datasets demonstrate that the proposed approach outperforms the existing state-of-the-art (SOTA) techniques.
Dong Wang 0022, Chanyue Wu, Yunpeng Bai, Ying Li 0017, Changjing Shang, Qiang Shen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination Mechanism
abstract
Recently, some challenging tasks in multi-agent systems have been solved by some hierarchical reinforcement learning methods. Inspired by the intra-level and inter-level coordination in the human nervous system, we propose a novel value decomposition framework HAVEN based on hierarchical reinforcement learning for fully cooperative multi-agent problems. To address the instability arising from the concurrent optimization of policies between various levels and agents, we introduce the dual coordination mechanism of inter-level and inter-agent strategies by designing reward functions in a two-level hierarchy. HAVEN does not require domain knowledge and pre-training, and can be applied to any value decomposition variant. Our method achieves desirable results on different decentralized partially observable Markov decision process domains and outperforms other popular multi-agent hierarchical reinforcement learning algorithms.
Zhiwei Xu 0005, Yunpeng Bai, Bin Zhang 0052, Dapeng Li 0001
AAAI2
2023 High-fidelity Facial Avatar Reconstruction from Monocular Video with Generative Priors
abstract
High-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has been considered for facial avatar reconstruction. However, the complex facial dynamics and missing 3D information in monocular videos raise significant challenges for faithful facial reconstruction. In this work, we propose a new method for NeRF-based facial avatar reconstruction that utilizes 3D-aware generative prior. Different from existing works that depend on a conditional deformation field for dynamic modeling, we propose to learn a personalized generative prior, which is formulated as a local and low dimensional subspace in the latent space of 3D-GAN. We propose an efficient method to construct the personalized generative prior based on a small set of facial images of a given individual. After learning, it allows for photo-realistic rendering with novel views, and the face reenactment can be realized by performing navigation in the latent space. Our proposed method is applicable for different driven signals, including RGB images, 3DMM coefficients, and audio. Compared with existing works, we obtain superior novel view synthesis results and faithfully face reenactment performance. The code is available here https://github.com/bbaaii/HFA-GP.
Yunpeng Bai, Yanbo Fan, Xuan Wang 0009, Yong Zhang 0034, Jingxiang Sun, Chun Yuan 0003, Ying Shan
CVPR1
2023 SAR Image Despeckling with Residual-in-Residual Dense Generative Adversarial Network
abstract
Deep convolutional neural networks have delivered remarkable aptitude in performing Synthetic Aperture Radar (SAR) image speckle removal tasks. Such approaches are nevertheless constrained in balancing speckle removal and preservation of spatial information, particularly with respect to strong speckle noise. In this paper, a novel residual-in-residual dense generative adversarial network is proposed to effectively suppress SAR image speckle while retaining rich spatial information. A despeckling sub-network composed of residual-in-residual dense blocks with an encoder-decoder structure is devised to learn end-to-end mapping of noisy images onto noise-free images, where the combination of residual-in-residual structure and dense connection significantly enhances the feature representation capability. In addition, a discriminator sub-network with a fully convolutional structure is introduced, and the adversarial learning strategy is adopted to continuously refine the quality of despeckled results. Systematic experimental results on simulated and real SAR images demonstrate that the novel approach offers superior performance in both quantitative and visual evaluation as compared to state-of-the-art methods.
Yunpeng Bai, Yayuan Xiao, Ying Li 0017, Changjing Shang, Qiang Shen 0001
ICASSP1
2023 HSR-Diff: Hyperspectral Image Super-Resolution via Conditional Diffusion Models
abstract
Despite the proven significance of hyperspectral images (HSIs) in performing various computer vision tasks, its potential is adversely affected by the low-resolution (LR) property in the spatial domain, resulting from multiple physical factors. Inspired by recent advancements in deep generative models, we propose an HSI Super-resolution (SR) approach with Conditional Diffusion Models (HSR-Diff) that merges a high-resolution (HR) multispectral image (MSI) with the corresponding LR-HSI. HSR-Diff generates an HR-HSI via repeated refinement, in which the HR-HSI is initialized with pure Gaussian noise and iteratively refined. At each iteration, the noise is removed with a Conditional Denoising Transformer (CDFormer) that is trained on denoising at different noise levels, conditioned on the hierarchical feature maps of HR-MSI and LR-HSI. In addition, a progressive learning strategy is employed to exploit the global information of full-resolution images. Systematic experiments have been conducted on four public datasets, demonstrating that HSR-Diff outperforms state-of-the-art methods.
Chanyue Wu, Dong Wang 0022, Yunpeng Bai, Hanyu Mao, Ying Li 0017, Qiang Shen 0001
ICCV3
2023 PS-NeRV: Patch-Wise Stylized Neural Representations for Videos
abstract
We study how to represent a video with implicit neural representations (INRs). Classical INRs methods generally utilize MLPs to map input coordinates to output pixels. While some recent works have tried to directly reconstruct the whole image with CNNs. However, we argue that both the above pixel-wise and image-wise strategies are not favorable to video data. Instead, we propose a patch-wise solution, PS-NeRV, which represents videos as a function of patches and the corresponding patch coordinate. It naturally inherits the advantages of image-wise methods, and achieves excellent reconstruction performance with fast decoding speed. The whole method includes conventional modules, like positional embedding, MLPs and CNNs. We also introduce AdaIN to enhance intermediate features. Extensive experiments have demonstrated its effectiveness in several video-related tasks, such as video compression and video inpainting.
Yunpeng Bai, Chao Dong 0005, Cairong Wang, Chun Yuan 0003
ICIP1
2023 Chest X-ray Image Classification: A Causal Perspective
Weizhi Nie, Dan Song 0006, Yunpeng Bai, Keliang Xie, Anan Liu
MICCAI (3)4
2023 SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin Regularization
abstract
Mixed-precision quantization (MPQ) suffers from the time-consuming process of searching the optimal bit-width allocation (i.e., the policy) for each layer, especially when using large-scale datasets such as ISLVRC-2012. This limits the practicality of MPQ in real-world deployment scenarios. To address this issue, this paper proposes a novel method for efficiently searching for effective MPQ policies using a small proxy dataset instead of the large-scale dataset used for training the model. Deviating from the established norm of employing a consistent dataset for both model training and MPQ policy search stages, our approach, therefore, yields a substantial enhancement in the efficiency of MPQ exploration. Nonetheless, using discrepant datasets poses challenges in searching for a transferable MPQ policy. Driven by the observation that quantization noise of sub-optimal policy exerts a detrimental influence on the discriminability of feature representations---manifesting as diminished class margins and ambiguous decision boundaries---our method aims to identify policies that uphold the discriminative nature of feature representations, i.e., intra-class compactness and inter-class separation. This general and dataset-independent property makes us search for the MPQ policy over a rather small-scale proxy dataset and then the policy can be directly used to quantize the model trained on a large-scale dataset. Our method offers several advantages, including high proxy data utilization, no excessive hyper-parameter tuning, and high searching efficiency. We search high-quality MPQ policies with the proxy dataset that has only 4% of the data scale compared to the large-scale target dataset, achieving the same accuracy as searching directly on the latter, improving MPQ searching efficiency by up to 300×.
Kai Ouyang, Zenghao Chai, Yunpeng Bai, Zhi Wang 0001, Wenwu Zhu 0001
ACM Multimedia4
2023 Instrumental Variable Learning for Chest X-ray Classification
abstract
The chest X-ray (CXR) is commonly employed to diagnose thoracic illnesses, but the challenge of achieving accurate automatic diagnosis through this method persists due to the complex relationship between pathology. In the past few years, numerous approaches based on deep learning have been proposed to address this issue but confounding factors such as image resolution or noise problems often damage model performance. In this paper, we focus on the chest X-ray classification task and proposed an interpretable instrumental variable (IV) learning framework, to eliminate the spurious association and obtain accurate causal representation. Specifically, we first construct a structural causal model (SCM) for our task and learn the confounders and the preliminary representations of IV, we then leverage electronic health record (EHR) as auxiliary information and we fuse the above feature with our transformer-based semantic fusion module, so the IV has the medical semantic. Meanwhile, the reliability of IV is further guaranteed via the constraints of mutual information between related causal variables. Finally, our approach's performance is demonstrated using the MIMIC-CXR, NIH ChestX-ray 14, and CheXpert datasets, and we achieve competitive results.
Weizhi Nie, Dan Song 0006, Yunpeng Bai, Keliang Xie, Anan Liu
SMC4
2023 SharpEye: Identify mKCP Camouflage Traffic through Feature Optimization
abstract
As a new self-developed protocol of V2Ray, mKCP disguises users’ network access as communication of four network applications by forging application layer headers to evade traffic-based detection. The emergence of mKCP has received widespread attention. Whether mKCP can provide secure network access that protects user privacy is the focus. Traditional methods cannot identify mKCP camouflage traffic, but machine learning (ML)-based traffic identification is considered a promising direction. Unlike the previous network traffic classification, mKCP camouflage traffic identification introduces new challenges. First, existing work has neither published any dataset containing mKCP camouflage traffic nor designed specific traffic features. Second, no researchers have optimized the identification scheme for deployment on network devices. Aiming at the shortcomings, we propose SharpEye, an ML-based mKCP camouflage traffic identification scheme. The novelty of our work lies in three points. Firstly, a complete dataset containing mKCP camouflage traffic is constructed through long-term traffic collection. Secondly, by analyzing the communication patterns of mKCP traffic, a feature set mFS is designed to improve identification accuracy. Finally, a two-stage feature selection method mGBFS is proposed to improve the operation efficiency. The experimental results show that mFS can enhance the performance of classifiers in identifying mKCP camouflage traffic, and mGBFS reduces the running time and overhead while ensuring high accuracy. Therefore, SharpEye achieves accurate and efficient mKCP camouflage traffic identification.
Yuwei Xu 0001, Zizhi Zhu, Yunpeng Bai, Lilanyi Wu, Kehui Song, Guang Cheng 0001
TrustCom3
2023 Deep collaborative learning with class-rebalancing for semi-supervised change detection in SAR images
abstract
Deep learning reveals excellent potential for accomplishing change detection in SAR imagery. Yet, it suffers from the problem of requiring large amounts of labeled samples, whilst labeling SAR imagery for change detection requires experts to label individual images at the pixel level , which is extremely tedious and time-consuming. Also, sample imbalance continues to present a serious challenge for the existing change detection techniques. To tackle these problems, in this study, a Deep Collaborative semi-supervised learning Framework with Class-Rebalancing (DCF-CRe) is proposed for SAR imagery change detection, by exploiting Convolutional Neural Network (CNN) and deep clustering. In particular, a Siamese Difference Fusion Network (SDFNet) is devised to implement change detection while effectively reducing the information loss due to the generation of difference images and highlighting features of the changed regions.In so doing, only a tiny batch of labeled samples is utilized to train SDFNet in order to obtain predicted change map and deep features. In addition, the Approximate Rank-Order Clustering (AROC) algorithm is employed to cluster the deep features, generating pseudo-labels for abundant unlabeled samples . DCF-CRe is then applied to select appropriate pseudo-labels and to add labeled samples to train SDFNet. Experimental results evaluated on six challenging datasets show that this proposed approach can achieve performance superior to state-of-the-art change detection methods for SAR imagery.
Yunpeng Bai, Yefan Xie, Huibin Ge, Ying Li 0017, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.2
2023 HMF-Former: Spatio-Spectral Transformer for Hyperspectral and Multispectral Image Fusion
abstract
The key to hyperspectral image (HSI) and multispectral image (MSI) fusion is to take advantage of the properties of interspectra self-similarities of HSIs and spatial correlations of MSIs. However, leading convolutional neural network (CNN)-based methods show shortcomings in capturing long-range dependencies and self-similarity prior. To this end, we propose a simple yet efficient Transformer-based network, hyperspectral and multispectral image fusion (HMF)-Former, for the HSI/MSI fusion. The HMF-Former adopts a U-shaped architecture with a spatio-spectral Transformer block (SSTB) as the basic unit. In the SSTB, embedded spatial-wise multihead self-attention (Spa-MSA) and spectral-wise multihead self-attention (Spe-MSA) effectively capture interactions of spatial regions and interspectra dependencies, respectively. They are consistent with the properties of spatial correlations of MSIs and interspectra self-similarities of HSIs. In addition, specially designed SSTB enables the HMF-Former to capture both local and global features while maintaining linear complexity. Extensive experiments on four benchmark datasets show that our method significantly outperforms state-of-the-art methods.
Tengfei You, Chanyue Wu, Yunpeng Bai, Dong Wang 0022, Huibin Ge, Ying Li 0017
IEEE Geosci. Remote. Sens. Lett.3
2023 Boundary-Aware Network With Two-Stage Partial Decoders for Salient Object Detection in Remote Sensing Images
abstract
Salient object detection is a binary pixel-wise classification to distinguish objects in an image, and also have attracted many research interests in the optical Remote Sensing Images (RSIs). The existing state-of-the-art method exploits the full encoder-decoder architecture to predict salient objects in the optical RSIs, suffering from the problem of unsmooth edges and incomplete structures. To address these problems, in this paper, we propose a Boundary-Aware Network (BANet) with two-stage partial decoders sharing the same encoders for salient object detection in RSIs. Specifically, a Boundary-Aware Partial Decoder (BAD) is introduced at the first stage to focus on learning clear edges of salient objects. To solve the pixel-imbalance problem between boundary and background, an edge-aware loss is proposed to guide learning the BAD network. The resulting features are then employed in turn to enhance high-level features. Afterwards, the Structure-Aware Partial Decoder (SAD) is further introduced at the second stage to improve the structure integrity of salient objects. To alleviate the problem of incomplete structures, the structural similarity loss is further proposed to supervise learning the SAD network. In a consequence, our proposed BANet can predict salient objects with clear edges and complete structure, while reducing model parameters due to the discardment of low-level features. Besides, training a deep neural network requires a large amount of images, and the current benchmark datasets for optical remote sensing images are not large enough. Therefore, we also create a large-scale challenging dataset for salient object detection in RSIs. Extensive experiments demonstrate that our proposed BANet outperforms previous RSI SOD models on all existing benchmark datasets and our new presented dataset available at https://github.com/QingpingZheng/RSISOD.
Qingping Zheng, Yunpeng Bai, Jiankang Deng, Ying Li 0017
IEEE Trans. Geosci. Remote. Sens.3
2022 Curiosity-Driven and Victim-Aware Adversarial Policies
abstract
Recent years have witnessed great potential in applying Deep Reinforcement Learning (DRL) in various challenging applications, such as autonomous driving, nuclear fusion control, complex game playing, etc. However, recently researchers have revealed that deep reinforcement learning models are vulnerable to adversarial attacks: malicious attackers can train adversarial policies to tamper with the observations of a well-trained victim agent, the latter of which fails dramatically when faced with such an attack. Understanding and improving the adversarial robustness of deep reinforcement learning is of great importance in enhancing the quality and reliability of a wide range of DRL-enabled systems.
Chen Gong 0005, Zhou Yang 0003, Yunpeng Bai, Jieke Shi, Arunesh Sinha, David Lo 0001, Xinwen Hou
ACSAC3
2022 Semantic-Sparse Colorization Network for Deep Exemplar-Based Colorization
Yunpeng Bai, Chao Dong 0005, Zenghao Chai, Andong Wang, Zhengzhuo Xu, Chun Yuan 0003
ECCV (6)1
2022 Coarse-To-Fine Unsupervised Change Detection for Remote Sensing Images Via Object-Based MRF and Inception UNET
abstract
With the rapid development of various satellite sensor techniques, remote sensing imagery has been an important source of data in change detection applications. This paper aims to propose an unsupervised change detection method based on Object-based Markov Random Filed (OMRF) and Inception UNet (IUNet). Our method first utilizes a difference image (DI) obtained from two bi-temporal images as the initial feature, and proposes the OMRF algorithm based on homogeneous region to pre-classify the DI thus derive the coarse change map. The IUNet is then constructed to extract the points with high confidence from the coarse change map for training. Eventually, the trained model is fed to classify the original feature, then the final change map is obtained. Experimental results indicate that our method yields great detection results even without supervision.
Yunpeng Bai, Ying Li 0017
ICASSP2
2022 CMS-LSTM: Context Embedding and Multi-Scale Spatiotemporal Expression LSTM for Predictive Learning
abstract
Spatiotemporal predictive learning (ST-PL) is a hotspot with numerous applications, such as object movement and mete-orological prediction. It aims at predicting the subsequent frames via observed sequences. However, inherent uncer-tainty among consecutive frames exacerbates the difficulty in long-term prediction. To tackle the increasing ambigu-ity during forecasting, we design CMS-LSTM to focus on context correlations and multi-scale spatiotemporal flow with details on fine-grained locals, containing two elaborate de-signed blocks: Context Embedding (CE) and Spatiotemporal Expression (SE) blocks. CE is designed for abundant context interactions, while SE focuses on multi-scale spatiotemporal expression in hidden states. The newly introduced blocks also facilitate other spatiotemporal models (e.g., PredRNN, SA-ConvLSTM) to produce representative implicit features for ST-PL and improve prediction quality. Qualitative and quanti-tative experiments demonstrate the effectiveness and flexibil-ity of our proposed method. With fewer params, CMS-LSTM outperforms state-of-the-art methods in numbers of metrics on two representative benchmarks and scenarios. Code is available at https://github.com/czh-98/CMS-LSTM.
Zenghao Chai, Zhengzhuo Xu, Yunpeng Bai, Zhihui Lin, Chun Yuan 0003
ICME3
2022 Multi-Agent Hyper-Attention Policy Optimization
Bin Zhang 0052, Zhiwei Xu 0005, Yiqun Chen 0004, Dapeng Li 0001, Yunpeng Bai, Lijuan Li 0002
ICONIP (1)5
2022 Efficient Policy Generation in Multi-agent Systems via Hypergraph Neural Network
Bin Zhang 0052, Yunpeng Bai, Zhiwei Xu 0005, Dapeng Li 0001
ICONIP (2)2
2022 Cooperative Multi-Agent Reinforcement Learning with Hypergraph Convolution
abstract
Recent years have witnessed the great success of multi-agent systems (MAS). Value decomposition, which decom-poses joint action values into individual action values, has been an important work in MAS. However, many value decomposition methods ignore the coordination among different agents, leading to the notorious “lazy agents” problem. To enhance the coordination in MAS, this paper proposes HyperGraph CoNvo-lution MIX (HGCN-MIX), a method that incorporates hyper-graph convolution with value decomposition. HGCN-MIX models agents as well as their relationships as a hypergraph, where agents are nodes and hyperedges among nodes indicate that the corresponding agents can coordinate to achieve larger rewards. Then, it trains a hypergraph that can capture the collaborative relationships among agents. Leveraging the learned hypergraph to consider how other agents' observations and actions affect their decisions, the agents in a MAS can better coordinate. We evaluate HGCN-MIX in the StarCraft II multi-agent challenge benchmark. The experimental results demonstrate that HGCN-MIX can train joint policies that outperform or achieve a similar level of performance as the current state-of-the-art techniques. We also observe that HGCN-MIX has an even more significant improvement of performance in the scenarios with a large amount of agents. Besides, we conduct additional analysis to emphasize that when the hypergraph learns more relationships, HGCN-MIX can train stronger joint policies.
Yunpeng Bai, Chen Gong 0005, Bin Zhang 0052, Xinwen Hou, Yu Liu 0078
IJCNN1
2022 Contrastive Learning in Wavelet Domain for Image Dehazing
abstract
Image dehazing remains a challenging problem because it is hard to restore a clean scene from a severely degraded hazy image. However, existing learning-based dehazing methods mostly ignore the fact that the interference of haze to an image is mainly concentrated in the low-frequency components. If all image components are processed indiscriminately, it is difficult to achieve a good restoration and accurate details cannot be guaranteed. In order to process the hazy images hierarchically, we propose a low-frequency sub-band contrastive regularization (LSCR) in the wavelet domain to ensure that the components of the restored image mainly affected by haze are pulled closer to the clear image and pushed far away from the hazy image. In addition, a high-frequency sub-band loss is also introduced to make high-frequency components of the restored image consistent with the clear image. Our method can better restore the haze-free image and achieve more accurate and rich details. The extensive experiments on synthetic and real-world datasets verify that the proposed method outperforms previous approaches.
Yunpeng Bai, Chun Yuan 0003
IJCNN1
2022 Mingling Foresight with Imagination: Model-Based Cooperative Multi-Agent Reinforcement Learning
abstract
Recently, model-based agents have achieved better performance than model-free ones using the same computational budget and training time in single-agent environments. However, due to the complexity of multi-agent systems, it is tough to learn the model of the environment. The significant compounding error may hinder the learning process when model-based methods are applied to multi-agent tasks. This paper proposes an implicit model-based multi-agent reinforcement learning method based on value decomposition methods. Under this method, agents can interact with the learned virtual environment and evaluate the current state value according to imagined future states in the latent space, making agents have the foresight. Our approach can be applied to any multi-agent value decomposition method. The experimental results show that our method improves the sample efficiency in different partially observable Markov decision process domains.
Zhiwei Xu 0005, Dapeng Li 0001, Bin Zhang 0052, Yuan Zhan, Yunpeng Bai
NeurIPS5
2022 Locality-Aware Rotated Ship Detection in High-Resolution Remote Sensing Imagery Based on Multiscale Convolutional Network
abstract
Ship detection has been an active and vital topic in the field of remote sensing for a decade, but it is still a challenging problem due to the large-scale variations, the high aspect ratios, the intensive and rotated arrangement, and the background clutter disturbance. In this letter, we propose a locality-aware rotated ship detection (LARSD) framework based on a multiscale convolutional neural network (CNN) to tackle these issues. The proposed framework applies a UNet-like multiscale CNN to generate multiscale feature maps with high-level semantic information in high resolution. Then, an anchor-based rotated bounding box regression is applied for directly predicting the probability, the edge distances, and the angle of ships. Finally, a locality-aware score alignment (LASA) is proposed to fix the mismatch between classification results and location results caused by the independence of each subnet. Furthermore, to enlarge the data sets of ship detection, we build a new high-resolution ship detection (HRSD) data set, where 2499 images and 9269 instances were collected from Google Earth with different resolutions. Experiments based on public data set high-resolution ship collection 2016 (HRSC2016) and our HRSD data set demonstrate that our detection method achieves the state-of-the-art performance.
Lingyi Liu, Yunpeng Bai, Ying Li 0017
IEEE Geosci. Remote. Sens. Lett.2
2022 MetaPan: Unsupervised Adaptation With Meta-Learning for Multispectral Pansharpening
abstract
Multispectral (MS) pansharpening aims to improve the spatial resolution of MS images (MSI) using the spatial details of panchromatic (PAN) images. Due to the gap of prior knowledge between the simulated data and real-world cases, unsupervised learning-based approaches have grown increasing interest. However, some key hyper-parameters, such as the initial weights of the networks, are set manually, which significantly impacts the fusion performance. To tackle this problem, we propose a novel unsupervised adaptation method with meta-learning for MS pansharpening (MetaPan), in which the meta-learning aims to automatically learn the initial parameters of a three-stream fusion network (TSFNet) for unsupervised adaptation learning (UAL). Specifically, the TSFNet consists of a PAN stream, an MS stream, and a fusion stream, where the fusion stream implicitly leverages domain-specific knowledge of input image pairs while the other two streams explicitly inject spatial details and spectral information into the fusion stream. The MetaPan consists of a pre-training stage, a meta-learning stage, and a UAL stage. At the pre-training stage, the TSFNet is trained with the supervision of simulated ground truth such that it is universal for all image pairs. Then, the process of meta-learning optimizes for an internal representation of network parameters that can adapt to a specific image pair with UAL through only a few steps. Finally, the learned internal representation is fine-tuned to a real-world image pair (a test image pair) with UAL. Experiments on two datasets show that our method performs better than state-of-the-art methods in both quantitative metrics and visual appearance.
Dong Wang 0022, Yunpeng Bai, Ying Li 0017
IEEE Geosci. Remote. Sens. Lett.3
2022 Convolutional LSTM-Based Hierarchical Feature Fusion for Multispectral Pan-Sharpening
abstract
Multispectral (MS) pan-sharpening aims at producing high-resolution (HR) MS images in both spatial and spectral domains, by merging single-band panchromatic (PAN) images and corresponding MS images with low spatial resolution. The intuitive way to accomplish such MS pan-sharpening tasks, or to reconstruct ideal HR-MS images, is to extract feature pairs from the given PAN and MS images and to fuse the results. Therefore, feature extraction and feature fusion are two key components for MS pan-sharpening. This article presents a novel MS pan-sharpening network (MPNet), including a heterogeneous pair of feature extraction pathways (FEPs) and a convolutional long short-term memory (ConvLSTM)-based hierarchical feature fusion module (HFFM). Specifically, we design a PAN FEP to extract 2-D feature maps via 2-D convolutions and dual attention, while an MS FEP is introduced in an effort to obtain 3-D representations of MS image by 3-D convolutions and triple attention. To merge the resulting hierarchical features, the ConvLSTM-based HFFM is developed, leveraging intralevel fusion, interlevel fusion, and information exchange within one single framework. Here, the interlevel fusion is implemented with the ConvLSTM to capture the dependencies among hierarchical features, reduce redundant information, and effectively integrate them via its recurrent architecture. The information exchange between different FEPs helps enhance the representations for subsequent processing. Systematic comparative experiments have been conducted on three publicly available datasets at both reduced resolution and full resolution, demonstrating that the proposed MPNet outperforms state-of-the-art methods in the literature.
Dong Wang 0022, Yunpeng Bai, Chanyue Wu, Ying Li 0017, Changjing Shang, Qiang Shen 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 3-D-ANAS: 3-D Asymmetric Neural Architecture Search for Fast Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) provide abundant spectral and spatial information, playing an irreplaceable role in land-cover classification. Recently, based on deep learning (DL) technologies, an increasing number of HSI classification approaches have been proposed, which demonstrate promising performance. However, previous studies suffer from two major drawbacks: 1) the architecture of most DL models is manually designed, relies on specialized knowledge, and is relatively tedious. Moreover, in HSI classifications, datasets captured by different sensors have different physical properties. Correspondingly, different models need to be designed for different datasets, which further increases the workload of designing architectures and 2) the mainstream framework is a patch-to-pixel framework. The overlap regions of patches of adjacent pixels are calculated repeatedly, which increases computational cost and time cost. In addition, the classification accuracy is sensitive to the patch size, which is artificially set based on extensive investigation experiments. To overcome the issues mentioned above, we first propose a 3-D asymmetric neural network search algorithm and leverage it to automatically search for efficient architectures for HSI classifications. By analyzing the characteristics of HSIs, we specifically build a 3-D asymmetric decomposition search space, where spectral and spatial information is processed with different decomposition convolutions. Furthermore, we propose a new fast classification framework, i.e., pixel-to-pixel classification framework, which has no repetitive operations and reduces the overall cost. Experiments on three public HSI datasets captured by different sensors demonstrate the networks designed by our 3-D asymmetric neural architecture search (3-D-ANAS) achieve competitive performance compared to several state-of-the-art methods, while having a much faster inference speed. Code is available at:https://github.com/hkzhang91/3D-ANAS.
Haokui Zhang, Chengrong Gong, Yunpeng Bai, Zongwen Bai, Ying Li 0017
IEEE Trans. Geosci. Remote. Sens.3
2021 Heterogeneous two-Stream Network with Hierarchical Feature Prefusion for Multispectral Pan-Sharpening
abstract
Multispectral (MS) pan-sharpening aims at producing a high spatial resolution (HR) MS image by fusing a single-band HR panchromatic (PAN) image and a corresponding MS image with low spatial resolution. In this paper, we propose a heterogeneous two-stream network (HTSNet) with hierarchical feature prefusion for MS pan-sharpening. The HTSNet employs a heterogeneous group of spatial and spectral streams for spatial and spectral information extraction, respectively. The spatial stream utilizes a 2D CNN for spatial information extraction from the PAN images, and the spectral stream obtains spectral feature cubes from the MS images by a 3D CNN. At the same time, a prefusion module is introduced to prefuse the spatial details with spectral information and transfer information between different streams, which can enhance later processing. In the experiment, the Gaofen-2 satellite dataset is utilized to compare the proposed method with the state-of-the-art MS pan-sharpening methods. Experimental results demonstrate the superiority of our HTSNet in terms of visual effect and quantitative qualities.
Dong Wang 0022, Yunpeng Bai, Bendu Bai, Chanyue Wu, Ying Li 0017
ICASSP2
2021 A Meta-Learning Framework for Few-Shot Classification of Remote Sensing Scene
abstract
While achieving remarkable success in remote sensing (RS) scene classification for the past few years, convolutional neural network (CNN) based methods suffer from the demand for large amounts of training data. The bottleneck in prediction accuracy has shifted from data processing limits toward a lack of ground truth samples, usually collected manually by experienced experts. In this work, we provide a metalearning framework for few-shot classification of RS scene. Under the umbrella of meta-learning, we show it is possible to learn much information about a new category from only 1 or 5 samples. The proposed method is based on Prototypical Networks with a pre-trained stage and a learnable similarity metric. The experimental results show that our method outperforms three state-of-the-art few-shot algorithms and one typical CNN-based method, D-CNN, on two challenging datasets: NWPU-RESISC45 and RSD46-WHU.
Yunpeng Bai, Dong Wang 0022, Bendu Bai, Ying Li 0017
ICASSP2
2021 Wide-Sense Stationary Policy Optimization with Bellman Residual on Video Games
abstract
Deep Reinforcement Learning (DRL) has an increasing application in video games. However, it usually suffers from unstable training, low sampling efficiency, etc. Under the assumption that Bellman residual follows a stationary random process when the training process is convergent, we propose the Wide-sense Stationary Policy Optimization (WSPO) framework, which leverages the Wasserstein distance from the Bellman Residual Distribution (BRD) between two adjacent time steps, to stabilize the training stage and improve the sampling efficiency. We minimize the Wasserstein distance with Quantile Regression, where the specific form of BRD is not needed. Finally, we combine WSPO with Advantage Actor-Critic (A2C) algorithm and Deep Deterministic Policy Gradient (DDPG) algorithm. We evaluate WSPO on Atari 2600 video games and continuous control tasks, illustrating that WSPO compares or outperforms the state-of-the-art algorithms we tested.
Chen Gong 0005, Yunpeng Bai, Xinwen Hou, Yu Liu 0078
ICME3
2021 Learning to Coordinate via Multiple Graph Neural Networks
Zhiwei Xu 0005, Bin Zhang 0052, Yunpeng Bai, Dapeng Li 0001
ICONIP (3)3
2021 MMD-MIX: Value Function Factorisation with Maximum Mean Discrepancy for Cooperative Multi-Agent Reinforcement Learning
abstract
In the real world, many tasks require multiple agents to cooperate with each other under the condition of local observations. To solve such problems, many multi-agent reinforcement learning methods based on Centralized Training with Decentralized Execution have been proposed. One representative class of work is value decomposition, which decomposes the global joint Q-value Qjtinto individual Q-values Qato guide individuals' behaviors, e.g. VDN (Value-Decomposition Networks) and QMIX. However, these baselines often ignore the randomness in the situation. We propose MMD-MIX, a method that combines distributional reinforcement learning and value decomposition to alleviate the above weaknesses. Besides, to improve data sampling efficiency, we were inspired by REM (Random Ensemble Mixture) which is a robust RL algorithm to explicitly introduce randomness into the MMD-MIX. The experiments demonstrate that MMD-MIX outperforms prior baselines in the StarCraft Multi-Agent Challenge (SMAC) environment.
Zhiwei Xu 0005, Dapeng Li 0001, Yunpeng Bai
IJCNN3
2021 Hyperspectral Image Classification With Spatial Consistence Using Fully Convolutional Spatial Propagation Network
abstract
In recent years, deep convolutional neural networks (CNNs) have demonstrated impressive ability to represent hyperspectral images (HSIs) and achieved encouraging results in HSI classification. However, the existing CNN-based models operate at the patch level, in which a pixel is separately classified into classes using a patch of images around it. This patch-level classification will lead to a large number of repeated calculations, and it is hard to identify the appropriate patch size that is beneficial to classification accuracy. In addition, the conventional CNN models operate convolutions with local receptive fields, which cause the failure of contextual spatial information modeling. To overcome these aforementioned limitations, we propose a novel end-to-end, pixel-to-pixel, fully convolutional spatial propagation network (FCSPN) for HSI classification. Our FCSPN consists of a 3-D fully convolution network (3D-FCN) and a convolutional spatial propagation network (CSPN). Specifically, the 3D-FCN is first introduced for reliable preliminary classification, in which a novel dual separable residual (DSR) unit is proposed to effectively capture spectral and spatial information simultaneously with fewer parameters. Moreover, the channel-wise attention mechanism is adapted in the 3D-FCN to grasp the most informative channels from redundant channel information. Finally, the CSPN is introduced to capture the spatial correlations of HSIs via learning a local linear spatial propagation, which allows maintaining the HSI spatial consistency and further refining the classification results. Experimental results on three HSI benchmark data sets demonstrate that the proposed FCSPN achieves state-of-the-art performance on HSI classification.
Yenan Jiang, Ying Li 0017, Shanrong Zou, Haokui Zhang, Yunpeng Bai
IEEE Trans. Geosci. Remote. Sens.5
2020 Stable Training of Bellman Error in Reinforcement Learning
Chen Gong 0005, Yunpeng Bai, Xinwen Hou, Xiaohui Ji
ICONIP (5)2