Fanyang Meng

dblp:166/6774 · DBLP profile ↗
← Back
56ranked-venue papers
4as first author
45since 2021 · last 2026
0000-0001-5725-2178ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 4 first-author · 33 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Learned-PPR Decoding Scheme for Partial Packet Recovery in Network Coding
abstract
Network coding (NC) has proven to offer significant benefits in long-distance and broadcast transmissions, enhancing both throughput and energy efficiency. Recent studies have incorporated partial packet recovery (PPR) into packet-level NC, using syndromes from coded packets to correct bit errors and thereby reduce completion delay. Motivated by recent breakthroughs in deep learning, this paper introduces a novel neural networkbased decoding framework for packet-level NC, referred to as Learned-PPR. The proposed framework incorporates a Bilateral Efficient Self-Attention Network (Bi-ESANet) architecture, which leverages a bilateral network structure to effectively capture both inter- and intra-packet information. Furthermore, we introduce an ESA module to mitigate the GPU memory overhead compared with traditional Transformer attention modules. To handle rateless NC, we propose a “rateless masking” training strategy that enables efficient decoding of rateless codes within the Bi-ESANet framework. Simulation results across various transmission scenarios demonstrate that the proposed approach significantly outperforms existing PPR schemes, achieving lower completion delay. Specifically, compared to existing methods, the proposed approach reduces completion delay by more than 25%. However, the introduced framework incurs higher computational complexity due to the integration of the Bi-ESANet architecture.
Qifu Tyler Sun, Zongpeng Li, Yangxuan Cheng, Fanyang Meng, Ye Wang 0002, Yongsheng Liang 0001
IEEE Internet Things J.6
2026 Turbo principles meet compression: Rethinking nonlinear transformations in learned image compression
Chao Li 0071, Wen Tan 0001, Fanyang Meng, Runwei Ding, Ye Wang 0002, Wei Liu 0065, Yongsheng Liang 0001
J. Vis. Commun. Image Represent.3
2026 Hierarchical quality-aware guidance for blind JPEG artifacts removal
Shuai Liu 0022, Qingyu Mao, Binqiang Liu, Fanyang Meng, Shuangyan Yi, Yongsheng Liang 0001
J. Vis. Commun. Image Represent.5
2026 DBML-Font :Double-branch multi-level feature fusion based on diffusion model for few-shot font generation
Yueyue Fang, Haipeng Xiao, Wenyi Zhou, Lixin Guan, Mengshan Li, Fanyang Meng, Jihong Zhu 0003
Neural Networks6
2026 Entropy-aware image representation via 2D Gaussian splatting
Jiacong Chen, Qingyu Mao, Shuai Liu 0022, Chao Li 0071, Jierun Lin, Xiandong Meng, Fanyang Meng, Yongsheng Liang 0001
Signal Process.7
2026 Blind JPEG Artifacts Removal via Inverse JPEG Compression
abstract
Quantization and chroma downsampling are two primary operations that introduce distortions in the JPEG compression. However, most existing blind methods treat artifacts removal as a direct mapping from compressed images to clean ones. They fail to explicitly model the underlying degradation process or design targeted compensation mechanisms. As a result, these methods can only partially remove compression artifacts and struggle to generalize to diverse or unseen degradation scenarios. In this work, we present a novel perspective that formulates artifacts removal as an approximate inversion of the lossy steps in JPEG. Based on this view, we propose an Inverse JPEG Compression Network (IJCN), which aims to progressively compensate for quantization errors and color distortions. Specifically, we first design a Learnable Offset Guidance Module (LOGM) to approximate inverse quantization by modeling both intra-block and inter-block coefficient correlations for predicting rounding offsets. In addition, we propose a Quantization Table Guidance Module (QTGM) that leverages the quantization tables to guide the reconstruction network in mitigating color distortions. By modeling compensation mechanisms under the guidance of quantization tables, IJCN effectively eliminates artifacts across varying compression levels. Extensive experiments demonstrate that IJCN outperforms existing methods in both quantitative metrics and visual quality.
Shuai Liu 0022, Binqiang Liu, Qingyu Mao, Jiacong Chen, Fanyang Meng, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Superimposed Pilot-Based Adaptive Semantic Communications for Wireless Image Transmission
abstract
The non-orthogonal superimposed pilot (NOSIP) scheme significantly improves the spectrum efficiency of semantic communication (SemCom) systems by effectively reducing pilot overhead. However, existing SemCom systems based on NOSIP still face challenges, including strong model-channel coupling and limited adaptability to heterogeneous channels. To address these issues, this paper proposes a flexible, channel-adaptive digital SemCom (D-SemCom) architecture based on the NOSIP scheme. Specifically, we design a lightweight semantic codec, termed ShiftViT, and a semantic receiver, termed ShiftRx, which employ time- and frequency-domain shift mechanisms to decouple pilot and data and suppress multi-user interference, thereby enabling image transmission under complex channel conditions. Furthermore, a lightweight channel adaptation algorithm based on first-order meta-learning is proposed to facilitate rapid adaptation and mitigate the strong coupling between semantic models and channel environments. Numerical results demonstrate that the proposed D-SemCom system achieves approximately 25.14% and 1.16% goodput improvements over traditional receivers and existing state-of-the-art methods, respectively, while reducing the computational complexity in terms of FLOPs by approximately 40.37%. In addition, the proposed channel adaptation algorithm is shown to rapidly adapt to diverse channel scenarios within the D-SemCom system.
Jian Xiao 0003, Wenwu Xie, Fanyang Meng, Renhai Feng, Liang Yang 0001, Yongsheng Liang 0001
IEEE Trans. Wirel. Commun.4
2025 Adaptive Modulation Inference via Input Skipping and Budget-Efficient Exiting
abstract
Automatic Modulation Recognition (AMR) is essential for efficient spectrum utilization, cognitive radio, and secure wireless communications. However, deploying accurate AMR models on resource-constrained devices remains challenging due to substantial computational overhead. Importantly, minimizing inference computational cost—distinct from conventional neural network lightweighting—is critical for meeting strict runtime constraints on resource-limited platforms. This paper proposes the Adaptive Modulation Inference (AMI) framework, a dynamic inference solution for efficient and adaptive AMR on limitedresource platforms such as satellites and unmanned aerial vehicles (UAVs). AMI integrates Adaptive Input Skipping (AIS) and Budget-Efficient Exiting (BEE) to dynamically tailor computation based on signal difficulty and real-time resource budgets. AIS employs Layer and Channel Gates for coarse-grained skipping and fine-grained pruning, while BEE adjusts early-exit thresholds based on entropy and Top-1 confidence. Implemented on a lightweight 1D MobileNetV2 backbone, AMI achieves up to 56% reduction in average computational cost with less than 1% accuracy loss on both RML22 and HisarMod2019.1 datasets, outperforming existing dynamic inference strategies applied in the AMR domain.
Kehan Xiang, Xingjian Zhang 0001, Xiqiao Zheng, Fanyang Meng, Qinyu Zhang 0001
GLOBECOM5
2025 Robust Deep Joint Source-Channel Coding for Video Transmission over Multipath Fading Channel
abstract
To address the challenges of wireless video transmission over multipath fading channels, we propose a robust deep joint source-channel coding (DeepJSCC) framework by effectively exploiting temporal redundancy and incorporating robust innovations at the modulation, coding, and decoding stages. At the modulation stage, tailored orthogonal frequency division multiplexing (OFDM) for robust video transmission is employed, decomposing wideband signals into orthogonal frequency-flat sub-channels to effectively mitigate frequency-selective fading. At the coding stage, conditional contextual coding with multi-scale Gaussian warped features is introduced to efficiently model temporal redundancy, significantly improving reconstruction quality under strict bandwidth constraints. At the decoding stage, a lightweight denoising module is integrated to robustly simplify signal restoration and accelerate convergence, addressing the suboptimality and slow convergence typically associated with simultaneously performing channel estimation, equalization, and semantic reconstruction. Experimental results demonstrate that the proposed robust framework significantly outperforms state-of-the-art video DeepJSCC methods, which achieves an average reconstruction quality gain of 5.13 dB under challenging multipath fading channel conditions1.
Bohuai Xiao, Fanyang Meng, Wei Liu 0065, Yongsheng Liang 0001
GLOBECOM3
2025 Multimodal-Guided Perceptual Image Compression via Joint Text and Audio
Genhong Wang, Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
ICIC (3)4
2025 DMSO: A Dynamic Momentum-Smoothing Optimizer for Learned Image Compression
abstract
Learned Image Compression (LIC) has rapidly evolved and recently surpassed traditional methods in Rate-Distortion (R-D) performance. However, most LIC approaches improve network architectures while increasing computational overhead and overlooking using optimizers tailored specifically for LIC. This paper proposes a Dynamic Momentum-Smoothing Optimizer (DMSO) tailored for LIC to achieve faster convergence and better R-D performance. Specifically, DMSO leverages historical gradient information to smooth the optimization process dynamically, thereby reducing in-stability from gradient oscillations. Furthermore, DMSO introduces a novel Enhanced Second-order Momentum mechanism to mitigate cumulative noise and align momentum updates more closely with the true gradient. Experimental results demonstrate that DMSO operates as a universal optimization plugin for LIC methods, achieving faster and more stable convergence while improving R-D performance to varying degrees without additional parameter count or computational cost.
Chao Li 0071, Chuanmin Jia, Fanyang Meng, Siwei Ma 0001, Yongsheng Liang 0001
ICIP3
2025 Towards Robust Text-Guided Image Compression Under Modality Missing
abstract
Text-guided image compression aims to enhance the perceptual quality of reconstructed images by leveraging textual semantic information. However, existing methods struggle to effectively integrate text information and suffer from significant performance degradation when text is unavailable. To address these issues, we propose a Robust Text-Guided Image Compression (RobustTGIC) network that fully utilizes text semantics when available and mitigates performance loss when absent. Specifically, we introduce a Dual-Dimensional Text Modulation (DDTM) module to enhance perceptual quality by accurately fusing textual information. Building on this, we further propose an Intermediate Feature Modulation (IFM) module, which compensates for missing semantics through lightweight adaptation, improving robustness. Experimental results indicate that, at low bitrates (e.g., 0.07 bpp), our method achieves superior perceptual quality reconstruction while significantly reducing bitrates (e.g., 0.5× HiFiC and 0.4× Bpg). Moreover, our method effectively mitigates performance degradation caused by missing text with a parameter increase of less than 1% of the total parameters.
Genhong Wang, Wen Tan 0001, Fanyang Meng, Yongsheng Liang 0001
ICIP3
2025 Grouped Transform for Ultra-Low-Complexity Learned Image Compression
abstract
Existing learned image compression (LIC) methods have shown strong performance advantages but also bring high computational complexity, making it challenging to deploy them on resource-constrained devices. To reduce the high computational and storage cost, we propose a fully grouped image compression network by introducing spatial and channel grouping operations. Grouping operation is helpful to obtain compact representations by aggregating similar features and reducing redundancy between features in LIC task. Specifically, our proposed network consists of two efficient parts, one is the spatial grouping transform for spatial resolution sampling, and the other is the channel grouping transform for nonlinear representation capability enhancement. Moreover, convolutional kernel factorization and inverted bottleneck are used to reduce redundancy and enrich information of each group in the channel grouping transform, which achieve a good balance between computational complexity and network performance. Experimental results show that our method not only achieves competitive rate-distortion performance with fewer KMACs/pixel and model parameters, but also reduces the real-world runtime. In particular, our proposed models provide at least over 84.6% computational complexity reduction when compared with several advanced LIC methods.
Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
ISCAS3
2025 Structured Sparsity Learning for Efficient Learned Image Compression
abstract
Existing learned image compression (LIC) methods have achieved outstanding performance, but their deployment on resource-constrained devices is hindered by the high computational complexity and large model storage. Sparsity learning can obtain sparse neural networks by applying regularization term and further achieve model compression by pruning. However, it is difficult to achieve a lightweight LIC network by directly applying sparsity learning and pruning due to unstructured sparsity and limitations of entropy model. In this paper, we propose to add L2,1regularization during the network training for image compression task, which generates structured sparsity at both filter and channel level. We further analyze the effect of entropy model capacity, and adopt filter/channel fixing to achieve the alignment of entropy estimation for actual pruning. Moreover, we utilize incremental regularization to improve the sparsity of network and training stability. Experimental results show that our pruned lightweight model can effectively reduce network parameters by an average of 59.73% at the cost of 1.57% BD-rate increase compared with original hyperprior model.
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Lihan Zhu, Yongsheng Liang 0001
ISCAS3
2025 Motion Matters: Compact Gaussian Streaming for Free-Viewpoint Video Reconstruction
abstract
3D Gaussian Splatting (3DGS) has emerged as a high-fidelity and efficient paradigm for online free-viewpoint video (FVV) reconstruction, offering viewers rapid responsiveness and immersive experiences. However, existing online methods face challenge in prohibitive storage requirements primarily due to point-wise modeling that fails to exploit the motion properties. To address this limitation, we propose a novel Compact Gaussian Streaming (ComGS) framework, leveraging the locality and consistency of motion in dynamic scene, that models object-consistent Gaussian point motion through keypoint-driven motion representation. By transmitting only the keypoint attributes, this framework provides a more storage-efficient solution. Specifically, we first identify a sparse set of motion-sensitive keypoints localized within motion regions using a viewspace gradient difference strategy. Equipped with these keypoints, we propose an adaptive motion-driven mechanism that predicts a spatial influence field for propagating keypoint motion to neighboring Gaussian points with similar motion. Moreover, ComGS adopts an error-aware correction strategy for key frame reconstruction that selectively refines erroneous regions and mitigates error accumulation without unnecessary overhead. Overall, ComGS achieves a remarkable storage reduction of over 159 × compared to 3DGStream and 14 × compared to the SOTA method QUEEN, while maintaining competitive visual fidelity and rendering speed. Project page: https://chenjiacong-1005.github.io/ComGS/.
Jiacong Chen, Qingyu Mao, Youneng Bao, Xiandong Meng, Fanyang Meng, Ronggang Wang, Yongsheng Liang 0001
NeurIPS5
2025 Progressive Diffusion-Based Low Rate Perceptual Image Compression with Discrete Gaussian Codebooks for Remote Sensing
Yangxuan Cheng, Fanyang Meng, Runwei Ding, Ye Wang 0002, Yongsheng Liang 0001
PRCV (9)2
2025 Boosting Neural Video Representation via Online Structural Reparameterization
Qingyu Mao, Shuai Liu 0022, Qilei Li, Fanyang Meng, Yongsheng Liang 0001
PRCV (6)5
2025 Stable successive Neural Image Compression via coherent demodulation-based transformation
Youneng Bao, Wen Tan 0001, Mu Li 0005, Fanyang Meng, Yongsheng Liang 0001
Signal Process.4
2025 Adaptive cross-channel transformation based on self-modulation for learned image compression
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Yongsheng Liang 0001
Signal Process. Image Commun.3
2025 Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
abstract
Recognizing interactive actions, including hand-to-hand interaction and human-to-human interaction, has attracted increasing attention for various applications in the field of video analysis and human–robot interaction. Considering the success of graph convolution in modeling topology-aware features from skeleton data, recent methods commonly operate graph convolution on separate entities and use late fusion for interactive action recognition, which can barely model the mutual semantic relationships between pairwise entities. To this end, we propose a mutual excitation graph convolutional network (me-GCN) by stacking mutual excitation graph convolution (me-GC) layers. Specifically, me-GC uses a mutual topology excitation module to firstly extract adjacency matrices from individual entities and then adaptively model the mutual constraints between them. Moreover, me-GC extends the above idea and further uses a mutual feature excitation module to extract and merge deep features from pairwise entities. Compared with graph convolution, our proposed me-GC gradually learns mutual information in each layer and each stage of graph convolution operations. Extensive experiments on a challenging hand-to-hand interaction dataset, i.e., the Assembely101 dataset, and two large-scale human-to-human interaction datasets, i.e., NTU60-Interaction and NTU120-Interaction consistently verify the superiority of our proposed method, which outperforms the state-of-the-art GCN-based and Transformer-based methods.
Mengyuan Liu 0001, Chen Chen 0015, Songtao Wu, Fanyang Meng, Hong Liu 0008
IEEE Trans. Hum. Mach. Syst.4
2025 One is All: A Unified Rate-Distortion-Complexity Framework for Learned Image Compression Under Energy Concentration Criteria
abstract
The learned image compression (LIC) technique has surpassed the state-of-the-art traditional codecs (H.266/VVC) in case of rate-distortion (R-D) performance. Its real-time deployments are far advanced. In order to achieve more flexible deployments, an LIC technique should be flexible in adjusting its computational complexity and rate as demanded by a situation and its environment. In this paper, we propose a unified Rate-Distortion-Complexity (R-D-C) framework for LIC under channel energy concentration criteria. Specifically, we first introduce an Energy Asymptotic Nonlinear Transformation (EANT) designed to directly concentrate on the channel energy of latent representations, thus laying the groundwork for a scalable entropy coding. Next, leveraging this energy concentration characteristic, we propose a corresponding Heterogeneous Scalable Entropy Model (HSEM) for flexibly scaling bitstreams as needed. Finally, utilizing the proposed EANT, we construct a fine-grained scalable codec for formulating, in combination with HSEM, a comprehensive scalable R-D-C framework under the energy concentration criteria. The obtained experimental results demonstrate that the proposed method could enable seamless transitions between 13 different widths of sub-models within a single network, allowing for fine-grained control over the model bitrate, complexity, and hardware inference time. Additionally, the proposed method exhibits competitive R-D performance compared to many existing methods.
Chao Li 0071, Fanyang Meng, Qingyu Mao, Youneng Bao, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Multim.3
2024 Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image Compression
abstract
Deep neural network-based image compression (NIC) has achieved excellent performance, but NIC method models have been shown to be susceptible to backdoor attacks. Adversarial training has been validated in image compression models as a common method to enhance model robustness. However, the improvement effect of adversarial training on model robustness is limited. In this paper, we propose a prior knowledge-guided adversarial training framework for image compression models. Specifically, first, we propose a gradient regularization constraint for training robust teacher models. Subsequently, we design a knowledge distillation-based strategy to generate a priori knowledge from the teacher model to the student model for guiding adversarial training. Experimental results show that our method improves the reconstruction quality by about 9dB when the Kodak dataset is elected as the backdoor attack object for psnr attack. Compared with Ma2023 [1], our method has a 5dB higher PSNR output at high bitrate points.
Youneng Bao, Fanyang Meng, Chao Li 0071, Wen Tan 0001, Genhong Wang, Yongsheng Liang 0001
ICASSP3
2024 Leveraging Redundancy in Feature for Efficient Learned Image Compression
abstract
In recent years, with the development of the field of learned image compression, numerous models with excellent rate-distortion performance have emerged. However, the considerable computational complexity inherent in these models poses challenges for their practical deployment. In this paper, we investigate feature redundancy in learned image compression (LIC) algorithms for efficient feature extraction and introduce an efficient and lightweight LIC framework. Specifically, we explore the existence of a large number of similar features in the network. Subsequently, we design effective feature extraction modules across various levels, such as layer and block. In addition, based on the fact that the role of the codec’s encoder is to remove redundancy and the decoder is to reconstruct, we propose an asynchronous feature fusion block. This fusion block incorporates an "edge smoothing" operator in the encoder and an "edge enhancement" operator in the decoder. Our methodology strikes an ideal balance between rate-distortion performance and efficiency. The experimental results indicate that our approach necessitates only 310KMac/pixel computation and 9.5M parameters, while in terms of performance, our method achieves a 20.7% BD-rate advantage over BPG on Kodak data, mirroring VVC’s performance. Compared to other learned image compression algorithms with SOTA performance, our method has a great advantage in terms of computation/parameter count.
Youneng Bao, Fanyang Meng, Wen Tan 0001, Chao Li 0071, Genhong Wang, Yongsheng Liang 0001
ICASSP3
2024 Enhanced Interpretability in Learned Image Compression via Convolutional Sparse Coding
abstract
Compared to traditional image compression methods, learned image compression (LIC) methods have demonstrated increasingly superior rate-distortion performance. However, LIC networks are often regarded as black boxes, still lacking a theoretical understanding. Sparse coding provides the sparse and interpretable modeling for analyzing or synthesizing natural images in various signal and image processing applications. Therefore, we introduce convolutional sparse coding (CSC) into transform network for enhancing the interpretability of LIC methods. In this paper, we first employ CSC layers to achieve certain theoretical modeling for LIC network, and adopt a weight sharing strategy in encoder-decoder pair and attention mechanism to balance the complexity and performance. Additionally, we analyze the model robustness against data input perturbations and consider the impact of sparsity trade-off parameter in the CSC layer optimization process. Experimental results demonstrate that our method achieves comparable performance with the corresponding baseline, and our model is more robust.
Yiwen Tu, Wen Tan 0001, Youneng Bao, Genhong Wang, Fanyang Meng, Yongsheng Liang 0001
ICME5
2024 Fine-Grained Adjustable Entropy Models for Rate-Complexity Jointly Adjustable Image Compression
Chao Li 0071, Shanzhi Yin, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
PRCV (9)5
2024 Multirate Progressive Entropy Model for Learned Image Compression
abstract
This paper proposes a unified and efficient entropy coding method for learned image compression (LIC) from the perspective of traditional signal processing. First, the consistency of structures and optimization objectives are used to interpret the existing split-coded-then-merge entropy coding strategies in LIC as a particular filter banks framework, with feature separation and feature aggregation representing the analysis filter bank and synthesis filter bank, respectively. Thus, we borrow the design from the multirate filter banks and proposed Multirate Progressive Entropy Model (MPEM) to enhance the rate-distortion performance and decoding speed. In particular, we create an analysis filter bank that divides compact features into a few nonuniform subsets based on various spatial and channel sampling rates. Then multi-scale detail and mean coefficients within the current subset are used as prior representations to help generate the prediction parameters of the next subset, and the carefully designed synthetic filter bank performs a near-perfect reconstruction of the features. In addition, we propose a Multi-level Edge Attention Moudal (MEAM) to increase the edge and texture information’s contribution and reduce the high-frequency information loss brought on by MPEM’s inherent multi-rate spatial sampling, which leverages the edge operator and structural reparameterization principles. The results of the experiments show that, in comparison to the effective LIC methods and traditional code, the proposed MPEM can decode data at a cutting-edge speed while also offering comparable rate-distortion performance.
Chao Li 0071, Shanzhi Yin, Chuanmin Jia, Fanyang Meng, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 BCAN: Bidirectional Correct Attention Network for Cross-Modal Retrieval
abstract
As a fundamental topic in bridging the gap between vision and language, cross-modal retrieval purposes to obtain the correspondences' relationship between fragments, i.e., subregions in images and words in texts. Compared with earlier methods that focus on learning the visual semantic embedding from images and sentences to the shared embedding space, the existing methods tend to learn the correspondences between words and regions via cross-modal attention. However, such attention-based approaches invariably result in semantic misalignment between subfragments for two reasons: 1) without modeling the relationship between subfragments and the semantics of the entire images or sentences, it will be hard for such approaches to distinguish images or sentences with multiple same semantic fragments and 2) such approaches focus attention evenly on all subfragments, including nonvisual words and a lot of redundant regions, which also will face the problem of semantic misalignment. To solve these problems, this article proposes a bidirectional correct attention network (BCAN), which introduces a novel concept of the relevance between subfragments and the semantics of the entire images or sentences and designs a novel correct attention mechanism by modeling the local and global similarity between images and sentences to correct the attention weights focused on the wrong fragments. Specifically, we introduce a concept about the semantic relationship between subfragments and entire images or sentences and use this concept to solve the semantic misalignment from two aspects. In our correct attention mechanism, we design two independent units to correct the weight of attention focused on the wrong fragments. Global correct unit (GCU) with modeling the global similarity between images and sentences into the attention mechanism to solve the semantic misalignment problem caused by focusing attention on relevant subfragments in irrelevant pairs (RI) and the local correct unit (LCU) consider the difference in the attention weights between fragments among two steps to solve the semantic misalignment problem caused by focusing attention on irrelevant subfragments in relevant pairs (IR). Extensive experiments on large-scale MS-COCO and Flickr30K show that our proposed method outperforms all the attention-based methods and is competitive to the state-of-the-art. Our code and pretrained model are publicly available at: https://github.com/liuyyy111/BCAN.
Yang Liu 0264, Hong Liu 0008, Huaqiu Wang, Fanyang Meng, Mengyuan Liu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Cloth Interactive Transformer for Virtual Try-On
abstract
The 2D image-based virtual try-on has aroused increased interest from the multimedia and computer vision fields due to its enormous commercial value. Nevertheless, most existing image-based virtual try-on approaches directly combine the person-identity representation and the in-shop clothing items without taking their mutual correlations into consideration. Moreover, these methods are commonly established on pure convolutional neural networks (CNNs) architectures which are not simple to capture the long-range correlations among the input pixels. As a result, it generally results in inconsistent results. To alleviate these issues, in this article, we propose a novel two-stage cloth interactive transformer (CIT) method for the virtual try-on task. During the first stage, we design a CIT matching block, aiming at precisely capturing the long-range correlations between the cloth-agnostic person information and the in-shop cloth information. Consequently, it makes the warped in-shop clothing items look more natural in appearance. In the second stage, we put forth a CIT reasoning block for establishing global mutual interactive dependencies among person representation, the warped clothing item, and the corresponding warped cloth mask. The empirical results, based on mutual dependencies, demonstrate that the final try-on results are more realistic. Substantial empirical results on a public fashion dataset illustrate that the suggested CIT attains competitive virtual try-on performance.
Bin Ren 0005, Hao Tang 0005, Fanyang Meng, Runwei Ding, Philip Torr 0001, Nicu Sebe
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Novel Motion Patterns Matter for Practical Skeleton-Based Action Recognition
abstract
Most skeleton-based action recognition methods assume that the same type of action samples in the training set and the test set share similar motion patterns. However, action samples in real scenarios usually contain novel motion patterns which are not involved in the training set. As it is laborious to collect sufficient training samples to enumerate various types of novel motion patterns, this paper presents a practical skeleton-based action recognition task where the training set contains common motion patterns of action samples and the test set contains action samples that suffer from novel motion patterns. For this task, we present a Mask Graph Convolutional Network (Mask-GCN) to focus on learning action-specific skeleton joints that mainly convey action information meanwhile masking action-agnostic skeleton joints that convey rare action information and suffer more from novel motion patterns. Specifically, we design a policy network to learn layer-wise body masks to construct masked adjacency matrices, which guide a GCN-based backbone to learn stable yet informative action features from dynamic graph structure. Extensive experiments on our newly collected dataset verify that Mask-GCN outperforms most GCN-based methods when testing with various novel motion patterns.
Mengyuan Liu 0001, Fanyang Meng, Chen Chen 0001, Songtao Wu
AAAI2
2023 Multi-Stream Facial Adaptive Network for Expression Recognition from a Single Image
abstract
Facial expression recognition from a single image has potential applications in fields including human-computer interaction and medical diagnosis. Most recent methods use deep neural networks to directly learn from a roughly cropped facial image which is usually detected from a whole image by face detection algorithms. We observe that unrelated surrounding regions in the rough facial image prevent deep neural networks from learning facial-related discriminate features. To solve this problem, we present a Facial Adaptive Network (FAN) which is able to adaptively select an interest region from the given facial image, thus suffering less from the effect of unrelated regions. Based on the selected interest region, we further apply the self-attention mechanism to learn discriminate facial features. Moreover, we introduce a multi-stream FAN (ms-FAN) that learns richer facial features from multiple interest regions that are selected from pose-augmented facial images. Extensive experiments on Oulu-CASIA, CK+, and RAF-DB datasets consistently verify the effect of our proposed MS-FAN by achieving comparable results with state-of-the-art methods. Our code is available at https://github.com/zhangbc12/DAtt-ViT.
Baichuan Zhang, Fanyang Meng, Runwei Ding, Mengyuan Liu 0001
ICASSP2
2023 A Decoupled Spatial-Channel Inverted Bottleneck For Image Compression
abstract
Residual block has achieved great success in deep networks to eliminate accuracy degradation, and there emerges a large number of variants with more competitive performance. However, these blocks are introduced for high-level tasks that only encode the input image into semantic and struc¬tural features but do not need to reconstruct. So for the low-level task like image compression where reconstruction quality contributes significantly to the rate-distortion perfor¬mance, the structure of the residual block needs modification for more suitable implementation. In this paper, we revisit the existing residual blocks and discover two key principles summarized as two decouplings: spatial-channel decoupling and linear-nonlinear decoupling. We propose an efficient nonlinear transform based on the principles dubbed decou¬pled spatial-channe inverted bottleneck(DSCIB), which has a linear-spatial branch for rough reconstruction and a nonlinear¬channel branch to provide detailed featrues. We employ the DSCIB module in the joint auto regression model to build an overall network. Experimental results show that our method achieves comparable performance with the existing learning-based image compression methods at high bitrate while re¬ducing 38% FLOPs.
Wen Tan 0001, Fanyang Meng, Yongsheng Liang 0001
ICIP3
2023 Dynamic Compositional Graph Convolutional Network for Efficient Composite Human Motion Prediction
abstract
With potential applications in fields including intelligent surveillance and human-robot interaction, the human motion prediction task has become a hot research topic and also has achieved high success, especially using the recent Graph Convolutional Network (GCN). Current human motion prediction task usually focuses on predicting human motions for atomic actions. Observing that atomic actions can happen at the same time and thus formulating the composite actions, we propose the composite human motion prediction task. To handle this task, we first present a Composite Action Generation (CAG) module to generate synthetic composite actions for training, thus avoiding the laborious work of collecting composite action samples. Moreover, we alleviate the effect of composite actions on demand for a more complicated model by presenting a Dynamic Compositional Graph Convolutional Network (DC-GCN). Extensive experiments on the Human3.6M dataset and our newly collected CHAMP dataset consistently verify the efficiency of our DC-GCN method, which achieves state-of-the-art motion prediction accuracies and meanwhile needs few extra computational costs than traditional GCN-based human motion methods.
Fanyang Meng, Songtao Wu, Mengyuan Liu 0001
ACM Multimedia3
2023 A Complex-Valued Neural Network Based Robust Image Compression
Can Luo, Youneng Bao, Wen Tan 0001, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001
PRCV (10)5
2023 Taylor series based dual-branch transformation for learned image compression
Youneng Bao, Wen Tan 0001, Linfeng Zheng, Fanyang Meng, Wei Liu 0065, Yongsheng Liang 0001
Signal Process.4
2023 Nonlinear Transforms in Learned Image Compression From a Communication Perspective
abstract
Recently, remarkable progress has been made in learned image compression (LIC), in which nonlinear transforms (NTs) play a crucial role. Although there are many NT methods for improving the rate distortion performance, all the existing methods sacrifice the computational complexity and the number of parameters of the transformation. This paper provides a fundamental novel viewpoint on nonlinear transforms from a communication perspective, and shows how this idea can be extended to design efficient NT methods. In particular, the nonlinear transforms are inferred as signal modulation modules. Under this extrapolation, the current NTs are generalized as amplitude modulation that only varies the amplitude of the carrier wave. Therefore, a nonlinear modulation-like transform (NMLT) which varies the phase angle of the carrier is proposed. Moreover, this concept is extended by introducing In-phase/Quadrature (IQ) modulation, which is a boosting technique in communication field, in order to enhance NMLT. Furthermore, the Bit-interleaved technique in communication is used to guide the optimization of NTML with IQ. The experimental results on different datasets and backbone architectures verify the efficiency and robustness of the proposed methods. For example, when backbone architecture is hyperprior model, our method achieves 19.37% BD-rate reduction over GDN on the Kodak dataset. In addition, our method with channel wise autoregressive model leads to the state-of-the-art rate-distortion performance.
Youneng Bao, Fanyang Meng, Chao Li 0071, Siwei Ma 0001, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 AdderIC: Towards Low Computation Cost Image Compression
abstract
Recently, learned image compression methods have shown their outstanding rate-distortion performance when compared to traditional frameworks. Although numerous progress has been made in learned image compression, the computation cost is still at a high level. To address this problem, we propose AdderIC, which utilizes adder neural networks (AdderNet) to construct an image compression framework. According to the characteristics of image compression, we introduce several strategies to improve the performance of AdderNet in this field. Specifically, Haar Wavelet Transform is adopted to make AdderIC learn high-frequency information efficiently. In addition, implicit deconvolution with the kernel size of 1 is applied after each adder layer to reduce spatial redundancies. Moreover, we develop a novel Adder-ID-PixelShuffle cascade upsampling structure to remove checkerboard artifacts. Experiments demonstrate that our AdderIC model can largely outperform conventional AdderNet when applied in image compression and achieve comparable rate-distortion performance to that of its CNN baseline with about 80% multiplication FLOPs and 30% energy consumption reduction.
Xin Yao 0001, Chao Li 0071, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
ICASSP5
2022 Universal Efficient Variable-Rate Neural Image Compression
abstract
Recently, Learning-based image compression has reached comparable performance with traditional image codecs(such as JPEG, BPG, WebP). However, computational complexity and rate flexibility are still two major challenges for its practical deployment. To tackle these problems, this paper proposes two universal modules named Energy-based Channel Gating(ECG) and Bit-rate Modulator(BM), which can be directly embedded into existing end-to-end image compression models. ECG uses dynamic pruning to reduce FLOPs for more than 50% in convolution layers, and a BM pair can modulate the latent representation to control the bit-rate in a channel-wise manner. By implementing these two modules, existing learning-based image codecs can obtain ability to output arbitrary bit-rate with a single model and reduced computation.
Shanzhi Yin, Chao Li 0071, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng, Wei Liu 0065
ICASSP5
2022 Exploring Structural Sparsity in Neural Image Compression
abstract
The performance of neural image compression have reached or suppressed traditional methods (such as JPEG, BPG, WebP). However, their sophisticated network structures with cascaded convolution layers bring heavy computational burden for practical deployment. In this paper, we explore structural sparsity in neural image compression network to obtain real-time acceleration without any specialized hardware design or algorithm. We propose a simple plug-in adaptive binary channel masking(ABCM) to judge the importance of each convolution channel and introduce sparsity during training. During inference, the unimportant channels are pruned to obtain slimmer network and less computation. We implement our method into three neural image compression networks with different entropy models to verify its effectiveness and generalization, the experiment results show that up to 7× computation reduction and 3× acceleration can be achieved with negligible performance drop.
Shanzhi Yin, Chao Li 0071, Fanyang Meng, Wen Tan 0001, Youneng Bao, Yongsheng Liang 0001, Wei Liu 0065
ICIP3
2022 STCDesc: Learning deep local descriptor using similar triangle constraint
Qiao Liu 0001, Fanyang Meng, Zhenyu He 0001
Knowl. Based Syst.3
2022 Spatial-Temporal Asynchronous Normalization for Unsupervised 3D Action Representation Learning
abstract
Unsupervised 3D action representation learning from skeleton sequences has attracted increasing attention in recent years. Existing methods have successfully applied autoencoder network to learn 3D action representation by reconstructing original skeleton sequence. However, these methods ignore motion cues thus suffer from distinguishing actions especially with similar shape information and slightly different motion information. Instead of reconstructing original skeleton sequence, we learn distinctive 3D action representation with autoencoder network by reconstructing normalized motion sequence extracted from original input. To obtain the normalized motion sequence, we specifically design a novel spatial-temporal asynchronous normalization (STAN) method, which normalizes original skeleton sequence in two steps. First, STAN reduces redundant temporal information and extracts motion sequence by subtracting mean value along the temporal dimension. Second, STAN further normalizes the motion sequence along the spatial dimension and generates normalized motion sequence that suffers less from the effect of different human body shapes. Extensive experiments on large scale NTU RGB+D 60 and NTU RGB+D 120 datasets verify the effectiveness of our proposed STAN method, which achieves comparative results with state-of-the-art methods, and also outperforms alternative normalization methods.
Mengyuan Liu 0001, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng
IEEE Signal Process. Lett.4
2022 TCDesc: Learning Topology Consistent Descriptors for Image Matching
abstract
The triplet loss is widely used in learning the local descriptors for image matching. However, existing triplet loss-based methods, like HardNet and DSM, employ the point-to-point distance metric, which neglects the neighborhood information of descriptors. Considering the fact that local neighborhood structures of matching descriptors would be similar under the ideal condition, this paper aims to learn the neighborhood topology-consistent descriptors (TCDesc). To this end, we first propose the linear combination weight as the topology weight to depict the neighborhood topology for each descriptor, where the difference between the center descriptor and the linear combination of its neighbors is minimized. For the global comparison, we then define a global topology vector by using the local topology weights. Next, beyond the Euclidean distance, we define a topology distance with the topology vectors to indicate the topological difference between the matching descriptors. Furthermore, we propose an adaptive weighting strategy to jointly minimize the topology distance and Euclidean distance in triplet loss. Experimental results on four widely-used datasets, i.e., UBC PhotoTourism, HPatches, W1BS and Oxford, demonstrate that our method can effectively improve the performance of both HardNet and DSM.
Honghu Pan, Yongyong Chen, Zhenyu He 0001, Fanyang Meng, Nana Fan
IEEE Trans. Circuits Syst. Video Technol.4
2021 Mbb: A Multi-Scale Method For Data Based On Bit Plane Slicing
abstract
Multi-scale methodology can enhance the performance of the model in deep learning. The current multi-scale methodology focuses on changing the formation, which will increase the parameters and calculations of the network. This paper offers a multi-scale method for data based on bit plane slicing(MBB). This expands the receptive field of valid information in image data. It is done by multi-level fusing image with high bit planes. Our experimentation shows that by adding MBB in front of the backbone network, one can achieve a significant performance improvement. The MBB approach is widely applicable because it does not require changes to the structure of the backbone network.
Youneng Bao, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001, Wei Liu 0065, Kaiyu Liu
ICIP3
2021 Attend, Correct And Focus: A Bidirectional Correct Attention Network For Image-Text Matching
abstract
Image-text matching task aims to learn the fine-grained correspondences between images and sentences. Existing methods use attention mechanism to learn the correspondences by attending to all fragments without considering the relationship between fragments and global semantics, which inevitably lead to semantic misalignment among irrelevant fragments. To this end, we propose a Bidirectional Correct Attention Network (BCAN), which leverages global similarities and local similarities to reassign the attention weight, to avoid such semantic misalignment. Specifically, we introduce a global correct unit to correct the attention focused on relevant fragments in irrelevant semantics. A local correct unit is used to correct the attention focused on irrelevant fragments in relevant semantics. Experiments on Flickr30K and MSCOCO datasets verify the effectiveness of our proposed BCAN by outperforming both previous attention-based methods and state-of-the-art methods. Code can be found at: https://github.com/liuyyy111/BCAN.
Yang Liu 0264, Huaqiu Wang, Fanyang Meng, Mengyuan Liu 0001, Hong Liu 0008
ICIP3
2021 Improving Convolutional Networks with Boosting Attention Convolutions
abstract
Convolutional neural networks (CNNs) have been widely used in a range of tasks because of its robust convolutional feature transformation ability. In this paper, we propose a novel type of convolution called Boosting Attention Convolution (BAC) to improve the basic convolutional feature transformation process of CNNs. The proposed method is designed based on two principles, boosting and attention mechanism. Specifically, we design a set of simple yet effective Boosting Attention Modules (BAM) within grouped convolution, which progressively recalibrate distribution of feature map and enable the future filters nested in a convolution layer to focus more on the feature regions that are unactivated by previous filters. Thus, it can help CNNs generate more discriminative representations by explicitly incorporating richer information. The experimental results on various datasets verify that BAC outperforms state-of-the-art methods. More importantly, the proposed BAC is a general convolution that can be deployed to various modern networks without introducing much parameters and computational complexity.
Chao Li 0071, Yongsheng Liang 0001, Huo-Xiang Yang, Fanyang Meng, Wei Liu 0065, Handong Wang
ICME4
2021 Bi-Directional Exponential Angular Triplet Loss for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared person re-identification (RGB-IR Re-ID) is a cross-modality matching problem, where the modality discrepancy is a big challenge. Most existing works use Euclidean metric based constraints to resolve the discrepancy between features of images from different modalities. However, these methods are incapable of learning angularly discriminative feature embedding because Euclidean distance cannot measure the included angle between embedding vectors effectively. As an angularly discriminative feature space is important for classifying the human images based on their embedding vectors, in this paper, we propose a novel ranking loss function, named Bi-directional Exponential Angular Triplet Loss, to help learn an angularly separable common feature space by explicitly constraining the included angles between embedding vectors. Moreover, to help stabilize and learn the magnitudes of embedding vectors, we adopt a common space batch normalization layer. The quantitative and qualitative experiments on the SYSU-MM01 and RegDB dataset support our analysis. On SYSU-MM01 dataset, the performance is improved from 7.40% / 11.46% to 38.57% / 38.61% for rank-1 accuracy / mAP compared with the baseline. The proposed method can be generalized to the task of single-modality Re-ID and improves the rank-1 accuracy / mAP from 92.0% / 81.7% to 94.7% / 86.6% on the Market-1501 dataset, from 82.6% / 70.6% to 87.6% / 77.1% on the DukeMTMC-reID dataset.
Hanrong Ye, Hong Liu 0008, Fanyang Meng, Xia Li 0005
IEEE Trans. Image Process.3
2020 CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of Modality
abstract
Previous studies in multimodal sentiment analysis have used limited datasets, which only contain unified multimodal annotations.However, the unified annotations do not always reflect the independent sentiment of single modalities and limit the model to capture the difference between modalities.In this paper, we introduce a Chinese single-and multimodal sentiment analysis dataset, CH-SIMS, which contains 2,281 refined video segments in the wild with both multimodal and independent unimodal annotations.It allows researchers to study the interaction between modalities or use independent unimodal annotations for unimodal sentiment analysis.Furthermore, we propose a multi-task learning framework based on late fusion as the baseline.Extensive experiments on the CH-SIMS show that our methods achieve state-of-the-art performance and learn more distinctive unimodal representations.
Wenmeng Yu, Fanyang Meng, Jiele Wu, Jiyun Zou
ACL3
2020 Spatio-Temporal and Geometry Constrained Network for Automobile Visual Odometry
abstract
Visual odometry (VO) is an essence of vision-based localization and mapping system where existing learning-based approaches utilize CNN and RNN to model camera motion and gain promising results. However, these methods lack full use of the relationship between spatial characteristics and temporal clues, as well as geometry constraints in VO. To overcome these deficiencies, an end-to-end framework that leverages spatio-temporal relevance and geometrical knowledge is proposed. In particular, a spatial response module (SRM) is designed to extract the visual motion features by emphasizing the most interconnected regions while suppressing the irrelevant areas. A module named temporal response module (TRM) is used to regress the camera motion via adopting the optimal motion features. Moreover, a geometry constrained (GC) loss that minimizes the estimated inter-frame pose errors and the accumulated pose errors within a local period is introduced. Actually, the GC loss utilizes adaptive learnable balance factors for balancing losses. Experimental results on KITTI and Malaga datasets demonstrate that the proposed model outperforms state-of-the-art monocular methods.
Hong Liu 0008, Weibo Huang, Guoliang Hua, Fanyang Meng
ICASSP5
2020 Unsupervised Monocular Visual-inertial Odometry Network
abstract
Recently, unsupervised methods for monocular visual odometry (VO), with no need for quantities of expensive labeled ground truth, have attracted much attention. However, these methods are inadequate for long-term odometry task, due to the inherent limitation of only using monocular visual data and the inability to handle the error accumulation problem. By utilizing supplemental low-cost inertial measurements, and exploiting the multi-view geometric constraint and sequential constraint, an unsupervised visual-inertial odometry framework (UnVIO) is proposed in this paper. Our method is able to predict the per-frame depth map, as well as extracting and self-adaptively fusing visual-inertial motion features from image-IMU stream to achieve long-term odometry task. A novel sliding window optimization strategy, which consists of an intra-window and an inter-window optimization, is introduced for overcoming the error accumulation and scale ambiguity problem. The intra-window optimization restrains the geometric inferences within the window through checking the photometric consistency. And the inter-window optimization checks the 3D geometric consistency and trajectory consistency among predictions of separate windows. Extensive experiments have been conducted on KITTI and Malaga datasets to demonstrate the superiority of UnVIO over other state-of-the-art VO / VIO methods. The codes are open-source.
Guoliang Hua, Weibo Huang, Fanyang Meng, Hong Liu 0008
IJCAI4
2019 Joint Dynamic Pose Image and Space Time Reversal for Human Action Recognition from Videos
abstract
Human action recognition aims to classify a given video according to which type of action it contains. Disturbance brought by clutter background and unrelated motions makes the task challenging for video frame-based methods. To solve this problem, this paper takes advantage of pose estimation to enhance the performances of video frame features. First, we present a pose feature called dynamic pose image (DPI), which describes human action as the aggregation of a sequence of joint estimation maps. Different from traditional pose features using sole joints, DPI suffers less from disturbance and provides richer information about human body shape and movements. Second, we present attention-based dynamic texture images (att-DTIs) as pose-guided video frame feature. Specifically, a video is treated as a space-time volume, and DTIs are obtained by observing the volume from different views. To alleviate the effect of disturbance on DTIs, we accumulate joint estimation maps as attention map, and extend DTIs to attention-based DTIs (att-DTIs). Finally, we fuse DPI and att-DTIs with multi-stream deep neural networks and late fusion scheme for action recognition. Experiments on NTU RGB+D, UTD-MHAD, and Penn-Action datasets show the effectiveness of DPI and att-DTIs, as well as the complementary property between them.
Mengyuan Liu 0001, Fanyang Meng, Chen Chen 0001, Songtao Wu
AAAI2
2019 Sample Fusion Network: An End-to-End Data Augmentation Network for Skeleton-Based Human Action Recognition
abstract
Data augmentation is a widely used technique for enhancing the generalization ability of deep neural networks for skeleton-based human action recognition (HAR) tasks. Most existing data augmentation methods generate new samples by means of handcrafted transforms. However, these methods often cannot be trained and then are discarded during testing because of the lack of learnable parameters. To solve those problems, a novel type of data augmentation network called a sample fusion network (SFN) is proposed. Instead of using handcrafted transforms, an SFN generates new samples via a long short-term memory (LSTM) autoencoder (AE) network. Therefore, an SFN and HAR network can be cascaded together to form a combined network that can be trained in an end-to-end manner. Moreover, an adaptive weighting strategy is employed to improve the complementarity between a sample and the new sample generated from it by an SFN, thus allowing the SFN to more efficiently improve the performance of the HAR network during testing. The experimental results on various datasets verify that the proposed method outperforms state-of-the-art data augmentation methods. More importantly, the proposed SFN architecture is a general framework that can be integrated with various types of networks for HAR. For example, when a baseline HAR model with three LSTM layers and one fully connected (FC) layer was used, the classification accuracy was increased from 79.53% to 90.75% on the NTU RGB+D dataset using a cross-view protocol, thus outperforming most other methods.
Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Juanhui Tu, Mengyuan Liu 0001
IEEE Trans. Image Process.1
2018 Instance Enhancing Loss: Deep Identity-Sensitive Feature Embedding for Person Search
abstract
Person search, which is vital for intelligent surveillance, aims at detecting and re-identifying pedestrians from whole monitoring images. However, due to the inaccurate pedestrian detections and extremely few instances per training identity, it remains challenging to learn discriminative representations only by labeled identities for person search. To this end, this paper proposes a novel loss function called instance enhancing loss (IEL) to learn deep identity-sensitive features by introducing unlabeled identity information. Specifically, the proposed IEL can selectively annotate unlabeled identities with similar appearances to labeled identities, and utilize these unlabeled identities in conjunction with labeled identities to train the person search network. The amount of unlabeled identities used as labeled instances can be quantitatively adjusted. Moreover, the proposed IEL is trainable and easy to optimize by back propagation algorithms. Extensive experiments on two benchmark datasets, namely CUHK-SYSU and PRW, show that our method outperforms state-of-the-arts for person search.
Wei Shi 0009, Hong Liu 0008, Fanyang Meng, Weipeng Huang
ICIP3
2018 Spatial-Temporal Data Augmentation Based on LSTM Autoencoder Network for Skeleton-Based Human Action Recognition
abstract
Data augmentation is known to be of crucial importance for the generalization of RNN-based methods of skeleton-based human action recognition. Traditional data augmentation methods artificially adopt various transformations merely in spatial domain, which lack effective temporal representation. This paper extends traditional Long Short-Term Memory (LSTM) and presents a novel LSTM autoencoder network (LSTM-AE) for spatial-temporal data augmentation. In the LSTM-AE, the LSTM network preserves the temporal information of skeleton sequences, and the autoencoder architecture can automatically eliminate irrelevant and redundant information. Meanwhile, a regularized cross-entropy loss is defined to guide the LSTM-AE to learn more suitable representations of skeleton data. Experimental results on the currently largest NTU RGB+D dataset and public SmartHome dataset verify that the proposed model outperforms the state-of-the-art methods, and can be integrated with most of the RNN-based action recognition models easily.
Juanhui Tu, Hong Liu 0008, Fanyang Meng, Mengyuan Liu 0001, Runwei Ding
ICIP3
2018 Hierarchical Dropped Convolutional Neural Network for Speed Insensitive Human Action Recognition
abstract
Human action recognition using skeleton data has lots of potential applications in content-based action retrieval and intelligent surveillance, with wide usage of depth sensors and robust skeleton estimation algorithms. Previous methods describe spatial temporal skeleton joints as a compact color image and then use Convolutional Neural Network (CNN) to extract more discriminative deep features. However, these methods ignore the effect of speed variation, which is a common phenomenon and can bring severe intra-varieties to same types of actions. To solve this problem, this paper presents a novel hierarchical dropped CNN architecture, which is constructed in two stages. Dropped CNN (d-CNN) is firstly developed to extract deep features from a probabilistic speed insensitive color image. This image expresses both spatial distributions and temporal evolutions of skeleton joints meanwhile avoids the effect of speed variations. To enhance the temporal discriminative power, we extend d-CNN to a hierarchical structure (h-CNN), where multiple scales of temporal information are encoded. Extensive experiments on benchmark MSRC-12 dataset and the largest NTU RGB+D dataset verify the effectiveness and robustness of the proposed method.
Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Mengyuan Liu 0001, Wei Liu 0065
ICME1
2018 Adaptive Weighted Sparse Principal Component Analysis
abstract
In this paper, we propose an unsupervised feature selection method from the perspective of optimal reconstruction. The features selected by the proposed method can well represent the original data, and the effectiveness of the selected features is demonstrated by robust reconstruction and clustering. The proposed method emphasizes the joint l2, 1-norms minimization on both reconstruction term and regularization term to make them be column-sparse. Relying on the column-sparse property of reconstruction term and regularization term, the proposed method is able to improve the robustness to outliers and select the effective features. The proposed objective function is nonconvex. Fortunately, it can be equivalently reformulated as a convex form (with change of variables) to capture a global optimization solution. In fact, the proposed method is related to the optimal mean robust principal component analysis (OMRPCA) because the proposed method is a sparse self-contained regression type of OMRPCA. Since OMRPCA essentially adds the adaptive weights for data samples, we call the proposed method adaptive weighted sparse principal component analysis (AW-SPCA). Experimental results demonstrate the effectiveness of AW-SPCA.
Shuangyan Yi, Yongsheng Liang 0001, Wei Liu 0065, Fanyang Meng
ICME4
2017 A bidirectional adaptive bandwidth mean shift strategy for clustering
abstract
The bandwidth of a kernel function is a crucial parameter in the mean shift algorithm. This paper proposes a novel adaptive bandwidth strategy which contains three main contributions. (1) The differences among different adaptive bandwidth are analyzed. (2) A new mean shift vector based on bidirectional adaptive bandwidth is defined, which combines the advantages of different adaptive bandwidth strategies. (3) A bidirectional adaptive bandwidth mean shift (BAMS) strategy is proposed to improve the ability to escape from the local maximum density. Compared with contemporary adaptive bandwidth mean shift strategies, experiments demonstrate the effectiveness of the proposed strategy.
Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Liu Wei, Jihong Pei
ICIP1
2015 A Feature Point Matching Based on Spatial Order Constraints Bilateral-Neighbor Vote
abstract
Feature point matching is a fundamental and challenging problem in many computer vision applications. In this paper, a robust feature point matching algorithm named spatial order constraints bilateral-neighbor vote (SOCBV) is proposed to remove outliers for a set of matches (including outliers) between two images. A directed k nearest neighbor (knn) graph of match sets is generated, and the problem of feature point matching is formulated as a binary discrimination problem. In the discrimination process, the class labeled matrix is built via the spatial order constraints defined on the edges that connect a point to its knn. Then, the posterior inlier class probability of each match is estimated with the knn density estimation and spatial order constraints. The vote of each match is determined by averaging all posterior class probabilities that originate from its associative inliers set and is used for removing outliers. The algorithm iteratively removes outliers from the directed graph and recomputes the votes until the stopping condition is satisfied. Compared with other popular algorithms, such as RANSAC, RSOC, GTM, SOC and WGTM, experiments under various testing data sets demonstrate strong robustness for the proposed algorithm.
Fanyang Meng, Xia Li 0006, Jihong Pei
IEEE Trans. Image Process.1