VLDB 2026 Research / reviewers in the wild / expert
Yongsheng Liang 0001
dblp:74/5653-1
· DBLP profile ↗
66ranked-venue papers
0as first author
55since 2021 · last 2026
0000-0002-0891-5577ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 42 since 2021Artificial intelligence and machine learning · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image CompressionabstractPrevailing quantization techniques in Learned Image Compression (LIC) typically employ a static, uniform bit-width across all layers, failing to adapt to the highly diverse data distributions and sensitivity characteristics inherent in LIC models. This leads to a suboptimal trade-off between performance and efficiency. In this paper, we introduce DynaQuant, a novel framework for dynamic mixed-precision quantization that operates on two complementary levels. First, we propose content-aware quantization, where learnable scaling and offset parameters dynamically adapt to the statistical variations of latent features. This fine-grained adaptation is trained end-to-end using a novel Distance-aware Gradient Modulator (DGM), which provides a more informative learning signal than the standard Straight-Through Estimator. Second, we introduce a data-driven, dynamic bit-width selector that learns to assign an optimal bit precision to each layer, dynamically reconfiguring the network's precision profile based on the input data. Our fully dynamic approach offers substantial flexibility in balancing rate-distortion (R-D) performance and computational cost. Experiments demonstrate that DynaQuant achieves R-D performance comparable to full-precision models while significantly reducing computational and storage requirements, thereby enabling the practical deployment of advanced LIC on diverse hardware platforms. Youneng Bao, Yulong Cheng, Mu Li 0005, Yongsheng Liang 0001 |
AAAI | 7 |
| 2026 | A Learned-PPR Decoding Scheme for Partial Packet Recovery in Network CodingabstractNetwork coding (NC) has proven to offer significant benefits in long-distance and broadcast transmissions, enhancing both throughput and energy efficiency. Recent studies have incorporated partial packet recovery (PPR) into packet-level NC, using syndromes from coded packets to correct bit errors and thereby reduce completion delay. Motivated by recent breakthroughs in deep learning, this paper introduces a novel neural networkbased decoding framework for packet-level NC, referred to as Learned-PPR. The proposed framework incorporates a Bilateral Efficient Self-Attention Network (Bi-ESANet) architecture, which leverages a bilateral network structure to effectively capture both inter- and intra-packet information. Furthermore, we introduce an ESA module to mitigate the GPU memory overhead compared with traditional Transformer attention modules. To handle rateless NC, we propose a “rateless masking” training strategy that enables efficient decoding of rateless codes within the Bi-ESANet framework. Simulation results across various transmission scenarios demonstrate that the proposed approach significantly outperforms existing PPR schemes, achieving lower completion delay. Specifically, compared to existing methods, the proposed approach reduces completion delay by more than 25%. However, the introduced framework incurs higher computational complexity due to the integration of the Bi-ESANet architecture. Qifu Tyler Sun, Zongpeng Li, Yangxuan Cheng, Fanyang Meng, Ye Wang 0002, Yongsheng Liang 0001 |
IEEE Internet Things J. | 8 |
| 2026 | Turbo principles meet compression: Rethinking nonlinear transformations in learned image compression
Chao Li 0071, Wen Tan 0001, Fanyang Meng, Runwei Ding, Ye Wang 0002, Wei Liu 0065, Yongsheng Liang 0001 |
J. Vis. Commun. Image Represent. | 7 |
| 2026 | Hierarchical quality-aware guidance for blind JPEG artifacts removal
Shuai Liu 0022, Qingyu Mao, Binqiang Liu, Fanyang Meng, Shuangyan Yi, Yongsheng Liang 0001 |
J. Vis. Commun. Image Represent. | 7 |
| 2026 | Entropy-aware image representation via 2D Gaussian splatting
Jiacong Chen, Qingyu Mao, Shuai Liu 0022, Chao Li 0071, Jierun Lin, Xiandong Meng, Fanyang Meng, Yongsheng Liang 0001 |
Signal Process. | 8 |
| 2026 | Blind JPEG Artifacts Removal via Inverse JPEG CompressionabstractQuantization and chroma downsampling are two primary operations that introduce distortions in the JPEG compression. However, most existing blind methods treat artifacts removal as a direct mapping from compressed images to clean ones. They fail to explicitly model the underlying degradation process or design targeted compensation mechanisms. As a result, these methods can only partially remove compression artifacts and struggle to generalize to diverse or unseen degradation scenarios. In this work, we present a novel perspective that formulates artifacts removal as an approximate inversion of the lossy steps in JPEG. Based on this view, we propose an Inverse JPEG Compression Network (IJCN), which aims to progressively compensate for quantization errors and color distortions. Specifically, we first design a Learnable Offset Guidance Module (LOGM) to approximate inverse quantization by modeling both intra-block and inter-block coefficient correlations for predicting rounding offsets. In addition, we propose a Quantization Table Guidance Module (QTGM) that leverages the quantization tables to guide the reconstruction network in mitigating color distortions. By modeling compensation mechanisms under the guidance of quantization tables, IJCN effectively eliminates artifacts across varying compression levels. Extensive experiments demonstrate that IJCN outperforms existing methods in both quantitative metrics and visual quality. Shuai Liu 0022, Binqiang Liu, Qingyu Mao, Jiacong Chen, Fanyang Meng, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Dual Tensor Low-Rank Representation for Subspace ClusteringabstractBenefiting from the powerful tensor techniques, the tensor low-rank representation has been proposed to construct sophisticated subspace clustering models. Existing tensor low-rank representation methods predominantly rely on a single low-rank prior to reconstruct the row space, which is instrumental in determining the subspace membership of samples by the row space information. However, this strategy neglects the column space and would lead to a subspace information loss. To address this issue, we propose a Dual Tensor Low-Rank Representation method (DTLRR), the first subspace clustering framework to theoretically recover both row and column subspaces simultaneously. Particularly, not simply formulating a dual self-representation model, we instead prove the recovery of both row and column spaces via a unified theoretical framework. Then, we impose low-rank constraints on the two corresponding affinity tensors to effectively capture high-order correlations. Meanwhile, we theoretically demonstrate the existence of compact dictionary tensors within the dual self-representation framework, which effectively eliminates the null spaces of the affinity tensors and significantly reduces computational complexity. Furthermore, an efficient Alternating Direction Method of Multipliers (ADMM) algorithm is designed to solve the proposed DTLRR model with guaranteed convergence. Extensive experiments validate the superior performance of the proposed DTLRR in data clustering, hyperspectral image denoising, and hyperspectral anomaly detection. Qiangqiang Shen, Yin-Ping Zhao, Yongyong Chen, Yongsheng Liang 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Anchor-Induced Serial Tensor Representation for Multi-View ClusteringabstractMulti-view clustering (MVC) has emerged as a powerful approach for integrating diverse sources of information from complex datasets. Nevertheless, existing methods struggle to accurately capture the global correlations and high-order structures in the data, and employ anchor-based techniques within a single dimension, limiting their representation. To address these issues, we propose an Anchor-induced Serial Tensor Representation (ASTR) framework, which effectively harnesses serial tensor representation to capture comprehensive multi-view information while reducing approximation errors and enhancing clustering performance. Specifically, ASTR begins with projection learning to explore low-dimensional latent spaces in multi-view data. Then, we introduce multi-anchor learning, where multiple anchor configurations are generated within the latent spaces, yielding a set of corresponding bipartite graphs. Besides, we organize these bipartite graphs into a sequence of global tensors, forming the serial tensor representation that encapsulates high-order inter- and intra-view relationships. Furthermore, we introduce the Laplace function to achieve a more accurate tensor rank approximation, complemented by a thorough theoretical analysis. Finally, a one-step clustering process, guided by adaptive weights, directly fuses the learned graphs to produce the final clustering indicator matrix. Experimental results demonstrate that ASTR possesses superior clustering accuracy and comparable efficiency. Zonglin Liu 0001, Zhiwei Zhong 0001, Qiangqiang Shen, Yongsheng Liang 0001, Yongyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Superimposed Pilot-Based Adaptive Semantic Communications for Wireless Image TransmissionabstractThe non-orthogonal superimposed pilot (NOSIP) scheme significantly improves the spectrum efficiency of semantic communication (SemCom) systems by effectively reducing pilot overhead. However, existing SemCom systems based on NOSIP still face challenges, including strong model-channel coupling and limited adaptability to heterogeneous channels. To address these issues, this paper proposes a flexible, channel-adaptive digital SemCom (D-SemCom) architecture based on the NOSIP scheme. Specifically, we design a lightweight semantic codec, termed ShiftViT, and a semantic receiver, termed ShiftRx, which employ time- and frequency-domain shift mechanisms to decouple pilot and data and suppress multi-user interference, thereby enabling image transmission under complex channel conditions. Furthermore, a lightweight channel adaptation algorithm based on first-order meta-learning is proposed to facilitate rapid adaptation and mitigate the strong coupling between semantic models and channel environments. Numerical results demonstrate that the proposed D-SemCom system achieves approximately 25.14% and 1.16% goodput improvements over traditional receivers and existing state-of-the-art methods, respectively, while reducing the computational complexity in terms of FLOPs by approximately 40.37%. In addition, the proposed channel adaptation algorithm is shown to rapidly adapt to diverse channel scenarios within the D-SemCom system. Jian Xiao 0003, Wenwu Xie, Fanyang Meng, Renhai Feng, Liang Yang 0001, Yongsheng Liang 0001 |
IEEE Trans. Wirel. Commun. | 7 |
| 2025 | Robust Deep Joint Source-Channel Coding for Video Transmission over Multipath Fading ChannelabstractTo address the challenges of wireless video transmission over multipath fading channels, we propose a robust deep joint source-channel coding (DeepJSCC) framework by effectively exploiting temporal redundancy and incorporating robust innovations at the modulation, coding, and decoding stages. At the modulation stage, tailored orthogonal frequency division multiplexing (OFDM) for robust video transmission is employed, decomposing wideband signals into orthogonal frequency-flat sub-channels to effectively mitigate frequency-selective fading. At the coding stage, conditional contextual coding with multi-scale Gaussian warped features is introduced to efficiently model temporal redundancy, significantly improving reconstruction quality under strict bandwidth constraints. At the decoding stage, a lightweight denoising module is integrated to robustly simplify signal restoration and accelerate convergence, addressing the suboptimality and slow convergence typically associated with simultaneously performing channel estimation, equalization, and semantic reconstruction. Experimental results demonstrate that the proposed robust framework significantly outperforms state-of-the-art video DeepJSCC methods, which achieves an average reconstruction quality gain of 5.13 dB under challenging multipath fading channel conditions1. Bohuai Xiao, Fanyang Meng, Wei Liu 0065, Yongsheng Liang 0001 |
GLOBECOM | 5 |
| 2025 | Deep Receiver for Multi-Layer Data Transmission with Superimposed PilotsabstractWe investigate a multi-layer data transmission scheme with superimposed pilots (SIPs) to enhance the throughput of multiple-input multiple-output orthogonal frequency-division multiplexing systems. However, in multi-layer data transmission scenarios, signal coupling between different antennas and layers causes severe interference issues, posing significant challenges for receiver design. To address this issue, we propose a deep learning-based receiver architecture, named SANet, which leverages the parallel processing capabilities of the multi-head self-attention (MHSA) mechanism. Specifically, each head of the MHSA mechanism is used to extract local features from each layer of the received signal, enabling the separation and reception of multi-layer bitstream information. Additionally, a flexible and diverse data augmentation strategy is designed to enhance the generalization capability of the deep receiver. Numerical results show that, compared to traditional schemes, the proposed SANet with orthogonal pilots can improve throughput by 7.01%, while the proposed SANet with SIPs can improve throughput by 37.15%. Jian Xiao 0003, Qingyu Mao, Shuai Liu 0022, Bohuai Xiao, Yongsheng Liang 0001 |
ICASSP | 6 |
| 2025 | Dataset Distillation as Data Compression: A Rate-Utility PerspectiveabstractDriven by the ``scale-is-everything'' paradigm, modern machine learning increasingly demands ever-larger datasets and models, yielding prohibitive computational and storage requirements. Dataset distillation mitigates this by compressing an original dataset into a small set of synthetic samples, while preserving its full utility. Yet, existing methods either maximize performance under fixed storage budgets or pursue suitable synthetic data representations for redundancy removal, without jointly optimizing both objectives. In this work, we propose a joint rate-utility optimization method for dataset distillation. We parameterize synthetic samples as optimizable latent codes decoded by extremely lightweight networks. We estimate the Shannon entropy of quantized latents as the rate measure and plug any existing distillation loss as the utility measure, trading them off via a Lagrange multiplier. To enable fair, cross-method comparisons, we introduce bits per class (bpc), a precise storage metric that accounts for sample, label, and decoder parameter costs. On CIFAR-10, CIFAR-100, and ImageNet-128, our method achieves up to $170\times$ greater compression than standard distillation at comparable accuracy. Across diverse bpc budgets, distillation losses, and backbone architectures, our approach consistently establishes better rate-utility trade-offs. Youneng Bao, Yongsheng Liang 0001, Mu Li 0005, Kede Ma |
ICCV | 4 |
| 2025 | Multimodal-Guided Perceptual Image Compression via Joint Text and Audio
Genhong Wang, Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
ICIC (3) | 5 |
| 2025 | DMSO: A Dynamic Momentum-Smoothing Optimizer for Learned Image CompressionabstractLearned Image Compression (LIC) has rapidly evolved and recently surpassed traditional methods in Rate-Distortion (R-D) performance. However, most LIC approaches improve network architectures while increasing computational overhead and overlooking using optimizers tailored specifically for LIC. This paper proposes a Dynamic Momentum-Smoothing Optimizer (DMSO) tailored for LIC to achieve faster convergence and better R-D performance. Specifically, DMSO leverages historical gradient information to smooth the optimization process dynamically, thereby reducing in-stability from gradient oscillations. Furthermore, DMSO introduces a novel Enhanced Second-order Momentum mechanism to mitigate cumulative noise and align momentum updates more closely with the true gradient. Experimental results demonstrate that DMSO operates as a universal optimization plugin for LIC methods, achieving faster and more stable convergence while improving R-D performance to varying degrees without additional parameter count or computational cost. Chao Li 0071, Chuanmin Jia, Fanyang Meng, Siwei Ma 0001, Yongsheng Liang 0001 |
ICIP | 6 |
| 2025 | Towards Robust Text-Guided Image Compression Under Modality MissingabstractText-guided image compression aims to enhance the perceptual quality of reconstructed images by leveraging textual semantic information. However, existing methods struggle to effectively integrate text information and suffer from significant performance degradation when text is unavailable. To address these issues, we propose a Robust Text-Guided Image Compression (RobustTGIC) network that fully utilizes text semantics when available and mitigates performance loss when absent. Specifically, we introduce a Dual-Dimensional Text Modulation (DDTM) module to enhance perceptual quality by accurately fusing textual information. Building on this, we further propose an Intermediate Feature Modulation (IFM) module, which compensates for missing semantics through lightweight adaptation, improving robustness. Experimental results indicate that, at low bitrates (e.g., 0.07 bpp), our method achieves superior perceptual quality reconstruction while significantly reducing bitrates (e.g., 0.5× HiFiC and 0.4× Bpg). Moreover, our method effectively mitigates performance degradation caused by missing text with a parameter increase of less than 1% of the total parameters. Genhong Wang, Wen Tan 0001, Fanyang Meng, Yongsheng Liang 0001 |
ICIP | 4 |
| 2025 | Grouped Transform for Ultra-Low-Complexity Learned Image CompressionabstractExisting learned image compression (LIC) methods have shown strong performance advantages but also bring high computational complexity, making it challenging to deploy them on resource-constrained devices. To reduce the high computational and storage cost, we propose a fully grouped image compression network by introducing spatial and channel grouping operations. Grouping operation is helpful to obtain compact representations by aggregating similar features and reducing redundancy between features in LIC task. Specifically, our proposed network consists of two efficient parts, one is the spatial grouping transform for spatial resolution sampling, and the other is the channel grouping transform for nonlinear representation capability enhancement. Moreover, convolutional kernel factorization and inverted bottleneck are used to reduce redundancy and enrich information of each group in the channel grouping transform, which achieve a good balance between computational complexity and network performance. Experimental results show that our method not only achieves competitive rate-distortion performance with fewer KMACs/pixel and model parameters, but also reduces the real-world runtime. In particular, our proposed models provide at least over 84.6% computational complexity reduction when compared with several advanced LIC methods. Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
ISCAS | 4 |
| 2025 | Structured Sparsity Learning for Efficient Learned Image CompressionabstractExisting learned image compression (LIC) methods have achieved outstanding performance, but their deployment on resource-constrained devices is hindered by the high computational complexity and large model storage. Sparsity learning can obtain sparse neural networks by applying regularization term and further achieve model compression by pruning. However, it is difficult to achieve a lightweight LIC network by directly applying sparsity learning and pruning due to unstructured sparsity and limitations of entropy model. In this paper, we propose to add L2,1regularization during the network training for image compression task, which generates structured sparsity at both filter and channel level. We further analyze the effect of entropy model capacity, and adopt filter/channel fixing to achieve the alignment of entropy estimation for actual pruning. Moreover, we utilize incremental regularization to improve the sparsity of network and training stability. Experimental results show that our pruned lightweight model can effectively reduce network parameters by an average of 59.73% at the cost of 1.57% BD-rate increase compared with original hyperprior model. Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Lihan Zhu, Yongsheng Liang 0001 |
ISCAS | 6 |
| 2025 | Motion Matters: Compact Gaussian Streaming for Free-Viewpoint Video Reconstructionabstract3D Gaussian Splatting (3DGS) has emerged as a high-fidelity and efficient paradigm for online free-viewpoint video (FVV) reconstruction, offering viewers rapid responsiveness and immersive experiences. However, existing online methods face challenge in prohibitive storage requirements primarily due to point-wise modeling that fails to exploit the motion properties. To address this limitation, we propose a novel Compact Gaussian Streaming (ComGS) framework, leveraging the locality and consistency of motion in dynamic scene, that models object-consistent Gaussian point motion through keypoint-driven motion representation. By transmitting only the keypoint attributes, this framework provides a more storage-efficient solution. Specifically, we first identify a sparse set of motion-sensitive keypoints localized within motion regions using a viewspace gradient difference strategy. Equipped with these keypoints, we propose an adaptive motion-driven mechanism that predicts a spatial influence field for propagating keypoint motion to neighboring Gaussian points with similar motion. Moreover, ComGS adopts an error-aware correction strategy for key frame reconstruction that selectively refines erroneous regions and mitigates error accumulation without unnecessary overhead. Overall, ComGS achieves a remarkable storage reduction of over 159 × compared to 3DGStream and 14 × compared to the SOTA method QUEEN, while maintaining competitive visual fidelity and rendering speed. Project page: https://chenjiacong-1005.github.io/ComGS/. Jiacong Chen, Qingyu Mao, Youneng Bao, Xiandong Meng, Fanyang Meng, Ronggang Wang, Yongsheng Liang 0001 |
NeurIPS | 7 |
| 2025 | Progressive Diffusion-Based Low Rate Perceptual Image Compression with Discrete Gaussian Codebooks for Remote Sensing
Yangxuan Cheng, Fanyang Meng, Runwei Ding, Ye Wang 0002, Yongsheng Liang 0001 |
PRCV (9) | 6 |
| 2025 | Boosting Neural Video Representation via Online Structural Reparameterization
Qingyu Mao, Shuai Liu 0022, Qilei Li, Fanyang Meng, Yongsheng Liang 0001 |
PRCV (6) | 6 |
| 2025 | Stable successive Neural Image Compression via coherent demodulation-based transformation
Youneng Bao, Wen Tan 0001, Mu Li 0005, Fanyang Meng, Yongsheng Liang 0001 |
Signal Process. | 5 |
| 2025 | Adaptive cross-channel transformation based on self-modulation for learned image compression
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Yongsheng Liang 0001 |
Signal Process. Image Commun. | 5 |
| 2025 | ShiftLIC: Lightweight Learned Image Compression With Spatial-Channel Shift OperationsabstractLearned Image Compression (LIC) has attracted considerable attention due to their outstanding rate-distortion (R-D) performance and flexibility. However, the substantial computational cost poses challenges for practical deployment. The issue of feature redundancy in LIC is rarely addressed. Our findings indicate that many features within the LIC backbone network exhibit similarities. This paper introduces ShiftLIC, a novel and efficient LIC framework that employs parameter-free shift operations to replace large-kernel convolutions, significantly reducing the model’s computational burden and parameter count. Specifically, we propose the Spatial Shift Block (SSB), which combines shift operations with small-kernel convolutions to replace large-kernel. This approach maintains feature extraction efficiency while reducing both computational complexity and model size. To further enhance the representation capability in the channel dimension, we propose a channel attention module based on recursive feature fusion. This module enhances feature interaction while minimizing computational overhead. Additionally, we introduce an improved entropy model integrated with the SSB module, making the entropy estimation process more lightweight and thereby comprehensively reducing computational costs. Experimental results demonstrate that ShiftLIC outperforms leading compression methods, such as VVC Intra and GMM, in terms of computational cost, parameter count, and decoding latency. Additionally, ShiftLIC sets a new SOTA benchmark with a BD-rate gain per MACs/pixel of −102.6%, showcasing its potential for practical deployment in resource-constrained environments. The code is released athttps://github.com/baoyu2020/ShiftLIC. Youneng Bao, Wen Tan 0001, Chuanmin Jia, Mu Li 0005, Yongsheng Liang 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Reliable Entropy-Induced Anchor Learning for Incomplete Multi-View Subspace ClusteringabstractUnder large-scale data with missing views, fast incomplete multi-view clustering (IMVC) with anchor learning is of critical importance due to its linear complexity$\mathcal {O}(n)$. However, existing anchor-based methods only explore the column orthogonality of anchor points, where their arbitrary column orthogonal basis vectors have weak constraint relationships with real samples and significant deviations from more representative anchors, thereby impeding the precise representation of sample similarities. To solve this issue, we propose a Reliable Entropy-induced anchor learning for incomplete Multi-view subspace Clustering (REMC), which performs an entropy approximation term to learn more representative anchors, and we prove that the information entropy minimization can be relaxed into the$\ell _{2,1}$-norm paradigm. Specifically, the proposed REMC first integrates anchor learning and subspace clustering to produce multiple view-specific bipartite graphs and capture the high-order correlations by imposing these bipartite graphs with the tensor nuclear norm. Then, we fuse all the view-specific bipartite graphs to build a consensus bipartite graph with entropy approximation regularization, and hence the proposed REMC can produce a more discriminative similarity graph, preserving each non-zero element in its column close to 1, while the other elements are approaching 0. Besides, an efficient algorithm is designed to solve the proposed REMC. Numerous results show the superior performance of our method on both the complete and incomplete data. Qiangqiang Shen, Zihou Guo, Yanhui Xu, Yongyong Chen, Shiqi Wang 0001, Yongsheng Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | One is All: A Unified Rate-Distortion-Complexity Framework for Learned Image Compression Under Energy Concentration CriteriaabstractThe learned image compression (LIC) technique has surpassed the state-of-the-art traditional codecs (H.266/VVC) in case of rate-distortion (R-D) performance. Its real-time deployments are far advanced. In order to achieve more flexible deployments, an LIC technique should be flexible in adjusting its computational complexity and rate as demanded by a situation and its environment. In this paper, we propose a unified Rate-Distortion-Complexity (R-D-C) framework for LIC under channel energy concentration criteria. Specifically, we first introduce an Energy Asymptotic Nonlinear Transformation (EANT) designed to directly concentrate on the channel energy of latent representations, thus laying the groundwork for a scalable entropy coding. Next, leveraging this energy concentration characteristic, we propose a corresponding Heterogeneous Scalable Entropy Model (HSEM) for flexibly scaling bitstreams as needed. Finally, utilizing the proposed EANT, we construct a fine-grained scalable codec for formulating, in combination with HSEM, a comprehensive scalable R-D-C framework under the energy concentration criteria. The obtained experimental results demonstrate that the proposed method could enable seamless transitions between 13 different widths of sub-models within a single network, allowing for fine-grained control over the model bitrate, complexity, and hardware inference time. Additionally, the proposed method exhibits competitive R-D performance compared to many existing methods. Chao Li 0071, Fanyang Meng, Qingyu Mao, Youneng Bao, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image CompressionabstractDeep neural network-based image compression (NIC) has achieved excellent performance, but NIC method models have been shown to be susceptible to backdoor attacks. Adversarial training has been validated in image compression models as a common method to enhance model robustness. However, the improvement effect of adversarial training on model robustness is limited. In this paper, we propose a prior knowledge-guided adversarial training framework for image compression models. Specifically, first, we propose a gradient regularization constraint for training robust teacher models. Subsequently, we design a knowledge distillation-based strategy to generate a priori knowledge from the teacher model to the student model for guiding adversarial training. Experimental results show that our method improves the reconstruction quality by about 9dB when the Kodak dataset is elected as the backdoor attack object for psnr attack. Compared with Ma2023 [1], our method has a 5dB higher PSNR output at high bitrate points. Youneng Bao, Fanyang Meng, Chao Li 0071, Wen Tan 0001, Genhong Wang, Yongsheng Liang 0001 |
ICASSP | 7 |
| 2024 | Leveraging Redundancy in Feature for Efficient Learned Image CompressionabstractIn recent years, with the development of the field of learned image compression, numerous models with excellent rate-distortion performance have emerged. However, the considerable computational complexity inherent in these models poses challenges for their practical deployment. In this paper, we investigate feature redundancy in learned image compression (LIC) algorithms for efficient feature extraction and introduce an efficient and lightweight LIC framework. Specifically, we explore the existence of a large number of similar features in the network. Subsequently, we design effective feature extraction modules across various levels, such as layer and block. In addition, based on the fact that the role of the codec’s encoder is to remove redundancy and the decoder is to reconstruct, we propose an asynchronous feature fusion block. This fusion block incorporates an "edge smoothing" operator in the encoder and an "edge enhancement" operator in the decoder. Our methodology strikes an ideal balance between rate-distortion performance and efficiency. The experimental results indicate that our approach necessitates only 310KMac/pixel computation and 9.5M parameters, while in terms of performance, our method achieves a 20.7% BD-rate advantage over BPG on Kodak data, mirroring VVC’s performance. Compared to other learned image compression algorithms with SOTA performance, our method has a great advantage in terms of computation/parameter count. Youneng Bao, Fanyang Meng, Wen Tan 0001, Chao Li 0071, Genhong Wang, Yongsheng Liang 0001 |
ICASSP | 7 |
| 2024 | Bandwidth-Efficient Inference for Nerual Image CompressionabstractWith neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottle-neck in implementing network inference on mobile and edge devices. In this paper, we propose an end-to-end differentiable bandwidth efficient neural inference method with the activation compressed by neural data compression method. Specifically, we propose a transform-quantization-entropy coding pipeline for activation compression with symmetric exponential Golomb coding and a data-dependent Gaussian entropy model for arithmetic coding. Optimized with existing model quantization methods, low-level task of image compression can achieve up to 19× bandwidth reduction with 6.21× energy saving. The code implementation is available at https://github.com/xyzysz/Bandwidth_efficient_nic. Shanzhi Yin, Tongda Xu, Yongsheng Liang 0001, Yanghao Li, Yan Wang 0002 |
ICASSP | 3 |
| 2024 | Enhanced Interpretability in Learned Image Compression via Convolutional Sparse CodingabstractCompared to traditional image compression methods, learned image compression (LIC) methods have demonstrated increasingly superior rate-distortion performance. However, LIC networks are often regarded as black boxes, still lacking a theoretical understanding. Sparse coding provides the sparse and interpretable modeling for analyzing or synthesizing natural images in various signal and image processing applications. Therefore, we introduce convolutional sparse coding (CSC) into transform network for enhancing the interpretability of LIC methods. In this paper, we first employ CSC layers to achieve certain theoretical modeling for LIC network, and adopt a weight sharing strategy in encoder-decoder pair and attention mechanism to balance the complexity and performance. Additionally, we analyze the model robustness against data input perturbations and consider the impact of sparsity trade-off parameter in the CSC layer optimization process. Experimental results demonstrate that our method achieves comparable performance with the corresponding baseline, and our model is more robust. Yiwen Tu, Wen Tan 0001, Youneng Bao, Genhong Wang, Fanyang Meng, Yongsheng Liang 0001 |
ICME | 6 |
| 2024 | Fine-Grained Adjustable Entropy Models for Rate-Complexity Jointly Adjustable Image Compression
Chao Li 0071, Shanzhi Yin, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
PRCV (9) | 6 |
| 2024 | Multirate Progressive Entropy Model for Learned Image CompressionabstractThis paper proposes a unified and efficient entropy coding method for learned image compression (LIC) from the perspective of traditional signal processing. First, the consistency of structures and optimization objectives are used to interpret the existing split-coded-then-merge entropy coding strategies in LIC as a particular filter banks framework, with feature separation and feature aggregation representing the analysis filter bank and synthesis filter bank, respectively. Thus, we borrow the design from the multirate filter banks and proposed Multirate Progressive Entropy Model (MPEM) to enhance the rate-distortion performance and decoding speed. In particular, we create an analysis filter bank that divides compact features into a few nonuniform subsets based on various spatial and channel sampling rates. Then multi-scale detail and mean coefficients within the current subset are used as prior representations to help generate the prediction parameters of the next subset, and the carefully designed synthetic filter bank performs a near-perfect reconstruction of the features. In addition, we propose a Multi-level Edge Attention Moudal (MEAM) to increase the edge and texture information’s contribution and reduce the high-frequency information loss brought on by MPEM’s inherent multi-rate spatial sampling, which leverages the edge operator and structural reparameterization principles. The results of the experiments show that, in comparison to the effective LIC methods and traditional code, the proposed MPEM can decode data at a cutting-edge speed while also offering comparable rate-distortion performance. Chao Li 0071, Shanzhi Yin, Chuanmin Jia, Fanyang Meng, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Self-Completed Bipartite Graph Learning for Fast Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC), excavating diversity and consistency from multiple incomplete views, has aroused widespread research enthusiasm. Nevertheless, most existing methods still encounter the following issues: 1) they generally concentrate on pair-wise instance correlation, which consumes at least a quadratic complexity and precludes them from applying at large scales; 2) they only concentrate on pair-wise instance relevance, whereas ignoring the discriminative correlation hidden across views. To overcome these drawbacks, we propose the Self-Completed Bipartite Graph Learning (SCBGL) method for fast IMVC, which adaptively learns a self-completed consensus bipartite graph with the guidance of global information. Specifically, SCBGL learns the consensus anchor matrix shared among diverse views and further constructs a consensus intra-view bipartite graph with missing instances to explore the diversity and complementarity underlying different views. Meanwhile, we concatenate all the multiple features with projection learning to learn global anchors that would be employed to construct an inter-view bipartite graph. Furthermore, SCBGL dexterously utilizes the abundant inter-view information to tutor the self-completion of the consensus intra-view bipartite graph. By devising an alternatively iterative strategy, we present an efficient algorithm, which enjoys a linear time complexity, to solve the proposed SCBGL model. Numerous experiments conducted on large-scale datasets substantiate the superior performance of the SCBGL beyond the state-of-the-arts. Xiaojia Zhao, Qiangqiang Shen, Yongyong Chen, Yongsheng Liang 0001, Junxin Chen 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Pick-and-Place Transform Learning for Fast Multi-View ClusteringabstractTo manipulate large-scale data, anchor-based multi-view clustering methods have grown in popularity owing to their linear complexity in terms of the number of samples. However, these existing approaches pay less attention to two aspects. 1) They target at learning a shared affinity matrix by using the local information from every single view, yet ignoring the global information from all views, which may weaken the ability to capture complementary information. 2) They do not consider the removal of feature redundancy, which may affect the ability to depict the real sample relationships. To this end, we propose a novel fast multi-view clustering method via pick-and-place transform learning named PPTL, which could capture insightful global features to characterize the sample relationships quickly. Specifically, PPTL first concatenates all the views along the feature direction to produce a global matrix. Considering the redundancy of the global matrix, we design a pick-and-place transform with ℓ2,p-norm regularization to abandon the poor features and consequently construct a compact global representation matrix. Thus, by conducting anchor-based subspace clustering on the compact global representation matrix, PPTL can learn a consensus skinny affinity matrix with a discriminative clustering structure. Numerous experiments performed on small-scale to large-scale datasets demonstrate that our method is not only faster but also achieves superior clustering performance over state-of-the-art methods across a majority of the datasets. Qiangqiang Shen, Yongyong Chen, Changqing Zhang 0002, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Robust Tensor Recovery for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering is gaining increased attention owing to its great success in mining underlying information from the missing views. However, the existing approaches still encounter two issues: 1) They generally do not give sufficient consideration to the robustness of incomplete multi-view data with noise; 2) They only exploit the low-rank structures in the intra-view graphs, while the low-rank priors embedded in inter-view graphs are ignored. To this end, we propose a Robust Tensor Recovery for Incomplete Multi-view Clustering (RIMC) method, which transforms the view-missing problem into the tensor graph recovery problem by manipulating the comprehensive low-rank priors. Specifically, RIMC first employs a marginalized denoising operation to construct robust graphs and further builds a tensor graph by stacking these robust graphs. Then, we develop a novel tensor completion to recover the tensor graph by performing comprehensive low-rank priors: low-rank structures in the inter-view graphs (i.e., horizontal and lateral slices); low-rank structures in the intra-view graphs (i.e., frontal slices). Meanwhile, we integrate the tensor completion and spectral clustering to learn a unified indicator matrix. Extensive experiments show the promising performance of our method. Qiangqiang Shen, Yongsheng Liang 0001, Yongyong Chen, Zhenyu He 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | A Decoupled Spatial-Channel Inverted Bottleneck For Image CompressionabstractResidual block has achieved great success in deep networks to eliminate accuracy degradation, and there emerges a large number of variants with more competitive performance. However, these blocks are introduced for high-level tasks that only encode the input image into semantic and struc¬tural features but do not need to reconstruct. So for the low-level task like image compression where reconstruction quality contributes significantly to the rate-distortion perfor¬mance, the structure of the residual block needs modification for more suitable implementation. In this paper, we revisit the existing residual blocks and discover two key principles summarized as two decouplings: spatial-channel decoupling and linear-nonlinear decoupling. We propose an efficient nonlinear transform based on the principles dubbed decou¬pled spatial-channe inverted bottleneck(DSCIB), which has a linear-spatial branch for rough reconstruction and a nonlinear¬channel branch to provide detailed featrues. We employ the DSCIB module in the joint auto regression model to build an overall network. Experimental results show that our method achieves comparable performance with the existing learning-based image compression methods at high bitrate while re¬ducing 38% FLOPs. Wen Tan 0001, Fanyang Meng, Yongsheng Liang 0001 |
ICIP | 4 |
| 2023 | A Complex-Valued Neural Network Based Robust Image Compression
Can Luo, Youneng Bao, Wen Tan 0001, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001 |
PRCV (10) | 6 |
| 2023 | Improving generalization of double low-rank representation using Schatten-p norm
Jiaoyan Zhao, Yongsheng Liang 0001, Shuangyan Yi, Qiangqiang Shen, Xiaofeng Cao 0002 |
Pattern Recognit. | 2 |
| 2023 | Taylor series based dual-branch transformation for learned image compression
Youneng Bao, Wen Tan 0001, Linfeng Zheng, Fanyang Meng, Wei Liu 0065, Yongsheng Liang 0001 |
Signal Process. | 6 |
| 2023 | Nonlinear Transforms in Learned Image Compression From a Communication PerspectiveabstractRecently, remarkable progress has been made in learned image compression (LIC), in which nonlinear transforms (NTs) play a crucial role. Although there are many NT methods for improving the rate distortion performance, all the existing methods sacrifice the computational complexity and the number of parameters of the transformation. This paper provides a fundamental novel viewpoint on nonlinear transforms from a communication perspective, and shows how this idea can be extended to design efficient NT methods. In particular, the nonlinear transforms are inferred as signal modulation modules. Under this extrapolation, the current NTs are generalized as amplitude modulation that only varies the amplitude of the carrier wave. Therefore, a nonlinear modulation-like transform (NMLT) which varies the phase angle of the carrier is proposed. Moreover, this concept is extended by introducing In-phase/Quadrature (IQ) modulation, which is a boosting technique in communication field, in order to enhance NMLT. Furthermore, the Bit-interleaved technique in communication is used to guide the optimization of NTML with IQ. The experimental results on different datasets and backbone architectures verify the efficiency and robustness of the proposed methods. For example, when backbone architecture is hyperprior model, our method achieves 19.37% BD-rate reduction over GDN on the Kodak dataset. In addition, our method with channel wise autoregressive model leads to the state-of-the-art rate-distortion performance. Youneng Bao, Fanyang Meng, Chao Li 0071, Siwei Ma 0001, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Semi-Supervised CT Lesion Segmentation Using Uncertainty-Based Data Pairing and SwapMixabstractSemi-supervised learning (SSL) methods show their powerful performance to deal with the issue of data shortage in the field of medical image segmentation. However, existing SSL methods still suffer from the problem of unreliable predictions on unannotated data due to the lack of manual annotations for them. In this paper, we propose an unreliability-diluted consistency training (UDiCT) mechanism to dilute the unreliability in SSL by assembling reliable annotated data into unreliable unannotated data. Specifically, we first propose an uncertainty-based data pairing module to pair annotated data with unannotated data based on a complementary uncertainty pairing rule, which avoids two hard samples being paired off. Secondly, we develop SwapMix, a mixed sample data augmentation method, to integrate annotated data into unannotated data for training our model in a low-unreliability manner. Finally, UDiCT is trained by minimizing a supervised loss and an unreliability-diluted consistency loss, which makes our model robust to diverse backgrounds. Extensive experiments on three chest CT datasets show the effectiveness of our method for semi-supervised CT lesion segmentation. Pengchong Qiao, Guoli Song, Hu Han 0001, Yonghong Tian 0001, Yongsheng Liang 0001, Xi Li 0011, Shaohua Kevin Zhou, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Bilateral Fast Low-Rank Representation With Equivalent Transformation for Subspace ClusteringabstractIn recent years, low-rank representation (LRR) has received increasing attention on subspace clustering. Due to inevitable matrix inversion and singular value decomposition in each iteration, however, most of existing LRR algorithms may suffer from high computational complexity, and hence can not cope with the large-scale sample data commendably. To overcome this problem, in this paper, we propose a bilateral fast low-rank representation (BFLRR), which has a linear time complexity with respect to the number of samples. Specifically, we introduce the equivalent transformation method to remove the null spaces of both the columns and rows of the coefficient matrix so that a hypercompact coefficient matrix can be learned. Furthermore, the proposed BFLRR is embedded into a distributed framework as DFC-BFLRR to make it more efficient, which utilizes a combination of the global and local projection matrices. Extensive experiments are carried out on real datasets, and the results testify that the proposed methods not only perform faster-computing speed but also obtain favorable clustering accuracy in comparison with the competing methods among large-scale sample data. Qiangqiang Shen, Shuangyan Yi, Yongsheng Liang 0001, Yongyong Chen, Wei Liu 0065 |
IEEE Trans. Multim. | 3 |
| 2022 | AdderIC: Towards Low Computation Cost Image CompressionabstractRecently, learned image compression methods have shown their outstanding rate-distortion performance when compared to traditional frameworks. Although numerous progress has been made in learned image compression, the computation cost is still at a high level. To address this problem, we propose AdderIC, which utilizes adder neural networks (AdderNet) to construct an image compression framework. According to the characteristics of image compression, we introduce several strategies to improve the performance of AdderNet in this field. Specifically, Haar Wavelet Transform is adopted to make AdderIC learn high-frequency information efficiently. In addition, implicit deconvolution with the kernel size of 1 is applied after each adder layer to reduce spatial redundancies. Moreover, we develop a novel Adder-ID-PixelShuffle cascade upsampling structure to remove checkerboard artifacts. Experiments demonstrate that our AdderIC model can largely outperform conventional AdderNet when applied in image compression and achieve comparable rate-distortion performance to that of its CNN baseline with about 80% multiplication FLOPs and 30% energy consumption reduction. Xin Yao 0001, Chao Li 0071, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
ICASSP | 6 |
| 2022 | Universal Efficient Variable-Rate Neural Image CompressionabstractRecently, Learning-based image compression has reached comparable performance with traditional image codecs(such as JPEG, BPG, WebP). However, computational complexity and rate flexibility are still two major challenges for its practical deployment. To tackle these problems, this paper proposes two universal modules named Energy-based Channel Gating(ECG) and Bit-rate Modulator(BM), which can be directly embedded into existing end-to-end image compression models. ECG uses dynamic pruning to reduce FLOPs for more than 50% in convolution layers, and a BM pair can modulate the latent representation to control the bit-rate in a channel-wise manner. By implementing these two modules, existing learning-based image codecs can obtain ability to output arbitrary bit-rate with a single model and reduced computation. Shanzhi Yin, Chao Li 0071, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng, Wei Liu 0065 |
ICASSP | 4 |
| 2022 | Exploring Structural Sparsity in Neural Image CompressionabstractThe performance of neural image compression have reached or suppressed traditional methods (such as JPEG, BPG, WebP). However, their sophisticated network structures with cascaded convolution layers bring heavy computational burden for practical deployment. In this paper, we explore structural sparsity in neural image compression network to obtain real-time acceleration without any specialized hardware design or algorithm. We propose a simple plug-in adaptive binary channel masking(ABCM) to judge the importance of each convolution channel and introduce sparsity during training. During inference, the unimportant channels are pruned to obtain slimmer network and less computation. We implement our method into three neural image compression networks with different entropy models to verify its effectiveness and generalization, the experiment results show that up to 7× computation reduction and 3× acceleration can be achieved with negligible performance drop. Shanzhi Yin, Chao Li 0071, Fanyang Meng, Wen Tan 0001, Youneng Bao, Yongsheng Liang 0001, Wei Liu 0065 |
ICIP | 6 |
| 2022 | Robust active representation via ℓ2, p-norm constraints
Jiaoyan Zhao, Shuangyan Yi, Yongsheng Liang 0001, Wei Liu 0065, Xiaofeng Cao 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Weighted Schatten p-norm minimization with logarithmic constraint for subspace clustering
Qiangqiang Shen, Yongyong Chen, Yongsheng Liang 0001, Shuangyan Yi, Wei Liu 0065 |
Signal Process. | 3 |
| 2022 | Spatial-Temporal Asynchronous Normalization for Unsupervised 3D Action Representation LearningabstractUnsupervised 3D action representation learning from skeleton sequences has attracted increasing attention in recent years. Existing methods have successfully applied autoencoder network to learn 3D action representation by reconstructing original skeleton sequence. However, these methods ignore motion cues thus suffer from distinguishing actions especially with similar shape information and slightly different motion information. Instead of reconstructing original skeleton sequence, we learn distinctive 3D action representation with autoencoder network by reconstructing normalized motion sequence extracted from original input. To obtain the normalized motion sequence, we specifically design a novel spatial-temporal asynchronous normalization (STAN) method, which normalizes original skeleton sequence in two steps. First, STAN reduces redundant temporal information and extracts motion sequence by subtracting mean value along the temporal dimension. Second, STAN further normalizes the motion sequence along the spatial dimension and generates normalized motion sequence that suffers less from the effect of different human body shapes. Extensive experiments on large scale NTU RGB+D 60 and NTU RGB+D 120 datasets verify the effectiveness of our proposed STAN method, which achieves comparative results with state-of-the-art methods, and also outperforms alternative normalization methods. Mengyuan Liu 0001, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng |
IEEE Signal Process. Lett. | 3 |
| 2022 | Unified Information Fusion Network for Multi-Modal RGB-D and RGB-T Salient Object DetectionabstractThe use of complementary information, namely depth or thermal information, has shown its benefits to salient object detection (SOD) during recent years. However, the RGB-D or RGB-T SOD problems are currently only solved independently, and most of them directly extract and fuse raw features from backbones. Such methods can be easily restricted by low-quality modality data and redundant cross-modal features. In this work, a unified end-to-end framework is designed to simultaneously analyze RGB-D and RGB-T SOD tasks. Specifically, to effectively tackle multi-modal features, we propose a novel multi-stage and multi-scale fusion network (MMNet), which consists of a cross-modal multi-stage fusion module (CMFM) and a bi-directional multi-scale decoder (BMD). Similar to the visual color stage doctrine in the human visual system (HVS), the proposed CMFM aims to explore important feature representations in feature response stage, and integrate them into cross-modal features in adversarial combination stage. Moreover, the proposed BMD learns the combination of multi-level cross-modal fused features to capture both local and global information of salient objects, and can further boost the multi-modal SOD performance. The proposed unified cross-modality feature analysis framework based on two-stage and multi-scale information fusion can be used for diverse multi-modal SOD tasks. Comprehensive experiments ($\sim 92\text{K}$image-pairs) demonstrate that the proposed method consistently outperforms the other 21 state-of-the-art methods on nine benchmark datasets. This validates that our proposed method can work well on diverse multi-modal SOD tasks with good generalization and robustness, and provides a good multi-modal SOD benchmark. Wei Gao 0003, Guibiao Liao, Siwei Ma 0001, Ge Li 0002, Yongsheng Liang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Fast Universal Low Rank RepresentationabstractAs well known, low rank representation method (LRR) has obtained promising performance for subspace clustering, and many LRR variants have been developed, which mainly solve the three problems existing in LRR: 1) Problem of mean calculation; 2) Problem of deviating from the real low rank solution; 3) Problem of high computation cost on the large-scale data. In this paper, we first propose a universal LRR method referring to the first two problems. More specifically, we introduce the ability of removing the optimal mean automatically into LRR and extend it to a Schatten$p$-norm minimization problem. Then, referring to the third problem, we reformulate the universal LRR version as an equivalent fast optimization version by removing the null space of data. More importantly, the effective theory proof is proposed to guarantee that the fast optimization method can dramatically improve the algorithmic computation efficiency on the large-scale data but without any loss of information. Finally, experimental results demonstrate the effectiveness and efficiency of the proposed method. Qiangqiang Shen, Yongsheng Liang 0001, Shuangyan Yi, Jiaoyan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Fast Extended Inductive Robust Principal Component Analysis With Optimal MeanabstractInspired by the mean calculation of RPCA_OM and inductiveness of IRPCA, we first propose an inductive robust principal component analysis method with removing the optimal mean automatically, which is shorted as IRPCA_OM. Furthermore, IRPCA_OM is extended to Schatten-$p$norm and a more general framework (i.e., EIRPCA_OM) is presented. The objective function of EIRPCA_OM includes two terms, the first term is a robust reconstruction error term constrained by an$\ell _{2,1}$-norm and the second term is a regularization term constrained by a Schatten-$p$norm. The proposed EIRPCA_OM method is robust, inductive and accurate. However, on the high-dimensional data, it would spend a large computation cost in training stage. To this end, a fast version of EIRPCA_OM called as FEIRPCA_OM is proposed, and its basic idea is to eliminate the zero eigenvalues of data matrix. More importantly, an effective theoretical proof is presented to ensure that FEIRPCA_OM has faster processing speed than EIRPCA_OM when processing high-dimensional data, but without any performance loss. Based on it, we also can exchange the less performance loss for the higher computation efficiency by removing the small eigenvalues of data matrix. Experimental results on the public datasets demonstrate that FEIRPCA_OM works efficiently on the high-dimensional data. Shuangyan Yi, Feiping Nie 0001, Yongsheng Liang 0001, Wei Liu 0065, Zhenyu He 0001, Qingmin Liao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Mbb: A Multi-Scale Method For Data Based On Bit Plane SlicingabstractMulti-scale methodology can enhance the performance of the model in deep learning. The current multi-scale methodology focuses on changing the formation, which will increase the parameters and calculations of the network. This paper offers a multi-scale method for data based on bit plane slicing(MBB). This expands the receptive field of valid information in image data. It is done by multi-level fusing image with high bit planes. Our experimentation shows that by adding MBB in front of the backbone network, one can achieve a significant performance improvement. The MBB approach is widely applicable because it does not require changes to the structure of the backbone network. Youneng Bao, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001, Wei Liu 0065, Kaiyu Liu |
ICIP | 4 |
| 2021 | Improving Convolutional Networks with Boosting Attention ConvolutionsabstractConvolutional neural networks (CNNs) have been widely used in a range of tasks because of its robust convolutional feature transformation ability. In this paper, we propose a novel type of convolution called Boosting Attention Convolution (BAC) to improve the basic convolutional feature transformation process of CNNs. The proposed method is designed based on two principles, boosting and attention mechanism. Specifically, we design a set of simple yet effective Boosting Attention Modules (BAM) within grouped convolution, which progressively recalibrate distribution of feature map and enable the future filters nested in a convolution layer to focus more on the feature regions that are unactivated by previous filters. Thus, it can help CNNs generate more discriminative representations by explicitly incorporating richer information. The experimental results on various datasets verify that BAC outperforms state-of-the-art methods. More importantly, the proposed BAC is a general convolution that can be deployed to various modern networks without introducing much parameters and computational complexity. Chao Li 0071, Yongsheng Liang 0001, Huo-Xiang Yang, Fanyang Meng, Wei Liu 0065, Handong Wang |
ICME | 2 |
| 2021 | Learnable Oriented-Derivative Network for Polyp Segmentation
Mengjun Cheng, Zishang Kong, Guoli Song, Yonghong Tian 0001, Yongsheng Liang 0001, Jie Chen 0001 |
MICCAI (1) | 5 |
| 2021 | Joint segmentation and detection of COVID-19 via a sequential region generation network
Jipeng Wu, Shengchuan Zhang, Xi Li 0011, Jie Chen 0001, Jiawen Zheng, Yue Gao 0002, Yonghong Tian 0001, Yongsheng Liang 0001, Rongrong Ji |
Pattern Recognit. | 9 |
| 2021 | Diverse part attentive network for video-based person re-identification
Xiujun Shu, Ge Li 0002, Longhui Wei, Jia-Xing Zhong, Xianghao Zang, Shiliang Zhang, Yaowei Wang 0001, Yongsheng Liang 0001, Qi Tian 0001 |
Pattern Recognit. Lett. | 8 |
| 2020 | Multi-Task Driven Feature Models for Thermal Infrared TrackingabstractExisting deep Thermal InfraRed (TIR) trackers usually use the feature models of RGB trackers for representation. However, these feature models learned on RGB images are neither effective in representing TIR objects nor taking fine-grained TIR information into consideration. To this end, we develop a multi-task framework to learn the TIR-specific discriminative features and fine-grained correlation features for TIR tracking. Specifically, we first use an auxiliary classification network to guide the generation of TIR-specific discriminative features for distinguishing the TIR objects belonging to different classes. Second, we design a fine-grained aware module to capture more subtle information for distinguishing the TIR objects belonging to the same class. These two kinds of features complement each other and recognize TIR objects in the levels of inter-class and intra-class respectively. These two feature models are learned using a multi-task matching framework and are jointly optimized on the TIR tracking task. In addition, we develop a large-scale TIR training dataset to train the network for adapting the model to the TIR domain. Extensive experimental results on three benchmarks show that the proposed algorithm achieves a relative gain of 10% over the baseline and performs favorably against the state-of-the-art methods. Codes and the proposed TIR dataset are available at https://github.com/QiaoLiuHit/MMNet. Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Wei Liu 0065, Yongsheng Liang 0001 |
AAAI | 7 |
| 2020 | Binary Representation and High Efficient Compression of 3D CNN Features for Action RecognitionabstractA common framework of the action recognition is to collect the videos from different cameras into a cloud center firstly, and then perform the 3D CNN on the cloud server. Although directly, this framework will bring a huge burden to the cloud server and video transmission. To handle this challenge, the "front-cloud" collaborative processing architecture can be used. The most import issue is to compress the feature from 3D CNN effectively without significant loss of accuracy. We propose logarithmic quantization with a maximum value threshold and HEVC inter encoding for 3D CNN features. Experimental results on ResNet-50 and InceptionV1 show that the features can be represented by only 1 bit without significant loss of accuracy. The compression ratio of the quantized 1 bit features using HEVC inter coding can reach to 5000 times and the loss of accuracy is less than 1%. Peiyin Xing, Peixi Peng, Yongsheng Liang 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
DCC | 3 |
| 2020 | BCData: A Large-Scale Dataset and Benchmark for Cell Detection and Counting
Yao Ding 0006, Guoli Song, Lin Wang 0026, Ruizhe Geng, Yonghong Tian 0001, Yongsheng Liang 0001, Shaohua Kevin Zhou, Jie Chen 0001 |
MICCAI (5) | 10 |
| 2020 | Adaptation-Oriented Feature Projection for One-Shot Action RecognitionabstractOne-shot action recognition aims at recognizing actions in unseen classes in cases where only one training video is provided. Compared with one-shot image recognition, one-shot learning on videos is more difficult due to the fact that the temporal dimension of video may lead to greater variation. To handle this variation, it is important to conduct further adaptation in the one-shot training process, despite the scarcity of the training data. While meta-learning is an option for facilitating this adaptation, it cannot be directly applied for two reasons: first, deep networks for action recognition can make current meta-learning methods infeasible to run because of their high computational complexity; second, due to the greater variation in actions, the adapted performance may not be higher than the un-adapted one, making it difficult to train the model by means of meta-learning. To address these problems and facilitate the adaptation, we propose the Adaptation-Oriented Feature (AOF) projection for one-shot action recognition. We first pre-train the base network on seen classes. The output of the network is projected to the adaptation-oriented feature space by fusing the important feature dimensions that are sensitive to adaptation. Subsequently, a small dataset (a.k.a. task) is sampled from seen classes to simulate the unseen-class training and testing settings. The feature adaptation is performed on the training data of this task to integrate the distribution information of the adapted feature. In order to reduce over-fitting, the triplet loss is applied to handle temporal variation with fewer parameters during the adaptation. On the testing data of this task, the losses on both adapted and un-adapted features are calculated to train the projection matrix. This sampling-adaptation-training procedure is then repeated on seen classes until convergence. Extensive experimental results on two challenging one-shot action recognition datasets demonstrate that our proposed method outperforms state-of-the-art methods. Yixiong Zou, Yemin Shi 0001, Daochen Shi, Yaowei Wang 0001, Yongsheng Liang 0001, Yonghong Tian 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | Sample Fusion Network: An End-to-End Data Augmentation Network for Skeleton-Based Human Action RecognitionabstractData augmentation is a widely used technique for enhancing the generalization ability of deep neural networks for skeleton-based human action recognition (HAR) tasks. Most existing data augmentation methods generate new samples by means of handcrafted transforms. However, these methods often cannot be trained and then are discarded during testing because of the lack of learnable parameters. To solve those problems, a novel type of data augmentation network called a sample fusion network (SFN) is proposed. Instead of using handcrafted transforms, an SFN generates new samples via a long short-term memory (LSTM) autoencoder (AE) network. Therefore, an SFN and HAR network can be cascaded together to form a combined network that can be trained in an end-to-end manner. Moreover, an adaptive weighting strategy is employed to improve the complementarity between a sample and the new sample generated from it by an SFN, thus allowing the SFN to more efficiently improve the performance of the HAR network during testing. The experimental results on various datasets verify that the proposed method outperforms state-of-the-art data augmentation methods. More importantly, the proposed SFN architecture is a general framework that can be integrated with various types of networks for HAR. For example, when a baseline HAR model with three LSTM layers and one fully connected (FC) layer was used, the classification accuracy was increased from 79.53% to 90.75% on the NTU RGB+D dataset using a cross-view protocol, thus outperforming most other methods. Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Juanhui Tu, Mengyuan Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Hierarchical Dropped Convolutional Neural Network for Speed Insensitive Human Action RecognitionabstractHuman action recognition using skeleton data has lots of potential applications in content-based action retrieval and intelligent surveillance, with wide usage of depth sensors and robust skeleton estimation algorithms. Previous methods describe spatial temporal skeleton joints as a compact color image and then use Convolutional Neural Network (CNN) to extract more discriminative deep features. However, these methods ignore the effect of speed variation, which is a common phenomenon and can bring severe intra-varieties to same types of actions. To solve this problem, this paper presents a novel hierarchical dropped CNN architecture, which is constructed in two stages. Dropped CNN (d-CNN) is firstly developed to extract deep features from a probabilistic speed insensitive color image. This image expresses both spatial distributions and temporal evolutions of skeleton joints meanwhile avoids the effect of speed variations. To enhance the temporal discriminative power, we extend d-CNN to a hierarchical structure (h-CNN), where multiple scales of temporal information are encoded. Extensive experiments on benchmark MSRC-12 dataset and the largest NTU RGB+D dataset verify the effectiveness and robustness of the proposed method. Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Mengyuan Liu 0001, Wei Liu 0065 |
ICME | 3 |
| 2018 | Adaptive Weighted Sparse Principal Component AnalysisabstractIn this paper, we propose an unsupervised feature selection method from the perspective of optimal reconstruction. The features selected by the proposed method can well represent the original data, and the effectiveness of the selected features is demonstrated by robust reconstruction and clustering. The proposed method emphasizes the joint l2, 1-norms minimization on both reconstruction term and regularization term to make them be column-sparse. Relying on the column-sparse property of reconstruction term and regularization term, the proposed method is able to improve the robustness to outliers and select the effective features. The proposed objective function is nonconvex. Fortunately, it can be equivalently reformulated as a convex form (with change of variables) to capture a global optimization solution. In fact, the proposed method is related to the optimal mean robust principal component analysis (OMRPCA) because the proposed method is a sparse self-contained regression type of OMRPCA. Since OMRPCA essentially adds the adaptive weights for data samples, we call the proposed method adaptive weighted sparse principal component analysis (AW-SPCA). Experimental results demonstrate the effectiveness of AW-SPCA. Shuangyan Yi, Yongsheng Liang 0001, Wei Liu 0065, Fanyang Meng |
ICME | 2 |
| 2017 | A bidirectional adaptive bandwidth mean shift strategy for clusteringabstractThe bandwidth of a kernel function is a crucial parameter in the mean shift algorithm. This paper proposes a novel adaptive bandwidth strategy which contains three main contributions. (1) The differences among different adaptive bandwidth are analyzed. (2) A new mean shift vector based on bidirectional adaptive bandwidth is defined, which combines the advantages of different adaptive bandwidth strategies. (3) A bidirectional adaptive bandwidth mean shift (BAMS) strategy is proposed to improve the ability to escape from the local maximum density. Compared with contemporary adaptive bandwidth mean shift strategies, experiments demonstrate the effectiveness of the proposed strategy. Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Liu Wei, Jihong Pei |
ICIP | 3 |
| 2012 | A discrete-time switching neural network for quadratic programmingabstractThis paper presents a discrete-time neural network with a switching structure to solve a general quadratic programming problem in real time. Compared with existing ones for solving quadratic programming problems, the proposed neural network model has a simple architecture and uses a limited number of neurons to solve the problem, irrespective of the dimension of the decision variables or the number of constraints. The global convergence of the model is proven using contraction theory. Simulations are performed to demonstrate the effectiveness of the proposed method. Sanfeng Chen, Shuai Li 0002, Yongsheng Liang 0001, Y. Lou |
IJCNN | 3 |
| 2012 | A recurrent neural network for inter-localization of mobile phonesabstractThe fact that most mobile phones are equipped with short-range communication devices, such as Bluetooth, etc., and a portion of mobile phones have GPS embedded enables us to envision to roughly localize GPS-free phones recursively and progressively by exploiting the available information. With the position of GPS equipped phones as beacons, and with the Bluetooth connection between neighbor phones as proximity constraints, we formulate the problem as an inequality problem defined on the Bluetooth network. A recurrent neural network is developed to solve the problem distributively in real time. The convergence of the neural network and the solution feasibility to the defined problem are both theoretically proven. Two applications examples are considered and simulated. Simulations demonstrate the effectiveness of the proposed method. Shuai Li 0002, Sanfeng Chen, Y. Lou, B. Lu, Yongsheng Liang 0001 |
IJCNN | 5 |
| 2012 | Decentralized kinematic control of a class of collaborative redundant manipulators via recurrent neural networks
Shuai Li 0002, Sanfeng Chen, Bo Liu 0006, Yangming Li, Yongsheng Liang 0001 |
Neurocomputing | 5 |