VLDB 2026 Research / reviewers in the wild / expert
Guangming Shi
dblp:97/3742
· DBLP profile ↗
392ranked-venue papers
11as first author
203since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 238 · 5 first-author · 99 since 2021Artificial intelligence and machine learning · 108 · 4 first-author · 72 since 2021Computer networks · 41 · 39 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 14 since 2021Systems, architecture and hardware · 13 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Spatio-Temporal Compression Ratio Learning and Frequency-Aware Semantic Compression for Video ImagingabstractSnapshot compressive imaging (SCI) and video compressive sensing (VCS) typically use fixed, globally uniform compression ratios that ignore spatio-temporal heterogeneity. We present D-STCRL, a reinforcement-learned framework that unifies adaptive sensing and semantic transmission under an explicit rate-distortion-energy objective. The pipeline comprises: (i) a Ratio Generation Network predicts per-patch ratio maps via spatio-temporal attention and 3D frequency cues; (ii) a Programmable Sensing Model emulates pixel-wise variable exposure through differentiable binary gating under a global budget; and (iii) a Frequency-Aware Swin decoder with a low-rank prior restores temporally consistent frames. A multi-objective policy gradient couples the ratio policy with reconstruction and JSCC, yielding stable training. On the NFS benchmark, D-STCRL improves PSNR by$2-3 ~\text{dB}$over fixed-ratio SCI at the same sampling budget; under 10 dB AWGN it surpasses a CRL baseline by$0.6-1.0 ~\text{dB}$while reducing transmitted symbols by up to 15 %. These results unify content-adaptive sensing and efficient transmission for next-generation cameras. Our code, configs and reproducible pipelines will be released upon acceptance. Haixiong Li, Dahua Gao, Xiaodan Song, Guangming Shi |
DCC | 4 |
| 2026 | SMART: A Sparse MoE-Transformer Framework for Environment-aware Channel Prediction
Shihao Xie, Xubo Li, Yong Xiao 0001, Yingyu Li, Guangming Shi |
ICC | 6 |
| 2026 | Multitask Semantic-Coded Image Communication for UAV Integrated Sensing and Communication: A Channel-wise Feature Enhancement Approach
Chen Mao, Shuhang Zhang, Shuai Ma 0002, Guangming Shi, Zhihua Yang |
INFOCOM | 5 |
| 2026 | Poster: DiSC2: Bandwidth-Efficient Distributed Semantic Communication for Cross-Modal Redundancy Suppression in the Internet of Bodies
Dahua Gao, Minxi Yang, Guangming Shi |
SECON | 4 |
| 2026 | Task-Oriented Semantic Communication for Satellite Remote Sensing Images with Channel Feedback
Qi Qiu, Chen Mao, Shuhang Zhang, Guangming Shi |
WCNC | 5 |
| 2026 | A lightweight semantic decoding network with group to individual transfer learning for EEG-based visual recognition
Xiaotian Wang 0001, Doudou Zhang, Qimin Xu, Rongkai Zhang 0008, Yiming Jiang 0023, Fu Li 0002, Yang Li 0019, Guangming Shi |
Neurocomputing | 9 |
| 2026 | On the Rate-Distortion-Complexity Tradeoff for Semantic CommunicationabstractSemantic communication is a novel communication paradigm that focuses on conveying the user's intended meaning rather than the bit-wise transmission of source signals. One of the key challenges is to effectively represent and extract the semantic meaning of any given source signals. While deep learning (DL)-based solutions have shown promising results in extracting implicit semantic information from a wide range of sources, existing work often overlooks the high computational complexity inherent in both model training and inference for the DL-based encoder and decoder. To bridge this gap, this paper proposes a rate-distortion-complexity (RDC) framework which extends the classical rate-distortion theory by incorporating the constraints on semantic distance, including both the traditional bit-wise distortion metric and statistical difference-based divergence metric, and complexity measure, adopted from the theory of minimum description length and information bottleneck. We derive the closed-form theoretical results of the minimum achievable rate under given constraints on semantic distance and complexity for both Gaussian and binary semantic sources. Our theoretical results show a fundamental three-way tradeoff among achievable rate, semantic distance, and model complexity. Extensive experiments on real-world image and video datasets validate this tradeoff and further demonstrate that our information-theoretic complexity measure effectively correlates with practical computational costs, guiding efficient system design in resource-constrained scenarios. Jingxuan Chai, Yong Xiao 0001, Guangming Shi |
IEEE Internet Things J. | 3 |
| 2026 | Adaptive Semantic Compression and Transmission for Cognitive Knowledge Coordination in a Hierarchical LLM-Agents SystemabstractThe rapid development of Large Language Model (LLM) agents has facilitated the advancement of multi-agent systems, where cognitive knowledge sharing is crucial for the execution of complex tasks. However, achieving the synchronization of the cognitive Knowledge base (KB) among agents under restricted wireless resources remains a challenge, especially in dynamic real-time environments. Therefore, we propose a hierarchical LLM agent system that consists of a high-level Cluster Brain (CB) and multiple Lower-level LLM Agents (LLAs). The cognitive KB of each LLA is represented in the form of a Knowledge Graph (KG). To improve the efficiency of transmitting cognitive KB updates from LLAs to CB, a KG compression framework named MED-EmPress is proposed, which adaptively compresses the semantic features of cognitive KB by applying dimensionality reduction and binary quantization, and then a joint optimization problem of Semantic Compression and Resource Allocation (SCRA) is formulated to maximize semantic fidelity of the cognitive KB being transmitted. To solve this problem, a hierarchical SCRA algorithm is designed to decouple the SCRA problem into two subproblems, which involve dynamically allocating wireless resources and rationally choosing the semantic compression ratio of cognitive KB. The goal is to maximize the system’s semantic fidelity. The evaluation results demonstrate that the MED-EmPress framework reduces the size of the cognitive updates that need to be transmitted by 96%, with only a loss of 3.6% in the entity alignment task. Furthermore, the proposed adaptive compression and transmission scheme improves semantic fidelity by 94% compared to existing methods when wireless resources are severely limited. Xinju He, Jiayi Liu 0001, Xuemei Xie, Guangming Shi |
IEEE Internet Things J. | 5 |
| 2026 | Hypergraph Information Bottleneck-Based Implicit Semantic CommunicationabstractSemantic communication is a novel communication paradigm focusing on the transmission of meaningful, task-oriented information. Recent results have shown that graphical structures represent the most robust and structurally faithful formalism for modeling the semantic knowledge within a wide range of source signals. However, previous solutions focus primarily on the pairwise relational graphs, which inherently lack the capacity to encapsulate complex, higher-order interactions fundamental to specific semantic contexts. In contrast, hypergraphs provide a more flexible mathematical framework that can more accurately map the multidimensional dependencies found in intricate data sources. In this paper, we investigate hypergraph-based semantic representation for the semantic communication system. We propose a novel hypergraph information bottleneck-based implicit semantic communication framework (HIB-SC), in which the semantic encoder is developed and optimized to extract the minimally sufficient representation of the hypergraph-based semantic information source, that maximizes the mutual information between the encoded representation and the implicit high-order semantic relations that are intended for the receiver. We theoretically prove that the proposed framework is able to extract the most informative subgraphs or motifs that are significantly more robust against adversarial attacks with improved generalization performance. Extensive experiments verify that the proposed HIB-SC achieves superior semantic compression efficiency, higher accuracy in implicit semantic inference, and enhanced resilience against noise, compared to the state-of-the-art solutions. Yiwei Liao, Shurui Tu, Yong Xiao 0001, Yingyu Li, Guangming Shi |
IEEE Internet Things J. | 6 |
| 2026 | Optimal and Robust Beamforming Design for Digital Semantic Communication System Under QoS ConstraintsabstractDriven by the demand for high transmission efficiency in 6G networks, semantic communication has attracted significant interest recently. While most existing works focus on optimizing semantic communication system design to enhance end-to-end transmission performance, they often overlook the integration of quality of service (QoS) requirements. To address this gap, we employ the Alpha-Beta-Gamma (ABG) formula to empirically approximate the relationship between end-to-end transmission quality and signal-to-noise ratio (SNR). Based on this model, we first design an optimal beamforming scheme that minimizes transmission power while ensuring real-time QoS guarantees. Furthermore, to account for inevitable channel state information (CSI) estimation errors in practical scenarios, we propose robust beamforming design schemes under QoS constraints for both bounded and unbounded CSI estimation errors. These optimization problems are efficiently solved using semidefinite relaxation (SDR),S-lemma, and Bernstein-type inequalities. Finally, experimental results demonstrate that our proposed optimal beamforming design scheme outperforms conventional beamforming methods, while the robust beamforming schemes achieve superior performance in handling CSI estimation errors compared to existing computational approaches. Shuai Ma 0002, Hang Li 0003, Yunlong Cai, Hailiang Xiong, Shiyin Li, Guangming Shi |
IEEE Internet Things J. | 7 |
| 2026 | A Channel Adaptive Encoding and Decoding Method for Unmanned Aerial Vehicle Image TransmissionabstractUnmanned Aerial Vehicles (UAVs) are an indispensable core component of low altitude economic networks. It is very critical for achieving efficient UAV image transmission of air-to-ground communication. Due to the changes in flight area and unstable channel conditions, the signal-to-noise and transmission rate change rapidly. To adapt to these changes, we propose a channel adaptive encoding and decoding method for UAV image transmission. The proposed method includes a lightweight feature extraction module, a channel-wise feature enhance module, a transmission rate adaptive module, and the corresponding decoding module. The lightweight feature extraction module can quickly extract local detailed features and long-range spatial dependencies via residual block and mobile mamba. The channel-wise feature enhance module can enhance channel wise useful features via the involution operation and the SNR adjustment block according to channel state SNRs. The transmission rate adaptive module can further adaptively adjust the size of transmission features according to the transmission rate via the rate adjustment block and the rate mask block. The extensive experimental results on the SIRI-WHU, WHU-RS19, AID, and UCMerced Land Use datasets demonstrate that our method obtains higher PSNR, MS-SSMI, LPIPS and ACC than state-of-the-art methods. Shuhang Zhang, Qi Qiu, Bin Li 0012, Chiya Zhang, Guangming Shi |
IEEE Internet Things J. | 6 |
| 2026 | Multi-layer graph constraint dictionary pair learning for image classification
Guangming Shi, Weisheng Dong, Xuemei Xie |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | CoCoFR: Collaborative codebooks learning with soft matching strategy for blind face restoration
Teng Feng, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi |
Neural Networks | 8 |
| 2026 | Parse graph-based visual-language interaction for human pose estimation
Shibang Liu, Xuemei Xie, Guangming Shi |
Pattern Recognit. | 3 |
| 2026 | Implicit Semantic-Aware Communication Based on Hypergraph Reasoning
Yiwei Liao, Shurui Tu, Yong Xiao 0001, Yingyu Li, Guangming Shi |
IEEE Trans. Commun. | 5 |
| 2026 | SITP: A High-Reliability Semantic Information Transport Protocol Without Retransmission for Semantic CommunicationabstractWith the evolution of 6G networks, modern communication systems are facing unprecedented demands for high reliability and low latency. However, conventional transport protocols are designed for bit-level reliability, failing to meet the semantic robustness requirements. To address this limitation, this paper proposes a novel Semantic Information Transport Protocol (SITP), which achieves TCP-level reliability and UDP level latency by verifying only packet headers while retaining potentially corrupted payloads for semantic decoding. Building upon SITP, a cross-layer analytical model is established to quantify packet-loss probability across the physical, data-link, network, transport, and application layers. The model provides a unified probabilistic formulation linking signal noise rate (SNR) and packet-loss rate, offering theoretical foundation into end-to-end semantic transmission. Furthermore, a cross-image feature interleaving mechanism is developed to mitigate consecutive burst losses by redistributing semantic features across multiple correlated images, thereby enhancing robustness in burst-fade channels. Extensive experiments show that SITP offers lower latency than TCP with comparable reliability at low SNRs, while matching UDP-level latency and delivering superior reconstruction quality. In addition, the proposed cross-image semantic interleaving mechanism further demonstrates its effectiveness in mitigating degradation caused by bursty packet losses. Shuai Ma 0002, Youlong Wu, Guangming Shi, Xiang Cheng 0001 |
IEEE Trans. Commun. | 4 |
| 2026 | ARGP: Adaptive and Recoverable 3D Gaussian Splatting Pruning for Efficient Real-Time Scene Reconstructionabstract3D Gaussian Splatting (3DGS) has emerged as an efficient explicit representation for real-time novel view synthesis. However, repeated densification during training often leads to uncontrolled Gaussian growth, causing heavy memory overhead and slow convergence. This redundancy stems from two limitations: (1) opacity-based sparsification that fails to maintain consistent pruning pressure under evolving distributions; (2) irreversible pruning strategies that risk eliminating essential Gaussians. To address these, we proposeAdaptive and Recoverable Gaussian Pruning (ARGP), a training-integrated framework that suppresses redundant growth while preserving reconstruction quality. During densification, we employAdaptive Opacity Pruning (AOP), which removes a fixed proportion of low-opacity Gaussians based on quantiles, effectively suppressing redundancy. In the fine-tuning stage, we introduceIterative Recovery Pruning (IRP), which selectively reinstates critical Gaussians using a gradient-informed recovery score, thus preventing over-pruning and preserving reconstruction quality. Extensive experimental results on Mip-NeRF 360, Tanks & Temples, and Deep Blending validate the effectiveness of our proposed ARGP. For instance, ARGP achieves a 90.2% reduction in Gaussian counts and a 1.33× training speedup, while maintaining competitive reconstruction quality on Mip-NeRF 360. Our code is available at https://github.com/Sinyo-Liu/ARGP. Dahua Gao, Wenglong Wang, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | SceneReasoner: Sufficient Embodied Scene Understanding From Limited Perception by Explicit Functional AssociationabstractSufficient embodied scene understanding serves as the foundation for embodied agents to perceive, interpret, and solve scene-related questions in a scene. Such understanding is often constrained by limited perception, which can be summarized into two aspects: (1) The lack of perceptual abilities in the agent. (2) The scene itself is incomplete. Existing large models-based methods attempt to leverage implicit knowledge to overcome these limitations, but this lacks interpretability and controllability. Inspired by human explainable associative thinking, we propose SceneReasoner framework, which imposes explicit functional associative rules on LLMs to guide the process of the scene understanding. This framework mines deeper functional relationships between objects, enabling the agent to gain sufficient scene understanding from limited perception in a controllable manner. Specifically, SceneReasoner employs an associative knowledge base to provide such rules from two aspects. (1) Functional complementarity of objects in a scene. For instance, in a computer workspace, if the agent perceives a monitor and other unclear objects, it can first analyze the function of the monitor in this area (content display), and then infer the presence of other related objects (e.g., a mouse and keyboard for content input). (2) Commonality of objects in a scene. For example, in a picture-hanging task, when the hammer in the scene is missing, the agent needs to first identify the hammer’s attributes (hard and applying force), and then associate a suitable substitute in the scene (e.g., a hard wrench). Due to such explicit functional association, the agent can rapidly form a sufficient scene understanding and effectively solve scene-related questions. Experimental results demonstrate that such explicit association augmented with functional reasoning can significantly enhance agents’ scene understanding under limited perception. It improves perceptual quality by 9.75% and scene reasoning ability by 21.42% compared with other methods. Xiukun Liu, Ning Lan, Xuemei Xie, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Learning Retinex Prior for Compressive Hyperspectral Image ReconstructionabstractImage reconstruction in coded aperture snapshot spectral compressive imaging (CASSI) aims to recover high-fidelity hyperspectral images (HSIs) from compressed 2D measurements. While deep unfolding networks have shown promising performance, the degradation induced by the CASSI degradation model often introduces global illumination discrepancies in the reconstructions, creating artifacts similar to those in low-light images. To address these challenges, we propose a novel Retinex Prior-Driven Unfolding Network (RPDUN), which unfolds the optimization incorporating the Retinex prior as a regularization term into a multi-stage network. This design provides global illumination adjustment for compressed measurements, effectively compensating for spatial-spectral degradation according to physical modulation and capturing intrinsic spectral characteristics. To the best of our knowledge, this is the first application of the Retinex prior in hyperspectral image reconstruction. Furthermore, to mitigate the noise in the reflectance domain, which can be amplified during decomposition, we introduce an Adaptive Token Selection Transformer (ATST). This module adaptively filters out weakly correlated tokens before the self-attention computation, effectively reducing noise and artifacts within the recovered reflectance map. Extensive experiments on both simulated and real-world datasets demonstrate that RPDUN achieves new state-of-the-art performance, significantly improving reconstruction quality while maintaining computational efficiency. The code is available at https://github.com/ZUGE0312/RPDUN. Mengzu Liu, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2026 | Double Nonconvex Tensor Robust Kernel Principal Component Analysis and Its Visual ApplicationsabstractTensor robust principal component analysis (TRPCA), as a popular linear low-rank method, has been widely applied to various visual tasks. The mathematical process of the low-rank prior is derived from the linear latent variable model. However, for nonlinear tensor data with rich information, their nonlinear structures may break through the assumption of low-rankness and lead to the large approximation error for TRPCA. Motivated by the latent low-dimensionality of nonlinear tensors, the general paradigm of the nonlinear tensor plus sparse tensor decomposition problem, called tensor robust kernel principal component analysis (TRKPCA), is first established in this paper. To efficiently tackle TRKPCA problem, two novel nonconvex regularizers the kernelized tensor Schatten- $p$ norm (KTSPN) and generalized nonconvex regularization are designed, where the former KTSPN with tighter theoretical support adequately captures nonlinear features (i.e., implicit low-rankness) and the latter ensures the sparser structural coding, guaranteeing more robust separation results. Then by integrating their strengths, we propose a double nonconvex TRKPCA (DNTRKPCA) method to achieve our expectation. Finally, we develop an efficient optimization framework via the alternating direction multiplier method (ADMM) to implement the proposed nonconvex kernel method. Experimental results on synthetic data and several real databases show the higher competitiveness of our method compared with other state-of-the-art regularization methods. The code has been released in our ResearchGate homepage: https://www.researchgate.net/publication/397181729_DNTRKPCA_code. Jianjun Wang 0003, Wei-Shi Zheng 0001, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2026 | SANet: A Semantic-Aware Agentic AI Networking Framework for Cross-Layer Optimization in 6GabstractAgentic AI networking (AgentNet) is a novel AI-native networking paradigm in which a large number of specialized AI agents collaborate to perform autonomous decisions, dynamic environmental adaptation, and complex missions. AgentNet has the potential to facilitate real-time network management and optimization functions, including self-configuration, self-optimization, and self-adaptation across diverse and complex environments, laying the foundation for fully autonomous networking systems. Despite its promise, AgentNet is still in the early stages of development and still lacks an effective networking framework to support automatic goal discovery, multi-agent self-orchestration, and task assignment. This paper proposes SANet, a novel semantic-aware AgentNet architecture for wireless networks. SANet can infer the semantic goal of the user and automatically assign agents associated with different layers of the network stack to fulfill the inferred goal. Motivated by the fact that AgentNet is a decentralized framework in which collaborating agents may generally have different and even conflicting objectives, we formulate the decentralized optimization of SANet as a multi-agent multi-objective problem, and focus on finding the Pareto-optimal solution for agents with distinct and potentially conflicting objectives. We propose three novel metrics for evaluating SANet: (the agents' objective) optimization error, (dynamic environment) generalization error, and (multi-objective) conflicting error. Furthermore, we develop a model partition and sharing (MoPS) framework in which large models, e.g., deep learning models, of different agents can be partitioned into shared and agent-specific parts that are jointly constructed and deployed according to agents' local computational resources. Two decentralized optimization algorithms, static-weighting and dynamic-weighting algorithms, are introduced to optimize the above three metrics. A bandwidth-adaptive compression framework is also proposed to enable different agents to perform in situ compression of their intermediate embeddings, dynamically adjusting to localized resource constraints and task requirements. We derive theoretical bounds for all these performance metrics and prove that there exists a three-way tradeoff among optimization, generalization, and conflicting errors. Finally, to validate our theoretical results, we develop an open-source Radio Access Network (RAN) and core network-based hardware prototype that implements three Transformer-based time-series prediction agents to interact with three different layers of the network. Experimental results show that the proposed MoPS framework achieves performance gains of up to$14.61\%$while requiring only$44.37\%$of the Floating-Point Operations (FLOPs) for inference at each agent compared to state-of-the-art algorithms. Also, compared to the static-weighting algorithm, the dynamic-weighting algorithm achieves up to$83.81\%$reduction in training errors caused by conflicting objectives. Yong Xiao 0001, Xubo Li, Yingyu Li, Yayu Gao, Guangming Shi, Ping Zhang 0003, Marwan Krunz |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Hierarchical Scene Graph Generation With Coarse-to-Fine ReasoningabstractScene graph generation denotes the process of parsing a visual input into a graph representation for downstream reasoning tasks. Most existing research models input scenes as flat scene graphs, capturing solely on horizontal relationships between two objects, while disregarding vertical dependencies. Such flat scene graphs merely describe direct relationships between objects, and cannot represent higher-level semantics. The lack of hierarchical relationships also limits the performance of scene graphs in downstream reasoning tasks. In this work, we propose hierarchical scene graph generation, a new problem task that requires the model to generate scene graphs with clear hierarchy. To achieve this goal, a hierarchical scene graph generation with coarse-to-fine reasoning framework is proposed. It contains three stages: first, summarize the main caption of the scene; second, describe the main visual elements related to the theme; and finally, add secondary elements with background. To facilitate training, a high-quality hierarchical scene graph dataset that comprises 40 K image-scene description pairs is constructed. Utilize rich knowledge and powerful multi-turn conversation capabilities of multi-modal large language models, the proposed framework can not only achieve better semantic understanding and relational modeling, but also seamlessly integrating its relational modeling capabilities to enhance visual-language tasks. Extensive experimental results demonstrate that our method constructs an effective scene representation, outperforming the state-of-the-art by 3.35% in recall while generating significantly fewer average triplets 8.7vs. 9.2. The results on downstream tasks also indicate that hierarchical scene graph significantly contributes to visual-language interactive reasoning. Zefang Han, Xiukun Liu, Xuemei Xie, Jianxiu Yang, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2026 | CMANet: Context-Aware Mutual Attention Network for Referring Image Segmentation
Xiong Pan, Xuemei Xie, Jianxiu Yang, Xiaodan Song, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2026 | Improved Spontaneous EEG Signal Decoding Efficiency by Function Predefined Convolutional Neural NetworkabstractA spontaneous electroencephalogram (EEG)-based brain-computer interface (BCI) is an ideal form of brain-computer interaction. The classical decoding methods can achieve classification by using meaningful manual features, but their performance is poor. The neural network (NN) methods have significantly improved the performance, but their interpretability and computational efficiency are much lower than those of the classical methods. This is because NN abandons the strong a priori knowledge of neuroscience and completely relies on training to extract EEG features. How to integrate the characteristics of neural signals into the design of the basic operator of the NNs while retaining its learning ability is the focus of this work. In this work, we proposed a function predefined convolutional NN (FPCNN) to search for the best frequency points and channel weights to decode spontaneous EEG signals. Among the FPCNN, a novel function predefined convolutional (FPC) layer adopts a learnable way to search for the key spatial-frequency parameters of spontaneous EEG, making its parameters have clear physical meanings. Furthermore, a trainable quadrature detector (TQD) based on FPC was constructed, and the quadrature characteristic was utilized to ensure the capture of complex phase change signals. The core contribution of our method lies in the proposal of a novel NN operator for decoding spontaneous EEG, and a quadrature scheme for handling the phase changes of signals. The experimental results show that the proposed FPCNN significantly improves the performance by 2.09% ( ${}^{\ast } $ ), 3.08% ( ${}^{\ast } $ ), and 3.41% ( ${}^{\ast \ast }$ ), respectively, compared with the state-of-the-art (SOTA) methods on three spontaneous EEG datasets. Moreover, the training and testing time cost of FPCNN in a non-GPU environment only takes 67.96 and 19.36 s per epoch. Its savings in computing resources and time are very beneficial for EEG processing in diverse environments. In addition, visualization experiments demonstrated the interpretability and stability of the proposed FPCNN. The experimental results show that our method is efficient, stable, and interpretable. This work has effectively improved the decoding efficiency of spontaneous EEG signals and demonstrated the power of combining traditional signal processing methods with NNs. Boxun Fu, Fu Li 0002, Junkai Li, Youshuo Ji, Yang Li 0019, Yinghui Quan, Lijian Zhang, Guangming Shi |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Asymmetric Hierarchical Difference-aware Interaction Network for Event-guided Motion DeblurringabstractEvent cameras are bio-inspired sensors that are capable of capturing motion information with high temporal resolution, which show potential in aiding image motion deblurring recently. Most existing methods indiscriminately handle feature fusion of two modalities with symmetric unidirectional/bidirectional interactions at different-level layers in feature encoder, while ignoring the different dependencies between cross-modal hierarchical features. To tackle these limitations, we propose a novel Asymmetric Hierarchical Difference-aware Interaction Network (AHDINet) for event-based motion deblurring, which explores the complementarity of two modalities with differential dependence modeling of cross-modal hierarchical features. Thereby, an event-assisted edge complement module is designed to leverage event modality to enhance the edge details of the image features in low-level encoder stage, and an image-assisted semantic complement module is developed to transfer contextual semantics of image features to event branch in high-level encoder stage. Benefiting from the proposed differentiated interaction mode, the respective advantages of image and event modalities are fully exploited. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance. Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
AAAI | 5 |
| 2025 | Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object DetectionabstractTiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture to enhance the extinguished regions. We for the first time exploit the regions to be enhanced from the perspective of pixel-wise amount of information. Specifically, we model the entire image pixels feature information by minimizing Information Entropy loss, generating an information map to attentively highlight weak activated regions in an unsupervised way. To effectively assist the above phase with more attention to tiny objects, we next introduce the Position Gaussian Distribution Map, explicitly modeled using a Gaussian Mixture distribution, where each Gaussian component's parameters depend on the position and size of object instance labels, serving as supervision for further feature enhancement. Taking the information map as prior knowledge guidance, we construct a multi-scale position gaussian distribution map prediction module, simultaneously modulating the information map and distribution map to focus on tiny objects during training. Extensive experiments on three public tiny object datasets demonstrate the superiority of our method over current state-of-the-art competitors. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
CVPR | 7 |
| 2025 | Parameterized Blur Kernel Prior Learning for Local Motion Deblurring
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Guangming Shi |
CVPR | 7 |
| 2025 | Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesabstractRecent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impact of distributional shifts. In this paper, we observe that classification errors arising from distribution shifts tend to cluster near the true values, suggesting that misclassifications commonly occur in semantically similar, neighboring categories. Furthermore, robust advanced vision foundation models maintain larger inter-class distances while preserving semantic consistency, making them less vulnerable to such shifts. Building on these findings, we propose a new method called GFN (Gain From Neighbors), which uses gradient priors from neighboring classes to perturb input images and incorporates an inter-class distance-weighted loss to improve class separation. This approach encourages the model to learn more resilient features from data prone to errors, enhancing its robustness against shifts in diverse settings. In extensive experiments across various model architectures and benchmark datasets, GFN consistently demonstrated superior performance. For instance, compared to the current state-of-the-art TAPADL method, our approach achieved a higher corruption robustness of 41.4% on ImageNet-C (+2.3%), without requiring additional parameters and using only minimal data. Mingtao Feng, Weisheng Dong, Xin Li 0005, Guangming Shi |
CVPR | 7 |
| 2025 | Simultaneous Denoising and Compression for DVS with Partitioned Cache-Like Spatiotemporal FilterabstractDynamic vision sensor (DVS) is a novel neuromorphic imaging device that asynchronously generates event data corresponding to changes in light intensity at each pixel. However, the differential imaging paradigm of DVS renders it highly sensitive to background noise. Additionally, the substantial volume of event data produced in a very short time presents significant challenges for data transmission and processing. In this work, we present a novel spatiotemporal filter design, named PCLF, to achieve simultaneous denoising and compression for the first time. The PCLF employs a hierarchical memory structure that utilizes symmetric multi-bank cache-like row and column memories to store event data from a partitioned pixel array, which exhibits low memory complexity of O(m + n) for an$\mathrm{m}\times \mathrm{n}$DVS. Furthermore, we propose a probability-based criterion to effectively control the compression ratio. We have implemented our design on an FPGA, demonstrating capabilities for real-time operation$(\leq 60\ \text{ns})$and low power consumption$(< 200\text{mW})$. Extensive experiments conducted on real-world DVS data across various tasks indicate that our design enables a reduction of event data by 30% to 68%, while maintaining or even enhancing the performance of the tasks. Qinghang Zhao, Yixi Ji, Jinjian Wu, Guangming Shi |
DATE | 5 |
| 2025 | Affine Transformation-Based Generative Face Video CompressionabstractIn this paper, we propose a generative face video compression framework based on affine transformations to better represent large movements without parameter transmission. It mainly consists of an encoder and decoder, and our encoder is similar to the one in [1]. Intra frame are compressed by the existing encoder, while subsequent inter frames are compressed into compact inter frame features. In the decoder, feature alignment is first established to map the decoded intra frame and inter frame features into the same domain. The aligned features are then combined with the appearance features extracted by the appearance encoder from the intra frame and fed into the coarse-fine affine transform module to establish motion estimation and compensation. The coarse affine transform focuses on global motion, while the fine affine transform deals with local motion, such as lip motion. Finally, the transformed features are fed into the image generation module to obtain the final reconstruction results. Xihua Lin, Xiaodan Song, Xuguang Zuo, Dahua Gao, Xuemei Xie, Guangming Shi |
DCC | 7 |
| 2025 | Skillsets on the Chain: A Blockchain-based Trustworthy Agentic AI Networking FrameworkabstractAgentic AI networking (AgentNet) has attracted significant interest due to its promising potential to move traditional AI-based networking solutions beyond closed-loop and passive learning to proactive interaction and goal-driven action, offering a path to self-learning and generally intelligent networking systems. Despite its promise, ensuring the security and trustworthiness of such systems presents significant challenges, particularly concerning identity management, agent capability verification, and data integrity during collaborative learning. To address these issues, this paper proposes TrustAgentNet, a novel consortium blockchain-based framework for unified and trusted agent identification, traceable skillset and tag descriptions, and secure on-chain collaborative learning in AgentNet. In TrustAgentNet, a chain of skillset (CoS) is introduced, consisting of a skillset chain to distributedly store all the verified skillsets and associated tags, and a dedicated training chain for each distinct skillset can be jointly constructed and maintained by the authorized agents using a collaborative learning-based approach. Theoretical analysis suggests that there exists a three-way trade-off among the security level, skillset performance, and resource cost. This tradeoff is also empirically validated by the experimental results obtained from a hardware prototype implemented based on a Hyperledger Fabric-based consortium blockchain. To verify the practical performance of TrustAgentNet, we consider a real-world scenario of multi-agent collaborative learning under malicious attack. Experimental results suggest that TrustAgentNet can effectively guarantee the security of skillset training and enable rapid response and recovery from potential attacks within seconds. Yayu Gao, Yong Xiao 0001, Xubo Li, Aoyu Hu, Yingyu Li, Guangming Shi, Ping Zhang 0003 |
GLOBECOM | 8 |
| 2025 | NOC4SC: A Novel Paradigm of Multi-User Semantic Communication Enabled by Non-Orthogonal CodewordsabstractThe limited bandwidth of 6G networks presents a significant challenge in accommodating the growing number of user connections. To tackle this issue, we propose a novel multi-user access paradigm for SemCom, termed non-orthogonal code-words for semantic communication (NOC4SC). By leveraging the Swin Transformer, NOC4SC enables users to independently extract their semantic features while maintaining network parameter sharing. NOC4SC ensures that the user’s data remains protected from unauthorized decoding while allowing simultaneous multi-user transmissions over shared time-frequency resources without the need for spread spectrum. Specifically, we introduce an adaptive NOC and SNR Modulation (NSM) block, which employs deep learning techniques to adaptively regulate the SNR mechanism. The NSM block facilitates the generation of approximately orthogonal semantic features across separable feature subspaces, thereby mitigating inter-user interference. Extensive experimental results demonstrate the superiority of NOC4SC, achieving a 5.02 dB improvement in PSNR at low SNR and 2.12dB at high SNR in the Rayleigh fading channel. Shuai Ma 0002, Dahua Gao, Guangming Shi |
GLOBECOM | 5 |
| 2025 | SANNet: A Semantic-Aware Agentic AI Networking Framework for Multi-Agent Cross-Layer CoordinationabstractAgentic AI networking (AgentNet) is a novel AI-native networking paradigm that relies on a large number of specialized AI agents to collaborate and coordinate for autonomous decision-making, dynamic environmental adaptation, and complex goal achievement. It has the potential to facilitate real-time network management alongside capabilities for self-configuration, self-optimization, and self-adaptation across diverse and complex networking environments, laying the foundation for fully autonomous networking systems in the future. Despite its promise, AgentNet is still in the early stage of development, and there still lacks an effective networking framework to support automatic goal discovery and multi-agent self-orchestration and task assignment. This paper proposes SANNet, a novel semantic-aware agentic AI networking architecture that can infer the semantic goal of the user and automatically assign agents associated with different layers of a mobile system to fulfill the inferred goal. Motivated by the fact that one of the major challenges in AgentNet is that different agents may have different and even conflicting objectives when collaborating for certain goals, we introduce a dynamic weighting-based conflict-resolving mechanism to address this issue. We prove that SANNet can provide theoretical guarantee in both conflict-resolving and model generalization performance for multi-agent collaboration in dynamic environment. We develop a hardware prototype of SANNet based on the open RAN and 5GS core platform. Our experimental results show that SANNet can significantly improve the performance of multi-agent networking systems, even when agents with conflicting objectives are selected to collaborate for the same goal. Yong Xiao 0001, Xubo Li, Yayu Gao, Guangming Shi, Ping Zhang 0003 |
GLOBECOM | 5 |
| 2025 | Multi-User Information Bottleneck for Semantic-Aware CommunicationabstractSemantic-aware communication (SAC) has attracted significant interest recently due to its potential to revolutionize the traditional communication framework by focusing on delivering the meaning of information, enabling more efficient, reliable, and intelligent communication. Most existing AI-based solutions for multi-user SAC focus on developing encoding models for all users and a decoding model for a specific receiver. While powerful, this single-task-oriented codec design faces significant challenges when deployed in more general multi-task scenarios. In this paper, we investigate codec design problems in multi-user multitask SAC based on distributed information bottleneck (DIB) theory. We propose a novel task-aware DIB scheme (TADIB) for joint codec designing. In TADIB, the receiver, when performing a specific task, first estimates the relevance between users' datasets and the task based on their mutual information, which is incorporated into the training phase of both encoders and decoders. Then most task-relevant users are selected to send the encoded signals for task inference during the deployment phase. Furthermore, we employ variational approximation to derive tractable upper bounds for the DIB-based objective, which would otherwise be computationally prohibitive for high-dimensional data. This approach also reduces the computational complexity of the user selection procedures. Extensive results show that the proposed TADIB can achieve up to 3.38% improvements in inference accuracy of classification tasks, compared to existing solutions for task-oriented SAC. Rulong Wang, Yong Xiao 0001, Yingyu Li, Guangming Shi |
ICC | 6 |
| 2025 | Triad: Empowering LMM-Based Anomaly Detection with Expert-Guided Region-of-Interest Tokenizer and Manufacturing Process
Yuanze Li, Shihao Yuan, Haolin Wang 0004, Qizhang Li, Ming Liu 0018, Guangming Shi, Wangmeng Zuo |
ICCV | 7 |
| 2025 | Redesigning Upsampling in Decoders with Aligned Feature Aggregation for Semantic SegmentationabstractHierarchical decoders are widely employed in semantic segmentation architectures, where feature upsampling and fusion are critical at each layer. However, inaccuracies in feature position restoration during upsampling and misalignments in multi-scale feature fusion limit model performance. To address these issues, we propose the Aligned Feature Aggregation Upsampler (AFAU), a novel upsampling module. AFAU consists of two types of Aligned Feature Aggregation Modules (AFAM), which partition feature maps from both the encoder and decoder into multiple aligned patches. By applying attention mechanisms between corresponding patches, AFAM facilitates more effective contextual information fusion. Additionally, AFAM corrects feature position inaccuracies by calculating the similarity between linear embeddings of low-level (query) features and high-level features. This process allows semantic information aggregation guided by fine-grained high-resolution features. We conducted extensive experiments across multiple advanced semantic segmentation frameworks, demonstrating the universality and efficiency of AFAU. Notably, our results indicate that decoders incorporating AFAU significantly outperform SegFormer across all sizes of MiT encoders on the ADE20K and Cityscapes datasets. Qinjie Hu, Fei Qi 0001, Kaiwen Fu, Chengyuan Chang, Xiaotian Wang 0001, Kun Liu 0015, Guangming Shi |
ICME | 7 |
| 2025 | PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote SensingabstractRemote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patterncom in RSCIR has significant flaws, including the lack of query text sentences and paired triplets, thus making it unable to evaluate the latest methods. To address this, we propose the Zero-Shot Query Text Generator (ZS-QTG) that can generate full query text sentences based on attributes, and then, by capitalizing on ZS-QTG, we develop the PatternCIR benchmark. Pattern CIR rectifies Patterncom’s deficiencies and enables the evaluation of existing methods. Additionally, we explore zero-shot composed image retrieval methods that do not rely on massive pre-collected triplets for training. Existing methods use only the text during retrieval, performing poorly in RSCIR. To improve this, we propose Text-image Sequential Training of Composed Image Retrieval (TisCIR). TisCIR undergoes sequential training of multiple self-masking projection and fine-grained image attention modules, which endows it with the capacity to filter out conflicting information between the image and text, enhancing the retrieval by utilizing both modalities in harmony. TisCIR outperforms existing methods by 12.40% to 62.03% on PatternCIR, achieving state-of-the-art performance in RSCIR. The data and code are available here. Zhechun Liang, Shiwen Xue, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 8 |
| 2025 | Beyond Visual Quality: Fidelity-Oriented Diffusion Model for Real-world Image Super-ResolutionabstractAlthough existing diffusion-based image super-resolution methods have achieved remarkable visual quality, they often struggle with fidelity issues, particularly in preserving consistency with the original input image. This issue arises because using low-quality images as conditional inputs introduces substantial errors in the diffusion backward denoising process, making the restored features deviate from target features and thus degrade image fidelity. To improve the accuracy of noise estimation, we propose a dual-memory module to reinforce the input low-quality conditional features, which consists of a pre-trained high-quality memory bank to enrich the structural information and a degradation memory to remove the degradation components. Furthermore, we develop an uncertainty-aware noise estimation framework, utilizing an extra branch in the denoising network to predict pixel-wise uncertainty values, thus dynamically adjust the optimization weights for high-uncertainty regions. This adaptive strategy effectively improves the accuracy of noise estimation in challenging reconstruction areas. Experimental results demonstrate that our method significantly enhances the fidelity while preserving high visual quality of diffusion-based super-resolution, improving the reliability of diffusion applications. Zhenxuan Fang, Shuaibo Wang, Weisheng Dong, Xin Li 0005, Guangming Shi |
ACM Multimedia | 7 |
| 2025 | Multi-Scale Adaptive Skeleton Transformer for action recognition
Xiaotian Wang 0001, Zhifu Zhao, Guangming Shi, Xuemei Xie, Xiang Jiang 0011 |
Comput. Vis. Image Underst. | 4 |
| 2025 | Compressing Vision Transformer from the View of Model Property in Frequency Domain
Zhenyu Wang 0008, Xuemei Xie, Hao Luo 0004, Weisheng Dong, Yongxu Liu 0001, Fan Wang 0019, Guangming Shi |
Int. J. Comput. Vis. | 10 |
| 2025 | Searching efficient network with lightweight optimization for image super-resolution
Ruina Shen, Zhangheng Peng, Weisheng Dong, Guangming Shi |
Neurocomputing | 6 |
| 2025 | Distributed Network Slicing for Time-Sensitive Edge Learning in Edge Computing-Supported IoT NetworksabstractNetwork slicing is one of the key enablers for B5G and 6G to support diversified IoT services and application scenarios. In this paper, the problem of network slicing for supporting time-sensitive edge learning in massive-scale IoT networks is studied. In particular, a novel distributed network slicing framework based on a new control plane entity, called D-orchestrator, is proposed. This framework can jointly optimize the allocation and orchestration of communication and edge computational resources without requiring exchanges of the local data or resource information between base stations (BSs) and edge servers. A distributed joint resource allocation algorithm is developed based on the alternating direction method of multipliers with partial variable splitting (DistADMM-PVS) that minimizes the average service response-time of a set of service instances when the coordination among the D-orchestrator, BSs, and edge servers is perfectly synchronized. Motivated by the observation that the synchronization of coordination may result in high coordination delay that can be intolerable in many practical scenarios, particularly for large IoT networks, a novel asynchronized ADMM (AsyncADMM) algorithm is proposed. In AsyncADMM, the D-orchestrator, BSs, and edge servers can be coordinated asynchronously. AsyncADMM is then shown to converge to the global optimal solution with improved scalability and negligible coordination delay. The performance of the proposed framework is evaluated using two-month of traffic data collected in an in-campus smart transportation system supported by a 5G network. Extensive simulations are conducted for both pedestrian and vehicular-related services during peak and non-peak hours. Simulation results show that the proposed distributed network slicing framework offers a significant reduction in the service response time for both supported services. Yingyu Li, Yong Xiao 0001, Xiaohu Ge, Guangming Shi, Walid Saad 0001 |
IEEE Internet Things J. | 5 |
| 2025 | Distributed Optimization of Resource Efficiency for Federated Edge Intelligence in AAV-Enabled IoT NetworksabstractAutonomous aerial vehicles (AAV)-enabled IoT networks have shown promising potential in a range of novel applications and service scenarios, such as extending the network coverage, extending the battery lifetime of IoT networks, and also supporting temporary data collecting and processing needs in various emergency situations. This article studies a federated edge intelligence (FEI) network based on the data collected and uploaded by a AAV-enabled IoT network. More specifically, a set of AAVs has periodically collected and uploaded the data generated by an IoT network to a set of edge servers. Edge servers will then collaboratively construct shared models based on the uploaded datasets. The data uploading performance of a AAV-enabled IoT network and the computational capacity of edge servers are entangled with each other in influencing the overall model training process. We propose a new framework called AAV-enabled IoT network for FEI (U-FEI). This framework enables edge servers to assess how many data samples need to be collected based on the energy costs of the AAV-enabled IoT network. It also considers the local data processing capacity of the edge servers. As a result, the edge servers can request just the right amount of data from the AAVs, which is enough to train a satisfactory model. We evaluate the energy cost for data uploading of AAVs when the data can be uploaded from two different types of frequencies: 1) licensed bands (e.g., using 5G) and 2) unlicensed bands (e.g., using Wi-Fi, ZigBee, or 5G NR-U). We prove that the cost minimization problem of the entire AAV-enabled IoT network is separable and can be divided into a set of subproblems, each of which can be solved by an individual edge server. We also introduce a mapping function to quantify the computational load of edge servers under the combinations of three key parameters: 1) size of the dataset; 2) local batch size; and 3) number of local training passes. Finally, we adopt an alternative direction method of multipliers (ADMM)-based approach to jointly optimize the energy cost of the AAV-enabled IoT network and average resource utilization of edge servers. We prove that our proposed algorithm does not cause any data leakage nor disclose any topological information of the AAV-enabled IoT networks. Simulation results show that our proposed framework significantly improves the resource efficiency of both the AAV-enabled IoT network and edge servers. Yingyu Li, Jiangying Rao, Yong Xiao 0001, Xiaohu Ge, Guangming Shi |
IEEE Internet Things J. | 6 |
| 2025 | Semantic Feature Division Multiple Access for Digital Semantic Broadcast ChannelsabstractIn this article, we propose a digital semantic feature division multiple access (SFDMA) paradigm in multiuser broadcast (broadcast communication (BC)) networks for the inference and the image reconstruction tasks. In this SFDMA scheme, the multiuser semantic information is encoded into discrete approximately orthogonal representations, and the encoded semantic features of multiple users can be simultaneously transmitted in the same time-frequency resource. Specifically, for inference tasks, we design a SFDMA digital BC network based on robust information bottleneck (RIB), which can achieve a tradeoff between inference performance, data compression and multiuser interference. Moreover, for image reconstruction tasks, we develop a SFDMA digital BC network by utilizing a Swin Transformer, which significantly reduces multiuser interference. More importantly, SFDMA can protect the privacy of users’ semantic information, in which each receiver can only decode its own semantic information. Furthermore, we establish a relationship between performance and signal to interference plus noise ratio (SINR), which is fitted by an Alpha-Beta–Gamma (ABG) function. Furthermore, an optimal power allocation method is developed for the inference and reconstruction tasks. Extensive simulations verify the effectiveness and superiority of our proposed SFDMA scheme. Shuai Ma 0002, Zhiye Sun, Youlong Wu, Hang Li 0003, Guangming Shi, Shiyin Li, Naofal Al-Dhahir |
IEEE Internet Things J. | 6 |
| 2025 | Optimal and Robust Beamforming Design for Multiuser Semantic Interference Networks
Shuai Ma 0002, Chuanhui Zhang, Hang Li 0003, Nan Li 0011, Jinjin Chai, Chuan Huang 0001, Shiyin Li, Guangming Shi |
IEEE Internet Things J. | 9 |
| 2025 | Recognizing human-object interactions in videos with the supervision of natural language
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
Neural Networks | 4 |
| 2025 | Fast Window-Based Event Denoising With Spatiotemporal Correlation EnhancementabstractPrevious deep learning-based event denoising methods mostly suffer from poor interpretability and difficulty in real-time processing due to their complex architecture designs. In this paper, we propose window-based event denoising, which simultaneously deals with a stack of events while existing element-based denoising focuses on one event each time. Besides, we give the theoretical analysis based on probability distributions in both temporal and spatial domains to improve interpretability. In temporal domain, we use timestamp deviations between processing events and central event to judge the temporal correlation and filter out temporal-irrelevant events. In spatial domain, we choose maximum a posteriori (MAP) to discriminate real-world event and noise and use the learned convolutional sparse coding to optimize the objective function. Based on the theoretical analysis, we build Temporal Window (TW) module and Soft Spatial Feature Embedding (SSFE) module to process temporal and spatial information separately, and construct a novel multi-scale window-based event denoising network, named WedNet. The high denoising accuracy and fast running speed of our WedNet enables us to achieve real-time denoising in complex scenes. Extensive experimental results verify the effectiveness and robustness of our WedNet. Our algorithm can remove event noise effectively and efficiently and improve the performance of downstream tasks. Huachen Fang, Jinjian Wu, Qibin Hou, Weisheng Dong, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Multi-Modality Multi-Attribute Contrastive Pre-Training for Image Aesthetics ComputingabstractIn the Image Aesthetics Computing (IAC) field, most prior methods leveraged the off-the-shelf backbones pre-trained on the large-scale ImageNet database. While these pre-trained backbones have achieved notable success, they often overemphasize object-level semantics and fail to capture the high-level concepts of image aesthetics, which may only achieve suboptimal performances. To tackle this long-neglected problem, we propose a multi-modality multi-attribute contrastive pre-training framework, targeting at constructing an alternative to ImageNet-based pre-training for IAC. Specifically, the proposed framework consists of two main aspects. 1) We build a multi-attribute image description database with human feedback, leveraging the competent image understanding capability of the multi-modality large language model to generate rich aesthetic descriptions. 2) To better adapt models to aesthetic computing tasks, we integrate the image-based visual features with the attribute-based text features, and map the integrated features into different embedding spaces, based on which the multi-attribute contrastive learning is proposed for obtaining more comprehensive aesthetic representation. To alleviate the distribution shift encountered when transitioning from the general visual domain to the aesthetic domain, we further propose a semantic affinity loss to restrain the content information and enhance model generalization. Extensive experiments demonstrate that the proposed framework sets new state-of-the-arts for IAC tasks. Yipo Huang, Leida Li, Pengfei Chen 0003, Haoning Wu 0001, Weisi Lin, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Self-Supervised Learning of LiDAR 3D Point Clouds via 2D-3D Neural CalibrationabstractThis paper introduces a novel self-supervised learning framework for enhancing 3D perception in autonomous driving scenes. Specifically, our approach, namely NCLR, focuses on 2D-3D neural calibration, a novel pretext task that estimates the rigid pose aligning camera and LiDAR coordinate systems. First, we propose the learnable transformation alignment to bridge the domain gap between image and point cloud data, converting features into a unified representation space for effective comparison and matching. Second, we identify the overlapping area between the image and point cloud with the fused features. Third, we establish dense 2D-3D correspondences to estimate the rigid pose. The framework not only learns fine-grained matching from points to pixels but also achieves alignment of the image and point cloud at a holistic level, understanding the LiDAR-to-camera extrinsic parameters. We demonstrate the efficacy of NCLR by applying the pre-trained backbone to downstream tasks, such as LiDAR-based 3D semantic segmentation, object detection, and panoptic segmentation. Comprehensive experiments on various datasets illustrate the superiority of NCLR over existing self-supervised methods. The results confirm that joint learning from different modalities significantly enhances the network's understanding abilities and effectiveness of learned representation. Yifan Zhang 0036, Junhui Hou, Jinjian Wu, Yixuan Yuan, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | An event-based motion scene feature extraction framework
Zhaoxin Liu, Jinjian Wu, Guangming Shi, Wen Yang 0008, Jupo Ma |
Pattern Recognit. | 3 |
| 2025 | Growing-before-pruning: A progressive neural architecture search strategy via group sparsity and deterministic annealingabstractNetwork pruning is a widely studied technique of obtaining compact representations from over-parameterized deep convolutional neural networks . Existing pruning methods are based on finding an optimal combination of pruned filters in the fixed search space . However, the optimality of those methods is often questionable due to limited search space and pruning choices - e.g., the difficulty with removing the entire layer and the risk of unexpected performance degradation . Inspired by the exploration vs. exploitation trade-off in reinforcement learning, we propose to reconstruct the filter space without increasing the model capacity and prune them by exploiting group sparsity . Our approach challenges the conventional wisdom by advocating the strategy of Growing-before-Pruning (GbP), which allows us to explore more space before exploiting the power of architecture search. Meanwhile, to achieve more efficient pruning, we propose to measure the importance of filters by global group sparsity , which extends the existing Gaussian scale mixture model. Such global characterization of sparsity in the filter space leads to a novel deterministic annealing strategy for progressively pruning the filters. We have evaluated our method on several popular datasets and network architectures. Our extensive experiment results have shown that the proposed method advances the current state-of-the-art. Xiaotong Lu, Weisheng Dong, Zhenxuan Fang, Jie Lin 0008, Xin Li 0005, Guangming Shi |
Pattern Recognit. | 6 |
| 2025 | Modeling and Performance Analysis for Semantic Communications Based on Empirical ResultsabstractDue to the black-box characteristics of deep learning based semantic encoders and decoders, finding a tractable method for the performance analysis of semantic communications is a challenging problem. In this paper, we propose an Alpha-Beta-Gamma (ABG) formula to model the relationship between the end-to-end measurement and SNR, which can be applied for both image reconstruction tasks and inference tasks. Specifically, for image reconstruction tasks, the proposed ABG formula can well fit the commonly used DL networks, such as SCUNet, and Vision Transformer, for semantic encoding with the multi scale-structural similarity index measure (MS-SSIM) measurement. Furthermore, we find that the upper bound of the MS-SSIM depends on the number of quantized output bits of semantic encoders, and we also propose a closed-form expression to fit the relationship between the MS-SSIM and quantized output bits. To the best of our knowledge, this is the first theoretical expression between end-to-end performance metrics and SNR for semantic communications. Based on the proposed ABG formula, we investigate an adaptive power control scheme for semantic communications over random fading channels, which can effectively guarantee quality of service (QoS) for semantic communications, and then design the optimal power allocation scheme to maximize the energy efficiency of the semantic communication system. Furthermore, by exploiting the bisection algorithm, we develop the power allocation scheme to maximize the minimum QoS of multiple users for OFDMA downlink semantic communication Extensive simulations verify the effectiveness and superiority of the proposed ABG formula and power allocation schemes. Shuai Ma 0002, Chuanhui Zhang, Youlong Wu, Hang Li 0003, Shiyin Li, Guangming Shi, Naofal Al-Dhahir |
IEEE Trans. Commun. | 7 |
| 2025 | Visual Environment-Interactive Planning for Embodied Complex-Question AnsweringabstractThis study focuses on Embodied Complex-Question Answering task, which means the embodied robot need to understand human questions with intricate structures and abstract semantics. The core of this task lies in making appropriate plans based on the perception of the visual environment. Existing methods often generate plans in a once-for-all manner,$i.e$., one-step planning. Such approach rely on large models, without sufficient understanding of the environment. Considering multi-step planning, the framework for formulating plans in a sequential manner is proposed in this paper. To ensure the ability of our framework to tackle complex questions, we create a structured semantic space, where hierarchical visual perception and chain expression of the question essence can achieve iterative interaction. This space makes sequential task planning possible. Within the framework, we first parse human natural language based on a visual hierarchical scene graph, which can clarify the intention of the question. Then, we incorporate external rules to make a plan for current step, weakening the reliance on large models. Every plan is generated based on feedback from visual perception, with multiple rounds of interaction until an answer is obtained. This approach enables continuous feedback and adjustment, allowing the robot to optimize its action strategy. To test our framework, we contribute a new dataset with more complex questions. Experimental results demonstrate that our approach performs excellently and stably on complex tasks. And also, the feasibility of our approach in real-world scenarios has been established, indicating its practical applicability. Ning Lan, Baoshan Ou, Xuemei Xie, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Scene Prior Constrained Self-Paced Learning for Unsupervised Satellite Video Vehicle DetectionabstractRecently, deep learning has significantly advanced the satellite object detection. However, the effectiveness of these methods heavily relies on abundant and accurate annotations, which are extremely labor-intensive for satellite videos. Meanwhile, the robustness of traditional difference-based methods is limited by the hand-craft feature from the satellite videos with low-resolution and frame misalignment. To address this problem, an unsupervised deep satellite video vehicle detection framework based on scene prior constrained and self-paced learning (S-SPL) is proposed in this paper. S-SPL obtains the initial pseudo label by the difference-based methods, and employs a deep learning-based detector and refiner to detect objects and update labels respectively. In the train phase, to alleviate the deviation of feature expression caused by noise samples, a novel cooperative self-paced learning scheme is designed to improve the label quality and model accuracy in an alternating optimization manner. Furthermore, considering the semantic relationship between the scene and the object distribution, multi-cue prior knowledge is introduced to provide scene-level constraints, with which samples in high-confidence scenes are emphasized to improve the self-paced learning process. The experimental results on Jilin-1 and SkySat satellite videos demonstrate the superiority of S-SPL. Yuping Liang, Guangming Shi, Jinjian Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Distilling Hierarchical Knowledge From Multimodal Fusion for Unimodal Image SegmentationabstractThe application of multimodal image fusion has become increasingly widespread across various fields in the era of deep learning. Existing fusion methods integrate infrared and visible images to provide complementary content and enhance the robustness of complex real-world scenes for high-level visual tasks, such as semantic segmentation and object detection. In return, high-level visual tasks facilitate the fusion of infrared and visible by providing mid-level semantic information. However, such frameworks rely heavily on multimodal data and require strict registration of images from different modalities before fusion, seriously limiting their practical applications due to the common realistic situations of missing modalities or misregistration. To move beyond this limitation, we propose a novel hierarchical knowledge distillation (HKD) framework tailored for unimodal image segmentation with the guidance of multi-modality. This framework aims to retain as much diverse information from multimodal image fusion as possible, thereby enhancing downstream high-level visual tasks when only the unimodal images are available during the inference phase. Our proposed method is two-stage, and we construct a robust multimodal fusion and segmentation interaction network in the first stage as a powerful teacher model. In the second stage, we design a hierarchical distillation method to transfer the fused and segmented multi-layer knowledge from the multimodal teacher model to the unimodal student model. Extensive experimental results on two public datasets, i.e., MFNet and FMB, demonstrate that the proposed hierarchical knowledge distillation framework can effectively transfuse multimodal knowledge into the unimodal student model for image enhancement and segmentation under incomplete multimodal conditions, and achieves considerably competitive results compared to multimodal image fusion and segmentation models. Weisheng Dong, Shuaibo Wang, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | CDS-Net: Contextual Difference Sensitivity Network for Pixel-Wise Road Crack DetectionabstractRoad crack detection is a key computer vision task that identifies and locates cracks in road surface images, which usually have an irregular shape and contain only a few pixels in width. Generative and unsupervised methods are popular these years, but generative methods require a lot of training data and computational power while unsupervised methods are not so satisfactory in pixel-level segmentation. The process is challenged by the irregularity of crack shapes and complex road image backgrounds. To alleviate these problems, we propose a novel method in this paper, CDS-Net, that significantly improves road crack detection performance through multiple practical modules, including the Multi-Directional Hierarchical Attention (MDHA) module and the Difference Sensitivity Reconstruction Block (DSRB). Specifically, the MDHA module employs a multi-directional feature extraction strategy to capture detailed information of cracks, thereby enhancing the discriminative power of the features. The DSRB module, designed to address the inefficiency of traditional skip-connections, utilizes masked convolution and graph convolution attention to reconstruct and refine feature representations. Additionally, we propose an improved weighted cross-entropy loss function to address the inherent class imbalance problem in road crack detection. Extensive experiments on five public datasets demonstrate that CDS-Net achieves superior performance compared to other state-of-the-art methods, showcasing its effectiveness and robustness in road crack detection. It also has a stronger generalization ability compared with other methods. Code is available athttps://github.com/ttttqz/CDS-Net/tree/master. Qinzhong Tan, Weisheng Dong, Xin Li 0005, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Oriented Vehicle Joint Detection and Tracking in Satellite Video via Identifier-Free Point SupervisionabstractOriented vehicle detection and tracking play a crucial role in various real-world applications. Yet, existing advanced models heavily rely on abundant and accurate oriented bounding box and tracking identifier annotations, which are extremely labor-intensive for satellite videos. In this paper, we endeavor to employ identifier-free point annotation to achieve the competitive performance while minimizing annotation costs. Specifically, each instance across video frames are labeled by single points, without providing its instance identifier. Building upon this setting, we introduce an oriented vehicle joint detection and tracking framework for satellite video, focusing on enhancing model performance by carefully-designed sample acquisition and robust learning processes. Firstly, we leverage temporal and visual information to generate sequence-aligned pseudo-labels and visually-aligned synthetic objects, which complement each other during training by providing both exact appearance and annotation information. Secondly, a novel spatio-temporal consistency metric is developed to assess sample quality, which is then incorporated into a curriculum learning schedule. This strategy facilitates a gradual learning progression from high-quality data to low-quality or noisy examples, thereyby minimizing interference from potentially misleading samples. Finally, an end-to-end oriented object joint detection and tracking network is constructed to enable effective oriented vehicle dynamic analysis. Extensive ablation and experimental results on two satellite video datasets demonstrate the superiority of our proposed method. Yuping Liang, Jinjian Wu, Junpeng Zhang 0002, Yuxuan Chang, Jie Feng 0003, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Enhanced Clutter Suppression and GMTIm Algorithm With Modified DKP and NCS for Single-Channel Spaceborne- Maneuvering BFSARabstractSingle-channel spaceborne-maneuvering bistatic forward-looking synthetic aperture radar (SS-BFSAR) enables the maneuvering platform to achieve high-resolution forward-looking imaging without deploying additional antennas or transmitting the radar signal. Nevertheless, achieving the suppression of spatial variant clutter with only the single-channel configuration remains a critical challenge for the ground moving target imaging (GMTIm) mission of SS-BFSAR. This study proposes an enhanced clutter suppression and GMTIm algorithm for SS-BFSAR. First, the proposed modified deramp-keystone processing (DKP) completely decouples the echo signal in range and azimuth while avoiding the signal-to-clutter ratio (SCR) degradation caused by the azimuth spectrum aliasing. Subsequently, two pre-focusing approaches are developed with nonlinear chirp scaling (NCS), i.e., global NCS and block NCS, to achieve deep focusing of spatial variant clutter while introducing differences of Doppler frequency position (DFP) between the clutter and GMT. These approaches preserve the clutter consistency of the pre-focus results, thereby ensuring that the GMT signal exhibits high SCR following the image domain cancellation. Finally, a matched filter and the proposed monostatic-equivalent model are used to refocus the GMT and estimate its velocity. The proposed algorithm can simultaneously obtain the well-focused ground scene image and the GMTIm result without DFP drift caused by the target’s motion. Comparative experiments using simulation and real data demonstrate the effectiveness and superiority of the proposed algorithm. Xuan Song 0002, Yachao Li 0001, Yanhong Guo, Pei Ye, Xuanqi Wang, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Locally Aware Visual State Space for Small Defect Segmentation in Complex Component ImagesabstractSegmenting small defects within large imaging fields remains challenging in industrial scenarios due to the difficulty in distinguishing defects from complex component backgrounds and identifying defects comprising only a few pixels in high-resolution images. To address these issues, we propose a novel dual-branch feature extraction architecture, the locally aware visual state space block, which captures global contextual information while maintaining locally aware perception. In addition, we introduce the parallel quad-directional scanning fusion module to extract multiscale information, aggregating high-level features at different scales for enhanced global information fusion. To avoid losing small target details when upsampling the global segmentation mask to high-resolution input size, we develop progressive location refinement modules to incrementally refine small defect localization from the bottom up. Extensive experiments on our proposed small defect segmentation dataset and a public PCB dataset demonstrate that our method outperforms existing state-of-the-art methods in both performance and efficiency. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | CrossEI: Boosting Motion-Oriented Object Tracking With an Event CameraabstractWith the differential sensitivity and high time resolution, event cameras can record detailed motion clues, which form a complementary advantage with frame-based cameras to enhance the object tracking, especially in challenging dynamic scenes. However, how to better match heterogeneous event-image data and exploit rich complementary cues from them still remains an open issue. In this paper, we align event-image modalities by proposing a motion adaptive event sampling method, and we revisit the cross-complementarities of event-image data to design a bidirectional-enhanced fusion framework. Specifically, this sampling strategy can adapt to different dynamic scenes and integrate aligned event-image pairs. Besides, we design an image-guided motion estimation unit for extracting explicit instance-level motions, aiming at refining the uncertain event clues to distinguish primary objects and background. Then, a semantic modulation module is devised to utilize the enhanced object motion to modify the image features. Coupled with these two modules, this framework learns both the high motion sensitivity of events and the full texture of images to achieve more accurate and robust tracking. The proposed method is easily embedded in existing tracking pipelines, and trained end-to-end. We evaluate it on four large benchmarks, i.e. FE108, VisEvent, FE240hz and CoeSot. Extensive experiments demonstrate our method achieves state-of-the-art performance, and large improvements are pointed as contributions by our sampling strategy and fusion concept. Zhiwen Chen 0002, Jinjian Wu, Weisheng Dong, Leida Li, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2025 | Alternating Direction Unfolding With a Cross Spectral Attention Prior for Dual-Camera Compressive Hyperspectral ImagingabstractCoded Aperture Snapshot Spectral Imaging (CASSI) multiplexes 3D Hyperspectral Images (HSIs) into a 2D sensor to capture dynamic spectral scenes, which, however, sacrifices the spatial information. Dual-Camera Compressive Hyperspectral Imaging (DCCHI) enhances CASSI by incorporating a Panchromatic (PAN) camera to compensate for the loss of spatial information in CASSI. However, the dual-camera structure of DCCHI disrupts the diagonal property of the product of the sensing matrix and its transpose, making it difficult to efficiently and accurately solve the data subproblem in closed-form and thereby hindering the application of model-based methods and Deep Unfolding Networks (DUNs) that rely on such a closed-form solution. To address this issue, we propose an Alternating Direction DUN, named ADRNN, which decouples the imaging model of DCCHI into a CASSI subproblem and a PAN subproblem. The ADRNN alternately solves data terms analytically and a joint prior term in these subproblems. Additionally, we propose a Cross Spectral Transformer (XST) to exploit the joint prior. The XST utilizes cross spectral attention to exploit the correlation between the compressed HSI and the PAN image, and incorporates Grouped-Query Attention (GQA) to alleviate the burden of parameters and computational cost brought by impartially treating the compressed HSI and the PAN image. Furthermore, we built a real DCCHI system and captured large-scale indoor and outdoor scenes for future academic research. Extensive experiments on both simulation and real datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. The code and datasets have been open-sourced at: https://github.com/ShawnDong98/ADRNN-XST. Yubo Dong, Dahua Gao, Danhua Liu, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2025 | Incomplete Modalities Restoration via Hierarchical Adaptation for Robust Multimodal SegmentationabstractMultimodal semantic segmentation has significantly advanced the field of semantic segmentation by integrating data from multiple sources. However, this task often encounters missing modality scenarios due to challenges such as sensor failures or data transmission errors, which can result in substantial performance degradation. Existing approaches to addressing missing modalities predominantly involve training separate models tailored to specific missing scenarios, typically requiring considerable computational resources. In this paper, we propose a Hierarchical Adaptation framework to Restore Missing Modalities for Multimodal segmentation (HARM3), which enables frozen pretrained multimodal models to be directly applied to missing-modality semantic segmentation tasks with minimal parameter updates. Central to HARM3 is a text-instructed missing modality prompt module, which learns multimodal semantic knowledge by utilizing available modalities and textual instructions to generate prompts for the missing modalities. By incorporating a small set of trainable parameters, this module effectively facilitates knowledge transfer between high-resource domains and low-resource domains where missing modalities are more prevalent. Besides, to further enhance the model's robustness and adaptability, we introduce adaptive perturbation training and an affine modality adapter. Extensive experimental results demonstrate the effectiveness and robustness of HARM3 across a variety of missing modality scenarios. Weisheng Dong, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Image Process. | 7 |
| 2025 | Local Uncertainty Energy Transfer for Active Domain AdaptationabstractActive Domain Adaptation (ADA) improves knowledge transfer efficiency from the labeled source domain to the unlabeled target domain by selecting a few target sample labels. However, most existing active sampling methods ignore the local uncertainty of neighbors in the target domain,making it easier to pick out anomalous samples that are detrimental to the model. To address this problem, we present a new approach to active domain adaptation called Local Uncertainty Energy Transfer (LUET), which integrates active learning of local uncertainty confusion and energy transfer alignment constraints into a unified framework. First, in the active learning module, the uncertainty difficult and representative samples from the target domain are selected through local uncertainty energy selection and entropy-weighted class confusion selection. And the active learning strategy based on local uncertainty energy will avoid selecting anomalous samples in the target domain. Second, for the discrimination issue caused by domain shift, we use a global and local energy-transfer alignment constraint module to eliminate the domain gap and improve accuracy. Finally, we used negative log-likelihood loss for supervised learning of source domains and query samples. With the introduction of sample-based energy metrics, the active learning strategy is more closely with the domain alignment. Experiments on multiple domain-adaptive datasets have demonstrated that our LUET can achieve outstanding results and outperform existing state-of-the-art approaches. Guangming Shi, Weisheng Dong, Xin Li 0005, Xuemei Xie |
IEEE Trans. Image Process. | 2 |
| 2025 | Joint Spatial and Frequency Domain Learning for Lightweight Spectral Image DemosaicingabstractConventional spectral image demosaicing algorithms rely on pixels' spatial or spectral correlations for reconstruction. Due to the missing data in the multispectral filter array (MSFA), the estimation of spatial or spectral correlations is inaccurate, leading to poor reconstruction results, and these algorithms are time-consuming. Deep learning-based spectral image demosaicing methods directly learn the nonlinear mapping relationship between 2D spectral mosaic images and 3D multispectral images. However, these learning-based methods focused only on learning the mapping relationship in the spatial domain, but neglected valuable image information in the frequency domain, resulting in limited reconstruction quality. To address the above issues, this paper proposes a novel lightweight spectral image demosaicing method based on joint spatial and frequency domain information learning. First, a novel parameter-free spectral image initialization strategy based on the Fourier transform is proposed, which leads to better initialized spectral images and eases the difficulty of subsequent spectral image reconstruction. Furthermore, an efficient spatial-frequency transformer network is proposed, which jointly learns the spatial correlations and the frequency domain characteristics. Compared to existing learning-based spectral image demosaicing methods, the proposed method significantly reduces the number of model parameters and computational complexity. Extensive experiments on simulated and real-world data show that the proposed method notably outperforms existing spectral image demosaicing methods. Xun Cao, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 7 |
| 2025 | SANSee: A Physical-Layer Semantic-Aware Networking Framework for Distributed Wireless SensingabstractContactless device-free wireless sensing has recently attracted significant interest due to its potential to support a wide range of immersive human-machine interactive applications using ubiquitously available radio frequency (RF) signals. Traditional approaches focus on developing a single global model based on a combined dataset collected from different locations. However, wireless signals are known to be location and environment specific. Thus, a global model results in inconsistent and unreliable sensing results. It is also unrealistic to construct individual models for all the possible locations and environmental scenarios. Motivated by the observation that signals recorded at different locations are closely related to a set of physical-layer semantic features, in this paper we propose SANSee, a semantic-aware networking-based framework for distributed wireless sensing. SANSee allows models constructed in one or a limited number of locations to be transferred to new locations without requiring any locally labeled data or model training. SANSee is built on the concept of physical-layer semantic-aware network (pSAN), which characterizes the semantic similarity and the correlations of sensed data across different locations. A pSAN-based zero-shot transfer learning solution is introduced to allow receivers in new locations to obtain location-specific models by directly aggregating the models trained by other receivers. We theoretically prove that models obtained by SANSee can approach the locally optimal models. Experimental results based on real-world datasets are used to verify that the accuracy of the transferred models obtained by SANSee matches that of the models trained by the locally labeled data based on supervised learning approaches. Huixiang Zhu, Yong Xiao 0001, Yingyu Li, Guangming Shi, Marwan Krunz |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Learning Pyramid-Structured Long-Range Dependencies for 3D Human Pose EstimationabstractAction coordination in human structure is indispensable for the spatial constraints of 2D joints to recover 3D pose. Usually, action coordination is represented as a long-range dependence among body parts. However, there are two main challenges in modeling long-range dependencies. First, joints should not only be constrained by other individual joints but also be modulated by the body parts. Second, existing methods make networks deeper to learn dependencies between non-linked parts. They introduce uncorrelated noise and increase the model size. In this paper, we utilize a pyramid structure to better learn potential long-range dependencies. It can capture the correlation across joints and groups, which complements the context of the human sub-structure. In an effective cross-scale way, it captures the pyramid-structured long-range dependence. Specifically, we propose a novel Pyramid Graph Attention (PGA) module to capture long-range cross-scale dependencies. It concatenates information from various scales into a compact sequence, and then computes the correlation between scales in parallel. Combining PGA with graph convolution modules, we develop a Pyramid Graph Transformer (PGFormer) for 3D human pose estimation, which is a lightweight multi-scale transformer architecture. It encapsulates human sub-structures into self-attention by pooling. Extensive experiments show that our approach achieves lower error and smaller model size than state-of-the-art methods on Human3.6 M and MPI-INF-3DHP datasets. Xuemei Xie, Yutong Zhong, Guangming Shi |
IEEE Trans. Multim. | 4 |
| 2025 | S4DL: Shift-Sensitive Spatial-Spectral Disentangling Learning for Hyperspectral Image Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) techniques, extensively studied in hyperspectral image (HSI) classification, aim to use labeled source domain data and unlabeled target domain data to learn domain invariant features for cross-scene classification. Compared to natural images, numerous spectral bands of HSIs provide abundant semantic information, but they also increase the domain shift significantly. In most existing methods, both explicit alignment and implicit alignment simply align feature distribution, ignoring domain information in the spectrum. We noted that when the spectral channel between source and target domains is distinguished obviously, the transfer performance of these methods tends to deteriorate. Additionally, their performance fluctuates greatly owing to the varying domain shifts across various datasets. To address these problems, a novel shift-sensitive spatial-spectral disentangling learning (S4DL) approach is proposed. In S4DL, gradient-guided spatial-spectral decomposition (GSSD) is designed to separate domain-specific and domain-invariant representations by generating tailored masks under the guidance of the gradient from domain classification. A shift-sensitive adaptive monitor is defined to adjust the intensity of disentangling according to the magnitude of domain shift. Furthermore, a reversible neural network is constructed to retain domain information that lies not only in semantic but also the shallow-level detailed information. Extensive experimental results on several cross-scene HSI datasets consistently verified that S4DL is better than the state-of-the-art UDA methods. Our source code will be available athttps://github.com/xdu-jjgs/IEEE_TNNLS_S4DL. Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Weisheng Dong, Guangming Shi, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Reconfiguring Satellite CDNs With Dynamic Uncertain User Requests Based on Multi-Agent DRLabstractThe Satellite-Terrestrial Integrated Network (STIN) is a key paradigm to achieve global coverage and ubiquitous connection in the 6G era. Integrating the network slice technology based on Software Defined Networking (SDN) and Network Function Virtualization (NFV) into STIN is recognized as an effective solution to achieve a rapid flexible service provisioning. Specifically, the Content Delivery Network (CDN) service, which is storage resource intensive, is suitable to be deployed on STIN to provide a global range content service suppply. Most existing research on the resource deployment of STIN focuses on static requests, ignoring the dynamic changes in requests caused by the high-speed movement of satellites. Especially, the reconfiguration of CDN slices and the variation of STIN are asynchronous in different time scales: this makes the reconfiguration of CDN slices to cope with faster changing user requests a challenging task. In this paper, we adopt a periodic reconfiguration strategy to configure CDN slices on the edge LEO satellite network of STIN in discrete time intervals. Within each time interval, we formulate the reconfiguration optimization problem to cope with the dynamically changing user requests, wherein the Stochastic Network Calculus (SNC) is used to measure the deployment performance within this time interval. Then, we describe the optimization problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), and propose a reconfiguration algorithm based on Multi-agent Deep Reinforcement Learning (MADRL) to determine the optimal adjustment strategy. Finally, intensive simulations are implemented to verify the performance of the algorithm. Compared with the baselines, the QoS of the proposed algorithm is increased by about 5.8%, while the operation cost and reconfiguration cost are decreased by about 8.5% and 35.3% respectively. Jiayi Liu 0001, Xuemei Xie, Guangming Shi |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | Inverse Weight-Balancing for Deep Long-Tailed LearningabstractThe performance of deep learning models often degrades rapidly when faced with imbalanced data characterized by a long-tailed distribution. Researchers have found that the fully connected layer trained by cross-entropy loss has large weight-norms for classes with many samples, but not for classes with few samples. How to address the data imbalance problem with both the encoder and the classifier seems an under-researched problem. In this paper, we propose an inverse weight-balancing (IWB) approach to guide model training and alleviate the data imbalance problem in two stages. In the first stage, an encoder and classifier (the fully connected layer) are trained using conventional cross-entropy loss. In the second stage, with a fixed encoder, the classifier is finetuned through an adaptive distribution for IWB in the decision space. Unlike existing inverse image frequency that implements a multiplicative margin adjustment transformation in the classification layer, our approach can be interpreted as an adaptive distribution alignment strategy using not only the class-wise number distribution but also the sample-wise difficulty distribution in both encoder and classifier. Experiments show that our method can greatly improve performance on imbalanced datasets such as CIFAR100-LT with different imbalance factors, ImageNet-LT, and iNaturelists2018. Wenqi Dang, Weisheng Dong, Xin Li 0005, Guangming Shi |
AAAI | 5 |
| 2024 | BVT-IMA: Binary Vision Transformer with Information-Modified AttentionabstractAs a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized models suffer from serious performance drops. In this paper, an attention shifting is observed in the binary multi-head self-attention module, which can influence the information fusion between tokens and thus hurts the model performance. From the perspective of information theory, we find a correlation between attention scores and the information quantity, further indicating that a reason for such a phenomenon may be the loss of the information quantity induced by constant moduli of binarized tokens. Finally, we reveal the information quantity hidden in the attention maps of binary vision transformers and propose a simple approach to modify the attention values with look-up information tables so that improve the model performance. Extensive experiments on CIFAR-100/TinyImageNet/ImageNet-1k demonstrate the effectiveness of the proposed information-modified attention on binary vision transformers. Zhenyu Wang 0008, Hao Luo 0004, Xuemei Xie, Fan Wang 0019, Guangming Shi |
AAAI | 5 |
| 2024 | Motion Deblurring via Spatial-Temporal Collaboration of Frames and EventsabstractMotion deblurring can be advanced by exploiting informative features from supplementary sensors such as event cameras, which can capture rich motion information asynchronously with high temporal resolution. Existing event-based motion deblurring methods neither consider the modality redundancy in spatial fusion nor temporal cooperation between events and frames. To tackle these limitations, a novel spatial-temporal collaboration network (STCNet) is proposed for event-based motion deblurring. Firstly, we propose a differential-modality based cross-modal calibration strategy to suppress redundancy for complementarity enhancement, and then bimodal spatial fusion is achieved with an elaborate cross-modal co-attention mechanism to weight the contributions of them for importance balance. Besides, we present a frame-event mutual spatio-temporal attention scheme to alleviate the errors of relying only on frames to compute cross-temporal similarities when the motion blur is significant, and then the spatio-temporal features from both frames and events are aggregated with the custom cross-temporal coordinate attention. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/STCNet. Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Guangming Shi |
AAAI | 5 |
| 2024 | Segment Any Event Streams via Weighted Adaptation of Pivotal TokensabstractIn this paper, we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data, with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is the precise alignment and calibration of embeddings derived from event-centric data such that they harmoniously coincide with those originating from RGB imagery. Capitalizing on the vast repositories of datasets with paired events and RGB images, our proposition is to harness and extrapolate the profound knowledge encapsulated within the pretrained SAM framework. As a cornerstone to achieving this, we introduce a multi-scale feature distillation methodology. This methodology rigorously optimizes the alignment of token embeddings originating from event data with their RGB image counterparts, thereby preserving and enhancing the robustness of the overall architecture. Considering the distinct significance that token embeddings from intermediate layers hold for higher-level embeddings, our strategy is centered on accurately calibrating the pivotal token embeddings. This targeted calibration is aimed at effectively managing the discrepancies in high-level embeddings originating from both the event and image domains. Extensive experiments on different datasets demonstrate the effectiveness of the proposed distillation method. Code in https://github.com/happychenpipi/EventSAM. Zhiwen Chen 0002, Yifan Zhang 0036, Junhui Hou, Guangming Shi, Jinjian Wu |
CVPR | 5 |
| 2024 | Clustered Federated Learning for Distributed Wireless SensingabstractRF-based wireless sensing is a promising technology for enabling applications such as human activity recognition, intelligent healthcare, and robotics. However, the growing concern in data privacy and also the complexity in modeling and keeping track of the statistical heterogeneity of RF signal hinders its wide application, especially in large-scale wireless networking systems. In this paper, we introduce a hierarchical clustering-based federated learning framework, called Uniform Manifold Clustering Federated Learning (UMCFL), that has the potential to address the above challenges. UMCFL first divides all the RF signal receivers into different clusters according to the similarity of their data distributions and then construct an individual model for receivers within each cluster. To measure the data distribution similarity between receivers in a computationally efficient way, UMCFL first adopts a uniform manifold approximation and projection (UMAP)-based solution to convert data samples at each receiver into a low-dimensional representation and then use maximum mean discrepancy (MMD) to calculate the similarity score between datasets of different receivers. We prove the convergence of UMCFL and perform extensive experiments to evaluate its performance. Experimental results show that UMCFL achieves up to 63% improvement in sensing accuracy, compared to the traditional FedAvg-based wireless sensing solution. Zijian Sun, Yong Xiao 0001, Haohui Cai, Huixiang Zhu, Yingyu Li, Guangming Shi |
GLOBECOM | 7 |
| 2024 | Mosic: Multimodal Semantic Integrated Communication for Health Monitoring in Iot ScenariosabstractMonitoring multimodal signals provides a more comprehensive understanding of health conditions compared to singlemode monitoring. In the face of the significant volumes of multimodal signals, existing IoT health monitoring systems primarily focus on high-fidelity signal transmission by encoding multimodal signals separately. However, due to the lack of consideration for the downstream applications and correlation between multimodal signals, a portion of bandwidth resources is wasted on task-irrelevant information and intermodal redundancy. To address this issue, we propose the Multimodal Semantic Integration Communication (MoSIC) framework composed of three levels: At the sensor level, multiple wearable sensors collect and send different modal signals to a mobile terminal; at the mobile terminal level, the terminal employs deep source-channel joint encoding for the received multimodal signals, extracting single-modal embedded features using a backbone network, and obtaining cross-modal features through a feature fusion network with contrastive constrain; at the cloud level, a decoding network symmetric to the encoding network reconstructs the multimodal signals, which are then used for downstream applications such as human activity recognition. MoSIC focuses on semantically integrating multimodal signals for downstream applications, resulting in improved encoding and transmission efficiency. It also reduces the radio-frequency power consumption and bandwidth requirements. Minxi Yang, Dahua Gao, Xiaodan Song, Guangming Shi |
ICASSP | 6 |
| 2024 | SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image TransmissionabstractIn recent years, semantic communication based on deep learning for source-channel joint encoding has garnered significant attention. It utilizes network models trained end-to-end to represent signals as embedding vectors and has demonstrated superior performance compared to traditional methods. However, due to the significant disparity between embedding vectors and human language, it can be challenging to succinctly capture abstract semantics such as scenes. In this paper, we introduce the Scene Graph-based Generative Semantic Communication (SG2SC) framework, built upon structured semantics and conditional generative models for image transmission. SG2SC aims to faithfully convey abstract semantics like scenes. It begins by detecting object categories, spatial attributes, and inter-category relationships in the image, representing scene semantics in a graph structure. Subsequently, it employs graph neural networks for scene graph encoding, decoding, and transmission, and finally utilizes a conditional diffusion model for semantic decoding. Benefiting from its concise graph structure semantics, SG2SC outperforms traditional method, semantic communication based on deep joint source-channel coding, and segmentation-based generative semantic communication in terms of noise resistance and encoding efficiency. Minxi Yang, Dahua Gao, Feng Xie 0009, Xiaodan Song, Guangming Shi |
ICASSP | 6 |
| 2024 | An O(m+n)-Space Spatiotemporal Denoising Filter with Cache-Like Memories for Dynamic Vision SensorsabstractDynamic vision sensor (DVS) is novel neuromorphic imaging device that generates asynchronous events. Despite the high temporal resolution and high dynamic range features, DVS is faced with background noise problem. Spatiotemporal filter is an effective and hardware-friendly solution for DVS denoising but previous designs have large memory overhead or degraded performance issues. In this paper, we present a lightweight and real-time spatiotemporal denoising filter with set-associative cache-like memories, which has low space complexity of O(m+n) for DVS of m×n resolution. A two-stage pipeline for memory access with read cancellation feature is proposed to reduce power consumption. Further the bitwidth redundancy for event storage is exploited to minimize the memory footprint. We implemented our design on FPGA and experimental results show that it achieves state-of-the-art performance compared with previous spatiotemporal filters while maintaining low resource utilization and low power consumption of about 125mW to 210mW at 100MHz clock frequency. Qinghang Zhao, Yixi Ji, Jinjian Wu, Guangming Shi |
ICCAD | 5 |
| 2024 | Swin-UMamba: Mamba-Based UNet with ImageNet-Based Pretraining
Jiarun Liu, Hao Yang 0026, Yan Xi, Lequan Yu, Cheng Li 0008, Yong Liang 0001, Guangming Shi, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (9) | 8 |
| 2024 | AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics PerceptionabstractThe highly abstract nature of image aesthetics perception (IAP) poses a significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma, resulting in MLLMs falling short of aesthetics perception capabilities. To address the above challenge, we first introduce a comprehensively annotated Aesthetic Multi-Modality Instruction Tuning (AesMMIT) dataset, which serves as the footstone for building multi-modality aesthetics foundation models. Specifically, to align MLLMs with human aesthetics perception, we construct a corpus-rich aesthetic critique database with 21,904 diverse-sourced images and 88K human natural language feedbacks, which are collected via progressive questions, ranging from coarse-grained aesthetic grades to fine-grained aesthetic descriptions. To ensure that MLLMs can handle diverse queries, we further prompt GPT to refine the aesthetic critiques and assemble the large-scale aesthetic instruction tuning dataset, i.e. AesMMIT, which consists of 409K multi-typed instructions to activate stronger aesthetic capabilities. Based on the AesMMIT database, we fine-tune the open-sourced general foundation models, achieving multi-modality Aesthetic Expert models, dubbed AesExpert. Extensive experiments demonstrate that the proposed AesExpert models deliver significantly better aesthetic perception performances than the state-of-the-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision. Project Page: https://yipoh.github.io/aes-expert/. Yipo Huang, Xiangfei Sheng, Zhichao Yang 0013, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin, Guangming Shi |
ACM Multimedia | 9 |
| 2024 | E-Motion: Future Motion Simulation via Event Sequence DiffusionabstractForecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict future motion with a level of detail and precision previously unachievable. Inspired by that, we propose to integrate the strong learning capacity of the video diffusion model with the rich motion information of an event camera as a motion simulation framework. Specifically, we initially employ pre-trained stable video diffusion models to adapt the event sequence dataset. This process facilitates the transfer of extensive knowledge from RGB videos to an event-centric domain. Moreover, we introduce an alignment mechanism that utilizes reinforcement learning techniques to enhance the reverse generation trajectory of the diffusion model, ensuring improved performance and accuracy. Through extensive testing and validation, we demonstrate the effectiveness of our method in various complex scenarios, showcasing its potential to revolutionize motion flow prediction in computer vision applications such as autonomous vehicle guidance, robotic navigation, and interactive media. Our findings suggest a promising direction for future research in enhancing the interpretative power and predictive accuracy of computer vision systems. The source code is
publicly available at https://github.com/p4r4mount/E-Motion. Junhui Hou, Guangming Shi, Jinjian Wu |
NeurIPS | 4 |
| 2024 | STDM-transformer: Space-time dual multi-scale transformer network for skeleton-based action recognition
Zhifu Zhao, Jianan Li 0003, Xuemei Xie, Xiaotian Wang 0001, Guangming Shi |
Neurocomputing | 7 |
| 2024 | TSUDepth: Exploring temporal symmetry-based uncertainty for unsupervised monocular depth estimation
Yufan Zhu, Weisheng Dong, Xin Li 0005, Guangming Shi |
Neurocomputing | 5 |
| 2024 | Self-supervised discriminative model prediction for visual tracking
Di Yuan 0002, Gu Geng, Xiu Shu, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001, Guangming Shi |
Neural Comput. Appl. | 7 |
| 2024 | A novel hybrid decoding neural network for EEG signal representation
Youshuo Ji, Fu Li 0002, Boxun Fu, Yijin Zhou, Yang Li 0019, Xiaoli Li 0002, Guangming Shi |
Pattern Recognit. | 8 |
| 2024 | Learning real-world heterogeneous noise models with a benchmark dataset
Jie Lin 0008, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
Pattern Recognit. | 6 |
| 2024 | Motion-Oriented Hybrid Spiking Neural Networks for Event-Based Motion DeblurringabstractImage deblurring based only on the blurry image is challenging as motion information is lost while imaging. Event cameras capture the texture of moving objects in high temporal resolution with asynchronous events. In this paper, we extract motion features from events and fuse them with background features from the image for event-based image deblurring. Spiking neural network (SNN), a widely recognized event feature extractor, is well suited for motion feature extraction due to its high temporal resolution. However, extracting motion information from events exclusively with SNN is challenging. We propose a novel Temporal-local-Spatio Spiking Transformer (TSST) to extract motion intensity and motion attention regions in the spatio-temporal domain. Motion intensity extracted from spiking features is represented as a high temporal resolution motion attention map to guide the fusion of the two networks. In the temporal domain, motion intensity maps spiking features to CNN features as motion features to avoid blurring. In the spatial domain, the motion intensity shows the motion regions and gives the weight of the motion feature during fusion. Moreover, a hybrid feature extraction encoder (HFEE) is introduced, which fully fuses the motion and background features for deblurring. The gradient is back-propagated from CNN to SNN, and the hybrid deblurring network is jointly optimized. We evaluated the performance of our model on the public dataset GoPro and a real event dataset we captured. Codes and pretrained models are available athttps://github.com/XDULzx/MotionSNN. Zhaoxin Liu, Jinjian Wu, Guangming Shi, Wen Yang 0008, Weisheng Dong, Qinghang Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Human Pose Estimation via Parse Graph of Body StructureabstractWhen observing a person’s body, humans can extract the structured representation of the body called a parse graph, which includes the hierarchical decompositions from the entire body to parts and primitives and the context relations by horizontal links between the body parts. This ability helps humans better locate body structures at different levels. In order for the model to have this ability for single-person pose estimation, we design a hierarchical network to model the context relations and hierarchical structure in the parse graph of body structure by convolutional neural networks. It overcomes the problem that most methods ignore one of the context relations and hierarchical structure in the parse graph. Our network contains bottom-up and top-down stages. In the bottom-up stage, the structural features of the hierarchy are captured from primitives to parts and the entire body. Then in the top-down stage, with the context information of each body part, the structural features of the body parts are refined separately rather than together from the entire body to parts and primitives. Experiments show that our model enhances the reasonableness of predictions and achieves superior results on the CrowdPose, COCO keypoint detection and MPII human pose datasets. Shibang Liu, Xuemei Xie, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Uncertainty Modeling of the Transmission Map for Single Image DehazingabstractDespite rapid progress of end-to-end optimization for single-image dehazing, a long-standing open problem is the non-homogenous haze, at the core of the differences between synthetic hazy images and real hazy images. The atmospheric scattering model (ASM) has been widely adopted to model the degradation process of haze images but based on the assumption of homogeneous haze. In realistic scenarios, non-homogeneous haze often makes it more difficult to estimate the transmission map in ASM, resulting in undesired artifacts in the restored images. To address the issue of non-homogeneous haze, we propose to model the uncertainty in the estimation of the transmission map and develop a spatially adaptive learning module for ASM correction. Specifically, we present an approach to enhancing the well-known Dark Channel prior (DCP) by relaxing the constraint with the transmission map in the DCP-net. Assuming the availability of paired training data, we have developed a strategy to address vulnerability in the DCP, leading to a more accurate estimation of the transmission map. Then, we explore the uncertainty between the estimated transmission map and target transmission map (Ground Truth) to reformulate the ASM for the presence of non-homogeneous haze. A robust and accurate estimated transmission map can boost the final dehazing performance of our DCP-net. Experiments on three popular synthetic and real non-homogeneous datasets show that our proposed approach has achieved better results on both synthetic scenes and real non-homogeneous scenes. The code is available athttps://see.xidian.edu.cn/faculty/wsdong/Projects/Projects/project_dehazing_TCSVT2024.htm Bokang Wang, Qian Ning, Xin Li 0005, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Glimpse and Zoom: Spatio-Temporal Focused Dynamic Network for Skeleton-Based Action RecognitionabstractGCN-based methods have achieved remarkable performance in skeleton-based action recognition. However, existing methods have not explicitly attempted to remove temporal and spatial redundancy that might introduce additional computational costs. Inspired by the fact that humans always tend to glimpse at overall motion and then zoom into the most important spatio-temporal regions, we propose a Spatio Temporal Focused Dynamic Network (STFD-Net) trained with reinforcement learning for skeleton-based action recognition. Specifically, we first propose a global extractor with Skeleton Pooling Module (SPM) to enable the network to focus on overall motion information with a refined skeleton structure. Then, a local extractor, containing pair-wise part partition, tubelet proposal network, and Partition-Grouped Module (PGM), is proposed to extract local motion details as a complement to the overall motion information. Finally, the dynamic classifier utilizes a recurrent neural network to dynamically terminate the process once the network is adequately confident. Extensive experiments have demonstrated that the proposed network achieves SOTA level performance with lower computational cost on the NTU 60 and NTU 120 dataset. Zhifu Zhao, Jianan Li 0003, Xiaotian Wang 0001, Xuemei Xie, Wanxin Zhang, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Degradation Estimation Recurrent Neural Network With Local and Non-Local Priors for Compressive Spectral ImagingabstractIn the coded aperture snapshot spectral imaging (CASSI) system, deep unfolding networks (DUNs) have demonstrated excellent performance in recovering 3-D hyperspectral images (HSIs) from 2-D measurements. However, some noticeable gaps exist between the imaging model used in DUNs and the real CASSI imaging process, such as the sensing error as well as photon and dark current noise, compromising the accuracy of solving the data subproblem and the prior subproblem in DUNs. To address this issue, we propose a degradation estimation network (DEN) to correct the imaging model used in DUNs by simultaneously estimating the sensing error and the noise level, thereby improving the performance of DUNs. Additionally, we propose an efficient local and non-local transformer (LNLT) to solve the prior subproblem, which not only effectively models local and non-local similarities but also reduces the computational cost of the window-based global multihead self-attention (MSA). Furthermore, we transform the DUN into a recurrent neural network (RNN) by sharing parameters of DNNs across stages, which not only allows DNN to be trained more adequately but also significantly reduces the number of parameters. The proposed DERNN-LNLT achieves state-of-the-art (SOTA) performance with fewer parameters on both simulation and real datasets (code:https://github.com/ShawnDong98/DERNN-LNLT). Yubo Dong, Dahua Gao, Guangming Shi, Danhua Liu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Fast Universal Azimuth Signal Modeling for Maneuvering-Platform BFSAR ImagingabstractAppropriate modeling and processing of echoes are the foundations for high-precision frequency-domain synthetic aperture radar (SAR) imaging. The complex geometry makes it challenging to accurately characterize and cope with the 2-D spatial variation of Doppler modulation in maneuvering-platform translational-variant bistatic forward-looking SAR (MTV-BFSAR), resulting in that the azimuth processing method by means of setting reference points on the Cartesian coordinate axis significantly impairs the performance in terms of robustness, accuracy, and efficiency. This article proposes a comprehensive MTV-BFSAR imaging algorithm based on universal frequency-domain azimuth signal modeling (UFDASM). The presented methodology utilizes the bistatic bisector to form the azimuth reference line (ARL) and develops two expeditious ARL and range isoline (RIL) coordinates’ solving techniques, which substantially augments the reliability and efficiency of the space-variant Doppler modulation coefficient (DMC) representation, reduces the order, and improves the robustness of the entire algorithm. Subsequently, thanks to UFDASM, a modified nonlinear chirp scaling (NLCS) method with orthogonal impulse response function (IRF) is derived to eliminate the spatial variation of the quadratic DMC in the range-Doppler domain while omitting that of the cubic one. Furthermore, one may find that the equalization of the second-order spatial variation introduces a cubic phase error (CPE) term. However, boundary analyses manifest that this error is small enough not to affect the imaging performance in MTV-BFSAR. Finally, the accuracy, robustness, and efficiency of the approach are validated through numerical simulation and raw data processing. Xuanqi Wang, Yachao Li 0001, Xuan Song 0002, Baixiao Chen, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | TransVQA: Transferable Vector Quantization Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the unlabeled target domain. Most existing domain adaptation methods are based on convolutional neural networks (CNNs) to learn cross-domain invariant features. Inspired by the success of transformer architectures and their superiority to CNNs, we propose to combine the transformer with UDA to improve their generalization properties. In this paper, we present a novel model named Trans ferable V ector Q uantization A lignment for Unsupervised Domain Adaptation (TransVQA), which integrates the Transferable transformer-based feature extractor (Trans), vector quantization domain alignment (VQA), and mutual information weighted maximization confusion matrix (MIMC) of intra-class discrimination into a unified domain adaptation framework. First, TransVQA uses the transformer to extract more accurate features in different domains for classification. Second, TransVQA, based on the vector quantization alignment module, uses a two-step alignment method to align the extracted cross-domain features and solve the domain shift problem. The two-step alignment includes global alignment via vector quantization and intra-class local alignment via pseudo-labels. Third, for intra-class feature discrimination problem caused by the fuzzy alignment of different domains, we use the MIMC module to constrain the target domain output and increase the accuracy of pseudo-labels. The experiments on several datasets of domain adaptation show that TransVQA can achieve excellent performance and outperform existing state-of-the-art methods. Weisheng Dong, Xin Li 0005, Guangming Shi, Xuemei Xie |
IEEE Trans. Image Process. | 5 |
| 2024 | Multi-Scale Spatio-Temporal Memory Network for Lightweight Video DenoisingabstractDeep learning-based video denoising methods have achieved great performance improvements in recent years. However, the expensive computational cost arising from sophisticated network design has severely limited their applications in real-world scenarios. To address this practical weakness, we propose a multiscale spatio-temporal memory network for fast video denoising, named MSTMN, aiming at striking an improved trade-off between cost and performance. To develop an efficient and effective algorithm for video denoising, we exploit a multiscale representation based on the Gaussian-Laplacian pyramid decomposition so that the reference frame can be restored in a coarse-to-fine manner. Guided by a model-based optimization approach, we design an effective variance estimation module, an alignment error estimation module and an adaptive fusion module for each scale of the pyramid representation. For the fusion module, we employ a reconstruction recurrence strategy to incorporate local temporal information. Moreover, we propose a memory enhancement module to exploit the global spatio-temporal information. Meanwhile, the similarity computation of the spatio-temporal memory network enables the proposed network to adaptively search the valuable information at the patch level, which avoids computationally expensive motion estimation and compensation operations. Experimental results on real-world raw video datasets have demonstrated that the proposed lightweight network outperforms current state-of-the-art fast video denoising algorithms such as FastDVDnet, EMVD, and ReMoNet with fewer computational costs. Xin Li 0005, Jie Lin 0008, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 7 |
| 2024 | Learning Frame-Event Fusion for Motion DeblurringabstractMotion deblurring is a highly ill-posed problem due to the significant loss of motion information in the blurring process. Complementary informative features from auxiliary sensors such as event cameras can be explored for guiding motion deblurring. The event camera can capture rich motion information asynchronously with microsecond accuracy. In this paper, a novel frame-event fusion framework is proposed for event-driven motion deblurring (FEF-Deblur), which can sufficiently explore long-range cross-modal information interactions. Firstly, different modalities are usually complementary and also redundant. Cross-modal fusion is modeled as complementary-unique features separation-and-aggregation, avoiding the modality redundancy. Unique features and complementary features are first inferred with parallel intra-modal self-attention and inter-modal cross-attention respectively. After that, a correlation-based constraint is designed to act between unique and complementary features to facilitate their differentiation, which assists in cross-modal redundancy suppression. Additionally, spatio-temporal dependencies among neighboring inputs are crucial for motion deblurring. A recurrent cross attention is introduced to preserve inter-input attention information, in which the current spatial features and aggregated temporal features are attending to each other by establishing the long-range interaction between them. Extensive experiments on both synthetic and real-world motion deblurring datasets demonstrate our method outperforms state-of-the-art event-based and image/video-based methods. The code will be made publicly available. Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 6 |
| 2024 | Superimposed Semantic Communication for IoT-Based Real-Time ECG MonitoringabstractReal-time electrocardiogram (ECG) monitoring and diagnosis through Internet of Things (IoT) are crucial for addressing the severity and timely treatment of cardiovascular diseases, enabling timely intervention and preventing life-threatening complications. However, current ECG monitoring research predominantly focuses on individual aspects such as signal compression, diagnostic analysis, or secure transmission, lacking joint optimization of various modules in IoT scenarios. To address this gap, this work proposes a novel framework based on superimposed semantic communication for real-time ECG monitoring in IoT. The framework comprises three hierarchical levels: the edge level for data collection and processing, the relay level for signal compression and coding, and the cloud level for data analysis and reconstruction. The proposed framework offers several unique advantages. By employing semantic encoding guided by ECG classification tasks, it selectively extracts crucial features within and between signals, improving compression ratio and adaptability to channel noise. The superimposed semantic encoding achieves content encryption without requiring any additional operations. Moreover, the framework utilizes lightweight anomaly detection neural networks, reducing edge device power consumption and conserving communication resources. Simulation and real experimental results demonstrate that the proposed method achieves real-time encoding and transmission of ECG signals with a compression ratio of 0.019 on the MIT-BIH dataset. Furthermore, it attains a heartbeat classification accuracy of 0.988 and a reconstruction error of 0.061. Minxi Yang, Dahua Gao, Guangming Shi |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Distributed Traffic Synthesis and Classification in Edge Networks: A Federated Self-Supervised Learning ApproachabstractWith the rising demand for wireless services and increased awareness of the need for data protection, existing network traffic analysis and management architectures are facing unprecedented challenges in classifying and synthesizing the increasingly diverse services and applications. This paper proposes FS-GAN, a federated self-supervised learning framework to support automatic traffic analysis and synthesis over a large number of heterogeneous datasets. FS-GAN is composed of multiple distributed Generative Adversarial Networks (GANs), with a set of generators, each being designed to generate synthesized data samples following the distribution of an individual service traffic, and each discriminator being trained to differentiate the synthesized data samples and the real data samples of a local dataset. A federated learning-based framework is adopted to coordinate local model training processes of different GANs across different datasets. FS-GAN can classify data of unknown types of service and create synthetic samples that capture the traffic distribution of the unknown types. We prove that FS-GAN can minimize the Jensen-Shannon Divergence (JSD) between the distribution of real data across all the datasets and that of the synthesized data samples. FS-GAN also maximizes the JSD among the distributions of data samples created by different generators, resulting in each generator producing synthetic data samples that follow the same distribution as one particular service type. Extensive simulation results show that the classification accuracy of FS-GAN achieves over$20\%$improvement in average compared to the state-of-the-art clustering-based traffic analysis algorithms. FS-GAN also has the capability to synthesize highly complex mixtures of traffic types without requiring any human-labeled data samples. Yong Xiao 0001, Rong Xia, Yingyu Li, Guangming Shi, Diep N. Nguyen, Dinh Thai Hoang, Dusit Niyato, Marwan Krunz |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Time-Sensitive Learning for Heterogeneous Federated Edge IntelligenceabstractReal-time machine learning (ML) has recently attracted significant interest due to its potential to support instantaneous learning, adaptation, and decision making in a wide range of application domains, including self-driving vehicles, intelligent transportation, and industry automation. In this paper, we investigate real-time ML in a federated edge intelligence (FEI) system, an edge computing system that implements federated learning (FL) solutions based on data samples collected and uploaded from decentralized data networks, e.g., Internet-of-Things (IoT) and/or wireless sensor networks. FEI systems often exhibit heterogenous communication and computational resource distribution, as well as non-i.i.d. data samples arrived at different edge servers, resulting in long model training time and inefficient resource utilization. Motivated by this fact, we propose a time-sensitive federated learning (TS-FL) framework to minimize the overall run-time for collaboratively training a shared ML model with desirable accuracy. Training acceleration solutions for both TS-FL with synchronous coordination (TS-FL-SC) and asynchronous coordination (TS-FL-ASC) are investigated. To address the straggler effect in TS-FL-SC, we develop an analytical solution to characterize the impact of selecting different subsets of edge servers on the overall model training time. A server dropping-based solution is proposed to allow some slow-performance edge servers to be removed from participating in the model training if their impact on the resulting model accuracy is limited. A joint optimization algorithm is proposed to minimize the overall time consumption of model training by selecting participating edge servers, the local epoch number (the number of model training iterations per coordination), and the data batch size (the number of data samples for each model training iteration). Motivated by the fact that data samples at the slowest edge server may exhibit special characteristics that cannot be removed from model training, we develop an analytical expression to characterize the impact of both staleness effect of asynchronous coordination and straggler effect of FL on the time consumption of TS-FL-ASC. We propose a load forwarding-based solution that allows a slow edge server to offload part of its training samples to trusted edge servers with higher processing capability. We develop a hardware prototype to evaluate the model training time of a heterogeneous FEI system. Experimental results show that our proposed TS-FL-SC and TS-FL-ASC can provide up to 63% and 28% of reduction, in the overall model training time, respectively, compared with traditional FL solutions. Yong Xiao 0001, Yingyu Li, Guangming Shi, Marwan Krunz, Diep N. Nguyen, Dinh Thai Hoang |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Coarse-to-Fine Image Aesthetics Assessment With Dynamic Attribute SelectionabstractImage aesthetics assessment (IAA) is an interesting but challenging task, owing to the ineffable nature of human sense of beauty. The study of IAA has evolved from simple binary classification to more complex score regression and distribution prediction. It is effortless for people to perform aesthetic binary classification,i.e., aesthetically pleasing or not. However, further judgment on the fine-level scalar aesthetic score is complex and typically determined by aesthetic attributes presented in the image, such as content, lighting and color. Motivated by the above facts, this paper presents a Coarse-to-fine image Aesthetics assessment model guided by Dynamic Attribute Selection, dubbed CADAS. The underlying idea is to simulate the process of human aesthetic perception by performing coarse-to-fine aesthetic reasoning. Specifically, a hierarchical AttributeNet is first pre-trained by imitating the staged mechanism of human aesthetic experience, producing the candidate aesthetic attributes. Then, an AestheticNet is introduced to perform the coarse-level binary classification, based on which a confidence-based attribute selection strategy is designed to dynamically pick out the dominant aesthetic attributes from the candidate ones. Finally, a self-attention-based FusionNet is designed to explore the interaction between dominant aesthetic attributes and aesthetic features, producing the fine-level aesthetic prediction. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-arts. Furthermore, CADAS is also able to output the dominant aesthetic attributes in images, facilitating model explainability. Yipo Huang, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Guangming Shi |
IEEE Trans. Multim. | 7 |
| 2024 | Blind Image Quality Assessment Based on Perceptual ComparisonabstractBlind image quality assessment (BIQA) is a regression task with continuous label space, the feature space of which is expected to have a corresponding continuity in the target space. However, existing approaches typically learn quality score regression directly in an end-to-end fashion, which leaves networks susceptible to interference from task-agnostic information, and fails to capture the continuity of BIQA. In this work, by explicitly establishing inter-sample associations, a simple yet effective BIQA framework based on perceptual comparison is proposed to capture the continuity. To this end, besides the basic quality score regression, the relative quality scores between images are predicted to exploit the relative quality relationships between samples for optimizing the representation of image perceptual quality. In addition, based on the human perceptual characteristic, we derive a novel sample weighting strategy to dynamically adjust the weights for different samples in the network learning process for further improving the robustness of the model. The performances on both single-database and cross-database experiments achieve state-of-the-art, indicating the effectiveness of the proposed method. Besides, the proposed framework is model-agnostic, which can effectively improve the performance of the benchmark model with no extra inference cost. Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Multim. | 6 |
| 2024 | Active Learning for Deep Visual TrackingabstractConvolutional neural networks (CNNs) have been successfully applied to the single target tracking task in recent years. Generally, training a deep CNN model requires numerous labeled training samples, and the number and quality of these samples directly affect the representational capability of the trained model. However, this approach is restrictive in practice, because manually labeling such a large number of training samples is time-consuming and prohibitively expensive. In this article, we propose an active learning method for deep visual tracking, which selects and annotates the unlabeled samples to train the deep CNN model. Under the guidance of active learning, the tracker based on the trained deep CNN model can achieve competitive tracking performance while reducing the labeling cost. More specifically, to ensure the diversity of selected samples, we propose an active learning method based on multiframe collaboration to select those training samples that should be and need to be annotated. Meanwhile, considering the representativeness of these selected samples, we adopt a nearest-neighbor discrimination method based on the average nearest-neighbor distance to screen isolated samples and low-quality samples. Therefore, the training samples' subset selected based on our method requires only a given budget to maintain the diversity and representativeness of the entire sample set. Furthermore, we adopt a Tversky loss to improve the bounding box estimation of our tracker, which can ensure that the tracker achieves more accurate target states. Extensive experimental results confirm that our active-learning-based tracker (ALT) achieves competitive tracking accuracy and speed compared with state-of-the-art trackers on the seven most challenging evaluation benchmarks. Project website: https://sites.google.com/view/altrack/. Di Yuan 0002, Xiaojun Chang, Qiao Liu 0001, Yi Yang 0001, Minglei Shu, Zhenyu He 0001, Guangming Shi |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Mobility-Aware MEC Planning With a GNN-Based Graph Partitioning FrameworkabstractMobile service continuity is essential important to ensure that user sessions and services will survive user mobility. The 5G enhances its mobility management by providing the flexibility and offering three types of Session and Service Continuity (SSC) modes to address various service continuity requirements. Multi-access edge computing (MEC) is a type of widely adopted network architecture that delivers network services from the boundary of the mobile network by provisioning a set of edge servers. Determining an optimum planning of MEC edge servers, which involves determining edge servers appropriate geographical positions and their serving areas, is a precondition for more efficient service provisioning and better usage of network resources. In this work, we investigate the MEC servers planning problem by considering the management cost for maintaining MEC service continuity. The problem is formulated as a graph partitioning problem to partition the RAN graph with minimum SSC management costs and balanced MEC servers workloads. Then, we adapt a generalizable approximate Graph Partitioning framework which leverages on Graph Neural Network (GNN) to embed the RAN network spacial feature and on Multilayer Perceptron (MLP) for graph partitioning. Based on the framework, we propose a MEC server planning algorithm named MECP-GAP. Finally, we evaluate MECP-GAP with extensive simulations and real network data. Comparing to several baselines, MECP-GAP achieves better performance with lower running time. Jiayi Liu 0001, Zhongyi Xu, Xuefang Liu, Xuemei Xie, Guangming Shi |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2024 | Graph Transformer and LSTM Attention for VNF Multi-Step Workload Prediction in SFCabstractThe knowledge on a Service Function Chain’s (SFC’s) resource requirements is an indispensable prerequisite for proactive resource provisioning and run-time management of the SFC. However, due to the intrinsic dynamics in network environment, accurate resource requirements and workloads prediction for the Virtual Network Functions (VNFs) of a SFC, especially in a large time scale, is a non-trivial challenge. In the literature, existing works largely neglect the application-level relationship of VNFs in improving prediction accuracy, and few work investigates the multi-step prediction. In this work, we propose a deep-learning-based multi-step prediction model for accurate workload prediction for SFC VNFs in a dynamic network environment. We first demonstrate that predictability can be improved by taking into account application-level dependency by calculating the spatial conditional entropy of adjacent VNFs workloads. Then, the prediction model, named Graph Transformer Networks and sequence-to-sequence LSTM with Attention (GTN-LA) is introduced, which utilizes the Graph Transformer as the encoder to capture the application-level dependencies among VNFs, and the LSTM with attention as the decoder to extract the temporal dependencies within the time varying load information. Finally, GTN-LA is validated through intensive evaluation with a real SFC workload dataset by comparing towards several baselines. Jiayi Liu 0001, Xuemei Xie, Guangming Shi |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | Semantic Feature Division Multiple Access for Multi-User Digital Interference NetworksabstractWith the ever-increasing user density and quality of service (QoS) demand, 5G networks with limited spectrum resources are facing massive access challenges. To address these challenges, in this paper, we propose a novel discrete semantic feature division multiple access (SFDMA) paradigm for multi-user digital interference networks. Specifically, by utilizing deep learning technology, SFDMA extracts multi-user semantic information into discrete representations in distinguishable semantic subspaces, which enables multiple users to transmit simultaneously over the same time-frequency resources. Furthermore, based on a robust information bottleneck, we design a SFDMA based multi-user digital semantic interference network for inference tasks, which can achieve approximate orthogonal transmission. Moreover, we propose a SFDMA based multi-user digital semantic interference network for image reconstruction tasks, where the discrete outputs of the semantic encoders of the users are approximately orthogonal, which significantly reduces multi-user interference. Furthermore, we propose an Alpha-Beta-Gamma (ABG) formula for semantic communications, which is the first theoretical relationship between inference accuracy and transmission power. Then, we derive adaptive power control methods with closed-form expressions for inference tasks. Extensive simulations verify the effectiveness and superiority of the proposed SFDMA. Shuai Ma 0002, Chuanhui Zhang, Youlong Wu, Hang Li 0003, Shiyin Li, Guangming Shi, Naofal Al-Dhahir |
IEEE Trans. Wirel. Commun. | 7 |
| 2024 | Features Disentangled Semantic Broadcast Communication NetworksabstractSingle-user semantic communications have attracted extensive research recently, but multi-user semantic broadcast communication (BC) is still in its infancy. In this paper, we propose a practical robust features-disentangled multi-user semantic BC framework, where the transmitter includes a feature selection module and each user has a feature completion module. Instead of broadcasting all extracted features, the semantic encoder extracts the disentangled semantic features, and then only the users’ intended semantic features are selected for broadcasting, which can further improve the transmission efficiency. Within this framework, we further investigate two information-theoretic metrics, including the ultimate compression rate under both the distortion and perception constraints, and the achievable rate region of the semantic BC. Furthermore, to realize the proposed semantic BC framework, we design a lightweight robust semantic BC network by exploiting a supervised autoencoder (AE), which can controllably disentangle sematic features. Moreover, we design the first hardware proof-of-concept prototype of the semantic BC network, where the proposed semantic BC network can be implemented in real time. Simulations and experiments demonstrate that the proposed robust semantic BC network can significantly improve transmission efficiency. Shuai Ma 0002, Zhi Zhang 0003, Youlong Wu, Hang Li 0003, Guangming Shi, Dahua Gao, Yuanming Shi, Shiyin Li, Naofal Al-Dhahir |
IEEE Trans. Wirel. Commun. | 5 |
| 2024 | Reasoning Over the Air: A Reasoning- Based Implicit Semantic-Aware Communication FrameworkabstractSemantic-aware communication is a novel paradigm that draws inspiration from human communication focusing on the delivery of the meaning of messages. It has attracted significant interest recently due to its potential to improve the efficiency and reliability of communication and enhance users’ quality-of-experience (QoE). Most existing works focus on transmitting and delivering the explicit semantic meaning that can be directly identified from the source signal. This paper investigates the implicit semantic-aware communication in which the hidden information, e.g., hidden relations, concepts and implicit reasoning mechanisms of users, that cannot be directly observed from the source signal must be recognized and interpreted by the intended users. To this end, a novel implicit semantic-aware communication (iSAC) architecture is proposed for representing, communicating, and interpreting the implicit semantic meaning between source and destination users. A graph-inspired structure is first developed to represent the complete semantics, including both explicit and implicit, of a message. A projection-based semantic encoder is then proposed to convert the high-dimensional graphical representation of explicit semantics into a low-dimensional semantic constellation space for efficient physical channel transmission. To enable the destination user to learn and imitate the implicit semantic reasoning process of source user, a generative adversarial imitation learning-based solution, called G-RML, is proposed. Different from existing communication solutions, the source user in G-RML does not focus only on sending as much of the useful messages as possible; but, instead, it tries to guide the destination user to learn a reasoning mechanism to map any observed explicit semantics to the corresponding implicit semantics that are most relevant to the semantic meaning. By applying G-RML, we prove that the destination user can accurately imitate the reasoning process of the source user and automatically generate a set of implicit reasoning paths following the same probability distribution as the expert paths. Compared to the existing solutions, our proposed G-RML requires much less communication and computational resources and scales well to the scenarios involving the communication of rich semantic meanings consisting of a large number of concepts and relations. Numerical results show that the proposed solution achieves up to 92% accuracy of implicit meaning interpretation. Yong Xiao 0001, Yiwei Liao, Yingyu Li, Guangming Shi, H. Vincent Poor, Walid Saad 0001, Mérouane Debbah, Mehdi Bennis |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Gradient Corner Pooling for Keypoint-Based Object DetectionabstractDetecting objects as multiple keypoints is an important approach in the anchor-free object detection methods while corner pooling is an effective feature encoding method for corner positioning. The corners of the bounding box are located by summing the feature maps which are max-pooled in the x and y directions respectively by corner pooling. In the unidirectional max pooling operation, the features of the densely arranged objects of the same class are prone to occlusion. To this end, we propose a method named Gradient Corner Pooling. The spatial distance information of objects on the feature map is encoded during the unidirectional pooling process, which effectively alleviates the occlusion of the homogeneous object features. Further, the computational complexity of gradient corner pooling is the same as traditional corner pooling and hence it can be implemented efficiently. Gradient corner pooling obtains consistent improvements for various keypoint-based methods by directly replacing corner pooling. We verify the gradient corner pooling algorithm on the dataset and in real scenarios, respectively. The networks with gradient corner pooling located the corner points earlier in the training process and achieve an average accuracy improvement of 0.2%-1.6% on the MS-COCO dataset. The detectors with gradient corner pooling show better angle adaptability for arrayed objects in the actual scene test. Xuemei Xie, Mingxuan Yu, Jiakai Luo, Chengwei Rao, Guangming Shi |
AAAI | 6 |
| 2023 | Improving Robotic Tactile Localization Super-resolution via Spatiotemporal Continuity Learning and Overlapping Air ChambersabstractHuman hand has amazing super-resolution ability in sensing the force and position of contact and this ability can be strengthened by practice. Inspired by this, we propose a method for robotic tactile super-resolution enhancement by learning spatiotemporal continuity of contact position and a tactile sensor composed of overlapping air chambers. Each overlapping air chamber is constructed of soft material and seals the barometer inside to mimic adapting receptors of human skin. Each barometer obtains the global receptive field of the contact surface with the pressure propagation in the hyperelastic seal overlapping air chambers. Neural networks with causal convolution are employed to resolve the pressure data sampled by barometers and to predict the contact position. The temporal consistency of spatial position contributes to the accuracy and stability of positioning. We obtain an average super-resolution (SR) factor of over 2500 with only four physical sensing nodes on the rubber surface (0.1 mm in the best case on 38 × 26 mm²), which outperforms the state-of-the-art. The effect of time series length on the location prediction accuracy of causal convolution is quantitatively analyzed in this article. We show that robots can accomplish challenging tasks such as haptic trajectory following, adaptive grasping, and human-robot interaction with the tactile sensor. This research provides new insight into tactile super-resolution sensing and could be beneficial to various applications in the robotics field. Xuemei Xie, Guangming Shi |
AAAI | 5 |
| 2023 | Residual Degradation Learning Unfolding Framework with Mixing Priors Across Spectral and Spatial for Compressive Spectral ImagingabstractTo acquire a snapshot spectral image, coded aperture snapshot spectral imaging (CASSI) is proposed. A core problem of the CASSI system is to recover the reliable and fine underlying 3D spectral cube from the 2D measurement. By alternately solving a data subproblem and a prior subproblem, deep unfolding methods achieve good performance. However, in the data subproblem, the used sensing matrix is ill-suited for the real degradation process due to the device errors caused by phase aberration, distortion; in the prior subproblem, it is important to design a suitable model to jointly exploit both spatial and spectral priors. In this paper, we propose a Residual Degradation Learning Unfolding Framework (RDLUF), which bridges the gap between the sensing matrix and the degradation process. Moreover, a MixS2Transformer is designed via mixing priors across spectral and spatial to strengthen the spectral-spatial representation capability. Finally, plugging the MixS2Transformer into the RDLUF leads to an end-to-end trainable neural network RDLUF-MixS2. Experimental results establish the superior performance of the proposed method over existing ones. Code is available: https://github.com/ShawnDong98/RDLUF_MixS2 Yubo Dong, Dahua Gao, Minxi Yang, Guangming Shi |
CVPR | 6 |
| 2023 | Self-supervised Non-uniform Kernel Estimation with Flow-based Motion Prior for Blind Image DeblurringabstractMany deep learning-based solutions to blind image deblurring estimate the blur representation and reconstruct the target image from its blurry observation. However, these methods suffer from severe performance degradation in real-world scenarios because they ignore important prior information about motion blur (e.g., real-world motion blur is diverse and spatially varying). Some methods have attempted to explicitly estimate non-uniform blur kernels by CNNs, but accurate estimation is still challenging due to the lack of ground truth about spatially varying blur kernels in real-world images. To address these issues, we propose to represent the field of motion blur kernels in a latent space by normalizing flows, and design CNNs to predict the latent codes instead of motion kernels. To further improve the accuracy and robustness of non-uniform kernel estimation, we introduce uncertainty learning into the process of estimating latent codes and propose a multi-scale kernel attention module to better integrate image features with estimated kernels. Extensive experimental results, especially on real-world blur datasets, demonstrate that our method achieves state-of-the-art results in terms of both subjective and objective quality as well as excellent generalization performance for non-uniform image deblurring. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UFPNet.htm. Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
CVPR | 6 |
| 2023 | Vector Quantization with Self-Attention for Quality-Independent Representation LearningabstractRecently, the robustness of deep neural networks has drawn extensive attention due to the potential distribution shift between training and testing data (e.g., deep models trained on high-quality images are sensitive to corruption during testing). Many researchers attempt to make the model learn invariant representations from multiple corrupted data through data augmentation or image-pair-based feature distillation to improve the robustness. Inspired by sparse representation in image restoration, we opt to address this issue by learning image-quality-independent feature representation in a simple plug-and-play manner, that is, to introduce discrete vector quantization (VQ) to remove redundancy in recognition models. Specifically, we first add a codebook module to the network to quantize deep features. Then we concatenate them and design a self-attention module to enhance the representation. During training, we enforce the quantization of features from clean and corrupted images in the same discrete embedding space so that an invariant quality-independentfeature representation can be learned to improve the recognition robustness of low-quality images. Qualitative and quantitative experimental results show that our method achieved this goal effectively, leading to a new state-of-the-art result of 43.1 % mCE on ImageNet-C with ResNet50 as the backbone. On other robustness benchmark datasets, such as ImageNet-R, our method also has an accuracy improvement of almost 2%. The source code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/VQSA.htm Weisheng Dong, Xin Li 0005, Mengluan Huang, Guangming Shi |
CVPR | 6 |
| 2023 | Adversarial Learning for Implicit Semantic-Aware CommunicationsabstractSemantic communication is a novel communication paradigm that focuses on recognizing and delivering the desired meaning of messages to the destination users. Most existing works in this area focus on delivering explicit semantics, labels or signal features that can be directly identified from the source signals. In this paper, we consider the implicit semantic communication problem in which hidden relations and closely related semantic terms that cannot be recognized from the source signals need to also be delivered to the destination user. We develop a novel adversarial learning-based implicit semantic-aware communication (iSAC) architecture in which the source user, instead of maximizing the total amount of information transmitted to the channel, aims to help the recipient learn an inference rule that can automatically generate implicit semantics based on limited clue information. We prove that by applying iSAC, the destination user can always learn an inference rule that matches the true inference rule of the source messages. Experimental results show that the proposed iSAC can offer up to a 19.69 dB improvement over existing non-inferential communication solutions, in terms of symbol error rate at the destination user. Zhimin Lu, Yong Xiao 0001, Zijian Sun, Yingyu Li, Guangming Shi, Xianfu Chen, Mehdi Bennis, H. Vincent Poor |
ICC | 5 |
| 2023 | Low-Light Image Enhancement with Multi-stage Residue Quantization and Brightness-aware AttentionabstractLow-light image enhancement (LLIE) aims to recover illumination and improve the visibility of low-light images. Conventional LLIE methods often produce poor results because they neglect the effect of noise interference. Deep learning-based LLIE methods focus on learning a mapping function between low-light images and normal-light images that outperforms conventional LLIE methods. However, most deep learning-based LLIE methods cannot yet fully exploit the guidance of auxiliary priors provided by normal-light images in the training dataset. In this paper, we propose a brightness-aware network with normal-light priors based on brightness-aware attention and residual-quantized codebook. To achieve a more natural and realistic enhancement, we design a query module to obtain more reliable normal-light features and fuse them with low-light features by a fusion branch. In addition, we propose a brightness-aware attention module to further improve the robustness of the network to the brightness. Extensive experimental results on both real-captured and synthetic data show that our method outperforms existing state-of-the-art methods. Weisheng Dong, Xin Li 0005, Guangming Shi |
ICCV | 6 |
| 2023 | Rate-Distortion-Perception Theory for Semantic CommunicationabstractSemantic communication has attracted significant interest recently due to its capability to meet the fast growing demand on user-defined and human-oriented communication services such as holographic communications, eXtended reality (XR), and human-to-machine interactions. Unfortunately, recent study suggests that the traditional Shannon information theory, focusing mainly on delivering semantic-agnostic symbols, will not be sufficient to investigate the semantic-level perceptual quality of the recovered messages at the receiver. In this paper, we study the achievable data rate of semantic communication under the symbol distortion and semantic perception constraints. Motivated by the fact that the semantic information generally involves rich intrinsic knowledge that cannot always be directly observed by the encoder, we consider a semantic information source that can only be indirectly sensed by the encoder. Both encoder and decoder can access to various types of side information that may be closely related to the user's communication preference. We derive the achievable region that characterizes the tradeoff among the data rate, symbol distortion, and semantic perception, which is then theoretically proved to be achievable by a stochastic coding scheme. We derive a closed-form achievable rate for binary semantic information source under any given distortion and perception constraints. We observe that there exists cases that the receiver can directly infer the semantic information source satisfying certain distortion and perception constraints without requiring any data communication from the transmitter. Experimental results based on the image semantic source signal have been presented to verify our theoretical observations. Jingxuan Chai, Yong Xiao 0001, Guangming Shi, Walid Saad 0001 |
ICNP | 3 |
| 2023 | Physical-Layer Semantic-Aware Network for Zero-Shot Wireless SensingabstractDevice-free wireless sensing has recently attracted significant interest due to its potential to support a wide range of immersive human-machine interactive applications. However, data heterogeneity in wireless signals and data privacy regulation of distributed sensing have been considered as the major challenges that hinder the wide applications of wireless sensing in large area networking systems. Motivated by the observation that signals recorded by wireless receivers are closely related to a set of physical-layer semantic features, in this paper we propose a novel zero-shot wireless sensing solution that allows models constructed in one or a limited number of locations to be directly transferred to other locations without any labeled data. We develop a novel physical-layer semantic-aware network (pSAN) framework to characterize the correlation between physical-layer semantic features and the sensing data distributions across different receivers. We then propose a pSAN-based zero-shot learning solution in which each receiver can obtain a location-specific gesture recognition model by directly aggregating the already constructed models of other receivers. We theoretically prove that models obtained by our proposed solution can approach the optimal model without requiring any local model training. Experimental results once again verify that the accuracy of models derived by our proposed solution matches that of the models trained by the real labeled data based on supervised learning approach. Huixiang Zhu, Yong Xiao 0001, Yingyu Li, Guangming Shi, Walid Saad 0001 |
ICNP | 4 |
| 2023 | Learning Primitive-Aware Discriminative Representations for Few-Shot Learning
Jianpeng Yang, Yuhang Niu, Xuemei Xie, Guangming Shi |
ICONIP (2) | 4 |
| 2023 | Exploring Correlations in Degraded Spatial Identity Features for Blind Face RestorationabstractBlind face restoration aims to recover high-quality face images from low-quality ones with complex and unknown degradation. Existing approaches have achieved promising performance by leveraging pre-trained dictionaries or generative priors. However, these methods may fail to exploit the full potential of degraded inputs and facial identity features due to complex degradation. To address this issue, we propose a novel method that explores the correlation of degraded spatial identity features by learning a general representation using memory network. Specifically, our approach enhances degraded features with more identity by leveraging similar facial features retrieved from memory network. We also propose a fusion approach that fuses memorized spatial features with GAN prior features via affine transformation and blending fusion to improve fidelity and realism. Additionally, the memory network is updated online in an unsupervised manner along with other modules, which obviates the requirement for pre-training. Experimental results on synthetic and popular real-world datasets demonstrate the effectiveness of our proposed method, which achieves at least comparable and often better performance than other state-of-the-art approaches. Qian Ning, Weisheng Dong, Xin Li 0005, Guangming Shi |
ACM Multimedia | 5 |
| 2023 | AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentabstractImage aesthetics assessment (IAA) aims at predicting the aesthetic quality of images. Recently, large pre-trained vision-language models, like CLIP, have shown impressive performances on various visual tasks. When it comes to IAA, a straightforward way is to finetune the CLIP image encoder using aesthetic images. However, this can only achieve limited success without considering the uniqueness of multimodal data in the aesthetics domain. People usually assess image aesthetics according to fine-grained visual attributes, e.g., color, light and composition. However, how to learn aesthetics-aware attributes from CLIP-based semantic space has not been addressed before. With this motivation, this paper presents a CLIP-based multi-attribute contrastive learning framework for IAA, dubbed AesCLIP. Specifically, AesCLIP consists of two major components, i.e., aesthetic attribute-based comment classification and attribute-aware learning. The former classifies the aesthetic comments into different attribute categories. Then the latter learns an aesthetic attribute-aware representation by contrastive learning, aiming to mitigate the domain shift from the general visual domain to the aesthetics domain. Extensive experiments have been done by using the pre-trained AesCLIP on four popular IAA databases, and the results demonstrate the advantage of AesCLIP over the state-of-the-arts. The source code will be public at https://github.com/OPPOMKLab/AesCLIP. Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong, Yuzhe Yang 0001, Liwu Xu, Guangming Shi |
ACM Multimedia | 9 |
| 2023 | OCSKB: An Object Component Sketch Knowledge Base for Fast 6D Pose Estimationabstract6D pose estimation from a single RGB image is a fundamental task in computer vision. In most methods of instance-level or category-level 6D pose estimation, accurate CAD models or point cloud models are indispensable part. It is not easy to quickly obtain the models of these everyday objects. To address this issue, we present a part-level object component sketch knowledge base which consists of 270 real-world object sketch models of 30 categories. Objects are disassembled into geometry components with spatial relationship according to their functions and structures, and convert them into three basic spatial structures: frustum, circular truncated cone, and sphere. We present a fast pipeline for sketch modeling with our tool. The average time for this method to build a simple model for everyday objects is about 2 minutes. Additionally, we leverage the geometric information and spatial relationships inherent in the multiple viewpoint projection maps of these sketch bases to develop a rapid inference framework for 6D pose estimation. The interpretable steps in our framework gradually retrieve and activate valid solutions in the discrete 6D pose space. Extensive experiments in real-world environments have demonstrated that our method can reliably and robustly estimate the 6D pose of objects, even without access to accurate CAD or point cloud models. Furthermore, our method achieves state-of-the-art performance, operating at a speed of 90 frames per second using parallel computing on GPU. Guangming Shi, Xuemei Xie, Mingxuan Yu, Chengwei Rao, Jiakai Luo |
ACM Multimedia | 1 |
| 2023 | Event-based Motion Deblurring with Modality-Aware Decomposition and RecompositionabstractEvent camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events. Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 5 |
| 2023 | Memory Based Temporal Fusion Network for Video Deblurring
Chaohua Wang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
Int. J. Comput. Vis. | 6 |
| 2023 | Anchor-based knowledge embedding for image aesthetics assessment
Leida Li, Tianwu Zhi, Guangming Shi, Yuzhe Yang 0001, Liwu Xu, Yandong Guo |
Neurocomputing | 3 |
| 2023 | Progressive graph convolution network for EEG emotion recognition
Yijin Zhou, Fu Li 0002, Yang Li 0019, Youshuo Ji, Guangming Shi, Wenming Zheng, Lijian Zhang, Yuanfang Chen, Rui Cheng 0010 |
Neurocomputing | 5 |
| 2023 | Imitation Learning-Based Implicit Semantic-Aware Communication Networks: Multi-Layer Representation and Collaborative ReasoningabstractSemantic communication has recently attracted significant interest from both industry and academia due to its potential to transform the existing data-focused communication architecture towards a more generally intelligent and goal-oriented semantic-aware networking system. Despite its promising potential, semantic communications and semantic-aware networking are still in their infancy. Most existing works focus on transporting and delivering the explicit semantic information, e.g., labels or features of objects, that can be directly identified from the source signal. The original definition of semantics as well as recent results in cognitive neuroscience suggest that it is the implicit semantic information, in particular the hidden relations connecting different concepts and feature items that play the fundamental role in recognizing, communicating, and delivering the real semantic meanings of messages. Motivated by this observation, we propose a novel reasoning-based implicit semantic-aware communication network architecture that allows destination users to directly learn a reasoning mechanism that can automatically generate complex implicit semantic information based on a limited clue information sent by the source users. Our proposed architecture can be implemented in a multi-tier cloud/edge computing networks in which multiple tiers of cloud data center (CDC) and edge servers can collaborate and support efficient semantic encoding, decoding, and implicit semantic interpretation for multiple end-users. We introduce a new multi-layer representation of semantic information taking into consideration both the hierarchical structure of implicit semantics as well as the personalized inference preference of individual users. We model the semantic reasoning process as a reinforcement learning process and then propose an imitation-based semantic reasoning mechanism learning (iRML) solution to learning a reasoning policy that imitates the inference behavior of the source user. A federated graph convolutional network (GCN)-based collaborative reasoning solution is proposed to allow multiple edge servers to jointly construct a shared semantic interpretation model based on decentralized semantic message samples. Extensive experiments have been conducted based on real-world datasets to evaluate the performance of our proposed architecture. Numerical results confirm that iRML offers up to 25.8 dB improvement on the semantic symbol error rate, compared to the semantic-irrelevant communication solutions. Yong Xiao 0001, Zijian Sun, Guangming Shi, Dusit Niyato |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | MADPL-net: Multi-layer attention dictionary pair learning network for image classification
Guangming Shi, Weisheng Dong, Xuemei Xie |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Few-shot human-object interaction video recognition with transformers
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
Neural Networks | 4 |
| 2023 | Deep Gaussian Scale Mixture Prior for Image ReconstructionabstractImage reconstruction from partial observations has attracted increasing attention. Conventional image reconstruction methods with hand-crafted priors often fail to recover fine image details due to the poor representation capability of the hand-crafted priors. Deep learning methods attack this problem by directly learning mapping functions between the observations and the targeted images can achieve much better results. However, most powerful deep networks lack transparency and are nontrivial to design heuristically. This paper proposes a novel image reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Unlike existing unfolding methods that only estimate the image means (i.e., the denoising prior) but neglected the variances, we propose characterizing images by the GSM models with learned means and variances through a deep network. Furthermore, to learn the long-range dependencies of images, we develop an enhanced variant based on the Swin Transformer for learning GSM models. All parameters of the MAP estimator and the deep network are jointly optimized through end-to-end training. Extensive simulation and real data experimental results on spectral compressive imaging and image super-resolution demonstrate that the proposed method outperforms existing state-of-the-art methods. Xin Yuan 0002, Weisheng Dong, Jinjian Wu, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Adaptive Search-and-Training for Robust and Efficient Network PruningabstractBoth network pruning and neural architecture search (NAS) can be interpreted as techniques to automate the design and optimization of artificial neural networks. In this paper, we challenge the conventional wisdom of training before pruning by proposing a joint search-and-training approach to learn a compact network directly from scratch. Using pruning as a search strategy, we advocate three new insights for network engineering: 1) to formulate adaptive search as a cold start strategy to find a compact subnetwork on the coarse scale; and 2) to automatically learn the threshold for network pruning; 3) to offer flexibility to choose between efficiency and robustness. More specifically, we propose an adaptive search algorithm in the cold start by exploiting the randomness and flexibility of filter pruning. The weights associated with the network filters will be updated by ThreshNet, a flexible coarse-to-fine pruning method inspired by reinforcement learning. In addition, we introduce a robust pruning strategy leveraging the technique of knowledge distillation through a teacher-student network. Extensive experiments on ResNet and VGGNet have shown that our proposed method can achieve a better balance in terms of efficiency and accuracy and notable advantages over current state-of-the-art pruning methods in several popular datasets, including CIFAR10, CIFAR100, and ImageNet. The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/AST-NP.htm. Xiaotong Lu, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | GMSS: Graph-Based Multi-Task Self-Supervised Learning for EEG Emotion RecognitionabstractPrevious electroencephalogram (EEG) emotion recognition relies on single-task learning, which may lead to overfitting and learned emotion features lacking generalization. In this paper, a graph-based multi-task self-supervised learning model (GMSS) for EEG emotion recognition is proposed. GMSS has the ability to learn more general representations by integrating multiple self-supervised tasks, including spatial and frequency jigsaw puzzle tasks, and contrastive learning tasks. By learning from multiple tasks simultaneously, GMSS can find a representation that captures all of the tasks thereby decreasing the chance of overfitting on the original task, i.e., emotion recognition task. In particular, the spatial jigsaw puzzle task aims to capture the intrinsic spatial relationships of different brain regions. Considering the importance of frequency information in EEG emotional signals, the goal of the frequency jigsaw puzzle task is to explore the crucial frequency bands for EEG emotion recognition. To further regularize the learned features and encourage the network to learn inherent representations, contrastive learning task is adopted in this work by mapping the transformed data into a common feature space. The performance of the proposed GMSS is compared with several popular unsupervised and supervised methods. Experiments on SEED, SEED-IV, and MPED datasets show that the proposed model has remarkable advantages in learning more discriminative and general features for EEG emotional signals. Yang Li 0019, Fu Li 0002, Boxun Fu, Youshuo Ji, Yijin Zhou, Guangming Shi, Wenming Zheng |
IEEE Trans. Affect. Comput. | 9 |
| 2023 | Uncertainty-Driven Knowledge Distillation for Language Model CompressionabstractDespite the remarkable performance on various Natural Language Processing (NLP) tasks, the parametric complexity of pretrained language models has remained a major obstacle due to limited computational resources in many practical applications. Techniques such as knowledge distillation, network pruning, and quantization have been developed for language model compression. However, it has remained challenging to achieve an optimal tradeoff between model size and inference accuracy. To address this issue, we propose a novel and efficient uncertainty-driven knowledge distillation compression method for transformer-based pretrained language models. Specifically, we design a method of parameter retention and feedforward network parameter distillation to compress N-stacked transformer modules into one module in the fine-tuning stage. A key innovation of our approach is to add the uncertainty estimation module (UEM) into the student network such that it can guide the student network's feature reconstruction in the latent space (similar to the teacher's). Across multiple datasets in the natural language inference tasks of GLUE, we have achieved more than 95% accuracy of the original BERT, while only using about 50% of the parameters. Weisheng Dong, Xin Li 0005, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Dynamic Expert-Knowledge Ensemble for Generalizable Video Quality AssessmentabstractDespite the impressive progress of supervised methods in quality assessment for in- the-wild videos, models trained on one domain often fail to generalize well to others due to the domain shifts caused by distortion diversity and content variation. Domain generalizable video quality assessment (VQA) methods that can work across domains remain an open research challenge. Although combining more data following the mixed-domain training strategy can improve the generalization performance to a certain extent, the specific knowledge from each source domain, which could potentially be useful for improving unseen domain generalization, is ignored in this principle. Motivated by this, we propose a domain generalizable VQA method named Dynamic Ensemble of Expert-Knowledge (DEEK), a novel framework that dynamically exploits the expert-knowledge from each source domain to achieve a generalizable ensemble prediction. Specifically, based on the multiple experts each trained to specialize in a particular source domain, we aim to exploit complementary information provided by the expert-knowledge. We effectively train an ensemble model by proposing a quality-sensitive InfoNCE loss to regularize the collaborative training of all experts in the contrastive learning formulation, aiming to exploit complementary information provided by the expert-knowledge when forming the ensemble. By dynamically integrating the experts according to their relevances to the target data, these expert-knowledge could be leveraged for better generalization. Experiments on five VQA datasets verify that our approach outperforms the state-of-the-arts by large margins. Pengfei Chen 0003, Leida Li, Haoliang Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | ECSNet: Spatio-Temporal Feature Learning for Event CameraabstractThe neuromorphic event cameras can efficiently sense the latent geometric structures and motion clues of a scene by generating asynchronous and sparse event signals. Due to the irregular layout of the event signals, how to leverage their plentiful spatio-temporal information for recognition tasks remains a significant challenge. Existing methods tend to treat events as dense image-like or point-serie representations. However, they either suffer from severe destruction on the sparsity of event data or fail to encode robust spatial cues. To fully exploit their inherent sparsity with reconciling the spatio-temporal information, we introduce a compact event representation, namely 2D-1T event cloud sequence (2D-1T ECS). We couple this representation with a novel light-weight spatio-temporal learning framework (ECSNet) that accommodates both object classification and action recognition tasks. The core of our framework is a hierarchical spatial relation module. Equipped with specially designed surface-event-based sampling unit and local event normalization unit to enhance the inter-event relation encoding, this module learns robust geometric features from the 2D event clouds. And we propose a motion attention module for efficiently capturing long-term temporal context evolving with the 1T cloud sequence. Empirically, the experiments show that our framework achieves par or even better state-of-the-art performance. Importantly, our approach cooperates well with the sparsity of event data without any sophisticated operations, hence leading to low computational costs and prominent inference speeds. Zhiwen Chen 0002, Jinjian Wu, Junhui Hou, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Theme-Aware Visual Attribute Reasoning for Image Aesthetics AssessmentabstractPeople usually assess image aesthetics according to visual attributes, e.g., interesting content, good lighting and vivid color, etc. Further, the perception of visual attributes depends on the image theme. Therefore, the inherent relationship between visual attributes and image theme is crucial for image aesthetics assessment (IAA), which has not been comprehensively investigated. With this motivation, this paper presents a new IAA model based on Theme-Aware Visual Attribute Reasoning (TAVAR). The underlying idea is to simulate the process of human perception in image aesthetics by performing bilevel reasoning. Specifically, a visual attribute analysis network and a theme understanding network are first pre-trained to extract aesthetic attribute features and theme features, respectively. Then, the first level Attribute-Theme Graph (ATG) is built to investigate the coupling relationship between visual attributes and image theme. Further, a flexible aesthetics network is introduced to extract general aesthetic features, based on which we built the second level Attribute-Aesthetics Graph (AAG) to mine the relationship between theme-aware visual attributes and aesthetic features, producing the final aesthetic prediction. Extensive experiments on four public IAA databases demonstrate the superiority of the proposed TAVAR model over the state-of-the-arts. Furthermore, TAVAR features better explainability due to the use of visual attributes. Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Quality Assessment of UGC Videos Based on Decomposition and RecompositionabstractThe prevalence of short-video applications imposes more requirements for video quality assessment (VQA). User-generated content (UGC) videos are captured under an unprofessional environment, thus suffering from various dynamic degradations, such as camera shaking. To cover the dynamic degradations, existing recurrent neural network-based UGC-VQA methods can only provide implicit modeling, which is unclear and difficult to analyze. In this work, we consider explicit motion representation for dynamic degradations, and propose a motion-enhanced UGC-VQA method based on decomposition and recomposition. In the decomposition stage, a dual-stream decomposition module is built, and VQA task is decomposed into single frame-based quality assessment problem and cross frames-based motion understanding. The dual streams are well grounded on the two-pathway visual system during perception, and require no extra UGC data due to knowledge transfer. Hierarchical features from shallow to deep layers are gathered to narrow the gaps from tasks and domains. In the recomposition stage, a progressively residual aggregation module is built to recompose features from the dual streams. Representations with different layers and pathways are interacted and aggregated in a progressive and residual manner, which keeps a good trade-off between representation deficiency and redundancy. Extensive experiments on UGC-VQA databases verify that our method achieves the state-of-the-art performance and keeps a good capability of generalization. The source code will be available inhttps://github.com/Sissuire/DSD-PRO. Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | View-Normalized and Subject-Independent Skeleton Generation for Action RecognitionabstractSkeleton-based action recognition has attracted great interest in computer vision. For this task, a challenging problem concerns the large intraclass variances of skeleton data, which are mainly caused by diverse viewpoints and subjects, and greatly increase the difficulty of modeling actions through a network. To address the above problem, we propose a variance reduction (VaRe) framework for skeleton-based action recognition, which consists of a view-normalization generative adversarial network (VN-GAN), a subject-independent network (SINet) and a classification network. First, the VN-GAN is responsible for reducing view-induced intraclass variances. Specifically, this network, comprising a generator and a discriminator, is aimed at learning a mapping from a diverse-view skeleton distribution to a unified-view skeleton distribution in an unsupervised manner, thereby generating a view-normalized skeleton. Second, taking the view-normalized skeleton as input, the SINet focuses on reducing the influences of the personal habits of subjects on action recognition. To generate SI skeleton data, the SINet automatically adjusts the human pose according to the human kinematic structure under a classification loss constraint. Finally, without the interference of view- and subject-induced variances, the classification network can concentrate more on learning discriminative action features to predict classes. Furthermore, by combining the joint and bone modalities, the proposed framework achieves competitive performance on three benchmarks: NTU RGB+D, NTU-120 RGB+D and Northwestern-UCLA Multiview Action 3D. Qingzhe Pan, Zhifu Zhao, Xuemei Xie, Jianan Li 0003, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Deep Unfolding Network for Efficient Mixed Video Noise RemovalabstractExisting image and video denoising algorithms have focused on removing homogeneous Gaussian noise. However, this assumption with noise modeling is often too simplistic for the characteristics of real-world noise. Moreover, the design of network architectures in most deep learning-based video denoising methods is heuristic, ignoring valuable domain knowledge. In this paper, we propose a model-guided deep unfolding network for the more challenging and realistic mixed noise video denoising problem, named DU-MVDnet. First, we develop a novel observation model/likelihood function based on the correlations among adjacent degraded frames. In the framework of Bayesian deep learning, we introduce a deep image denoiser prior and obtain an iterative optimization algorithm based on the maximum a posterior (MAP) estimation. To facilitate end-to-end optimization, the iterative algorithm is transformed into a deep convolutional neural network (DCNN)-based implementation. Furthermore, recognizing the limitations of traditional motion estimation and compensation methods, we propose an efficient multistage recursive fusion strategy to exploit temporal dependencies. Specifically, we divide video frames into several overlapping groups and progressively integrate these frames into one frame. Toward this objective, we implement a multiframe adaptive aggregation operation to integrate feature maps of intragroup with those of intergroup frames. Extensive experimental results on different video test datasets have demonstrated that the proposed model-guided deep network outperforms current state-of-the-art video denoising algorithms such as FastDVDnet and MAP-VDNet. Xin Li 0005, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Depth Perception Assessment of 3D Videos Based on Stereoscopic and Spatial Orientation Structural FeaturesabstractDepth quality of stereoscopic three-dimensional (S3D) videos is a significant factor which directly affects the quality of experience (QoE) associated with 3D video applications and services. Nevertheless, there remain limited reports on the investigation of depth perception and depth quality evaluation of S3D videos, which impedes further advancement and deployment of 3D video technology. This paper reports a series of subjective experiments which have been conducted to investigate the depth perception and its related properties of the human visual system (HVS) using S3D video compressed by the H.264/AVC standard. The experimental results reveal that the HVS response in depth perception varies at different frequencies and in varying orientations, and the distortions introduced by video coding can cause the loss of and/or variation in depth perception. By integration of binocular and monocular features (BM) extracted from left and right views of S3D video with respect to depth perception, a depth quality assessment model, herein referred to as BM-DQAM, is devised by training these stereoscopic and spatial orientation structural features with a support vector regression model. It is shown that the BM-DQAM provides a novel no-reference metric for evaluation of the depth quality of S3D videos. Based on two publicly available 3D video databases and the proposed depth perception assessment database, the experimental results show that the BM-DQAM has demonstrated better performance in assessing the depth quality in S3D video viewing than that of other metrics reported in the published literatures, correlating well with the HVS response in the depth perception assessment experiment. Wenfei Wan, Dengjia Huang, Bin Shang, Shengyu Wei, Hong Ren Wu, Jinjian Wu, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Filter Clustering for Compressing CNN Model With Better Feature DiversityabstractAs a practical approach for compressing convolutional neural networks (CNNs), network pruning has been rapidly developed in recent years. The conventional methods prune inactive filters permanently from models to reduce the width of each layer and then train the pruned model until convergence. However, such methods have limitations in that: (1) The activation-based pruning criteria ignore the correlation between filters, leading to attenuation in types of features; (2) The permanent filter removal restricts the architecture of models in the subsequent training so that reducing the chances of learning more features; (3) The single-width compression may generate narrow layers that block the information flow, resulting in limited feature capacity in the next layers and hard optimization. These limitations reduce the feature diversity in the pruned model and thus lead to sub-optimal model quality. In this paper, a compression method named filter clustering is proposed to rectify the problem of poor feature diversity in traditional pruning and achieve better model quality from three perspectives. Firstly, to maintain the variety of features after pruning, we treat the model compression as a clustering task and merge filters with similar outputs, rather than removing inactive filters. Specifically, a handy estimation approach is designed to convert the similarity of the output into filter similarity, which liberates the measurement from sampling numerous images. Secondly, to increase the probability of learning more features during training, we propose a periodic training and clustering pipeline, which creates a larger optimization space by dynamically exploring different sub-model architectures. Finally, to prevent the feature capacity from being influenced by the narrow layers, we introduce and leverage a fusible anti-blocking branch to smoothly remove such layers. Extensive experiments demonstrate that the proposed method can achieve compact models with better feature diversity and reduce 1%~15% more calculations than the previous methods while maintaining performance. Zhenyu Wang 0008, Xuemei Xie, Qinghang Zhao, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Supervised Contrastive Learning Based on Fusion of Global and Local Features for Remote Sensing Image RetrievalabstractWith the rapid development of remote sensing sensor technology, the number of remote sensing images (RSIs) has exploded. How to effectively retrieve and manage these massive data has become an urgent problem. At present, content-based image retrieval (CBIR) methods have become a mainstream method due to their excellent performance. However, most of the existing retrieval methods only consider the global features of images, which lacks the ability to discriminate images with the same semantic information but different visual representations. To alleviate this issue, supervised contrastive learning based on the fusion of global and local features method is proposed in this paper, named SCFR. Firstly, a fusion module is designed to combine global and local features to enhance the ability of image expression. Secondly, supervised contrastive learning is introduced into the retrieval task to effectively improve the feature distribution, so that the positive sample pairs are close to each other, and the negative sample pairs are far away from each other in the feature space. Furthermore, to make the distribution of features of the same class more compact, the center contrastive loss is added to the constraints, and combines the class centers that change iteratively with the network. Experimental results on three RSI datasets show that our proposed method has more effective retrieval performance than the state-of-the-art methods. The code and models are available at https://github.com/xdplay17/SCFR. Mengluan Huang, Weisheng Dong, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Spatially Varying Prior Learning for Blind Hyperspectral Image FusionabstractIn recent years, researchers have become more interested in hyperspectral image fusion (HIF) as a potential alternative to expensive high-resolution hyperspectral imaging systems, which aims to recover a high-resolution hyperspectral image (HR-HSI) from two images obtained from low-resolution hyperspectral (LR-HSI) and high-spatial-resolution multispectral (HR-MSI). It is generally assumed that degeneration in both the spatial and spectral domains is known in traditional model-based methods or that there existed paired HR-LR training data in deep learning-based methods. However, such an assumption is often invalid in practice. Furthermore, most existing works, either introducing hand-crafted priors or treating HIF as a black-box problem, cannot take full advantage of the physical model. To address those issues, we propose a deep blind HIF method by unfolding model-based maximum a posterior (MAP) estimation into a network implementation in this paper. Our method works with a Laplace distribution (LD) prior that does not need paired training data. Moreover, we have developed an observation module to directly learn degeneration in the spatial domain from LR-HSI data, addressing the challenge of spatially-varying degradation. We also propose to learn the uncertainty (mean and variance) of LD models using a novel Swin-Transformer-based denoiser and to estimate the variance of degraded images from residual errors (rather than treating them as global scalars). All parameters of the MAP estimation algorithm and the observation module can be jointly optimized through end-to-end training. Extensive experiments on both synthetic and real datasets show that the proposed method outperforms existing competing methods in terms of both objective evaluation indexes and visual qualities. Xin Li 0005, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 6 |
| 2023 | Knowledge-Guided Blind Image Quality Assessment With Few Training SamplesabstractBlind image quality assessment (BIQA) for in-the-wild images has achieved great progress by training advanced deep neural networks. However, the current BIQA models are suffering the generalization challenge, meaning that a well-trained BIQA model is still very limited in evaluating images with different distributions. Deep BIQA models are data-intensive, but the annotation of image quality labels is extremely expensive. To design a generalizable BIQA model with few training samples is highly desired. Motivated by the above fact, this paper presents a knowledge-guided BIQA (KG-IQA) framework by integrating domain knowledge from the human visual system (HVS) and natural scene statistics (NSS). Specifically, the quality-aware HVS and NSS features are first extracted as prior knowledge. Then, we embed the two types of knowledge into the conventional deep neural network by learning to predict the HVS and NSS features, producing the knowledge-enhanced quality features, based on which the final image quality score is obtained. We conduct extensive experiments and comparisons on five authentically distorted IQA datasets. The experimental results demonstrate that the introduction of knowledge greatly reduces the requirement on the amount of training images, and the proposed KG-IQA model achieves superior performance in terms of both prediction accuracy and generalization ability. Tianshu Song, Leida Li, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi |
IEEE Trans. Multim. | 7 |
| 2023 | Task-Oriented Explainable Semantic CommunicationsabstractSemantic communications utilize the transceiver computing resources to alleviate scarce transmission resources, such as bandwidth and energy. Although the conventional deep learning (DL) based designs may achieve certain transmission efficiency, the uninterpretability issue of extracted features is the major challenge in the development of semantic communications. In this paper, we propose an explainable and robust semantic communication framework by incorporating the well-established bit-level communication system, which not only extracts and disentangles features into independent and semantically interpretable features, but also only selects task-relevant features for transmission, instead of all extracted features. Based on this framework, we derive the optimal input for rate-distortion-perception theory, and derive both lower and upper bounds on the semantic channel capacity. Furthermore, based on the$\beta $-variational autoencoder ($\beta $-VAE), we propose a practical explainable semantic communication system design, which simultaneously achieves semantic features selection and is robust against semantic channel noise. We further design a real-time wireless mobile semantic communication proof-of-concept prototype. Our simulations and experiments demonstrate that our proposed explainable semantic communications system can significantly improve transmission efficiency, and also verify the effectiveness of our proposed robust semantic transmission scheme. Shuai Ma 0002, Weining Qiao, Youlong Wu, Hang Li 0003, Guangming Shi, Dahua Gao, Yuanming Shi, Shiyin Li, Naofal Al-Dhahir |
IEEE Trans. Wirel. Commun. | 5 |
| 2022 | Robust Depth Completion with Uncertainty-Driven Loss FunctionsabstractRecovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumulated outliers in the synthesized ground truth. In this work, we introduce uncertainty-driven loss functions to improve the robustness of depth completion and handle the uncertainty in depth completion. Specifically, we propose an explicit uncertainty formulation for robust depth completion with Jeffrey's prior. A parametric uncertain-driven loss is introduced and translated to new loss functions that are robust to noisy or missing data. Meanwhile, we propose a multiscale joint prediction model that can simultaneously predict depth and uncertainty maps. The estimated uncertainty map is also used to perform adaptive prediction on the pixels with high uncertainty, leading to a residual map for refining the completion results. Our method has been tested on KITTI Depth Completion Benchmark and achieved the state-of-the-art robustness performance in terms of MAE, IMAE, and IRMSE metrics. Yufan Zhu, Weisheng Dong, Leida Li, Jinjian Wu, Xin Li 0005, Guangming Shi |
AAAI | 6 |
| 2022 | Uncertainty Learning in Kernel Estimation for Multi-stage Blind Image Super-Resolution
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
ECCV (18) | 6 |
| 2022 | Self-feature Distillation with Uncertainty Modeling for Degraded Image Recognition
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
ECCV (24) | 6 |
| 2022 | Supervised Contrastive Learning-Based Deep Hash Retrieval for Remote Sensing ImageabstractWith the development of remote sensing technology, the earth observation data shows a blowout growth. How to quickly and accurately retrieve the target content from the massive data has become a noteworthy task. Recently, the methods based on convolutional neural networks (CNN) have been far ahead in remote sensing retrieval. However, due to the diversity of sources acquired, remote sensing images with the same semantic information may have great visual differences. The pre-trained CNN is not always able to successfully extract representative and distinguishing features to recognize the target content. To address this problem, a supervised contrastive learning-based deep hash retrieval method (SCLDHR) is introduced in this paper, which effectively uses label information to gather image features belonging to the same class and separate image features of different classes in the embedding space. Furthermore, a quantization function in the hash space is designed to improve hashing quality. The experimental results conducted on three datasets show that SCLDHR can achieve competitive retrieval performance compared with state-of-the-art methods. Mengluan Huang, Weisheng Dong, Guangming Shi |
IGARSS | 4 |
| 2022 | Learning Degradation Uncertainty for Unsupervised Real-world Image Super-resolutionabstractAcquiring degraded images with paired high-resolution (HR) images is often challenging, impeding the advance of image super-resolution in real-world applications. By generating realistic low-resolution (LR) images with degradation similar to that in real-world scenarios, simulated paired LR-HR data can be constructed for supervised training. However, most of the existing work ignores the degradation uncertainty of the generated realistic LR images, since only one LR image has been generated given an HR image. To address this weakness, we propose learning the degradation uncertainty of generated LR images and sampling multiple LR images from the learned LR image (mean) and degradation uncertainty (variance) and construct LR-HR pairs to train the super-resolution (SR) networks. Specifically, uncertainty can be learned by minimizing the proposed loss based on Kullback-Leibler (KL) divergence. Furthermore, the uncertainty in the feature domain is exploited by a novel perceptual loss; and we propose to calculate the adversarial loss from the gradient information in the SR stage for stable training performance and better visual quality. Experimental results on popular real-world datasets show that our proposed method has performed better than other unsupervised approaches. Qian Ning, Jingzhu Tang, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 6 |
| 2022 | Rate-Distortion Theory for Strategic Semantic CommunicationabstractThis paper analyzes the fundamental limit of the strategic semantic communication problem in which a transmitter obtains a limited number of indirect observations of an intrinsic semantic information source and can then influence the receiver’s decoding by sending a limited number of messages over an imperfect channel. The transmitter and the receiver can have different distortion measures and can make rational decisions about their encoding and decoding strategies, respectively. The decoder can also have some side information (e.g., background knowledge and/or information obtained from previous communications) about the semantic source to assist its interpretation of the semantic information. We focus particularly on the case that the transmitter can commit to an encoding strategy and study the impact of the strategic decision making on the rate distortion of semantic communication. Three equilibrium solution concepts including the optimal Stackelberg equilibrium, robust Stackelberg equilibrium, as well as Nash equilibrium are studied and compared. The optimal encoding and decoding strategy profiles under various equilibrium solutions are derived. We prove that committing to an encoding strategy cannot always bring benefit to the encoder. We provide a feasible condition under which committing to an encoding strategy can always reduce the distortion of semantic communication. We consider an example with a dictionary-based semantic information source to verify our observation. Yong Xiao 0001, Yingyu Li, Guangming Shi, Tamer Basar |
ITW | 4 |
| 2022 | AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular DataabstractDynamic Vision Sensor (DVS) is a compelling neuromorphic camera compared to conventional camera, but it suffers from fiercer noise. Due to the nature of irregular format and asynchronous readout, DVS data is always transformed into a regular tensor (e.g., 3D voxel or image) for deep learning method, which corrupts its own asynchronous properties. To maintain asynchronous, we establish an innovative asynchronous event denoise neural network, named AEDNet, which directly consumes the correlation of the irregular signal in spatial-temporal range without destroying its original structural property. Based on the property of continuation in temporal domain and discreteness in spatial domain, we decompose the DVS signal into two parts, i.e., temporal correlation and spatial affinity, and separately process these two parts. Our spatial feature embedding unit is a unique feature extraction module that extracts feature from event-level, which perfectly maintains its spatial-temporal correlation. To test effectiveness, we build a novel dataset named DVSCLEAN containing both simulated and real-world data. The experimental results of AEDNet achieve SOTA. Huachen Fang, Jinjian Wu, Leida Li, Junhui Hou, Weisheng Dong, Guangming Shi |
ACM Multimedia | 6 |
| 2022 | Bayesian based Re-parameterization for DNN Model PruningabstractFilter pruning, as an effective strategy to obtain efficient compact structures from over-parametric deep neural networks(DNN), has attracted a lot of attention. Previous pruning methods select channels for pruning by developing different criteria, yet little attention has been devoted to whether these criteria can represent correlations between channels. Meanwhile, most existing methods generally ignore the parameters being pruned and only perform additional training on the retained network to reduce accuracy loss. In this paper, we present a novel perspective of re-parametric pruning by Bayesian estimation. First, we estimate the probability distribution of different channels based on Bayesian estimation and indicate the importance of the channels by the discrepancy in the distribution before and after channel pruning. Second, to minimize the variation in distribution after pruning, we re-parameterize the pruned network based on the probability distribution to pursue optimal pruning. We evaluate our approach on popular datasets with some typical network architectures, and comprehensive experimental results validate that this method illustrates better performance compared to the state-of-the-art approaches. Xiaotong Lu, Teng Xi, Baopu Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 6 |
| 2022 | Learning for Motion Deblurring with Hybrid Frames and EventsabstractEvent camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events. Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 6 |
| 2022 | Semantic Attribute Guided Image Aesthetics AssessmentabstractImage aesthetics assessment (IAA) measures the perceived beauty of images using a computational approach. People usually assess the aesthetics of an image according to semantic attributes, e.g., lighting, color, object emphasis, etc. However, the state-of-the-art IAA approaches usually follow the data-driven framework without considering the rich attributes contained in images. With this motivation, this paper presents a new semantic attribute guided IAA model, where the attention maps of semantic attributes are employed to enhance the representation ability of general aesthetic features for more effective aesthetics assessment. Specifically, we first design an attribute attention generation network to obtain the attention maps for different semantic attributes, which are utilized to weight the general aesthetic features, producing the semantic attribute-enhanced feature representations. Then, the Graph Convolutional Network (GCN) is employed to further investigate the inherent relationship among the enhanced aesthetic features, producing the final image aesthetics prediction. Extensive experiments and comparisons on three public IAA databases demonstrate the effectiveness of the proposed method. Jiachen Duan, Pengfei Chen 0003, Leida Li, Jinjian Wu, Guangming Shi |
VCIP | 5 |
| 2022 | Robust Dynamic Background Modeling for Foreground EstimationabstractSeparating the background and foreground components from video frames is important to many tasks in computer vision and multimedia. As of today, robust principal component analysis (RPCA) has shown highly promising performance with the assumption that the background is low-rank and the foreground is sparse. However, existing RPCA-based methods have overlooked the uncertainty that some parts of the background (e.g., moving leaves in a dynamic background) or even the whole background (e.g., camera jittering) can be moving, which violates the low-rank assumption. To address this issue, we propose a novel enhanced RPCA framework (called ERPCA) by robustly modeling the dynamic background. Different from traditional RPCA framework, the background is decomposed into a low-rank component and a sparse component in the proposed ERPCA framework. Specifically, the sparse parts including foreground and dynamic parts of the background are modeled by Gaussian scale mixture (GSM) model. Moreover, those sparse components are further constrained by temporal consistency using nonzeromeans Gaussian models; the correspondences between sparse pixels in adjacent frames are explored by optical flow. Experimental results on 40 real videos demonstrate the superiority of our proposed method, with better average results than current state-of-the-art foreground estimation methods. Qian Ning, Weisheng Dong, Jinjian Wu, Guangming Shi, Xin Li 0005 |
VCIP | 5 |
| 2022 | ACO: lossless quality score compression based on adaptive coding orderabstractBACKGROUND: With the rapid development of high-throughput sequencing technology, the cost of whole genome sequencing drops rapidly, which leads to an exponential growth of genome data. How to efficiently compress the DNA data generated by large-scale genome projects has become an important factor restricting the further development of the DNA sequencing industry. Although the compression of DNA bases has achieved significant improvement in recent years, the compression of quality score is still challenging. RESULTS: In this paper, by reinvestigating the inherent correlations between the quality score and the sequencing process, we propose a novel lossless quality score compressor based on adaptive coding order (ACO). The main objective of ACO is to traverse the quality score adaptively in the most correlative trajectory according to the sequencing process. By cooperating with the adaptive arithmetic coding and an improved in-context strategy, ACO achieves the state-of-the-art quality score compression performances with moderate complexity for the next-generation sequencing (NGS) data. CONCLUSIONS: The competence enables ACO to serve as a candidate tool for quality score compression, ACO has been employed by AVS(Audio Video coding Standard Workgroup of China) and is freely available at https://github.com/Yoniming/ACO. Mingming Ma, Fu Li 0002, Guangming Shi |
BMC Bioinform. | 5 |
| 2022 | Adaptive feature denoising based deep convolutional network for single image super-resolutionabstractRecently, the feature map recalibration (FMR) mechanism has been widely explored in single image super-resolution (SISR) and obtained remarkable performances. However, the existing FMR-based SISR methods directly incorporate the attention module into a deeper network structure (e.g. EDSR), while neglecting the differences between the low-level and high-level vision problems. In this paper, we design a low-level specific FMR mechanism for SISR task based on a new observation by examining current SISR methods, which all demonstrate a solid correlation between the SISR performance and the convolutional feature noise. Inspired by this, we extend the classic soft thresholding technique in the way of deep network, and develop an Adaptive Soft Thresholding (AST) module for feature noise suppression. Comparing to existing attention modules, AST is light-weighted and can be taken as an easy plug-in module in any SISR networks. To this end, we construct a adaptive Feature Denoising Super-Resolution (FDSR) network by combining the baseline EDSR and the proposed AST. Extensive experimental results show that the proposed FDSR network could achieve the state-of-the-art performances on SISR benchmarks, and significantly reduce the parameter (28.8% for EDSR, 74.6% for RCAN,s 76.0% for SAN) with respect to FMR module. Rui Cheng 0010, Jia Wang 0018, Mingming Ma, Guangming Shi |
Comput. Vis. Image Underst. | 6 |
| 2022 | Correlation filters based on spatial-temporal Gaussion scale mixture modelling for visual tracking
Guangming Shi, Weisheng Dong, Tianzhu Zhang 0001, Jinjian Wu, Xuemei Xie, Xin Li 0005 |
Neurocomputing | 2 |
| 2022 | Blind image quality assessment based on progressive multi-task learning
Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi |
Neurocomputing | 6 |
| 2022 | Detecting human-object interactions in videos by modeling the trajectory of objects and human skeleton
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
Neurocomputing | 5 |
| 2022 | Soft focal loss: Evaluating sample quality for dense object detection
Zhenyuan Wang, Xuemei Xie, Jianxiu Yang, Guangming Shi |
Neurocomputing | 4 |
| 2022 | Modeling content-attribute preference for personalized image esthetics assessment
Yuanyang Wang, Yihua Huang 0003, Xiumin Chen, Leida Li, Guangming Shi |
Image Vis. Comput. | 5 |
| 2022 | Language-guided graph parsing attention network for human-object interaction recognition
Qiyue Li 0004, Xuemei Xie, Guangming Shi |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | A novel intrinsically explainable model with semantic manifolds established via transformed priors
Guangming Shi, Minxi Yang, Dahua Gao |
Knowl. Based Syst. | 1 |
| 2022 | SAR Imaging and Despeckling Based on Sparse, Low-Rank, and Deep CNN PriorsabstractSynthetic aperture radar (SAR) generally suffers from enormous strains from large quantities of sampling data and serious interferences from the speckle noise. This letter proposes a novel deep network to address these problems. By utilizing the prior knowledge in a more reasonable way, the proposed network could realize SAR imaging and despeckling with down-sampled data simultaneously. Specifically, we decompose the SAR image in the SAR imaging-despeckling observation model into a sparse matrix and a low-rank matrix, and then establish an optimization problem with the corresponding sparse and low-rank priors. Moreover, the deep convolutional neural networks (CNN) denoiser prior is also introduced to further improve the speckle reduction capability. Then, we devise a deep network called SLRCP-Net to solve this problem. Experiments conducted on real Radarsat-1 down-sampled data demonstrate the validity of SLRCP-Net in SAR imaging and speckle suppression. Guanghui Zhao 0003, Yingbin Wang, Guangming Shi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Bayesian Correlation Filter Learning With Gaussian Scale Mixture Model for Visual TrackingabstractCorrelation filters (CF), a popular tool for visual tracking, suffer from unwanted boundary effects due to the periodic assumption needed for FFT implementation. To address this issue, spatially regularized discriminative correlation filters (SRDCF) have been proposed by introducing a weighting matrix to the regularization term. However, the existing design of spatial weighting matrix is often heuristic and non-adaptive. Inspired by recent advances in joint discrimination and reliability learning for correlation tracking, we propose a principled Bayesian correlation filter learning method using Gaussian scale mixture (GSM) model. The key idea is to decompose each CF coefficient into the product of a positive scalar multiplier and a Gaussian random variable. Treating positive multipliers as weighting coefficients, GSM-based modeling of CFs leads to a spatially adaptive regularization strategy with improved capability of handling various appearance-related uncertainty factors (e.g., scale variation, out-of-plane rotation, and motion blur). Moreover, by imposing a sparse prior over the multipliers, we can jointly learn multipliers and CFs under a unified Bayesian estimation framework. Structured GSM model allows us to better exploit the spatial correlations among CFs and further improve the tracking performance. Experimental results on OTB-2013, OTB-2015, Temple Color-128, VOT-2016, and VOT-2017 show that our tracking method performs favorably when compared with current state-of-the-art methods. Guangming Shi, Tianzhu Zhang 0001, Weisheng Dong, Jinjian Wu, Xuemei Xie, Xin Li 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Blind Image Quality Index for Authentic Distortions With Local and Global Deep Feature AggregationabstractBlind image quality assessment (BIQA) for authentic distortions is still a great challenge, even in today’s deep learning era. It has been widely acknowledged that local and global features are both indispensable for IQA, which play complementary roles. While combining local and global features is straightforward in traditional handcrafted feature-based IQA metrics, it is not an easy task in the deep learning framework. This is mainly due to the fact that deep neural networks typically require input images with a fixed size. Current metrics either resize the image or use local patches as input, which are problematic in that they cannot integrate local and global aspects as well as their interactions to achieve comprehensive quality evaluation. Motivated by the above facts, this paper presents a new BIQA metric for authentic distortions by aggregating local and global deep features in a Vision-Transformer framework. In the proposed metric, selective local regions and global content are simultaneously input for complementary feature extraction, and the Vision-Transformer is employed to build the relationship between different local patches and image quality. Self-attention mechanism is further adopted to explore the interaction between local and global deep features, producing the final image quality score. Extensive experiments on five authentically distorted IQA databases demonstrate that the proposed metric outperforms the state-of-the-arts in terms of both prediction performance and generalization ability. Leida Li, Tianshu Song, Jinjian Wu, Weisheng Dong, Jiansheng Qian, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Spatiotemporal Representation Learning for Blind Video Quality AssessmentabstractBlind video quality assessment (BVQA) is of great importance for video-related applications, yet still challenging even in this deep learning era. The difficulty lies in the shortage of large-scale labeled data, thus making it hard to train a robust spatiotemporal encoder for BVQA. To relieve such difficulty, we first build a video dataset, which contains over 320K samples suffering from various compression and transmission artifacts. While manually annotating the dataset with subjective perception is much labor-intensive and time-consuming, we adopt reference-based VQA algorithms to weakly label the data automatically. We consider that single weak label is derived from single knowledge, which is deficient and incomplete for VQA. To alleviate the bias from single weak label (i.e., single knowledge) in the weakly labeled dataset, we propose HEterogeneous Knowledge Ensemble (HEKE) for spatiotemporal representation learning. Compared to learning from single knowledge, learning with HEKE is thought to achieve a lower infimum theoretically, and obtain richer representation. On the basis of the built dataset and the HEKE methodology, a feature encoder specific to BVQA is formed, and directly extract spatiotemporal representation from videos. Then, the video quality can be either acquired in a completely BVQA manner without ground truth, or via a finetuning-based regressor with labels. Extensive experiments on various VQA databases show that our BVQA model with the pretrained encoder achieves the state-of-the-art performance. More surprisingly, even trained on the synthetic data, our model still shows competitive performance on authentic databases. The data and source code will be available athttps://github.com/Sissuire/BVQA-HEKE. Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Generalizable No-Reference Image Quality Assessment via Deep Meta-LearningabstractRecently, researchers have shown great interest in using convolutional neural networks (CNNs) for no-reference image quality assessment (NR-IQA). Due to the lack of big training data, the efforts of existing metrics in optimizing CNN-based NR-IQA models remain limited. Furthermore, the diversity of distortions in images result in the generalization problem of NR-IQA models when trained with known distortions and tested on unseen distortions, which is an easy task for human. Hence, we propose a NR-IQA metric via deep meta-learning, which is highly generalizable in the face of unseen distortions. The fundamental idea is to learn the meta-knowledge shared by human when evaluating the quality of images with diversified distortions. Specifically, we define NR-IQA of different distortions as a series of tasks and propose a task selection strategy to build two task sets, which are characterized by synthetic to synthetic and synthetic to authentic distortions, respectively. Based on these two task sets, an optimization-based meta-learning is proposed to learn the generalized NR-IQA model, which can be directly used to evaluate the quality of images with unseen distortions. Extensive experiments demonstrate that our NR-IQA metric outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability. Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Personalized Image Aesthetics Assessment via Meta-Learning With Bilevel Gradient OptimizationabstractTypical image aesthetics assessment (IAA) is modeled for the generic aesthetics perceived by an "average" user. However, such generic aesthetics models neglect the fact that users' aesthetic preferences vary significantly depending on their unique preferences. Therefore, it is essential to tackle the issue for personalized IAA (PIAA). Since PIAA is a typical small sample learning (SSL) problem, existing PIAA models are usually built by fine-tuning the well-established generic IAA (GIAA) models, which are regarded as prior knowledge. Nevertheless, this kind of prior knowledge based on "average aesthetics" fails to incarnate the aesthetic diversity of different people. In order to learn the shared prior knowledge when different people judge aesthetics, that is, learn how people judge image aesthetics, we propose a PIAA method based on meta-learning with bilevel gradient optimization (BLG-PIAA), which is trained using individual aesthetic data directly and generalizes to unknown users quickly. The proposed approach consists of two phases: 1) meta-training and 2) meta-testing. In meta-training, the aesthetics assessment of each user is regarded as a task, and the training set of each task is divided into two sets: 1) support set and 2) query set. Unlike traditional methods that train a GIAA model based on average aesthetics, we train an aesthetic meta-learner model by bilevel gradient updating from the support set to the query set using many users' PIAA tasks. In meta-testing, the aesthetic meta-learner model is fine-tuned using a small amount of aesthetic data of a target user to obtain the PIAA model. The experimental results show that the proposed method outperforms the state-of-the-art PIAA metrics, and the learned prior model of BLG-PIAA can be quickly adapted to unseen PIAA tasks. Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, Guangming Shi |
IEEE Trans. Cybern. | 6 |
| 2022 | Lq-SPB-Net: A Real-Time Deep Network for SAR Imaging and DespecklingabstractLarge quantities of sampling data and speckle noise are two serious problems existing in synthetic aperture radar (SAR). The former puts enormous strain on data measurement, transmission, and storage. The latter deteriorates imaging quality, disturbing the subsequent processing in SAR systems. This study proposes a real-time deep network to address these issues. The proposed network is able to concurrently achieve SAR imaging and despeckling with down-sampled data. Specifically, to fit more diverse imaging regions, we suppose that the noise in the SAR imaging-despeckling observation model follows a universal complex generalized Gaussian distribution. Based on this assumption, an optimization problem with a convex$L_{q}$-norm ($q >$1) fidelity term is constructed through the maximum$a$posteriori(MAP) estimation. We employ an$L_{1}$-norm sparse constraint and a convolutional neural networks (CNNs) projection-based detail preservation constraint to further promote the imaging and speckle reduction capabilities. Then, the complex-valued split Bregman method (CV-SBM) is applied to convert the proposed problem into an equivalent sequence of sub-problems. We devise a computationally efficient solution for the fidelity term-related sub-problem, due to the specific down-sampled strategy in SAR. A substitutive cost function and a CNN structure are introduced to solve the projection-related sub-problem. Finally, the iterative steps of CV-SBM are cast into a deep network-dubbed$L_{q}$-split Bregman (SPB)-Net to yield a desirable imaging and despeckling result within a small number of iterations. Numerical experiments based on the real Radarsat-1 data validate the efficiency and feasibility of the proposed$L_{q}$-SPB-Net in real-time imaging and despeckling with down-sampled data. Guanghui Zhao 0003, Yingbin Wang, Guangming Shi, Shuxuan Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | S³Net: Spectral-Spatial-Semantic Network for Hyperspectral Image Classification With the Multiway Attention MechanismabstractIn hyperspectral image (HSI) classification, it is a great challenge on how to extract key informative spectral–spatial features efficiently and suppress useless features from abundant spectral–spatial information. In this article, inspired by the attention mechanism of the human visual system, we propose a novel spectral–spatial–semantic network (S3Net) with the multiway attention mechanism for HSI classification. The S3Net consists of a spectral branch, a spatial branch, and a multiscale semantic module. The spectral branch extracts the multiway spectral features with a dense spectral block and a multiway spectral attention module. The spatial branch extracts the multiway spatial features with a multiscale spatial block and a multiway spatial attention module. The multiscale semantic module extracts spectral–spatial–semantic features that are used for classification. In the proposed S3Net, the multiway attention modules in spectral and spatial branches are built to enhance the extraction ability of the key informative spectral–spatial features. The multiscale spatial block in the spatial branch is designed to learn strong complementary and related information. The Res2Net in the multiscale semantic module is used to learn multiscale semantic features at a granular level. A large number of experimental results demonstrate that, on the University of Pavia, Kennedy Space Center, and Pavia Center data sets, the proposed S3Net achieves higher classification accuracy than state-of-the-art methods on the limited training samples. Remarkably, our S3Net also achieves the best performance on the mineral exploration HSI data set, called Huoshaoyun, which is collected by the GaoFen-5 (GF-5) satellite. Danhua Liu, Dahua Gao, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Contrastive Self-Supervised Pre-Training for Video Quality AssessmentabstractVideo quality assessment (VQA) task is an ongoing small sample learning problem due to the costly effort required for manual annotation. Since existing VQA datasets are of limited scale, prior research tries to leverage models pre-trained on ImageNet to mitigate this kind of shortage. Nonetheless, these well-trained models targeting on image classification task can be sub-optimal when applied on VQA data from a significantly different domain. In this paper, we make the first attempt to perform self-supervised pre-training for VQA task built upon contrastive learning method, targeting at exploiting the plentiful unlabeled video data to learn feature representation in a simple-yet-effective way. Specifically, we implement this idea by first generating distorted video samples with diverse distortion characteristics and visual contents based on the proposed distortion augmentation strategy. Afterwards, we conduct contrastive learning to capture quality-aware information by maximizing the agreement on feature representations of future frames and their corresponding predictions in the embedding space. In addition, we further introduce distortion prediction task as an additional learning objective to push the model towards discriminating different distortion categories of the input video. Solving these prediction tasks jointly with the contrastive learning not only provides stronger surrogate supervision signals, but also learns the shared knowledge among the prediction tasks. Extensive experiments demonstrate that our approach sets a new state-of-the-art in self-supervised learning for VQA task. Our results also underscore that the learned pre-trained model can significantly benefit the existing learning based VQA models. Source code is available at https://github.com/cpf0079/CSPT. Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2022 | Visual Cluster Grounding for Image CaptioningabstractAttention mechanisms have been extensively adopted in vision and language tasks such as image captioning. It encourages a captioning model to dynamically ground appropriate image regions when generating words or phrases, and it is critical to alleviate the problems of object hallucinations and language bias. However, current studies show that the grounding accuracy of existing captioners is still far from satisfactory. Recently, much effort is devoted to improving the grounding accuracy by linking the words to the full content of objects in images. However, due to the noisy grounding annotations and large variations of object appearance, such strict word-object alignment regularization may not be optimal for improving captioning performance. In this paper, to improve the performance of both grounding and captioning, we propose a novel grounding model which implicitly links the words to the evidence in the image. The proposed model encourages the captioner to dynamically focus on informative regions of the objects, which could be either discriminative parts or full object content. With slacked constraints, the proposed captioning model can capture correct linguistic characteristics and visual relevance, and then generate more grounded image captions. In addition, we propose a novel quantitative metric for evaluating the correctness of the soft attention mechanism by considering the overall contribution of all object proposals when generating certain words. The proposed grounding model can be seamlessly plugged into most attention-based architectures without introducing inference complexity. We conduct extensive experiments on Flickr30k (Young et al., 2014) and MS COCO datasets (Lin et al., 2014), demonstrating that the proposed method consistently improves image captioning in both grounding and captioning. Besides, the proposed attention evaluation metric shows better consistency with the captioning performance. Wenhui Jiang 0001, Minwei Zhu, Yuming Fang 0001, Guangming Shi, Yang Liu 0293 |
IEEE Trans. Image Process. | 4 |
| 2022 | From Whole Video to Frames: Weakly-Supervised Domain Adaptive Continuous-Time QoE EvaluationabstractDue to the rapid increase in video traffic and relatively limited delivery infrastructure, end users often experience dynamically varying quality over time when viewing streaming videos. The user quality-of-experience (QoE) must be continuously monitored to deliver an optimized service. However, modern approaches for continuous-time video QoE estimation require densely annotating the continuous-time QoE labels, which is labor-intensive and time-consuming. To cope with such limitations, we propose a novel weakly-supervised domain adaptation approach for continuous-time QoE evaluation, by making use of a small amount of continuously labeled data in the source domain and abundant weakly-labeled data (only containing the retrospective QoE labels) in the target domain. Specifically, given a pair of videos from source and target domains, effective spatiotemporal segment-level feature representation is first learned by a combination of 2D and 3D convolutional networks. Then, a multi-task prediction framework is developed to simultaneously achieve continuous-time and retrospective QoE predictions, where a quality attentive adaptation approach is investigated to effectively alleviate the domain discrepancy without hampering the prediction performance. This approach is enabled by explicitly attending to the video-level discrimination and segment-level transferability in terms of the domain discrepancy. Experiments on benchmark databases demonstrate that the proposed method significantly improves the prediction performance under the cross-domain setting. Leida Li, Pengfei Chen 0003, Weisi Lin, Mai Xu, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2022 | NL-CALIC Soft Decoding Using Strict Constrained Wide-Activated Recurrent Residual Network
Chang Liu 0047, Mingming Ma, Fu Li 0002, Zhiwen Chen 0002, Guangming Shi |
IEEE Trans. Image Process. | 6 |
| 2022 | Fine-Grained Image Quality Caption With Hierarchical Semantics DegradationabstractBlind image quality assessment (BIQA), which is capable of precisely and automatically estimating human perceived image quality with no pristine image for comparison, attracts extensive attention and is of wide applications. Recently, many existing BIQA methods commonly represent image quality with a quantitative value, which is inconsistent with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where humans are able to extract image contents with local to global semantics as they need. The mediocre quality value represents coarse or holistic image quality and fails to reflect degradation on hierarchical semantics. In this paper, to comply with human cognition, a novel quality caption model is inventively proposed to measure fine-grained image quality with hierarchical semantics degradation. Research on human visual system indicates there are hierarchy and reverse hierarchy correlations between hierarchical semantics. Meanwhile, empirical evidence shows that there are also bi-directional degradation dependencies between them. Thus, a novel bi-directional relationship-based network (BDRNet) is proposed for semantics degradation description, through adaptively exploring those correlations and degradation dependencies in a bi-directional manner. Extensive experiments demonstrate that our method outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability. Wen Yang 0008, Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 6 |
| 2022 | Video Quality Assessment With Serial Dependence ModelingabstractVideo quality assessment (VQA) is much more challenging than image quality assessment, due to the difficulty of modeling temporal influence among frames. Most of the existing VQA methods usually isolate each moment within the video (i.e., it neglects the sequential nature), leading to a large gap from the subjective perception. Recent research on neuroscience suggests a serially dependent perception (SDP) mechanism in the human visual system (HVS). Namely, the HVS tends to incorporate the recent past visual experience to predict the present perception. Inspired by the SDP, we suggest that the HVS prefers stable and continuous degradations in videos due to their predictability, and exhibits less tolerance to interrupted and unpredictable disturbances. Thus, we introduce a novel serial dependence modeling (SDM) framework for full-reference VQA in this paper. Firstly, the instantaneous degradation is measured on both the static appearance and motion information for each glimpse of scenes. Since motion plays an important role in videos, two types of structures are extracted for motion representation, namely, an explicit content-based 3D structure and an implicit feature-based 2D structure. Next, an assessment-directed long-short term memory (A-LSTM) is proposed to capture the serial dependence among instantaneous degradations. With the consideration of the perceptual effect from the previous moment on the current one, especially the effect from the perceptually worst moment, the serially dependent degradation is characterized. Finally, by mimicking the subjective rating for video-viewing, an attention-based quality decision procedure is presented to acquire the final video quality. Experimental results on publicly available VQA databases demonstrate that the proposed method maintains good consistency with the subjective perception. Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin |
IEEE Trans. Multim. | 6 |
| 2021 | PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering NetworkabstractThe reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level annotations. In this paper, to address the above problems, we propose a novel fully convolutional Point Gathering Network (PGNet) for reading arbitrarily-shaped text in real-time. The PGNet is a single-shot text spotter, where the pixel-level character classification map is learned with proposed PG-CTC loss avoiding the usage of character-level annotations. With a PG-CTC decoder, we gather high-level character classification vectors from two-dimensional space and decode them into text symbols without NMS and RoI operations involved, which guarantees high efficiency. Additionally, reasoning the relations between each character and its neighbors, a graph refinement module (GRM) is proposed to optimize the coarse recognition and improve the end-to-end performance. Experiments prove that the proposed method achieves competitive accuracy, meanwhile significantly improving the running speed. In particular, in Total-Text, it runs at 46.7 FPS, surpassing the previous spotters with a large margin. Chengquan Zhang, Fei Qi 0001, Xiaoqiang Zhang 0006, Pengyuan Lv, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi |
AAAI | 10 |
| 2021 | Deep Gaussian Scale Mixture Prior for Spectral Compressive ImagingabstractIn coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achieved limited success due to the poor representation capability of these hand-crafted priors. Deep learning based methods learning the mappings between the compressive images and the HSIs directly achieved much better results. Yet, it is nontrivial to design a powerful deep network heuristically for achieving satisfied results. In this paper, we propose a novel HSI reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Different from existing GSM models using hand-crafted scale priors (e.g., the Jeffrey’s prior), we propose to learn the scale prior through a deep convolutional neural network (DCNN). Furthermore, we also propose to estimate the local means of the GSM models by the DCNN. All the parameters of the MAP estimation algorithm and the DCNN parameters are jointly optimized through end-to-end training. Extensive experimental results on both synthetic and real datasets demonstrate that the proposed method outperforms existing state-of-the-art methods. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/DGSM-SCI.htm. Weisheng Dong, Xin Yuan 0002, Jinjian Wu, Guangming Shi |
CVPR | 5 |
| 2021 | EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion SegmentationabstractThis paper addresses the challenging unsupervised scene flow estimation problem by jointly learning four low-level vision sub-tasks: optical flow F, stereo-depth D, camera pose P and motion segmentation S. Our key insight is that the rigidity of the scene shares the same inherent geometrical structure with object movements and scene depth. Hence, rigidity from S can be inferred by jointly coupling F, D and P to achieve more robust estimation. To this end, we propose a novel scene flow framework named EffiScene with efficient joint rigidity learning, going beyond the existing pipeline with independent auxiliary structures. In EffiScene, we first estimate optical flow and depth at the coarse level and then compute camera pose by Perspective-n-Points method. To jointly learn local rigidity, we design a novel Rigidity From Motion (RfM) layer with three principal components: (i) correlation extraction; (ii) boundary learning; and (iii) outlier exclusion. Final outputs are fused based on the rigid map MRfrom RfM at finer levels. To efficiently train EffiScene, two new losses ℒbndand ℒuncare designed to prevent trivial solutions and to regularize the flow boundary discontinuity. Extensive experiments on scene flow benchmark KITTI show that our method is effective and significantly improves the state-of-the-art approaches for all sub-tasks, i.e. optical flow (5.19→4.20), depth estimation (3.78→3.46), visual odometry (0.012→0.011) and motion segmentation (0.57→ 0.62). Trac D. Tran, Guangming Shi |
CVPR | 3 |
| 2021 | Optimizing Intelligent Reflecting Surface-Base Station Association for Mobile NetworksabstractThis paper studies a multi-Intelligent Reflecting Surfaces (IRSs)-assisted wireless network consisting of multiple base stations (BSs) serving a set of mobile users. We focus on the IRS-BS association problem in which multiple BSs compete with each other for controlling the phase shifts of a limited number of IRSs to maximize the long-term downlink data rate for the associated users. We propose MDLBI, a Multi-agent Deep Reinforcement Learning-based BS-IRS association scheme that optimizes the BS-IRS association as well as the phase-shift of each IRS when being associated with different BSs. MDLBI does not require information exchanging among BSs. Simulation results show that MDLBI achieves significant performance improvement and is scalable for large networking systems. Dongzi Jin, Yong Xiao 0001, Yingyu Li, Guangming Shi, Dusit Niyato |
ICC | 4 |
| 2021 | Spatio-temporal Modeling for Large-scale Vehicular Networks Using Graph Convolutional NetworksabstractThe effective deployment of connected vehicular networks is contingent upon maintaining a desired performance across spatial and temporal domains. In this paper, a graph-based framework, called SMART, is proposed to model and keep track of the spatial and temporal statistics of vehicle-to-infrastructure (V2I) communication latency across a large geographical area. SMART first formulates the spatio-temporal performance of a vehicular network as a graph in which each vertex corresponds to a subregion consisting of a set of neighboring location points with similar statistical features of V2I latency and each edge represents the spatio-correlation between latency statistics of two connected vertices. Motivated by the observation that the complete temporal and spatial latency performance of a vehicular network can be reconstructed from a limited number of vertices and edge relations, we develop a graph reconstruction-based approach using a graph convolutional network integrated with a deep Q-networks algorithm in order to capture the spatial and temporal statistic of feature map pf latency performance for a large-scale vehicular network. Extensive simulations have been conducted based on a five-month latency measurement study on a commercial LTE network. Our results show that the proposed method can significantly improve both the accuracy and efficiency for modeling and reconstructing the latency performance of large vehicular networks. Juntong Liu, Yong Xiao 0001, Yingyu Li, Guangming Shi, Walid Saad 0001, H. Vincent Poor |
ICC | 4 |
| 2021 | Federated Traffic Synthesizing and Classification Using Generative Adversarial NetworksabstractWith the fast growing demand on new services and applications as well as the increasing awareness of data protection, traditional centralized traffic classification approaches are facing unprecedented challenges. This paper introduces a novel framework, Federated Generative Adversarial Networks and Automatic Classification (FGAN-AC), which integrates decentralized data synthesizing with traffic classification. FGAN-AC is able to synthesize and classify multiple types of service data traffic from decentralized local datasets without requiring a large volume of manually labeled dataset or causing any data leakage. Two types of data synthesizing approaches have been proposed and compared: computation-efficient FGAN (FGAN-I) and communication-efficient FGAN (FGAN-II). The former only implements a single CNN model for processing each local dataset and the later only requires coordination of intermediate model training parameters. An automatic data classification and model updating framework has been proposed to automatically identify unknown traffic from the synthesized data samples and create new pseudo-labels for model training. Numerical results show that our proposed framework has the ability to synthesize highly mixed service data traffic and can significantly improve the traffic classification performance compared to existing solutions. Chenxin Xu, Rong Xia, Yong Xiao 0001, Yingyu Li, Guangming Shi, Kwang-Cheng Chen |
ICC | 5 |
| 2021 | Unsupervised Curriculum Domain Adaptation for No-Reference Video Quality AssessmentabstractDuring the last years, convolutional neural networks (C-NNs) have triumphed over video quality assessment (VQA) tasks. However, CNN-based approaches heavily rely on annotated data which are typically not available in VQA, leading to the difficulty of model generalization. Recent advances in domain adaptation technique makes it possible to adapt models trained on source data to unlabeled target data. However, due to the distortion diversity and content variation of the collected videos, the intrinsic subjectivity of VQA tasks hampers the adaptation performance. In this work, we propose a curriculum-style unsupervised domain adaptation to handle the cross-domain no-reference VQA problem. The proposed approach could be divided into two stages. In the first stage, we conduct an adaptation between source and target domains to predict the rating distribution for target samples, which can better reveal the subjective nature of VQA. From this adaptation, we split the data in target domain into confident and uncertain subdomains using the proposed uncertainty-based ranking function, through measuring their prediction confidences. In the second stage, by regarding samples in confident subdomain as the easy tasks in the curriculum, a fine-level adaptation is conducted between two subdomain-s to fine-tune the prediction model. Extensive experimental results on benchmark datasets highlight the superiority of the proposed method over the competing methods in both accuracy and speed. The source code is released at https://github.com/cpf0079/UCDA. Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
ICCV | 5 |
| 2021 | Optical Flow Estimation Via Motion Feature RecoveryabstractOptical flow estimation with occlusion or large displacement is a problematic challenge due to the lost of corresponding pixels between consecutive frames. In this paper, we discover that the lost information is related to a large quantity of motion features (more than 40%) computed from the popular discriminative cost-volume feature would completely vanish due to invalid sampling, leading to the low efficiency of optical flow learning. We call this phenomenon the Vanishing Cost Volume Problem. Inspired by the fact that local motion tends to be highly consistent within a short temporal window, we propose a novel iterative Motion Feature Recovery (MFR) method to address the vanishing cost volume via modeling motion consistency across multiple frames. In each MFR iteration, invalid entries from original motion features are first determined based on the current flow. Then, an efficient network is designed to adaptively learn the motion correlation to recover invalid features for lost-information restoration. The final optical flow is then decoded from the recovered motion features. Experimental results on Sintel and KITTI show that our method achieves state-of-the-art performances. In fact, MFR currently ranks second on Sintel public website. Guangming Shi, Trac D. Tran |
ICIP | 2 |
| 2021 | View-normalized Skeleton Generation for Action RecognitionabstractSkeleton-based action recognition has attracted great interest due to low cost of skeleton data acquisition and high robustness to external conditions. A challenging problem of skeleton-based action recognition is the large intra-class gap caused by various viewpoints of skeleton data, which makes the action modeling difficult for network. To alleviate this problem, a feasible solution is to utilize label supervised methods to learn a view-normalization model. However, since the skeleton data in real scenes is acquired from diverse viewpoints, it is difficult to obtain the corresponding view-normalized skeleton as label. Therefore, how to learn a view-normalization model without the supervised label is the key to solving view-variance problem. To this end, we propose a view normalization-based action recognition framework, which is composed of view-normalization generative adversarial network (VN-GAN) and classification network. For VN-GAN, the model is designed to learn the mapping from diverse-view distribution to normalized-view distribution. In detail, it is implemented by graph convolution, where the generator predicts the transformation angles for view normalization and discriminator classifies the real input samples from the generated ones. For classification network, view-normalized data is processed to predict the action class. Without the interference of view variances, classification network can extract more discriminative feature of action. Furthermore, by combining the joint and bone modalities, the proposed method reaches the state-of-the-art performance on NTU RGB+D and NTU-120 RGB+D datasets. Especially in NTU-120 RGB+D, the accuracy is improved by 3.2% and 2.3% under cross-subject and cross-set criteria, respectively. Qingzhe Pan, Zhifu Zhao, Xuemei Xie, Jianan Li 0003, Guangming Shi |
ACM Multimedia | 6 |
| 2021 | No-Reference Video Quality Assessment with Heterogeneous Knowledge EnsembleabstractBlind assessment of video quality is still challenging even in this deep learning era. The limited number of samples in existing databases is insufficient to learn a good feature extractor for video quality assessment (VQA), while manually labeling a larger database with subjective perception is very labor-intensive and time-consuming. To relieve such difficulty, we first collect 3589 high-quality video clips as the reference and build a large VQA dataset. The dataset contains more than 300K samples degraded by various distortion types due to compression and transmission error, and provides weak labels for each distorted sample with several full-reference VQA algorithms. To learn effective representation from the weakly labeled data, we alleviate the bias of single weak label (i.e., single knowledge) via learning from multiple heterogeneous knowledge. To this end, we propose a novel no-reference VQA (NR-VQA) method with HEterogeneous Knowledge Ensemble (HEKE). Comparing to learning from single knowledge, HEKE can theoretically reach a lower infimum, and learn richer representation due to the heterogeneity. Extensive experimental results show that the proposed HEKE outperforms existing NR-VQA methods, and achieves the state-of-the-art performance. The source code will be available at https://github.com/Sissuire/BVQA-HEKE. Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 5 |
| 2021 | Image Quality Caption with Attentive and Recurrent Semantic Attractor NetworkabstractIn this paper, a novel quality caption model is inventively developed to assess the image quality with hierarchical semantics. Existing image quality assessment (IQA) methods usually represent image quality with a quantitative value, resulting in inconsistency with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where hierarchical semantics are extracted. The mediocre quality value fails to reflect degradations on hierarchical semantics. Therefore, a new IQA framework is proposed to describe the quality for needs-oriented cognition. A novel quality caption procedure is firstly introduced, in which the quality is represented as patterns of activation distributed across the diverse degradations on hierarchical semantics. Then, an attentive and recurrent semantic attractor network (ARSANet) is designed to activate the distributed patterns for image quality description. Experiments demonstrate that our method achieves superior performance and is highly compliant with human cognition. Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 5 |
| 2021 | Uncertainty-Driven Loss for Single Image Super-ResolutionabstractIn low-level vision such as single image super-resolution (SISR), traditional MSE or L1 loss function treats every pixel equally with the assumption that the importance of all pixels is the same. However, it has been long recognized that texture and edge areas carry more important visual information than smooth areas in photographic images. How to achieve such spatial adaptation in a principled manner has been an open problem in both traditional model-based and modern learning-based approaches toward SISR. In this paper, we propose a new adaptive weighted loss for SISR to train deep networks focusing on challenging situations such as textured and edge pixels with high uncertainty. Specifically, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis into SISR solutions so the targeted pixels in a high-resolution image (mean) and their corresponding uncertainty (variance) can be learned simultaneously. Moreover, uncertainty estimation allows us to leverage conventional wisdom such as sparsity prior for regularizing SISR solutions. Ultimately, pixels with large certainty (e.g., texture and edge pixels) will be prioritized for SISR according to their importance to visual quality. For the first time, we demonstrate that such uncertainty-driven loss can achieve better results than MSE or L1 loss for a wide range of network architectures. Experimental results on three popular SISR networks show that our proposed uncertainty-driven loss has achieved better PSNR performance than traditional loss functions without any increased computation during testing. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UDL-SR.htm Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
NeurIPS | 5 |
| 2021 | Deep Maximum a Posterior Estimator for Video Denoising
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
Int. J. Comput. Vis. | 6 |
| 2021 | A novel transferability attention neural network model for EEG emotion recognition
Yang Li 0019, Boxun Fu, Fu Li 0002, Guangming Shi, Wenming Zheng |
Neurocomputing | 4 |
| 2021 | Knowledge embedded GCN for skeleton-based two-person interaction recognition
Jianan Li 0003, Xuemei Xie, Qingzhe Pan, Zhifu Zhao, Guangming Shi |
Neurocomputing | 6 |
| 2021 | Knowledge-guided semantic computing network
Guangming Shi, Dahua Gao, Jie Lin 0008, Xuemei Xie, Danhua Liu |
Neurocomputing | 1 |
| 2021 | Domain-aware Stacked AutoEncoders for zero-shot learning
Jianqiang Song, Guangming Shi, Xuemei Xie, Qingtao Wu, Mingchuan Zhang |
Neurocomputing | 2 |
| 2021 | Multi-Scale and Single-Scale Fully Convolutional Networks for Sound Event Detection
Yingbin Wang, Guanghui Zhao 0003, Guangming Shi |
Neurocomputing | 4 |
| 2021 | Toward blind joint demosaicing and denoising of raw color filter array data
Weisheng Dong, Guangming Shi, Zhonglong Zheng, Xin Li 0005 |
Neurocomputing | 4 |
| 2021 | Blind image quality prediction with hierarchical feature aggregation
Jinjian Wu, Wen Yang 0008, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin |
Inf. Sci. | 5 |
| 2021 | Conditional generative adversarial network for EEG-based emotion fine-grained estimation and visualization
Boxun Fu, Fu Li 0002, Yang Li 0019, Guangming Shi |
J. Vis. Commun. Image Represent. | 6 |
| 2021 | Attention-shift based deep neural network for fine-grained visual categorization
Guangming Shi |
Pattern Recognit. | 3 |
| 2021 | Robust subspace clustering network with dual-domain regularization
Guangming Shi, Xin Li 0005, Weisheng Dong, Jinjian Wu |
Pattern Recognit. Lett. | 3 |
| 2021 | Hybrid sparsity learning for image restoration: An iterative and trainable approach
Weisheng Dong, Guangming Shi, Shaoyuan Cheng, Xin Li 0005 |
Signal Process. | 4 |
| 2021 | SPB-Net: A Deep Network for SAR Imaging and Despeckling With Downsampled DataabstractSynthetic aperture radar (SAR) typically faces both large-scale data and speckle noise problems, which, respectively, induce enormous strains on transmission and storage and interfere with the analysis and interpretation of SAR images. To tackle these, we present a real-time processing deep network, called SPB-Net, to implement imaging and speckle suppression simultaneously. First, a novel imaging-despeckling observation model with the nonlogarithmic additive speckle noise is established. Subsequently, guided by the statistical properties of noise and sparse and detail-preserved requirements in SAR imaging and despeckling, we formulate an$L_{2}$along with two$L_{1}$regularizations as the fidelity, sparse, and image detail-preserved constraints, respectively. Convolution layers are employed to improve the feature representation capability in the latter$L_{1}$term as well. Based on this, we construct a corresponding convex optimization problem. Then, the complex-valued split Bregman method, focusing on the complex-variable convex problem, is unfolded into a parameter-learnable and architecture-fixed SPB-Net to solve the proposed problem effectively and efficiently. Experimental results with the downsampled Radarsat-1 raw data demonstrate the validity in imaging and speckle suppression and the real-time processing capability of the proposed SPB-Net. Guanghui Zhao 0003, Yingbin Wang, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Model-Guided Deep Hyperspectral Image Super-ResolutionabstractThe trade-off between spatial and spectral resolution is one of the fundamental issues in hyperspectral images (HSI). Given the challenges of directly acquiring high-resolution hyperspectral images (HR-HSI), a compromised solution is to fuse a pair of images: one has high-resolution (HR) in the spatial domain but low-resolution (LR) in spectral-domain and the other vice versa. Model-based image fusion methods including pan-sharpening aim at reconstructing HR-HSI by solving manually designed objective functions. However, such hand-crafted prior often leads to inevitable performance degradation due to a lack of end-to-end optimization. Although several deep learning-based methods have been proposed for hyperspectral pan-sharpening, HR-HSI related domain knowledge has not been fully exploited, leaving room for further improvement. In this paper, we propose an iterative Hyperspectral Image Super-Resolution (HSISR) algorithm based on a deep HSI denoiser to leverage both domain knowledge likelihood and deep image prior. By taking the observation matrix of HSI into account during the end-to-end optimization, we show how to unfold an iterative HSISR algorithm into a novel model-guided deep convolutional network (MoG-DCN). The representation of the observation matrix by subnetworks also allows the unfolded deep HSISR network to work with different HSI situations, which enhances the flexibility of MoG-DCN. Extensive experimental results are reported to demonstrate that the proposed MoG-DCN outperforms several leading HSISR methods in terms of both implementation cost and visual quality. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/MoG-DCN.htm. Weisheng Dong, Chen Zhou 0005, Jinjian Wu, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 5 |
| 2021 | Blind Image Quality Assessment With Active InferenceabstractBlind image quality assessment (BIQA) is a useful but challenging task. It is a promising idea to design BIQA methods by mimicking the working mechanism of human visual system (HVS). The internal generative mechanism (IGM) indicates that the HVS actively infers the primary content (i.e., meaningful information) of an image for better understanding. Inspired by that, this paper presents a novel BIQA metric by mimicking the active inference process of IGM. Firstly, an active inference module based on the generative adversarial network (GAN) is established to predict the primary content, in which the semantic similarity and the structural dissimilarity (i.e., semantic consistency and structural completeness) are both considered during the optimization. Then, the image quality is measured on the basis of its primary content. Generally, the image quality is highly related to three aspects, i.e., the scene information (content-dependency), the distortion type (distortion-dependency), and the content degradation (degradation-dependency). According to the correlation between the distorted image and its primary content, the three aspects are analyzed and calculated respectively with a multi-stream convolutional neural network (CNN) based quality evaluator. As a result, with the help of the primary content obtained from the active inference and the comprehensive quality degradation measurement from the multi-stream CNN, our method achieves competitive performance on five popular IQA databases. Especially in cross-database evaluations, our method achieves significant improvements. Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie, Guangming Shi, Weisi Lin |
IEEE Trans. Image Process. | 6 |
| 2021 | Quality Index for View Synthesis by Measuring Instance Degradation and Global AppearanceabstractVirtual view synthesis plays a vital role in the application of multi-view and free-viewpoint videos. Depth-image-based rendering (DIBR) is the most commonly used approach in view synthesis, and many DIBR algorithms have been proposed. However, how to evaluate the quality of DIBR-synthesized images and benchmark the DIBR algorithms are still very challenging, which may hinder the further development of the view synthesis technique. Hence, an effective quality metric for evaluating the distortions in view synthesis is urgently needed. With this motivation, this paper presents a quality index for view synthesis by simultaneously measuring local Instance DEgradation and global Appearance (IDEA). Due to the imperfection of rendering algorithms, local geometric distortions are easily introduced around instance contours, causing instance degradation, which is the dominant distortion in synthesized views. In this work, image instances are first detected and local instance degradation is measured based on discrete orthogonal moments. Meantime, we propose to measure the global appearance of synthesized images based on the superpixel representation. By integrating both local and global aspects of the distortions, a more accurate quality model is built for view synthesis. Extensive experiments and comparisons have demonstrated the superiority of the proposed method in evaluating the quality of DIBR-synthesized images and benchmarking the performance of view synthesis algorithms. Leida Li, Yu Zhou 0009, Jinjian Wu, Fu Li 0002, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2021 | Probabilistic Undirected Graph Based Denoising Method for Dynamic Vision SensorabstractDynamic Vision Sensor (DVS) is a new type of neuromorphic event-based sensor, which has an innate advantage in capturing fast-moving objects. Due to the interference of DVS hardware itself and many external factors, noise is unavoidable in the output of DVS. Different from frame/image with structural data, the output of DVS is in the form of address-event representation (AER), which means that the traditional denoising methods cannot be used for the output (i.e., event stream) of the DVS. In this paper, we propose a novel event stream denoising method based on probabilistic undirected graph model (PUGM). The motion of objects always shows a certain regularity/trajectory in space and time, which reflects the spatio-temporal correlation between effective events in the stream. Meanwhile, the event stream of DVS is composed by the effective events and random noise. Thus, a probabilistic undirected graph model is constructed to describe such priori knowledge (i.e., spatio-temporal correlation). The undirected graph model is factorized into the product of the cliques energy function, and the energy function is defined to obtain the complete expression of the joint probability distribution. Better denoising effect means a higher probability (lower energy), which means the denoising problem can be transfered into energy optimization problem. Thus, the iterated conditional modes (ICM) algorithm is used to optimize the model to remove the noise. Experimental results on denoising show that the proposed algorithm can effectively remove noise events. Moreover, with the preprocessing of the proposed algorithm, the recognition accuracy on AER data can be remarkably promoted. Jinjian Wu, Chuanwei Ma, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2020 | Spatial-Temporal Gaussian Scale Mixture Modeling for Foreground EstimationabstractSubtracting the backgrounds from the video frames is an important step for many video analysis applications. Assuming that the backgrounds are low-rank and the foregrounds are sparse, the robust principle component analysis (RPCA)-based methods have shown promising results. However, the RPCA-based methods suffered from the scale issue, i.e., the ℓ1-sparsity regularizer fails to model the varying sparsity of the moving objects. While several efforts have been made to address this issue with advanced sparse models, previous methods cannot fully exploit the spatial-temporal correlations among the foregrounds. In this paper, we proposed a novel spatial-temporal Gaussian scale mixture (STGSM) model for foreground estimation. In the proposed STGSM model, a temporal consistent constraint is imposed over the estimated foregrounds through nonzero-means Gaussian models. Specifically, the estimates of the foregrounds obtained in the previous frame are used as the prior for these of the current frame, and nonzero means Gaussian scale mixture models (GSM) are developed. To better characterize the temporal correlations, the optical flow has been used to model the correspondences between foreground pixels in adjacent frames. The spatial correlations have also been exploited by considering that local correlated pixels should be characterized by the same STGSM model, leading to further performance improvements. Experimental results on real video datasets show that the proposed method performs comparably or even better than current state-of-the-art background subtraction methods. Qian Ning, Weisheng Dong, Jinjian Wu, Jie Lin 0008, Guangming Shi |
AAAI | 6 |
| 2020 | MetaIQA: Deep Meta-Learning for No-Reference Image Quality AssessmentabstractRecently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training DCNNs heavily relies on massive annotated data. Unfortunately, IQA is a typical small sample problem. Therefore, most of the existing DCNN-based IQA metrics operate based on pre-trained networks. However, these pre-trained networks are not designed for IQA task, leading to generalization problem when evaluating different types of distortions. With this motivation, this paper presents a no-reference IQA metric based on deep meta-learning. The underlying idea is to learn the meta-knowledge shared by human when evaluating the quality of images with various distortions, which can then be adapted to unknown distortions easily. Specifically, we first collect a number of NR-IQA tasks for different distortions. Then meta-learning is adopted to learn the prior knowledge shared by diversified distortions. Finally, the quality prior model is fine-tuned on a target NR-IQA task for quickly obtaining the quality model. Extensive experiments demonstrate that the proposed metric outperforms the state-of-the-arts by a large margin. Furthermore, the meta-model learned from synthetic distortions can also be easily generalized to authentic distortions, which is highly desired in real-world applications of IQA metrics. Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
CVPR | 5 |
| 2020 | Denoising of Event-Based Sensors with Spatial-Temporal CorrelationabstractAs a novel asynchronous-driven cameras, event-based sensors are with high sensitivity, fast speed, low power consumption and low data volume, but with abundant noise. Since the output of event-based sensors is in the form of address-event-representation (AER), the traditional frame-based denoising method cannot be used. In this paper, we introduce a novel event stream denoising method for such sensors. Effective events tend to show temporal and spatial regularity, while noise events show a kind of randomness. Thus, we build a probabilistic undirected graph model to describe this difference, with which the denoising problem is converted to a probability maximization problem. Then, the model is decomposed into the product of the energy function on the maximum cliques, and the iterated condition model (ICM) is used for energy minimization to obtain the denoised event stream. Experiments show that our method can effectively remove noise events directly from the event stream and significantly improve event recognition rate. Jinjian Wu, Chuanwei Ma, Xiaojie Yu, Guangming Shi |
ICASSP | 4 |
| 2020 | Channel-Grouping Based Patch Swap For Arbitrary Style TransferabstractThe basic principle of the patch-matching based style transfer is to substitute the patches of the content image feature maps by the closest patches from the style image feature maps. Since the finite features harvested from one single aesthetic style image are inadequate to represent the rich textures of the content natural image, existing techniques treat the full-channel style feature patches as simple signal tensors and create new style feature patches via signal-level fusion. In this paper, we propose a channel-grouping based patch swap technique to group the style feature maps into surface and texture channels, and the new features are created by the combination of these two groups, which can be regarded as a semantic-level fusion of the raw style features. Experimental results demonstrate that the proposed method outperforms the existing techniques in providing more style-consistent textures while keeping the content fidelity. Fu Li 0002, Chunbo Zou, Guangming Shi |
ICIP | 5 |
| 2020 | Beyond Network Pruning: a Joint Search-and-Training ApproachabstractNetwork pruning has been proposed as a remedy for alleviating the over-parameterization problem of deep neural networks. However, its value has been recently challenged especially from the perspective of neural architecture search (NAS). We challenge the conventional wisdom of pruning-after-training by proposing a joint search-and-training approach that directly learns a compact network from the scratch. By treating pruning as a search strategy, we present two new insights in this paper: 1) it is possible to expand the search space of networking pruning by associating each filter with a learnable weight; 2) joint search-and-training can be conducted iteratively to maximize the learning efficiency. More specifically, we propose a coarse-to-fine tuning strategy to iteratively sample and update compact sub-network to approximate the target network. The weights associated with network filters will be accordingly updated by joint search-and-training to reflect learned knowledge in NAS space. Moreover, we introduce strategies of random perturbation (inspired by Monte Carlo) and flexible thresholding (inspired by Reinforcement Learning) to adjust the weight and size of each layer. Extensive experiments on ResNet and VGGNet demonstrate the superior performance of our proposed method on popular datasets including CIFAR10, CIFAR100 and ImageNet. Xiaotong Lu, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 5 |
| 2020 | RIRNet: Recurrent-In-Recurrent Network for Video Quality AssessmentabstractVideo quality assessment (VQA), which is capable of automatically predicting the perceptual quality of source videos especially when reference information is not available, has become a major concern for video service providers due to the growing demand for video quality of experience (QoE) by end users. While significant advances have been achieved from the recent deep learning techniques, they often lead to misleading results in VQA tasks given their limitations on describing 3D spatio-temporal regularities using only fixed temporal frequency. Partially inspired by psychophysical and vision science studies revealing the speed tuning property of neurons in visual cortex when performing motion perception (i.e., sensitive to different temporal frequencies), we propose a novel no-reference (NR) VQA framework named Recurrent-In-Recurrent Network (RIRNet) to incorporate this characteristic to prompt an accurate representation of motion perception in VQA task. By fusing motion information derived from different temporal frequencies in a more efficient way, the resulting temporal modeling scheme is formulated to quantify the temporal motion effect via a hierarchical distortion description. It is found that the proposed framework is in closer agreement with quality perception of the distorted videos since it integrates concepts from motion perception in human visual system (HVS), which is manifested in the designed network structure composed of low- and high- level processing. A holistic validation of our methods on four challenging video quality databases demonstrates the superior performances over the state-of-the-art methods. Pengfei Chen 0003, Leida Li, Lei Ma 0003, Jinjian Wu, Guangming Shi |
ACM Multimedia | 5 |
| 2020 | Multi-level Prediction with Graphical Model for Human Pose Estimation
Xuemei Xie, Lihua Ma, Jiang Du 0011, Guangming Shi |
PRCV (2) | 5 |
| 2020 | A Multi-level Equilibrium Clustering Approach for Unsupervised Person Re-identification
Fangyu Wang, Zhenyu Wang 0008, Xuemei Xie, Guangming Shi |
PRCV (3) | 4 |
| 2020 | Network pruning using sparse learning and genetic algorithm
Zhenyu Wang 0008, Fu Li 0002, Guangming Shi, Xuemei Xie, Fangyu Wang |
Neurocomputing | 3 |
| 2020 | No-reference quality index of depth images based on statistics of edge profiles for view synthesis
Leida Li, Jinjian Wu, Shiqi Wang 0001, Guangming Shi |
Inf. Sci. | 5 |
| 2020 | SGM-Net: Skeleton-guided multimodal network for action recognition
Jianan Li 0003, Xuemei Xie, Qingzhe Pan, Zhifu Zhao, Guangming Shi |
Pattern Recognit. | 6 |
| 2020 | On Estimation of Time-Varying Variances of Source and Noise for Sensor Array ProcessingabstractEstimation of time-varying variances of signals for beamforming in sensor arrays is a challenging problem. Based on the assumption that the array manifold vector and the noise pseudo-coherence matrix are known a priori or are well estimated, we present in this paper two estimators for estimating the time-varying variances of the source signal of interest and the noise. These two estimators are then extended to deal with the following situations: 1) there are multiple candidates of the noise pseudo-coherence matrix or the noise pseudo-coherence matrix is a linear combination of some base pseudo-coherence matrices, and 2) the estimation variance is large and smoothing is needed. Simulations for speech enhancement applications are performed and the results show that the proposed estimators can well track the time-varying variances of both the speech and noise signals. It is also demonstrated that the optimal beamformer using the variance parameters estimated with the presented estimators outperforms the widely used traditional optimal beamformers in terms of improvement in both the signal-to-noise ratio (SNR) and the log-spectral distortion (LSD). Chao Pan 0001, Jingdong Chen, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | End-to-End Blind Image Quality Prediction With Cascaded Deep Neural NetworkabstractThe deep convolutional neural network (CNN) has achieved great success in image recognition. Many image quality assessment (IQA) methods directly use recognition-oriented CNN for quality prediction. However, the properties of IQA task is different from image recognition task. Image recognition should be sensitive to visual content and robust to distortion, while IQA should be sensitive to both distortion and visual content. In this paper, an IQA-oriented CNN method is developed for blind IQA (BIQA), which can efficiently represent the quality degradation. CNN is large-data driven, while the sizes of existing IQA databases are too small for CNN optimization. Thus, a large IQA dataset is firstly established, which includes more than one million distorted images (each image is assigned with a quality score as its substitute of Mean Opinion Score (MOS), abbreviated as pseudo-MOS). Next, inspired by the hierarchical perception mechanism (from local structure to global semantics) in human visual system, a novel IQA-orientated CNN method is designed, in which the hierarchical degradation is considered. Finally, by jointly optimizing the multilevel feature extraction, hierarchical degradation concatenation (HDC) and quality prediction in an end-to-end framework, the Cascaded CNN with HDC (named as CaHDC) is introduced. Experiments on the benchmark IQA databases demonstrate the superiority of CaHDC compared with existing BIQA methods. Meanwhile, the CaHDC (with about 0.73M parameters) is lightweight comparing to other CNN-based BIQA models, which can be easily realized in the microprocessing system. The dataset and source code of the proposed method are available at https://web.xidian.edu.cn/wjj/paper.html. Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi, Weisi Lin |
IEEE Trans. Image Process. | 5 |
| 2019 | Zero-Shot Learning Using Stacked Autoencoder with Manifold RegularizationsabstractZero-shot learning (ZSL), which focuses on transferring the knowledge from the seen classes to unseen ones, has attracted more and more attention in the computer vision community. Exploring the relationships among the spaces of visual representation, semantic description and label information is a key to the success of ZSL. In this paper, we propose a novel approach by using a two-layer Stacked AutoEncoder (StAE) with manifold regularizations to construct the tight relations of different spaces, where the first-layer encoder aims to project a visual feature vector into the semantic space, and the second-layer encoder connects the semantic description of a sample with its label directly. Meanwhile, the decoders seek to reconstruct the visual representation from label information and semantic description successively. Besides, two manifold regularizers are integrated in the stacked autoencoder, which captures the manifold structures residing in the different spaces effectively. Compared with the previous related works, the proposed approach is a more general framework and has stronger transfer ability from seen classes to unseen classes. Extensive experiments on the benchmark datasets clearly demonstrate that our StAE performs significantly better than the state-of-the-arts. Jianqiang Song, Guangming Shi, Xuemei Xie, Dahua Gao |
ICIP | 2 |
| 2019 | End-to-End Blind Image Quality Assessment with Cascaded Deep FeaturesabstractThe convolutional neural network (CNN) has achieved great success in many visual tasks. However, it has limited progress on image quality assessment (IQA) due to the lacking of IQA-oriented CNN framework which can efficiently represent the hierarchical quality degradation. In this paper, inspired by the hierarchical perception mechanism (from local structure to global semantics) in the human visual system, we design an end-to-end cascaded CNN framework for blind IQA (BIQA), in which multilevel features are extracted and concatenated to represent the hierarchical quality degradation. By jointly optimizing the feature extraction, hierarchical degradation integration, and quality prediction in an end-to-end manner, the novel cascaded CNN with hierarchical feature integration (CaHFI) for BIQA is designed. Experimental results on five benchmark IQA databases demonstrate that the proposed CaHFI achieves the state-of-the-art. And experiments on cross-database evaluation further prove the high generalization ability of the proposed CaHFI. Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi |
ICME | 5 |
| 2019 | Co-Representation Network for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning is a significant topic but faced with bias problem, which leads to unseen classes being easily misclassified into seen classes. Hence we propose a embedding model called co-representation network to learn a more uniform visual embedding space that effectively alleviates the bias problem and helps with classification. We mathematically analyze our model and find it learns a projection with high local linearity, which is proved to cause less bias problem. The network consists of a cooperation module for representation and a relation module for classification, it is simple in structure and can be easily trained in an end-to-end manner. Experiments show that our method outperforms existing generalized zero-shot learning methods on several benchmark datasets. Guangming Shi |
ICML | 2 |
| 2019 | A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task LearningabstractDetecting scene text of arbitrary shapes has been a challenging task over the past years. In this paper, we propose a novel segmentation-based text detector, namely SAST, which employs a context attended multi-task learning framework based on a Fully Convolutional Network (FCN) to learn various geometric properties for the reconstruction of polygonal representation of text regions. Taking sequential characteristics of text into consideration, a Context Attention Block is introduced to capture long-range dependencies of pixel information to obtain a more reliable segmentation. In post-processing, a Point-to-Quad assignment method is proposed to cluster pixels into text instances by integrating both high-level object knowledge and low-level pixel information in a single shot. Moreover, the polygonal representation of arbitrarily-shaped text can be extracted with the proposed geometric properties much more effectively. Experiments on several benchmarks, including ICDAR2015, ICDAR2017-MLT, SCUT-CTW1500, and Total-Text, demonstrate that SAST achieves better or comparable performance in terms of accuracy. Furthermore, the proposed algorithm runs at 27.63 FPS on SCUT-CTW1500 with a Hmean of 81.0% on a single NVIDIA Titan Xp graphics card, surpassing most of the existing segmentation-based methods. Chengquan Zhang, Fei Qi 0001, Zuming Huang, Mengyi En, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi |
ACM Multimedia | 9 |
| 2019 | Channel Feature Enhanced Detector for Small Ball Detection
Shambel Ferede, Xuemei Xie, Jiang Du 0011, Guangming Shi |
PRCV (1) | 5 |
| 2019 | CG Animation Creator: Auto-rendering of Motion Stick Figure Based on Conditional Adversarial Learning
Jie Lin 0008, Guangming Shi, Danhua Liu |
PRCV (3) | 3 |
| 2019 | A Real-Time Rock-Paper-Scissor Hand Gesture Recognition System Based on FlowNet and Event Camera
Xuemei Xie, Jinjian Wu, Guangming Shi |
PRCV (1) | 5 |
| 2019 | Facial Attention based Convolutional Neural Network for 2D+3D Facial Expression RecognitionabstractDiscriminative facial parts are essential for facial expression recognition (FER) tasks because of small inter-class differences and large intra-class variations in expression images. Existing methods localize discriminative regions with the aid of extra facial landmarks, such as action units (AU). However, it consumes a lot of manpower in manually labeling. To address this problem, in this paper, we propose an advanced facial attention based convolutional neural network (FA-CNN) for 2D+3D FER. The main contribution of FA-CNN is the facial attention mechanism, which enables the network to localize the discriminative regions automatically from multi-modality expression images without dense landmark annotations. Experimental results conducted on BU-3DFE demonstrate that FA-CNN achieves state-of-the-art performance comparing with the existing 2D+3D FER techniques, and the discriminative facial parts estimated by the facial attention mechanism are highly interpretable and consistent with human perception. Fu Li 0002, Chunbo Zou, Guangming Shi |
VCIP | 6 |
| 2019 | Survey of visual just noticeable difference estimation
Jinjian Wu, Guangming Shi, Weisi Lin |
Frontiers Comput. Sci. | 2 |
| 2019 | Fully convolutional measurement network for compressive sensing image reconstruction
Jiang Du 0011, Xuemei Xie, Chenye Wang, Guangming Shi |
Neurocomputing | 4 |
| 2019 | SISRSet: Single image super-resolution subjective evaluation test and objective quality assessment
Guangming Shi, Wenfei Wan, Jinjian Wu, Xuemei Xie, Weisheng Dong, Hong Ren Wu |
Neurocomputing | 1 |
| 2019 | Visualizing and understanding of learned compressive sensing with residual network
Zhifu Zhao, Xuemei Xie, Chenye Wang, Wan Liu 0001, Guangming Shi, Jiang Du 0011 |
Neurocomputing | 5 |
| 2019 | No-reference image quality assessment with visual pattern degradation
Jinjian Wu, Man Zhang 0007, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin |
Inf. Sci. | 5 |
| 2019 | Blind image quality assessment with semantic information
Weiping Ji, Jinjian Wu, Guangming Shi, Wenfei Wan, Xuemei Xie |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | A multiscale dilated dense convolutional network for saliency prediction with instance-level attention competition
Hao Li 0051, Fei Qi 0001, Guangming Shi, Chunhuan Lin |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Blind image quality assessment with hierarchy: Degradation from local structure to deep semantics
Jinjian Wu, Jichen Zeng, Weisheng Dong, Guangming Shi, Weisi Lin |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Denoising Prior Driven Deep Neural Network for Image RestorationabstractDeep neural networks (DNNs) have shown very promising results for various image restoration (IR) tasks. However, the design of network architectures remains a major challenging for achieving further improvements. While most existing DNN-based methods solve the IR problems by directly mapping low quality images to desirable high-quality images, the observation models characterizing the image degradation processes have been largely ignored. In this paper, we first propose a denoising-based IR algorithm, whose iterative steps can be computed efficiently. Then, the iterative process is unfolded into a deep neural network, which is composed of multiple denoisers modules interleaved with back-projection (BP) modules that ensure the observation consistencies. A convolutional neural network (CNN) based denoiser that can exploit the multi-scale redundancies of natural images is proposed. As such, the proposed network not only exploits the powerful denoising ability of DNNs, but also leverages the prior of the observation model. Through end-to-end training, both the denoisers and the BP modules can be jointly optimized. Experimental results on several IR tasks, e.g., image denoisig, super-resolution and deblurring show that the proposed method can lead to very competitive and often state-of-the-art results on several IR tasks, including image denoising, deblurring, and super-resolution. Weisheng Dong, Wotao Yin, Guangming Shi, Xiaotong Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | High-Speed Hyperspectral Video Acquisition By Combining Nyquist and Compressive SamplingabstractWe propose a novel hybrid imaging system to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. The proposed system consists of two branches: one branch performs Nyquist sampling in the temporal dimension while integrating the whole spectrum, resulting in a high-frame-rate panchromatic video; the other branch performs compressive sampling in the spectral dimension with longer exposures, resulting in a low-frame-rate hyperspectral video. Owing to the high light throughput and complementary sampling, these two branches jointly provide reliable measurements for recovering the underlying HSHS video. Moreover, the panchromatic video can be used to learn an over-complete 3D dictionary to represent each band-wise video sparsely, thanks to the inherent structural similarity in the spectral dimension. Based on the joint measurements and the self-adaptive dictionary, we further propose a simultaneous spectral sparse (3S) model to reinforce the structural similarity across different bands and develop an efficient computational reconstruction algorithm to recover the HSHS video. Both simulation and hardware experiments validate the effectiveness of the proposed approach. To the best of our knowledge, this is the first time that hyperspectral videos can be acquired at a frame rate up to 100fps with commodity optical elements and under ordinary indoor illumination. Lizhi Wang 0001, Zhiwei Xiong, Hua Huang 0001, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Multi-layer discriminative dictionary learning with locality constraint for image classification
Jianqiang Song, Xuemei Xie, Guangming Shi, Weisheng Dong |
Pattern Recognit. | 3 |
| 2019 | Depth acquisition with the combination of structured light and deep learning stereo matching
Fu Li 0002, Quanlu Li, Guangming Shi |
Signal Process. Image Commun. | 5 |
| 2019 | ROI-CSNet: Compressive sensing network for ROI-aware image recovery
Zhifu Zhao, Xuemei Xie, Chenye Wang, Siying Mao, Wan Liu 0001, Guangming Shi |
Signal Process. Image Commun. | 6 |
| 2019 | Development of a Human-Robot Hybrid Intelligent System Based on Brain Teleoperation and Deep Learning SLAMabstractTo achieve the better navigation performance of a mobile robot in the unknown environments, a novel human-robot hybrid system incorporating a motor-imagery (MI)-based brain teleoperation control is presented in this paper, where a deep-learning-based active perception is developed in the simultaneous localization and mapping (SLAM) framework. Using the deep-learning-based object recognition in the red-green-blue-depth (RGB-D) data acquisition process, the designed SLAM approach can select the valid feature points effectively, and the speed of displacement tracking can be improved by combining the oriented FAST and rotated BRIEF (ORB) SLAM algorithm with the optical flow method. The global trajectory map can also be mended using graph-based nonlinear error optimization. In addition, to build the connection between human intentions and the robot control commands flexibly in the developed mobile robot, a common spatial pattern (CSP)-based support vector machine (SVM) classification algorithm is proposed so that the control commands can be obtained directly from the human electroencephalograph (EEG) signals, which are preanalyzed and classified using the phenomena of event-related synchronization/desynchronization (ERS/ERD). Experiments involving several operators have verified the effectiveness of the proposed framework in the actual unstructured environments. Zhijun Li 0001, Yiliang Liu, Guangming Shi |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2019 | On the Design of Target Beampatterns for Differential Microphone ArraysabstractDifferential microphone arrays (DMAs) have many interesting properties and have been widely used in acoustic, audio, and speech applications. A critical part of a DMA is the differential beamformer, which is generally designed in two important steps: 1) specifying a target beampattern based on what differential sound pressure field the DMA is expected to respond to and 2) designing the differential beamforming filter so that the resulting beampattern matches the target one. Most efforts in the study of DMAs so far have focused on the second step while choosing one of the limited patterns available in the literature as the target beampattern. Since it governs how the array performs, how to design the target beampattern is an important problem, which this paper addresses. The major contributions of this paper consists of the following four aspects. First, a positive superposition theorem is presented, which shows that the linear combination of effective beampatterns with non-negative coefficients is always an effective beampattern. Second, we propose a general approach to the design of target DMA beampatterns based on the positive superposition theorem. Third, an overview of the classical target beampatterns is provided and discussion is made on how to form effective base patterns. Fourth, we show that the smallest first null of a DMA is π(2N) with N being the DMA order, which provides the rule of setting nulls in practice. Finally, with examples, we show that with the use of the alternating-direction-method-of-multipliers algorithm, the proposed approach is able to generate useful DMA target beampatterns. Chao Pan 0001, Jingdong Chen, Jacob Benesty, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Predicting Human Saccadic Scanpaths Based on Iterative Representation LearningabstractVisual attention is a dynamic process of scene exploration and information acquisition. However, existing research on attention modeling has concentrated on estimating static salient locations. In contrast, dynamic attributes presented by saccade have not been well explored in previous attention models. In this paper, we address the problem of saccadic scanpath prediction by introducing an iterative representation learning framework. Within the framework, saccade can be interpreted as an iterative process of predicting one fixation according to the current representation and updating the representation based on the gaze shift. In the predicting phase, we propose a Bayesian definition of saccade to combine the influence of perceptual residual and spatial location on the selection of fixations. In implementation, we compute the representation error of an autoencoder-based network to measure perceptual residuals of each area. Simultaneously, we integrate saccade amplitude and center-weighted mechanism to model the influence of spatial location. Based on estimating the influence of two parts, the final fixation is defined as the point with the largest posterior probability of gaze shift. In the updating phase, we update the representation pattern for the subsequent calculation by retraining the network with samples extracted around the current fixation. In the experiments, the proposed model can replicate the fundamental properties of psychophysics in visual search. In addition, it can achieve superior performance on several benchmark eye-tracking data sets. Chen Xia, Junwei Han 0001, Fei Qi 0001, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2019 | Quality Assessment for Video With Degradation Along Salient TrajectoriesabstractWith the rapid growth of digital video through the Internet, a reliable objective video-quality assessment (VQA) algorithm is in great demand for video management. Motion information plays a dominant role for video perception, and the human visual system (HVS) is able to track moving objects effectively with eye movement. Moreover, the middle temporal area of the brain is selective for moving objects with particular velocities. In other words, visual contents that are along the motion trajectories will automatically attract our attention for dedicated processing. Inspired by the motion-related process in the HVS, we suggest analyzing the degradation along attended motion trajectories for VQA. The characteristic of motion velocity along each trajectory is analyzed for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial-quality degradation from each frame, a novel full-reference assessor along salient trajectories (FAST) for VQA (which combines the spatial, temporal, and joint spatial-temporal quality degradations) is introduced. Experimental results on five publicly available VQA databases demonstrate that the proposed FAST VQA model performs consistently with the subjective perception. The source code of the proposed method is available at http://web.xidian.edu.cn/wjj/paper.html. Jinjian Wu, Yongxu Liu 0001, Weisheng Dong, Guangming Shi, Weisi Lin |
IEEE Trans. Multim. | 4 |
| 2018 | Super-Resolution Quality Assessment: Subjective Evaluation Database and Quality Index Based on Perceptual Structure MeasurementabstractWith the outstanding performance of deep learning based single image super-resolution (SISR) methods, the traditional SISR evaluation metrics (e.g., PSNR and SSIM, which measure the per-pixel differences and simple structure similarities respectively) are facing great challenges. When assessing SISR algorithms, they generally are hardly consistent with the human visual system (HVS). According to the psychological studies, the HVS presents different sensitivities to the plain, edge and texture regions, which are difficult to be accurately identified and measured with the existing quality indexes, especially for SR images. To deal with this problem, we firstly build a SISR subjective assessment database including several major deep learning based SR methods. Then we propose a more accurate perception structure measurement and use their similarity comparisons to evaluate the SR algorithms. Experimental results on the databases demonstrate that the proposed method performs well consistent with the human visual perception. Wenfei Wan, Jinjian Wu, Guangming Shi, Weisheng Dong |
ICME | 3 |
| 2018 | Full Image Recover for Block-Based Compressive SensingabstractCompressive sensing (CS) theory is able to acquire measurements of a scene at sub-Nyquist rate and recover the scene image from these under-sampled measurements. Recent years, CS has been improved greatly for the application of deep learning technology. In conventional methods, block-based mechanism is used to recover images from measurements, which usually causes block effect in reconstructed images. In this paper, we propose a novel CNN-based network for CS to solve this problem. In the measurement part, the input is measured block by block to acquire the measurements. While in the recovery part, all the measurements from one image are used simultaneously to reconstruct the full image. Different from previous methods recovering images block by block, the proposed framework rebuilds the structure information destroyed in the measurement part. Block effect is removed accordingly. Experiments show that there is no block effect at all in the reconstructed images. On a standard dataset our method has significant improvements in reconstruction results compared with existing state-of-the-art methods. Xuemei Xie, Chenye Wang, Jiang Du 0011, Guangming Shi |
ICME | 4 |
| 2018 | Color Image Reconstruction with Perceptual Compressive SensingabstractWe propose a novel compressive sensing framework for color images. Recently, compressive sensing (CS) has gain its popularity with the development of deep learning. To our best knowledge, existing methods all deal with RGB images channel by channel. This brings redundancy of measurements. In this paper, we do a breakthrough work. Instead of recovering RGB images channel by channel uniformly, we adopt non-uniform sampling in different channels in YCbCr color space. The luminance component takes up more measurements while the other channels take up less in the proposed framework. It greatly enhances the performance on CS for color images. Moreover, perceptual loss gives a powerful ability to better capture the structure information. We give the measurement rate at 2% as an example in the experiments, and the results show the proposed method outperforms all the existing methods with better structure of images. Jiang Du 0011, Xuemei Xie, Chenye Wang, Guangming Shi |
ICPR | 4 |
| 2018 | Lightweight Deep Residue Learning for Joint Color Image Demosaicking and DenoisingabstractColor demosaicking and image denoising each plays an important role in digital cameras. Conventional model-based methods often fail around the areas of strong textures and produce disturbing visual artifacts such as aliasing and zippering. Recently developed deep learning based methods were capable of obtaining images of better qualities though at the price of high computational cost, which make them not suitable for real-time applications. In this paper, we propose a lightweight convolutional neural network for joint demosaicking and denoising (JDD) problem with the following salient features. First, the densely connected network is trained in an end-to-end manner to learn the mapping from the noisy low-resolution space (CFA image) to the clean high-resolution space (color image). Second, the concept of deep residue learning and aggregated residual transformations are extended from image denoising and classification to JDD supporting more efficient training. Third, the design of our end-to-end network architecture is inspired by a rigorous analysis of JDD using sparsity models. Experimental results conducted for both demosaicking-only and JDD tasks have shown that the proposed method performs much better than existing state-of-the-art methods (i.e., higher visual quality, smaller training set and lower computational cost). Weisheng Dong, Guangming Shi, Xin Li 0005 |
ICPR | 4 |
| 2018 | Perceptual Compressive Sensing
Jiang Du 0011, Xuemei Xie, Chenye Wang, Guangming Shi |
PRCV (3) | 4 |
| 2018 | Motion Trajectory based Spatial-Temporal Degradation Measurement for Video Quality AssessmentabstractWith the rapid growth of digital video through the Internet, a reliable video quality assessment (VQA) technology is greatly demanded for video management. Motion information plays a dominant role for video perception, however it is too difficult to be accurately analyzed for VQA. The human visual system (HVS) is highly adaptive to track moving objects with pursuit eye movement. Inspired by the motion process in the HVS, we suggest to analyze the degradation along attended motion trajectories for VQA. As a convenient representation of motion, optical flow is calculated for motion trajectory searching. Next, the degradation on the motion velocity along each trajectory is analyzed with the optical flow for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial quality degradation from each frame, a novel VQA model is introduced. Experimental results on the public available VQA databases demonstrate that the proposed VQA model performs highly consistency with the subjective perception. Jinjian Wu, Yongxu Liu 0001, Guangming Shi |
VCIP | 3 |
| 2018 | Exploiting class-wise coding coefficients: Learning a discriminative dictionary for pattern classification
Jianqiang Song, Xuemei Xie, Guangming Shi, Weisheng Dong |
Neurocomputing | 3 |
| 2018 | Stereoscopic saliency estimation with background priors based deep reconstruction
Chen Xia, Fei Qi 0001, Guangming Shi, Chunhuan Lin |
Neurocomputing | 3 |
| 2018 | Depth sensing with coding-free pattern based on topological constraint
Guangming Shi, Ruodai Li, Fu Li 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Variable Block-Sized Signal-Dependent Transform for Video CodingabstractTransform, as one of the most important modules of mainstream video coding systems, seems very stable over the past several decades. However, recent developments indicate that bringing more options for transform can lead to coding efficiency benefits. In this paper, we go further to investigate how the coding efficiency can be improved over the state-of-the-art method by adapting a transform for each block. We present a variable block-sized signal-dependent transforms (SDTs) design based on the High Efficiency Video Coding (HEVC) framework. For a coding block ranged from $4\times4$ to $32\times32$ , we collect a quantity of similar blocks from the reconstructed area and use them to derive the Karhunen-Loève transform. We avoid sending overhead bits to denote the transform by performing the same procedure at the decoder. In this way, the transform for every block is tailored according to its statistics, to be signal-dependent. To make the large block-sized SDTs feasible, we present a fast algorithm for transform derivation. Experimental results show the effectiveness of the SDTs for different block sizes, which leads to up to 23.3% bit-saving. On average, we achieve BD-rate saving of 2.2%, 2.4%, 3.3%, and 7.1% under AI-Main10, RA-Main10, RA-Main10, and LP-Main10 configurations, respectively, compared with the test model HM-12 of HEVC. The proposed scheme has also been adopted into the joint exploration test model for the exploration of potential future video coding standard. Cuiling Lan, Jizheng Xu, Wenjun Zeng 0001, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Simultaneous Depth and Spectral Imaging With a Cross-Modal Stereo SystemabstractThis letter presents a novel approach for simultaneous depth and spectral imaging with a cross-modal stereo system. Two images of the target scene are captured at the same time: one compressively sampled hyperspectral measurement and one panchromatic measurement. The underlying hyperspectral cube is first reconstructed by leveraging the compressive sensing theory, during which a self-adaptive dictionary is learned from the panchromatic measurement to facilitate the reconstruction. The depth information of the scene is then recovered by estimating a disparity map between the hyperspectral cube and the panchromatic measurement through stereo matching. This disparity map, once obtained, is used to align the hyperspectral and panchromatic measurements to boost the hyperspectral reconstruction in an iterative manner. Through hardware experiments, for the first time to our knowledge, we demonstrate a snapshot system that allows for simultaneous depth and spectral imaging. The proposed system is capable of recording depth and spectral videos of dynamic scenes. Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Weighted Rate-Distortion Optimization for Screen Content CodingabstractUnlike camera-captured video, screen content (SC) often contains a lot of repeating patterns, which makes some blocks used as references much more important than others. However, conventional rate-distortion optimization (RDO) schemes in video coding do not consider the dependence among image blocks, which often leads to a locally optimal parameter selection, especially for SC. In this paper, we present a weighted RDO scheme for SC coding (SCC), in which the repeating characteristics are taken into account when deciding RD tradeoff for each block. For one block, the number being referenced by the current picture and following pictures is estimated and based on the number, we set a proper weight in the RDO process to reflect its importance from a global point of view. To estimate the number being referenced, we propose a hash-based method to approximate the results to avoid the complexity of direct search. Experimental results show that compared with the High Efficiency Video Coding SCC reference software, 10.1%, 14.5%, and 2.2% on average and up to 25.7%, 39.8%, and 4.6% bit saving can be achieved by considering weights provided by our scheme for hierarchical-B, IBBB, and all intra coding structures, respectively. Thanks to our hash-based design, the complexity increase brought by the proposed scheme is marginal. Bin Li 0012, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Fast Hash-Based Inter-Block Matching for Screen Content CodingabstractIn the latest High Efficiency Video Coding (HEVC) development, i.e., HEVC screen content coding extensions (HEVC-SCC), a hash-based inter-motion search/block matching scheme is adopted in the reference test model, which brings significant coding gains to code screen content. However, the hash table generation itself may take up to half the encoding time and is thus too complex for practical usage. In this paper, we propose a hierarchical hash design and the corresponding block matching scheme to significantly reduce the complexity of hash-based block matching. The hierarchical structure in the proposed scheme allows large block calculation to use the results of small blocks. Thus, we avoid redundant computation among blocks with different sizes, which greatly reduces complexity without compromising coding efficiency. The experimental results show that compared with the hash-based block matching scheme in the HEVC-SCC test model (SCM)-6.0, the proposed scheme reduces about 77% of hash processing time, which leads to 12% and 16% encoding time savings in random access (RA) and low-delay B coding structures. The proposed scheme has been adopted into the latest SCM. A parallel implementation of the proposed hash table generation on graphics processing unit (GPU) is also presented to show the high parallelism of the proposed scheme, which achieves more than 30 frames/s for 1080p sequences and 60 frames/s for 720p sequences. With the fast hash-based block matching integrated into x265 and the hash table generated on GPU, the encoder can achieve 11.8% and 14.0% coding gains on average for RA and low-delay P coding structures, respectively, for real-time encoding. Guangming Shi, Bin Li 0012, Jizheng Xu, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Dynamic Range Reduction of SAR Image via Global Optimum Entropy Maximization With Reflectivity-Distortion ConstraintabstractThe visualization of synthetic aperture radar (SAR) images plays a critical role in remote sensing applications. To effectively obtain the image suitable for human observation, this paper introduces a new SAR image visualization algorithm to map the high dynamic range SAR amplitude values to low dynamic range displays via reflectivity distortion preserved entropy maximization. Its designed objective is to present the maximal amount of information content in the displayed image, and being optimal in an information theoretical sense, as well as restricting the upper bound of the reflection distortion caused by tone mapping. The resulting optimization problem can be graph theoretically modeled as a K-edges maximum weight path problem in a directed acyclic graph, and it can be solved efficiently by dynamic programming in real time. Empirical evidences are provided to demonstrate the superior visual quality obtained by our new visualization technique. Guanghui Zhao 0003, Guangming Shi, Fu Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Image Super-Resolution With Parametric Sparse Model LearningabstractRecovering a high-resolution (HR) image from its low-resolution (LR) version is an ill-posed inverse problem. Learning accurate prior of HR images is of great importance to solve this inverse problem. Existing super-resolution (SR) methods either learn a non-parametric image prior from training data (a large set of LR/HR patch pairs) or estimate a parametric prior from the LR image analytically. Both methods have their limitations: the former lacks flexibility when dealing with different SR settings; while the latter often fails to adapt to spatially varying image structures. In this paper, we propose to take a hybrid approach toward image SR by combining those two lines of ideas - that is, a parametric sparse prior of HR images is learned from the training set as well as the input LR image. By exploiting the strengths of both worlds, we can more accurately recover the sparse codes and therefore HR image patches than conventional sparse coding approaches. Experimental results show that the proposed hybrid SR method significantly outperforms existing model-based SR methods and is highly competitive to current state-of-the-art learning-based SR methods in terms of both subjective and objective image qualities. Weisheng Dong, Xuemei Xie, Guangming Shi, Jinjian Wu, Xin Li 0005 |
IEEE Trans. Image Process. | 4 |
| 2018 | Robust Foreground Estimation via Structured Gaussian Scale Mixture ModelingabstractRecovering the background and foreground parts from video frames has important applications in video surveillance. Under the assumption that the background parts are stationary and the foreground are sparse, most of existing methods are based on the framework of robust principal component analysis (RPCA), i.e., modeling the background and foreground parts as a low-rank and sparse matrices, respectively. However, in realistic complex scenarios, the conventional norm sparse regularizer often fails to well characterize the varying sparsity of the foreground components. How to select the sparsity regularizer parameters adaptively according to the local statistics is critical to the success of the RPCA framework for background subtraction task. In this paper, we propose to model the sparse component with a Gaussian scale mixture (GSM) model. Compared with the conventional norm, the GSM-based sparse model has the advantages of jointly estimating the variances of the sparse coefficients (and hence the regularization parameters) and the unknown sparse coefficients, leading to significant estimation accuracy improvements. Moreover, considering that the foreground parts are highly structured, a structured extension of the GSM model is further developed. Specifically, the input frame is divided into many homogeneous regions using superpixel segmentation. By characterizing the set of sparse coefficients in each homogeneous region with the same GSM prior, the local dependencies among the sparse coefficients can be effectively exploited, leading to further improvements for background subtraction. Experimental results on several challenging scenarios show that the proposed method performs much better than most of existing background subtraction methods in terms of both performance and speed. Guangming Shi, Weisheng Dong, Jinjian Wu, Xuemei Xie |
IEEE Trans. Image Process. | 1 |
| 2018 | Unequal Error Protection for Scalable Video Storage in the CloudabstractRedundancy is necessary for a storage system to achieve reliability. Frequent errors in large-scale storage systems, for example, cloud, make it desirable to reduce the cost of recovery. Among all types of data in cloud storage, videos generally occupy significant amounts of space due to high volumes and the rapid development of video sharing and video-on-demand services. Unlike general data, videos can tolerate a certain level of quality degradation. This paper investigates multilayer video representations, such as scalable videos and simulcast streaming, and proposes an unequal error protection scheme based on local reconstruction codes (LRC) for video storage. By providing less protection for less important layers or video copies, a better tradeoff between storage and repair cost is achieved. Both theoretical and simulation results show that such a tradeoff can be achieved over the LRC with equal error protection, though the recovered video quality might be slightly lower in rare cases. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Blind Quality Index for Multiply Distorted Images Using Biorder Structure Degradation and Nonlocal StatisticsabstractIn the past decade, extensive image quality metrics have been proposed. The majority of them are tailored for the images that contain a specific type of distortion. However, in practice, the images are usually degraded by different types of distortions simultaneously. This poses great challenges to the existing quality metrics. Motivated by this, this paper proposes a no-reference quality index for the multiply distorted images using the biorder structure degradation and the nonlocal statistics. The design philosophy is inspired by the fact that the human visual system (HVS) is highly sensitive to the degradations of both the spatial contrast and the spatial distribution, which are prone to be changed by the joint effects of the multiple distortions. Specifically, the multiresolution representation of the image is first built by downsampling to simulate the hierarchical property of the HVS. Then, the structure degradation is calculated to measure the spatial contrast. Considering the fact that the human visual cortex has the separate mechanisms to perceive the first- and second-order structures, dubbed biorder structures, the degradations of biorder structures are calculated to account for the spatial contrast, producing the first group of the quality-aware features. Furthermore, the nonlocal self-similarity statistics is calculated to measure the spatial distribution, producing the second group of features. Finally, all the features are fed into the random forest regression model to learn the quality model for the multiply distorted images. Extensive experimental results conducted on the three public databases demonstrate the superiority of the proposed metric to the state-of-the-art metrics. Moreover, the proposed metric is also advantageous over the existing metrics in terms of the generalization ability. Yu Zhou 0009, Leida Li, Jinjian Wu, Ke Gu 0001, Weisheng Dong, Guangming Shi |
IEEE Trans. Multim. | 6 |
| 2017 | No-reference image quality assessment with orientation selectivity mechanismabstractNo-reference (NR) image quality assessment (IQA) technology is greatly required in quality-orientated visual signal processing systems. However, without the guidance of the reference information, it is still a great challenge for NR IQA to perform consistent with the subjective perception. Researches on cognitive neuroscience state that the human visual system (HVS) presents substantially orientation selectivity mechanism, within which the visual structures are extracted in the local receptive fields for scene understanding. Inspired by this mechanism, a set of orientation selectivity based visual patterns are designed. By analyzing the quality degradation on those patterns, a novel visual pattern degradation based NR IQA method is proposed. Experimental results on large databases demonstrate that the proposed method outperforms the existing NR IQA methods. Jinjian Wu, Man Zhang 0007, Guangming Shi, Xuemei Xie, Weisi Lin |
ICIP | 3 |
| 2017 | An iterative representation learning framework to predict the sequence of eye fixationsabstractVisual attention is a dynamic search process of acquiring information. However, most previous studies have focused on the prediction of static attended locations. Without considering the temporal relationship of fixations, these models usually cannot explain the dynamic saccadic behavior well. In this paper, an iterative representation learning framework is proposed to predict the saccadic scanpath. Within the proposed framework, saccade can be explained as an iterative process of finding the most uncertain area and updating the representation of scenes. In implementation, a deep autoencoder is employed for representation learning. The current fixation is predicted to be the most salient pixel, with saliency estimated by the reconstruction residual of the deep network. Image patches around this fixation are then sampled to update the network for the selection of subsequent fixations. Compared with existing models, the proposed model shows the state-of-the-art performance on several public data sets. Chen Xia, Fei Qi 0001, Guangming Shi |
ICME | 3 |
| 2017 | Saliency change based reduced reference image quality assessmentabstractThe image quality assessment (IQA) technique, which aims to perform coherently with subjective perception, is useful in quality-orientated image processing systems. In this paper, we suggest to take the saliency change into account for reduced reference (RR) IQA model. Generally, a saliency region will attract more attention, and our human vision is more sensitive to quality degradation on such region. Inspired by this, saliency values are firstly used to highlight these sensitive regions, and a local saliency weighted histogram (LSWH) based on visual orientation pattern is generated for visual feature extraction. Next, strong distortion may change the saliency from the reference to the distorted images. Thus, the saliency of each visual orientation pattern is measured, and a global saliency based histogram (GSBH) is created. Finally, by combining the LSWH and GSBH, a novel IQA model for reduced reference is introduced. Experimental results on five publicly available databases demonstrate that the proposed model uses only several values (9 values) as reference information, and performs consistently with subjective perception. Jinjian Wu, Yongxu Liu 0001, Guangming Shi, Weisi Lin |
VCIP | 3 |
| 2017 | Object-dependent sparse representation for extracellular spike detection
Guangming Shi, Zuozhi Liu, Xiaotian Wang 0001, Chengyu T. Li |
Neurocomputing | 1 |
| 2017 | Bag-of-words feature representation for blind image quality assessment with local quantized pattern
Xuemei Xie, Yazhong Zhang, Jinjian Wu, Guangming Shi, Weisheng Dong |
Neurocomputing | 4 |
| 2017 | Using 3D face priors for depth recovery
Chongyu Chen, Hai Xuan Pham, Vladimir Pavlovic 0001, Jianfei Cai 0001, Guangming Shi, Yuefang Gao |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | Single-shot dense depth sensing with frequency-division multiplexing fringe projection
Fu Li 0002, Zhiwei Xiong, Guangming Shi, Ruodai Li |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | A hierarchical multiplier-free architecture for HEVC transform
Chunxiao Fan 0002, Fu Li 0002, Guangming Shi, Fei Qi 0001, Xuemei Xie, Dandan Jiao |
Multim. Tools Appl. | 3 |
| 2017 | An AR based fast mode decision for H.265/HEVC intra coding
Fu Li 0002, Dandan Jiao, Guangming Shi, Chunxiao Fan 0002, Xuemei Xie |
Multim. Tools Appl. | 3 |
| 2017 | Adaptive Nonlocal Sparse Representation for Dual-Camera Compressive Hyperspectral ImagingabstractLeveraging the compressive sensing (CS) theory, coded aperture snapshot spectral imaging (CASSI) provides an efficient solution to recover 3D hyperspectral data from a 2D measurement. The dual-camera design of CASSI, by adding an uncoded panchromatic measurement, enhances the reconstruction fidelity while maintaining the snapshot advantage. In this paper, we propose an adaptive nonlocal sparse representation (ANSR) model to boost the performance of dual-camera compressive hyperspectral imaging (DCCHI). Specifically, the CS reconstruction problem is formulated as a 3D cube based sparse representation to make full use of the nonlocal similarity in both the spatial and spectral domains. Our key observation is that, the panchromatic image, besides playing the role of direct measurement, can be further exploited to help the nonlocal similarity estimation. Therefore, we design a joint similarity metric by adaptively combining the internal similarity within the reconstructed hyperspectral image and the external similarity within the panchromatic image. In this way, the fidelity of CS reconstruction is greatly enhanced. Both simulation and hardware experimental results show significant improvement of the proposed method over the state-of-the-art. Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Estimation of directions of arrival of multiple distributed sources for nested array
Guangming Shi, Xuemei Xie |
Signal Process. | 2 |
| 2017 | Off-grid DOA estimation under nonuniform noise via variational sparse Bayesian learning
Xuemei Xie, Guangming Shi |
Signal Process. | 3 |
| 2017 | Phase Retrieval From Multiple-Window Short-Time Fourier MeasurementsabstractIn this paper, we introduce two undirected graphs depending on supports of signals and windows, and we show that the connectivity of those graphs provides either necessary or sufficient conditions to phase retrieval of a signal from magnitude measurements of its multiple-window short-time Fourier transform. Also, we propose an algebraic reconstruction algorithm, and provide an error estimate to our algorithm when magnitude measurements are corrupted by deterministic/random noises. Cheng Cheng 0003, Deguang Han, Qiyu Sun, Guangming Shi |
IEEE Signal Process. Lett. | 5 |
| 2017 | Mixed Noise Removal via Laplacian Scale Mixture Modeling and Nonlocal Low-Rank ApproximationabstractRecovering the image corrupted by additive white Gaussian noise (AWGN) and impulse noise is a challenging problem due to its difficulties in an accurate modeling of the distributions of the mixture noise. Many efforts have been made to first detect the locations of the impulse noise and then recover the clean image with image in painting techniques from an incomplete image corrupted by AWGN. However, it is quite challenging to accurately detect the locations of the impulse noise when the mixture noise is strong. In this paper, we propose an effective mixture noise removal method based on Laplacian scale mixture (LSM) modeling and nonlocal low-rank regularization. The impulse noise is modeled with LSM distributions, and both the hidden scale parameters and the impulse noise are jointly estimated to adaptively characterize the real noise. To exploit the nonlocal self-similarity and low-rank nature of natural image, a nonlocal low-rank regularization is adopted to regularize the denoising process. Experimental results on synthetic noisy images show that the proposed method outperforms existing mixture noise removal methods. Weisheng Dong, Xuemei Xie, Guangming Shi, Xiang Bai |
IEEE Trans. Image Process. | 4 |
| 2017 | Enhanced Just Noticeable Difference Model for Images With Pattern ComplexityabstractThe just noticeable difference (JND) in an image, which reveals the visibility limitation of the human visual system (HVS), is widely used for visual redundancy estimation in signal processing. To determine the JND threshold with the current schemes, the spatial masking effect is estimated as the contrast masking, and this cannot accurately account for the complicated interaction among visual contents. Research on cognitive science indicates that the HVS is highly adapted to extract the repeated patterns for visual content representation. Inspired by this, we formulate the pattern complexity as another factor to determine the total masking effect: the interaction is relatively straightforward with a limited masking effect in a regular pattern, and is complicated with a strong masking effect in an irregular pattern. From the orientation selectivity mechanism in the primary visual cortex, the response of each local receptive field can be considered as a pattern; therefore, in this paper, the orientation that each pixel presents is regarded as the fundamental element of a pattern, and the pattern complexity is calculated as the diversity of the orientation in a local region. Finally, considering both pattern complexity and luminance contrast, a novel spatial masking estimation function is deduced, and an improved JND estimation model is built. Experimental results on comparing with the latest JND models demonstrate the effectiveness of the proposed model, which performs highly consistent with the human perception. The source code of the proposed model is publicly available at http://web.xidian.edu.cn/wjj/en/index.html. Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 4 |
| 2017 | Color-Guided Depth Recovery via Joint Local Structural and Nonlocal Low-Rank RegularizationabstractHigh-quality depth recovery from RGB-D data has received increasingly more attention in recent years due to their wide applications from depth-based image rendering to three-dimensional imaging and video. Sharp contrast between high-quality color images and low-quality depth maps presents severe challenges to the development of color-guided depth recovery techniques. Previous works have emphasized either locally varying characteristics of color-depth dependence or nonlocal similarities around the discontinuities of the scene geometry. Therefore, it is desirable to exploit both local and nonlocal structural constraints for optimizing the performance of color-guided depth recovery. In this work, we propose a unified variational approach via joint local and nonlocal regularization. The local regularization term consists of two complementary parts-one characterizing the color-depth dependence in the gradient domain and the other in the spatial domain; nonlocal regularization involves a low-rank constraint suitable for large-scale depth discontinuities. Extensive experimental results are reported to show that our approach outperforms several existing state-of-the-art depth recovery methods on both synthetic and real-world data sets. Weisheng Dong, Guangming Shi, Xin Li 0005, Kefan Peng, Jinjian Wu, Zhenhua Guo 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Distributed Compressive Sensing for Cloud-Based Wireless Image TransmissionabstractWe consider efficient image transmission via time-varying channels. To improve the performance, we propose a new distributed compressive sensing (CS) scheme that can leverage similar images in the cloud. It is featured by channel SNR and bandwidth scalability, high efficiency, and low encoding complexity. For each image, a compressed thumbnail is first transmitted after forward error correction (FEC) and modulation to retrieve similar images and generate a side information (SI) in the cloud. The residual image after subtracting the decompressed thumbnail is then coded and transmitted by CS through a very dense constellation without FEC. The linearly and ratelessly generated CS measurements make it capable of achieving both graceful quality degradation (GD) with the channel SNR and bandwidth scalability in a universal scheme. A mode decision and transform-domain power allocation are introduced for better bandwidth usage and protection against channel errors. At the decoder, a two-step CS decoding is performed to recover the residual signal, where both the local and nonlocal correlations within the image and that with the SI are exploited. Simulations on landmark images and an AWGN channel show that the received image quality gracefully increases with the channel SNR and bandwidth. Furthermore, it outperforms existing schemes both subjectively and objectively by up to 11 dB gains compared with the state-of-the-art transmission scheme with GD, i.e. SoftCast. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Enhanced just noticeable difference model with visual regularity considerationabstractJust noticeable difference (JND) reveals the visibility of our human visual system (HVS), below which changes cannot be perceived by the human. Though dozens of JND estimation models have been introduced during the past decade, how to accurately estimate the JND thresholds for different content regions (e.g., edge and texture region) is still an open problem. Research on cognitive science indicates that the HVS is adaptive to extract the visual regularities from an input scene for content perception and understanding. Thus, we analyze the effect of content regularity on visual sensitivity, and suggest that the visual regularity is another important factor that determines the JND threshold. According to the orientation distributions of local regions, the content regularities are firstly calculated. Then, by considering the effect from content regularity, luminance adaptation, and contrast masking, a novel JND model is proposed. Experimental results demonstrate that the proposed model can effectively estimate the JND thresholds of regions with different visual contents. Jinjian Wu, Guangming Shi, Weisi Lin, C.-C. Jay Kuo |
ICASSP | 2 |
| 2016 | Perceptual CU Size Decision and Fast Prediction Mode Decision Algorithm for HEVC Intra CodingabstractIntra coding has been significantly improved in HEVC over H.264/AVC with quad-tree based coding unit (CU) structure from size 64×64 to 8×8 and more prediction modes. However, these techniques cause a dramatic increase in computational complexity. In this paper, a novel intra coding algorithm is proposed consists of perceptual CU size decision algorithm and fast intra prediction mode decision algorithm. Firstly, based on the visual saliency detection, an adaptive and perceptual CU size decision method is proposed to alleviate intra encoding complexity. Furthermore, a fast intra prediction mode decision algorithm with step halving rough mode decision method is presented to selectively check the potential modes and effectively reduce the complexity of computation. Experimental results show that our proposed method reduces the computational complexity of the current HM to about 54.18% in encoding time with only 0.36% increases in BD rate and reasonable peak signal-to-noise ratio losses. Xin Zhou 0001, Guangming Shi, Wei Zhou 0020 |
ISM | 2 |
| 2016 | The 3-POCs Structure Based GPU Acceleration in Computational Spectral ImagingabstractIn this paper, the computational spectral imaging system is reexamined, and the most time-consuming module, namely the two-step iterative shrinkage/thresholding (TwIST) reconstruction, is accelerated by GPU. The acceleration can be roughly divided into two level: (1) data parallelization and (2)operation parallelization. Data parallelization: by discovering that the observation data array in computational spectral imaging is independent with each other in the norm of the dispersion direction, we propose a strategy to decompose the large scale 2-D observation data array into several small scale overlapped stripes. Considering that the complexity for TwIST is highly non-linear, this divide-conquer strategy can significantly reduce the executing time of TwIST. (2)operation parallelization: we proposed a new 3-projection onto convex sets (POCs) structure for the GPU implementation of TwIST. All the operations are classified into 3 sets which only consist of independent matrix/vector operations. With the above two contributions, the proposed 3-POCs structure based parallel implementation achieves more than 100 times ratio than the CPU based version. Xuewen Geng, Guangming Shi |
ISPDC | 5 |
| 2016 | Parallel Implementation of the Range-Doppler Radar Processing on a GPU ArchitectureabstractGraphic processing units (GPUs) is widely used to accelerate the processing speed of the radar detection procedure, including the range compression, coherent integration and constant false alarm rate. Specifically, detailed parallel design of the radar algorithm and the thread programming are shown. The experimental results show that, by engaging the parallel technology into the radar processing procedure, much high speedup ratio can be obtained. Furthermore, precise target detection can be guaranteed. Guanghui Zhao 0003, Yongfei Liu, Shuping Zhang, Fangfang Shen, Yaohai Lin, Guangming Shi |
ISPDC | 6 |
| 2016 | Learning Parametric Sparse Models for Image Super-ResolutionabstractLearning accurate prior knowledge of natural images is of great importance for single image super-resolution (SR). Existing SR methods either learn the prior from the low/high-resolution patch pairs or estimate the prior models from the input low-resolution (LR) image. Specifically, high-frequency details are learned in the former methods. Though effective, they are heuristic and have limitations in dealing with blurred LR images; while the latter suffers from the limitations of frequency aliasing. In this paper, we propose to combine those two lines of ideas for image super-resolution. More specifically, the parametric sparse prior of the desirable high-resolution (HR) image patches are learned from both the input low-resolution (LR) image and a training image dataset. With the learned sparse priors, the sparse codes and thus the HR image patches can be accurately recovered by solving a sparse coding problem. Experimental results show that the proposed SR method outperforms existing state-of-the-art methods in terms of both subjective and objective image qualities. Weisheng Dong, Xuemei Xie, Guangming Shi, Xin Li 0005, Donglai Xu |
NIPS | 4 |
| 2016 | Compressive hyperspectral imaging with complementary RGB measurementsabstractCoded aperture snapshot spectral imaging (CASSI) has been demonstrated as a feasible solution to recover a 3D hyperspectral image by using a single 2D measurement. In this paper, we propose a new hybrid camera design for CASSI to capture high quality hyperspectral images while maintaining the snapshot advantage. Specifically, we employ a complementary RGB camera in conjunction with the CASSI system. The recorded RGB image can provide reliable spectral clue of the scene. By combining the coded hyperspectral information from the CASSI branch and the uncoded color information from the RGB branch, hyperspectral images can be reconstructed with high fidelity. Furthermore, by conducting demosaicing on the raw RGB image as a preprocessing procedure, even better performance can be achieved. Both theoretical analysis and simulation results show improved accuracy of the proposed method compared to the state-of-the-arts. Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001 |
VCIP | 3 |
| 2016 | Visual information measurement with quality assessmentabstractThe quantity of visual information measurement is significant for many perception-oriented signal processing system. The classical Shannon theory, which based on the probability of signal, is useful to measure the quantity of channel information. However, it fails to accurately measure the quantity of visual information of a given image. Image quality refers to the subjective perception on the visual information that an image carried. An image with high quality carries more information than that of a low quality image. Thus, the quality of an image can effectively represent its quantity of visual information. In this paper, we propose a novel visual information measurement, and verify it with quality assessment. Firstly, a dictionary is learned from natural images, in which the content change of each atom is calculated to present its quantity of visual information. Then, a testing image is represented by the dictionary, and the sparse coefficients for each local block in the image are acquired. Finally, according to the sparse coefficients, the quantity of visual information is measured. The information measurement result is verified with the subjective quality score. Experimental results on a large amount of images demonstrate the accuracy of the proposed method for quantity of visual information measurement. Jinjian Wu, Guangming Shi, Man Zhang 0007, Guanmi Chen |
VCIP | 2 |
| 2016 | High quality impulse noise removal via non-uniform sampling and autoregressive modelling based super-resolutionabstractThe challenge of image impulse noise removal is to restore spatial details from damaged pixels using remaining ones in random locations. Most existing methods use all uncontaminated pixels within a local window to estimate the centred noisy one via a statistic way. These kinds of methods have two defects. First, all noisy pixels are treated as independent individuals and estimated by their neighbours one by one, with the correlation between their true values ignored. Second, the image structure as a natural feature is usually ignored. This study proposes a new denoising framework, in which all noisy pixels are jointly restored via non‐uniform sampling and supervised piecewise autoregressive modelling based super‐resolution. In this method, the noisy pixels are jointly estimated in groups through solving a well‐designed optimisation problem, in which image structure feature is considered as an important constraint. Another contribution is that piecewise autoregressive model is not simply adopted but carefully designed so that all noise‐free pixels can be used to supervise the model training and optimisation problem solving for higher accuracy. The experimental results demonstrate that the proposed method exhibits good denoising performance in a large noise density range (10–90%). Xiaotian Wang 0001, Guangming Shi, Jinjian Wu, Fu Li 0002, Yantao Wang |
IET Image Process. | 2 |
| 2016 | Orientation selectivity based visual pattern for reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Guangming Shi, Leida Li, Yuming Fang 0001 |
Inf. Sci. | 3 |
| 2016 | Energy Efficient Resource Allocation for Wireless Power Transfer Enabled Collaborative Mobile CloudsabstractIn order to fully enjoy high rate broadband multimedia services, prolonging the battery lifetime of user equipment is critical for mobile users, especially for smartphone users. In this paper, the problem of distributing cellular data via a wireless power transfer enabled collaborative mobile cloud (WeCMC) in an energy efficient manner is investigated. WeCMC is formed by a group of users who have both functionalities of information decoding and energy harvesting, and are interested for cooperating in downloading content from the operators. Through device-to-device communications, the users inside WeCMC are able to cooperate during the downloading procedure and offload data from the base station to other WeCMC members. When considering multi-input multi-output wireless channel and wireless power transfer, an efficient algorithm is presented to optimally schedule the data offloading and radio resources in order to maximize energy efficiency as well as fairness among mobile users. Specifically, the proposed framework takes energy minimization and quality of service requirement into consideration. Performance evaluations demonstrate that a significant energy saving gain can be achieved by the proposed schemes. Zheng Chang 0001, Jie Gong 0003, Yingyu Li, Zhenyu Zhou 0001, Tapani Ristaniemi, Guangming Shi, Zhu Han 0001, Zhisheng Niu |
IEEE J. Sel. Areas Commun. | 6 |
| 2016 | Iterative non-local means filter for salt and pepper noise removal
Xiaotian Wang 0001, Shanshan Shen, Guangming Shi, Yuan-nan Xu |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Visual structural degradation based reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Yuming Fang 0001, Leida Li, Guangming Shi, S. Issac Niwas |
Signal Process. Image Commun. | 5 |
| 2016 | Hybrid Distortion Ranking Tuned Bitstream-Layer Video Quality AssessmentabstractNo-reference bitstream-layer video quality assessment is very important and practical for monitoring the perceptual experience of end users and facilitating network maintenance. For pervasive Internet Protocol Television and mobile streaming services, in addition to quality degradation due to lossy compression, the unreliable transmission mechanism (i.e., User Datagram Protocol/IP) often leads to quality degradation due to packet loss. Different technical solutions bring in different types of visual artifacts. In this paper, we proposed a hybrid distortion ranking (HDR)-based bitstream-layer quality assessment model, whose artifact combination framework is based on the ranked linear combination operation. The model can predict the perceived quality of a video with sufficient accuracy when the video is distorted by compression artifacts, slicing artifacts, freezing (with frame skipping) artifacts, or their combinations. The core algorithms of the model were adopted into ITU-T Recommendations, P.1202.1 and P.1202.2. Furthermore, with respect to the three different types of artifacts, we compared the proposed no-reference HDR model with some state-of-the-art full-reference perceptual quality assessment models including Video Quality Model (i.e., ITU-T Rec. J.144), structural similarity (SSIM), multiscale SSIM, visual information fidelity, and the widely used metric, peak signal-to-noise ratio. We also compared our HDR model with the top performing no-reference models including Blind/Referenceless Image Spatial Quality Evaluator and video Blind Prediction of Natural Video Quality. The experiment results demonstrate the efficiency of our HDR model. Zhibo Chen 0001, Ning Liao, Xiaodong Gu 0005, Feng Wu 0001, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Hyperspectral Image Super-Resolution via Non-Negative Structured Sparse RepresentationabstractHyperspectral imaging has many applications from agriculture and astronomy to surveillance and mineralogy. However, it is often challenging to obtain high-resolution (HR) hyperspectral images using existing hyperspectral imaging techniques due to various hardware limitations. In this paper, we propose a new hyperspectral image super-resolution method from a low-resolution (LR) image and a HR reference image of the same scene. The estimation of the HR hyperspectral image is formulated as a joint estimation of the hyperspectral dictionary and the sparse codes based on the prior knowledge of the spatial-spectral sparsity of the hyperspectral image. The hyperspectral dictionary representing prototype reflectance spectra vectors of the scene is first learned from the input LR image. Specifically, an efficient non-negative dictionary learning algorithm using the block-coordinate descent optimization technique is proposed. Then, the sparse codes of the desired HR hyperspectral image with respect to learned hyperspectral basis are estimated from the pair of LR and HR reference images. To improve the accuracy of non-negative sparse coding, a clustering-based structured sparse coding method is proposed to exploit the spatial correlation among the learned sparse codes. The experimental results on both public datasets and real LR hypspectral images suggest that the proposed method substantially outperforms several existing HR hyperspectral image recovery techniques in the literature in terms of both objective quality metrics and computational efficiency. Weisheng Dong, Fazuo Fu, Guangming Shi, Xun Cao, Jinjian Wu, Xin Li 0005 |
IEEE Trans. Image Process. | 3 |
| 2016 | Image Enhancement by Entropy Maximization and Quantization Resolution UpconversionabstractThis article introduces a new contrast enhancement algorithm of tone-preserving entropy maximization. Its design objective is to present the maximal amount of information content in the enhanced image, or being optimal in an information theoretical sense, while preventing the loss of tone continuity. The resulting optimization problem can be graph theoretically modeled as the construction of the K-edges maximum-weight path, and it can be solved efficiently by dynamic programming. Moreover, the proposed algorithm is made more effective by being combined with a preprocess of image restoration that aims to correct quantization errors caused by the analog-to-digital conversion of image signals. Empirical evidences are provided to demonstrate the superior visual quality obtained by the new image enhancement algorithm. Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2016 | Bottom-Up Visual Saliency Estimation With Deep Autoencoder-Based Sparse ReconstructionabstractResearch on visual perception indicates that the human visual system is sensitive to center-surround (C-S) contrast in the bottom-up saliency-driven attention process. Different from the traditional contrast computation of feature difference, models based on reconstruction have emerged to estimate saliency by starting from original images themselves instead of seeking for certain ad hoc features. However, in the existing reconstruction-based methods, the reconstruction parameters of each area are calculated independently without taking their global correlation into account. In this paper, inspired by the powerful feature learning and data reconstruction ability of deep autoencoders, we construct a deep C-S inference network and train it with the data sampled randomly from the entire image to obtain a unified reconstruction pattern for the current image. In this way, global competition in sampling and learning processes can be integrated into the nonlocal reconstruction and saliency estimation of each pixel, which can achieve better detection results than the models with separate consideration on local and global rarity. Moreover, by learning from the current scene, the proposed model can achieve the feature extraction and interaction simultaneously in an adaptive way, which can form a better generalization ability to handle more types of stimuli. Experimental results show that in accordance with different inputs, the network can learn distinct basic features for saliency modeling in its code layer. Furthermore, in a comprehensive evaluation on several benchmark data sets, the proposed method can outperform the existing state-of-the-art algorithms. Chen Xia, Fei Qi 0001, Guangming Shi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | High-speed hyperspectral video acquisition with a dual-camera architectureabstractWe propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hyperspectral video at a low frame rate, which jointly provide reliable projections for the underlying HSHS video. Second, we exploit the panchromatic video to learn an over-complete 3D dictionary to represent each band-wise video sparsely, and a robust computational reconstruction is then employed to recover the HSHS video based on the joint videos and the self-learned dictionary. Experimental results demonstrate that, for the first time to our knowledge, the hyperspectral video frame rate reaches up to 100fps with decent quality, even when the incident light is not strong. Lizhi Wang 0001, Zhiwei Xiong, Dahua Gao, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001 |
CVPR | 4 |
| 2015 | Low-Rank Tensor Approximation with Laplacian Scale Mixture Modeling for Multiframe Image DenoisingabstractPatch-based low-rank models have shown effective in exploiting spatial redundancy of natural images especially for the application of image denoising. However, two-dimensional low-rank model can not fully exploit the spatio-temporal correlation in larger data sets such as multispectral images and 3D MRIs. In this work, we propose a novel low-rank tensor approximation framework with Laplacian Scale Mixture (LSM) modeling for multi-frame image denoising. First, similar 3D patches are grouped to form a tensor of d-order and high-order Singular Value Decomposition (HOSVD) is applied to the grouped tensor. Then the task of multiframe image denoising is formulated as a Maximum A Posterior (MAP) estimation problem with the LSM prior for tensor coefficients. Both unknown sparse coefficients and hidden LSM parameters can be efficiently estimated by the method of alternating optimization. Specifically, we have derived closed-form solutions for both subproblems. Experimental results on spectral and dynamic MRI images show that the proposed algorithm can better preserve the sharpness of important image structures and outperform several existing state-of-the-art multiframe denoising methods (e.g., BM4D and tensor dictionary learning). Weisheng Dong, Guangming Shi, Xin Li 0005, Yi Ma 0001 |
ICCV | 3 |
| 2015 | Learning Parametric Distributions for Image Super-Resolution: Where Patch Matching Meets Sparse CodingabstractExisting approaches toward Image super-resolution (SR) is often either data-driven (e.g., based on internet-scale matching and web image retrieval) or model-based (e.g., formulated as an Maximizing a Posterior estimation problem). The former is conceptually simple yet heuristic, while the latter is constrained by the fundamental limit of frequency aliasing. In this paper, we propose to develop a hybrid approach toward SR by combining those two lines of ideas. More specifically, the parameters underlying sparse distributions of desirable HR image patches are learned from a pair of LR image and retrieved HR images. Our hybrid approach can be interpreted as the first attempt of reconciling the difference between parametric and nonparametric models for low-level vision tasks. Experimental results show that the proposed hybrid SR method performs much better than existing state-of-the-art methods in terms of both subjective and objective image qualities. Weisheng Dong, Guangming Shi, Xuemei Xie |
ICCV | 3 |
| 2015 | Unequal error protection for scalable video storage in the cloudabstractRedundancy is necessary for a storage system to recover from errors. The frequent errors in large-scale systems, e.g. cloud, make it desired to reduce the recovery cost. Among all kinds of data stored in the cloud, video takes a large portion due to its large data volume. The other characteristic of video is that a certain distortion can be tolerated. This paper investigates using scalable video representation and unequal error protection scheme to reduce the storage and recovery costs in the cloud. By introducing more protection for the base layer and less on the enhancement layers, it can achieve a better tradeoff between storage and reconstruction costs although the reliability for the enhancement layer sacrifices a little. Simulation results based on local reconstruction codes (LRC) show that comparing with the existing (12, 2, 2) LRC code in Windows Azure Storage, the reconstruction cost can be reduced from 6x to 3x at the same storage cost at the expense of possible video quality loss. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
ICME | 4 |
| 2015 | Reduced-reference image quality assessment based on entropy differences in DCT domainabstractReduced-reference image quality assessment (RR-IQA) algorithm aims to automatically evaluate the image quality using only partial information about the reference image. In this paper, we propose a new RR-IQA metric by employing the entropy features of each frequency band in the DCT domain. It is well known that human eyes have different sensitivity to different bands, and distortions on each band result in individual quality degradations. Therefore, we suggest to separately compute the visual information degradations on different band for quality assessment. The degradations on each DCT band are firstly analyzed according to the entropy difference. And then, the quality score is obtained using the weighted sum of the entropy difference of each band from low frequency to high frequency. Experimental results on several public image databases show that the proposed method uses limited reference data (8 values) and performs highly consistent with human perception. Yazhong Zhang, Jinjian Wu, Guangming Shi, Xuemei Xie |
ISCAS | 3 |
| 2015 | Compressive sensing based image transmission with side information at the decoderabstractThis paper proposes a distributed compressive sensing (CS) scheme for robust image transmission over unknown or time-varying channels with highly correlated images at the decoder. A compressed thumbnail is first transmitted after digital forward error correction (FEC) and modulation to retrieve highly correlated images and generate a side information (SI) at the decoder. The current residual image after subtracting the decompressed thumbnail is then coded and transmitted by CS through a very dense constellation without FEC. The linear representation of the residual signal by CS measurements and rateless sampling makes it able to achieve graceful degradation and bandwidth scalability without channel feedback. Moreover, a transform-domain power allocation is employed before random sampling to protect against channel errors. At the decoder, both the nonlocal correlations within the original image and the correlation with the SI are exploited in CS decoding via a low-rank regulation on similar patches. After CS decoding, a block-wise minimum-mean-square-error (MMSE) reconstruction using the SI is further performed in the spatial domain to enhance the reconstruction quality. Simulations on landmark images and an unknown Gaussian channel show that an up to 10 dB gain is achieved at low channel SNRs compared with the state-of-the-art uncoded image transmission scheme, i.e. SoftCast, when highly correlated images are available at the decoder. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
VCIP | 4 |
| 2015 | Weighted rate-distortion optimization for screen content intra codingabstractScreen content videos often have mixed content consisting of various types such as natural content, text and graphics in the same picture. To achieve high coding efficiency for this content, new coding tools are developed in the High Efficiency Video Coding (HEVC) standard Screen Content Coding (SCC) extension. Among them, intra block copy (IntraBC) allows a nonlocal intra prediction from the coded region of the same picture. However, the rate-distortion optimization (RDO) scheme of the screen content coding still follows that of the HEVC reference software and each block is optimized locally. The screen content characteristics are not fully utilized in the current RDO scheme. This paper presents a weighted RDO scheme for the intra coding of screen content videos, which measures the importance of each block within the picture first and then larger distortion weights are applied to the blocks with larger importance in the RDO. In this way, a better rate-distortion trade-off can be achieved for the picture instead of the blocks themselves. The experimental results show that up to 5.5% coding efficiency gain can be achieved compared with the reference software. Bin Li 0012, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
VCIP | 4 |
| 2015 | Image Restoration via Simultaneous Sparse Coding: Where Structured Sparsity Meets Gaussian Scale Mixture
Weisheng Dong, Guangming Shi, Yi Ma 0001, Xin Li 0005 |
Int. J. Comput. Vis. | 2 |
| 2015 | Incremental low-rank and sparse decomposition for compressing videos captured by fixed cameras
Chongyu Chen, Jianfei Cai 0001, Weisi Lin, Guangming Shi |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | Nonlocal center-surround reconstruction-based bottom-up saliency estimation
Chen Xia, Fei Qi 0001, Guangming Shi, Pengjin Wang |
Pattern Recognit. | 3 |
| 2015 | Nonbinary LDPC Codes on Cages: Structural Property and Code OptimizationabstractA (v,g)-cage is a (not necessarily unique) smallest v-regular graph of girth g. On such a graph, a nonbinary (2,v)-regular low-density parity-check (LDPC) code can be defined such that the Tanner graph has girth 2g and the code length achieves the minimum possible. In this paper, we focus on two aspects of this class of codes, structural property and code optimization. We find that, in addition to those found previously, many cages can be used to construct structured LDPC codes. We show that all cages with even girth can be structured as protograph-based codes, many of which have block-circulant Tanner graphs. We also find that four cages with odd girth can be structured as protograph-based codes with block-circulant Tanner graphs. For code optimization, we develop an ontology-based approach. All possible inter-connected cycle patterns that lead to low symbol-weight codewords are identified to put together the ontology. By doing so, it becomes handleable to estimate and optimize distance spectrum of equivalent binary image codes. We further analyze some known codes from the Consultative Committee for Space Data Systems recommendation and design several new codes. Numerical results show that these codes have reasonably good minimum bit distance and perform well under iterative decoding. Chao Chen 0013, Baoming Bai, Guangming Shi, Xiaotian Wang 0001, Xiaopeng Jiao |
IEEE Trans. Commun. | 3 |
| 2015 | Cloud-Based Distributed Image CodingabstractWith multimedia flourishing on the Web, it is easy to find similar images for a query, especially landmark images. Traditional image coding, such as JPEG, cannot exploit correlations with external images. Existing vision-based approaches are able to exploit such correlations by reconstructing from local descriptors but cannot ensure the pixel-level fidelity of the reconstruction. In this paper, a cloud-based distributed image coding (Cloud-DIC) scheme is proposed to exploit external correlations for mobile photo uploading. For each input image, a thumbnail is transmitted to retrieve correlated images and reconstruct it in the cloud by geometrical and illumination registrations. Such a reconstruction serves as the side information (SI) in the Cloud-DIC. The image is then compressed by a transform-domain syndrome coding to correct the disparity between the original image and the SI. Once a bitplane is received in the cloud, an iterative refinement process is performed between the final reconstruction and the SI. Moreover, a joint encoder/decoder mode decision at block, frequency, and bitplane levels is proposed to adapt to different correlations. Experimental results on a landmark image database show that the Cloud-DIC can largely enhance the coding efficiency both subjectively and objectively, with up to 5-dB gains and 70% bits saving over JPEG with arithmetic coding, and perform comparably at low bitrates with the intra coding of the High Efficiency Video Coding standard with a much lower encoder complexity. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | HEVC Encoding Optimization Using Multicore CPUs and GPUsabstractAlthough the High Efficiency Video Coding (HEVC) standard significantly improves the coding efficiency of video compression, it is unacceptable even in offline applications to spend several hours compressing 10 s of high-definition video. In this paper, we propose using a multicore central processing unit (CPU) and an off-the-shelf graphics processing unit (GPU) with 3072 streaming processors (SPs) for HEVC fast encoding, so that the speed optimization does not result in loss of coding efficiency. There are two key technical contributions in this paper. First, we propose an algorithm that is both parallel and fast for the GPU, which can utilize 3072 SPs in parallel to estimate the motion vector (MV) of every prediction unit (PU) in every combination of the coding unit (CU) and PU partitions. Furthermore, the proposed GPU algorithm can avoid coding efficiency loss caused by the lack of a MV predictor (MVP). Second, we propose a fast algorithm for the CPU, which can fully utilize the results from the GPU to significantly reduce the number of possible CU and PU partitions without any coding efficiency loss. Our experimental results show that compared with the reference software, we can encode high-resolution video that consumes 1.9% of the CPU time and 1.0% of the GPU time, with only a 1.4% rate increase. Bin Li 0012, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Retrieval of Atmospheric Aerosol and Surface Properties Over Land Using Satellite ObservationsabstractIt is difficult to retrieve aerosol and surface reflectance properties over land, simultaneously, from satellite observations. In the Moderate Resolution Imaging Spectroradiometer (MODIS) aerosol algorithm, 2.1-μm observations are essential to determine land surface reflectances. However, observations at 2.1 μm are not widely available for other sensors. To get rid of 2.1-μm observations in aerosol retrievals, we suggest a new method that makes use of the priori relationships among surface reflectances at different wavelengths in terms of the surface normalized difference vegetation index (sNDVI), which is defined by the surface reflectances at 0.87 and 0.67 μm. It has advantages to parameterize surface reflectances due to its independence from atmospheric condition. Based on these newly established surface reflectance relationships and the MODIS observations at 0.47, 0.67, 0.87, and 1.64 μm, we solve four variables simultaneously with the new method: aerosol optical depth (AOD) at 0.55 μm, proportion of fine mode AOD, surface reflectance at 0.67 μm, and sNDVI. The preliminary retrieval tests indicate that this method can successfully retrieve the aerosol properties and land cover features over East Asia. A comparison of the retrieved AODs at 0.55 μm with the AERONET observations shows a root-mean-square error (RMSE) of about 0.16, at the Beijing and Xianghe sites, in August 2007. The accuracy is the same order of magnitude as the MODIS aerosol algorithm (RMSE = 0.14) during the same period. The results demonstrate that the quality of aerosol retrieval in the new method is comparable to that of the MODIS aerosol algorithm over the testing area. Guangming Shi, Chengcai Li, Tong Ren, Yefang Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Visual Orientation Selectivity Based Structure DescriptionabstractThe human visual system is highly adaptive to extract structure information for scene perception, and structure character is widely used in perception-oriented image processing works. However, the existing structure descriptors mainly describe the luminance contrast of a local region, but cannot effectively represent the spatial correlation of structure. In this paper, we introduce a novel structure descriptor according to the orientation selectivity mechanism in the primary visual cortex. Research on cognitive neuroscience indicate that the arrangement of excitatory and inhibitory cortex cells arise orientation selectivity in a local receptive field, within which the primary visual cortex performs visual information extraction for scene understanding. Inspired by the orientation selectivity mechanism, we compute the correlations among pixels in a local region based on the similarities of their preferred orientation. By imitating the arrangement of the excitatory/inhibitory cells, the correlations between a central pixel and its local neighbors are binarized, and the spatial correlation is represented with a set of binary values, which is named the orientation selectivity-based pattern. Then, taking both the gradient magnitude and the orientation selectivity-based pattern into account, a rotation invariant structure descriptor is introduced. The proposed structure descriptor is applied in texture classification and reduced reference image quality assessment, as two different application domains to verify its generality and robustness. Experimental results demonstrate that the orientation selectivity-based structure descriptor is robust to disturbance, and can effectively represent the structure degradation caused by different types of distortion. Jinjian Wu, Weisi Lin, Guangming Shi, Yazhong Zhang, Weisheng Dong, Zhibo Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Kinect Depth Recovery Using a Color-Guided, Region-Adaptive, and Depth-Selective FrameworkabstractConsidering that the existing depth recovery approaches have different limitations when applied to Kinect depth data, in this article, we propose to integrate their effective features including adaptive support region selection, reliable depth selection, and color guidance together under an optimization framework for Kinect depth recovery. In particular, we formulate our depth recovery as an energy minimization problem, which solves the depth hole filling and denoising simultaneously. The energy function consists of a fidelity term and a regularization term, which are designed according to the Kinect characteristics. Our framework inherits and improves the idea of guided filtering by incorporating structure information and prior knowledge of the Kinect noise model. Through analyzing the solution to the optimization framework, we also derive a local filtering version that provides an efficient and effective way of improving the existing filtering techniques. Quantitative evaluations on our developed synthesized dataset and experiments on real Kinect data show that the proposed method achieves superior performance in terms of recovery accuracy and visual quality. Chongyu Chen, Jianfei Cai 0001, Jianmin Zheng, Tat-Jen Cham, Guangming Shi |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2014 | Depth Recovery with Face Priors
Chongyu Chen, Hai Xuan Pham, Vladimir Pavlovic 0001, Jianfei Cai 0001, Guangming Shi |
ACCV (4) | 5 |
| 2014 | Sparsity fine tuning in wavelet domain with application to compressive image reconstructionabstractIn compressive sensing, wavelet space is widely used to generate sparse signal (image signal in particular) representations. In this work, we propose a novel approach of statistical context modeling to increase the level of sparsity of wavelet image representations. It is shown, contrary to a widely held assumption, that high-frequency wavelet coefficients have non-zero mean distributions if conditioned on local image structures. Removing this bias can make wavelet image representations sparser, i.e., having a greater number of zero and close-to-zero coefficients. The resulting unbiased probability models can significantly improve the performance of existing wavelet-based compressive image reconstruction methods in both PSNR and visual quality. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
ICASSP | 3 |
| 2014 | Image restoration via Bayesian structured sparse codingabstractIn this work, we propose a Bayesian structured sparse coding (BSSC) framework containing a nonlocal extension of Gaussian scale mixture (GSM) model by exploiting structured sparsity. It is shown that the variances of sparse coefficients (the field of Gaussian scalars) - if treated as a latent variable - can besparse coefficients jointly estimated along with the unknown sparse coefficients via the the method of alternative optimization. When applied to image restoration, BSSC leads to closed-form solutions involving iterative shrinkage/filtering and therefore admits computationally efficient implementation. Our experimental results have shown that BSSC-based image restoration often delivers reconstructed images with higher subjective/objective qualities than other competing approaches including IDD-BM3D and NCSR. Weisheng Dong, Xin Li 0005, Yi Ma 0001, Guangming Shi |
ICIP | 4 |
| 2014 | Image enhancement by entropy maximization and quantization resolution upconversionabstractThis article introduces a new contrast enhancement algorithm of tone-preserving entropy maximization. Its design objective is to present the maximal amount of information content in the enhanced image, or being optimal in an information theoretical sense, while preventing the loss of tone continuity. The resulting optimization problem can be graph-theoretically modeled as the construction of K-edge maximum-weight path, and it can be solved efficiently by dynamic programming. Moreover, the proposed algorithm is made more effective by being combined with a preprocess of image restoration that aims to correct quantization errors caused by the analog-to-digital conversion of image signals. Empirical evidences are provided to demonstrate the superior visual quality obtained by the new image enhancement algorithm. Xiaolin Wu 0001, Guangming Shi |
ICIP | 3 |
| 2014 | Reduced-reference image quality assessment with local binary structural patternabstractReduced-reference (RR) image quality assessment (IQA) aims to use less reference data and achieve higher quality prediction accuracy. Recent researches confirm that the human visual system (HVS) is adapted to extract structural information and is sensitive to structure degradation. Therefore, in this paper, we try to represent image contents with several structural patterns, and measure image quality according to the structural degradation on these patterns. The classic local binary patterns (LBPs) are firstly employed to extract image structures and create LBP based structural histogram. And then, the structural degradation is computed as the histogram distance between the reference and distorted images. Experimental results on three large databases demonstrate that the proposed RR IQA method greatly improved the quality prediction accuracy. Jinjian Wu, Weisi Lin, Guangming Shi, Long Xu 0001 |
ISCAS | 3 |
| 2014 | Correlation based universal image/video coding loss recovery
Jinjian Wu, Weisi Lin, Guangming Shi, Jimin Xiao |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | A linear combination-based weighted least square approach for target localization with noisy range measurements
Fei Qi 0001, Guangming Shi, Jingbo Ren |
Signal Process. | 3 |
| 2014 | Image Quality Assessment with Degradation on Spatial StructureabstractIn this letter, we introduce an improved structural degradation based image quality assessment (IQA) method. Most of the existing structural similarity based IQA metrics mainly consider the spatial contrast degradation but have not fully considered the changes on the spatial distribution of structures. Since the human visual system (HVS) is sensitive to degradations on both spatial contrast and spatial distribution, both factors need to be considered for IQA. In order to measure the structural degradation on spatial distribution, the local binary patterns (LBPs) are first employed to extract structural information. And then, the LBP shift between the reference and distorted images is computed, because noise distorts structural patterns. Finally, the spatial contrast degradation on each pair of LBP shifts is calculated for quality assessment. Experimental results on three large benchmark databases confirm that the proposed IQA method is highly consistent with the subjective perception. Jinjian Wu, Weisi Lin, Guangming Shi |
IEEE Signal Process. Lett. | 3 |
| 2014 | Nonlocal Sparse and Low-Rank Regularization for Optical Flow EstimationabstractDesigning an appropriate regularizer is of great importance for accurate optical flow estimation. Recent works exploiting the nonlocal similarity and the sparsity of the motion field have led to promising flow estimation results. In this paper, we propose to unify these two powerful priors. To this end, we propose an effective flow regularization technique based on joint low-rank and sparse matrix recovery. By grouping similar flow patches into clusters, we effectively regularize the motion field by decomposing each set of similar flow patches into a low-rank component and a sparse component. For better enforcing the low-rank property, instead of using the convex nuclear norm, we use the log det(·) function as the surrogate of rank, which can also be efficiently minimized by iterative singular value thresholding. Experimental results on the Middlebury benchmark show that the performance of the proposed nonlocal sparse and low-rank regularization method is higher than (or comparable to) those of previous approaches that harness these same priors, and is competitive to current state-of-the-art methods. Weisheng Dong, Guangming Shi, Xiaocheng Hu, Yi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Compressive Sensing via Nonlocal Low-Rank RegularizationabstractSparsity has been widely exploited for exact reconstruction of a signal from a small number of random measurements. Recent advances have suggested that structured or group sparsity often leads to more powerful signal reconstruction techniques in various compressed sensing (CS) studies. In this paper, we propose a nonlocal low-rank regularization (NLR) approach toward exploiting structured sparsity and explore its application into CS of both photographic and MRI images. We also propose the use of a nonconvex log det ( X) as a smooth surrogate function for the rank instead of the convex nuclear norm and justify the benefit of such a strategy using extensive experiments. To further improve the computational efficiency of the proposed algorithm, we have developed a fast implementation using the alternative direction multiplier method technique. Experimental results have shown that the proposed NLR-CS algorithm can significantly outperform existing state-of-the-art CS techniques for image recovery. Weisheng Dong, Guangming Shi, Xin Li 0005, Yi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Sparsity Fine Tuning in Wavelet Domain With Application to Compressive Image ReconstructionabstractIn compressive sensing, wavelet space is widely used to generate sparse signal (image signal in particular) representations. In this paper, we propose a novel approach of statistical context modeling to increase the level of sparsity of wavelet image representations. It is shown, contrary to a widely held assumption, that high-frequency wavelet coefficients have nonzero mean distributions if conditioned on local image structures. Removing this bias can make wavelet image representations sparser, i.e., having a greater number of zero and closeto-zero coefficients. The resulting unbiased probability models can significantly improve the performance of existing wavelet-based compressive image reconstruction methods in both PSNR and visual quality. An efficient algorithm is presented to solve the compressive image recovery (CIR) problem using the refined models. Experimental results on both simulated compressive sensing (CS) image data and real CS image data show that the new CIR method significantly outperforms existing CIR methods in both PSNR and visual quality. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2013 | Nonlocal center-surround reconstruction-based bottom-up saliency estimationabstractThe center-surround comparison principle is widely used in existing bottom-up saliency estimation models. However, most of them are based on local image processing techniques which are hard to handle texture regions well as a relatively large neighborhood is required to represent textures. In this paper, we propose a nonlocal patch-based reconstruction approach to reformulate the center-surround comparison. In the proposed approach, the saliency is measured by the reconstruction residual of representing the central patch with a linear combination of its surrounding patches. As a generalization of Itti et al.'s classical center-surround comparison scheme, the proposed approach performs well on images with symmetric structures where Itti et al.'s method fails, as well as on general natural images. Numerical experiments show the proposed approach produces better results compared to the state-of-the-art algorithms on several public databases. Chen Xia, Pengjin Wang, Fei Qi 0001, Guangming Shi |
ICIP | 4 |
| 2013 | Compressive modulation in digital communicationabstractBandwidth efficiency is one of the most important indicators to measure different modulation schemes in digital communication systems. The waveforms of existing modulation schemes are all separated in time domain, making it difficult for them to improve in bandwidth efficiency. Compressive Sensing (CS) theory shows that it is possible to reconstruct original signals in aliasing measurements. In this paper, we propose a Compressive Modulation scheme by combining CS theory and traditional BPSK pattern. Theoretic analysis and experimental results show that the bandwidth efficiency can be highly improved by using the proposed scheme. Yingyu Li, Guangming Shi, Xuemei Xie, Chongyu Chen |
ISCAS | 2 |
| 2013 | Visual masking estimation based on structural uncertaintyabstractA model of visual masking, which reveals the visible threshold of human perception, is useful in perceptual based image/video processing. The existing visual masking formulation, which mainly considers luminance contrast, cannot accurately estimate the visible threshold. Recent researches indicate that human perception is highly adaptive to extract orderly structures and is insensitive to disorderly structures. Therefore, we suggest that the structural characteristic is another determining factor for visual masking, and deduce a novel visual masking function based on structural uncertainty. Experimental results demonstrate that the proposed model is more consistent with human perception than the existing visual masking model. Jinjian Wu, Weisi Lin, Guangming Shi |
ISCAS | 3 |
| 2013 | A color-guided, region-adaptive and depth-selective unified framework for Kinect depth recoveryabstractConsidering the existing depth recovery approaches that have different limitations when applying to Kinect depth data, in this paper, we propose to integrate their effective features including adaptive support region selection, reliable depth selection and color guidance together under a unified framework for Kinect depth recovery. In particular, we formulate our depth recovery as an energy minimization problem, which solves the depth hole-filling and denoising simultaneously. The energy function consists of a fidelity term and a regularization term. The fidelity term takes into account the characteristics of Kinect data. The regularization term is designed to incorporate the joint bilateral filtering (JBF) kernel and the joint trilateral filtering (JTF) kernel so as to facilitate both depth hole-filling and denoising. Moreover, the JBF kernel is modified to incorporate the structure information. Both simulations on the benchmark Middlebury dataset and experiments on real Kinect data show that our proposed method achieves state-of-the-art performance in terms of recovery accuracy and visual quality. Chongyu Chen, Jianfei Cai 0001, Jianmin Zheng, Tat-Jen Cham, Guangming Shi |
MMSP | 5 |
| 2013 | Dense depth acquisition via one-shot stripe structured lightabstractDepth acquisition for moving objects becomes increasingly critical for some applications such as human facial expression recognition. This paper presents a method for capturing the depth maps of moving objects that uses a one-shot black-and-white stripe pattern with the features of simplicity and easily generation. Considering the accuracy of a matching is crucial for a precise depth map but the matching of variant-width stripes is sparse and rough, the phase differences extracted by Gabor filter to achieve a pixel-wise matching with sub-pixel accuracy are used. The details of the derivation are presented to prove that this method based on the phase difference calculated by Gabor filter is valid. In addition, the periodic ambiguity of the encoded stripe is eliminated by the epipolar segment covering a given depth range at a camera-projector calibrating stage to decrease the calculation complexity. Experimental results show that our method can get a dense and accurate depth map of a moving object. Fu Li 0002, Guangming Shi, Fei Qi 0001, Yuexin Shi |
VCIP | 3 |
| 2013 | A high quality image reconstruction method based on nonconvex decoding
Guanghui Zhao 0003, Fangfang Shen, Guangming Shi, Danhua Liu |
Sci. China Inf. Sci. | 5 |
| 2013 | Cauchy diversity measures: a novel methodology for enhancing sparsity in compressed sensingabstractAs a new enchanting theory, compressed sensing (CS) demonstrates that a sparse signal can be recovered through a surprisingly small number of linear measurements by solving a problem of ℓ 1 norm minimisation (which can be thought as a special case of the signomial diversity measures). However, the traditional CS model with ℓ 1 norm minimisation can not fully exploit the sparsity especially when the degree of sparsity increases or the measurements number reduces. In this study, the Cauchy diversity measures is incorporated into the proposed model to deal with the above difficulties. The simulation results demonstrate that under the same condition, this new model offers a superior reconstruction precision compared with the common used signomial diversity measures. Guanghui Zhao 0003, Fangfang Shen, Guangming Shi |
IET Signal Process. | 4 |
| 2013 | A learning-based method for compressive image recovery
Weisheng Dong, Guangming Shi, Xiaolin Wu 0001, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Structure guided fusion for depth map inpainting
Fei Qi 0001, Junyu Han, Pengjin Wang, Guangming Shi, Fu Li 0002 |
Pattern Recognit. Lett. | 4 |
| 2013 | Nonlocal Image Restoration With Bilateral Variance Estimation: A Low-Rank ApproachabstractSimultaneous sparse coding (SSC) or nonlocal image representation has shown great potential in various low-level vision tasks, leading to several state-of-the-art image restoration techniques, including BM3D and LSSC. However, it still lacks a physically plausible explanation about why SSC is a better model than conventional sparse coding for the class of natural images. Meanwhile, the problem of sparsity optimization, especially when tangled with dictionary learning, is computationally difficult to solve. In this paper, we take a low-rank approach toward SSC and provide a conceptually simple interpretation from a bilateral variance estimation perspective, namely that singular-value decomposition of similar packed patches can be viewed as pooling both local and nonlocal information for estimating signal variances. Such perspective inspires us to develop a new class of image restoration algorithms called spatially adaptive iterative singular-value thresholding (SAIST). For noise data, SAIST generalizes the celebrated BayesShrink from local to nonlocal models; for incomplete data, SAIST extends previous deterministic annealing-based solution to sparsity optimization through incorporating the idea of dictionary learning. In addition to conceptual simplicity and computational efficiency, SAIST has achieved highly competent (often better) objective performance compared to several state-of-the-art methods in image denoising and completion experiments. Our subjective quality results compare favorably with those obtained by existing techniques, especially at high noise levels and with a large amount of missing data. Weisheng Dong, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 2 |
| 2013 | Sparse Representation Based Image Interpolation With Nonlocal Autoregressive ModelingabstractSparse representation is proven to be a promising approach to image super-resolution, where the low-resolution (LR) image is usually modeled as the down-sampled version of its high-resolution (HR) counterpart after blurring. When the blurring kernel is the Dirac delta function, i.e., the LR image is directly down-sampled from its HR counterpart without blurring, the super-resolution problem becomes an image interpolation problem. In such cases, however, the conventional sparse representation models (SRM) become less effective, because the data fidelity term fails to constrain the image local structures. In natural images, fortunately, many nonlocal similar patches to a given patch could provide nonlocal constraint to the local structure. In this paper, we incorporate the image nonlocal self-similarity into SRM for image interpolation. More specifically, a nonlocal autoregressive model (NARM) is proposed and taken as the data fidelity term in SRM. We show that the NARM-induced sampling matrix is less coherent with the representation dictionary, and consequently makes SRM more effective for image interpolation. Our extensive experimental results demonstrate that the proposed NARM-based image interpolation method can effectively reconstruct the edge structures and suppress the jaggy/ringing artifacts, achieving the best image interpolation results so far in terms of PSNR as well as perceptual quality metrics such as SSIM and FSIM. Weisheng Dong, Lei Zhang 0006, Rastislav Lukac, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2013 | Nonlocally Centralized Sparse Representation for Image RestorationabstractSparse representation models code an image patch as a linear combination of a few atoms chosen out from an over-complete dictionary, and they have shown promising results in various image restoration applications. However, due to the degradation of the observed image (e.g., noisy, blurred, and/or down-sampled), the sparse representations by conventional models may not be accurate enough for a faithful reconstruction of the original image. To improve the performance of sparse representation-based image restoration, in this paper the concept of sparse coding noise is introduced, and the goal of image restoration turns to how to suppress the sparse coding noise. To this end, we exploit the image nonlocal self-similarity to obtain good estimates of the sparse coding coefficients of the original image, and then centralize the sparse coding coefficients of the observed image to those estimates. The so-called nonlocally centralized sparse representation (NCSR) model is as simple as the standard sparse representation model, while our extensive experiments on various types of image restoration problems, including denoising, deblurring and super-resolution, validate the generality and state-of-the-art performance of the proposed NCSR algorithm. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 3 |
| 2013 | Perceptual Quality Metric With Internal Generative MechanismabstractObjective image quality assessment (IQA) aims to evaluate image quality consistently with human perception. Most of the existing perceptual IQA metrics cannot accurately represent the degradations from different types of distortion, e.g., existing structural similarity metrics perform well on content-dependent distortions while not as well as peak signal-to-noise ratio (PSNR) on content-independent distortions. In this paper, we integrate the merits of the existing IQA metrics with the guide of the recently revealed internal generative mechanism (IGM). The IGM indicates that the human visual system actively predicts sensory information and tries to avoid residual uncertainty for image perception and understanding. Inspired by the IGM theory, we adopt an autoregressive prediction algorithm to decompose an input scene into two portions, the predicted portion with the predicted visual content and the disorderly portion with the residual content. Distortions on the predicted portion degrade the primary visual information, and structural similarity procedures are employed to measure its degradation; distortions on the disorderly portion mainly change the uncertain information and the PNSR is employed for it. Finally, according to the noise energy deployment on the two portions, we combine the two evaluation results to acquire the overall quality score. Experimental results on six publicly available databases demonstrate that the proposed metric is comparable with the state-of-the-art quality metrics. Jinjian Wu, Weisi Lin, Guangming Shi, Anmin Liu |
IEEE Trans. Image Process. | 3 |
| 2013 | Pattern Masking Estimation in Image With Structural UncertaintyabstractA model of visual masking, which reveals the visibility of stimuli in the human visual system (HVS), is useful in perceptual based image/video processing. The existing visual masking function mainly considers luminance contrast, which always overestimates the visibility threshold of the edge region and underestimates that of the texture region. Recent research on visual perception indicates that the HVS is sensitive to orderly regions that possess regular structures and insensitive to disorderly regions that possess uncertain structures. Therefore, structural uncertainty is another determining factor on visual masking. In this paper, we introduce a novel pattern masking function based on both luminance contrast and structural uncertainty. Through mimicking the internal generative mechanism of the HVS, a prediction model is firstly employed to separate out the unpredictable uncertainty from an input image. In addition, an improved local binary pattern is introduced to compute the structural uncertainty. Finally, combining luminance contrast with structural uncertainty, the pattern masking function is deduced. Experimental result demonstrates that the proposed pattern masking function outperforms the existing visual masking function. Furthermore, we extend the pattern masking function to just noticeable difference (JND) estimation and introduce a novel pixel domain JND model. Subjective viewing test confirms that the proposed JND model is more consistent with the HVS than the existing JND models. Jinjian Wu, Weisi Lin, Guangming Shi, Xiaotian Wang 0001, Fu Li 0002 |
IEEE Trans. Image Process. | 3 |
| 2013 | Reduced-Reference Image Quality Assessment With Visual Information FidelityabstractReduced-reference (RR) image quality assessment (IQA) aims to use less data about the reference image and achieve higher evaluation accuracy. Recent research on brain theory suggests that the human visual system (HVS) actively predicts the primary visual information and tries to avoid the residual uncertainty for image perception and understanding. Therefore, the perceptual quality relies to the information fidelities of the primary visual information and the residual uncertainty. In this paper, we propose a novel RR IQA index based on visual information fidelity. We advocate that distortions on the primary visual information mainly disturb image understanding, and distortions on the residual uncertainty mainly change the comfort of perception. We separately compute the quantities of the primary visual information and the residual uncertainty of an image. Then the fidelities of the two types of information are separately evaluated for quality assessment. Experimental results demonstrate that the proposed index uses few data (30 bits) and achieves high consistency with human perception. Jinjian Wu, Weisi Lin, Guangming Shi, Anmin Liu |
IEEE Trans. Multim. | 3 |
| 2013 | Just Noticeable Difference Estimation for Images With Free-Energy PrincipleabstractIn this paper, we introduce a novel just noticeable difference (JND) estimation model based on the unified brain theory, namely the free-energy principle. The existing pixel-based JND models mainly consider the orderly factors and always underestimate the JND threshold of the disorderly region. Recent research indicates that the human visual system (HVS) actively predicts the orderly information and avoids the residual disorderly uncertainty for image perception and understanding. Thus, we suggest that there exists disorderly concealment effect which results in high JND threshold of the disorderly region. Beginning with the Bayesian inference, we deduce an autoregressive model to imitate the active prediction of the HVS. Then, we estimate the disorderly concealment effect for the novel JND model. Experimental results confirm that the proposed JND model outperforms the relevant existing ones. Furthermore, we apply the proposed JND model in image compression, and around 15% of bit rate can be reduced without jeopardizing the perceptual quality. Jinjian Wu, Guangming Shi, Weisi Lin, Anmin Liu, Fei Qi 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Joint rate-distortion optimization for H.264/AVC intra coding based on cluster computingabstractThis paper presents a joint rate-distortion optimization coding scheme for H.264/AVC intra coding, which uses cluster computing technologies to jointly optimize coding parameters among different image blocks. We identify two kinds of dependencies, inter sub-block dependency and inter macroblock dependency, ignored in the ordinary rate-distortion optimization and a joint RD cost is designed in the optimization process to consider those dependencies. We show that with our proposed scheme, a significant better coding performance can be achieved with the standard-compliant bitstreams. Experiment results show up to 5% bits saving compared to the H.264/AVC reference software. Jizheng Xu, Feng Wu 0001, Guangming Shi |
ISCAS | 4 |
| 2012 | Surveillance video coding via low-rank and sparse decompositionabstractSurveillance videos are usually with a static or gradually changed background. The state-of-the-art block-based codec, H.264/AVC, is not sufficiently efficient for encoding surveillance videos since it cannot exploit the strong background temporal redundancy in a global manner. In this paper, motivated by the recent advance on low-rank and sparse decomposition (LRSD), we propose to apply it for the compression of surveillance videos. In particular, the LRSD is employed to decompose a surveillance video into the low-rank component, representing the background, and the sparse component, representing the moving objects. Then, we design different coding methods for the two different components. We represent the frames of the background by very few independent frames based on their linear dependency, which dramatically removes the temporal redundancy. Experimental results show that, for the compression of surveillance videos, the proposed scheme can significantly outperform H.264/AVC, up to 3 dB PSNR gain, especially at relatively low bit rates. Chongyu Chen, Jianfei Cai 0001, Weisi Lin, Guangming Shi |
ACM Multimedia | 4 |
| 2012 | Color demosaicking with an image formation model and adaptive PCA
Dahua Gao, Xiaolin Wu 0001, Guangming Shi, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Model-based adaptive resolution upconversion of degraded images
Xiaolin Wu 0001, Xiangjun Zhang, Guangming Shi |
J. Vis. Commun. Image Represent. | 4 |
| 2012 | Self-similarity based structural regularity for just noticeable difference estimation
Jinjian Wu, Fei Qi 0001, Guangming Shi |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Non-local spatial redundancy reduction for bottom-up saliency estimation
Jinjian Wu, Fei Qi 0001, Guangming Shi, Yongheng Lu |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Robust ISAR imaging based on compressive sensing from noisy measurements
Guanghui Zhao 0003, Guangming Shi, Fangfang Shen |
Signal Process. | 4 |
| 2012 | Image reconstruction with locally adaptive sparsity and nonlocal robust regularization
Weisheng Dong, Guangming Shi, Xin Li 0005, Lei Zhang 0006, Xiaolin Wu 0001 |
Signal Process. Image Commun. | 2 |
| 2012 | Edge-Based Perceptual Image CodingabstractWe develop a novel psychovisually motivated edge-based low-bit-rate image codec. It offers a compact description of scale-invariant second-order statistics of natural images, the preservation of which is crucial to the perceptual quality of coded images. Although being edge based, the codec does not explicitly code the edge geometry. To save bits on edge descriptions, a background layer of the image is first coded and transmitted, from which the decoder estimates the trajectories of significant edges. The edge regions are then refined by a residual coding technique based on edge dilation and sequential scanning in the edge direction. Experimental results show that the new image coding technique outperforms the existing ones in both objective and perceptual quality, particularly at low bit rates. Xiaolin Wu 0001, Guangming Shi, Xiaotian Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Binned Progressive Quantization for Compressive SensingabstractCompressive sensing (CS) has been recently and enthusiastically promoted as a joint sampling and compression approach. The advantages of CS over conventional signal compression techniques are architectural: the CS encoder is made signal independent and computationally inexpensive by shifting the bulk of system complexity to the decoder. While these properties of CS allow signal acquisition and communication in some severely resource-deprived conditions that render conventional sampling and coding impossible, they are accompanied by rather disappointing rate-distortion performance. In this paper, we propose a novel coding technique that rectifies, to a certain extent, the problem of poor compression performance of CS and, at the same time, maintains the simplicity and universality of the current CS encoder design. The main innovation is a scheme of progressive fixed-rate scalar quantization with binning that enables the CS decoder to exploit hidden correlations between CS measurements, which was overlooked in the existing literature. Experimental results are presented to demonstrate the efficacy of the new CS coding technique. Encouragingly, on some test images, the new CS technique matches or even slightly outperforms JPEG. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2012 | Model-Assisted Adaptive Recovery of Compressed Sensing with Imaging ApplicationsabstractIn compressive sensing (CS), a challenge is to find a space in which the signal is sparse and, hence, faithfully recoverable. Since many natural signals such as images have locally varying statistics, the sparse space varies in time/spatial domain. As such, CS recovery should be conducted in locally adaptive signal-dependent spaces to counter the fact that the CS measurements are global and irrespective of signal structures. On the contrary, existing CS reconstruction methods use a fixed set of bases (e.g., wavelets, DCT, and gradient spaces) for the entirety of a signal. To rectify this problem, we propose a new framework for model-guided adaptive recovery of compressive sensing (MARX) and show how a 2-D piecewise autoregressive model can be integrated into the MARX framework to make CS recovery adaptive to spatially varying second order statistics of an image. In addition, MARX offers a mechanism of characterizing and exploiting structured sparsities of natural images, greatly restricting the CS solution space. Simulation results over a wide range of natural images show that the proposed MARX technique can improve the reconstruction quality of existing CS methods by 2-7 dB. Xiaolin Wu 0001, Weisheng Dong, Xiangjun Zhang, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2011 | Enhancing Gradient Sparsity for Parametrized Motion EstimationabstractIn this paper, we propose a novel motion estimation framework based on the sparsity associated with gradients of the parametrized motion field. Beginning with Shen and Wu’s sparse model for optic flow estimation [15], we show the sparsity of the motion field can be enhanced by increasing the degree of freedom of the parametrized motion model. With such an enhancement, we formulate the motion estimation as an ‘0 optimization problem. Along with an ‘1 norm regularization to the instant constancy assumption, this problem is solved by a reweighted ‘1 optimization approach. Experiments on constant, pure translational, and affine motion models certify that the enhanced sparsity provides improved accuracy for motion estimation. Junyu Han, Fei Qi 0001, Guangming Shi |
BMVC | 3 |
| 2011 | Sparsity-based image denoising via dictionary learning and structural clusteringabstractWhere does the sparsity in image signals come from? Local and nonlocal image models have supplied complementary views toward the regularity in natural images - the former attempts to construct or learn a dictionary of basis functions that promotes the sparsity; while the latter connects the sparsity with the self-similarity of the image source by clustering. In this paper, we present a variational framework for unifying the above two views and propose a new denoising algorithm built upon clustering-based sparse representation (CSR). Inspired by the success of l1-optimization, we have formulated a double-header l1-optimization problem where the regularization involves both dictionary learning and structural structuring. A surrogate-function based iterative shrinkage solution has been developed to solve the double-header l1-optimization problem and a probabilistic interpretation of CSR model is also included. Our experimental results have shown convincing improvements over state-of-the-art denoising technique BM3D on the class of regular texture images. The PSNR performance of CSR denoising is at least comparable and often superior to other competing schemes including BM3D on a collection of 12 generic natural images. Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi |
CVPR | 4 |
| 2011 | Progressive Quantization of Compressive Sensing MeasurementsabstractCompressive sensing (CS) is recently and enthusiastically promoted as a joint sampling and compression approach. The advantages of CS over conventional signal compression techniques are architectural: the CS encoder is made signal independent and computationally inexpensive by shifting the bulk of system complexity to the decoder. While these properties of CS allow signal acquisition and communication in some severely resource-deprived conditions that render conventional sampling and coding impossible, they are accompanied by rather disappointing rate-distortion performance. In the present work we propose a novel coding technique that rectifies, to certain extent, the problem of poor compression performance of CS and at the same time maintains the simplicity and universality of the current CS encoder design. The main innovation is a scheme of progressive fixed-rate scalar quantization with binning that enables the CS decoder to exploit hidden correlations between CS measurements, which was overlooked in the existing literature. Experimental results are presented to demonstrate the efficacy of the new CS coding technique. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
DCC | 3 |
| 2011 | Centralized sparse representation for image restorationabstractThis paper proposes a novel sparse representation model called centralized sparse representation (CSR) for image restoration tasks. In order for faithful image reconstruction, it is expected that the sparse coding coefficients of the degraded image should be as close as possible to those of the unknown original image with the given dictionary. However, since the available data are the degraded (noisy, blurred and/or down-sampled) versions of the original image, the sparse coding coefficients are often not accurate enough if only the local sparsity of the image is considered, as in many existing sparse representation models. To make the sparse coding more accurate, a centralized sparsity constraint is introduced by exploiting the nonlocal image statistics. The local sparsity and the nonlocal sparsity constraints are unified into a variational framework for optimization. Extensive experiments on image restoration validated that our CSR model achieves convincing improvement over previous state-of-the-art methods. Weisheng Dong, Lei Zhang 0006, Guangming Shi |
ICCV | 3 |
| 2011 | Sparsity-based image deblurring with locally adaptive and nonlocally robust regularizationabstractImportant structures in photographic images such as edges and textures are jointly characterized by local variation and nonlocal invariance (similarity). Both of them provide valuable heuristics to the regularization of image restoration process. In this pa per, we propose to explore two sets of complementary ideas: 1) locally learn PCA-based dictionaries and estimate the sparsity regularization parameters for each coefficient; and 2) nonlocally enforce the invariance constraint by introducing a patch-similarity based term into the cost functional. The minimization of this new cost functional leads to an iterative thresholding-based image deblurring algorithm and its efficient implementation is discussed. Our experimental results have shown that the proposed scheme significantly outperforms several leading deblurring techniques in the literature on both objective and visual quality assessments. Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi |
ICIP | 4 |
| 2011 | Gradient sparsity for piecewise continuous optical flow estimationabstractThis paper introduces a new sparse model for robust and reliable optical flow estimation. We show that the sparsity is directly related to gradient fields of optical flow. According to theory on sparse signal recovery, we rigorously formulate the optical flow estimation as an ℓ0optimization problem. Considering the piecewise continuous nature of motion, the basic optical flow constraint is regularized by an ℓ1norm. Then, with convex relaxation, the solution is obtain via the ordinary or reweighted ℓ1optimization approach. Experimental results show the proposed method performs better than traditional methods and deals well with the discontinuities on motion boundaries. Junyu Han, Fei Qi 0001, Guangming Shi |
ICIP | 3 |
| 2011 | An efficient VLSI architecture for 4×4 intra prediction in the High Efficiency Video Coding (HEVC) standardabstractIntra prediction with fine directions is a critical feature in the new High Efficiency Video Coding (HEVC) standard because it provides significant performance gain. Different from the intra prediction in the H.264/AVC, this approach is more complicated in terms of computation and memory access, which makes the VLSI design very difficult. In this paper, we propose an efficient uniform architecture for all of the 4×4 intra directional modes. The architecture is implemented by a register array and a flexible reference sample selection technique. This novel architecture does not need to project the samples from the side reference to the main reference. Thus, it reduces the processing latency and the number of registers considerably. The proposed architecture has been implemented with TSMC 0.13μm CMOS technology. Simulation results show that the proposed architecture only needs 9020 logic gates for 17 directional modes and can run at 150 MHz operation frequency. Fu Li 0002, Guangming Shi, Feng Wu 0001 |
ICIP | 2 |
| 2011 | Adaptive patch matching for motion compensated predictionabstractMotion compensated prediction (MCP) plays an important role in video coding due to its great capability of reducing temporal redundancy. In this paper, we propose a new MCP scheme by adaptive patch matching with the full use of the reconstructed pixels surrounding the current block (referred to as the template inside the patch) aiming at achieving a more accurate prediction than conventional MCP. The proposed scheme not only takes advantage of the temporal correlation but also efficiently exploits the spatial correlation between the current block and its template inside the patch. An adaptive linear combination of the current block and its template in motion estimation is designed to generate an optimal prediction while maintaining the local variation of the current block. Accordingly, a modification of the rate-distortion criterion is introduced to select the combined prediction. Experimental results show that our proposed APM achieves improved coding performance compared with H.264/AVC. Tianmi Chen, Xiaoyan Sun 0001, Feng Wu 0001, Guangming Shi |
ISCAS | 4 |
| 2011 | A scheme of parallel arithmetic codingabstractThis paper presents a parallel arithmetic coding scheme in which supports a large degree of parallelism with a marginal cost in terms of the coding efficiency. The parallelism is brought by coding the bits using multiple arithmetic coders. We identify two types of losses in coding efficiency by breaking the dependency among the data: the loss by breaking the probability prediction process and the loss by breaking the information for context decision at the starting of each slice. We further analyze these losses quantitatively and find that the loss by breaking probability adaptation process takes the most of losses. A coding method is proposed in this paper to compensate such a loss by sending additional data. Experimental results show that our proposed method can compensate the loss well, especially for large scale parallel, whereas the overhead is moderate. Jizheng Xu, Guangming Shi |
ISCAS | 4 |
| 2011 | A pipelined architecture for 4×4 intra frame mode decision in the high efficiency video codingabstractMode decision in High Efficient Video Coding (HEVC) is occupied more than half of the computational complexity in intra frame coding. Block size of 4×4 is the most frequently used block in HM. In this paper, we proposed a pipelined architecture for the 4×4 intra frame mode decision in HEVC to improve the computational capability. This novel architecture consists of six-stage pipelines, and each of the pipelines can be accomplished within 24 clock cycles. In the pipeline of prediction procedure, we proposed a folded project-skip architecture for prediction. It can save the processing latency and the registers considerably. We also proposed a simplified CAVLC with low complexity in the pipeline of bits estimation procedure. The architecture for mode decision has been evaluated with TSMC 0.13μm CMOS technology. Synthesized results show that the proposed architecture only needs 99K logic gates for modes decision and can run at 165 MHz operation frequency. Fu Li 0002, Guangming Shi |
MMSP | 2 |
| 2011 | Observation quality guaranteed layout of camera networks via sparse representationabstractNodes layout is a critical issue affecting the quality of service of camera networks. The layout aims at adopting the least number of cameras to cover the whole scene with expected observation quality. This is generally formulated as a non-convex optimization problem which is hard to be solved in polynomial time. In this paper, we propose an efficient near-optimal convex solution for layout guaranteeing satisfactory observation quality based on an anisotropic sensing model of camera. We construct the sensing model based on the physical imaging process. This model provides an accurate measurement of observation quality for the directional nonuniform sensing field of cameras. The layout is treated as selecting the best subset of cameras from a redundant initial layout with numerous cameras. Based on this characteristic and the sensing model, we formulate the layout as a sparse ℓ0problem which is further relaxed to a convex ℓ1minimization. Therefore, the near-optimal layout is efficiently obtained via convex optimization. Simulation results confirm the effectiveness of the proposed approaches. Fei Qi 0001, Guangming Shi |
VCIP | 3 |
| 2011 | Multiple description video coding against both erasure and bit errors by compressive sensingabstractWe propose a novel multiple description video coding (MDVC) technique for robust video transmission via lossy networks of both packet erasure and bit errors. The new MDVC technique is designed to meet two objectives: ultra fine description granularity and low encoder complexity, which allow resource-deprived video transmitters (e.g., smart phones) to operate in time-varying adverse network conditions. These design goals are met by the signal acquisition and coding strategy of compressive sensing, and by locally adaptive sparse representation of video signals. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
VCIP | 3 |
| 2011 | Image Deblurring and Super-Resolution by Adaptive Sparse Domain Selection and Adaptive RegularizationabstractAs a powerful statistical image modeling technique, sparse representation has been successfully used in various image restoration applications. The success of sparse representation owes to the development of the l(1)-norm optimization techniques and the fact that natural images are intrinsically sparse in some domains. The image restoration quality largely depends on whether the employed sparse domain can represent well the underlying image. Considering that the contents can vary significantly across different images or different patches in a single image, we propose to learn various sets of bases from a precollected dataset of example image patches, and then, for a given patch to be processed, one set of bases are adaptively selected to characterize the local sparse domain. We further introduce two adaptive regularization terms into the sparse representation framework. First, a set of autoregressive (AR) models are learned from the dataset of example image patches. The best fitted AR models to a given patch are adaptively selected to regularize the image local structures. Second, the image nonlocal self-similarity is introduced as another regularization term. In addition, the sparsity regularization parameter is adaptively estimated for better image restoration performance. Extensive experiments on image deblurring and super-resolution validate that by using adaptive sparse domain selection and adaptive regularization, the proposed method achieves much better results than many state-of-the-art algorithms in terms of both PSNR and visual perception. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2011 | Nonuniform Directional Filter Banks With Arbitrary Frequency PartitioningabstractDirectional filter banks (DFBs) are highly desired in directional representation of images. In this correspondence, we propose a 2-D nonsubsampled nonuniform directional filter bank (NUDFB) and its design method. The proposed NUDFB has nonuniform wedge-shaped subbands and allows arbitrary frequency partitioning schemes. It can extract directional information according to the directional distribution of images. This attractive advantage cannot be achieved by the existing directional transforms. The design method of the proposed NUDFB is based upon the pseudopolar Fourier transform. By utilizing the geometry property of the pseudopolar grid, we employ a 1-D nonsubsampled nonuniform filter bank to obtain a set of nonuniform wedge-shaped subbands. During the design process, only 1-D operations are involved and, thus, the difficulty encountered in the design of 2-D fan filters is avoided. To demonstrate the potential of the proposed NUDFB, an example on image directional decomposition is given. Lili Liang, Guangming Shi, Xuemei Xie |
IEEE Trans. Image Process. | 2 |
| 2011 | High-Resolution Imaging Via Moving Random Exposure and Its SimulationabstractIn this correspondence, we introduce a new imaging method to obtain high-resolution (HR) images. The image acquisition is performed in two stages, compressive measurement and optimization reconstruction. In order to reconstruct HR images by a small number of sensors, compressive measurements are made. Specifically, compressive measurements are made by a low-resolution (LR) camera with randomly fluttering shutter, which can be viewed as a moving random exposure pattern. In the optimization reconstruction stage, the HR image is computed by different models according to the prior knowledge of scenes. The proposed imaging method offers a new way of acquiring HR images of essentially static scenes when the camera resolution is limited by severe constraints such as cost, battery capacity, memory space, transmission bandwidth, etc. and when the prior knowledge of scenes is available. The simulation results demonstrate the effectiveness of the proposed imaging method. Guangming Shi, Dahua Gao, Xiaoxia Song, Xuemei Xie, Danhua Liu |
IEEE Trans. Image Process. | 1 |
| 2010 | Image denoising based on translation invariant Directional LiftingabstractAdaptive Directional Lifting (ADL) has been successfully implemented in image compression and denoising due to the feature of simple structure and flexible directional selectivity. However, image denoising by means of ADL introduces many visual artifacts caused by Gibbs phenomena due to the lack of translation invariance. In this paper, we propose a translation invariant directional lifting (TI-DL) by employing the cycle-spinning based technique to reduce artifacts in denoising results. Moreover, the inefficiency and high computational complexity of the orientation estimation technique in ADL strongly influences the performance. In order to achieve better denoising results, in this paper, 2-D Gabor filters are adopted for orientation estimation to achieve better orientation estimation results with lower complexity. Experimental results demonstrate that the proposed method achieves state-of-art denoising performance in terms of both objective (PSNR) and subjective (SSIM) evaluation. Xiaotian Wang 0001, Guangming Shi, Lili Liang |
ICASSP | 2 |
| 2010 | An improved model of pixel adaptive just-noticeable difference estimationabstractA pixel-wise adaptive model for estimating the just-noticeable difference (JND) in spatial domain is proposed in this paper. As the human visual system (HVS) can be considered as a multichannel system, we assume that there exist two channels in the HVS, which deliver luminance adaption factor and texture masking factor, respectively. Both channels affect the JND threshold in a cooperative manner. The texture regions are with abundant redundancy and can tolerate much noise. The disorder degree and spatial masking of the texture are considered to estimate the texture masking effect, for deducing such JND threshold that coincides with the HVS. Finally, the luminance adaptation factor and texture masking factor are combined nonlinearly. Various experiments confirm the improved model has a better visual effect than models proposed before. Jinjian Wu, Fei Qi 0001, Guangming Shi |
ICASSP | 3 |
| 2010 | Intra frame coding with template matching prediction and adaptive transformabstractFor natural images, there are usually repeating similar contents but hard to be well predicted locally. Prediction using template matching is an effective technology to exploit such a non-local correlation. In this paper, we propose an alternative scheme to further exploit the non-local correlation. In the proposed scheme, template matching is also used to search for probable similar references to the current block to be coded. We then use these references to train an adaptive transform, which most likely reflects the statistical characteristic of the current block. The proposed scheme can further exploit correlation between the current block and more possible references. Compared to the scheme that only integrates prediction by template matching, the proposed scheme shows improvement about 0.45dB PSNR increase or 9.3% bit saving on average, which leads to 1dB's gain or 19.5% bit saving on average compared to the state-of-the-art scheme without using template matching. Cuiling Lan, Jizheng Xu, Feng Wu 0001, Guangming Shi |
ICIP | 4 |
| 2010 | Edge-based image coding at low bit-rateabstractWe propose a novel edge-based low bit-rate image codec. Although being edge based, the codec does not explicitly code the edge geometry. To avoid spending high bit budget on edges, a background layer of the image is first coded and transmitted, from which the decoder estimates the trajectories of significant edges. The edge regions are then refined by a residual coding technique based on edge dilation and sequential scanning in the edge direction. Experimental results show that the new image coding technique outperforms existing ones in both objective and perceptual quality, particularly at low bit rates. Xiaolin Wu 0001, Guangming Shi, Xiaotian Wang 0001 |
ICIP | 3 |
| 2010 | Color demosaicking with sparse representationsabstractColor demosaicking is an ill-posed inverse problem of image restoration. The performance of a color demosaicking algorithm depends on how thoroughly it can exploit domain knowledge to confine the solution space for the underlying true color image. We propose a sparsity-based ℓ1minimization technique for color demosaicking that exploits both interband and intra-band sparse representations of natural images. In some of most challenging cases of color demosaicking, the proposed technique outperforms those published in the literature by a significant margin in both PSNR and visual quality. Xiaolin Wu 0001, Dahua Gao, Guangming Shi, Danhua Liu |
ICIP | 3 |
| 2010 | Live demonstration: Spatial-temporal color video reproduction from noisy CFA sequence track: Digital signal processingabstractThis demonstration shows a spatial-temporal denoising and demosaicking scheme for noisy CFA videos. This scheme can significantly reduce the noise-caused color artifacts and effectively preserve the image edge structures. The experimental results showed that this scheme achieves promising color video reproduction in terms of both PSNR and visual perception. Lei Zhang 0006, Weisheng Dong, Chiu-Wai Hui, Xiaolin Wu 0001, Guangming Shi |
ISCAS | 5 |
| 2010 | Super-resolution with nonlocal regularized sparse representationabstractThe reconstruction of a high resolution (HR) image from its low resolution (LR) counterpart is a challenging problem. The recently developed sparse representation (SR) techniques provide new solutions to this inverse problem by introducing the l1-norm sparsity prior into the super-resolution reconstruction process. In this paper, we present a new SR based image super-resolution by optimizing the objective function under an adaptive sparse domain and with the nonlocal regularization of the HR images. The adaptive sparse domain is estimated by applying principal component analysis to the grouped nonlocal similar image patches. The proposed objective function with nonlocal regularization can be efficiently solved by an iterative shrinkage algorithm. The experiments on natural images show that the proposed method can reconstruct HR images with sharp edges from degraded LR images. Weisheng Dong, Guangming Shi, Lei Zhang 0006, Xiaolin Wu 0001 |
VCIP | 2 |
| 2010 | Efficient architecture for adaptive directional lifting-based wavelet transformabstractAdaptive direction lifting-based wavelet transform (ADL) has better performance than conventional lifting both in image compression and de-noising. However, no architecture has been proposed to hardware implement it because of its high computational complexity and huge internal memory requirements. In this paper, we propose a four-stage pipelined architecture for 2 Dimensional (2D) ADL with fast computation and high data throughput. The proposed architecture comprises column direction estimation, column lifting, row direction estimation and row lifting which are performed in parallel in a pipeline mode. Since the column processed data is transposed, the row processor can reuse the column processor which can decrease the design complexity. In the lifting step, predict and update are also performed in parallel. For an 8×8 image sub-block, the proposed architecture can finish the ADL forward transform within 78 clock cycles. The architecture is implemented on Xilinx Virtex5 device on which the frequency can achieve 367 MHz. The processed time is 212.5 ns, which can meet the request of real-time system. Zan Yin, Li Zhang 0004, Guangming Shi |
VCIP | 3 |
| 2010 | Two-stage image denoising by principal component analysis with local pixel grouping
Lei Zhang 0006, Weisheng Dong, David Zhang 0001, Guangming Shi |
Pattern Recognit. | 4 |
| 2010 | Lossy-to-lossless image compression based on multiplier-less reversible integer time domain lapped transform
Lei Wang 0018, Licheng Jiao, Jiaji Wu, Guangming Shi, Yanjun Gong |
Signal Process. Image Commun. | 4 |
| 2010 | Morphological dilation image coding with context weights prediction
Jiaji Wu, Anand Paul 0001, Yong Fang 0001, Jechang Jeong, Licheng Jiao, Guangming Shi |
Signal Process. Image Commun. | 7 |
| 2010 | Spatial-Temporal Color Video Reconstruction From Noisy CFA SequenceabstractSingle-sensor digital video cameras use a color filter array (CFA) to capture video and a color demosaicking (CDM) procedure to reproduce the full color sequence. The reproduced video frames suffer from the inevitable sensor noise introduced in the video acquisition process. This paper presents a spatial-temporal denoising and demosaicking scheme that works without explicit motion estimation. We first perform patch based denoising on the mosaic CFA video. For each CFA patch to be denoised, similar patches are selected within a local spatial-temporal neighborhood. The principal component analysis is performed on the selected patches to remove noise. We then apply an initial single-frame CDM to the denoised CFA data, and subsequently post-process the demosaicked frames by exploiting the spatial-temporal redundancy to reduce the color artifacts. The experimental results on simulated and real noisy CFA sequences demonstrate that the proposed spatial-temporal CFA video denoising and demosaicking scheme can significantly reduce the noise-caused color artifacts and effectively preserve the image edge structures. Lei Zhang 0006, Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Compress Compound Images in H.264/MPGE-4 AVC by Exploiting Spatial CorrelationabstractCompound images are a combination of text, graphics and natural image. They present strong anisotropic features, especially on the text and graphics parts. These anisotropic features often render conventional compression inefficient. Thus, this paper proposes a novel coding scheme from the H.264 intraframe coding. In the scheme, two new intramodes are developed to better exploit spatial correlation in compound images. The first is the residual scalar quantization (RSQ) mode, where intrapredicted residues are directly quantized and coded without transform. The second is the base colors and index map (BCIM) mode that can be viewed as an adaptive color quantization. In this mode, an image block is represented by several representative colors, referred to as base colors, and an index map to compress. Every block selects its coding mode from two new modes and the previous intramodes in H.264 by rate-distortion optimization (RDO). Experimental results show that the proposed scheme improves the coding efficiency even more than 10 dB at most bit rates for compound images and keeps a comparable efficient performance to H.264 for natural images. Cuiling Lan, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | Extracting regions of attention by imitating the human visual systemabstractDetecting and segmenting out the regions of interest (ROIs) is one of the foundations in image processing and analysis. Because the final information sink of images is human, for segmenting out the ROIs effectively, we need to study human visual system (HVS) and imitate the behaviors when human viewing a scene. Researchers have found several factors which affect human attentions by studying eye movements when one views an image. In this paper, a method is proposed to detect the ROIs automatically based on HVS. In the proposed algorithm, the properties of pixels such as the contrast, location and edges are analyzed, and the pixels are enhanced according to the sensitivity of HVS. Then these factors are combined to a salient map, which classifies each pixel of the image in relation to its perceptual importance. Finally, the ROIs are segmented according to the salient map. This algorithm is easy to work, and can segment the objects from complex background efficiently. Fei Qi 0001, Jinjian Wu, Guangming Shi |
ICASSP | 3 |
| 2009 | Context-based bias removal of statistical models of wavelet coefficients for image denoisingabstractExisting wavelet-based image denoising techniques all assume a probability model of wavelet coefficients that has zero mean, such as zero-mean Laplacian, Gaussian, or generalized Gaussian distributions. While such a zero-mean probability model fits a wavelet subband well, in areas of edges and textures the distribution of wavelet coefficients exhibits a significant bias. We propose a context modeling technique to estimate the expectation of each wavelet coefficient conditioned on the local signal structure. The estimated expectation is then used to shift the probability model of wavelet coefficient back to zero. This bias removal technique can significantly improve the performance of existing wavelet-based image denoisers. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi, Lei Zhang 0006 |
ICIP | 3 |
| 2009 | Nonlocal back-projection for adaptive image enlargementabstractThis paper presents a novel non-local iterative back-projection (NLIBP) algorithm for image enlargement. The iterative back-projection (IBP) technique iteratively reconstructs a high resolution (HR) image from its blurred and downsampled low resolution (LR) counterpart. However, the conventional IBP methods often produce many ¿jaggy¿ and ¿ringing¿ artifacts because the reconstruction errors are back projected into the reconstructed image isotropically and locally. In natural images, usually there exist many non-local redundancies which can be exploited to improve the image reconstruction quality. Therefore, we propose to incorporate adaptively the non-local information into the IBP process so that the reconstruction errors can be reduced. Experimental results demonstrated that the proposed NLBP can reconstruct faithfully the HR images with sharp edges and texture structures. It outperforms the state-of-the-art methods in both PSNR and visual perception. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xiaolin Wu 0001 |
ICIP | 3 |
| 2009 | LDA based color information fusion for visual objects trackingabstractIn this paper, an approach for object tracking is proposed based on the online color information fusion scheme. The fusion scheme, is performed by projecting multi-channel color images to a one-dimensional pseudo gray scale space. This dimensionality reduction simplifies the designation of the tracking algorithm. The fusing coefficients are determined by taking the Fisher linear discriminant analysis to maximize the discriminative capability after taking the projection. The robustness of the approach lies in exploiting the appearance discrimination of the object in cluttered scenarios with varying illumination conditions. This scheme is embedded into a mean-shift tracking system and the experimental results show our scheme can enhance the discriminant characters in the changing environment and hence producing robust tracking results. Fei Qi 0001, Guangming Shi |
ICIP | 3 |
| 2009 | High resoluton image reconstruction: A new imager via movable random exposureabstractRecently compressive sensor developed as an imager for capturing images effectively has been studied extensively. In this paper, we design a new imager to reconstruct high resolution image from a low resolution blurred image obtained by the intended movable random exposure. This imager grabs an image by moving a camera with a randomly fluttering shutter along a certain motion route. By analyzing this kind of movable random exposure process, we find it can be considered as compressive sampling described in the compressive sensing (CS) theory. Then according to the CS theory, the exposure result of this imager can be used to recover a high resolution image. Since this imager consists only a movable camera and a fluttered shutter, it is relatively simple and easy to implement. The simulation results show that the proposed imager can recover high even ultra-high resolution images with good reconstruction performance. Guangming Shi, Dahua Gao, Danhua Liu, Liangjun Wang |
ICIP | 1 |
| 2009 | 3D medical image compression based on multiplierless low-complexity RKLT and shape-adaptive wavelet transformabstractA multiplierless low complexity reversible integer Karhunen-Loe¿ve transform (Low-RKLT) is proposed based on matrix factorization. Conventional methods based on KLT suffer from high computational complexity and unability of applying in lossless medical image compression. To solve the two problems, multiplierless Low-RKLT is investigated using multi-lifting in this paper. Combined with ROI coding method, we have proposed a progressive lossy-to-lossless ROI compression method for three dimensional (3D) medical images with high performance. In our proposed method Low-RKLT is used for the inter-frame decorrelation after SA-DWT in the spatial domain. Simulation results show that, the proposed method performs much better in both lossless and lossy compression than 3D-DWT-based method. Lei Wang 0018, Jiaji Wu, Licheng Jiao, Guangming Shi |
ICIP | 4 |
| 2009 | Image compression with downsampling and overlapped transform at low bit ratesabstractThis paper proposes an image coding method based on adaptive downsampling which not only uses the pixel redundancy but also considers visual redundancy. At the encoder side, codec adaptively chooses some smooth regions of the original image to downsample, and then overlapped transform with selectivity, block DCT and adaptive-shape DCT (SA-DCT) are used against the image after being downsampled. For the incomplete transformed image, OB-SPECK is adopted to code. At the decoder side, in order to reduce the computational complexity, we use the simple cubic interpolation which not only is very suitable to the downsampled regions but also enhances greatly the real time of this coding system. Experimental results shows the proposed method outperforms JPEG2000, SPECK, SPIHT, and LT+SPECK at low bit rates. Jiaji Wu, Guangming Shi, Licheng Jiao |
ICIP | 3 |
| 2009 | Compress Compound Images in H.264/MPEG-4 AVC by Fully Exploiting Spatial CorrelationabstractCompound images consist of text, graphics and natural images, which present strong anisotropic features. It makes existing image coding standards inefficient on compressing them. To solve the problem, this paper proposes a novel coding scheme based on the H.264 intra-frame coding. Two new intra modes are proposed to better exploit spatial correlations in compound images. The first is residual scalar quantization (RSQ) mode, where intra-predicted residues are directly quantized and entropy coded. The second is base colors and index map (BCIM) mode that can be viewed as an adaptive vector quantization. In this mode, an image block is represented by several representative colors, called as base colors, and an index map to compress. Two new modes as well as previous intra modes in H.264 are selected by the rate-distortion optimization (RDO) method in each block. Experimental results show that the proposed scheme not only improves the coding efficiency even more than 10 dB for compound images but also keeps the similar performance as H.264 for natural images. Cuiling Lan, Feng Wu 0001, Guangming Shi |
ISCAS | 3 |
| 2009 | Edge-based dynamic ROI coding with standard complianceabstractAn edge-based region-of-interest (ROI) image coding technique is proposed for low bit-rate visual communication. An image is compactly coded into two semantic levels: background sketch and object textures. The background sketch is a down-sampled version of the input image. The decoder reconstructs the background first by adaptive interpolation and lets users identify the ROI, and then requests the ROI textures to be transmitted. A distinct advantage of the proposed ROI technique is a compact edge-based descriptor of natural object boundaries. The expensive ROI geometry computations are carried out at the encoder only on demand, keeping decoder complexity low to benefit wireless devices. The new system outperforms the dynamic ROI coding of JPEG 2000 in both visuality and PSNR. Furthermore, both background and texture coding can be made compliant with any existing compression standard. Xiaolin Wu 0001, Guangming Shi |
MMSP | 3 |
| 2009 | Learning-based recovery of compressive sensing with application in multiple description codingabstractThe recently proposed compressive sensing (CS) theory provides a new solution for multiple description coding (MDC) with fine granularity, by treating each random CS measurement as a description. The performance of CS-based MDC (CS-MDC) depends on the efficacy of the CS recovery algorithm. Existing CS recovery algorithms recover the signal in a fixed space (e.g., Wavelet, DCT, and gradient spaces) for the entire duration of the signal, even though a typical multimedia signal exhibits sparsity in time/space variant spaces. To rectify this problem and develop a better CS recovery algorithm for CSMDC, we propose a learning-based framework to conduct the CS recovery in locally adaptive spaces, and carry out a case study on image MDC. A set of prior image models are learned offline from a training set to facilitate the CS recovery in local adaptive bases. Experiments show that the learning-based CS recovery algorithm can significantly improve the performance of the previous CS-MDC technique in both PSNR and visual quality. Guangming Shi, Weisheng Dong, Xiaolin Wu 0001 |
MMSP | 2 |
| 2009 | Lossy-to-Lossless Hyperspectral Image Compression Based on Multiplierless Reversible Integer TDLT/KLTabstractWe proposed a new transform scheme of multiplierless reversible time-domain lapped transform and Karhunen-Loeve transform (RTDLT/KLT) for lossy-to-lossless hyperspectral image compression. Instead of applying discrete wavelet transform (DWT) in the spatial domain, RTDLT is applied for decorrelation. RTDLT can be achieved by existing discrete cosine transform and pre- and postfilters, while the reversible transform is guaranteed by a matrix factorization method. In the spectral direction, reversible integer low-complexity KLT is used for decorrelation. Owing to completely reversible transform, the proposed method can realize progressive lossy-to-lossless compression from a single embedded code-stream file. Numerical experiments on benchmark images show that the proposed transform scheme performs better than 5/3DWT-based methods in both lossy and lossless compressions, comparable with the optimal 9/7DWT-FloatKLT-based lossy compression method. Lei Wang 0018, Jiaji Wu, Licheng Jiao, Guangming Shi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2008 | Lossy to lossless image compression based on reversible integer DCTabstractA progressive image compression scheme is investigated using reversible integer discrete cosine transform (RDCT) which is derived from the matrix factorization theory. Previous techniques based on DCT suffer from bad performance in lossy image compression compared with wavelet image codec. And lossless compression methods such as IntDCT, I2I-DCT and so on could not compare with JPEG-LS or integer discrete wavelet transform (DWT) based codec. In this paper, lossy to lossless image compression can be implemented by our proposed scheme which consists of RDCT, coefficients reorganization, bit plane encoding, and reversible integer pre- and post-filters. Simulation results show that our method is competitive against JPEG-LS and JPEG2000 in lossless compression. Moreover, our method outperforms JPEG2000 (reversible 5/3 filter) for lossy compression, and the performance is even comparable with JPEG2000 which adopted irreversible 9/7 floating-point filter (9/7F filter). Lei Wang 0018, Jiaji Wu, Licheng Jiao, Li Zhang 0004, Guangming Shi |
ICIP | 5 |
| 2008 | Signal-adapted directional lifting scheme for image compressionabstractIn this paper, we propose a new adaptive lifting scheme that not only locally adapts the filtering directions to the orientations of image features, but also adapts the lifting filters to the statistic properties of image signal. The proposed approach refines previous adaptive directional lifting-based wavelet transform (ADL) by combining directional lifting and adaptive lifting filters to form a unified framework. The image signal is first segmented into regions of textures and edges with close directional features. The lifting filters are then effectively designed. The prediction step is designed to minimize the prediction error of the image signal, and the update step is designed to minimize the reconstruction error. Significant improvements on objective and subjective quality over conventional 2-D wavelet transform and previous ADL transform are achieved. Weisheng Dong, Guangming Shi, Jizheng Xu |
ISCAS | 2 |
| 2008 | A parallel sampling scheme for ultra-wideband signal based on the random projectionabstractHigh-rate analog-to-digital converters (ADC's) are ubiquitous and critical components in advance signal processing, such as software radio and ultra wideband (UWB) radar system. In this paper, we propose an alternative parallel sampling scheme, which is based on random projection, for an ultra- wideband analog to digital conversion. The sampling rate by using proposed scheme can be reduced to 1/m Nyquist sampling rate where m is the number of parallel channel. Compared with conventional ADCs based on pulse code modulation (PCM), the proposed method makes the sub-Nyquist sampling possible. And compared with other parallel sampling scheme methods, it can be easy implemented without any apriori knowledge about the signal. And it is very suitable for sampling UWB signals with short time support. Additionally, the detail implementation structure, performance analysis and simulation results of proposed method are given. Guangming Shi, Zhe Liu 0008, X. Y. Chen, Liangjun Wang |
ISCAS | 1 |
| 2008 | Trellis quantization for L∞-constrained compression with integer waveletsabstractAlthough integer wavelets have been successfully used in lossless signal compression, they have not been generalized to near-lossless (L∞ constrained) coding. This paper proposes a new technique of trellis quantization in the framework of lifting integer wavelet as a promising way of near-lossless signal compression. This new technique achieves continuous scalability of the wavelet code stream from highly lossy (low bit rates) to near-lossless (respecting a tight error bound on each decoded sample) reconstruction. Xiaolin Wu 0001, Guangming Shi |
ISIT | 3 |
| 2008 | A compressive sensing approach of multiple descriptions for network multimedia communicationabstractA new multiple description coding (MDC) approach is proposed based on the theory of compressive sensing (CS). The CS theory allows a signal to be reconstructed from a small number of its random measurements if the signal is sparse in some space. An attractive property of CS for MDC applications is that the reconstruction error only depends on the number but not on which of the transmitted measurements that are received. By treating each CS measurement as a description, we have a balanced MDC scheme with fine description granularity and low encoding complexity. Another advantage of the new MDC approach is that all signals can be coded the same but decoded in different spaces for better sparse reconstruction. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
MMSP | 3 |
| 2008 | Adaptive Nonseparable Interpolation for Image Compression With Directional Wavelet TransformabstractThe adaptive directional lifting-based wavelet transform (ADL) locally adapts the filtering directions to the local properties of the image. In this letter, instead of using the conventional interpolation filter for the directional prediction with fractional-pel accuracy, a new two-dimensional nonseparable adaptive interpolation filter is proposed. The adaptive filter is calculated for every fractional-pel direction so as to minimize the energy of the prediction error. The tradeoff between reducing the prediction error and the overhead to code the interpolation filter is discussed. This enables coding gains of up to 0.98 dB, compared to ADL coder, and up to 2.4 dB, compared to the JPEG 2000 for typical test images. Weisheng Dong, Guangming Shi, Jizheng Xu |
IEEE Signal Process. Lett. | 2 |
| 2007 | Immune memory clonal selection algorithms for designing stack filters
Weisheng Dong, Guangming Shi, Li Zhang 0004 |
Neurocomputing | 2 |