Guangming Shi

dblp:97/3742 · DBLP profile ↗
← Back
8ranked-venue papers in the field
0as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 4Big Data, Cloud & Distributed Data Systems · 3Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Dynamic Spatio-Temporal Compression Ratio Learning and Frequency-Aware Semantic Compression for Video Imaging
abstract
Snapshot compressive imaging (SCI) and video compressive sensing (VCS) typically use fixed, globally uniform compression ratios that ignore spatio-temporal heterogeneity. We present D-STCRL, a reinforcement-learned framework that unifies adaptive sensing and semantic transmission under an explicit rate-distortion-energy objective. The pipeline comprises: (i) a Ratio Generation Network predicts per-patch ratio maps via spatio-temporal attention and 3D frequency cues; (ii) a Programmable Sensing Model emulates pixel-wise variable exposure through differentiable binary gating under a global budget; and (iii) a Frequency-Aware Swin decoder with a low-rank prior restores temporally consistent frames. A multi-objective policy gradient couples the ratio policy with reconstruction and JSCC, yielding stable training. On the NFS benchmark, D-STCRL improves PSNR by$2-3 ~\text{dB}$over fixed-ratio SCI at the same sampling budget; under 10 dB AWGN it surpasses a CRL baseline by$0.6-1.0 ~\text{dB}$while reducing transmitted symbols by up to 15 %. These results unify content-adaptive sensing and efficient transmission for next-generation cameras. Our code, configs and reproducible pipelines will be released upon acceptance.
Haixiong Li, Dahua Gao, Xiaodan Song, Guangming Shi
DCC4
2025 Affine Transformation-Based Generative Face Video Compression
abstract
In this paper, we propose a generative face video compression framework based on affine transformations to better represent large movements without parameter transmission. It mainly consists of an encoder and decoder, and our encoder is similar to the one in [1]. Intra frame are compressed by the existing encoder, while subsequent inter frames are compressed into compact inter frame features. In the decoder, feature alignment is first established to map the decoded intra frame and inter frame features into the same domain. The aligned features are then combined with the appearance features extracted by the appearance encoder from the intra frame and fed into the coarse-fine affine transform module to establish motion estimation and compensation. The coarse affine transform focuses on global motion, while the fine affine transform deals with local motion, such as lip motion. Finally, the transformed features are fed into the image generation module to obtain the final reconstruction results.
Xihua Lin, Xiaodan Song, Xuguang Zuo, Dahua Gao, Xuemei Xie, Guangming Shi
DCC7
2021 Blind image quality prediction with hierarchical feature aggregation
Jinjian Wu, Wen Yang 0008, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
Inf. Sci.5
2020 No-reference quality index of depth images based on statistics of edge profiles for view synthesis
Leida Li, Jinjian Wu, Shiqi Wang 0001, Guangming Shi
Inf. Sci.5
2019 No-reference image quality assessment with visual pattern degradation
Jinjian Wu, Man Zhang 0007, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
Inf. Sci.5
2016 Orientation selectivity based visual pattern for reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Guangming Shi, Leida Li, Yuming Fang 0001
Inf. Sci.3
2015 Kinect Depth Recovery Using a Color-Guided, Region-Adaptive, and Depth-Selective Framework
abstract
Considering that the existing depth recovery approaches have different limitations when applied to Kinect depth data, in this article, we propose to integrate their effective features including adaptive support region selection, reliable depth selection, and color guidance together under an optimization framework for Kinect depth recovery. In particular, we formulate our depth recovery as an energy minimization problem, which solves the depth hole filling and denoising simultaneously. The energy function consists of a fidelity term and a regularization term, which are designed according to the Kinect characteristics. Our framework inherits and improves the idea of guided filtering by incorporating structure information and prior knowledge of the Kinect noise model. Through analyzing the solution to the optimization framework, we also derive a local filtering version that provides an efficient and effective way of improving the existing filtering techniques. Quantitative evaluations on our developed synthesized dataset and experiments on real Kinect data show that the proposed method achieves superior performance in terms of recovery accuracy and visual quality.
Chongyu Chen, Jianfei Cai 0001, Jianmin Zheng, Tat-Jen Cham, Guangming Shi
ACM Trans. Intell. Syst. Technol.5
2011 Progressive Quantization of Compressive Sensing Measurements
abstract
Compressive sensing (CS) is recently and enthusiastically promoted as a joint sampling and compression approach. The advantages of CS over conventional signal compression techniques are architectural: the CS encoder is made signal independent and computationally inexpensive by shifting the bulk of system complexity to the decoder. While these properties of CS allow signal acquisition and communication in some severely resource-deprived conditions that render conventional sampling and coding impossible, they are accompanied by rather disappointing rate-distortion performance. In the present work we propose a novel coding technique that rectifies, to certain extent, the problem of poor compression performance of CS and at the same time maintains the simplicity and universality of the current CS encoder design. The main innovation is a scheme of progressive fixed-rate scalar quantization with binning that enables the CS decoder to exploit hidden correlations between CS measurements, which was overlooked in the existing literature. Experimental results are presented to demonstrate the efficacy of the new CS coding technique.
Liangjun Wang, Xiaolin Wu 0001, Guangming Shi
DCC3