Jinjia Zhou

dblp:77/8116 · DBLP profile ↗
← Back
14ranked-venue papers in the field
0as first author
12since 2021 · last 2025
0000-0002-5078-0522ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 8Other / Interdisciplinary · 3Knowledge Engineering, Semantic Web & Information Systems · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Multi Functional Compressive Sensing Sampling-Based Trustworthy Compressive Learning
abstract
Compressive Learning (CL) is a framework that solves computer vision tasks directly from a small amount of acquired signals via compressive sensing (CS). Existing CL works have achieved accuracy comparable to classical image-domain methods, showcasing innovative potential [1]. However, they attempt to cover the incomplete information by applying a function to enhance the features of acquired signals. Moreover, they do not consider the possibility of producing uncertain predictions by inferring from insufficient information. In this paper, we propose a multi-functional compressive sensing sampling-based trustworthy compressive learning, dubbed MCSTCL as shown in Figure 1. First CS-sampling is performed and obtain measurement vectors. Then, we rearrange each element in a image structure following convolution operations algorithm, expecting it to enhance performance in subsequent tasks. The resulting image structure can be regarded as a feature map, enabling an efficient process that simultaneously performs signal acquisition, compression, and information extraction during the CS sampling process. Additionally, we propose a new loss function based on evidential deep learning (EDL) to appropriately quantify uncertainty and allocate more evidence to the target class. Experimental results show that MCSTCL can achieve an image classification accuracy of 96.83% at a low sampling rate of 6% while reducing the model size by 85.76% compared with state-of-the-art work on practical datasets such as the UC Merced Land Use Dataset.
Fuma Kimishima, Jinjia Zhou
DCC3
2025 Bidirectional Learned Facial Animation Codec for Low Bitrate Talking Head Videos
abstract
In this paper, we propose a novel bidirectional learned animation codec that generates natural facial videos by using past and future keyframes. First, we introduce a compact auxiliary stream for non-keyframes, which is enhanced by adaptively selecting one of two keyframes (past and future) in the BRG-ASE process. This stream improves video quality with a slight increase in bitrate. Then, we animate the adaptively selected keyframe and reconstruct the target frame using both the animated keyframe and the auxiliary frame in the BRG-VRec process. In our bidirectional frame reconstruction method, the future keyframe is used as the past keyframe in the next group of pictures. Therefore, it is temporarily stored in the decoder.
Riku Takahashi, Ryugo Morita, Fuma Kimishima, Kosuke Iwama, Jinjia Zhou
DCC5
2025 Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
abstract
Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial reconstructions. To address these problems, we propose a novel audio-visual driven video codec that integrates compact 3D motion features and audio signals. This approach robustly models significant head rotations and aligns lip movements with speech, improving both compression efficiency and reconstruction quality. Experiments on the CelebV-HQ dataset show that our method reduces bitrate by 22% compared to VVC and by 8.5% over state-of-the-art learning-based codec. Furthermore, it provides superior lip-sync accuracy and visual fidelity at comparable bitrates, highlighting its effectiveness in bandwidth-constrained scenarios.
Riku Takahashi, Ryugo Morita, Jinjia Zhou
ICMR3
2023 Temporal Down-sampling based Video Coding with Frame-Recurrent Enhancement
abstract
In many digital systems, the transmission bandwidth, as well as storage capacity, are usually very limited. This introduces challenges for both video transmission and video storage. To seek lower bit rates and further obtain high-quality up-sampled videos, this paper proposes a temporal down-sampling based video coding system and a frame-recurrent enhancement based video upsampling strategy. The structure of our proposed method is shown in Fig. 1. Unlike the existing work [1], instead of downsampling all video frames, only the intermediate frames are downsampled and two frames remain with high quality on the video coding system. Then, these two high-quality frames are used to iteratively enhance the quality of the low-bitrate low-quality frames through a deep-learned enhancement network. Compared to the latest video coding standard Versatile Video Coding (VVC), our work can obtain a BD-rate reduction from $39.261 {\%} \sim 85. 455$ % in All-Intra and Low-Delay-P configurations on the downsampled frames. A temporal down-sampling based video coding framework (TDS) is proposed. It can be combined with all the existing coding standards including HEVC/H.265 and VVC/H.266. A method of super-resolution with frame recurrent image enhancement (SRFR) is applied to up-sampling the frames by the neighboring high resolution frame. The temporal information from high resolution frames can be fully used to improve the video quality through frame recurrent.
Keren He, Chi Do-Kim Pham, Lu Zhang 0037, Jinjia Zhou
DCC5
2023 VCSL: Video Compressive Sensing with Low-complexity ROI Detection in Compressed Domain
abstract
By exploiting the potential of deep learning, video compressive sensing (CS) has achieved tremendous improvement recently. Due to the video CS is mainly served for the fixed scene in real life. In this paper, we propose a novel video compressive sensing with a low-complexity region-of-interest (ROI) detection method (VCSL). The ROI is located by calculating the difference between the reference frame and the following frames in our framework, which is compact without introducing any additional neural networks and parameters. Subsequently, only the detected ROIs are sampled and transmitted, except for the frame that is regarded as the background. The final re-constructed sequence would be attained by combining the ROIs and the background. Moreover, the proposed reference frame renewal method successfully solves the issue of background changing and achieves more accurate reconstructed results while reducing the sampling rate (SR) further. The specific testing results of VIRAT dataset are shown in Table. 1. We control the SR of our baseline method to be close to the other methods with fixed SR. As shown in Table. 1, our proposed VCSL achieves the best reconstruction performance among algorithms in the comparison while using the lowest average SR. Compared to the state-of-the-art counterparts, extensive experimental results have demonstrated that our proposed methods achieve superior performance while tackling more complex sequences and using a lower sampling rate. We believe that the proposed framework can be integrated into the other existing works to save the data of transmission.
Haixing Wang, Yibo Fan, Jinjia Zhou
DCC4
2023 Zigzag Ordered Walsh Matrix for Compressed Sensing Image Sensor
abstract
In compressed sensing (CS) based CMOS image sensors (CS-CIS), the ternary measurement matrix determines the compression performance in terms of decoded image quality versus sampling rate (data rate). Several studies have been carried out to investigate the effect of Hadamard and Walsh projection order selection on image reconstruction quality by simply reordering orthogonal matrices [1]. However, there is still room for improvement in the quality of reconstructed images from these works, especially at low SR. In this paper, we propose a structured measurement matrix called Zigzag ordered Walsh matrix (ZoW), which outperforms at low sampling rates. Firstly, the Walsh matrix is divided into several measurement patterns. Because the lower frequency component in an image plays a more critical role in determining the image quality, we arrange low-frequency patterns on the upper-left corner, and the frequency increases according to the zigzag scan order. Then, vectorize each pattern and stacking back into ZoW matrix. Hence, under various sampling rates, the proposed ZoW always remains the lowest frequency patterns which are the most critical patterns. Comparing with the existing measurement matrices, recovery errors are improved by 4.09dB in PSNR on average and provided significantly better image quality via visual perception when sampling rates are 5%~15%.In Table 1, we can observe that the reconstruction quality via PSNR is improved by 4.63dB, and SSIM is improved 69% on average when the sampling rate is 10%.
Jinyao Zhou, Jiayao Xu, Jirayu Peetakul, Jinjia Zhou
DCC4
2023 Block based Adaptive Compressive Sensing with Sampling Rate Control
abstract
Compressive sensing (CS), acquiring and reconstructing signals below the Nyquist rate, has great potential in image and video acquisition to exploit data redundancy and greatly reduce the amount of sampled data. To further reduce the sampled data while keeping the video quality, this paper explores the temporal redundancy in video CS and proposes a block based adaptive compressive sensing framework with a sampling rate (SR) control strategy. To avoid redundant compression of non-moving regions, we first incorporate moving block detection between consecutive frames, and only transmit the measurements of moving blocks. The non-moving regions are reconstructed from the previous frame. In addition, we propose a block storage system and a dynamic threshold to achieve adaptive SR allocation to each frame based on the area of moving regions and target SR for controlling the average SR within the target SR. Finally, to reduce blocking artifacts and improve reconstruction quality, we adopt a cooperative reconstruction of the moving and non-moving blocks by referring to the measurements of the non-moving blocks from the previous frame. Extensive experiments have demonstrated that this work is able to control SR and obtain better performance than existing works.
Kosuke Iwama, Ryugo Morita, Jinjia Zhou
MMAsia3
2023 Adaptive Sampling for Computer Vision-Oriented Compressive Sensing
abstract
Compressive sensing (CS) is renowned for its efficient signal data compression. However, due to its compressive nature, the accuracy of downstream computer vision (CV) tasks by reconstruction inevitably degrades as sampling rate decreases. This limitation significantly hinders the application of existing CS techniques. To overcome the drawback, this paper presents a novel CS technique that employs adaptive sampling rates based on saliency distribution. The goal of this work is to enhance the preservation of information necessary for classification while reducing the weight of non-essential information. Experimental results show the effectiveness of the proposed adaptive sampling technique, which outperforms existing sampling CS techniques on STL10 and Imagenette datasets. The average classification accuracy is maximally improved by 26.23% and 18.25%, respectively.
Hiroki Nishikawa, Jinjia Zhou, Ittetsu Taniguchi, Takao Onoye
MMAsia3
2023 NuclSeg: nuclei segmentation using semi-supervised stain deconvolution
abstract
Recently, deep learning-inferred stain deconvolution/separation-based nuclei segmentation works demonstrated significant results by translating low-cost and prevalent immunohistochemical (IHC) slides to more expensive-yet-informative multiplex immunofluorescence (mpIF) images. However, when the input stain style is changed to Hematoxylin and Eosin stain (H&E), which is one of the principal tissue stains used in histology, the stain deconvolution/separation based works can not achieve satisfactory performance because the features of the input stain image are greatly changed. To solve this problem, we integrate stain transfer (H&E->IHC) before stain deconvolution (IHC->mpIF) to revise the image style. Moreover, a new semi-supervised learning strategy collaborating supervised and unsupervised learning processes are employed to diversify training data content for stain deconvolution.Firstly, in the supervised learning process, generative adversarial network (GAN) based image-to-image mapping (stain deconvolution) and inverse mapping models are trained on the dataset of co-registered IHC staining and mpIF staining of the same slides to respectively convert IHC to mpIF (I2m), and mpIF to IHC (m2I). After that, in the unsupervised learning process, high-quality IHC images from m2I model are selected according to the Discriminator score on another unpaired dataset, and then, IHC images are paired with input mpIF as the unsupervised training data for further improving the I2m model. This semi-supervised scheme balances two supervised and unsupervised errors while optimizing to limit the effect of imperfect pseudo inputs but still enhance stain deconvolution. Furthermore, image enhancement is applied after the stain deconvolution model to obtain high-quality segmentation masks. We thoroughly evaluate our method on publicly available benchmark datasets. The results show the proposed model obtains significant improvement compared to the SOTAs.
Ryohei Katayama, Michiya Matsusaki, Tomoyuki Miyao, Jinjia Zhou
MMAsia6
2022 Cube-based Video Coding Framework for Block-based Compressive Imaging
abstract
Block-based compressive imaging enables new video acquisition methodology while reducing raw data size, theoretically eliminating the need for complex coding algorithms. However, the redundancy associated with random projection remains when transmitting raw data. This paper takes a fresh look at raw data structure by viewing it as cube made up of multiple downsampled images rather than a vector. As a result, each individual data point can be regarded as a pixel, allowing us to code with greater flexibility and versatility than current works. Following that, we propose a tailored video coding framework for this structure that includes directional 9 modes intra and inter prediction with block-matching motion estimation, transformation using DCT, and quantization with custom 4×4 quantization table as shown in Figure 1. We evaluated coding performance using various 4K datasets, resulting in 60-65% lower bit-per-pixels while maintaining visual quality compared to state-of-the-art works [1].
Jirayu Peetakul, Yibo Fan, Jinjia Zhou
DCC3
2022 Text-to-image synthesis: Starting composite from the foreground content
Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002
Inf. Sci.2
2021 Local Feature Normalization
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou
KSEM4
2020 Temporal Redundancy Reduction in Compressive Video Sensing by using Moving Detection and Inter-Coding
abstract
Summary form only given. Compressed sensing-based CMOS image sensor (CS-CIS) has gained significant interest in the past few years. It can greatly reduce on-chip processor complexity, transmission cost, and storage requirement compared to conventional sampling method using Nyquist-Shannon rate. CS-CIS performs acquisition and compression simultaneously and transfers all heavy computation burden components to decoder, where the measurement streams can be processed and analyzed with unlimited resources, resulting in a low-complexity encoder. Thus, it is very suitable for transmitting only applications, where computational resource and power is limited. However, spatial and temporal redundancy in measurement has become a primary concern which it is necessary to further compress. In this paper, we proposed temporal redundancy reduction in compressive video sensing by using moving detection and inter-coding. Firstly, the moving detection is performed coding area extraction using local adaptive threshold to classify the measurement with an association of error distinction. However, false-positive detection could be occurred randomly, which transmission cost can be increasing uncertainty. To reduce transmission costs, the adaptive quantization parameters are adjusted by how frequently the area is detected. Moreover, we further compress the detected area by encoding the difference of current measurement and the best-matched measurement in neighboring frames. Finally, an efficient recovery algorithm of sparse signal is performed by using -minimization via primal-dual interior-point algorithm and reconstructed by inverse fast Walsh-Hadamard transform with horizontal kernel filter to prevent staircase artifacts simultaneously. The experimental results show that our proposed can greatly reduce bandwidth usage in terms of BPP by 63.15%, improve in PSNR by 1.56dB, and SSIM by 14.81% on average when compared to the state-of-the-art works.
Jirayu Peetakul, Jinjia Zhou
DCC2
2019 A Measurement Coding System for Block-Based Compressive Sensing Images by Using Pixel-Domain Features
abstract
Compressive sensing (CS) is data acquiring and innovative mathematical approach that accelerate and efficient sampling from large into small volumes of data. Moreover, it could be dramatically reduced amounts of sensor, power consumption, storage size, and bandwidth which results in lower hardware costs [1]. In wireless cameras network for video surveillance, the large amount of data is produced. However, there is still a lot of redundant data in measurement domain. To solve this problem, coding techniques such as block-based CS (BCS), intra-prediction and quantization is applied to avoid higher rate-distortion than other CS frameworks. Therefore, new imaging architecture has been proposed to be sensed, removed redundant information, and compressed simultaneously, thus leading to the faster image acquisition system.
Jirayu Peetakul, Jinjia Zhou, Koichi Wada 0001
DCC2