Jiguo Li 0002

dblp:33/31-2 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-1447-4798ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Cross Modal Compression With Variable Rate Prompt
abstract
Traditional image/video compression compresses the highly redundant visual data while preserving signal fidelity. Recently, cross modal compression (CMC) is proposed to compress the data into a compact, human-comprehensible domain (such as text) with an ultra-high compression ratio while preserving semantic fidelity for machine analysis and semantic monitoring. CMC is with a constant rate because the CMC encoder can only represent the data with a fixed grain. But in practice, variable rate is necessary due to the complicated and dynamic transmission conditions, the different storage mediums, and the diverse levels of application requirements. To deal with this problem, in this paper, we propose variable rate cross modal compression (VR-CMC), where we introduce variable rate prompt to represent the data with different grains. Variable rate prompt is composed of three strategies. Specifically, 1) target length prompt (TLP) introduces the target length into the language prompt to guide the generation of the text representation; 2) decaying EOS probability (DEP) exponentially decays the probability of the EOS token with regard to the decoding step and target length, where the EOS (end-of-sequence) token is a special token indicating the end of the text; 3) text augmentation (TA) enriches the training data and makes the text representation length more balanced when training. Experimental results show that our proposed VR-CMC can effectively control the rate in the CMC framework and achieve state-of-the-art performance on MSCOCO and IM2P datasets.
Junlong Gao, Jiguo Li 0002, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.2
2023 Semantic-Aware Visual Decomposition for Image Coding
Jianhui Chang, Jian Zhang 0018, Jiguo Li 0002, Shiqi Wang 0001, Qi Mao 0002, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
Int. J. Comput. Vis.3
2023 Rethinking Semantic Image Compression: Scalable Representation With Cross-Modality Transfer
abstract
This article proposes the scalable cross-modality compression (SCMC) paradigm, in which the image compression problem is further cast into a representation task by hierarchically sketching the image with different modalities. Herein, we adopt the conceptual organization philosophy to model the overwhelmingly complicated visual patterns, based upon the semantic, structure, and signal level representation accounting for different tasks. The SCMC paradigm that incorporates the representation at different granularities supports diverse application scenarios, such as high-level semantic communication and low-level image reconstruction. The decoder, which enables the recovery of the visual information, benefits from the scalable coding based upon the semantic, structure, and signal layers. Qualitative and quantitative results demonstrate that the SCMC can convey accurate semantic and perceptual information of images, especially at low bitrates, and promising rate-distortion performance has been achieved compared to state-of-the-art methods. The code will be available onlinehttps://github.com/ppingzhang/SCMC.
Shiqi Wang 0001, Meng Wang 0017, Jiguo Li 0002, Xu Wang 0006, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2022 Consistency-Contrast Learning for Conceptual Coding
abstract
As an emerging compression scheme, conceptual coding usually encodes images into structural and textural representations and decodes them in a deep synthesis fashion. However, existing conceptual coding schemes ignore the structure of deep texture representation space, leading to a challenge of establishing efficient and faithful conceptual representations. In this paper, we firstly introduce contrastive learning into conceptual coding and propose Consistency-Contrast Learning (CCL) which optimizes the representation space by a consistency-contrast regularization. By modeling the original images and reconstructed images as "positive'' pairs and random images in a batch as "negative'' samples, CCL aims to align texture representation space with source images space relatively. Extensive experiments on diverse datasets demonstrate that: (1) the proposed CCL can achieve the best compression performance on the conceptual coding task; (2) CCL is superior to other popular regularization methods towards improving reconstruction quality; (3) CCL is general and can be applied to other tasks related to representation optimization and image reconstruction, such as GAN inversion.
Jianhui Chang, Jian Zhang 0018, Youmin Xu, Jiguo Li 0002, Siwei Ma 0001, Wen Gao 0001
ACM Multimedia4
2021 Cross Modal Compression: Towards Human-comprehensible Semantic Compression
abstract
Traditional image/video compression aims to reduce the transmission/storage cost with signal fidelity as high as possible. However, with the increasing demand for machine analysis and semantic monitoring in recent years, semantic fidelity rather than signal fidelity is becoming another emerging concern in image/video compression. With the recent advances in cross modal translation and generation, in this paper, we propose the cross modal compression~(CMC), a semantic compression framework for visual data, to transform the high redundant visual data~(such as image, video, etc.) into a compact, human-comprehensible domain~(such as text, sketch, semantic map, attributions, etc.), while preserving the semantic. Specifically, we first formulate the CMC problem as a rate-distortion optimization problem. Secondly, we investigate the relationship with the traditional image/video compression and the recent feature compression frameworks, showing the difference between our CMC and these prior frameworks. Then we propose a novel paradigm for CMC to demonstrate its effectiveness. The qualitative and quantitative results show that our proposed CMC can achieve encouraging reconstructed results with an ultrahigh compression ratio, showing better compression performance than the widely used JPEG baseline.
Jiguo Li 0002, Chuanmin Jia, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001
ACM Multimedia1
2021 Learning to Fool the Speaker Recognition
abstract
Due to the widespread deployment of fingerprint/face/speaker recognition systems, the risk in these systems, especially the adversarial attack, has drawn increasing attention in recent years. Previous researches mainly studied the adversarial attack to the vision-based systems, such as fingerprint and face recognition. While the attack for speech-based systems has not been well studied yet, although it has been widely used in our daily life. In this article, we attempt to fool the state-of-the-art speaker recognition model and present speaker recognition attacker , a lightweight multi-layer convolutional neural network to fool the well-trained state-of-the-art speaker recognition model by adding imperceptible perturbations onto the raw speech waveform. We find that the speaker recognition system is vulnerable to the adversarial attack, and achieve a high success rate on both the non-targeted attack and targeted attack. Besides, we present an effective method by leveraging a pretrained phoneme recognition model to optimize the speaker recognition attacker to obtain a tradeoff between the attack success rate and the perceptual quality. Experimental results on the TIMIT and LibriSpeech datasets demonstrate the effectiveness and efficiency of our proposed model. And the experiments for frequency analysis indicate that high-frequency attack is more effective than low-frequency attack, which is different from the conclusion drawn in previous image-based works. Additionally, the ablation study gives more insights into our model.
Jiguo Li 0002, Xinfeng Zhang 0001, Jizheng Xu, Siwei Ma 0001, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Learning to Fool the Speaker Recognition
abstract
Due to the widespread deployment of fingerprint/face/speaker recognition systems, attacking deep learning based biometric systems has drawn more and more attention. Previous research mainly studied the attack to the vision-based system, such as fingerprint and face recognition. While the attack for speaker recognition has not been investigated yet, although it has been widely used in our daily life. In this paper, we attempt to fool the state-of-the-art speaker recognition model and present speaker recognition attacker, a lightweight model to fool the deep speaker recognition model by adding imperceptible perturbations onto the raw speech waveform. We find that the speaker recognition system is also vulnerable to the attack, and we achieve a high success rate on the non-targeted attack. Besides, we also present an effective method to optimize the speaker recognition attacker to obtain a trade-off between the attack success rate with the perceptual quality. Experiments on the TIMIT dataset show that we can achieve a sentence error rate of 99.2% with an average SNR 57.2dB and PESQ 4.2 with speed rather faster than real-time.
Jiguo Li 0002, Xinfeng Zhang 0001, Jizheng Xu, Li Zhang 0006, Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001
ICASSP1
2020 Universal Adversarial Perturbations Generative Network For Speaker Recognition
abstract
Attacking deep learning based biometric systems has drawn more and more attention with the wide deployment of fingerprint/face/speaker recognition systems, given the fact that the neural networks are vulnerable to the adversarial examples, which have been intentionally perturbed to remain almost imperceptible for human. In this paper, we demonstrated the existence of the universal adversarial perturbations (UAPs) for the speaker recognition systems. We proposed a generative network to learn the mapping from the low-dimensional normal distribution to the UAPs subspace, then synthesize the UAPs to perturbe any input signals to spoof the well-trained speaker recognition model with high probability. Experimental results on TIMIT and LibriSpeech datasets demonstrate the effectiveness of our model.
Jiguo Li 0002, Xinfeng Zhang 0001, Chuanmin Jia, Jizheng Xu, Li Zhang 0006, Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001
ICME1