VLDB 2026 Research / reviewers in the wild / expert
Wuzhen Shi
dblp:136/2850
· DBLP profile ↗
52ranked-venue papers
20as first author
38since 2021 · last 2026
0000-0002-6819-0125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 17 first-author · 29 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View RestorationabstractMulti-view diabetic retinopathy (DR) grading has achieved remarkable performance by capturing more comprehensive pathological features than single-view methods. However, complete multi-view fundus images are often difficult to obtain in clinical practice, and the performance degrades significantly when fewer views are available. To overcome this limitation, we propose the first incomplete multi-view DR grading framework, aiming to provide accurate diagnosis regardless of the number of available views. It introduces two novel modules. First, cross-view spatial correlation attention (CSCA) captures region correlations across views, automatically identifying and fusing diagnostically relevant spatial features to improve feature representation. Second, self-supervised mask consistency learning (SMCL) formulates a novel pretext task of missing-view information reconstruction by strategically masking inter- and intra-view regions, enabling the model to infer complete features from incomplete views. Benefiting from CSCA and SMCL, our method enhances structural feature consistency across views and effectively compensates for missing information during DR grading. Extensive experiments demonstrate that our method achieves state-of-the-art grading performance, particularly under realistic conditions where some views are unavailable. Zhihao Wu 0002, Jie Wen 0001, Wuzhen Shi, LinLin Shen |
AAAI | 4 |
| 2026 | Quantization-Aware Diffusion Model for Variable-Rate Extreme Image CompressionabstractIn recent years, image compression based on diffusion models has achieved excellent perceptual reconstruction. However, their models typically support only fixed bitrates, which results in significant training costs and memory demands. In this paper, we propose a quantization-controllable diffusion-based image compression framework, which extends diffusion-based extreme compression to variable-rate scenarios by modulating quantization through a single control parameter. Furthermore, we introduce the Quantization Regulator Modulation RefineNet (QRMR), which dynamically modulates the diffusion decoding process according to different quantization levels, allowing the diffusion-based decoder to be quantization-aware and thereby significantly improving the rate-distortion performance. Experiments demonstrate that our method achieves good fidelity and perceptual quality in reconstructed images at extremely low bitrates. Yuran Zhang, Wuzhen Shi |
DCC | 2 |
| 2026 | Diff-KATKG: Diffusion-based talking head generation with joint keypoint and action unit guidance
Wuzhen Shi, Shuai Wang 0074, Zibang Xue |
Pattern Recognit. | 1 |
| 2026 | M2Net: Multimodal Multitask Mutual Learning for Anti-VEGF Efficacy PredictionabstractAge-related macular degeneration with abnormal blood vessel growth (neovascular AMD) is the leading cause of vision loss in elderly populations. While anti-VEGF injections are the standard treatment, they present financial burdens for patients and vary in effectiveness. Predicting treatment efficacy is therefore crucial for patient care. Current prediction methods fail to fully integrate information from different imaging techniques, typically focusing on either forecasting vision improvements or generating post-treatment images-but not both simultaneously. This approach overlooks the important relationship between these tasks. We present M2Net, a novel joint generation and classification network based on Multimodal Multitask Mutual learning, to simultaneously predict changes in visual acuity and generate post-treatment retinal images. M2Net employs a dual-branch structure that processes both fundus photographs and Optical Coherence Tomography (OCT) scans to improve prediction accuracy. Our framework includes two key innovations: the Multimodal Collaborative Treatment Efficacy Prediction module, which interacts the features between the two modalities and provides initial visual acuity change classification to guide the generation of post-treatment images; and the Pre-Post Treatment Image Joint Analysis module, which identifies both common and changing features between pre-treatment and post-treatment images to enhance prediction accuracy. To validate our approach, we created the dataset (MMPD) containing paired multimodal retinal images with corresponding visual acuity measurements. Experiments on the dataset demonstrate that M2Net achieves superior performance compared to existing methods, with a classification accuracy of 96.03%, an SSIM of 0.6377 on the OCT modality, and an SSIM of 0.8347 on the fundus modality. Our code will be available at https://github.com/zengying123/M2Net. Lei Bi 0001, Wuzhen Shi, Huazhu Fu, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Dynamic Interactive Bimodal Hypergraph Networks for Emotion Recognition in ConversationsabstractThe advancement in multimodal research has increased focus on Emotion Recognition in Conversations (ERC), targeting accurately identifying emotional changes. Methods based on graph convolution can better capture the dynamic changes of emotions and improve the accuracy and robustness of emotion recognition. However, existing methods do not distinguish the interaction patterns of a conversation, which results in limiting their ability to model contextual emotional relationships. In this paper, we propose a Dynamic Interactive Bimodal HyperGraph Convolutional Networks (DIB-HGCN), which creatively constructs two types of sub-hypergraphs, i.e., the monologic sub-hypergraph and the dialogic sub-hypergraph, for modeling emotion relationships of different interaction patterns. The monologic sub-hypergraph is used to explore the contextual consistent emotions during the speaker's monologue interactions, while the dialogic sub-hypergraph focuses on capturing the emotional transfers in the dialogic interactions. Meanwhile, the single window partitioning mechanism fails to accommodate the distinct emotional velocity variations across the two interaction patterns. Therefore, we set up dynamic windows in the monologic interactions to fully utilize the information of sentence nodes with consistent emotions, and we add fragment windows to the dialogic interactions to prevent information interference caused by frequent emotional transfers. The experimental results show that our proposed method outperforms existing methods on two benchmark multimodal ERC datasets. Xuping Chen, Wuzhen Shi |
AAAI | 2 |
| 2025 | PLNet: Entropy-Guided Pseudo Label Refinement with Sensitivity-Specificity Enhancement for Medical Image Segmentation
Huibin Weng, Zhiquan He, Wuzhen Shi |
CGI (1) | 6 |
| 2025 | Scalable Image Compression Based on Diffusion Models at Ultra-Low BitratesabstractImage codecs are typically optimized to trade-off between bitrate and distortion metrics. At low bitrates, they often result in compression artefacts. To solve this problem, we leverage the ability of diffusion models to produce high-quality images and design a scalable lossy image compression framework for ultra-low bitrates, as shown in Fig. 1. On the encoding side, we extract multiple conditional information from the image and perform scalable encoding on them. When the bandwidth is limited, only the text is encoded. We use Prompt Inversion (PI) to extract text and apply lossless compression using Lempel-Ziv (LZ) coding from the zlib library. As the bandwidth gradually increases, sketch and Spatial color palette can be encoded to supplement the conditional information. We use the edge prediction model PiDiNet to extract the sketch of the image. We further perform scalable palette compression to provide color information at different bitrates. We downsample the image by 64x using bicubic interpolation and upsample it to 1/32, 1/16, 1/4, and 1/2 of the original size with nearest-neighbor to obtain the spatial color palette at different bitrates. We use standard learned nonlinear transform codes (NTC) to compress the sketch and spatial color palette. On the decoding side, We use the T2I-Adapter [1] as our decoder. The text is processed through the CLIP text encoder to generate conditional text embeddings, while the sketch and spatial color palette are fed through adapters to obtain features at different scales. The decoding is performed in a scalable manner, where the transmitted information is processed as described above and input into the diffusion model to guide the image generation. Compared to JPEG and the latest diffusion-based methods (PIC and PICS [2]), our method achieves better perceptual quality, as shown in Table 1. It's important to note that t denotes using text as the sole condition, ts indicates using both text and sketch, and tsc1 to tsc4 refer to using text, sketch, and color palette at different bitrates as conditions. Experimental results show that as the number of conditional information increases and the accuracy of color information improves, the perceptual quality of the images improves. Wuzhen Shi, Yuran Zhang |
DCC | 1 |
| 2025 | Semantic Prior-Guided Scalable Image CodingabstractWe propose a semantic priori-guided scalable image coding method for simultaneously supporting fast machine vision analysis and high quality human visual experience. To obtain high-performance machine vision analysis results with a more compact base layer bitrate, our base layer directly encodes intermediate semantic features of the pre-trained machine vision task network, which effectively reduces the impact of the information needed for the human vision task. To improve model performance while quickly supporting machine vision tasks, we use structural re-parameterization technology to optimize the model. Considering that the base layer’s semantic features can effectively reflect the regions of important image content, we use the semantic prior provided by the base layer to guide the enhancement layer encoding and decoding, which allows us to pay more attention to the reconstruction of semantically important regions. In addition, we use the base layer features to predict the enhancement layer features for performing feature-domain residual coding, which further reduces the bitrate and also reduces the effect of noise compared to pixel-domain residual coding. Extensive experimental results show that our method achieves significant advantages in both object detection performance and image reconstruction tasks compared with BPG and state-of-the-art deep learning-based scalable image coding methods. Wuzhen Shi, Wennan Yin, Fei Tao 0005 |
ICASSP | 1 |
| 2025 | Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection
Yuqi Xiong, Wuzhen Shi, Ruhan Liu |
ICONIP (5) | 2 |
| 2025 | Edge-guided 3D reconstruction from multi-view sketches and RGB images
Wuzhen Shi, Aixue Yin, Yingxiang Li |
Pattern Recognit. | 1 |
| 2025 | Keypoints and Action Units Jointly Drive Talking Head Generation for Video ConferencingabstractThis paper introduces a high-quality talking head generation method that is jointly driven by keypoints and action units, aiming to strike a balance between low-bandwidth transmission and high-quality generation in video conference scenarios. Existing methods for talking head generation often face limitations: they either require an excessive amount of driving information or struggle with accuracy and quality when adapted to low-bandwidth conditions. To address this, we decompose the talking head generation task into two components: a driving task, focused on information-limited control, and an enhancement task, aimed at achieving high-quality, high-definition output. Our proposed method innovatively incorporates the joint driving of keypoints and action units, improving the accuracy of pose and expression generation while remaining suitable for low-bandwidth environments. Furthermore, we implement a multistep video quality enhancement process, targeting both the entire frame and key regions, while incorporating temporal consistency constraints. By leveraging attention mechanisms, we enhance the realism of the challenging-to-generate mouth regions and mitigate background jitter through background fusion. Finally, a prior-driven super-resolution network is employed to achieve high-quality display. Extensive experiments demonstrate that our method effectively supports low-resolution recording, low-bandwidth transmission, and high-definition display. Wuzhen Shi, Zibang Xue |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Identity and Modality Attributes Driven Multimodal Fusion Networks for Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC) is a crucial aspect of human-computer interaction and plays an important role in various domains, including healthcare, entertainment, and education. Since the conversation data in the form of multimodal sequences is well suited to be constructed into graphs, the methods based on graph convolutional network (GCN) show incomparable advantages. However, existing methods attempt to model the highly uncertain emotional relationships between different speakers, which is not an easy task and may even introduce interference information. Therefore, we propose an identity and modality attributes driven multimodality fusion network (dubbed IMDNet) for emotion recognition in conversations. Specifically, we construct a speaker-centric graph that only connects nodes of the same speaker within modalities to each other, reducing the interference between the emotions of different speakers. We also introduce the attribute embedding mechanism, which facilitates the correct calculation of correlations between nodes for better multimodal feature fusion. Considering that the emotional correlation between utterances will decrease over time, we present an utterance distance attention to make the fusion network pay more attention to the adjacent utterances. Furthermore, we explore the solution to the data imbalance problem suitable for conversation scenarios. Given the presence of possible anomalous samples in the dataset, we opt for the BoundaryFocalLoss. Experiments on the IEMOCAP and MELD datasets show that our IMDNet outperforms the state-of-the-art methods. Wuzhen Shi, Xuping Chen, Biyun Yao, Bin Sheng 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | SAT-Net: Structure-Aware Transformer-Based Attention Fusion Network for Low-Quality Retinal FunduImages EnhancementabstractIn ophthalmology diagnosis, high-fidelity fundus images are essential for disease diagnosis and intervention. However, many real-world clinical conditions may degrade the quality of the acquired images and thus affect clinical diagnostic accuracy. Traditional convolutional neural network-based retinal fundus image enhancement methods cannot always capture long-range dependencies, which reduces the overall visual quality of images, especially for real retinal fundus images. Furthermore, existing enhancement methods often fail to fully utilize low-resolution structural detail information, which potentially leads to inaccurate pivotal fundus vessel topology or capillary details. In this paper, we propose a novel Structure-Aware Transformer-based attention fusion Network (SAT-Net) for low-quality retinal fundus image enhancement. First, we introduce a Transformer-based attention fusion module which incorporates window-based self-attention and channel self-attention to capture global spatial dependencies and emphasize important feature channels simultaneously. This fusion significantly improves the overall perceptual quality of the image by enhancing both the local details and the uniformity of the non-vessel background regions. Second, we introduce a cross-quality knowledge distillation technique, which bridges the quality gap between high-quality and low-quality fundus images. By designing a high-performing teacher network to guide a lightweight student network, the student network enables to capture detailed features from low-quality fundus images, further preserving critical diagnostic information and fine topology structures. Moreover, we design a structure-aware multi-scale loss function by using a trainable subnetwork to obtain the edge structure from different scales to better constrain pivotal fundus vessel structure and capillary details. Comprehensive quantitative and qualitative experiments on both synthetic and real fundus image datasets robustly validate that our proposed SAT-Net outperforms other state-of-the-art methods for fundus image enhancement. In addition, extensive comparative experiments on both the vessel segmentation and Optic Disc/Cup detection tasks further validate the effectiveness and superiority of our proposed method. Wuzhen Shi, Jianhua Ji, Wenming Cao 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Video Compressed Sensing Via Wavelet Residual Sampling and Dual-Domain FusionabstractDeep learning-based compressed sensing (CS) technology attracts widespread attention owing to its remarkable reconstruction with only a few sampling measurements and low computational complexity. However, the existing video compressive sampling approaches cannot fully exploit the inherent interframe and intraframe correlations and sparsity of video sequences. To address this limitation, a novel sampling and reconstruction method for video CS (called WRDD) is proposed, which exploits the advantages of wavelet residual sampling and dual-domain fusion optimization. Specifically, in order to capture high-frequency details and achieve efficient and high-quality measurements, we propose a wavelet residual (WR) sampling strategy for the nonkeyframe sampling, which is achieved by the wavelet residuals between nonkeyframes and keyframes. Furthermore, a dual-domain (DD) fusion strategy is proposed, which fully combine intraframe and interframe to improve the reconstruction quality of nonkeyframes both in the pixel domain and multilevel feature domains. Extensive experiments demonstrate that our WRDD surpasses the state-of-the-art video and image CS methods in both subjective and objective evaluations. Besides, it exhibits outstanding antinoise capability and computational efficiency. Zhu Yin, ZhongCheng Wu, Wuzhen Shi, Guyue Hu 0001, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2025 | A Lightweight Depthwise Separable ConvNet with Frequency-domain Enhancement for Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is crucial in the diagnosis and treatment of various cardiovascular and eye diseases. Although current vessel segmentation methods have achieved impressive performance, some challenging issues still need to be addressed. For example, existing methods always cannot segment complex capillaries well because they may be interfered with or covered by other components in the retina, and they need to further improve the continuity and consistency of vessel segmentation results. Moreover, the excellent vessel segmentation methods are usually built on bulky and cumbersome models which greatly limit their application range. In this article, we propose a novel efficient depthwise separable convolution network with frequency-domain enhancement (dubbed RetiNeXt) for retinal vessel segmentation. Firstly, we design a lightweight vessel enhancement module to extract global fine topological structure features from the frequency domain to enhance the complex capillary vessel details. Secondly, we propose a global feature extraction block to fully capture the large-scale spatial information and global characterizations, which enables the model to maintain vessel structural coherence from a global perspective. Thirdly, we construct a local feature mixing block based on SimAM attention mechanism to highlight the tiny capillary topological structure features and optimize the segmentation of low-contrast blood vessels, thereby improving the integrity and continuity of complex capillaries. Comprehensive comparison experiments on three well-benchmarked retinal vessel segmentation datasets fully verify the effectiveness and superiority of the proposed RetiNeXt. To further demonstrate the universality of RetiNeXt for medical image segmentation, we also conduct sufficient comparative experiments on two classical coronary angiography datasets. Extensive quantitative and qualitative experiments fully show that RetiNeXt outperforms other state-of-the-art methods with only 0.4M of trainable parameters. Shunzhe Shen, Wuzhen Shi, Wenming Cao 0001, Lei Bi 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Joint super-resolution-based fast face image coding for human and machine vision
Wuzhen Shi, Fei Tao 0005 |
Vis. Comput. | 1 |
| 2025 | Cross-view Transformer for enhanced multi-view 3D reconstruction
Wuzhen Shi, Aixue Yin, Yingxiang Li |
Vis. Comput. | 1 |
| 2025 | MSPFM: Multi-Scale Pyramid Fusion Mamba for Medical Image Classification
Wuzhen Shi, Daquan Feng, Wenming Cao 0001 |
Vis. Comput. | 3 |
| 2025 | Multi-weather Image Restoration via Histogram-Based Transformer Feature Enhancement
Anyu Lai, Wuzhen Shi, Wenming Cao 0001 |
Vis. Comput. | 4 |
| 2025 | Multi-prior guided depth map super-resolution based on a diffusion model
Wuzhen Shi, Jianhua Ji, Wenming Cao 0001, Zhiquan He |
Vis. Comput. | 3 |
| 2025 | Dual prior guided depth image super-resolution with multi-scale transformer fusion network
Jianhua Ji, Wuzhen Shi, Wenming Cao 0001 |
Vis. Comput. | 4 |
| 2024 | Multiple Weather Images Restoration Using the Task Transformer and Adaptive Mixup Strategy
Anyu Lai, Hao Wang 0075, Wuzhen Shi, Wenming Cao 0001 |
CGI (1) | 5 |
| 2024 | Fusing Structure and Appearance Features in Facial Expression Recognition TransformerabstractFacial expression recognition (FER) methods are fundamental in various human-computer interaction scenarios. Although deep learning-based models have made substantial progress in the FER field, they primarily focus on capturing facial appearance features while neglecting the importance of structure features, which encompass the overall shape and structure details of the key facial regions. We propose a Structure and Appearance Feature Cross-fusion Transformer (SAFCT) network to leverage structure and appearance features. Specifically, we introduce the gradient-based structure feature to simultaneously capture the overall face shape and local organ variations. For appearance features, we extract both global and landmarks-guided local features to capture global texture and local details. Furthermore, we employ the structure-dominated cross-fusion transformer to integrate these three facial features. Through extensive experimental results, we evaluate the state-of-the-art recognition performance of SAFCT on widely used FER datasets. Siwei Meng, Wuzhen Shi |
ICASSP | 2 |
| 2024 | Speaker-Centric Multimodal Fusion Networks for Emotion Recognition in ConversationsabstractExisting emotion recognition methods in conversations (ERC) focus on using different utterances information between speakers to improve emotion recognition performance, but they ignore the differential contributions of different utterances to emotion recognition. In this paper, we propose a speaker-centric multimodal fusion network for ERC, in which bidirectional gated recurrent units (BiGRU) is used for intra-modal feature fusion and graph convolution is used for speaker-centric cross-modal feature fusion. We construct a speaker-centric graph based on the differences between one speaker’s utterances and that of the other speakers. This graph enhances the network’s focus on each speaker’s own utterance information, effectively reducing interference from other speakers. Simultaneously, we employ a Utterance Distance Attention (UDA) module, tailoring the attention allocation to mitigate the impact of distant utterances on the current utterance. Experimental results on IEMOCAP and MELD demonstrate the effectiveness of our approach. Biyun Yao, Wuzhen Shi |
ICASSP | 2 |
| 2024 | A Transformer-Assisted Cascade Learning Network for Choroidal Vessel Segmentation
Lei Bi 0001, Wuzhen Shi, Yupeng Xu, Wenming Cao 0001, David Dagan Feng |
J. Comput. Sci. Technol. | 4 |
| 2024 | Scalable compressive sampling network with progressive hierarchical subspace learning
Zhu Yin, ZhongCheng Wu, Wuzhen Shi |
Pattern Recognit. | 3 |
| 2024 | Co-salient object detection with iterative purification and predictive optimizationabstractCo-salient object detection (Co-SOD) aims to identify and segment commonly salient objects in a set of related images. However, most current Co-SOD methods encounter issues with the inclusion of irrelevant information in the co-representation. These issues hamper their ability to locate co-salient objects and significantly restrict the accuracy of detection. To address this issue, this study introduces a novel Co-SOD method with iterative purification and predictive optimization (IPPO) comprising a common salient purification module (CSPM), predictive optimizing module (POM), and diminishing mixed enhancement block (DMEB). These components are designed to explore noise-free joint representations, assist the model in enhancing the quality of the final prediction results, and significantly improve the performance of the Co-SOD algorithm. Furthermore, through a comprehensive evaluation of IPPO and state-of-the-art algorithms focusing on the roles of CSPM, POM, and DMEB, our experiments confirmed that these components are pivotal in enhancing the performance of the model, substantiating the significant advancements of our method over existing benchmarks. Experiments on several challenging benchmark co-saliency datasets demonstrate that the proposed IPPO achieves state-of-the-art performance. Yuhuan Wang, Hao Wang 0075, Wuzhen Shi, Wenming Cao 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2023 | Occlusion-Aware Graph Neural Networks for Skeleton Action RecognitionabstractDue to its important applications in industry, the research of action recognition has attracted many attention. In particular, skeleton-based human action recognition methods have evolved to become more and more competitive. Nevertheless, occlusion is still a very challenging task in human action recognition as yet. Off-the-shelf works are usually based on complete skeleton data, few people consider action recognition in occlusion. For the sake of improving the recognition accuracy on occluded skeleton data, we put forward an occlusion-aware multistream fusion graph convolutional network (dubbed MSFGCN). Multiple streams are comprised in MSFGCN, and different occlusion cases can be disposed by different streams. Besides, in order to construct more discriminative features, multimodal features were extracted simultaneously, such as joint coordinates, relative coordinates, small-scale temporal differences, and large-scale temporal differences. In particular, it is the first time to take advantage of motion features at large and small scales at the same time, which helps to distinguish actions with different motion amplitudes. What is more, considering the different importance of different parts of the human body for different action recognition, the content adaptation operation is used to further optimize the recognition performance. Effective training strategies are also presented to further improve the performance of the model. Experimental results show that the proposed MSFGCN has a great advantage over other methods on occluded skeleton datasets. To showcase the effectiveness of different modules, extensive ablation experiments are performed on various skeleton action recognition datasets. Wuzhen Shi |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Detail Generation and Fusion Networks for Image InpaintingabstractRecent image inpainting methods based on end-to-end models have achieved great success with the help of generative adversarial training and structure generation. However, the generated conditional structure priors cannot support the model to reconstruct finer texture. This is mainly because these models lack good texture prior knowledge, which leads to their limited ability to generate finer texture. In this paper, we propose a novel detail generation and fusion network (DGFNet) to strengthen the generation of texture details for image inpainting, which includes a dual-stream texture generation network and a multi-scale difference perception fusion network. The dual-stream texture generation network can explicitly model the missing texture information and generate a texture map to compensate for the coarse result produced by the parallel network. Furthermore, to merge two different kinds of information effectively, a fusion network based on the differential perception fusion module (DPFM) is introduced for multi-scale perception fusion in feature level. Extensive qualitative and quantitative experiments on the benchmark dataset show that the proposed DGFNet achieves state-of-the-art performance. Wuzhen Shi |
ICASSP | 2 |
| 2022 | Multilevel wavelet-based hierarchical networks for image compressed sensing
Zhu Yin, Wuzhen Shi, ZhongCheng Wu, Jun Zhang 0034 |
Pattern Recognit. | 2 |
| 2022 | Hierarchical complementary learning for weakly supervised object localization
Sabrina Narimene Benassou, Wuzhen Shi, Feng Jiang 0001, Abdallah Benzine |
Signal Process. Image Commun. | 2 |
| 2022 | Spatio-Temporal Context Based Adaptive Camcorder Recording WatermarkingabstractVideo watermarking technology has attracted increasing attention in the past few years, and a great deal of traditional and deep learning-based methods have been proposed. However, these existing methods usually suffer from the following two challenges: First, most algorithms cannot resist camcorder recording attack, which limits their practical application. Second, watermark embedding may cause substantial degradation of video quality. Through analyzing the unique distortions presented in the camcorder recording process, including geometric distortion, temporal sampling distortion, sensor distortion and processing distortion, this paper proposes a novel spatio-temporal context based adaptive camcorder recording watermarking scheme STACR. In STACR, considering the geometric distortion and video visual quality, we embed the watermark by constructing a spatio-temporal histogram and incorporate a content features based adaptive locating algorithm to select embedding blocks and embedding strengths. As for the temporal sampling attack, we put forward a watermark correlation-based synchronization algorithm and combine it with cross-validation. Moreover, to resist the sensor distortion, we design a local matching-based algorithm to improve the extraction accuracy. In addition, grouped and repeated embedding strategies are combined to cope with the processing distortion. Experimental results compared with the state-of-the-art show that the proposed scheme achieves high video quality and is robust to geometric attacks, compression, scaling, transcoding, recoding, frame rate changes and especially for camcorder recording. Shaohui Liu, Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Hiding Message Using a Cycle Generative Adversarial NetworkabstractTraining an image steganography is an unsupervised problem, because it is impossible to obtain an ideal supervised steganographic image corresponding to the cover image and secret message. Inspired by the success of cycle generative adversarial networks in unsupervised tasks such as style transfer, this article proposes to use a cycle generative adversarial network to solve the problem of unsupervised image steganography. Specifically, this article jointly trains five networks, i.e., a steganographic network, an inverse steganographic network, a hidden message reconstruction network, and two discriminative networks, which together constitute a hidden message cycle generative adversarial network (HCGAN). Compared with the recent image steganography based on generative adversative network, HCGAN provides more accurate supervised information, which makes the training process of HCGAN converge faster and the performance of the trained image steganography network is better. In addition, this article introduces an image steganographic network based on residual learning and shows that residual learning can effectively improve the performance of steganography. Furthermore, to the best of our knowledge, we are the first to propose an inverse steganographic network for eliminating steganographic message from steganographic images, which can be used to avoid steganographic message being discovered or acquired by a third party. The experimental results show that compared with the steganography based on generative adversarial network, the proposed HCGAN has a higher correct decoding rate, better visual quality of steganographic image, and higher secrecy. Wuzhen Shi, Shaohui Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Partially Occluded Skeleton Action Recognition Based on Multi-stream Fusion Graph Convolutional Networks
Wuzhen Shi |
CGI | 2 |
| 2021 | Small Object Recognition Using a Spatio-Temporal Neural NetworkabstractObject recognition at different scales has been a fundamental problem in computer vision. In particular, small object recognition attracts increasing attention recently. However, because of working on a single frame only, many recognizers’ performances become unacceptable in many practical application scenarios: very low resolutions, invisible small targets, extremely similar appearances etc. Motivated by the way humans deal with these challenging scenarios of object recognition, this paper introduces frame sequence and attention mechanism to compensate for mutilated information. Specifically, this paper proposes a spatiotemporal neural network (dubbed STNet) for small object recognition. STNet fixes the regions of interest with a super-resolution module, and focuses on the discriminative region with a spatio-temporal attention module. In addition, STNet applies a double layer long short-term memory subnet to make full use of the inter-frame information. Furthermore, this paper presents a challenging air-target recognition dataset ATSETC4 for evaluating the performance of each method in identifying small targets. Our model outperforms many state-of-the-art models on ATSETC4, including MobileNetV2 and SENet. In particular, STNet surpasses VGG11 at an average of 3.67%, even reaches 87.50% and 82.50% on 28 scale and 14 scale on AT-SETC4 respectively. Zhibo Liang, Shaohui Liu, Wuzhen Shi, Feng Jiang 0001 |
ICME | 3 |
| 2021 | Entropy guided adversarial model for weakly supervised object localization
Sabrina Narimene Benassou, Wuzhen Shi, Feng Jiang 0001 |
Neurocomputing | 2 |
| 2021 | Combining Fields of Experts (FoE) and K-SVD methods in pursuing natural image priors
Feng Jiang 0001, Zhiyuan Chen 0007, Amril Nazir, Wuzhen Shi, Wei Xiang Lim, Shaohui Liu, Seungmin Rho |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Video Compressed Sensing Using a Convolutional Neural NetworkabstractRecently, a few image compressed sensing (CS) methods based on deep learning have been developed, which achieve remarkable reconstruction quality with low computational complexity. However, these existing deep learning-based image CS methods focus on exploring intraframe correlation while ignoring interframe cues, resulting in inefficiency when directly applied to video CS. In this paper, we propose a novel video CS framework based on a convolutional neural network (dubbed VCSNet) to explore both intraframe and interframe correlations. Specifically, VCSNet divides the video sequence into multiple groups of pictures (GOPs), of which the first frame is a keyframe that is sampled at a higher sampling ratio than the other nonkeyframes. In a GOP, the block-based framewise sampling by a convolution layer is proposed, which leads to the sampling matrix being automatically optimized. In the reconstruction process, the framewise initial reconstruction by using a linear convolutional neural network is first presented, which effectively utilizes the intraframe correlation. Then, the deep reconstruction with multilevel feature compensation is proposed, which compensates the nonkeyframes with the keyframe in a multilevel feature compensation manner. Such multilevel feature compensation allows the network to better explore both intraframe and interframe correlations. Extensive experiments on six benchmark videos show that VCSNet provides better performance over state-of-the-art video CS methods and deep learning-based image CS methods in both objective and subjective reconstruction quality. Wuzhen Shi, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Image Compressed Sensing Using Convolutional Neural NetworkabstractIn the study of compressed sensing (CS), the two main challenges are the design of sampling matrix and the development of reconstruction method. On the one hand, the usually used random sampling matrices (e.g. GRM) are signal independent, which ignore the characteristics of the signal. On the other hand, the state-of-the-art image CS methods (e.g. GSR and MH) achieve quite good performance, however with much higher computational complexity. To deal with the two challenges, we propose an image CS framework using convolutional neural network (dubbed CSNet) that includes a sampling network and a reconstruction network, which are optimized jointly. The sampling network adaptively learns the sampling matrix from the training images, which makes the CS measurements retain more image structural information for better reconstruction. Specifically, three types of sampling matrices are learned, i.e. floating-point matrix, {0,1}-binary matrix, and {-1,+1}-bipolar matrix. The last two matrices are specially designed for easy storage and hardware implementation. The reconstruction network, which contains a linear initial reconstruction network and a non-linear deep reconstruction network, learns an end-to-end mapping between the CS measurements and the reconstructed images. Experimental results demonstrate that CSNet offers state-of-the-art reconstruction quality, while achieving fast running speed. In addition, CSNet with {0,1}-binary matrix, and {-1,+1}-bipolar matrix gets comparable performance with the existing deep learning based CS methods, and outperforms the traditional CS methods. What's more, the experimental results further suggest that the learned sampling matrices can improve the traditional image CS reconstruction methods significantly. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
IEEE Trans. Image Process. | 1 |
| 2020 | VINet: A Visually Interpretable Image Diagnosis NetworkabstractRecently, due to the black box characteristics of deep learning techniques, the deep network-based computer-aided diagnosis (CADx) systems have encountered many difficulties in practical applications. The crux of the problem is that these models should be explainable the model should give doctors rationales that can explain the diagnosis. In this paper, we propose a visually interpretable network (VINet) which can generate diagnostic visual interpretations while making accurate diagnoses. VINet is an end-to-end model consisting of an importance estimation network and a classification network. The former produces a diagnostic visual interpretation for each case, and the classifier diagnoses the case. In the classifier, by exploring the information in the diagnostic visual interpretation, the irrelevant information in the feature maps is eliminated by our proposed feature destruction process. This allows the classification network to concentrate on the important features and use them as the primary references for classification. Through a joint optimization of higher classification accuracy and eliminating as many irrelevant features as possible, a precise, fine-grained diagnostic visual interpretation, along with an accurate diagnosis, can be produced by our proposed network simultaneously. Based on a computed tomography image dataset (LUNA16) on pulmonary nodule, extensive experiments have been conducted, demonstrating that the proposed VINet can produce state-of-the-art diagnostic visual interpretations compared with all baseline methods. Donghao Gu, Feng Jiang 0001, Zhaojing Wen, Shaohui Liu, Wuzhen Shi, Guangming Lu 0001, Changsheng Zhou |
IEEE Trans. Multim. | 6 |
| 2019 | Scalable Convolutional Neural Network for Image Compressed SensingabstractRecently, deep learning based image Compressed Sensing (CS) methods have been proposed and demonstrated superior reconstruction quality with low computational complexity. However, the existing deep learning based image CS methods need to train different models for different sampling ratios, which increases the complexity of the encoder and decoder. In this paper, we propose a scalable convolutional neural network (dubbed SCSNet) to achieve scalable sampling and scalable reconstruction with only one model. Specifically, SCSNet provides both coarse and fine granular scalability. For coarse granular scalability, SCSNet is designed as a single sampling matrix plus a hierarchical reconstruction network that contains a base layer plus multiple enhancement layers. The base layer provides the basic reconstruction quality, while the enhancement layers reference the lower reconstruction layers and gradually improve the reconstruction quality. For fine granular scalability, SCSNet achieves sampling and reconstruction at any sampling ratio by using a greedy method to select the measurement bases. Compared with the existing deep learning based image CS methods, SCSNet achieves scalable sampling and quality scalable reconstruction at any sampling ratio with only one model. Experimental results demonstrate that SCSNet has the state-of-the-art performance while maintaining a comparable running speed with the existing deep learning based image CS methods. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
CVPR | 1 |
| 2019 | Hierarchical residual learning for image denoising
Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Rui Wang 0093, Debin Zhao, Huiyu Zhou 0001 |
Signal Process. Image Commun. | 1 |
| 2018 | Multi-Scale Deep Networks for Image Compressed SensingabstractAs a successful deep model applied in image compressed sensing, the Compressed Sensing Network (CSNet) has demonstrated superior performance to the previous handcrafted models in both running speed and reconstruction quality. However, CSNet trains different models for different sampling rates that hinders it from practical usage since too many models need to store. In this paper, we propose multi-scale deep network for image compressed sensing. We still use a sampling network to learn the sampling operator and implement the compressed sampling process. Given the compressed measurements, the reconstruction network directly maps them to the desired reconstructed images. There are three main differences in comparison with CSNet. Firstly, this paper proposes to use an unified deep reconstruction network for all sampling rates that decreases large amount of storage requirements. Secondly, we redesign a better deep reconstruction network using the popular residual learning technology. Finally, we investigate an image local smooth prior based loss function to enhance image structural information. Extensive experimental results show that the proposed multi-scale deep network based image compressed sensing method outperforms many other state-of-the-art methods. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICIP | 1 |
| 2017 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractSummary form only given. Traditional image coding standards (such as JPEG and JPEG2000) make the decoded image suffer from many blocking artifacts or noises since the use of big quantization steps. To overcome this problem, we proposed an end-to-end compression framework based on two CNNs, as shown in Figure 1, which produce a compact representation for encoding using a third party coding standard and reconstruct the decoded image, respectively. To make two CNNs effectively collaborate, we develop a unified end-to-end learning framework to simultaneously learn CrCNN and ReCNN such that the compact representation obtained by CrCNN preserves the structural information of the image, which facilitates to accurately reconstruct the decoded image using ReCNN and also makes the proposed compression framework compatible with existing image coding standards. Wen Tao, Feng Jiang 0001, Shengping Zhang, Jie Ren 0016, Wuzhen Shi, Wangmeng Zuo, Xun Guo 0002, Debin Zhao |
DCC | 5 |
| 2017 | Single image super-resolution with dilated convolution based multi-scale information learning inception moduleabstractTraditional works have shown that patches in a natural image tend to redundantly recur many times inside the image, both within the same scale, as well as across different scales. Make full use of these multi-scale information can improve the image restoration performance. However, the current proposed deep learning based restoration methods do not take the multi-scale information into account. In this paper, we propose a dilated convolution based inception module to learn multi-scale information and design a deep network for single image super-resolution. Different dilated convolution learns different scale feature, then the inception module concatenates all these features to fuse multi-scale information. In order to increase the reception field of our network to catch more contextual information, we cascade multiple inception modules to constitute a deep network to conduct single image super-resolution. With the novel dilated convolution based inception module, the proposed end-to-end single image super-resolution network can take advantage of multi-scale information to improve image super-resolution performance. Experimental results show that our proposed method outperforms many state-of-the-art single image super-resolution methods. Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ICIP | 1 |
| 2017 | Deep networks for compressed image sensingabstractThe compressed sensing (CS) theory has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been recently proposed and obtained superior performance. However, there still exist two important challenges within the CS theory. The first one is how to design a sampling mechanism to achieve an optimal sampling efficiency, and the second one is how to perform the reconstruction to get the highest quality to achieve an optimal signal recovery. In this paper, we try to deal with these two problems with a deep network. First of all, we train a sampling matrix via the network training instead of using a traditional manually designed one, which is much appropriate for our deep network based reconstruct process. Then, we propose a deep network to recover the image, which imitates traditional compressed sensing reconstruction processes. Experimental results demonstrate that our deep networks based CS reconstruction method offers a very significant quality improvement compared against state-of-the-art ones. Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Debin Zhao |
ICME | 1 |
| 2017 | A weighted full-reference image quality assessment based on visual saliency
Wuzhen Shi, Jiawei Chen 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Group-based sparse representation for low lighting image enhancementabstractThe Group-based Sparse Representation (GSR) is able to sparsely represent natural images in the domain of group, which enforces the intrinsic local sparsity and nonlocal self-similarity of images simultaneously in a unified framework. And the GSR-driven L0 minimization method for image restoration has been proposed. This paper expands the application of GSR from image restoration to low lighting image enhancement. The GSR is not used to represent the natural images anymore, but representing the transmission map of the haze image and recovering it. Because the transmission map is very important for the low lighting image enhancement, the dark channel prior based enhancement method with the enhanced transmission map can get a better enhanced results. Different from other methods, we evaluate the quality of the enhanced images not only by qualitative analysis but also by quantitative results. Extensive experiments on low lighting image show that the GSR-based method gets a better enhancement result than many current state-of-the-art ones. Wuzhen Shi, Feng Jiang 0001, Debin Zhao, Weizheng Shen |
ICIP | 1 |
| 2016 | Image Entropy of Primitive and visual quality assessmentabstractRecently, the concept of Entropy of Primitive (EoP) has been proposed to measure the image visual information. Some successful EoP based application also be developed. In this paper, we further explore the concept of EoP and propose an improved version: the L1 norm based EoP. Our EoP takes full account of the properties of a dictionary's layered structure and the characteristic of a basis pursuit method. Experimental results show that the L1 norm based EoP is superior to the L0 norm based one in measuring the image visual information. The curve of L1 norm based EoP holds a more consistent monotonicity with SSIM, its values is not trapped in the local convergence and the convergence value is less than that of the L0 norm based one. With the convergence characteristics of EoP, we further explore its application in stereoscopic image quality assessment (SIQA). With EoP as monocular cue and mutual information of primitive (MIP) as binocular cue, the relative entropy between the original stereoscopic image and the distorted one is used to compute the quality score by a prediction function which is trained using support vector regression (SVR). Extensive experimental results show that our new EoP based SIQA outperforms many state-of-the-art on the LIVE phase II databases. Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ICIP | 1 |
| 2015 | Reference image based method of region of interest enhancement for haze imageabstractDifferent from general algorithms of haze removal and low lighting image enhancement, which only use the information of image to process, this paper adds a reference image to get more information for the algorithm and focuses on enhancing region of interest of an image based on the reference one. With the reference image, the haze one can be divided into Region of Interest (RoI) and Region of no Interest (non-RoI). Furthermore, the reference image can provide more useful information for computing the transmission map and atmospheric light. For the non-RoI region, a more robust transmission map and minimizing reconstruction error cost function based method to estimate atmospheric light has been proposed. Because the atmospheric light is a global variable, the optimized one is also suitable for the RoI region. With the global optimized atmospheric light, an optimized transmission map can be got for the RoI region. The RoI region can be enhanced via the optimal transmission map and atmosphere light. Theoretical analysis gives eloquent proof proving that the proposed method is definitely better than the traditional dark-channel-prior-based methods due to our better transmission map and atmosphere light. Extensive experiments also show the expected results. Wuzhen Shi, Xinwei Gao, Boqi Chen, Feng Jiang 0001, Debin Zhao |
ICIP | 1 |
| 2014 | Improved appearance updating method in multiple instance learning trackingabstractMultiple instance learning (MIL) tracker becomes recently very popular because of their great success in complex scenes. Dynamically reflecting the appearance changes of the tracked object, the appearance updating plays an important role on tracking. In the original MIL tracker, the appearance model is assumed to obey normal distribution and its updating rule consists of a simple linearly weighted sum of the original and the current target distributions in the current frame. However, this updating method is not proved theoretically. In this work, the authors deduce a novel appearance updating method by estimating the mean and the variance of the sum of two normal distributions being merged in maximum likelihood estimation. The method can be naturally extended to multivariable distributions, useful to track colour object. Experimental results on some benchmark video sequences show that the method achieve higher precision and reliability than the three state‐of‐art trackers. Jifeng Ning, Wuzhen Shi, Shuqin Yang, Paul Yanne |
IET Comput. Vis. | 2 |
| 2013 | Visual tracking based on Distribution Fields and online weighted multiple instance learning
Jifeng Ning, Wuzhen Shi, Shuqin Yang, Paul Yanne |
Image Vis. Comput. | 2 |