Mengzhu Yu

dblp:264/5353 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A two-stage multimodal learning framework based on text-driven vision pretraining and cross-modal feature fusion for thyroid ultrasound diagnosis
Mengzhu Yu, Tianwei Yan 0004, Zihan Xi, Junchao Zeng, Mingyue Ding
Expert Syst. Appl.1
2024 Unifying Pictorial and Textual Features for Screen Content Image Quality Evaluation
abstract
Dividing a Screen Content Image (SCI) with complex components into pictorial and textual regions for predicting scores is one of the common Screen Content Image Quality Assessment (SCIQA) methods. However, how to efficiently leverage pictorial and textual features to predict quality scores for no-reference SCIQA still needs to be explored. In addition, statistical analysis reveals that labels of SCIs present a distribution. Therefore, both the distribution of quality scores of SCIQA and the distribution of labels need to be considered in the SCIQA. This paper proposes a no-reference SCIQA method unifying pictorial and textual features. One contribution is the proposed dual-branch extraction module with the parameter-free attention convolution block and the joint prediction module. The proposed method employs the dual-branch extraction module to generate efficient pictorial and textual features and then uses the joint prediction module to predict quality scores. Another contribution is the joint distribution loss. It makes the distribution of the quality scores as close as possible to the distribution of labels. Experiments on the SCIQA datasets show that the proposed method achieves excellent SCIQA performance and generalization ability.
Yihua Chen 0001, Xiaoping Liang, Mengzhu Yu, Zhenjun Tang
ICMR3
2024 Robust Video Hashing with Non-negative Tensor Factorization for Copy Detection
abstract
Copy detection is a key task of video copyright protection. This paper presents a robust video hashing with non-negative tensor factorization (NTF) for copy detection. In the presented video hashing scheme, secondary frames are computed from the preprocessed video by assigning weights to all frames within a video group based on color entropy. Next, the secondary frames are fed into the pre-trained MobileNetV2 and then NTF is exploited to compress the three-order tensor constructed by stacking the output feature maps for hash construction. Experiments conducted on publicly available video datasets indicate that the presented hashing scheme outperforms the evaluated hashing schemes in the performances of classification and copy detection.
Mengzhu Yu, Zhenjun Tang, Huijiang Zhuang, Xiaoping Liang, Zhixin Li 0001, Xianquan Zhang
ICMR1
2024 Video Hashing with Tensor Robust PCA and Histogram of Optical Flow for Copy Detection
abstract
Abstract This paper proposes a novel video hashing with tensor robust Principal Component Analysis (PCA) and Histogram of Optical Flow (HOF) for copy detection. In the proposed hashing, a video is divided into some video groups. For each video group, a low-rank secondary frame is constructed from the low-rank component decomposed by applying tensor robust PCA to the video group. Since the low-rank component can well indicate spatial-temporal intrinsic structure of the video group and it is slightly disturbed by digital operations, feature extraction from the low-rank secondary frames is discriminative and stable. Next, spatial features and temporal features are extracted from low-rank secondary frames by Charlier moments and HOF, respectively. Since the Charlier moments are robust to geometric transform and they can efficiently distinguish video frames with different contents, the use of Charlier moments can make robust and discriminative spatial features. As the HOF can measure the distribution of motion information between frames, the temporal features formed by HOFs can provide good discrimination. Hash is ultimately determined by quantizing the spatial and temporal features and concatenating the quantized results. Numerous experiments on open video datasets indicate that the proposed hashing is superior to some hashing baseline schemes in terms of classification and copy detection.
Mengzhu Yu, Zhenjun Tang, Hanyun Zhang, Xiaoping Liang, Xianquan Zhang
Comput. J.1
2024 Robust Hashing With Local Tangent Space Alignment for Image Copy Detection
abstract
Robust hashing is a useful technique for the image applications of watermarking, authentication, quality assessment and copy detection. This paper proposes a new robust hashing for image copy detection by using local tangent space alignment (LTSA). A key contribution is the weighted visual map computation based on the difference of Gaussian (DOG) and visual attention model. The weighted visual map can provide the proposed method with good robustness. Another contribution is the feature learning via LTSA from the feature matrix of the weighted visual map in discrete cosine transform domain. As it can maintain the local geometric relationships within image, the learned features can make the proposed method discriminative. Extensive experiments on public databases are conducted to validate the proposed robust hashing method. Compared with some famous robust hashing methods, the proposed robust hashing method demonstrates preferable classification performance in terms of discrimination and robustness. Copy detection performance is tested and the result verifies effectiveness of the proposed robust hashing method.
Xiaoping Liang, Zhenjun Tang, Xianquan Zhang, Mengzhu Yu, Xinpeng Zhang 0001
IEEE Trans. Dependable Secur. Comput.4
2024 Robust Hashing via Global and Local Invariant Features for Image Copy Detection
abstract
Robust hashing is a powerful technique for processing large-scale images. Currently, many reported image hashing schemes do not perform well in balancing the performances of discrimination and robustness, and thus they cannot efficiently detect image copies, especially the image copies with multiple distortions. To address this, we exploit global and local invariant features to develop a novel robust hashing for image copy detection. A critical contribution is the global feature calculation by gray level co-occurrence moment learned from the saliency map determined by the phase spectrum of quaternion Fourier transform, which can significantly enhance discrimination without reducing robustness. Another essential contribution is the local invariant feature computation via Kernel Principal Component Analysis (KPCA) and vector distances. As KPCA can maintain the geometric relationships within image, the local invariant features learned with KPCA and vector distances can guarantee discrimination and compactness. Moreover, the global and local invariant features are encrypted to ensure security. Finally, the hash is produced via the ordinal measures of the encrypted features for making a short length of hash. Numerous experiments are conducted to show efficiency of our scheme. Compared with some well-known hashing schemes, our scheme demonstrates a preferable classification performance of discrimination and robustness. The experiments of detecting image copies with multiple distortions are tested and the results illustrate the effectiveness of our scheme.
Xiaoping Liang, Zhenjun Tang, Zhixin Li 0001, Mengzhu Yu, Hanyun Zhang, Xianquan Zhang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Robust Hashing with Deep Features and Meixner Moments for Image Copy Detection
abstract
Copy detection is a key task of image copyright protection. Most robust hashing schemes do not make satisfied performance of image copy detection yet. To address this, a robust hashing scheme with deep features and Meixner moments is proposed for image copy detection. In the proposed hashing, global deep features are extracted by applying tensor Singular Value Decomposition (t-SVD) to the three-order tensor constructed in the DWT domain of the feature maps calculated by the pre-trained VGG16. Since the feature maps in the DWT domain are slightly disturbed by digital operations, the constructed three-order tensor is stable and thus the desirable robustness is guaranteed. Moreover, since t-SVD can decompose a three-order tensor into multiple low-dimensional matrices reflecting intrinsic structure, the global deep feature calculation from the low-dimensional matrices can provide good discrimination. Local features are calculated by the block-based Meixner moments. As the Meixner moments are resistant to geometric transformation and can efficiently discriminate various images, the use of the block-based Meixner moments can make discriminative and robust local features. Hash is ultimately determined by quantifying and combining global deep features and local features. The results of extensive experiments on open image datasets demonstrate that the proposed robust hashing outperforms some state-of-the-art robust hashing schemes in terms of classification and copy detection performances.
Mengzhu Yu, Zhenjun Tang, Xiaoping Liang, Xianquan Zhang, Zhixin Li 0001, Xinpeng Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Dual-Feature Aggregation Network for No-Reference Image Quality Assessment
Yihua Chen 0001, Mengzhu Yu, Zhenjun Tang
MMM (1)3
2023 Robust Image Hashing With Saliency Map And Sparse Model
abstract
Abstract Image hashing is an effective technology for extensive image applications, such as retrieval, authentication and copy detection. This paper designs a new image hashing scheme based on saliency map and sparse model. The major contributions are twofold. The first contribution is the construction of a weighted image representation by combining a visual attention model called Itti model and the matrix of color vector angle (CVA). Since the Itti model can efficiently detect saliency map and CVA fully captures color information of image, they contribute to a visually robust and discriminative image representation. The second contribution is the hash extraction from the weighted image representation via sparse model. A classical sparse model called robust principal component analysis is exploited to decompose the weighted image representation into a low-rank component and a sparse component. As the low-rank component can describe intrinsic structure of image, hash calculation with low-rank component can achieve good discrimination. The efficiencies of the proposed scheme are validated by extensive experiments with open databases. The results demonstrate that the proposed scheme is superior to some state-of-the-art schemes in terms of classification performance between robustness and discrimination.
Mengzhu Yu, Zhenjun Tang, Zhixin Li 0001, Xiaoping Liang, Xianquan Zhang
Comput. J.1
2022 Perceptual Hashing With Complementary Color Wavelet Transform and Compressed Sensing for Reduced-Reference Image Quality Assessment
abstract
Image quality assessment (IQA) is an important task of image processing and has diverse applications, such as image super-resolution reconstruction, image transmission and monitoring systems. This paper proposes a perceptual hashing algorithm with complementary color wavelet transform (CCWT) and compressed sensing (CS) for reduced-reference (RR) IQA. The CCWT is exploited to decompose input color image into different sub-bands. Since the calculation of CCWT uses all color channels without discarding any information, the distortions introduced by digital operations on color channels are preserved in the CCWT sub-bands. The block-based CS is used to extract features from the CCWT sub-bands. As the Euclidean distance between the block-based CS features is slightly influenced by content-preserving operations, perceptual features constructed by Euclidean distances are robust, discriminative and compact. Hash sequence is finally determined by quantifying the perceptual features. Effectiveness of the proposed hashing is verified by various experiments on four open image databases. Experimental results demonstrate that the proposed hashing is superior to some state-of-the-art algorithms in terms of classification and RR IQA application.
Mengzhu Yu, Zhenjun Tang, Xianquan Zhang, Bineng Zhong 0001, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Robust Image Hashing With Singular Values Of Quaternion SVD
abstract
Abstract Image hashing is an efficient technique of many multimedia systems, such as image retrieval, image authentication and image copy detection. Classification between robustness and discrimination is one of the most important performances of image hashing. In this paper, we propose a robust image hashing with singular values of quaternion singular value decomposition (QSVD). The key contribution is the innovative use of QSVD, which can extract stable and discriminative image features from CIE L*a*b* color space. In addition, image features of a block are viewed as a point in the Cartesian coordinates and compressed by calculating the Euclidean distance between its point and a reference point. As the Euclidean distance requires smaller storage than the original block features, this technique helps to make a discriminative and compact hash. Experiments with three open image databases are conducted to validate efficiency of our image hashing. The results demonstrate that our image hashing can resist many digital operations and reaches a good discrimination. Receiver operating characteristic curve comparisons illustrate that our image hashing outperforms some state-of-the-art algorithms in classification performance.
Zhenjun Tang, Mengzhu Yu, Heng Yao 0001, Hanyun Zhang, Chunqiang Yu, Xianquan Zhang
Comput. J.2
2020 Robust image hashing with visual attention model and invariant moments
abstract
Image hashing is an efficient technique of multimedia processing for many applications, such as image copy detection, image authentication, and social event detection. In this study, the authors propose a novel image hashing with visual attention model and invariant moments. An important contribution is the weighted DWT (discrete wavelet transform) representation by incorporating a visual attention model called Itti saliency model into LL sub‐band. Since the Itti saliency model can efficiently extract saliency map reflecting regions of attention focus, perceptual robustness of the proposed hashing is achieved. In addition, as invariant moments are robust and discriminative features, hash construction with invariant moments extracted from the weighted DWT representation ensures good classification performance between robustness and discrimination. Extensive experiments with open image datasets are done to validate the performances of the proposed hashing. The results demonstrate that the proposed hashing is robust and discriminative. Performance comparisons with some hashing algorithms are also conducted, and the receiver operating characteristic results illustrate that the proposed hashing outperforms the compared hashing algorithms in classification performance between robustness and discrimination.
Zhenjun Tang, Hanyun Zhang, Chi-Man Pun, Mengzhu Yu, Chunqiang Yu, Xianquan Zhang
IET Image Process.4