VLDB 2026 Research / reviewers in the wild / expert
Zheming Lu 0001
dblp:148/1888-1 · also Zhe-Ming Lu 0001, Zhe-ming Lu 0001
· DBLP profile ↗
76ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0003-1785-7847ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 11 since 2021Security and privacy · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PSM: Prompt specialization module for prompt-based continual learning
Qinyue Tong, Zheming Lu 0001, Ziqian Lu |
Comput. Vis. Image Underst. | 3 |
| 2026 | Chat-MedGen: omni-adaptation of multi-modal large language models for diverse biomedical tasks
Qinyue Tong, Ziqian Lu, Zheming Lu 0001, Yunlong Yu 0001, Yang-Ming Zheng |
Knowl. Based Syst. | 3 |
| 2026 | Improving anomaly detection with foundation-model synthesis and wavelet-domain attention
Wensheng Wu, Zheming Lu 0001, Ziqian Lu, Zewei He, Xuecheng Sun, Jungong Han, Yunlong Yu 0001 |
Neural Networks | 2 |
| 2025 | UltraFusion: Ultra High Dynamic Imaging using Exposure FusionabstractCapturing high dynamic range (HDR) scenes is one of the most important issues in camera design. Majority of cameras use exposure fusion, which fuses images captured by different exposure levels, to increase dynamic range. However, this approach can only handle images with limited exposure difference, normally 3-4 stops. When applying to very high dynamic range scenes where a large exposure difference is required, this approach often fails due to incorrect alignment or inconsistent lighting between inputs, or tone mapping artifacts. In this work, we propose UltraFusion, the first exposure fusion technique that can merge inputs with 9 stops differences. The key idea is that we model exposure fusion as a guided inpainting problem, where the under-exposed image is used as a guidance to fill the missing information of over-exposed highlights in the over-exposed region. Using an under-exposed image as a soft guidance, instead of a hard constraint, our model is robust to potential alignment issue or lighting variations. Moreover, by utilizing the image prior of the generative model, our model also generates natural tone mapping, even for very high-dynamic range scenes. Our approach outperforms HDR-Transformer on latest HDR benchmarks. Moreover, to test its performance in ultra high dynamic range scenes, we capture a new real-world exposure fusion benchmark, UltraFusion dataset, with exposure differences up to 9 stops, and experiments show that UltraFusion can generate beautiful and high-quality fusion results under various scenarios. Code and data will be available at https://openimaginglab.github.io/UltraFusion. Yujin Wang 0001, Zhiyuan You, Zheming Lu 0001, Shi Guo, Tianfan Xue |
CVPR | 5 |
| 2025 | MediSee: Reasoning-Based Pixel-Level Perception in Medical ImagesabstractDespite progress in pixel-level medical image perception, existing methods remain task-specific or depend on precise prompts like bounding boxes or text. However, the need for medical knowledge limits accessibility for the general public, who are more likely to use logically reasoned oral queries than domain-specific inputs. In this paper, we introduce a novel medical vision task: Medical Reasoning Segmentation and Detection (MedSD), which aims to comprehend implicit queries about medical images and generate the corresponding segmentation mask and bounding box for the target object. To accomplish this task, we first introduce a Multi-perspective, Logic-driven Medical Reasoning Segmentation and Detection (MLMR-SD) dataset, which encompasses a substantial collection of medical entity targets along with their corresponding reasoning. Furthermore, we propose MediSee, an effective baseline model designed for MedSD. The experimental results indicate that the proposed method can effectively address MedSD with implicit colloquial queries and outperform traditional medical referring segmentation methods. The MediSee project can be found here. Qinyue Tong, Ziqian Lu, Yangming Zheng, Zheming Lu 0001 |
ACM Multimedia | 5 |
| 2025 | Progressive Multi-Prompt Learning for Vision-Language ModelsabstractRecently, methods that utilize prompt tuning to rapidly transfer pretrained vision-language models (VLMs) to downstream tasks have been proposed. Although these models have produced reasonable results, they typically learn a single prompt, which limits their ability to capture more diverse information. This ability is crucial for addressing fine-grained classification challenges and intraclass visual variability (e.g., color, pose, and size variations within the same category). However, learning multiple prompts provides a larger optimization space, which further exacerbates the overfitting phenomenon. This makes it more challenging balance the performances acienved for base and new categories. To address these issues, we propose progressive multi-prompt (PMP) learning method.ecently, methods that utilize prompt tuning to rapidly transfer pretrained vision-language models (VLMs) to downstream tasks have been proposed. Although these models have produced reasonable results, they typically learn a single prompt, which limits their ability to capture more diverse information. This ability is crucial for addressing fine-grained classification challenges and intraclass visual variability (e.g., color, pose, and size variations within the same category). However, learning multiple prompts provides a larger optimization space, which further exacerbates the overfitting phenomenon. This makes it more challenging balance the performances acienved for base and new categories. To address these issues, we propose progressive multiprompt (PMP) learning method.R Specifically, we introduce multiple prompts in a step-by-step manner to focus on various information. To reduce overfitting, we utilize alate attachingmechanism to defer the interactions of prompts and features to a deeper encoding layer. Furthermore, we balance the prompts for different layers with learnable weights to guide the optimal optimization procedure. We compared our method with several state-of-the-art approaches in base-to-new task settings and demonstrate superior base-new tradeoff performance. Additionally, we conducted cross-dataset transfer, domain generalization, and few-shot experiments to further validate the effectiveness of our method. Our code is available at https://github.com/JunLGeek/PMP.git. Ziqian Lu, Hao Luo 0001, Zheming Lu 0001, Yangming Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Protecting the Copyright of Intelligent Transportation Systems Based on Zernike MomentsabstractIntelligent transportation systems are at risk of data misuse. Watermarking data that needs to be opened and shared can mitigate this problem, i.e., embedding signals into generated images, which are imperceptible to humans. Regardless of the shape of the image or any potential attacks it may undergo, the watermark can be detected by algorithms when needed. However, many watermarking schemes fail to resist attacks like cropping and translation, limiting their applicability in the intelligent transportation domain. To tackle these issues, we propose a dual watermarking framework based on Zernike moments for intelligent transportation systems, where a robust watermark and a periodic watermark are embedded in different planes of the cover image. Specifically, we propose a single-circle model (SCM) where Zernike moments are locally computed based on a circle centered at the image center with a radius proportional to image size for embedding the robust watermark. Since SCM is determined by measuring its center and radius, SCM is applicable to images of various sizes and shapes. Then we employ a robust combination of discrete wavelet transform (DWT) and discrete cosine transform (DCT) watermarking algorithms to embed the periodic watermark. For watermark extraction, we propose an efficient adaptive correction mechanism (AC) to recognize attack types and automatically relocate the embedding position of the watermark. By combining the above strategies, our proposed scheme can adaptively resist various attacks (e.g., random cropping, translation), which addresses the shortcomings of most existing watermarking schemes. We test the watermark using over 300 images of different sizes and shapes, and the experimental results prove that our proposed scheme achieves stronger robustness against various distortions with better invisibility, outperforming the state-of-the-art (SOTA) methods. Jiale Meng, Zheming Lu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Prompt-Based Test-Time Real Image Dehazing: A Novel Pipeline
Zewei He, Ziqian Lu, Xuecheng Sun, Zheming Lu 0001 |
ECCV (76) | 5 |
| 2024 | Pixel Matching Network for Cross-Domain Few-Shot SegmentationabstractFew-Shot Segmentation (FSS) aims to segment the novel class images with a few annotated samples. In the past, numerous studies have concentrated on cross-category tasks, where the training and testing sets are derived from the same dataset, while these methods face significant difficulties in domain-shift scenarios. To better tackle the cross-domain tasks, we propose a pixel matching network (PMNet) to extract the domain-agnostic pixel-level affinity matching with a frozen backbone and capture both the pixel-to-pixel and pixel-to-patch relations in each support-query pair with the bidirectional 3D convolutions. Different from the existing methods that remove the support background, we design a hysteretic spatial filtering module (HSFM) to filter the background-related query features and retain the foreground-related query features with the assistance of the support background, which is beneficial for eliminating interference objects in the query background. We comprehensively evaluate our PMNet on ten benchmarks under cross-category, cross-dataset, and cross-domain FSS tasks. Experimental results demonstrate that PMNet performs very competitively under different settings with only 0.68M parameters, especially under cross-domain FSS tasks, showing its effectiveness and efficiency. Code will be released at: https://github.com/chenhao-zju/PMNet Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Jungong Han |
WACV | 3 |
| 2024 | Hierarchical contrastive representation for zero shot learning
Ziqian Lu, Zheming Lu 0001, Zewei He, Xuecheng Sun, Hao Luo 0001, Yangming Zheng |
Appl. Intell. | 2 |
| 2024 | Dense affinity matching for Few-Shot Segmentation
Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Yingming Li, Jungong Han, Zhongfei Zhang |
Neurocomputing | 3 |
| 2024 | Self-supervised graph representations with generative adversarial learning
Xuecheng Sun, Zonghui Wang, Zheming Lu 0001, Ziqian Lu |
Neurocomputing | 3 |
| 2024 | Learning Multiple Criteria Calibration for Generalized Zero-shot Learning
Ziqian Lu, Zheming Lu 0001, Yunlong Yu 0001, Zewei He, Hao Luo 0001, Yangming Zheng |
Knowl. Based Syst. | 2 |
| 2024 | Multipurpose video watermarking algorithm for copyright protection and tamper detection
Kai An, Zheming Lu 0001, Xue-Cheng Sun, Zong-Hui Wang |
Multim. Tools Appl. | 2 |
| 2024 | A novel feature for action recognition
Zheming Lu 0001, Jia-Lin Cui, Hao-Lai Li |
Multim. Tools Appl. | 2 |
| 2024 | Self-Prompting Perceptual Edge Learning for Dense PredictionabstractNumerous studies have employed prompt learning structures to enhance dense prediction tasks by integrating additional semantic or geometric information. While the inclusion of extra information has shown improvements in performance, it also poses challenges for applications that cannot provide extra input. To address this issue, this study evaluates the performance of different prompts and introduces an additional-input-free method, called self-prompting perceptual edge learning (SPPEL), which extracts edge-embedded semantic prompts directly from the image feature itself using trainable handcrafted edge operators within a plug-and-play module. To obtain the edge features, our approach incorporates an adversarial structure that compares the similarity between two edge features generated by the Hog and Kirsch operators, where the edge features are measured using multiplication, finetuned through a trainable all-one embedding, and enhanced with channel-to-channel attention. We conduct extensive evaluations of SPPEL on 7 tasks, utilizing 7 different backbones and applying 5 distinct methods. Our experimental results demonstrate that SPPEL achieves strong competitiveness in various settings with an average improvement of 1.7% across all 7 tasks, including ADE20K, COCO (Instance Segmentation), COCO (Object Detection), Pascal VOC2012, STARE, CHASE DB1, and HRF, while incurring a parameter increase of less than 3% (the detailed computation analysis of parameters and Gflops are shown in different experimental tables). Code will be released at: https://github.com/chenhao-zju/sppel. Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DEA-Net: Single Image Dehazing Based on Detail-Enhanced Convolution and Content-Guided AttentionabstractSingle image dehazing is a challenging ill-posed problem which estimates latent haze-free images from observed hazy images. Some existing deep learning based methods are devoted to improving the model performance via increasing the depth or width of convolution. The learning ability of Convolutional Neural Network (CNN) structure is still under-explored. In this paper, a Detail-Enhanced Attention Block (DEAB) consisting of Detail-Enhanced Convolution (DEConv) and Content-Guided Attention (CGA) is proposed to boost the feature learning for improving the dehazing performance. Specifically, the DEConv contains difference convolutions which can integrate prior information to complement the vanilla one and enhance the representation capacity. Then by using the re-parameterization technique, DEConv is equivalently converted into a vanilla convolution to reduce parameters and computational cost. By assigning the unique Spatial Importance Map (SIM) to every channel, CGA can attend more useful information encoded in features. In addition, a CGA-based mixup fusion scheme is presented to effectively fuse the features and aid the gradient flow. By combining above mentioned components, we propose our Detail-Enhanced Attention Network (DEA-Net) for recovering high-quality haze-free images. Extensive experimental results demonstrate the effectiveness of our DEA-Net, outperforming the state-of-the-art (SOTA) methods by boosting the PSNR index over 41 dB with only 3.653 M parameters. (The source code of our DEA-Net is available at https://github.com/cecret3350/DEA-Net.). Zewei He, Zheming Lu 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Multi-Content Interaction Network for Few-Shot SegmentationabstractFew-Shot Segmentation (FSS) poses significant challenges due to limited support images and large intra-class appearance discrepancies. Most existing approaches focus on aligning the support-query correlations from the same layer of the frozen backbone while neglecting the bias between different tasks and different layers. In this article, we propose a Multi-Content Interaction Network (MCINet) to remedy these issues by fully exploiting and interacting with the different contextual information contained in distinct branches. Specifically, MCINet improves FSS from three perspectives: (1) boosting the query representations through incorporating the independent information from another learnable branch into the features from the frozen backbone, (2) enhancing the support-query correlations by exploiting both the same-layer and adjacent-layer features, and (3) refining the predicted results with a multi-scale mask prediction strategy. Experiments on three benchmarks demonstrate that our approach reaches state-of-the-art performances and outperforms the best competitors with many desirable advantages, especially on the challenging COCO dataset. Code will be released on GitHub ( https://github.com/chenhao-zju/mcinet ). Hao Chen 0107, Yunlong Yu 0001, Yonghan Dong, Zheming Lu 0001, Yingming Li, Zhongfei Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Single image super-resolution based on progressive fusion of orientation-aware features
Zewei He, Yanpeng Cao, Jiangxin Yang, Yanlong Cao, Xin Li 0003, Siliang Tang, Yueting Zhuang, Zheming Lu 0001 |
Pattern Recognit. | 9 |
| 2022 | Learn more from less: Generalized zero-shot learning with severely limited labeled data
Ziqian Lu, Zheming Lu 0001, Yunlong Yu 0001, Zonghui Wang |
Neurocomputing | 2 |
| 2022 | Dual semantic-guided model for weakly-supervised zero-shot semantic segmentation
Zheming Lu 0001, Ziqian Lu, Zonghui Wang |
Multim. Tools Appl. | 2 |
| 2021 | A geometrically robust multi-bit video watermarking algorithm based on 2-D DFT
Xue-Cheng Sun, Zheming Lu 0001, Yong-Liang Liu |
Multim. Tools Appl. | 2 |
| 2020 | Acceleration of multi-task cascaded convolutional networksabstractMulti‐task cascaded convolutional neural network (MTCNN) is a human face detection architecture which uses a cascaded structure with three stages (P‐Net, R‐Net and O‐Net). The authors intend to reduce the computation time of the whole process of the MTCNN. They find that the non‐maximum suppression (NMS) processes after the P‐Net occupy over half of the computation time. Therefore, the authors propose a self‐fine‐tuning method which makes the control of computation time for the NMS process easier. Self‐fine‐tuning is a training trick which uses hard samples generated by P‐Net to retrain P‐Net. After self‐fine‐tuning, the distribution of human face probabilities generated by P‐Net is changed, and the tail of distribution becomes thinner. The control of the number of NMS input boxes can be made easier when the distribution has a thinner tail, and choosing a suitable threshold to filter the face boxes will generate less boxes. So the computation time can be reduced. In order to keep the performance of MTCNN, the authors still propose a landmark data set augmentation, which can enhance the performance of the self‐fine‐tuned MTCNN. From the experiments, it is found that the proposed scheme can significantly reduce the computation time of MTCNN. Longhua Ma, Hang-Yu Fan, Zheming Lu 0001, Dong Tian |
IET Image Process. | 3 |
| 2020 | Reinforcement Learning-Based Control for Nonlinear Discrete-Time Systems with Unknown Control Directions and Control Constraints
Miao Huang, Xiaoqi He, Longhua Ma, Zheming Lu 0001 |
Neurocomputing | 5 |
| 2020 | A low-frequency construction watermarking based on histogram
Hang-Yu Fan, Zheming Lu 0001, Yongxiang Liu |
Multim. Tools Appl. | 2 |
| 2020 | An attentional spatial temporal graph convolutional network with co-occurrence feature learning for action recognition
Dong Tian, Zheming Lu 0001, Longhua Ma |
Multim. Tools Appl. | 2 |
| 2020 | Face recognition based on local binary pattern and improved Pairwise-constrained Multiple Metric Learning
Lijian Zhou, Shanshan Lin, Siyuan Hao, Zheming Lu 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Face feature extraction and recognition via local binary pattern and two-dimensional locality preserving projection
Lijian Zhou, Wanquan Liu, Zheming Lu 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Fast codebook design method for image vector quantisationabstractVector quantisation (VQ) is a widely used method for data compression or data clustering. This study proposes a fast codebook design method for image VQ. General particle swarm optimisation (PSO)‐ K ‐means hybrid methods for codebook design take too much time. To deal with this problem, the authors present a new clustering method, which takes a very short time and generates a pretty good codebook. This method contains four parts, i.e. pooling, PSO algorithm, reverse cast, and K ‐means algorithm. Pooled images are used for the PSO process while original images are used for K ‐means fine‐tuning, and these two processes connected by the reverse cast process. In the authors’ experiment, this method can dramatically reduce the computation time by using big sizes of pooling window or enhance the codebook quality by using small sizes of pooling window. They can make the calculation time almost one‐tenth of that of the PSO‐ K ‐means hybrid method, and the scheme is also faster than the K ‐means algorithm. Experimental results demonstrate that the main advantages of the proposed algorithm lie in the fact that it can reduce the computation time and enhance the quality of codebooks. Hang-Yu Fan, Zheming Lu 0001 |
IET Image Process. | 2 |
| 2018 | Anomaly detection and localisation in the crowd scenes using a block-based social force modelabstractA novel approach to detect and localise anomalous events in crowed scenes by processing surveillance videos is introduced in this study. Unusual events are those that significantly differ from current dominated behaviours. The proposed approach both detects pixel‐level and block‐level anomalies. In pixel level, Gaussian mixture models are used to detect abnormalities. Block‐level detection segments the crowd into blocks according to pedestrian detection, and then anomalies are spotted and localised with a social force model. Experimental results using the USCD datasets Ped1 and Ped2 show that the proposed method performs favourably against state‐of‐the‐art methods. Qing-Ge Ji, Rui Chi, Zheming Lu 0001 |
IET Image Process. | 3 |
| 2018 | Hierarchical palmprint feature extraction and recognition based on multi-wavelets and complex networkabstractThis study presents a hierarchical palmprint feature extraction and recognition approach based on multi‐wavelet and complex network (CN) since they can effectively decrease redundant information and enhance key points of main lines and wrinkles. The palmprint is first pre‐filtered and decomposed once using multi‐wavelet. Three components (LL 1,2,3 ) corresponding to the pre‐filter except for diagonal component are extracted as the elementary features. Second, binary images (BLL 1,2,3 ) are obtained by the average window method using different thresholds. Third, three series of dynamic evolution CN models (the 1st, 2nd, 3rd CNs) are constructed from global to local, which is based on the mosaiced images obtained from BLL 1,2,3 , BLL 1 and four equally divided sub‐images of BLL 1 , respectively. Fourth, statistical features are extracted from complex networks, in which average degree and standard deviation of the degrees are extracted for the 1st CNs and average degrees are extracted for the 2nd and 3rd CNs. Fifth, the fisher feature is extracted using the linear discriminate analysis method. Finally, the nearest neighbourhood classifier is used to recognise palmprint. Based on the CASIA Palmprint Image Database, experimental results show that the proposed method can effectively recognise palmprint with good robustness and overcome the problem of small training samples number. Lijian Zhou, Zuowei Wang 0002, Zheming Lu 0001 |
IET Image Process. | 5 |
| 2018 | Optimal blind watermarking for color images based on the U matrix of quaternion singular value decomposition
Longhua Ma, Zheming Lu 0001 |
Multim. Tools Appl. | 4 |
| 2018 | VQ codebook design using modified K-means algorithm with feature classification and grouping based initialization
Zheming Lu 0001, Longhua Ma, Ya-Pei Feng |
Multim. Tools Appl. | 2 |
| 2018 | Inter-frame passive-blind forgery detection for video shot based on similarity analysis
Dongning Zhao, Ren-Kui Wang, Zheming Lu 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Real-time multi-feature based fire flame detection in videoabstractIn this study, the authors present a new approach to detect fire flame by processing and analysing the stationary camera videos. For a fire detection system, it is desired to be sensitive and reliable. The proposed method improves not only the sensitivity but also the reliability through reducing the susceptibility to false alarms. The proposed approach based on multi‐feature, i.e. chromatic features, dynamic features, texture features, and contour features, can both improve the sensitivity and reliability in fire detection. In their approach, the authors adopt a novel algorithm to extract the moving region and analyse the frequency of flickers. Experimental results show that the proposed method can run in real‐time and performs favourably against the state‐of‐the‐art methods with higher accuracy in fire videos, lower false alarm rates in non‐fire videos and faster response time. Rui Chi, Zheming Lu 0001, Qing-Ge Ji |
IET Image Process. | 2 |
| 2017 | Image coding based on classified vector quantisation using edge orientation patternsabstractVector quantisation (VQ) shows a good performance for image coding with high‐compression ratios. However, there are many difficulties for image coding with VQ, especially the edge degradation and high‐computational complexity. To resolve these two problems, the authors propose a new coding method based on edge orientation patterns (EOPs) by classifying image blocks into nine classes according to their edge orientations. For colour image coding, 27 codebooks (nine for each colour component) are pre‐designed based on a series training images. In the encoding stage, an input colour image is decomposed into Y, Cb, and Cr components, and each component image is divided into non‐overlapping 4 × 4 blocks. For each block, eight edge orientation templates of size 4 × 4 are performed to determine its edge orientation. According to the edge orientation, each block is compressed by using the corresponding codebook. Essentially, the authors’ scheme is a kind of classified VC (CVQ). Simulation results show that, their EOP‐based CVQ can largely improve the compression efficiency as well as speeding up the encoding process and it is sufficient to establish effectiveness of the authors’ algorithm as compared with the existing techniques. Ya-Pei Feng, Zheming Lu 0001 |
IET Image Process. | 2 |
| 2017 | Online and offline based load balance algorithm in cloud computing
Linlin Tang, Zuohua Li, Pingfei Ren, Jeng-Shyang Pan 0001, Zheming Lu 0001, Jing-Yong Su, Zhenyu Meng |
Knowl. Based Syst. | 5 |
| 2014 | Neighbourhood sensitive preserving embedding for pattern classificationabstractRecently, a large family of supervised or unsupervised manifold learning algorithms that stem from statistical or geometrical theory has been designed to solve the problem of pattern classification. In this study, consider the fact that the data are usually sampled from a low‐dimensional manifold space which resides in a high‐dimensional Euclidean space, the authors propose a novel two‐graph‐based supervised linear classification algorithm called neighbourhood sensitive preserving embedding (NSPE). Different from local linear embedding (LLE) (or neighbourhood preserving embedding (NPE)) which preserves the local neighbourhood structure with one graph, NSPE can discover both the intrinsic and discriminant structure of the data manifold by constructing two graphs, that is, the within‐class graph and the between‐class graph. Thus, the data are mapped into a subspace where the nearby points with the same label are close to each other, whereas the nearby points with different labels are far apart. As a classification method, besides being defined on training samples, NSPE is also defined on testing samples. Experiments carried on the real‐world face databases demonstrate that the results of all two‐graph‐based spectral methods are comparable and better than that of one‐graph‐based methods. Binghui Wang, Chuang Lin 0001, Xue-Feng Zhao, Zheming Lu 0001 |
IET Image Process. | 4 |
| 2014 | Face recognition based on curvelets and local binary pattern features via using local property preservation
Lijian Zhou, Wanquan Liu, Zheming Lu 0001, Tingyuan Nie |
J. Syst. Softw. | 3 |
| 2013 | An improved VLC-based lossless data hiding scheme for JPEG images
Yongjian Hu, Zheming Lu 0001 |
J. Syst. Softw. | 3 |
| 2013 | A high capacity lossless data hiding scheme for JPEG images
Zheming Lu 0001, Yongjian Hu |
J. Syst. Softw. | 2 |
| 2013 | Video abstraction based on the visual attention model and online clustering
Qing-Ge Ji, Zhi-Dang Fang, Zhen-Hua Xie, Zheming Lu 0001 |
Signal Process. Image Commun. | 4 |
| 2013 | Fast Video Shot Boundary Detection Based on SVD and Pattern MatchingabstractVideo shot boundary detection (SBD) is the first and essential step for content-based video management and structural analysis. Great efforts have been paid to develop SBD algorithms for years. However, the high computational cost in the SBD becomes a block for further applications such as video indexing, browsing, retrieval, and representation. Motivated by the requirement of the real-time interactive applications, a unified fast SBD scheme is proposed in this paper. We adopted a candidate segment selection and singular value decomposition (SVD) to speed up the SBD. Initially, the positions of the shot boundaries and lengths of gradual transitions are predicted using adaptive thresholds and most non-boundary frames are discarded at the same time. Only the candidate segments that may contain the shot boundaries are preserved for further detection. Then, for all frames in each candidate segment, their color histograms in the hue-saturation-value) space are extracted, forming a frame-feature matrix. The SVD is then performed on the frame-feature matrices of all candidate segments to reduce the feature dimension. The refined feature vector of each frame in the candidate segments is obtained as a new metric for boundary detection. Finally, cut and gradual transitions are identified using our pattern matching method based on a new similarity measurement. Experiments on TRECVID 2001 test data and other video materials show that the proposed scheme can achieve a high detection speed and excellent accuracy compared with recent SBD schemes. Zheming Lu 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | An oblivious fragile watermarking scheme for images utilizing edge transitions in BTC bitmaps
Yong Zhang 0023, Zheming Lu 0001, Dongning Zhao |
Sci. China Inf. Sci. | 2 |
| 2012 | Robust Image Hashing Based on Random Gabor Filtering and Dithered Lattice Vector QuantizationabstractIn this paper, we propose a robust-hash function based on random Gabor filtering and dithered lattice vector quantization (LVQ). In order to enhance the robustness against rotation manipulations, the conventional Gabor filter is adapted to be rotation invariant, and the rotation-invariant filter is randomized to facilitate secure feature extraction. Particularly, a novel dithered-LVQ-based quantization scheme is proposed for robust hashing. The dithered-LVQ-based quantization scheme is well suited for robust hashing with several desirable features, including better tradeoff between robustness and discrimination, higher randomness, and secrecy, which are validated by analytical and experimental results. The performance of the proposed hashing algorithm is evaluated over a test image database under various content-preserving manipulations. The proposed hashing algorithm shows superior robustness and discrimination performance compared with other state-of-the-art algorithms, particularly in the robustness against rotations (of large degrees). Yuenan Li 0001, Zheming Lu 0001, Ce Zhu, Xiamu Niu |
IEEE Trans. Image Process. | 2 |
| 2010 | A Novel Embedded Coding Algorithm Based on the Reconstructed DCT Coefficients
Linlin Tang, Jeng-Shyang Pan 0001, Zheming Lu 0001 |
ICCCI (3) | 3 |
| 2009 | Video Identification Using Spatio-temporal Salient PointsabstractAutomatic identification of video is a key technique to various applications such as content based retrieval, copyright protection and broadcast monitoring. In this paper, we present a novel video identification scheme via spatio-temporal salient points. To achieve higher robustness to content preserved manipulations, especially geometric distortions, salient points are detected from both spatial and temporal aspects. The salient points are found to be invariant in a series of distortions. In addition, we have compared our approach with a state-of-the-art one on diverse sequences and superior performances on detection accuracy have been observed in the proposed work. Yue-Nan Li 0002, Zheming Lu 0001 |
IAS | 2 |
| 2009 | A CELP-Speech Information Hiding Algorithm Based on Vector QuantizationabstractThis paper presents a speech information-hiding scheme that is integrated with the CELP (Coded Exited Linear Prediction) speech coding method. The G.729 codec (CS-ACELP) is used to test the efficiency of this scheme. The index-constrained method is applied for the secret information embedding. The selected bits of first-stage and residual vector indices during the Predictive Two-stage Vector Quantization procedure are modulated by the secret information bits. For the imperceptibility of this scheme, the quality of the watermarked speech signal can be controlled by adaptive embedding through referring to the original and watermarked speech. Experimental results show that this scheme can be effective and the distortion of the speech signal is trivial and imperceptible. Zheming Lu 0001, Hao Luo 0001 |
IAS | 2 |
| 2009 | A Novel Multiple Description Coding Frame Based on Reordered DCT Coefficients and SPIHT AlgorithmabstractA novel MDC algorithm based on discrete cosine transform (DCT) and the set partition in hierarchical trees (SPIHT) is proposed in this paper. Different from the commonly used DCT algorithm, all the transformed coefficients are reshaped into the wavelet decomposition structure to facilitate the use of the SPIHT algorithm. Then the direction-based information is used to form the three different channels. By using different bit rates to encode the information from three different orientations, i.e., vertical, horizontal and diagonal directions, the redundancy is introduced into the three channels. Every channel contains the hybrid information from three different directions. Experimental results show the advantages of this novel algorithm and the theoretical analysis has also been studied. Linlin Tang, Zheming Lu 0001, Faxin Yu |
IAS | 2 |
| 2009 | Fast video shot boundary detection framework employing pre-processing techniquesabstractVideo shot boundary detection is the initial and fundamental step towards video indexing, browsing and retrieval. Great efforts have been paid on developing accurate shot boundary detection algorithms. However, the high computational cost in shot detection becomes a bottleneck for real-time applications. The problem of making a balance between detection accuracy and speed is addressed in this paper, and a novel fast detection framework is presented. The general framework that employs pre-processing techniques can improve both detection speed and precision. In the pre-processing stage, adaptive local thresholding is adopted to classify non-boundary segments and candidate segments that may contain shot boundaries. The candidate segments are refined using bisection-based comparisons to eliminate non-boundary frames. Only refined candidate segments are preserved for further detections; hence, the speed of shot detection is improved by reducing detection scope. Moreover, prior knowledge about each possible shot boundary such as its type and duration can be obtained in the pre-processing stage, which can accelerate the consequent hard cut and gradual transition detections. Experimental results indicate that the proposed framework is effective in accelerating the shot detection process, and it can also achieve excellent detection accuracies. Yue-Nan Li 0002, Zheming Lu 0001, Xia-Mu Niu |
IET Image Process. | 2 |
| 2009 | Learning nonlinear manifolds based on mixtures of localized linear manifolds under a self-organizing framework
Huicheng Zheng, Qionghai Dai, Sanqing Hu, Zheming Lu 0001 |
Neurocomputing | 5 |
| 2009 | A path optional lossless data hiding scheme based on VQ joint neighboring coding
Jun-Xiang Wang, Zheming Lu 0001 |
Inf. Sci. | 2 |
| 2009 | An improved lossless data hiding scheme based on image VQ-index residual value coding
Zheming Lu 0001, Jun-Xiang Wang |
J. Syst. Softw. | 1 |
| 2009 | Kernel optimization-based discriminant analysis for face recognition
Junbao Li, Jeng-Shyang Pan 0001, Zheming Lu 0001 |
Neural Comput. Appl. | 3 |
| 2009 | Face recognition using Gabor-based complete Kernel Fisher Discriminant analysis with fractional power polynomial models
Junbao Li, Jeng-Shyang Pan 0001, Zheming Lu 0001 |
Neural Comput. Appl. | 3 |
| 2009 | Robust dual watermarking algorithm for AVS video
Yuan-Gen Wang, Zheming Lu 0001, Liang Fan |
Signal Process. Image Commun. | 2 |
| 2008 | Adaptive quasiconformal kernel discriminant analysis
Jeng-Shyang Pan 0001, Junbao Li, Zheming Lu 0001 |
Neurocomputing | 3 |
| 2007 | Spectrum Shaped Dither Modulation Watermarking for Correlated Host SignalabstractSpectrum shaping of watermark is necessary when the host signal is colored. This paper proposes a spectrum shaped dither modulation watermarking scheme. The covariance matrix of the watermark is scaled version of the host covariance matrix, so the corresponding spectrum of the watermark is proportional to the spectrum of the host signal. The performance of the proposed system is analyzed in terms of average probability of bit decoding error; closed form expression is derived. Theoretical analysis is finally validated by Monte Carlo simulation. Bin Yan 0001, Zheming Lu 0001, Sheng-He Sun |
ICME | 2 |
| 2007 | Watermarking-Based Transparency Authentication in Visual CryptographyabstractThis paper proposes two transparency authentication schemes used in visual cryptography, which are based on watermarking techniques. In the first scheme, a secret image can be perceptible when stacking two transparencies. In addition, another watermark image is visible when shifting one transparency to an appropriate position and stacking it with the other transparency. In the second scheme, both the secret image and the watermark are of the same size, while in the first scheme, the watermark is half of the size of the secret image. No computer aid is needed in decryption in the first scheme while simple computation is required in the second one. The secret image and the watermark are encrypted at the same time in the two schemes. Experimental results demonstrate the two schemes are effective and practical. Hao Luo 0001, Jeng-Shyang Pan 0001, Zheming Lu 0001, Bin-Yih Liao |
ISDA | 3 |
| 2007 | High Capacity Reversible Data Hiding for 3D Meshes in the PVQ Domain
Zheming Lu 0001 |
IWDW | 1 |
| 2007 | Multiple Watermarking in Visual Cryptography
Hao Luo 0001, Zheming Lu 0001, Jeng-Shyang Pan 0001 |
IWDW | 2 |
| 2007 | Hadamard transform based fast codeword search algorithm for high-dimensional VQ encoding
Shu-Chuan Chu 0001, Zheming Lu 0001, Jeng-Shyang Pan 0001 |
Inf. Sci. | 2 |
| 2006 | Complete Kernel Fisher discriminant analysis of Gabor features with fractional power polynomial models for face recognitionabstractThis paper presents a novel face recognition method based on complete Kernel Fisher discriminant (CKFD) analysis of Gabor features with power polynomial models. By integrating the Gabor wavelet representation of face images and the enhanced powerful discriminator named CKFD analysis, the method is robust to changes in illumination and facial expressions and poses. On the other hand, the extended polynomial Kernels, namely fractional power polynomial (FPP) models, are employed in CKFD analysis, which enhance face recognition performance. Comparing with existing PCA, LDA, KPCA, KFD and CKFD methods, the proposed method gives superior results in the ORL and Yale face databases. Its good performance in the two face databases gives the promising idea to solve the pose, illumination, and expression (PIE) problem of face recognition Junbao Li, Jeng-Shyang Pan 0001, Zheming Lu 0001, Jung-Chou Harry Chang |
ISCAS | 3 |
| 2006 | Reversible Watermarking for Error Diffused Halftone Images Using Statistical Features
Zheming Lu 0001, Hao Luo 0001, Jeng-Shyang Pan 0001 |
IWDW | 1 |
| 2006 | Security of autoregressive speech watermarking model under guessing attackabstractThe security of the "autoregressive (AR) watermark in AR host" signal model is investigated. It is demonstrated through analysis and Monte Carlo simulation that the AR watermarking model is asymptotically as secure as the "white watermark in white host" model under the guessing attack Bin Yan 0001, Zheming Lu 0001, Sheng-He Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2005 | Image Retrieval Based on a Multipurpose Watermarking Scheme
Zheming Lu 0001, Henrik Skibbe, Hans Burkhardt |
KES (2) | 1 |
| 2005 | Side-Match Predictive Vector Quantization
Yue-Nan Li 0002, Zheming Lu 0001 |
KES (3) | 3 |
| 2005 | Compressed Domain Video Watermarking in Motion Vector
Hao-Xian Wang, Yue-Nan Li 0002, Zheming Lu 0001, Sheng-He Sun |
KES (2) | 3 |
| 2005 | Speech Authentication by Semi-fragile Watermarking
Bin Yan 0001, Zheming Lu 0001, Sheng-He Sun, Jeng-Shyang Pan 0001 |
KES (3) | 2 |
| 2005 | Multipurpose image watermarking algorithm based on multistage vector quantizationabstractThe rapid growth of digital multimedia and Internet technologies has made copyright protection, copy protection, and integrity verification three important issues in the digital world. To solve these problems, the digital watermarking technique has been presented and widely researched. Traditional watermarking algorithms are mostly based on discrete transform domains, such as the discrete cosine transform, discrete Fourier transform (DFT), and discrete wavelet transform (DWT). Most of these algorithms are good for only one purpose. Recently, some multipurpose digital watermarking methods have been presented, which can achieve the goal of content authentication and copyright protection simultaneously. However, they are based on DWT or DFT. Lately, several robust watermarking schemes based on vector quantization (VQ) have been presented, but they can only be used for copyright protection. In this paper, we present a novel multipurpose digital image watermarking method based on the multistage vector quantizer structure, which can be applied to image authentication and copyright protection. In the proposed method, the semi-fragile watermark and the robust watermark are embedded in different VQ stages using different techniques, and both of them can be extracted without the original image. Simulation results demonstrate the effectiveness of our algorithm in terms of robustness and fragility. Zheming Lu 0001, Dianguo Xu 0001, Sheng-He Sun |
IEEE Trans. Image Process. | 1 |
| 2004 | Hadamard transform based equal-average equal-variance equal-norm nearest neighbor codeword search algorithmabstractThe work presents a novel, efficient, nearest-neighbor codeword search algorithm based on three elimination criteria in the Hadamard transform (HT) domain. Before the search process, all codewords in the codebook are Hadamard-transformed and sorted in the ascending order of their first elements. During the search process, we first perform the HT on the input vector and calculate its variance and norm, and secondly exploit three efficient elimination criteria to find the nearest codeword to the input vector using the up-down search mechanism near the initial best-match codeword. Experimental results demonstrate that the performance of the proposed algorithm is much better than that of most existing nearest neighbor codeword search algorithms, especially in the case of high dimension. Shu-Chuan Chu 0001, John F. Roddick, Zheming Lu 0001, Jeng-Shyang Pan 0001 |
ICME | 3 |
| 2003 | An efficient encoding algorithm for vector quantization based on subvector techniqueabstractIn this paper, a new and fast encoding algorithm for vector quantization is presented. This algorithm makes full use of two characteristics of a vector: the sum and the variance. A vector is separated into two subvectors: one is composed of the first half of vector components and the other consists of the remaining vector components. Three inequalities based on the sums and variances of a vector and its two subvectors components are introduced to reject those codewords that are impossible to be the nearest codeword, thereby saving a great deal of computational time, while introducing no extra distortion compared to the conventional full search algorithm. The simulation results show that the proposed algorithm is faster than the equal-average nearest neighbor search (ENNS), the improved ENNS, the equal-average equal-variance nearest neighbor search (EENNS) and the improved EENNS algorithms. Comparing with the improved EENNS algorithm, the proposed algorithm reduces the computational time and the number of distortion calculations by 2.4% to 6% and 20.5% to 26.8%, respectively. The average improvements of the computational time and the number of distortion calculations are 4% and 24.6% for the codebook sizes of 128 to 1024, respectively. Jeng-Shyang Pan 0001, Zheming Lu 0001, Sheng-He Sun |
IEEE Trans. Image Process. | 2 |
| 2001 | Vector quantization based on genetic simulated annealing
Hsiang-Cheh Huang, Jeng-Shyang Pan 0001, Zheming Lu 0001, Sheng-He Sun, Hsueh-Ming Hang |
Signal Process. | 3 |
| 2000 | VQ Image Coding Using Sub-Vector TechniquesabstractA fast codeword search algorithm for vector quantization is presented. Three inequalities based on the characteristics of sums and variances of a vector and its two sub-vectors are introduced to reject a larger number of codewords. Simulation results demonstrate the effectiveness of the proposed algorithm. Jeng-Shyang Pan 0001, Zheming Lu 0001, Sheng-He Sun |
ICIP | 2 |
| 2000 | A fast image coding algorithm using variable-rate mean-match correlation vector quantizationabstractA variable-rate mean-match correlation vector quantizer, MMCVQ, is presented as an alternative vector quantizer of side-match vector quantizer (SMVQ) for fast image encoding. Before encoding, the mean values of all codewords are computed, then a sorted codebook is obtained according to the ascending order of the mean values of codewords. During the encoding stage, high correlation of the adjacent image blocks and mean values of the current processing vector and codewords are considered to obtain the state-codebook of the current processing vector. The experimental result shows: (1) the encoding time of MMCVQ is almost 7 times faster than that of full-search VQ and almost 42 times faster than that of SMVQ, and (2) compared with SMVQ, the encoding quality will increase more than 1 dB in PSNR, and (3) the bit rate of MMCVQ is smaller than that of SMVQ by 0.03-0.05 b/pixel. Jeng-Shyang Pan 0001, Zheming Lu 0001, Sheng-He Sun |
KES | 2 |
| 1999 | Non-redundant VQ channel coding using modified tabu search approach with simulated annealingabstractCodeword Index Assignment (CIA) is a key issue to vector quantization (VQ). A new algorithm called Modified Tabu Search Algorithm (MTSA) is applied to codeword index assignment for noisy channels for the purpose of minimizing the distortion due to bit errors. Simulated annealing (SA) technique and a new parameter are introduced in the Tabu Search Approach (TSA) to improve the performance of the tabu search approach. Experimental tests show the modified tabu search algorithm is superior to the tabu search algorithm by evaluating the performance of channel distortion after the same number of iterations. Jeng-Shyang Pan 0001, Zheming Lu 0001, Shu-Chuan Chu 0001, Sheng-He Sun |
KES | 2 |