VLDB 2026 Research / reviewers in the wild / expert
Jie Liu 0028
dblp:03/2134-28
· DBLP profile ↗
21ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-8204-424XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Keypoint-enhanced image watermarking with spatial-frequency mapping and perceptual optimization
Fei Ge, Jie Liu 0028, Guixuan Zhang, Shuwu Zhang, Hu Guan |
Inf. Sci. | 4 |
| 2026 | Deep Compression on Segment Anything Model for Efficient Industrial ManufacturingabstractSegment Anything Model (SAM) is a popular vision foundation model that can segment data from any domain. Benefiting from its outstanding generalization ability, SAM has been widely adopted in many industrial scenarios. However, as SAM is built upon a heavy Vision Transformer (ViT), it suffers from memory-hungry and low latency, which restricts the deployment on edge devices. In this paper, we systematically explore how to compress SAM effectively, making it feasible to adapt edge devices with limited calculation abilities. Specifically, our method consists of three aspects: weight initialization, knowledge distillation, and model quantization. It is notable that all three aspects are not simply inherited from previous methods, but tactfully designed based on the teacher-student learning paradigm, considering the task-attributes of SAM pre-training. Firstly, we design a weight initialization method for fully using the pre-training knowledge implicitly contained in the teacher’s parameter space. Secondly, based on the weight initialization, we design a novel distillation method tailored to SAM pre-training, focusing on learning the semantic differences among areas. Lastly, we perform quantization on our distilled models. Unlike the previous method, we used both the teacher and the student to calibrate our model in the quantization process. We conduct systematic experiments on various teacher-student network pairs to validate the broad effectiveness of our method. By applying our method, our target models achieve over 64.5× speed increase compared to the original SAM. Core code is available at: https://github.com/ZG-ZZ/DC-SAM. Yang Zheng 0002, Jie Liu 0028, Qing Li 0015, Jiangyun Li, Zhenghao Xi |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Enhancing Semantic Information Representation in Multi-View Geo-Localization through Dual-Branch Network with Feature Consistency Enhancement and Multi-Level Feature MiningabstractABSTRACT Metric learning is fundamental to multi‐view geo‐localization, as it aims to establish a distance metric that minimizes the feature space distance between similar data points while maximizing the separation between dissimilar ones. However, in Siamese networks employed for metric learning, individual branches may exhibit discrepancies in their interpretation of semantic information from input data, resulting in semantically inconsistent feature representations. To address this issue, a method is designed to enhance significant region consistency within multi‐view spaces by integrating feature consistency enhancement (FCE) and multi‐level feature mining (MLFM) techniques into a dual‐branch network. The FCE method emphasizes critical components of the input data, ensuring feature consistency between the two branches. Additionally, the MLFM mechanism facilitates feature integration across multiple levels, thereby enabling a more comprehensive extraction of semantic information. This approach enhances semantic understanding and promotes feature consistency across branches. The proposed method achieves AP values of 82.38% for drone‐to‐satellite and 77.36% for satellite‐to‐drone image matching. Notably, the method maintains computational efficiency without significantly affecting inference time. Additionally, improvements are observed in R@1, R@5 and R@10 metrics. The experimental results show that integrating FCE and MLFM into the dual‐branch network improves semantic representation and outperforms existing methods. Yang Zheng 0002, Qing Li 0015, Jiangyun Li, Zhenghao Xi, Jie Liu 0028 |
IET Image Process. | 5 |
| 2024 | A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
Zeyu Zhao 0005, Nan Gao 0001, Guixuan Zhang, Jie Liu 0028, Shuwu Zhang |
MMAsia | 5 |
| 2024 | Degradation regression with uncertainty for blind super-resolution
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang |
Neurocomputing | 4 |
| 2023 | Learning Video Localization on Segment-Level Video Copy Detection with Transformer
Chi Zhang 0093, Jie Liu 0028, Shuwu Zhang |
ICANN (7) | 2 |
| 2023 | Adversarial Audio Watermarking: Embedding Watermark into Deep FeatureabstractAudio watermarking is a promising technology for copyright protection, yet traditional methods are limited that must be combined with auxiliary techniques against attacks. This article proposes a new audio watermarking method that embeds watermarks through a trained neural network. It adds small imperceptible perturbations to the original audio so that its deep features point to specific watermark features. Data augmentation and error correcting coding are employed to guarantee its practicable robustness. This method is robust against many attacks without auxiliary techniques and shows better performance than other deep learning-based methods. Shiqiang Wu, Jie Liu 0028, Hu Guan, Shuwu Zhang |
ICME | 2 |
| 2023 | Robust Texture-Aware Local Adaptive Image Watermarking With Perceptual GuaranteeabstractWatermarking involves embedding a watermark in an image and later extracting it to prove the image’s copyright. In most cases, a complete image contains both smooth and textured regions. As a rule of thumb, the visual quality of an image with a watermark embedded in its textured regions is better than that of the same image with a watermark in smooth regions. This paper, by taking advantage of the fact, proposes a texture-aware local adaptive watermarking algorithm to maximize the watermark’s robustness while maintaining its imperceptibility. To identify textured regions in an image, we introduce the texture value, an efficient and proper metric of the richness of image texture. It combines the texture correlation of the AC coefficients, the luminance masking of the DC coefficient, and the distribution of image texture. A watermark is embedded adaptively into multiple non-overlapping textured regions of an image under the specified SSIM condition. Its adaptiveness comes from a novel texture-aware adaptive parameter model derived by multivariate regression analysis. Correct extraction of watermarks from multiple textured regions can be done by the cooperation of embedding and extraction strategies, with the assistance of RS-based watermark coding model. They allow for greater robustness, faster extraction, and adjustable watermark capacity. The simulation experiments on 100 images demonstrate that our proposed algorithm outperforms state-of-the-art algorithms with respect to imperceptibility, robustness, and adaptability. Hu Guan, Jie Liu 0028, Shuwu Zhang, Baoning Niu, Guixuan Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | From general to specific: Online updating for blind super-resolution
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang |
Pattern Recognit. | 4 |
| 2021 | Approaching the Limit of Image Rescaling via Flow Guidance
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang |
BMVC | 4 |
| 2021 | Learning to predict more accurate text instances for scene text detection
Jie Liu 0028, Guixuan Zhang, Yang Zheng 0002, Shuwu Zhang |
Neurocomputing | 2 |
| 2020 | IBN-STR: A Robust Text Recognizer for Irregular Text in Natural ScenesabstractAlthough text recognition methods based on deep neural networks have promising performance, there are still challenges due to the variety of text styles, perspective distortion, text with large curvature, and so on. To obtain a robust text recognizer, we have improved the performance from two aspects: data aspect and feature representation aspect. In terms of data, we transform the input images into S-shape distorted images in order to increase the diversity of training data. Besides, we explore the effects of different training data. In terms of feature representation, the combination of instance normalization and batch normalization improves the model's capacity and generalization ability. This paper proposes a robust scene text recognizer IBN-STR, which is an attention-based model. Through extensive experiments, the model analysis and comparison have been carried out from the aspects of data and feature representation, and the effectiveness of IBN-STR on both regular and irregular text instances has been verified. Furthermore, IBN-STR is an end-to-end recognition system that can achieve state-of-the-art performance. Jie Liu 0028, Guixuan Zhang, Shuwu Zhang |
ICPR | 2 |
| 2020 | Single shot multi-oriented text detection based on local and non-local features
Jie Liu 0028, Shuwu Zhang, Guixuan Zhang, Yang Zheng 0002 |
Int. J. Document Anal. Recognit. | 2 |
| 2018 | Aspect-Level Sentiment Classification with Conv-Attention Mechanism
Jie Liu 0028, Guixuan Zhang, Shuwu Zhang |
ICONIP (4) | 2 |
| 2017 | A cascaded method for text detection in natural scene images
Yang Zheng 0002, Qing Li 0015, Jie Liu 0028, Heping Liu, Shuwu Zhang |
Neurocomputing | 3 |
| 2016 | Scene text detection with extremal region based cascaded filteringabstractIn this paper, we present a robust Extremal Region (ER) based scene text detection system. To eliminate the vast non-text components generated by ER operator, a three-stage cascaded filter is proposed. In the first stage, a powerful character classifier enhanced by recursive local search is introduced to separate text components from noises. Then, an efficient heuristic pruning method is designed to further clean overlapped duplicate characters. Finally, after text line construction, a cascaded text line classification model integrating word entropy and sliding window based CNN is proposed to remove false text lines. Experiments on benchmarks show that our method achieves state-of-the-art performance. Jie Liu 0028, Shuwu Zhang, Yang Zheng 0002 |
ICIP | 2 |
| 2011 | Hierarchical Latent Dirichlet Allocation models for realistic action recognitionabstractIt has always been very difficult to recognize realistic actions from unconstrained videos because there are tremendous variations from camera motion, background clutter, object appearance and so on. In this paper, a Single-Feature Hierarchical Latent Dirichlet Allocation model called SF-HLDA by extending Latent Dirichlet Allocation to the hierarchical one is first proposed for realistic action recognition. And then, by extending SF-HLDA, we present another model called Multi-Feature Hierarchical Latent Dirichlet Allocation model MF-HLDA which can effectively fuse several different features into one model for recognizing the realistic actions. Experiments demonstrate the effectiveness of our proposed models. Heping Li, Jie Liu 0028, Shuwu Zhang |
ICASSP | 2 |
| 2011 | A Novel Italic Detection and Rectification Method for Chinese Advertising ImagesabstractThe italic detection and slant rectification is a key step of optical character recognition (OCR). In this paper, a novel method is proposed to detect and rectify italic characters in Chinese advertising images. Based on observations on structures of many characters, the centroid angle is proposed and a statistical study on it is presented. According to the statistical results, the centroid angle of a Chinese character approximately obeys a Gaussian distribution with its slant angle. Moreover, a Markov Random Field (MRF) model, considering the font-face similarity of neighboring characters and the strong correlation between the centroid angle and the slant angle of a character, is then presented to estimate the slant angle of a character. The italic characters can be detected and rectified by the estimated angle. The experimental results demonstrate the proposed method is effective and applicable. Jie Liu 0028, Heping Li, Shuwu Zhang, Wei Liang 0009 |
ICDAR | 1 |
| 2011 | A Chinese Character Localization Method Based on Intergrating Structure and CC-Clustering for Advertising ImagesabstractIn this paper, a novel Chinese character localization method is proposed for texts in advertising images. To deal with the texts with gradient color, a color clustering method based on edge is introduced to separate the color image into homogeneous color layers. To solve the problem of locating characters varied in size, style and arranged in irregular direction, a novel character localization method is proposed, which integrates structure and CC-clustering to locate characters according to reliable features of characters. Finally, a new noise removal method based on stroke width histogram is employed to remove all non-characters connected components, and then all characters are located. The experimental results show that the proposed method can effectively locate characters in advertising images. Jie Liu 0028, Shuwu Zhang, Heping Li, Wei Liang 0009 |
ICDAR | 1 |
| 2011 | Automatic behavior model selection by iterative learning and abnormality recognitionabstractAutomatic behavior recognition is one important task of community security and surveillance system. In this paper, a novel method is proposed for automatic selection of behavior models by iterative learning and abnormality recognition. The method is mainly composed of the following two steps: (1) The models of normal behaviors are automatically selected and trained by combining Dynamic Time Warping based spectral clustering and iterative learning; (2) Maximum A Posteriori adaptation technique is used to estimate the parameters of abnormal behavior models from those of normal behavior models. Compared with the related works in the literature, our method has three advantages: (1) automatic selection of the class number of normal behaviors from large unlabeled video data according to the process of iterative learning, (2) semi-supervised learning of abnormal behavior models, and (3) avoidance of the running risk of over-fitting during learning the Hidden Markov Models of behaviors in case of sparse data. Experiments demonstrate the effectiveness of our proposed method. Heping Li, Jie Liu 0028, Shuwu Zhang |
ISI | 2 |
| 2010 | A Fast Image Inpainting Method Based on Hybrid Similarity-DistanceabstractA fast image in painting method based on hybrid similarity-distance is proposed in this paper. In Criminisi et al.'s work, similarity distance are not reliable enough in many cases and the algorithm performs inefficiently. To solve these problems, we propose a new searching strategy to accelerate the algorithm. In addition, we modify the confidence-updating rule to make more reasonable the distributions of the confidences in source region. Besides, taking account of the stationarity of texture and the reliability of the source regions, we present a hybrid similarity-distance, which combines the distance in color space with the distance in spatial space by weight coefficients related to the confidence value. A more reasonable patch will be found out by this hybrid similarity-distance. The experiments verify that the proposed method yields qualitative improvements compared to Criminisi et al.'s work. Jie Liu 0028, Shuwu Zhang, Wuyi Yang, Heping Li |
ICPR | 1 |