VLDB 2026 Research / reviewers in the wild / expert
Hongjian Zhan
dblp:187/9093
· DBLP profile ↗
32ranked-venue papers
9as first author
26since 2021 · last 2025
0000-0002-3906-658XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MSA2: Multi-Task Framework With Structure-Aware and Style-Adaptive Character Representation for Open-Set Chinese Text Recognition
Yangfu Li, Hongjian Zhan, Yujie Xiong, Yue Lu 0001 |
ICCV | 2 |
| 2025 | A Local Perceptual Approach for Few-Shot Text Effect Transfer
Hongjian Zhan, Tian Wei, Yue Lu 0001 |
ICIG (3) | 1 |
| 2025 | Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced CoveringabstractExisting visual token pruning methods target prompt alignment and visual preservation with static strategies, overlooking the varying relative importance of these objectives across tasks, which leads to inconsistent performance. To address this, we derive the first closed-form error bound for visual token pruning based on the Hausdorff distance, uniformly characterizing the contributions of both objectives. Moreover, leveraging $\epsilon$-covering theory, we reveal an intrinsic trade-off between these objectives and quantify their optimal attainment levels under a fixed budget. To practically handle this trade-off, we propose Multi-Objective Balanced Covering (MoB), which reformulates visual token pruning as a bi-objective covering problem. In this framework, the attainment trade-off reduces to budget allocation via greedy radius trading. MoB offers a provable performance bound and linear scalability with respect to the number of input visual tokens, enabling adaptation to challenging pruning scenarios. Extensive experiments show that MoB preserves 96.4\% of performance for LLaVA-1.5-7B using only 11.1\% of the original visual tokens and accelerates LLaVA-Next-7B by 1.3-1.5$\times$ with negligible performance loss. Additionally, evaluations on Qwen2-VL and Video-LLaVA confirm that MoB integrates seamlessly into advanced MLLMs and diverse vision-language tasks. The code will be made available soon. Yangfu Li, Hongjian Zhan, Yujie Xiong, Yue Lu 0001 |
NeurIPS | 2 |
| 2025 | Length-aware center loss for sequence to sequence Thai scene text recognition
Hongjian Zhan, Yue Lu 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | VGTS: Visually Guided Text Spotting for novel categories in historical manuscriptsabstractIn the field of historical manuscript research, scholars frequently encounter novel symbols in ancient texts, investing considerable effort in their identification and documentation. Although existing object detection methods achieve impressive performance on known categories, they struggle to recognize novel symbols without retraining. To address this limitation, we propose a Visually Guided Text Spotting (VGTS) approach that accurately spots novel characters using just one annotated support sample. The core of VGTS is a spatial alignment module consisting of a Dual Spatial Attention (DSA) block and a Geometric Matching (GM) block. The DSA block aims to identify, focus on, and learn discriminative spatial regions in the support and query images, mimicking the human visual spotting process. It first refines the support image by analyzing inter-channel relationships to identify critical areas, and then refines the query image by focusing on informative key points. The GM block, on the other hand, establishes the spatial correspondence between the two images, enabling accurate localization of the target character in the query image. To tackle the example imbalance problem in low-resource spotting tasks, we develop a novel torus loss function that enhances the discriminative power of the embedding space for distance metric learning . To further validate our approach, we introduce a new dataset featuring ancient Dongba hieroglyphics (DBH) associated with the Naxi minority of China. Extensive experiments on the DBH dataset and other public datasets, including Egyptian Hieroglyph (EGY), Historical Arabic Documents (HAD), Tripitaka Koreana in Han (TKH), and Notary Charters (NC), show that VGTS consistently surpasses state-of-the-art methods. The proposed framework exhibits great potential for application in historical manuscript text spotting, enabling scholars to efficiently identify and document novel symbols with minimal annotation effort. Wenbo Hu 0008, Hongjian Zhan, Xinchen Ma, Cong Liu 0006, Yue Lu 0001, Ching Y. Suen |
Expert Syst. Appl. | 2 |
| 2025 | Diff-TST: Diffusion model for one-shot text-image style transfer
Sizhe Pang, Yangchen Xie, Hongjian Zhan, Yue Lu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Spotting the Unseen: Reciprocal Consensus Network Guided by Visual ArchetypesabstractHumans often require only a few visual archetypes to spot novel objects. Based on this observation, we present a strategy rooted in ``spotting the unseen" by establishing dense correspondences between potential query image regions and a visual archetype, and we propose the Consensus Network (CoNet). Our method leverages relational patterns intra and inter images via Auto-Correlation Representation (ACR) and Mutual-Correlation Representation (MCR). Within each image, the ACR module is capable of encoding both local self-similarity and global context simultaneously. Between the query and support images, the MCR module computes the cross-correlation across two image representations and introduces a reciprocal consistency constraint, which can incorporate to exclude outliers and enhance model robustness. To overcome the challenges of low-resource training data, particularly in one-shot learning scenarios, we incorporate an adaptive margin strategy to better handle diverse instances. The experimental results indicate the effectiveness of the proposed method across diverse domains such as object detection in natural scenes, and text spotting in both historical manuscripts and natural scenes, which demonstrates its sparkling generalization ability. Our code is available at: https://github.com/infinite-hwb/conet. Wenbo Hu 0008, Hongjian Zhan, Xinchen Ma, Yue Lu 0001, Ching Y. Suen |
AAAI | 2 |
| 2024 | Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning NetworkabstractScene text recognition is inherently a vision-language task. However, previous works have predominantly focused either on extracting more robust visual features or designing better language modeling. How to effectively and jointly model vision and language to mitigate heavy reliance on a single modality remains a problem. In this paper, aiming to enhance vision-language reasoning in scene text recognition, we present a balanced, unified and synchronized vision-language reasoning network (BUSNet). Firstly, revisiting the image as a language by balanced concatenation along length dimension alleviates the issue of over-reliance on vision or language. Secondly, BUSNet learns an ensemble of unified external and internal vision-language model with shared weight by masked modality modeling (MMM). Thirdly, a novel vision-language reasoning module (VLRM) with synchronized vision-language decoding capacity is proposed. Additionally, BUSNet achieves improved performance through iterative reasoning, which utilizes the vision-language prediction as a new language input. Extensive experiments indicate that BUSNet achieves state-of-the-art performance on several mainstream benchmark datasets and more challenge datasets for both synthetic and real training data compared to recent outstanding methods. Code and dataset will be available at https://github.com/jjwei66/BUSNet. Jiajun Wei, Hongjian Zhan, Yue Lu 0001, Xiao Tu, Cong Liu 0006, Umapada Pal 0001 |
AAAI | 2 |
| 2024 | FaRE: A Feature-Aware Radical Encoding Strategy for Zero-Shot Chinese Character Recognition
Hongjian Zhan, Yangfu Li, Yujie Xiong, Yue Lu 0001 |
ACCV (1) | 1 |
| 2024 | RSTAN: Residual Spatio-Temporal Attention Network for End-to-End Human Fall Detection
Yaru Jiang, Shujing Lyu, Hongjian Zhan, Yue Lu 0001 |
ICPR (15) | 3 |
| 2024 | Learning to Detect Lithography Defects in SEM Images
Hu Lu, Botong Zhao, Jiwei Shen, Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICPR (5) | 4 |
| 2024 | LK-Net: Efficient Large Kernel ConvNet for Document Enhancement
Qijun Shi, Hongjian Zhan, Yangfu Li, Weijun Zou, Huasheng Li, Umapada Pal 0001, Yue Lu 0001 |
ICPR (21) | 2 |
| 2024 | Free Lunch: Frame-level Contrastive Learning with Text Perceiver for Robust Scene Text Recognition in Lightweight Models
Hongjian Zhan, Yangfu Li, Yujie Xiong, Umapada Pal 0001, Yue Lu 0001 |
ACM Multimedia | 1 |
| 2024 | LRATNet: Local-Relationship-Aware Transformer Network for Table Structure Recognition
Guangjie Yang, Dajian Zhong, Yujie Xiong, Hongjian Zhan |
MMM (2) | 4 |
| 2024 | TANet: Text region attention learning for vehicle re-identification
Wenbo Hu 0008, Hongjian Zhan, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Enhanced video clustering using multiple riemannian manifold-valued descriptors and audio-visual information
Wenbo Hu 0008, Hongjian Zhan, Yinghong Tian, Yujie Xiong, Yue Lu 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Weakly supervised scene text generation for low-resource languages
Yangchen Xie, Hongjian Zhan, Palaiahnakote Shivakumara, Cong Liu 0006, Yue Lu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | NDOrder: Exploring a novel decoding order for scene text recognition
Dajian Zhong, Hongjian Zhan, Shujing Lyu, Cong Liu 0006, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Adaptive watermarking with self-mutual check parameters in deep neural networks
Zhenzhe Gao, Zhao-Xia Yin, Hongjian Zhan, Yue Lu 0001 |
Pattern Recognit. Lett. | 3 |
| 2024 | DS-TDNN: Dual-Stream Time-Delay Neural Network With Global-Aware Filter for Speaker VerificationabstractConventional time-delay neural networks (TDNNs) struggle to handle long-range context, their ability to represent speaker information is therefore limited for long utterances. Existing solutions either depend on increasing model complexity or try to strike a balance between local features and global context to address this issue. To effectively leverage the long-term dependencies of audio signals and constrain model complexity, we introduce a novel module called Global-aware Filter layer (GF layer) in this work, which employs a set of learnable transform-domain filters between a 1D discrete Fourier transform and its inverse transform to capture global context. Additionally, we develop a dynamic filtering strategy and a sparse regularization method to enhance the performance of the GF layer and prevent overfitting. Based on the GF layer, we present a dual-stream TDNN architecture called DS-TDNN for automatic speaker verification (ASV), which utilizes two unique branches to extract both local and global features in parallel and employs an efficient strategy to fuse different-scale information. Experiments on the Voxceleb and SITW databases demonstrate that the DS-TDNN achieves a relative improvement of 10% together with a relative decline of 20% in computational cost over the ECAPA-TDNN in the speaker verification task. This improvement becomes more evident as the utterance's duration grows. Furthermore, the DS-TDNN also beats popular deep residual models and attention-based systems on utterances of arbitrary length. Yangfu Li, Jiapan Gan, Xiaodan Lin, Yingqiang Qiu, Hongjian Zhan, Hui Tian 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Scene Text Recognition with Image-Text Matching-Guided Dictionary
Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu 0001, Umapada Pal 0001 |
ICDAR (6) | 2 |
| 2023 | Table Structure Recognition of Historical Dongba Documents
Hongjian Zhan, Xiao Tu, Yue Lu 0001 |
ICIG (1) | 2 |
| 2023 | 2C2S: A two-channel and two-stream transformer based framework for offline signature verification
Jian-Xin Ren, Yujie Xiong, Hongjian Zhan, Bo Huang 0014 |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | SANet-SI: A new Self-Attention-Network for Script Identification in scene images
Hongjian Zhan, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Thai Scene Text Recognition with Character Combination
Hongjian Zhan, Yue Lu 0001 |
PRCV (3) | 2 |
| 2021 | DenseNet-CTC: An end-to-end RNN-free architecture for context-free string recognition
Hongjian Zhan, Shujing Lyu, Yue Lu 0001, Umapada Pal 0001 |
Comput. Vis. Image Underst. | 1 |
| 2019 | A Handwritten Chinese Text Recognizer Applying Multi-level Multimodal Fusion NetworkabstractHandwritten Chinese text recognition (HCTR) has received extensive attention from the community of pattern recognition in the past decades. Most existing deep learning methods consist of two stages, i.e., training a text recognition network on the base of visual information, followed by incorporating language constrains with various language models. Therefore, the inherent linguistic semantic information is often neglected when designing the recognition network. To tackle this problem, in this work, we propose a novel multi-level multimodal fusion network and properly embed it into an attention-based LSTM so that both the visual information and the linguistic semantic information can be fully leveraged when predicting sequential outputs from the feature vectors. Experimental results on the ICDAR-2013 competition dataset demonstrate a comparable result with the state-of-the-art approaches. Yuhuan Xiu, Hongjian Zhan, Man Lan, Yue Lu 0001 |
ICDAR | 3 |
| 2019 | Writing Style Adversarial Network for Handwritten Chinese Character Recognition
Shujing Lyu, Hongjian Zhan, Yue Lu 0001 |
ICONIP (4) | 3 |
| 2019 | Residual CRNN and Its Application to Handwritten Digit String Recognition
Hongjian Zhan, Shujing Lyu, Xiao Tu, Yue Lu 0001 |
ICONIP (5) | 1 |
| 2018 | Improving Off-Line Handwritten Chinese Character Recognition with Semantic Information
Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICONIP (5) | 1 |
| 2018 | Handwritten Digit String Recognition using Convolutional Neural NetworkabstractString recognition is one of the most important tasks in computer vision applications. Recently the combinations of convolutional neural network (CNN) and recurrent neural network (RNN) have been widely applied to deal with the issue of string recognition. However RNNs are not only hard to train but also time-consuming. In this paper, we propose a new architecture which is based on CNN only, and apply it to handwritten digit string recognition (HDSR). This network is composed of three parts from bottom to top: feature extraction layers, feature dimension transposition layers and an output layer. Motivated by its super performance of DenseNet, we utilize dense blocks to conduct feature extraction. At the top of the network, a CTC (connectionist temporal classification) output layer is used to calculate the loss and decode the feature sequence, while some feature dimension transposition layers are applied to connect feature extraction and output layer. The experiments have demonstrated that, compared to other methods, the proposed method obtains significant improvements on ORAND-CAR-A and ORAND-CAR-B datasets with recognition rates 92.2% and 94.02%, respectively. Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICPR | 1 |
| 2017 | Handwritten Digit String Recognition by Combination of Residual Network and RNN-CTC
Hongjian Zhan, Yue Lu 0001 |
ICONIP (6) | 1 |