VLDB 2026 Research / reviewers in the wild / expert
Yue Lu 0001
dblp:74/6493-1
· DBLP profile ↗
135ranked-venue papers
17as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 93 · 11 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 2 first-author · 34 since 2021Databases, data management, data science and information retrieval · 22 · 9 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LithoMamba: High-fidelity lithography simulation with State Space ModelsabstractLithography simulation is a critical technology in modern semiconductor manufacturing, yet existing deep learning models often fail to accurately model the complex, long-range optical physics due to the inherent locality of convolution. This limitation results in insufficient simulation fidelity and poses significant challenges for optimization tasks. To overcome this challenge, we introduce LithoMamba, the first generative framework to leverage Mamba for high-fidelity lithography simulation. Our architecture uses a Mamba Generator to model global and long-range optical interactions, while a local, MLP-free Discriminator provides precise, spatial feedback to ensure fine-grained pattern fidelity. This global-local design enables our model to achieve both physical realism and exceptional detail. Our experiments show that LithoMamba outperforms existing methods, both in quantitative and qualitative results. These findings demonstrate the promise of State Space Models for improving lithography simulation and suggest new possibilities for combining physics with generative AI in chip manufacturing. Daohui Wang, Shujing Lyu, Pourya Shamsolmoali, Jiwei Shen, Yue Lu 0001 |
DATE | 6 |
| 2026 | CalliNet: a triplet network for chinese calligraphy style classification
Weilun Zhang, Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
Int. J. Document Anal. Recognit. | 4 |
| 2026 | Context-aware contrastive learning via structural harmony preservation for generalized category discovery
Wenbo Hu 0008, Yue Lu 0001, Xinchen Ma, Ching Y. Suen |
Knowl. Based Syst. | 2 |
| 2026 | Feature Enhancement Module Based on Class-Centric Loss for Fine-Grained Visual ClassificationabstractWe propose a novel feature enhancement module designed for fine-grained visual classification tasks, which can be seamlessly integrated into various backbone architectures, including both convolutional neural network (CNN)-based and Transformer-based networks. The plug-and-play module outputs pixel-level feature maps and performs a weighted fusion of filtered features to enhance fine-grained feature representation. We introduce a class-centric loss function that optimizes the alignment of samples with their target class centers by pulling them toward the center of the target class while simultaneously pushing them away from the center of the most visually similar nontarget classes. Soft labels are employed to mitigate overfitting, ensuring the model generalizes well to unseen examples. Our approach consistently delivers significant improvements in accuracy across various mainstream backbone architectures, underscoring its versatility and robustness. Furthermore, we achieved the highest accuracy on the NABirds (NAB) and our proprietary lock cylinder datasets. We have released our source code and pretrained model on GitHub: https://github.com/Richard5413/FEM-CC.git. Daohui Wang, Shujing Lyu, Tian Wei, Yue Lu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | BiMAC: Bidirectional Multimodal Alignment in Contrastive LearningabstractAchieving robust performance in vision-language tasks requires strong multimodal alignment, where textual and visual data interact seamlessly. Existing frameworks often combine contrastive learning with image captioning to unify visual and textual representations. However, reliance on global representations and unidirectional information flow from images to text limits their ability to reconstruct visual content accurately from textual descriptions. To address this limitation, we propose BiMAC, a novel framework that enables bidirectional interactions between images and text at both global and local levels. BiMAC employs advanced components to simultaneously reconstruct visual content from textual cues and generate textual descriptions guided by visual features. By integrating a text-region alignment mechanism, BiMAC identifies and selects relevant image patches for precise cross-modal interaction, reducing information noise and enhancing mapping accuracy. BiMAC achieves state-of-the-art performance across diverse vision-language tasks, including image-text retrieval, captioning, and classification. Masoumeh Zareapoor, Pourya Shamsolmoali, Yue Lu 0001 |
AAAI | 3 |
| 2025 | TriFormer: Triple Branch Heterogeneous Transformers for UHD Image RestorationabstractUltra-high-definition (UHD) image restoration is becoming a critical research area due to the increasing demand for high-quality visual content in various applications, including autonomous driving, remote sensing, digital entertainment, etc. However, UHD image restoration tasks present notable challenges, including the significant computational burden posed by the large image size scale, the difficulty in preserving fine details and textures, and the increased susceptibility to noise and artifacts during the restoration process. In this paper, we propose a novel triple-branched optimized heterogeneous transformer for UHD image restoration, named TriFormer. Specifically, our method embeds a high-resolution CNN branch for high-frequency features, a low-resolution pixel-wise attention branch for low-frequency features, and a channel-wise attention branch for feature fusion. We perform experiments across two UHD image restoration tasks: enhancing low-light images and deblurring. Results demonstrate that our model outperforms state-of-the-art approaches both quantitatively and qualitatively. The code will be made available at https://github.com/Chloe-mxxxxc/TRIFORMER. Xinchen Ma, Yue Lu 0001 |
ICASSP | 2 |
| 2025 | MSA2: Multi-Task Framework With Structure-Aware and Style-Adaptive Character Representation for Open-Set Chinese Text Recognition
Yangfu Li, Hongjian Zhan, Yujie Xiong, Yue Lu 0001 |
ICCV | 6 |
| 2025 | A New Fourier-Attention Guided Approach for Domain-Agnostic Text Localization
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
ICDAR (3) | 5 |
| 2025 | A Local Perceptual Approach for Few-Shot Text Effect Transfer
Hongjian Zhan, Tian Wei, Yue Lu 0001 |
ICIG (3) | 3 |
| 2025 | Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced CoveringabstractExisting visual token pruning methods target prompt alignment and visual preservation with static strategies, overlooking the varying relative importance of these objectives across tasks, which leads to inconsistent performance. To address this, we derive the first closed-form error bound for visual token pruning based on the Hausdorff distance, uniformly characterizing the contributions of both objectives. Moreover, leveraging $\epsilon$-covering theory, we reveal an intrinsic trade-off between these objectives and quantify their optimal attainment levels under a fixed budget. To practically handle this trade-off, we propose Multi-Objective Balanced Covering (MoB), which reformulates visual token pruning as a bi-objective covering problem. In this framework, the attainment trade-off reduces to budget allocation via greedy radius trading. MoB offers a provable performance bound and linear scalability with respect to the number of input visual tokens, enabling adaptation to challenging pruning scenarios. Extensive experiments show that MoB preserves 96.4\% of performance for LLaVA-1.5-7B using only 11.1\% of the original visual tokens and accelerates LLaVA-Next-7B by 1.3-1.5$\times$ with negligible performance loss. Additionally, evaluations on Qwen2-VL and Video-LLaVA confirm that MoB integrates seamlessly into advanced MLLMs and diverse vision-language tasks. The code will be made available soon. Yangfu Li, Hongjian Zhan, Yujie Xiong, Yue Lu 0001 |
NeurIPS | 6 |
| 2025 | ShapeMorph: 3D Shape Completion via Blockwise Discrete Diffusion
Pourya Shamsolmoali, Yue Lu 0001, Masoumeh Zareapoor |
WACV | 3 |
| 2025 | Length-aware center loss for sequence to sequence Thai scene text recognition
Hongjian Zhan, Yue Lu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | VGTS: Visually Guided Text Spotting for novel categories in historical manuscriptsabstractIn the field of historical manuscript research, scholars frequently encounter novel symbols in ancient texts, investing considerable effort in their identification and documentation. Although existing object detection methods achieve impressive performance on known categories, they struggle to recognize novel symbols without retraining. To address this limitation, we propose a Visually Guided Text Spotting (VGTS) approach that accurately spots novel characters using just one annotated support sample. The core of VGTS is a spatial alignment module consisting of a Dual Spatial Attention (DSA) block and a Geometric Matching (GM) block. The DSA block aims to identify, focus on, and learn discriminative spatial regions in the support and query images, mimicking the human visual spotting process. It first refines the support image by analyzing inter-channel relationships to identify critical areas, and then refines the query image by focusing on informative key points. The GM block, on the other hand, establishes the spatial correspondence between the two images, enabling accurate localization of the target character in the query image. To tackle the example imbalance problem in low-resource spotting tasks, we develop a novel torus loss function that enhances the discriminative power of the embedding space for distance metric learning . To further validate our approach, we introduce a new dataset featuring ancient Dongba hieroglyphics (DBH) associated with the Naxi minority of China. Extensive experiments on the DBH dataset and other public datasets, including Egyptian Hieroglyph (EGY), Historical Arabic Documents (HAD), Tripitaka Koreana in Han (TKH), and Notary Charters (NC), show that VGTS consistently surpasses state-of-the-art methods. The proposed framework exhibits great potential for application in historical manuscript text spotting, enabling scholars to efficiently identify and document novel symbols with minimal annotation effort. Wenbo Hu 0008, Hongjian Zhan, Xinchen Ma, Cong Liu 0006, Yue Lu 0001, Ching Y. Suen |
Expert Syst. Appl. | 6 |
| 2025 | Diff-TST: Diffusion model for one-shot text-image style transfer
Sizhe Pang, Yangchen Xie, Hongjian Zhan, Yue Lu 0001 |
Expert Syst. Appl. | 6 |
| 2025 | MAED: Mask assignment encoder decoder solver for multiple patterning layout decomposition
Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | GCENet: A geometric correspondence estimation network for tracking and loop detection in visual-inertial SLAM
Jichao Zhou, Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Struck-out handwritten word detection and restoration for automatic descriptive answer evaluationabstract• Exploring the combination of ResNet50 and the diagonal lines for segmentation. • Proposing the combination of U-Net and Bi-LSTM for restoring text pixels. • Results show that the model is impressive for detection and restoration. Unlike objective type evaluation, descriptive answer evaluation is challenging due to unpredictable answers and free writing style of answers. Because of these, descriptive answer evaluation has received special attention from many researchers. Automatic answer evaluation is useful for the following situations. It can avoid human intervention for marking, eliminates bias marking and most important is that it can save huge manpower. To develop an efficient and accurate system, there are several open challenges. One such open challenge is cleaning the document, which includes struck-out words removal and restoring the struck-out words. In this paper, we have proposed a system for struck-out handwritten word detection and restoration for automatic descriptive answer evaluation. The work has two stages. In the first stage, we explore the combination of ResNet50 and the diagonal line (principal and secondary diagonal lines) segmentation module for detecting words and then classifying struck-out words using a classification network. In the second stage, we explore the combination of U-Net as a backbone and Bi-LSTM for predicting pixels that represent actual text information of the struck-out words based on the relationship between sequences of pixels for restoration. Experimental results on our dataset and standard datasets show that the proposed model is impressive for struck-out word detection and restoration. A comparative study with the state-of-the-art methods shows that the proposed approach outperforms the existing models in terms of struck-out word detection and restoration. Dajian Zhong, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Signal Process. Image Commun. | 4 |
| 2025 | Dongba Machine Translation with Transfer Learning: Leveraging Pre-trained Ancient Chinese ModelsabstractThe Dongba script, a logographic writing system used by the Naxi people in religious activities, faces challenges in translation due to the advanced age of Dongba script experts and the time-consuming nature of manual deciphering. This study focuses on translating the resource-scarce Dongba script into Modern Chinese using a novel approach based on cross-lingual transfer learning from Ancient Chinese. By examining translation patterns from Ancient Chinese to Modern Chinese, we determine the feasibility of transferring knowledge from Ancient Chinese to Dongba script translation. We propose the Dongba Machine Translation Model (DMTM), a pre-trained, low-resource machine translation model that utilizes the linguistic similarities between Ancient Chinese and Dongba script to improve translation quality. The model undergoes pre-training on a large-scale Ancient Chinese corpus and fine-tuning on a small-scale Dongba script corpus, enabling effective knowledge transfer. To address the scarcity of Dongba script translation resources, we present DongBa Corpus 1.0, a fine-grained parallel dataset of Dongba script and Modern Chinese. Experimental results demonstrate that our proposed DMTM achieves a translation score of 50.01% BLEU on the test set. As no prior methods exist for Dongba script translation, we compared various architectures commonly used in low-resource translation tasks, and DMTM exhibited the best performance with a 5.39% improvement over alternative architectures tested. The implementation codes and dataset for our approach are available at https://github.com/Chloe-mxxxxc/DMTM . Xinchen Ma, Man Lan, Wenbo Hu 0008, Yue Lu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2025 | Leveraging Predictions of Task-Related Latents for Interactive Visual NavigationabstractInteractive visual navigation (IVN) involves tasks where embodied agents learn to interact with the objects in the environment to reach the goals. Current approaches exploit visual features to train a reinforcement learning (RL) navigation control policy network. However, RL-based methods continue to struggle at the IVN tasks as they are inefficient in learning a good representation of the unknown environment in partially observable settings. In this work, we introduce predictions of task-related latents (PTRLs), a flexible self-supervised RL framework for IVN tasks. PTRL learns the latent structured information about environment dynamics and leverages multistep representations of the sequential observations. Specifically, PTRL trains its representation by explicitly predicting the next pose of the agent conditioned on the actions. Moreover, an attention and memory module is employed to associate the learned representation to each action and exploit spatiotemporal dependencies. Furthermore, a state value boost module is introduced to adapt the model to previously unseen environments by leveraging input perturbations and regularizing the value function. Sample efficiency in the training of RL networks is enhanced by modular training and hierarchical decomposition. Extensive evaluations have proved the superiority of the proposed method in increasing the accuracy and generalization capacity. Jiwei Shen, Yue Lu 0001, Shujing Lyu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Enhancing scene text script identification through multi-task self-supervised learning
Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
Vis. Comput. | 3 |
| 2024 | Spotting the Unseen: Reciprocal Consensus Network Guided by Visual ArchetypesabstractHumans often require only a few visual archetypes to spot novel objects. Based on this observation, we present a strategy rooted in ``spotting the unseen" by establishing dense correspondences between potential query image regions and a visual archetype, and we propose the Consensus Network (CoNet). Our method leverages relational patterns intra and inter images via Auto-Correlation Representation (ACR) and Mutual-Correlation Representation (MCR). Within each image, the ACR module is capable of encoding both local self-similarity and global context simultaneously. Between the query and support images, the MCR module computes the cross-correlation across two image representations and introduces a reciprocal consistency constraint, which can incorporate to exclude outliers and enhance model robustness. To overcome the challenges of low-resource training data, particularly in one-shot learning scenarios, we incorporate an adaptive margin strategy to better handle diverse instances. The experimental results indicate the effectiveness of the proposed method across diverse domains such as object detection in natural scenes, and text spotting in both historical manuscripts and natural scenes, which demonstrates its sparkling generalization ability. Our code is available at: https://github.com/infinite-hwb/conet. Wenbo Hu 0008, Hongjian Zhan, Xinchen Ma, Yue Lu 0001, Ching Y. Suen |
AAAI | 4 |
| 2024 | Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning NetworkabstractScene text recognition is inherently a vision-language task. However, previous works have predominantly focused either on extracting more robust visual features or designing better language modeling. How to effectively and jointly model vision and language to mitigate heavy reliance on a single modality remains a problem. In this paper, aiming to enhance vision-language reasoning in scene text recognition, we present a balanced, unified and synchronized vision-language reasoning network (BUSNet). Firstly, revisiting the image as a language by balanced concatenation along length dimension alleviates the issue of over-reliance on vision or language. Secondly, BUSNet learns an ensemble of unified external and internal vision-language model with shared weight by masked modality modeling (MMM). Thirdly, a novel vision-language reasoning module (VLRM) with synchronized vision-language decoding capacity is proposed. Additionally, BUSNet achieves improved performance through iterative reasoning, which utilizes the vision-language prediction as a new language input. Extensive experiments indicate that BUSNet achieves state-of-the-art performance on several mainstream benchmark datasets and more challenge datasets for both synthetic and real training data compared to recent outstanding methods. Code and dataset will be available at https://github.com/jjwei66/BUSNet. Jiajun Wei, Hongjian Zhan, Yue Lu 0001, Xiao Tu, Cong Liu 0006, Umapada Pal 0001 |
AAAI | 3 |
| 2024 | Brush Your Text: Synthesize Any Scene Text on Images via Diffusion ModelabstractRecently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images. To tackle this problem, we propose Diff-Text, which is a training-free scene text generation framework for any language. Our model outputs a photo-realistic image given a text of any language along with a textual description of a scene. The model leverages rendered sketch images as priors, thus arousing the potential multilingual-generation ability of the pre-trained Stable Diffusion. Based on the observation from the influence of the cross-attention map on object placement in generated images, we propose a localized attention constraint into the cross-attention layer to address the unreasonable positioning problem of scene text. Additionally, we introduce contrastive image-level prompts to further refine the position of the textual region and achieve more accurate scene text generation. Experiments demonstrate that our method outperforms the existing method in both the accuracy of text recognition and the naturalness of foreground-background blending. Lingjun Zhang, Yaohui Wang 0001, Yue Lu 0001, Yu Qiao 0001 |
AAAI | 4 |
| 2024 | FaRE: A Feature-Aware Radical Encoding Strategy for Zero-Shot Chinese Character Recognition
Hongjian Zhan, Yangfu Li, Yujie Xiong, Yue Lu 0001 |
ACCV (1) | 4 |
| 2024 | Autoregressive 3D Shape Completion via Sphere-Guided Disentangled RepresentationabstractThis paper introduces a novel 3D shape completion method based on sphere-guided disentangled representation. Utilizing an autoregressive transformer-based model, our approach efficiently constructs object completion distributions given incomplete point clouds. To enhance completion modeling, we propose sDVQ-DIF (sphere-guided disentangled vector quantized deep implicit function), a new approach using decoupled discrete variables to represent 3D shapes efficiently. Experimental results demonstrate our model’s superior performance in terms of completion quality and fidelity compared to state-of-the-art methods, applicable to various shape types and incomplete patterns. Pourya Shamsolmoali, Yue Lu 0001 |
ICASSP | 3 |
| 2024 | Enhanced Deep Reinforcement Learning for Parcel Singulation in Non-Stationary EnvironmentsabstractIn the rapidly expanding logistics sector, parcel singulation has emerged as a significant bottleneck. To address this, we propose an automated parcel singulator utilizing a sparse actuator array, which presents an optimal balance between cost and efficiency, albeit requiring a sophisticated control policy. In this study, we frame the parcel singulation issue as a Markov Decision Process with a variable state space dimension, addressed through a deep reinforcement learning (RL) algorithm complemented by a State Space Standardization Module (S3). Distinct from previous RL approaches, our methodology initially considers the non-stationary environment during the problem modeling phase. To counter this challenge, the S3 module standardizes the dynamic input state, thereby stabilizing the RL training process. We validate our method through simulation experiments in complex environments, comparing it with several baseline algorithms. Results indicate that our algorithm excels in parcel singulation tasks, achieving a higher success rate and enhanced efficiency. Jiwei Shen, Hu Lu, Shujing Lyu, Yue Lu 0001 |
ICASSP | 5 |
| 2024 | Enhancing Reinforcement Learning via Causally Correct Input Identification and Targeted InterventionabstractCausal confusion, characterized by the learning of spurious correlations, detrimentally affects the generalization and effectiveness of reinforcement learning (RL) algorithms, especially in environments without latent confounders often encountered in robot autonomous navigation tasks. This study addresses this gap by developing a causal structure within a Partially Observable Markov Decision Process (POMDP). Subsequently, we introduce a targeted intervention that mitigates the influence of spurious correlations by isolating causally significant state variables and discarding irrelevant inputs. Testing in three real-world scenarios confirms the approach’s feasibility and superiority in enhancing the RL algorithms’ performance and generalization ability, signifying a promising step towards more robust online RL frameworks. Jiwei Shen, Hu Lu, Shujing Lyu, Yue Lu 0001 |
ICASSP | 5 |
| 2024 | Fragile Model Watermark for integrity protection: leveraging boundary volatility and sensitive sample-pairingabstractNeural networks have increasingly influenced people’s lives. Ensuring the faithful deployment of neural networks as designed by their model owners is crucial, as they may be susceptible to various malicious or unintentional modifications, such as backdooring and poisoning attacks. Fragile model watermarks aim to prevent unexpected tampering that could lead DNN models to make incorrect decisions. They ensure the detection of any tampering with the model as sensitively as possible. However, prior watermarking methods suffered from inefficient sample generation and insufficient sensitivity, limiting their practical applicability. Our approach employs a sample-pairing technique, placing the model boundaries between pairs of samples, while simultaneously maximizing logits. This ensures that the model’s decision results of sensitive samples change as much as possible and the Top-1 labels easily alter regardless of the direction it moves. Experimental evaluations conducted across multiple models and datasets demonstrate the superior sensitivity and generation efficiency of our method compared to the current approaches. Zhenzhe Gao, Zhenjun Tang, Zhao-Xia Yin, Baoyuan Wu, Yue Lu 0001 |
ICME | 5 |
| 2024 | RSTAN: Residual Spatio-Temporal Attention Network for End-to-End Human Fall Detection
Yaru Jiang, Shujing Lyu, Hongjian Zhan, Yue Lu 0001 |
ICPR (15) | 4 |
| 2024 | Learning to Detect Lithography Defects in SEM Images
Hu Lu, Botong Zhao, Jiwei Shen, Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICPR (5) | 6 |
| 2024 | LK-Net: Efficient Large Kernel ConvNet for Document Enhancement
Qijun Shi, Hongjian Zhan, Yangfu Li, Weijun Zou, Huasheng Li, Umapada Pal 0001, Yue Lu 0001 |
ICPR (21) | 7 |
| 2024 | Fractional Correspondence Framework in Detection TransformerabstractThe Detection Transformer (DETR), by incorporating the Hungarian algorithm, has significantly simplified the matching process in object detection tasks. This algorithm facilitates optimal one-to-one matching of predicted bounding boxes to ground-truth annotations during training. While effective, this strict matching process does not inherently account for the varying densities and distributions of objects, leading to suboptimal correspondences such as failing to handle multiple detections of the same object or missing small objects. To address this, we propose the Regularized Transport Plan (RTP). RTP introduces a flexible matching strategy that captures the cost of aligning predictions with ground truths to find the most accurate correspondences between these sets. By utilizing the differentiable Sinkhorn algorithm, RTP allows for soft, fractional matching rather than strict one-to-one assignments. This approach enhances the model's capability to manage varying object densities and distributions effectively. Our extensive evaluations on the MS-COCO and VOC benchmarks demonstrate the effectiveness of our approach. RTP-DETR, surpassing the performance of the Deform-DETR and the recently introduced DINO-DETR, achieving absolute gains in mAP of +3.8% and +1.7%, respectively. Masoumeh Zareapoor, Pourya Shamsolmoali, Huiyu Zhou 0001, Yue Lu 0001, Salvador García 0001 |
ACM Multimedia | 4 |
| 2024 | Free Lunch: Frame-level Contrastive Learning with Text Perceiver for Robust Scene Text Recognition in Lightweight Models
Hongjian Zhan, Yangfu Li, Yujie Xiong, Umapada Pal 0001, Yue Lu 0001 |
ACM Multimedia | 5 |
| 2024 | TANet: Text region attention learning for vehicle re-identification
Wenbo Hu 0008, Hongjian Zhan, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Adaptive graph-based feature normalization for facial expression recognition
Yujie Xiong, Yangtao Du, Yue Lu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Enhanced video clustering using multiple riemannian manifold-valued descriptors and audio-visual information
Wenbo Hu 0008, Hongjian Zhan, Yinghong Tian, Yujie Xiong, Yue Lu 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Enhancing parcel singulation efficiency through transformer-based position attention and state space augmentation
Jiwei Shen, Hu Lu, Shujing Lyu, Yue Lu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Weakly supervised scene text generation for low-resource languages
Yangchen Xie, Hongjian Zhan, Palaiahnakote Shivakumara, Cong Liu 0006, Yue Lu 0001 |
Expert Syst. Appl. | 7 |
| 2024 | NDOrder: Exploring a novel decoding order for scene text recognition
Dajian Zhong, Hongjian Zhan, Shujing Lyu, Cong Liu 0006, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 8 |
| 2024 | Adaptive feature fusion for scene text script identification
Fuyou Peng, Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
Multim. Tools Appl. | 4 |
| 2024 | Distance-based Weighted Transformer Network for image completion
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Xuelong Li 0001, Yue Lu 0001 |
Pattern Recognit. | 5 |
| 2024 | Adaptive watermarking with self-mutual check parameters in deep neural networks
Zhenzhe Gao, Zhao-Xia Yin, Hongjian Zhan, Yue Lu 0001 |
Pattern Recognit. Lett. | 5 |
| 2024 | LithoPW: Leveraging Visual Memory Encoding and Defect-Aware Optimization for Precise Determination of the Lithography Process WindowsabstractLithography stands as a critical step in the manufacturing of integrated circuits, where the precise control of focus and exposure dose parameters is vital for optimal results. The conventional methodologies for defining lithography process windows often face difficulties with managing measurement errors, detecting printed defects, and exploiting visual features from Scanning Electron Microscope (SEM) images. This paper proposes LithoPW, a novel framework that utilizes visual features of SEM images for the determination of process windows. This approach is comprised of a denoising module, a Transformer-based visual memory encoder, and a defect-aware process window optimization module. The denoising module incorporates a Transformer architecture to mitigate the impact of noise, thereby enhancing the efficiency of downstream tasks in leveraging information embedded within SEM images. The transformer-based visual memory encoder discerns each SEM image as a Query, maintaining neighbouring SEM images in memory as Key and Value elements, thereby facilitating precise lithography quality classification associated with the query image. The defect-aware process window optimization module heightens the reliability of the results by adjusting the process window according to the defects identified within the SEM images. Experimental results confirm the efficacy of our framework, highlighting its promising application in lithography production for accurate process window determination. Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Feature fusion and decomposition: exploring a new way for Chinese calligraphy style classification
Li Liu 0010, Taorong Qiu, Yue Lu 0001, Ching Y. Suen |
Vis. Comput. | 5 |
| 2023 | Modeling Cross-layer Interaction for Chinese Calligraphy Style Classification
Li Liu 0010, Taorong Qiu, Yue Lu 0001, Ching Y. Suen |
ICDAR (4) | 4 |
| 2023 | Scene Text Recognition with Image-Text Matching-Guided Dictionary
Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu 0001, Umapada Pal 0001 |
ICDAR (6) | 4 |
| 2023 | Table Structure Recognition of Historical Dongba Documents
Hongjian Zhan, Xiao Tu, Yue Lu 0001 |
ICIG (1) | 4 |
| 2023 | CCH-YOLOX: Improved YOLOX for Challenging Vehicle Detection from UAV ImagesabstractIn this paper, we focus on vehicle detection algorithms from UAV images in intelligent transportation applications. We propose an improved YOLOX model called CCHYOLOX for the problems of dense distribution and drastic scale changes of vehicle targets from the UAV viewpoint. Firstly, we construct a novel plug-and-play module, called Correlation Extraction and Feature Fusion (CEFF), to process the multi-layer adaptive fusion of pyramidal features. It enables the adaptive fusion of features in adjacent layers by adding spatial-awareness and extracting global channels correlation of adjacent layers features to mitigate the effect of target size diversity. Then, we design a special cascade strategy, including a feature alignment module named DConv, for single-stage and anchor-free detectors by considering the feature offset problem to produce more accurate detection bounding boxes. The cascade strategy allows the model to improve performance with little increase in computational cost. Finally, a high-resolution branch is designed for the small-target detection task, which greatly improves the detection accuracy of the model. Experiments on the challenging Visdrone-vehicle and Drone Vehicle datasets show that the proposed method effectively tackle the detection accuracy decrease caused by the above problems. CCH-YOLOX achieves 43.2% mAP and 57.3% mAP on the above two datasets, respectively, which is about 3%-4% higher than the strong baseline model (YOLOX) and exceeds many current popular models. The code is available at https://github.com/lz06787/CCH-YOLOX Song Qiu, Mingsong Chen 0001, Dingding Han, Tiantian Qi, Qingli Li, Yue Lu 0001 |
IJCNN | 7 |
| 2023 | A New Lightweight Script Independent Scene Text Style Transfer NetworkabstractScene text style transfer without a language barrier is an open challenge for the video and scene text recognition community because this plays a vital role in poster, web design, augmenting character images, and editing characters to improve scene text recognition performance and usability. This work presents a new model, called Script Independent Scene Text Style Transfer Network (SISTSTNet), for extracting scene characters and transferring text style simultaneously. The SISTSTNet performs mapping in language-independent feature space for transferring style. It is designed based on a Style Parameter Network and Target Encoder Network through lightweight MobileNetv3 convolutional and residual blocks to capture the style and shape to generate target characters. Similarly, a generative model is explored through the Visual Geometry Group (VGG) network for character replacement. The SISTSTNet is flexible and works on different languages and arbitrary examples in a neat and unified fashion. The experimental results on images in various languages, namely, English, Chinese, Hindi, Russian, Japanese, Arabic, Greek, and Bengali and cross-language validation demonstrate the effectiveness of the proposed method. The performance of the method is superior compared to the state-of-the-art methods in terms of quality measures, language independence, shape-preserving, and efficiency. The code and dataset will be released to the public to support reproducibility. Palaiahnakote Shivakumara, Ayush Roy, Lokesh Nandanwar, Umapada Pal 0001, Yue Lu 0001, Cheng-Lin Liu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2023 | A multi-grained unsupervised domain adaptation approach for semantic segmentation
Tai Ma, Yue Lu 0001, Qingli Li, Lianghua He, Ying Wen 0003 |
Pattern Recognit. | 3 |
| 2023 | VQAPT: A New visual question answering model for personality traits in social media images
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001, Yue Lu 0001 |
Pattern Recognit. Lett. | 5 |
| 2023 | SANet-SI: A new Self-Attention-Network for Script Identification in scene images
Hongjian Zhan, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Pattern Recognit. Lett. | 5 |
| 2023 | A New Few-Shot Learning-Based Model for Prohibited Objects Detection in Cluttered Baggage X-Ray Images Through Edge Detection and Reverse ValidationabstractDetecting prohibited items via X-ray screening at airports and sensitive venues is essential for preventing smuggling and breaches of security. The difficulty in prohibited items inspection lies in accurately detecting prohibited items in complex X-ray images and limited access to X-ray images containing prohibited items. Few-shot detection aims at learning with limited examples and assigning a category label to each object. However, most few-shot learning methods do not focus on the edge information of the occluded object in X-ray images, which is crucial for the model to detect prohibited items in the X-ray images. In this paper, we presents a method (RVViT) for few-shot prohibited items detection tasks which fully acknowledges the significance of X-ray penetrability and increases the stability of few-shot learning model. Specifically, a Transformer encoder is firstly adopted for generating high-level semantic features that contain global information. At the same time, an edge detection module is devised for enhancing the edge information of prohibited items. Moreover, to further improve the stability of the few-shot learning model and ensure prototype consistency between the support and query samples, a reverse validation strategy is proposed to assist training. Extensive experiments demonstrate our method outperforms state-of-the-art approaches in terms of detection with a small number of samples. Shujing Lyu, Palaiahnakote Shivakumara, Michael Blumenstein, Yue Lu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Efficient Object Detection in Optical Remote Sensing Imagery via Attention-Based Feature DistillationabstractEfficient object detection methods have recently received great attention in remote sensing. Although deep convolutional networks often have excellent detection accuracy, their deployment on resource-limited edge devices is difficult. Knowledge distillation (KD) is a strategy for addressing this issue since it makes models lightweight while maintaining accuracy. However, existing KD methods for object detection have encountered two constraints. First, they discard potentially important background information and only distill nearby foreground regions. Second, they only rely on the global context, which limits the student detector’s ability to acquire local information from the teacher detector. To address the aforementioned challenges, we propose Attention-based Feature Distillation (AFD), a new KD approach that distills both local and global information from the teacher detector. To enhance local distillation, we introduce a multi-instance attention mechanism that effectively distinguishes between background and foreground elements. This approach prompts the student detector to focus on the pertinent channels and pixels, as identified by the teacher detector. Local distillation lacks global information, thus attention global distillation is proposed to reconstruct the relationship between various pixels and pass it from teacher to student detector. The performance of AFD is evaluated on two public aerial image benchmarks, and the evaluation results demonstrate that AFD in object detection can attain the performance of other state-of-the-art models while being efficient. Pourya Shamsolmoali, Jocelyn Chanussot, Huiyu Zhou 0001, Yue Lu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Few-Shot Segmentation for Prohibited Items Inspection With Patch-Based Self-Supervised Learning and Prototype Reverse ValidationabstractProhibited items inspection using X-ray screening is essential for reducing the risk of crime and terrorist attacks. The difficulty in prohibited items inspection lies in accurately detecting prohibited items in complex X-ray images and limited access to X-ray images containing prohibited items. Few-shot segmentation aims at learning with limited examples and assigning a category label to each image pixel. However, current few-shot methods are mostly full-supervised and less robust to the prohibited items categories that did not appear during training process. In this paper, we propose a method for few-shot prohibited items segmentation tasks which utilize unlabeled data and better leverage the representation of input samples during model training process. Specifically, a patch-based self-supervised embedding network is firstly devised as the base learner to learn an abstract representation of the observation from unlabeled samples. Then we apply few-shot learning and generate abstract representation related to prohibited items from support sample within the embedding space, which is followed by obtaining the corresponding class-specific prototype representations via masked average pooling. The distance between each pixel of query sample and prototypes are calculated to predict the label of each pixel. Moreover, prototype reverse validation strategy (PRV) is proposed to further exploit the support representation to assist training. Extensive experiments show that our proposed method outperforms the state-of-the-art by delivering a higher accuracy on automated prohibited items inspection and requiring less labeled samples. Shujing Lyu, Yue Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | SGBANet: Semantic GAN and Balanced Attention Network for Arbitrarily Oriented Scene Text Recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
ECCV (28) | 7 |
| 2022 | ScriptNet: A Two Stream CNN for Script Identification in Camera-Based Document Images
Minzhen Deng, Li Liu 0010, Taorong Qiu, Yue Lu 0001, Ching Y. Suen |
ICONIP (6) | 5 |
| 2022 | Thai Scene Text Recognition with Character Combination
Hongjian Zhan, Yue Lu 0001 |
PRCV (3) | 4 |
| 2022 | DSAM-GN: Graph Network Based on Dynamic Similarity Adjacency Matrices for Vehicle Re-identification
Yuejun Jiao, Song Qiu, Mingsong Chen 0001, Dingding Han, Qingli Li, Yue Lu 0001 |
PRICAI (1) | 6 |
| 2022 | Text proposals with location-awareness-attention network for arbitrarily shaped scene text detection and recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 5 |
| 2022 | Local Resultant Gradient Vector Difference and Inpainting for 3D Text Detection in the WildabstractThree-dimensional (3D) text appearing in natural scene images is common due to 3D cameras and the capture of text from different angles, which presents new problems for text detection. This is because of the presence of depth information, shadows, and decorative characters in the images. In this work, we consider those images where 3D text appears with depth, as well as shadow information for text detection. We propose a novel method based on local resultant gradient vector difference (LRGVD), inpainting and a deep learning model for detecting 3D as well as two-dimensional (2D) texts in natural scene images. The boundary of components that are invariant to the above challenges is detected by exploring LRGVD. The LRGVD uses gradient magnitude and direction in a novel way for detecting the boundary of the components. Further, we propose an inpainting method in a new way for restoring the character background information using boundaries. For a given region and the input image, the inpainting method divides the whole image into planes and then propagates the values in the planes into the missing region based on posterior probabilities and neighboring information. This results in text regions with false positives. Then, the differential binarization network (DB-Net) is proposed for detecting text irrespective of orientation, background, 3D or 2D, etc. Experiments conducted on our 3D text images and standard datasets of natural scene text images, namely ICDAR 2019 MLT, ICDAR 2019 ArT, DAST1500, Total-Text and SCUT-CTW1500, show that the proposed method is effective in detecting 3D and 2D texts in the images. Dajian Zhong, Palaiahnakote Shivakumara, Lokesh Nandanwar, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2022 | A New Deep Wavefront Based Model for Text Localization in 3D VideoabstractWith the evolution of electronic devices, such as 3D cameras, addressing the challenges of text localization in 3D video (e.g., for indexing) is increasingly drawing the attention of the multimedia and video processing community. Existing methods focus on 2D video and their performance in the presence of the challenges in 3D video, such as shadow areas associated with text and irregularly sized and shaped text, degrades. This paper proposes the first approach that successfully addresses the challenges of 3D video in addition to those of 2D. It employs a number of innovations, among which, the first is the Generalized Gradient Vector Flow (GGVF) for dominant points detection. The second is the Wavefront concept for text candidate point detection from those dominant points. In addition, an Adaptive B-Spline Polygon Curve Network (ABS-Net) is proposed for accurate text localization in 3D videos by constructing tight fitting bounding polygons using text candidate points. Extensive experiments on custom (3D video) and standard datasets (2D video and scene text) show that the proposed method is practical and useful, and overall outperforms existing state-of-the-art methods. Lokesh Nandanwar, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Apostolos Antonacopoulos, Yue Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | DG-Font: Deformable Generative Networks for Unsupervised Font GenerationabstractFont generation is a challenging problem especially for some writing systems that consist of a large number of characters and has attracted a lot of attention in recent years. However, existing methods for font generation are often in supervised learning. They require a large number of paired data, which is labor-intensive and expensive to collect. Besides, common image-to-image translation models often define style as the set of textures and colors, which cannot be directly applied to font generation. To address these problems, we propose novel deformable generative networks for unsupervised font generation (DG-Font). We introduce a feature deformation skip connection (FDSC) which predicts pairs of displacement maps and employs the predicted maps to apply deformable convolution to the low-level feature maps from the content encoder. The outputs of FDSC are fed into a mixer to generate the final results. Taking advantage of FDSC, the mixer outputs a high-quality character with a complete structure. To further improve the quality of generated images, we use three deformable convolution layers in the content encoder to learn style-invariant feature representations. Experiments demonstrate that our model generates characters in higher quality than state-of-art methods. The source code is available at https://github.com/ecnuycxie/DG-Font. Yangchen Xie, Yue Lu 0001 |
CVPR | 4 |
| 2021 | Scene Text Transfer for Cross-Language
Lingjun Zhang, Yangchen Xie, Yue Lu 0001 |
ICIG (1) | 4 |
| 2021 | Multi-loss Siamese Convolutional Neural Network for Chinese Calligraphy Style Classification
Li Liu 0010, Wenyan Cheng, Taorong Qiu, Chengying Tao, Qiu Chen, Yue Lu 0001, Ching Y. Suen |
ICONIP (6) | 6 |
| 2021 | A Coherent Cooperative Learning Framework Based on Transfer Learning for Unsupervised Cross-Domain Classification
Xinxin Shan, Ying Wen 0003, Qingli Li, Yue Lu 0001, Haibin Cai |
MICCAI (5) | 4 |
| 2021 | Towards Reasoning Ability in Scene Text Visual Question AnsweringabstractWorks on scene text visual question answering (TextVQA) always emphasize the importance of reasoning questions and image contents. However, we find current TextVQA models lack reasoning ability and tend to answer questions by exploiting dataset bias and language priors. Moreover, our observations indicate that recent accuracy improvement in TextVQA is mainly contributed by stronger OCR engines, better pre-training strategies and more Transformer layers, instead of newly proposed networks. In this work, towards the reasoning ability, we 1) conduct module-wise contribution analysis to quantitatively investigate how existing works improve accuracies in TextVQA; 2) design a gradient-based explainability method to explore why TextVQA models answer what they answer and find evidence for their predictions; 3) perform qualitative experiments to visually analyze models reasoning ability and explore potential reasons behind such a poor ability. Liqiang Xiao, Yue Lu 0001, Yaohui Jin, Hao He 0007 |
ACM Multimedia | 3 |
| 2021 | DenseNet-CTC: An end-to-end RNN-free architecture for context-free string recognition
Hongjian Zhan, Shujing Lyu, Yue Lu 0001, Umapada Pal 0001 |
Comput. Vis. Image Underst. | 3 |
| 2021 | See more than once: Kernel-sharing atrous convolution for semantic segmentation
Wenjing Jia, Yue Lu 0001, Xiangjian He |
Neurocomputing | 4 |
| 2021 | Document image classification: Progress over two decades
Li Liu 0010, Taorong Qiu, Qiu Chen, Yue Lu 0001, Ching Y. Suen |
Neurocomputing | 5 |
| 2021 | Guest Editorial
Yue Lu 0001, Nicole Vincent, Ching Y. Suen, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2021 | Model-Based Transfer Learning and Sparse Coding for Partial Face RecognitionabstractWith the growing needs of practical applications such as security monitoring, partial face recognition is a challenging but important issue, because the captured faces in real-world surveillance videos may be occluded or with variations. Though current face recognition methods perform well in relatively constrained scenes, they may suffer from degradation for partial faces. In this paper, we propose a framework of model-based transfer learning and sparse coding (MTLSC) for partial face recognition. First, due to less information in partial face image, we exploit the mirrored image of an original probe sample as sample augment to provide further information. Considering the inadequacy of training face samples, we obtain face features based on model-based transfer learning VGGNet that is pre-trained on VGGFace dataset. Then we reconstruct face features by sliding window in view of different sizes of partial face hard to extract the same feature dimension. Finally we carry out sparse coding with rectification and calculate the minimum score of the probe and mirrored samples among all classes to get the results. Thus, by model-based transfer learning, sliding window for feature reconstruction and sparse coding with rectification, the proposed framework improves partial face recognition performance. Experimental results on three face databases (LFW, AR and NIR), and two person re-identification databases (iLIDS-VID and PKU-Reid) demonstrate our method is effective for partial face recognition. Xinxin Shan, Yue Lu 0001, Qingli Li, Ying Wen 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | FACLSTM: ConvLSTM with focused attention for scene text recognition
Wenjing Jia, Xiangjian He, Michael Blumenstein, Shujing Lyu, Yue Lu 0001 |
Sci. China Inf. Sci. | 7 |
| 2020 | Combination of spatially enhanced bag-of-visual-words model and genuine difference subspace for fake coin detection
Li Liu 0010, Taorong Qiu, Yue Lu 0001, Qiu Chen, Ching Y. Suen |
Expert Syst. Appl. | 3 |
| 2020 | Gabor Feature-Based LogDemons With Inertial Constraint for Nonrigid Image RegistrationabstractNonrigid image registration plays an important role in the field of computer vision and medical application. The methods based on Demons algorithm for image registration usually use intensity difference as similarity criteria. However, intensity based methods can not preserve image texture details well and are limited by local minima. In order to solve these problems, we propose a Gabor feature based LogDemons registration method in this paper, called GFDemons. We extract Gabor features of the registered images to construct feature similarity metric since Gabor filters are suitable to extract image texture information. Furthermore, because of the weak gradients in some image regions, the update fields are too small to transform the moving image to the fixed image correctly. In order to compensate this deficiency, we propose an inertial constraint strategy based on GFDemons, named IGFDemons, using the previous update fields to provide guided information for the current update field. The inertial constraint strategy can further improve the performance of the proposed method in terms of accuracy and convergence. We conduct experiments on three different types of images and the results demonstrate that the proposed methods achieve better performance than some popular methods. Ying Wen 0003, Yue Lu 0001, Qingli Li, Haibin Cai, Lianghua He |
IEEE Trans. Image Process. | 3 |
| 2019 | DeepText: Detecting Text from the Wild with Multi-ASPP-Assembled DeepLababstractIn this paper, we address the issue of scene text detection in the way of direct regression and successfully adapt an effective semantic segmentation model, DeepLab v3+ [1], for this application. In order to handle texts with arbitrary orientations and sizes and improve the recall of small texts, we propose to extract features of multiple scales by inserting multiple Atrous Spatial Pyramid Pooling (ASPP) layers to the DeepLab after the feature maps with different resolutions. Then, we set multiple auxiliary IoU losses at the decoding stage and make auxiliary connections from the intermediate encoding layers to the decoder to assist network training and enhance the discrimination ability of lower encoding layers. Experiments conducted on the benchmark scene text dataset ICDAR2015 demonstrate the superior performance of our proposed network, named as DeepText, over the state-of-the-art approaches. Wenjing Jia, Xiangjian He, Yue Lu 0001, Michael Blumenstein, Shujing Lyu |
ICDAR | 4 |
| 2019 | A Handwritten Chinese Text Recognizer Applying Multi-level Multimodal Fusion NetworkabstractHandwritten Chinese text recognition (HCTR) has received extensive attention from the community of pattern recognition in the past decades. Most existing deep learning methods consist of two stages, i.e., training a text recognition network on the base of visual information, followed by incorporating language constrains with various language models. Therefore, the inherent linguistic semantic information is often neglected when designing the recognition network. To tackle this problem, in this work, we propose a novel multi-level multimodal fusion network and properly embed it into an attention-based LSTM so that both the visual information and the linguistic semantic information can be fully leveraged when predicting sequential outputs from the feature vectors. Experimental results on the ICDAR-2013 competition dataset demonstrate a comparable result with the state-of-the-art approaches. Yuhuan Xiu, Hongjian Zhan, Man Lan, Yue Lu 0001 |
ICDAR | 5 |
| 2019 | Change Detection via Graph Matching and Multi-View Geometric ConstraintsabstractChange detection is a critical preprocessing step of visual perception with broad prospects. Its primary challenge is to identify all the meaningful changes from a target image to the source image, which is observed of the same scene and has a different perspective as well. A robust change detection method involving graph matching and geometric constraints is proposed in this paper. Maximum common sub-graph matching is applied for alleviating the risk of suboptimal results and geometric constraints are used to remove the possible mistaken results. Detection results in different real-world scenes with respect to considerable textural moved objects show that the proposed method is more robust than the state-of-the-art methods. Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
ICIP | 4 |
| 2019 | Writing Style Adversarial Network for Handwritten Chinese Character Recognition
Shujing Lyu, Hongjian Zhan, Yue Lu 0001 |
ICONIP (4) | 4 |
| 2019 | Modified Adaptive Implicit Shape Model for Object Detection
Ziyan Xu, Shujing Lyu, Weiping Jin, Yue Lu 0001 |
ICONIP (5) | 4 |
| 2019 | Residual CRNN and Its Application to Handwritten Digit String Recognition
Hongjian Zhan, Shujing Lyu, Xiao Tu, Yue Lu 0001 |
ICONIP (5) | 4 |
| 2019 | Improving Text-Independent Chinese Writer Identification with the Aid of Character PairsabstractText-independent Chinese writer identification does not depend on the text content of the query and reference handwritings. In order to deal with the uncertainty of the text content, text-independent approaches usually give special attention to the global writing style of handwriting, rather than the properties of each individual character or word. Thanks to the existence of high-frequency characters, some characters probably appear in both the query and reference handwritings in most cases. If character images in the query handwriting are similar to those in the reference handwriting, this query handwriting and the corresponding reference handwriting are very likely to be written by the identical writer. In this paper, we exploit the above characteristic to improve the performance of Chinese writer identification. We first present an identification scheme using edge co-occurrence feature (ECF). Then, we detect the character pairs in the query and reference handwritings using a two-step framework and propose the displacement field-based similarity (DFS) to determine whether a character pair is written by the identical writer. The character pairs help to re-rank the candidate list obtained by text-independent ECF-based similarity and finally decide the writer of the query handwriting. The proposed method is evaluated on the HIT-MW and CASIA-2.1 datasets. Experimental results demonstrate that our proposed method outperforms the existing ones, and its Top-1 accuracy on the two datasets reaches 97.1% and 98.3%, respectively. Yujie Xiong, Li Liu 0010, Shujing Lyu, Patrick Shen-Pei Wang, Yue Lu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2018 | Improving Off-Line Handwritten Chinese Character Recognition with Semantic Information
Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICONIP (5) | 3 |
| 2018 | A Fusion Strategy for the Single Shot Text DetectorabstractIn this paper, we propose a new fusion strategy for scene text detection. The system is based on a single fully convolution network, which outputs the coordinates of text bounding boxes at multiple scales. We improve the performance of text detection by combining a fusion strategy. This strategy obtains precise text bounding boxes according to the confidence of candidate text boxes. It exhibits promising robustness and discriminative power by fusing text boxes. Experimental results on ICDAR2011 and ICDAR2013 datasets indicate the effectiveness and robustness of the proposed fusion strategy with an F-measure of 87%, which outperforms the base network 2%. Shujing Lyu, Yue Lu 0001, Patrick Shen-Pei Wang |
ICPR | 3 |
| 2018 | Handwritten Digit String Recognition using Convolutional Neural NetworkabstractString recognition is one of the most important tasks in computer vision applications. Recently the combinations of convolutional neural network (CNN) and recurrent neural network (RNN) have been widely applied to deal with the issue of string recognition. However RNNs are not only hard to train but also time-consuming. In this paper, we propose a new architecture which is based on CNN only, and apply it to handwritten digit string recognition (HDSR). This network is composed of three parts from bottom to top: feature extraction layers, feature dimension transposition layers and an output layer. Motivated by its super performance of DenseNet, we utilize dense blocks to conduct feature extraction. At the top of the network, a CTC (connectionist temporal classification) output layer is used to calculate the loss and decode the feature sequence, while some feature dimension transposition layers are applied to connect feature extraction and output layer. The experiments have demonstrated that, compared to other methods, the proposed method obtains significant improvements on ORAND-CAR-A and ORAND-CAR-B datasets with recognition rates 92.2% and 94.02%, respectively. Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICPR | 3 |
| 2018 | A platform of digital brain using crowd powerabstractA powerful platform of digital brain is proposed using crowd wisdom for brain research, based on the computational artificial intelligence model of synthesis reasoning and multi-source analogical generating. The design of the platform aims to make it a comprehensive brain database, a brain phantom generator, a brain knowledge base, and an intelligent assistant for research on neurological and psychiatric diseases and brain development. Using big data, crowd wisdom, and high performance computers may significantly enhance the capability of the platform. Preliminary achievements along this track are reported. Dongrong Xu, Yue Lu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | Compact MQDF classifiers using sparse coding for handwritten Chinese character recognition
Xiaohua Wei, Shujing Lu, Yue Lu 0001 |
Pattern Recognit. | 3 |
| 2017 | Similar Handwritten Chinese Character Recognition Using Hierarchical CNN ModelabstractWe propose a hierarchical CNN model for the recognition of confusable similar handwritten Chinese characters, which are automatically extracted from a large character set by utilizing a classifier's recognition result. The proposed hierarchical CNN model takes advantage of deep networks and traditional hierarchical methods, and consists of two stages, which are expected to differentiate inter-group characters and intra-group characters, respectively. Different from traditional ways of expanding depth and/or width of general sole classifier CNNs, we explore the way of designing multiple parallel CNN classifiers to capture critical regions of similar characters. Each classifier along with their feature extraction layers is trained only with a group of similar characters so that the subtle shape difference can be captured. Totally, 368 similar characters (categorized into 172 groups) are extracted from 3755 frequently used Chinese characters. Experimental results on these similar characters demonstrate the superiority of the proposed method to the expanded CNN models. Yue Lu 0001 |
ICDAR | 2 |
| 2017 | Handwritten Digit String Recognition by Combination of Residual Network and RNN-CTC
Hongjian Zhan, Yue Lu 0001 |
ICONIP (6) | 3 |
| 2017 | A Sequence Labeling Convolutional Network and Its Application to Handwritten String RecognitionabstractHandwritten string recognition has been struggling with connected patterns fiercely. Segmentation-free and over-segmentation frameworks are commonly applied to deal with this issue. For the past years, RNN combining with CTC has occupied the domain of segmentation-free handwritten string recognition, while CNN is just employed as a single character recognizer in the over-segmentation framework. The main challenges for CNN to directly recognize handwritten strings are the appropriate processing of arbitrary input string length, which implies arbitrary input image size, and reasonable design of the output layer. In this paper, we propose a sequence labeling convolutional network for the recognition of handwritten strings, in particular, the connected patterns. We properly design the structure of the network to predict how many characters present in the input images and what exactly they are at every position. Spatial pyramid pooling (SPP) is utilized with a new implementation to handle arbitrary string length. Moreover, we propose a more flexible pooling strategy called FSPP to adapt the network to the straightforward recognition of long strings better. Experiments conducted on handwritten digital strings from two benchmark datasets and our own cell-phone number dataset demonstrate the superiority of the proposed network. Yue Lu 0001 |
IJCAI | 2 |
| 2017 | Off-line Text-Independent Writer Recognition: A SurveyabstractWriter recognition is to identify a person on the basis of handwriting, and great progress has been achieved in the past decades. In this paper, we concentrate ourselves on the issue of off-line text-independent writer recognition by summarizing the state of the art methods from the perspectives of feature extraction and classification. We also exhibit some public datasets and compare the performance of the existing prominent methods. The comparison demonstrates that the performance of the methods based on frequency domain features decreases seriously when the number of writers becomes larger, and that spatial distribution features are superior to both frequency domain features and shape features in capturing the individual traits. Yujie Xiong, Yue Lu 0001, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2017 | An Image-Based Approach to Detection of Fake CoinsabstractWe propose a new approach to detect fake coins using their images in this paper. A coin image is represented in the dissimilarity space, which is a vector space constructed by comparing the image with a set of prototypes. Each dimension measures the dissimilarity between the image under consideration and a prototype. In order to obtain the dissimilarity between two coin images, the local keypoints on each image are detected and described. Based on the characteristics of the coin, the matched keypoints between the two images can be identified in an efficient manner. A post-processing procedure is further proposed to remove mismatched keypoints. Due to the limited number of fake coins in real life, one-class learning is conducted for fake coin detection, so only genuine coins are needed to train the classifier. Extensive experiments have been carried out to evaluate the proposed approach on different data sets. The impressive results have demonstrated its validity and effectiveness. Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | Multilingual scene character recognition with co-occurrence of histogram of oriented gradients
Shangxuan Tian, Ujjwal Bhattacharya, Shijian Lu, Bolan Su, Xiaohua Wei, Yue Lu 0001, Chew Lim Tan |
Pattern Recognit. | 7 |
| 2016 | Recognition of handwritten Chinese address with writing variations
Xiaohua Wei, Shujing Lu, Ying Wen 0003, Yue Lu 0001 |
Pattern Recognit. Lett. | 4 |
| 2015 | Using multiple sequence alignment and statistical language model to integrate multiple Chinese address recognition outputsabstractDifferent recognizers may result in different mistakes when they are used to recognize a Chinese address. In this paper, we present a method of combining multiple Chinese address recognition outputs to improve Chinese address recognition accuracy. The method first employs multiple sequence alignment to generate a lattice of candidate hypotheses from multiple different recognizer outputs and then applies statistical language model to choose the maximum likelihood candidate sequence. Taking the maximum as the final decision, the performance of our method is superior, compared to the single recognizers and Miyao's method. The experiments on the address images of real envelopes demonstrate that the proposed method increases the character recognition accuracy rate from 95.80% to 98.38%, with 61.30% error reduction. Furthermore, the corrected sorting rate of an automatic mail sorting system increases from 84.11% to 93.72%. Shengchang Chen, Shujing Lu, Ying Wen 0003, Yue Lu 0001 |
ICDAR | 4 |
| 2015 | Text detection in nature scene images using two-stage nontext filteringabstractWe present a text detection method in natural scene images based on two-stage nontext filtering. Firstly, we detect multi-channel maximally stable extremal regions (MSERs) as character candidates. To reduce the amount of repeating components, we merge the MSERs by choosing the most character-like ones when overlap happens. Then nontext components are filtered out by a two-stage labeling procedure, wherein we combine random forests with CRF. Finally, components labeled as text are grouped into words by an edge-cut strategy, and false positives are eliminated by a HOG-based classifier. The experimental results on the ICDAR2013 database show the effectiveness of the proposed method. Yue Lu 0001, Shiliang Sun |
ICDAR | 2 |
| 2015 | Text-independent writer identification using SIFT descriptor and contour-directional featureabstractThis paper presents a method for text-independent writer identification using SIFT descriptor and contour-directional feature (CDF). The proposed method contains two stages. In the first stage, a codebook of local texture patterns is constructed by clustering a set of SIFT descriptors extracted from images. Using this codebook, the occurrence histograms are calculated to determine the similarities between different images. For each image, we obtain a candidate list of reference images. The next stage is to refine the candidate list using the contour-directional feature and SIFT descriptor. The proposed method is evaluated with two datasets: the ICFHR2012-Latin dataset and the ICDAR2013 dataset. Experimental results show that the proposed method outperforms the state-of-the-art algorithms and archives the best performance. Yujie Xiong, Ying Wen 0003, Patrick Shen-Pei Wang, Yue Lu 0001 |
ICDAR | 4 |
| 2015 | HoG based two-directional Dynamic Time Warping for handwritten word spottingabstractWe present a Histogram of Oriented Gradient (HoG) based two-directional Dynamic Time Warping (DTW) matching method for handwritten word spotting. Firstly, we extract HoG descriptors from each cell in the normalized images. Then we connect the HoG descriptors in the same column and get a sequence of feature vectors. We do the same operation for the HoG descriptors in the same row. We then apply the two-directional DTW method to calculate the distance between the feature vectors sequences extracted from the query word and the candidate one. The experimental results show that the two-directional DTW is more robust to word deformation than the traditional DTW. And the local features such as HoG, LBP and SIFT combined with the two-directional DTW method outperform the method using the local feature descriptors directly. The HoG based two-directional DTW get the highest mean average precision on both the George Washington dataset and the CASIA-HWDB 2.1 dataset. Shunyi Yao, Ying Wen 0003, Yue Lu 0001 |
ICDAR | 3 |
| 2015 | Scene text detection using sequential nontext filteringabstractWe present a scene text detection method based on sequential nontext filtering. Firstly, we start our work with multi-channel maximally stable extremal region (MSER) detection. Then nontext components are eliminated by a four-stage sequential nontext filtering strategy which consists of inner-channel MSER pruning, between-channel MSER pruning, unary feature-based nontext filtering, and binary feature-based nontext filtering. Finally, text components are grouped into words and false positives are eliminated. The proposed method achieves the state-of-the-art on the ICDAR2013 database when compared with some existing methods. Yue Lu 0001, Ying Wen 0003 |
ICIP | 2 |
| 2015 | Integrating word embeddings and traditional NLP features to measure textual entailment and semantic relatedness of sentence pairsabstractRecent years the distributed representations of words (i.e., word embeddings) have been shown to be able to significantly improve performance in many natural language processing tasks, such as pos-of-tag tagging, chunking, named entity recognition and sentiment polarity judgement, etc. However, previous tasks only involve a single sentence. In contrast, this paper evaluates the effectiveness of word embeddings in sentence pair classification or regression problems. Specifically, we propose novel simple yet effective features based on word embeddings and extract many traditional linguistic features. Then these features serve as input of a classification/regression algorithm in isolation and in combination. Evaluations are conducted on three sentence pair classification/regression tasks, i.e., textual entailment, cross-lingual textual entailment and semantic relatedness estimation. Experiments on benchmark datasets provided by Semantic Evaluation 2013 and 2014 showed that using word embeddings is able to significantly improve the performance and our results outperform the best achieved results so far. Jiang Zhao, Man Lan, Zhengyu Niu, Yue Lu 0001 |
IJCNN | 4 |
| 2015 | Variable-Length Signature for Near-Duplicate Image MatchingabstractWe propose a variable-length signature for near-duplicate image matching in this paper. An image is represented by a signature, the length of which varies with respect to the number of patches in the image. A new visual descriptor, viz., probabilistic center-symmetric local binary pattern, is proposed to characterize the appearance of each image patch. Beyond each individual patch, the spatial relationships among the patches are captured. In order to compute the similarity between two images, we utilize the earth mover's distance which is good at handling variable-length signatures. The proposed image signature is evaluated in two different applications, i.e., near-duplicate document image retrieval and near-duplicate natural image detection. The promising experimental results demonstrate the validity and effectiveness of the proposed variable-length signature. Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
IEEE Trans. Image Process. | 2 |
| 2014 | Novel Global and Local Features for Near-Duplicate Document Image MatchingabstractA new near-duplicate document image matching approach is proposed. Globally, we model the spatial arrangements of objects in an image. Locally, the micro-patterns within each object are captured. To define a micro-pattern, the N-nary center-symmetric gray value differences in an image local neighborhood of a variable radius are exploited. A visual descriptor is proposed to characterize the appearance of the object based on micro-pattern distributions. By combining the global and local features, each document image is represented by a compact signature with a variable length. We employ Earth Mover's Distance for image dissimilarity computation, which stands out for its remarkable ability to tolerate the instability of object segmentation by allowing many-to-many correspondence among objects. Extensive experiments on two data sets demonstrate the effectiveness of the proposed approach. Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
ICPR | 2 |
| 2014 | Near-duplicate document image matching: A graphical perspective
Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
Pattern Recognit. | 2 |
| 2013 | Modeling Local Word Spatial Configurations for Near Duplicate Document Image RetrievalabstractThe issue of near duplicate document image retrieval is addressed in this paper, which is characterized by not only encoding each individual word in the image but also modeling its local spatial configuration. On representing each word in the image as a string in terms of its shape characteristics, a lexicon is first learnt from a training set. Then a word in an arbitrary document image can be soft assigned to a weighted combination of several nearest neighbors in the lexicon. The rationale behind soft-assignment is to tolerate the distortions induced by character segmentations which are error-prone in degraded document images. Most importantly, we look beyond the single word and capture the local spatial configuration for each word which plays a very important role in human perception. It provides much useful information in discriminating between different document images compared with the single word. A graph, benefitting from its great representative power, is built for each word to model its relationships with the neighborhoods locally. The local word spatial configurations are integrated within the inverted file index structure to achieve scalable retrieval. Thus the retrieval of near duplicate document images is formulated as a voting problem. Experimental results on 45,000 document images demonstrate that the proposed approach brings significant improvements in successful retrieval of near duplicate images. Li Liu 0010, Yue Lu 0001, Ching Y. Suen, Jinhua Xu |
ICDAR | 2 |
| 2012 | Discriminative common vectors based on the Gram-Schmidt reorthogonalization for the small sample size problemabstractThe discriminative common vectors (DCV) algorithm shows better face recognition effects than some commonly used linear discriminant algorithms, which uses the subspace methods and the Gram-Schmidt orthogonalization (GSO) procedure to obtain the DCV. However, the Gram-Schmidt technique may produce a set of vectors which is far from orthogonal so that sometimes the orthogonality may be lost completely. Hence, the effectiveness of the DCV is also decreased. In this paper, we proposed an improved DCV method based on the GSO. For obtaining an accurate projection onto the corresponding space, the orthogonal basis problem is usually solved with the Gram-Schmidt process with reorthogonalization. Thus, the effectiveness of the DCV can be improved and the experimental results show that the proposed method is better for the small sample size problem as compared to the DCV. Ying Wen 0003, Lianghua He, Yue Lu 0001 |
ICASSP | 3 |
| 2012 | Multitask multiclass privileged information support vector machines
You Ji, Shiliang Sun, Yue Lu 0001 |
ICPR | 3 |
| 2012 | Document image matching using probabilistic graphical models
Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
ICPR | 2 |
| 2012 | Connective prediction using machine learning for implicit discourse relation classificationabstractImplicit discourse relation classification is a challenge task due to missing discourse connective. Some work directly adopted machine learning algorithms and linguistically informed features to address this task. However, one interesting solution is to automatically predict implicit discourse connective. In this paper, we present a novel two-step machine learning-based approach to implicit discourse relation classification. We first use machine learning method to automatically predict the discourse connective that can best express the implicit discourse relation. Then the predicted implicit discourse connective is used to classify the implicit discourse relation. Experiments on Penn Discourse Treebank 2.0 (PDTB) and Biomedical Discourse Relation Bank (BioDRB) show that our method performs better than the baseline system and previous work. Man Lan, Yue Lu 0001, Zhengyu Niu, Chew Lim Tan |
IJCNN | 3 |
| 2012 | Cost-Sensitive Neural Network Classifiers for Postcode RecognitionabstractMost traditional postcode recognition systems implicitly assumed that the distribution of the 10 numerals (0–9) is balanced. However it is far from a reasonable setting because the distribution of 0–9 in postcodes of a country or a city is generally imbalanced. Some numerals appear in more postcodes, while some others do not. In this paper, we study cost-sensitive neural network classifiers to address the class imbalance problem in postcode recognition. Four methods, namely: cost-sampling, cost-convergence, rate-adapting and threshold-moving are considered in training neural networks. Cost-sampling adjusts the distribution of the training data such that the costs of classes are conveyed explicitly by the appearances of their instances. Cost-convergence and rate-adapting are carried out in training phase by modifying the architecture of training algorithms of the neural network. Threshold-moving tries to increase the probability estimations of expensive classes to avoid the samples with higher costs to be misclassified. 10,702 postcode images are experimented using five cost matrices based on the distribution of numerals in postcodes. The results suggest that cost-sensitive learning is indeed effective on class imbalanced postcode analysis and recognition. It also reveals that cost-sampling on a proper cost matrix outperforms others in this application. Shujing Lu, Li Liu 0010, Yue Lu 0001, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | Text Location in Scene Images using Visual Attention ModelabstractLocating text region from an image of nature scene is significantly helpful for better understanding the semantic meaning of the image, which plays an important role in many applications such as image retrieval, image categorization, social media processing, etc. Traditional approach relies on the low level image features to progressively locate the candidate text regions. However, these approaches often suffer for the cases of the clutter background since the adopted low level image features are fairly simple which may not reliably distinguish text region from the clutter background. Motivated by the recent popular research on attention model, salience detection is revisited in this paper. Based on the case of text detection on nature scene image, saliency map is further analyzed and is adjusted accordingly. Using the adjusted saliency map, the candidate text regions detected by the common low level features are further verified. Moreover, efficient low level text feature, Histogram of Edge-direction (HOE), is adopted in this paper, which statistically describes the edge direction information of the region of interest on the image. Encouraging experimental results have been obtained on the nature scene images with the text of various languages. Qiaoyu Sun, Yue Lu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | Retrieval of Envelope Images Using Graph MatchingabstractA graph matching approach is proposed to retrieve envelope images from a large image database. First, the graph representation of an envelop image is generated based on the image segmentation results, in which each node corresponds to one segmented region. The attributes of nodes and edges in the graph are described by characteristics of the envelope image. Second, a minimum weighted bipartite graph matching method is employed to compute the distance between two graphs. Finally, the whole retrieval system including two principal stages is presented, namely, rough matching and fine matching. The experiments on a database of envelope images captured from real-life mail pieces demonstrate that the proposed method achieves promising results. Li Liu 0010, Yue Lu 0001, Ching Y. Suen |
ICDAR | 2 |
| 2011 | The stochastic approximation method for adaptive Bayesian classifiers: towards online brain-computer interfaces
Shiliang Sun, Yue Lu 0001, Youguang Chen |
Neural Comput. Appl. | 2 |
| 2011 | An Algorithm for License Plate Recognition Applied to Intelligent Transportation SystemabstractAn algorithm for license plate recognition (LPR) applied to the intelligent transportation system is proposed on the basis of a novel shadow removal technique and character recognition algorithms. This paper has two major contributions. One contribution is a new binary method, i.e., the shadow removal method, which is based on the improved Bernsen algorithm combined with the Gaussian filter. Our second contribution is a character recognition algorithm known as support vector machine (SVM) integration. In SVM integration, character features are extracted from the elastic mesh, and the entire address character string is taken as the object of study, as opposed to a single character. This paper also presents improved techniques for image tilt correction and image gray enhancement. Our algorithm is robust to the variance of illumination, view angle, position, size, and color of the license plates when working in a complex environment. The algorithm was tested with 9026 images, such as natural-scene vehicle images using different backgrounds and ambient illumination particularly for low-resolution images. The license plates were properly located and segmented as 97.16% and 98.34%, respectively. The optical character recognition system is the SVM integration with different character features, whose performance for numerals, Kana, and address recognition reached 99.5%, 98.6%, and 97.8%, respectively. Combining the preceding tests, the overall performance of success for the license plate achieves 93.54% when the system is used for LPR in various complex conditions. Ying Wen 0003, Yue Lu 0001, Jingqi Yan, Karen M. von Deneen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2010 | A Visual Attention Based Approach to Text ExtractionabstractA visual attention based approach is proposed to extract texts from complicated background in camera-based images. First, it applies the simplified visual attention model to highlight the region of interest (ROI) in an input image and to yield a map, named the VA map, consisting of the ROIs. Second, an edge map of image containing the edge information of four directions is obtained by Sobel operators. Character areas are detected by connected component analysis and merged into candidate text regions. Finally, the VA map is employed to confirm the candidate text regions. The experimental results demonstrate that the proposed method can effectively extract text information and locate text regions contained in camera-based images. It is robust not only for font, size, color, language, space, alignment and complexity of background, but also for perspective distortion and skewed texts embedded in images. Qiaoyu Sun, Yue Lu 0001, Shiliang Sun |
ICPR | 2 |
| 2009 | Supervised and Traditional Term Weighting Methods for Automatic Text CategorizationabstractIn vector space model (VSM), text representation is the task of transforming the content of a textual document into a vector in the term space so that the document could be recognized and classified by a computer or a classifier. Different terms (i.e. words, phrases, or any other indexing units used to identify the contents of a text) have different importance in a text. The term weighting methods assign appropriate weights to the terms to improve the performance of text categorization. In this study, we investigate several widely-used unsupervised (traditional) and supervised term weighting methods on benchmark data collections in combination with SVM and kappa NN algorithms. In consideration of the distribution of relevant documents in the collection, we propose a new simple supervised term weighting method, i.e. tf.rf, to improve the terms' discriminating power for text categorization task. From the controlled experimental results, these supervised term weighting methods have mixed performance. Specifically, our proposed supervised term weighting method, tf.rf, has a consistently better performance than other term weighting methods while other supervised term weighting methods based on information theory or statistical metric perform the worst in all experiments. On the other hand, the popularly used tf.idf method has not shown a uniformly good performance in terms of different data sets. Man Lan, Chew Lim Tan, Jian Su 0002, Yue Lu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2008 | Adaptive EEG signal classification using stochastic approximation methodsabstractClassification of time-varying electrophysiological signals is an important problem in the development of brain-computer interfaces (BCIs). Designing adaptive classifiers is a potential way to address this task. In this paper, Bayesian classifiers with Gaussian mixture models (GMMs) are adopted as the decision rule to classify electroencephalogram (EEG) signals. The stochastic approximation method (SAM) is used as the specific gradient descent method for updating the parameters of mean values and covariance matrices in the distribution of GMMs, where the parameters are simultaneously updated in a batch mode. Experimental results using data from a BCI show that the stochastic approximation method is effective for EEG classification tasks. Shiliang Sun, Man Lan, Yue Lu 0001 |
ICASSP | 3 |
| 2008 | The random electrode selection ensemble for EEG signal classification
Shiliang Sun, Changshui Zhang, Yue Lu 0001 |
Pattern Recognit. | 3 |
| 2007 | Constructing Area Voronoi Diagram Based on Direct Calculation of the Freeman Code of Expanded ContoursabstractA Voronoi diagram of image elements provides an intuitive and appealing definition of proximity, which has been suggested as an effective tool for the description of relations among the neighboring objects in a digital image. In this paper, an implementation algorithm based on direct calculation of the Freeman code of expanded contours is proposed for generating area Voronoi diagram of connected components. A closed convex polygon is utilized to bound each connected component, as an approximate representation, and the contour is represented using Freeman chain coding, from which we can compute the corresponding Freeman chain coding of its expanded contour directly, without recourse to the operation on pixels. While the contours iteratively expand outwards, the Voronoi diagram is constructed by the intersections of the expanded contours from different connected components. The experimental results show that our proposed approach significantly improves the speed of constructing area Voronoi diagram in digital images. Yue Lu 0001, Chunyun Xiao, Chew Lim Tan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2007 | Handwritten Bangla numeral recognition system and its application to postal automation
Ying Wen 0003, Yue Lu 0001 |
Pattern Recognit. | 2 |
| 2006 | Bangla/English Script Identification Based on Analysis of Connected Component Profiles
Yue Lu 0001, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2005 | Constructing Area Voronoi Diagram in Document ImagesabstractVoronoi diagram of image elements provides an intuitive and appealing definition of proximity, which has been suggested as an effective tool for the description of relations among the neighboring objects in a digital image. In this paper, a fast implementation algorithm is proposed for generating area Voronoi diagram of connected components in document images. A closed convex polygon is utilized to bound each connected component, and the contour is represented using Freeman chain coding, from which we can compute the corresponding Freeman chain coding of its expanded contour directly, without recourse to the operation on pixels. While the contours iteratively expand outwards, the Voronoi diagram is constructed by the intersections of the expanded contours from different connected components. The experimental results show that our proposed approach significantly improves the speed of constructing area Voronoi diagram. Yue Lu 0001, Chew Lim Tan |
ICDAR | 1 |
| 2004 | Word Grouping in Document Images Based on Voronoi Tessellation
Yue Lu 0001, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2004 | A search engine for imaged documents in PDF filesabstractProceedings of Sheffield SIGIR - Twenty-Seventh Annual International ACM SIGIR Conference on Research and Development in Information Retrieval Yue Lu 0001, Li Zhang 0005, Chew Lim Tan |
SIGIR | 1 |
| 2004 | Chinese Word Searching In Imaged DocumentsabstractAn approach to searching for user-specified words in imaged Chinese documents, without the requirements of layout analysis and OCR processing of the entire documents, is proposed in this paper. A small number of Chinese characters that cannot be successfully bounded using connected component analysis due to larger gaps between elements within the characters are blacklisted. A suitable character that is not included in the blacklist is chosen from the user-specified word as the initial character to search for a matching candidate in the document. Once a matched candidate is found, the adjacent characters in the horizontal and vertical directions are examined for matching with other corresponding characters in the user-specified word, subject to the constraints of alignment (either horizontal or vertical direction) and size similarity. A weighted Hausdorff distance is proposed for the character matching. Experimental results show that the present method can effectively search the user-specified Chinese words from the document images with the format of either horizontal or vertical text lines, or both appearing on the same image. Yue Lu 0001, Chew Lim Tan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | Information Retrieval in Document Image DatabasesabstractWith the rising popularity and importance of document images as an information source, information retrieval in document image databases has become a growing and challenging problem. In this paper, we propose an approach with the capability of matching partial word images to address two issues in document image retrieval: word spotting and similarity measurement between documents. First, each word image is represented by a primitive string. Then, an inexact string matching technique is utilized to measure the similarity between the two primitive strings generated from two word images. Based on the similarity, we can estimate how a word image is relevant to the other and, thereby, decide whether one is a portion of the other. To deal with various character fonts, we use a primitive string which is tolerant to serif and font differences to represent a word image. Using this technique of inexact string matching, our method is able to successfully handle the problem of heavily touching characters. Experimental results on a variety of document image databases confirm the feasibility, validity, and efficiency of our proposed approach in document image retrieval. Yue Lu 0001, Chew Lim Tan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Keyword Searching in Compressed Document ImagesabstractSummary form only given. A compressed pattern matching method for searching keywords from the CCIT group 4-compressed document images, without explicit decompression, is presented. According to the CCIT Group 4 standards, each coded position indicates current pixel color is different from its previous pixel, except for the next coded positions of the pass mode. The changing elements from the compressed images are extracted and are then utilized to segment and bound the word objects and to measure the similarity of two word images. A two-stage matching strategy is constructed to measure the dissimilarity between the template image of the user's query word and the word extracted from document images. Experiments were conducted to verify the validity of the approach. The results show that the proposed approach was much faster than the traditional approach, because it avoids the pixel-level processing for analyzing the connected components and extracting word features. Yue Lu 0001, Chew Lim Tan |
DCC | 1 |
| 2003 | Word Searching in CCITT Group 4 Compressed Document ImagesabstractIn this paper, we present a compressed pattern matching method for searching user queried words in the CCITT Group 4 compressed document images, without decompressing. The feature pixels composed of black changing elements and white changing elements are extracted directly from the CCITT Group 4 compressed document images. The connected components are labeled based on a line-by-line strategy according to the relative positions between the changing elements of the current coding line and the changing elements of the reference line. Word boxes are bounded by merging the connected components. A two-stage matching strategy is constructed to measure the dissimilarity between the template image of the user's query word and the words extracted from document images. Experimental results confirmed the validity of the proposed approach. Yue Lu 0001, Chew Lim Tan |
ICDAR | 1 |
| 2003 | Improved Nearest Neighbor Based Approach to Accurate Document Skew EstimationabstractThe nearest-neighbor based document skew detection methods do not require the presence of a predominant text area, and are not subject to skew angle limitation. However, the accuracy of these methods is not perfect in general. In this paper, we present an improved nearest-neighbor based approach to perform accurate document skew estimation. Size restriction is introduced to the detection of nearest-neighbor pairs. Then the chains with a largest possible number of nearest-neighbor pairs are selected, and their slopes are computed to give the skew angle of document image. Experimental results on various types of documents containing different linguistic scripts and diverse layouts show that the proposed approach has achieved an improved accuracy for estimating document image skew angle and has an advantage of being language independent. Yue Lu 0001, Chew Lim Tan |
ICDAR | 1 |
| 2003 | Document retrieval from compressed images
Yue Lu 0001, Chew Lim Tan |
Pattern Recognit. | 1 |
| 2003 | A nearest-neighbor chain based approach to skew estimation in document images
Yue Lu 0001, Chew Lim Tan |
Pattern Recognit. Lett. | 1 |
| 2002 | Word Searching in Document Images Using Word Portion Matching
Yue Lu 0001, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2002 | Segmentation of Handwritten Chinese Characters from Destination Addresses of Mail PiecesabstractIn this paper, we illustrate a method to segment handwritten Chinese characters from destination addresses of mail pieces. Fast Hough transform is utilized to detect the reference lines preprinted on the mail piece. In the segmentation, subassemblies of Chinese characters are merged based on the structural features of Chinese characters and the subassemblies' topological relations, viz. upper–lower, inside–outside and left–right relations. The width of subassemblies and the spacing between neighboring subassemblies in the whole image of the destination address are analyzed to guide the merging of the left–right subassemblies. Experimental results with real mail piece images show that the proposed approach has achieved a promising performance for segmenting handwritten Chinese characters. Yue Lu 0001, Chew Lim Tan, Kehua Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2002 | Combination of multiple classifiers using probabilistic dictionary and its application to postcode recognition
Yue Lu 0001, Chew Lim Tan |
Pattern Recognit. | 1 |
| 2001 | An Approach to Word Image Matching Based on Weighted Hausforff DistanceabstractAn approach to word image matching based on weighted Hausdorff distance (WHD) is proposed in this paper to facilitate the detection and location of the user-specified words in the document images. Preprocessing such as eliminating the space between adjacent characters in the word images and scale normalization is first done before the WHD is utilized to measure the distance between the template image and the word image extracted from the document image. Experimental results in the application of detecting the user-specified words from both English and Chinese document images show that it is a promising approach for word image matching. Yue Lu 0001, Chew Lim Tan, Weihua Huang, Liying Fan |
ICDAR | 1 |
| 2001 | Similarity measure for CCITT Group 4 compressed document imagesabstractThe similarity measure of document images has a crucial role in the area of document image retrieval. A method of measuring the similarity of CCITT Group 4 compressed document images is proposed. The features are extracted directly from the changing elements of the compressed images. Weighted Hausdorff distance is utilized to assign all of the word objects from two document images to corresponding classes by an unsupervised classifier, whereas the possible stop words are excluded. Document vectors are built by the occurrence frequency of the word object classes, and the pair-wise similarity of two document images is represented by the scalar product of the document vectors. Five group articles relating to different domains are used to test the validity of the presented approach. Yue Lu 0001, Chew Lim Tan, Liying Fan, Weihua Huang |
ICIP (1) | 1 |