Qiufeng Wang 0001

dblp:86/7443-1 · also Qiu-Feng Wang 0001 · DBLP profile ↗
← Back
86ranked-venue papers
6as first author
61since 2021 · last 2026
0000-0002-0918-4606ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 5 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 21 since 2021Databases, data management, data science and information retrieval · 18 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning
Tong Chen 0005, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003, Shuihua Wang, Ling Chen 0006, Jionglong Su, Haiyang Zhang 0004, Wei Wang 0042
LREC9
2026 You look from old classes: Towards accurate few shot class-incremental learning
Yijie Hu, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
Pattern Recognit.5
2026 A benchmark and method for photographed table reasoning
Xiaoqiang Kang, Xiaochen Zi, Xiao-Bo Jin, Kaizhu Huang, Qiufeng Wang 0001
Pattern Recognit.7
2026 A comprehensive survey of oracle character recognition: Challenges, datasets, methodology, and beyond
Xueke Chi, Qiufeng Wang 0001, Kaizhu Huang, Dahan Wang, Yongge Liu
Pattern Recognit.3
2026 IDEA: Image description enhanced CLIP-adapter for image classification
Zhipeng Ye, Qiufeng Wang 0001, Kaizhu Huang
Pattern Recognit.3
2026 EmoSENSE: Modeling Sentiment-Semantic Knowledge With Hierarchical Reinforcement Learning for Emotional Image Generation
abstract
Emotional image generation aims to create images that effectively reflect target emotions. A fundamental challenge in this task is the affective gap, which refers to the discrepancy between visual content and emotional states perceived by users. Existing methods generally assume strong and explicit associations between target emotions and specific objects (e.g., “monster” and “fear”), which limits their generalization ability when encountering uncommon emotion-object pairs. This limitation stems from two main factors: 1) Most existing approaches primarily focus on semantic alignment without explicitly modeling how emotions influence visual attributes such as brightness and colorfulness; 2) diffusion-based image generation methods have limited capability in handling diverse sentiment-semantic pairs. To address these challenges, we propose EmoSENSE, a novel hierarchical fuzzy reinforcement learning framework for the emotional image generation task. EmoSENSE consists of a high-level module and a low-level module, working collaboratively in a hierarchical structure to inject sentiment-semantic knowledge into emotional images. The high-level module quantifies sentiment-semantic correlations within a unified emotional space, connecting emotions to visual attributes. The low-level module refines this connection by optimizing a fuzzy-logic-based mapping between emotions and visual attributes through reinforcement learning, enabling flexible adaptation to diverse emotion-object pairs. Extensive qualitative and quantitative experiments on public dataset demonstrate that EmoSENSE significantly enhances both the visual quality and emotional expression ability of the generated images, achieving a 12.21% higher EmoAccuracy-8 classes than the previous state-of-the-art methods.https://github.com/forever3600/EmoSENSE.
Junyi Guo, Qiufeng Wang 0001, Yaran Chen, Fangyu Wu 0001, Eng Gee Lim
IEEE Trans. Affect. Comput.3
2026 Diff-Oracle: Learning Styles and Contents to Augment Realistic Oracle Characters in Diffusion Model
abstract
Recognizing oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains because of the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel multi-modal conditional diffusion model that generates a diverse range of controllable oracle characters by inputting random combinations of references. Given the challenge of accurately describing oracle character styles using natural language, Diff-Oracle departs from traditional diffusion models that rely primarily on text prompts by introducing a style encoder. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pre-trained language-vision model. Additionally, given the lack of explicit content information for oracle characters, ensuring that generated characters accurately represent the intended glyphs is challenging. Therefore, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model, providing content information for the generation process. Meanwhile, Diff-Oracle integrates a content encoder designed to capture specific content details from content reference images. Extensive experiments on Oracle-241 and OBC306 datasets demonstrate that Diff-Oracle significantly outperforms existing generative methods in image quality and diversity. Moreover, Diff-Oracle substantially benefits downstream recognition tasks, outperforming all existing state-of-the-art methods by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle achieves a 7.70% accuracy gain in the zero-shot setting and reaches 84.62% accuracy for unseen oracle characters, setting a new benchmark for oracle character recognition. The code is available at https://github.com/JJJingLi/Diff-Oracle .
Jing Li 0049, Qiufeng Wang 0001, Siyuan Wang 0017, Rui Zhang 0012, Kaizhu Huang, Erik Cambria
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation
abstract
Solving tabular math word problems (TMWPs) has become a critical role in evaluating the mathematical reasoning ability of large language models (LLMs), where large-scale TMWP samples are commonly required for fine-tuning. Since the collection of high-quality TMWP datasets is costly and time-consuming, recent research has concentrated on automatic TMWP generation. However, current generated samples usually suffer from issues of either correctness or diversity. In this paper, we propose a Template-driven LLM-paraphrased (TeLL) framework for generating high-quality TMWP samples with diverse backgrounds and accurate tables, questions, answers, and solutions. To this end, we first extract templates from existing real samples to generate initial problems, ensuring correctness. Then, we adopt an LLM to extend templates and paraphrase problems, obtaining diverse TMWP samples. Furthermore, we find the reasoning annotation is important for solving TMWPs. Therefore, we propose to enrich each solution with illustrative reasoning steps. Through the proposed framework, we construct a high-quality dataset TabMWP-TeLL by adhering to the question types in the TabMWP dataset, and we conduct extensive experiments on a variety of LLMs to demonstrate the effectiveness of TabMWP-TeLL in improving TMWP-solving performance.
Xiaoqiang Kang, Xiao-Bo Jin, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001
AAAI6
2025 GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMs
abstract
With the outstanding capabilities of Large Language Models (LLMs), solving math word problems (MWP) has greatly progressed, achieving higher performance on several benchmark datasets. However, it is more challenging to solve plane geometry problems (PGPs) due to the necessity of understanding, reasoning and computation on two modality data including both geometry diagrams and textual questions, where Multi-Modal Large Language Models (MLLMs) have not been extensively explored. Previous works simply regarded a plane geometry problem as multi-modal QA task, which ignored the importance of explicit parsing geometric elements from problems. To tackle this limitation, we propose to solve plane Geometry problems by Neural-Symbolic reasoning with MLLMs (GNS). We first leverage an MLLM to understand PGPs through knowledge prediction and symbolic parsing, next perform mathematical reasoning to obtain solutions, last adopt a symbolic solver to compute answers. Correspondingly, we introduce the largest PGPs dataset GNS-260K with multiple annotations including symbolic parsing, understanding, reasoning and computation. In experiments, our Phi3-Vision-based MLLM wins the first place on the PGPs solving task of MathVista benchmark, outperforming GPT-4o, Gemini Ultra and other much larger MLLMs. While LLaVA-13B-based MLLM markedly exceeded other close-source and open-source MLLMs on the MathVerse benchmark and also achieved the new SOTA on GeoQA dataset.
Maizhen Ning, Qiufeng Wang 0001, Xiaowei Huang 0001, Kaizhu Huang
AAAI3
2025 Towards Better Robustness Against Natural Corruptions in Document Tampering Localization
abstract
Marvelous advances have been exhibited in recent document tampering localization (DTL) systems. However, confronted with corrupted tampered document images, their vulnerability is fatal in real-world scenarios. While robustness against adversarial attack has been extensively studied by adversarial training (AT), the robustness on natural corruptions remains under-explored for DTL. In this paper, to overcome forensic dependency, we propose the adversarial forensic regularization (AFR) based on min-max optimization to improve robustness. Specifically, we adopt mutual information (MI) to represent forensic dependency between two random variable over tampered and authentic pixels spaces, where the MI can be approximated by Jensen-Shannon-Divergence (JSD) with empirical sampling. To further enable a trade-off between predictive representations in clean tampered document pixels and robust ones in corrupted pixels, an additional regularization term is formulated with divergence between clean and perturbed pixels distribution (DDR). Following min-max optimization framework, our method can also work well against adversarial attacks. To evaluate our proposed method, we collect a dataset (i.e., TSorie-CRP) for evaluating robustness against natural corruptions in real scenarios. Extensive experiments demonstrate the effectiveness of our method against natural corruptions. Without any surprise, our method also achieves good performance against adversarial attack on DTL benchmark datasets.
Huiru Shao, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
AAAI5
2025 BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis
abstract
3D semantic segmentation plays a fundamental and crucial role to understand 3D scenes. While contemporary state-of-the-art techniques predominantly concentrate on elevating the overall performance of 3D semantic segmentation based on general metrics (e.g. mIoU, mAcc, and oAcc), they unfortunately leave the exploration of challenging regions for segmentation mostly neglected. In this paper, we revisit 3D semantic segmentation through a more granular lens, shedding light on subtle complexities that are typically overshadowed by broader performance metrics. Concretely, we have delineated 3D semantic segmentation errors into four comprehensive categories as well as corresponding evaluation metrics tailored to each. Building upon this categorical framework, we introduce an innovative 3D semantic segmentation network called BFANet that incorporates detailed analysis of semantic boundary features. First, we design the boundary-semantic module to decouple point cloud features into semantic and boundary features, and fuse their query queue to enhance semantic features with attention. Second, we introduce a more concise and accelerated boundary pseudo-label calculation algorithm, which is 3.9 times faster than the state-of-the-art, offering compatibility with data augmentation and enabling efficient computation in training. Extensive experiments on benchmark data indicate the superiority of our BFANet model, confirming the significance of emphasizing the four uniquely designed metrics. Code is available at https://github.com/weiguangzhao/BFANet.
Weiguang Zhao, Rui Zhang 0012, Qiufeng Wang 0001, Kaizhu Huang
CVPR3
2025 HiGarment: Cross-Modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment Image
abstract
Diffusion-based garment synthesis tasks primarily focus on the design phase in the fashion domain, while the garment production process remains largely underexplored. To bridge this gap, we introduce a new task: Flat Sketch to Realistic Garment Image (FS2RG), which generates realistic garment images by integrating flat sketches and textual guidance. FS2RG presents two key challenges: 1) fabric characteristics are solely guided by textual prompts, providing insufficient visual supervision for diffusion-based models, which limits their ability to capture fine-grained fabric details; 2) flat sketches and textual guidance may provide conflicting information, requiring the model to selectively preserve or modify garment attributes while maintaining structural coherence. To tackle this task, we propose HiGarment, a novel framework that comprises two core components: i) a multi-modal semantic enhancement mechanism that enhances fabric representation across textual and visual modalities, and ii) a harmonized cross-attention mechanism that dynamically balances information from flat sketches and text prompts, allowing controllable synthesis by generating either sketch-aligned (image-biased) or text-guided (text-biased) outputs. Furthermore, we collect Multi-modal Detailed Garment, the largest open-source dataset for garment generation. Experimental results and user studies demonstrate the effectiveness of HiGarment in garment synthesis. The code and dataset are available at https://github.com/Maple498/HiGarment.
Junyi Guo, Fangyu Wu 0001, Huanda Lu, Qiufeng Wang 0001, Wenmian Yang, Eng Gee Lim, Dongming Lu
ICCV5
2025 Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant Representation
Zhaorui Tan, Xi Yang 0008, Tan Pan, Chen Jiang 0006, Xin Guo 0010, Qiufeng Wang 0001, Anh Nguyen 0003, Yuan Qi 0001, Kaizhu Huang
ICCV7
2025 Towards Cross-Modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
Junyi Yuan, Jian Zhang 0002, Fangyu Wu 0001, Huanda Lu, Dongming Lu, Qiufeng Wang 0001
ICDAR (4)6
2025 The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
Kaizhu Huang, Qiufeng Wang 0001, Xiao-Bo Jin
ICDM5
2025 Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
abstract
Exceptional mathematical reasoning ability is one of the key features that demonstrate the power of large language models (LLMs). How to comprehensively define and evaluate the mathematical abilities of LLMs, and even reflect the user experience in real-world scenarios, has emerged as a critical issue. Current benchmarks predominantly concentrate on problem-solving capabilities, presenting a substantial risk of model overfitting and fails to accurately measure the genuine mathematical reasoning abilities. In this paper, we argue that if a model really understands a problem, it should be robustly and readily applied across a diverse array of tasks. To this end, we introduce MathCheck, a well-designed checklist for testing task generalization and reasoning robustness, as well as an automatic tool to generate checklists efficiently. MathCheck includes multiple mathematical reasoning tasks and robustness tests to facilitate a comprehensive evaluation of both mathematical reasoning ability and behavior testing. Utilizing MathCheck, we develop MathCheck-GSM and MathCheck-GEO to assess mathematical textual reasoning and multi-modal reasoning capabilities, respectively, serving as upgraded versions of benchmarks including GSM8k, GeoQA, UniGeo, and Geometry3K. We adopt MathCheck-GSM and MathCheck-GEO to evaluate over 26 LLMs and 17 multi-modal LLMs, assessing their comprehensive mathematical reasoning abilities. Our results demonstrate that while frontier LLMs like GPT-4o continue to excel in various abilities on the checklist, many other model families exhibit a significant decline. Further experiments indicate that, compared to traditional math benchmarks, MathCheck better reflects true mathematical abilities and represents mathematical intelligence more linearly, thereby supporting our design. Using MathCheck, we can also efficiently conduct informative behavior analysis to deeply investigate models. Finally, we show that our proposed checklist paradigm can easily extend to other reasoning tasks for their comprehensive evaluation.
Shudong Liu 0004, Maizhen Ning, Wei Liu 0131, Jindong Wang 0001, Derek F. Wong, Xiaowei Huang 0001, Qiufeng Wang 0001, Kaizhu Huang
ICLR8
2025 DvD: Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model
abstract
Document dewarping aims to rectify deformations in photographic document images, thus improving text readability, which has attracted much attention and made great progress, but it is still challenging to preserve document structures. Given recent advances in diffusion models, it is natural for us to consider their potential applicability to document dewarping. However, it is far from straightforward to adopt diffusion models in document dewarping due to their unfaithful control on highly complex document images (e.g., 2000 × 3000 resolution). In this paper, we propose DvD, the first generative model to tackle document Dewarping via a Diffusion framework. To be specific, DvD introduces a coordinate-level denoising instead of typical pixel-level denoising, generating a mapping for deformation rectification. In addition, we further propose a time-variant condition refinement mechanism to enhance the preservation of document structures. In experiments, we find that current document dewarping benchmarks can not evaluate dewarping models comprehensively. To this end, we present AnyPhotoDoc6300, a rigorously designed large-scale document dewarping benchmark comprising 6,300 real image pairs across three distinct domains, enabling fine-grained evaluation of dewarping models. Comprehensive experiments demonstrate that our proposed DvD can achieve state-of-the-art performance with acceptable computational efficiency on multiple metrics across various benchmarks, including DocUNet, DIR300, and AnyPhotoDoc6300. The new benchmark and code will be publicly available at https://github.com/hanquansanren/DvD.
Huangcheng Lu, Maizhen Ning, Xiaowei Huang 0001, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001
SIGGRAPH Asia7
2025 WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU Communications
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Chaoyi Pang
USENIX ATC4
2025 Covariance-Based Space Regularization for Few-Shot Class Incremental Learning
Yijie Hu, Guanyu Yang 0002, Zhaorui Tan, Xiaowei Huang 0001, Kaizhu Huang, Qiufeng Wang 0001
WACV6
2025 Revisiting 3D point cloud analysis with Markov process
Chenru Jiang, Wuwei Ma, Kaizhu Huang, Qiufeng Wang 0001, Xi Yang 0008, Weiguang Zhao, Junwei Wu 0001, Xinheng Wang 0001, Jimin Xiao, Zhenxing Niu
Pattern Recognit.4
2025 AlignMalloc: Warp-Aware Memory Rearrangement Aligned With UVM Prefetching for Large-Scale GPU Dynamic Allocations
abstract
As parallel computing tasks rapidly expand in both complexity and scale, the need for efficient GPU dynamic memory allocation becomes increasingly important. While progress has been made in developing dynamic allocators for substantial applications, their real-world applicability is still limited due to inefficient memory access behaviors. This paper introduces AlignMalloc, a novel memory management system that aligns with the Unified Virtual Memory (UVM) prefetching strategy, significantly enhancing both memory allocation and access performance in large-scale dynamic allocation scenarios. We analyze the fundamental inefficiencies in UVM access and first reveal the mismatch between memory access and UVM prefetching methods. To resolve this issue, AlignMalloc implements a warp-aware memory rearrangement strategy that exploits the regularity of warps to align with the UVM's static prefetching setup. Additionally, AlignMalloc introduces an OR tree-based structure within a host-co-managed framework to further optimize dynamic allocation. Comprehensive experiments demonstrate that AlignMalloc substantially outperforms current state-of-the-art systems, achieving up to$2.7 \times$improvement in dynamic allocation and$2.3 \times$in memory access. Additionally, eight real-world applications with diverse memory access patterns exhibit consistent performance enhancements, with average speedups$1.5 \times$.
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Eng Gee Lim, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.4
2024 Unraveling Batch Normalization for Realistic Test-Time Adaptation
abstract
While recent test-time adaptations exhibit efficacy by adjusting batch normalization to narrow domain disparities, their effectiveness diminishes with realistic mini-batches due to inaccurate target estimation. As previous attempts merely introduce source statistics to mitigate this issue, the fundamental problem of inaccurate target estimation still persists, leaving the intrinsic test-time domain shifts unresolved. This paper delves into the problem of mini-batch degradation. By unraveling batch normalization, we discover that the inexact target statistics largely stem from the substantially reduced class diversity in batch. Drawing upon this insight, we introduce a straightforward tool, Test-time Exponential Moving Average (TEMA), to bridge the class diversity gap between training and testing batches. Importantly, our TEMA adaptively extends the scope of typical methods beyond the current batch to incorporate a diverse set of class information, which in turn boosts an accurate target estimation. Built upon this foundation, we further design a novel layer-wise rectification strategy to consistently promote test-time performance. Our proposed method enjoys a unique advantage as it requires neither training nor tuning parameters, offering a truly hassle-free solution. It significantly enhances model robustness against shifted domains and maintains resilience in diverse real-world scenarios with various batch sizes, achieving state-of-the-art performance on several major benchmarks. Code is available at https://github.com/kiwi12138/RealisticTTA.
Zixian Su, Jingwei Guo 0001, Xi Yang 0008, Qiufeng Wang 0001, Kaizhu Huang
AAAI5
2024 MathAttack: Attacking Large Language Models towards Math Solving Ability
abstract
With the boom of Large Language Models (LLMs), the research of solving Math Word Problem (MWP) has recently made great progress. However, there are few studies to examine the robustness of LLMs in math solving ability. Instead of attacking prompts in the use of LLMs, we propose a MathAttack model to attack MWP samples which are closer to the essence of robustness in solving math problems. Compared to traditional text adversarial attack, it is essential to preserve the mathematical logic of original MWPs during the attacking. To this end, we propose logical entity recognition to identify logical entries which are then frozen. Subsequently, the remaining text are attacked by adopting a word-level attacker. Furthermore, we propose a new dataset RobustMath to evaluate the robustness of LLMs in math solving ability. Extensive experiments on our RobustMath and two another math benchmark datasets GSM8K and MultiAirth show that MathAttack could effectively attack the math solving ability of LLMs. In the experiments, we observe that (1) Our adversarial samples from higher-accuracy LLMs are also effective for attacking LLMs with lower accuracy (e.g., transfer from larger to smaller-size LLMs, or from few-shot to zero-shot prompts); (2) Complex MWPs (such as more solving steps, longer text, more numbers) are more vulnerable to attack; (3) We can improve the robustness of LLMs by using our adversarial samples in few-shot prompts. Finally, we hope our practice and observation can serve as an important attempt towards enhancing the robustness of LLMs in math solving ability. The code and dataset is available at: https://github.com/zhouzihao501/MathAttack.
Qiufeng Wang 0001, Mingyu Jin, Jianan Ye, Wei Liu 0131, Wei Wang 0042, Xiaowei Huang 0001, Kaizhu Huang
AAAI2
2024 Generating Valid and Natural Adversarial Examples with Large Language Models
abstract
Deep learning-based natural language processing (NLP) models, particularly pre-trained language models (PLMs), have been revealed to be vulnerable to adversarial attacks. However, the adversarial examples generated by many mainstream word-level adversarial attack models are neither valid nor natural, leading to the loss of semantic maintenance, grammaticality, and human imperceptibility. Based on the exceptional capacity of language understanding and generation of large language models (LLMs), we propose LLM-Attack, which aims at generating both valid and natural adversarial examples with LLMs. The method consists of two stages: word importance ranking (which searches for the most vulnerable words) and word synonym replacement (which substitutes them with their synonyms obtained from LLMs). Experimental results on the Movie Review (MR), IMDB, and Yelp Review Polarity datasets against the baseline adversarial attack models illustrate the effectiveness of LLM-Attack, and it outperforms the baselines in human and GPT-4 evaluation by a significant margin. The model can generate adversarial examples that are typically valid and natural, with the preservation of semantic meaning, grammaticality, and human imperceptibility.
Wei Wang 0042, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003
CSCWD4
2024 Delving into Adversarial Robustness on Document Tampering Localization
Huiru Shao, Zhuang Qian, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
ECCV (65)6
2024 Class Incremental Learning for Character String Recognition
Yijie Hu, Yan-Ming Zhang 0001, Kaizhu Huang, Qiufeng Wang 0001
ICDAR (5)4
2024 Coarse-to-Fine Document Image Registration for Dewarping
Qiufeng Wang 0001, Kaizhu Huang, Xiaomeng Gu, Fengjun Guo
ICDAR (4)2
2024 SyncMalloc: A Synchronized Host-Device Co-Management System for GPU Dynamic Memory Allocation across All Scales
abstract
Dynamic memory allocation on GPUs, increasingly crucial for applications with dynamic computational patterns, encounters significant challenges due to the complex calculations with intricate branches and substantial memory resources consumed by metadata from massive thread allocations. Despite the current research, there is a lack of a scalable and flexible solution that effectively manages dynamic memory allocation while minimizing memory usage on GPUs. This paper introduces SyncMalloc, a synchronized Host-Device Co-Management system that is specifically designed to adeptly handle dynamic memory allocations of diverse magnitudes. Through the integration of pipelining and producer-consumer mechanisms, SyncMalloc effectively reduces communication overhead and resolves architectural mismatches, further enhancing its capability through synergistic integration with CUDA’s unified memory to facilitate oversubscription. Moreover, SyncMalloc advances slab-based memory management to enhance the efficiency of small allocations, reducing conflict probabilities and overhead in high-activity scenarios. Finally, we present a comprehensive performance evaluation, expanding benchmarks and measurement dimensions to reflect the performance of real-world applications more accurately. The experimental results demonstrate the effectiveness of SyncMalloc in supporting dynamic GPU allocations scaled from 4B to 200GB from multiple perspectives. Our source code is available at https://github.com/jjZhang94/SyncMalloc.
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Genlang Chen, Qiufeng Wang 0001
ICPP6
2024 Document Registration: Towards Automated Labeling of Pixel-Level Alignment Between Warped-Flat Documents
Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001, Fengjun Guo, Xiaomeng Gu
ACM Multimedia2
2024 Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
abstract
Vision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A logical regularization termed L-Reg is derived which bridges a logical analysis framework to image classification. Our work reveals that L-Reg reduces the complexity of the model in terms of the feature distribution and classifier weights. Specifically, we unveil the interpretability brought by L-Reg, as it enables the model to extract the salient features, such as faces to persons, for classification. Theoretical analysis and experiments demonstrate that L-Reg enhances generalization across various scenarios, including multi-domain generalization and generalized category discovery. In complex real-world scenarios where images span unknown classes and unseen domains, L-Reg consistently improves generalization, highlighting its practical efficacy.
Zhaorui Tan, Xi Yang 0008, Qiufeng Wang 0001, Anh Nguyen 0003, Kaizhu Huang
NeurIPS3
2024 Discriminative Feature Enhancement Network for few-shot classification and beyond
Fangyu Wu 0001, Qiufeng Wang 0001, Qi Chen 0026, Eng Gee Lim
Expert Syst. Appl.2
2024 Inter-feature Relationship Certifies Robust Generalization of Adversarial Training
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Bin Gu 0001, Huan Xiong, Xinping Yi
Int. J. Comput. Vis.4
2024 Perturbation diversity certificates robust generalization
Zhuang Qian, Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Xinping Yi, Bin Gu 0001, Huan Xiong
Neural Networks4
2024 Prototype Guided Pseudo Labeling and Perturbation-based Active Learning for domain adaptive semantic segmentation
Junkun Peng, Mingjie Sun, Eng Gee Lim, Qiufeng Wang 0001, Jimin Xiao
Pattern Recognit.4
2024 SaliencyCut: Augmenting plausible anomalies for anomaly detection
Jianan Ye, Yijie Hu, Xi Yang 0008, Qiufeng Wang 0001, Kaizhu Huang
Pattern Recognit.4
2024 Scene Text Recognition via Dual-path Network with Shape-driven Attention Alignment
abstract
Scene text recognition (STR), one typical sequence-to-sequence problem, has drawn much attention recently in multimedia applications. To guarantee good performance, it is essential for STR to obtain aligned character-wise features from the whole-image feature maps. While most present works adopt fully data-driven attention-based alignment, such practice ignores specific character geometric information. In this article, built upon a group of learnable geometric points, we propose a novel shape-driven attention alignment method that is able to obtain character-wise features. Concretely, we first design a corner detector to generate a shape map to guide the attention alignments explicitly, where a series of points can be learned to represent character-wise features flexibly. We then propose a dual-path network with a mutual learning and cooperating strategy that successfully combines CNN with a ViT-based model, leading to further accuracy improvement. We conduct extensive experiments to evaluate the proposed method on various scene text benchmarks, including six popular regular and irregular datasets, two more challenging datasets (i.e., WordArt and OST), and three Chinese datasets. Experimental results indicate that our method can achieve superior performance with a comparable model size against many state-of-the-art models.
Yijie Hu, Bin Dong 0003, Kaizhu Huang, Lei Ding 0012, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Correction: ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim
Vis. Comput.2
2024 ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim
Vis. Comput.2
2023 Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation
abstract
Single-source domain generalization (SDG) in medical image segmentation is a challenging yet essential task as domain shifts are quite common among clinical image datasets. Previous attempts most conduct global-only/random augmentation. Their augmented samples are usually insufficient in diversity and informativeness, thus failing to cover the possible target domain distribution. In this paper, we rethink the data augmentation strategy for SDG in medical image segmentation. Motivated by the class-level representation invariance and style mutability of medical images, we hypothesize that unseen target data can be sampled from a linear combination of C (the class number) random variables, where each variable follows a location-scale distribution at the class level. Accordingly, data augmented can be readily made by sampling the random variables through a general form. On the empirical front, we implement such strategy with constrained Bezier transformation on both global and local (i.e. class-level) regions, which can largely increase the augmentation diversity. A Saliency-balancing Fusion mechanism is further proposed to enrich the informativeness by engaging the gradient information, guiding augmentation with proper orientation and magnitude. As an important contribution, we prove theoretically that our proposed augmentation can lead to an upper bound of the generalization risk on the unseen target domain, thus confirming our hypothesis. Combining the two strategies, our Saliency-balancing Location-scale Augmentation (SLAug) exceeds the state-of-the-art works by a large margin in two challenging SDG tasks. Code is available at https://github.com/Kaiseem/SLAug.
Zixian Su, Xi Yang 0008, Kaizhu Huang, Qiufeng Wang 0001, Jie Sun 0024
AAAI5
2023 Decoupled Learning for Long-Tailed Oracle Character Recognition
Jing Li 0049, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Rui Zhang 0012, Kaizhu Huang
ICDAR (4)3
2023 Context Does Matter: End-to-end Panoptic Narrative Grounding with Deformable Attention Refined Matching Network
abstract
Panoramic Narrative Grounding (PNG) is an emerging visual grounding task that aims to segment visual objects in images based on dense narrative captions. The current state-of-the-art methods first refine the representation of phrase by aggregating the most similar k image pixels, and then match the refined text representations with the pixels of the image feature map to generate segmentation results. However, simply aggregating sampled image features ignores the contextual information, which can lead to phrase-to-pixel mis-match. In this paper, we propose a novel learning framework called Deformable Attention Refined Matching Network (DRMN), whose main idea is to bring deformable attention in the iterative process of feature learning to incorporate essential context information of different scales of pixels. DRMN iteratively re-encodes pixels with the deformable attention network after updating the feature representation of the top-k most similar pixels. As such, DRMN can lead to accurate yet discriminative pixel representations, purify the top-k most similar pixels, and consequently alleviate the phrase-to-pixel mis-match substantially. Experimental results show that our novel design significantly improves the matching results between text phrases and image pixels. Concretely, DRMN achieves new state-of-the-art performance on the PNG benchmark with an average recall improvement 3.5%. The codes are available in: https://github.com/JaMesLiMers/DRMN.
Xiao-Bo Jin, Qiufeng Wang 0001, Kaizhu Huang
ICDM3
2023 Improving Handwritten Mathematical Expression Recognition via an Attention Refinement Network
Qiufeng Wang 0001, Jianghan Chen, Kaizhu Huang
ICONIP (13)2
2023 Diff-Writer: A Diffusion Model-Based Stylized Online Handwritten Chinese Character Generator
Minsi Ren, Yan-Ming Zhang 0001, Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICONIP (10)3
2023 Progressive Supervision for Tampering Localization in Document Images
Huiru Shao, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001
ICONIP (15)5
2023 A Symbolic Characters Aware Model for Solving Geometry Problems
abstract
AI has made significant progress in solving math problems, but geometry problems remain challenging due to their reliance on both text and diagrams. In the text description, symbolic characters such as "ABC" often serve as a bridge to connect the corresponding diagram. However, by simply tokenizing symbolic characters into individual letters (e.g., 'A', 'B' and 'C'), existing works fail to study them explicitly and thus lose the semantic relationship with the diagram. In this paper, we develop a symbolic character-aware model to fully explore the role of these characters in both text and diagram understanding and optimize the model under a multi-modal reasoning framework. In the text encoder, we propose merging individual symbolic characters to form one semantic unit along with geometric information from the corresponding diagram. For the diagram encoder, we pre-train it under a multi-label classification framework with the symbolic characters as labels. In addition, we enhance the geometry diagram understanding ability via a self-supervised learning method under the masked image modeling auxiliary task. By integrating the proposed model into a general encoder-decoder pipeline for solving geometry problems, we demonstrate its superiority on two benchmark datasets, including GeoQA and Geometry3K, with extensive experiments. Specifically, on GeoQA, the question-solving accuracy is increased from 60.0% to 64.1%, achieving a new state-of-the-art accuracy; on Geometry3K, we reduce the question average solving steps from 6.9 down to 6.0 with marginally higher solving accuracy.
Maizhen Ning, Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001
ACM Multimedia2
2023 Retrieval-based language model adaptation for handwritten Chinese text recognition
Shuying Hu, Qiufeng Wang 0001, Kaizhu Huang, Frans Coenen
Int. J. Document Anal. Recognit.2
2023 Towards better long-tailed oracle character recognition with adversarial data augmentation
abstract
Deciphering oracle bone script is of great significance to the study of ancient Chinese culture as well as archaeology. Although recent studies on oracle character recognition have made substantial progress, they still suffer from the long-tailed data situation that results in a noticeable performance drop on the tail classes. To mitigate this issue, we propose a generative adversarial framework to augment oracle characters in the problematic classes. In this framework, the generator produces synthetic data through convex combinations of all the available samples in the corresponding classes, and is further optimized through adversarial learning with the classifier and simultaneously the discriminator . Meanwhile, we introduce Repatch to generalize samples in the generator. Since tail classes do not have sufficient data for convex combinations , we propose the TailMix mechanism to generate suitable tail class samples from other classes. Experimental results show that our proposed algorithm obtains remarkable performance in oracle character recognition and achieves new state-of-the-art average (total) accuracy with 86.03% (89.46%), 86.54% (93.86%), 95.22% (96.17%) on the three datasets Oracle-AYNU, OBC306 and Oracle-20K, respectively.
Jing Li 0049, Qiufeng Wang 0001, Kaizhu Huang, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas
Pattern Recognit.2
2023 Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang
Pattern Recognit.4
2023 Mind the Gap: Alleviating Local Imbalance for Unsupervised Cross-Modality Medical Image Segmentation
abstract
Unsupervised cross-modality medical image adaptation aims to alleviate the severe domain gap between different imaging modalities without using the target domain label. A key in this campaign relies upon aligning the distributions of source and target domain. One common attempt is to enforce the global alignment between two domains, which, however, ignores the fatal local-imbalance domain gap problem, i.e., some local features with larger domain gap are harder to transfer. Recently, some methods conduct alignment focusing on local regions to improve the efficiency of model learning. While this operation may cause a deficiency of critical information from contexts. To tackle this limitation, we propose a novel strategy to alleviate the domain gap imbalance considering the characteristics of medical images, namely Global-Local Union Alignment. Specifically, a feature-disentanglement style-transfer module first synthesizes the target-like source images to reduce the global domain gap. Then, a local feature mask is integrated to reduce the 'inter-gap' for local features by prioritizing those discriminative features with larger domain gap. This combination of global and local alignment can precisely localize the crucial regions in segmentation target while preserving the overall semantic consistency. We conduct a series of experiments with two cross-modality adaptation tasks, i,e. cardiac substructure and abdominal multi-organ segmentation. Experimental results indicate that our method achieves state-of-the-art performance in both tasks.
Zixian Su, Xi Yang 0008, Qiufeng Wang 0001, Yuyao Yan, Jie Sun 0024, Kaizhu Huang
IEEE J. Biomed. Health Informatics4
2022 Towards Accurate Alignment and Sufficient Context in Scene Text Recognition
Yijie Hu, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Xiao-Bo Jin, Kaizhu Huang
ICONIP (3)3
2022 Rethinking Image Inpainting with Attention Feature Fusion
Shuyi Qu, Kaizhu Huang, Qiufeng Wang 0001, Bin Dong 0003
ICONIP (3)3
2022 Certifying Better Robust Generalization for Unsupervised Domain Adaptation
abstract
Recent studies explore how to obtain adversarial robustness for unsupervised domain adaptation (UDA). These efforts are however dedicated to achieving an optimal trade-off between accuracy and robustness on a given or seen target domain but ignore the robust generalization issue over unseen adversarial data. Consequently, degraded performance will be often observed when existing robust UDAs are applied to future adversarial data. In this work, we make a first attempt to address the robust generalization issue of UDA. We conjecture that the poor robust generalization of present robust UDAs may be caused by the large distribution gap among adversarial examples. We then provide an empirical and theoretical analysis showing that this large distribution gap is mainly owing to the discrepancy between feature-shift distributions. To reduce such discrepancy, a novel Anchored Feature-Shift Regularization (AFSR) method is designed with a certificated robust generalization bound. We conduct a series of experiments on benchmark UDA datasets. Experimental results validate the effectiveness of our proposed AFSR over many existing robust UDA methods.
Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Rui Zhang 0012, Chaoliang Zhong
ACM Multimedia4
2022 Sparse matrix factorization with L2, 1 norm for matrix completion
Xiao-Bo Jin, Jianyu Miao, Qiufeng Wang 0001, Guanggang Geng, Kaizhu Huang
Pattern Recognit.3
2022 A survey of robust adversarial training in pattern recognition: Fundamental, theory, and methodologies
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Xu-Yao Zhang
Pattern Recognit.3
2022 Unsupervised domain adaptation in homogeneous distance space for person re-identification
Dingyuan Zheng, Jimin Xiao, Yunchao Wei, Qiufeng Wang 0001, Kaizhu Huang, Yao Zhao 0001
Pattern Recognit.4
2022 Exploiting Attention-Consistency Loss For Spatial-Temporal Stream Action Recognition
abstract
Currently, many action recognition methods mostly consider the information from spatial streams. We propose a new perspective inspired by the human visual system to combine both spatial and temporal streams to measure their attention consistency. Specifically, a branch-independent convolutional neural network (CNN) based algorithm is developed with a novel attention-consistency loss metric, enabling the temporal stream to concentrate on consistent discriminative regions with the spatial stream in the same period. The consistency loss is further combined with the cross-entropy loss to enhance the visual attention consistency. We evaluate the proposed method for action recognition on two benchmark datasets: Kinetics400 and UCF101. Despite its apparent simplicity, our proposed framework with the attention consistency achieves better performance than most of the two-stream networks, i.e., 75.7% top-1 accuracy on Kinetics400 and 95.7% on UCF101, while reducing 7.1% computational cost compared with our baseline. Particularly, our proposed method can attain remarkable improvements on complex action classes, showing that our proposed network can act as a potential benchmark to handle complicated scenarios in industry 4.0 applications.
Xiao-Bo Jin, Qiufeng Wang 0001, Amir Hussain 0001, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Gradient Distribution Alignment Certificates Better Adversarial Domain Adaptation
abstract
The latest heuristic for handling the domain shift in un-supervised domain adaptation tasks is to reduce the data distribution discrepancy using adversarial learning. Recent studies improve the conventional adversarial domain adaptation methods with discriminative information by integrating the classifier’s outputs into distribution divergence measurement. However, they still suffer from the equilibrium problem of adversarial learning in which even if the discriminator is fully confused, sufficient similarity between two distributions cannot be guaranteed. To overcome this problem, we propose a novel approach named feature gradient distribution alignment (FGDA)1. We demonstrate the rationale of our method both theoretically and empirically. In particular, we show that the distribution discrepancy can be reduced by constraining feature gradients of two domains to have similar distributions. Meanwhile, our method enjoys a theoretical guarantee that a tighter error upper bound for target samples can be obtained than that of conventional adversarial domain adaptation methods. By integrating the proposed method with existing adversarial domain adaptation models, we achieve state-of-the-art performance on two real-world benchmark datasets.
Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Chaoliang Zhong
ICCV4
2021 Mix-Up Augmentation for Oracle Character Recognition with Imbalanced Data Distribution
Jing Li 0049, Qiufeng Wang 0001, Rui Zhang 0012, Kaizhu Huang
ICDAR (1)2
2021 Towards Better Robust Generalization with Shift Consistency Regularization
abstract
While adversarial training becomes one of the most promising defending approaches against adversarial attacks for deep neural networks, the conventional wisdom through robust optimization may usually not guarantee good generalization for robustness. Concerning with robust generalization over unseen adversarial data, this paper investigates adversarial training from a novel perspective of shift consistency in latent space. We argue that the poor robust generalization of adversarial training is owing to the significantly dispersed latent representations generated by training and test adversarial data, as the adversarial perturbations push the latent features of natural examples in the same class towards diverse directions. This is underpinned by the theoretical analysis of the robust generalization gap, which is upper-bounded by the standard one over the natural data and a term of feature inconsistent shift caused by adversarial perturbation {–} a measure of latent dispersion. Towards better robust generalization, we propose a new regularization method {–} shift consistency regularization (SCR) {–} to steer the same-class latent features of both natural and adversarial data into a common direction during adversarial training. The effectiveness of SCR in adversarial training is evaluated through extensive experiments over different datasets, such as CIFAR-10, CIFAR-100, and SVHN, against several competitive methods.
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Rui Zhang 0012, Xinping Yi
ICML4
2021 A Segment-Based Layout Aware Model for Information Extraction on Document Images
Maizhen Ning, Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001
ICONIP (5)2
2021 Residual attention-based multi-scale script identification in scene text images
Mengkai Ma, Qiufeng Wang 0001, Shen Huang, John Yannis Goulermas, Kaizhu Huang
Neurocomputing2
2020 Weakly Supervised Learning for Over-Segmentation Based Handwritten Chinese Text Recognition
abstract
In this paper, we proposed a weakly supervised learning method for string-level training of character classifier in over-segmentation based handwritten Chinese text recognition (HCTR). The over-segmentation based framework can easily integrate multiple context models and provide accurate character boundary and recognition confidence, but has not been implemented with string-level training for HCTR. We propose to optimize the character classifier by minimizing the marginal log-likelihood on a string-level annotated handwriting dataset, where the forward-backward algorithm is utilized in a segmentation-and-recognition lattice. Experimental results on the CASIA-HWDB and ICDAR-2013 competition datasets show that the proposed method improves the recognition performance significantly, which demonstrates its effectiveness.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICFHR2
2020 Adversarial Rectification Network for Scene Text Regularization
Jing Li 0049, Qiufeng Wang 0001, Rui Zhang 0012, Kaizhu Huang
ICONIP (2)2
2020 Multi-scale Attention Consistency for Multi-label Image Classification
Xiao-Bo Jin, Qiufeng Wang 0001, Kaizhu Huang
ICONIP (4)3
2020 Improving deep neural network performance by integrating kernelized Min-Max objective
Qiufeng Wang 0001, Rui Zhang 0012, Amir Hussain 0001, Kaizhu Huang
Neurocomputing1
2020 Generative adversarial classifier for handwriting characters super-resolution
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Jimin Xiao, Rui Zhang 0012
Pattern Recognit.3
2019 An Interactive and Generative Approach for Chinese Shanshui Painting Document
abstract
Chinese Shanshui is a landscape painting document mainly drawing mountain and water, which is popular in Chinese culture. However, it is very challenging to create this by general people. In this paper, we propose an interactive and generative approach to automatically generate the Chinese Shanshui painting documents based on users' input, where the users only need to sketch simple lines to represent their ideal landscape without any professional Shanshui painting skills. This sketch-to-Shanshui translation is optimized by the model of cycle Generative Adversarial Networks (GAN). To evaluate the proposed approach, we collected a large set of both sketch data and Chinese Shanshui painting data to train the model of cycle-GAN, and developed an interactive system called Shanshui-DaDA (i.e., Design and Draw with AI) to generate Chinese Shanshui painting documents in real-time. The experimental results show that this system can generate satisfied Chinese Shanshui painting documents by general users.
Aven-Le Zhou, Qiufeng Wang 0001, Kaizhu Huang, Cheng-Hung Lo
ICDAR2
2014 An over-segmentation method for single-touching Chinese handwriting with learning-based filtering
Qiufeng Wang 0001, Cheng-Lin Liu 0001
Int. J. Document Anal. Recognit.3
2014 Unsupervised language model adaptation for handwritten Chinese text recognition
Qiufeng Wang 0001, Cheng-Lin Liu 0001
Pattern Recognit.1
2013 ICDAR 2013 Chinese Handwriting Recognition Competition
abstract
This paper describes the Chinese handwriting recognition competition held at the 12th International Conference on Document Analysis and Recognition (ICDAR 2013). This third competition in the series again used the CASIA-HWDB/OLHWDB databases as the training set, and all the submitted systems were evaluated on closed datasets to report character-level correct rates. This year, 10 groups submitted 27 systems for five tasks: classification on extracted features, online/offline isolated character recognition, online/offline handwritten text recognition. The best results (correct rates) are 93.89% for classification on extracted features, 94.77% for offline character recognition, 97.39% for online character recognition, 88.76% for offline text recognition, and 95.03% for online text recognition, respectively. In addition to the test results, we also provide short descriptions of the recognition methods and brief discussions on the results.
Qiufeng Wang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICDAR2
2013 Style Consistent Perturbation for Handwritten Chinese Character Recognition
abstract
Perturbation-based recognition is effective to recover the deformation of handwritten characters and improve the recognition performance by generating multiple distortions and selecting a distortion that best restores character deformation. Considering that the characters in a field undergo similar deformation under a consistent style, we proposed style consistent perturbation for handwritten character recognition. By generating multiple distortions for the characters in a field, each distortion style is evaluated at the field level and the uniform distortion style of maximum recognition confidence is selected to give the final result. To overcome the slight deviation from uniform style, we also propose to search the neighborhood distortions from the optimal uniform distortion for higher confidence. The experiments of handwritten Chinese character recognition on multi-writer data show that style consistent perturbation in very short fields outperforms individual character recognition, and neighborhood distortion search yields further improvement.
Ming-Ke Zhou, Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICDAR3
2013 Online and offline handwritten Chinese character recognition: Benchmarking on new databases
Cheng-Lin Liu 0001, Dahan Wang, Qiufeng Wang 0001
Pattern Recognit.4
2013 Transcript mapping for handwritten Chinese documents by integrating character recognition model and geometric context
Qiufeng Wang 0001, Cheng-Lin Liu 0001
Pattern Recognit.2
2012 Improving Handwritten Chinese Text Recognition by Unsupervised Language Model Adaptation
abstract
This paper investigates the effects of unsupervised language model adaptation (LMA) in handwritten Chinese text recognition. For no prior information of recognition text is available, we use a two-pass recognition strategy. In the first pass, the generic language model (LM) is used to get a preliminary result, which is used to choose the best matched LMs from a set of pre-defined domains, then the matched LMs are used in the second pass recognition. Each LM is compressed to a moderate size via the entropy-based pruning, tree-structure formatting and fewer-byte quantization. We evaluated the LMA for five LM types, including both character-level and word-level ones. Experiments on the CASIA-HWDB database show that language model adaptation improves the performance for each LM type in all domains. The documents of ancient domain gained the biggest improvement of character-level correct rate of 5.87 percent up and accurate rate of 6.05 percent up.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
Document Analysis Systems1
2012 A Touching Character Database from Chinese Handwriting for Assessing Segmentation Algorithms
abstract
For assessing touching character segmentation algorithms, we present a database of touching characters collected from the Chinese handwriting database CASIA-HWDB, called CASIA-HWDB-T. It includes 56,469 two-character or multiple-character touching strings, among which 1,818 strings have multiple-touching characters. We also partition the touching strings into 50,157 all-Chinese strings, 2,788 all-digit ones, 328 all-letter ones, and 3,196 mixed-character ones. All the strings are annotated with the character classes, locations of touching points, and auxiliary values like string height and average stroke width. And last, we measure the segmentation performance of three existing algorithms on this database for reference.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICFHR3
2012 Off-Line Handwritten Arabic Word Recognition Using SVMs with Normalized Poly Kernel
Abdulrahman Alalshekmubarak, Amir Hussain 0001, Qiufeng Wang 0001
ICONIP (2)3
2012 Towards IMACA: Intelligent Multimodal Affective Conversational Agent
Amir Hussain 0001, Erik Cambria, Thomas Mazzocco, Marco Grassi, Qiufeng Wang 0001, Tariq S. Durrani
ICONIP (1)5
2012 Handwritten Chinese Text Recognition by Integrating Multiple Contexts
abstract
This paper presents an effective approach for the offline recognition of unconstrained handwritten Chinese texts. Under the general integrated segmentation-and-recognition framework with character oversegmentation, we investigate three important issues: candidate path evaluation, path search, and parameter estimation. For path evaluation, we combine multiple contexts (character recognition scores, geometric and linguistic contexts) from the Bayesian decision view, and convert the classifier outputs to posterior probabilities via confidence transformation. In path search, we use a refined beam search algorithm to improve the search efficiency and, meanwhile, use a candidate character augmentation strategy to improve the recognition accuracy. The combining weights of the path evaluation function are optimized by supervised learning using a Maximum Character Accuracy criterion. We evaluated the recognition performance on a Chinese handwriting database CASIA-HWDB, which contains nearly four million character samples of 7,356 classes and 5,091 pages of unconstrained handwritten texts. The experimental results show that confidence transformation and combining multiple contexts improve the text line recognition performance significantly. On a test set of 1,015 handwritten pages, the proposed approach achieved character-level accurate rate of 90.75 percent and correct rate of 91.39 percent, which are superior by far to the best results reported in the literature.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 CASIA Online and Offline Chinese Handwriting Databases
abstract
This paper introduces a pair of online and offline Chinese handwriting databases, containing samples of isolated characters and handwritten texts. The samples were produced by 1,020 writers using Anoto pen on papers for obtaining both online trajectory data and offline images. Both the online samples and offline samples are divided into six datasets, three for isolated characters (DB1.0-C1.2) and three for handwritten texts (DB2.0-C2.2). The (either online or offline) datasets of isolated characters contain about 3.9 million samples of 7,356 classes (7,185 Chinese characters and 171 symbols), and the datasets of handwritten texts contain about 5,090 pages and 1.35 million character samples. Each dataset is segmented and annotated at character level, and is partitioned into standard training and test subsets. The online and offline databases can be used for the research of various handwritten document analysis tasks.
Cheng-Lin Liu 0001, Dahan Wang, Qiufeng Wang 0001
ICDAR4
2011 ICDAR 2011 Chinese Handwriting Recognition Competition
abstract
In the Chinese handwriting recognition competition organized with the ICDAR 2011, four tasks were evaluated: offline and online isolated character recognition, offline and online handwritten text recognition. To enable the training of recognition systems, we announced the large databases CASIA-HWDB/OLHWDB. The submitted systems were evaluated on un-open datasets to report character-level correct rates. In total, we received 25 systems submitted by eight groups. On the test datasets, the best results (correct rates) are 92.18% for offline character recognition, 95.77% for online character recognition, 77.26% for offline text recognition, and 94.33% for online text recognition, respectively. In addition to the evaluation results, we provide short descriptions of the recognition methods and have brief discussions.
Cheng-Lin Liu 0001, Qiufeng Wang 0001, Dahan Wang
ICDAR3
2011 Improving Handwritten Chinese Text Recognition by Confidence Transformation
abstract
This paper investigates the effects of confidence transformation (CT) of the character classifier outputs in handwritten Chinese text recognition. The classifier outputs are transformed to confidence values in three confidence types, namely, sigmoid, soft max and Dempster-Shafer theory of evidence (D-S evidence). The confidence parameters are optimized by minimizing the cross-entropy (CE) loss function (both binary and multi-class) on a validation dataset, where we add non-character samples to enhance the outlier rejection capability in text recognition. Experimental results on the CASIA-HWDB database show that confidence transformation improves the handwritten text recognition performance significantly and adding non-characters for confidence parameter estimation is beneficial. Among the confidence types, the D-S evidence performs best.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICDAR1
2011 Touching Character Separation in Chinese Handwriting Using Visibility-Based Foreground Analysis
abstract
In offline handwritten text recognition, the separation of touching characters remains a challenge due to the variability of touching structures. This paper proposes a new touching character separation method for Chinese handwriting based on skeleton analysis and contour analysis incorporating the visibility of separating points. Separating points are detected from strokes that are common in both upper and lower skeleton tracing, and the profile visibility of strokes and separating points is analyzed to adjust and verify separating points. Our experiments on two large handwriting databases demonstrate the effectiveness of the proposed method.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICDAR3
2011 Transcript Mapping for Handwritten Text Lines Using Conditional Random Fields
abstract
This paper presents a conditional random field (CRF) model for aligning online handwritten Chinese/Japanese text lines (character strings) with the corresponding transcripts. The CRF model is defined on a lattice which contains all possible segmentation hypotheses. The feature functions characterize the shape and context dependences of characters, including the scores of character recognition and the geometric compatibilities between characters. The combining parameters are optimized by energy minimization. Experimental results on two online databases: CASIA-OLHWDB and TUAT Kondate demonstrate the effectiveness of the proposed method.
Dahan Wang, Qiufeng Wang 0001, Masaki Nakagawa, Cheng-Lin Liu 0001
ICDAR4
2010 Integrating Geometric Context for Text Alignment of Handwritten Chinese Documents
abstract
The alignment of text line images with text transcript is a crucial step of handwritten document annotation. Handwritten text alignment is prone to errors due to the difficulty of character segmentation and the variability of character shape, size and position. In this paper, we propose to incorporate the geometric context of character strings to improve the alignment accuracy for offline handwritten Chinese documents. We use four statistical models to evaluate the geometric features of single characters and between-character relationships. By combining the geometric models with a character recognizer, we have achieved a large improvement of alignment accuracy in our experiments on unconstrained handwritten Chinese text lines.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICFHR2
2009 Integrating Language Model in Handwritten Chinese Text Recognition
abstract
This paper describes a system for handwritten Chinese text recognition integrating language model. On a text line image, the system generates character segmentation and word segmentation candidates, and the candidate paths are evaluated by character recognition scores and language model. The optimal path, giving segmentation and recognition result, is found using a pruned dynamic programming search method. We evaluate various language models, including the character-based n-gram, word-based n-gram, and hybrid n-gram models. Experimental results on the HIT-HW database show that the language models improve the recognition performance remarkably.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICDAR1
2009 A Tool for Ground-Truthing Text Lines and Characters in Off-Line Handwritten Chinese Documents
abstract
Annotating the regions, text lines and characters of document images is an important, but tedious and expensive task. A ground-truthing tool may largely alleviate the human burden in this process. This paper describes an automated recognition-based tool GTLC for finding the best alignment between the text transcript and the connected components of unconstrained handwritten document image. The alignment process is formulated as an optimization problem involving candidate character segmentation and recognition. We have validated the effectiveness of this tool and have used it for annotating a large number of handwritten Chinese documents.
Qiufeng Wang 0001, Cheng-Lin Liu 0001
ICDAR2