Dahan Wang

dblp:69/7055 · also Da-Han Wang · DBLP profile ↗
← Back
105ranked-venue papers
9as first author
76since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 8 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 3 first-author · 40 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021
YearPublicationVenuePosition
2026 Fair Facial Attribute Recognition via Group-Decoupled Vision Transformer with Mask-Guided Correlation Suppression
abstract
Facial Attribute Recognition (FAR) holds significant potential for wide-ranging applications. However, traditionally trained FAR models exhibit unfairness, largely due to data bias—where certain sensitive attributes correlate statistically with target attributes. To address this, we propose a group-attention mechanism: first, each image is categorized into subgroups (e.g., Male/Female&short hair, Male/Female&long hair). Within the attention mechanism, distinct Query parameters are used for each group, with shared Key and Value parameters. As group-specific Query parameters are trained on subgrouped data, the noted bias is effectively mitigated. Consequently, integrating this Group-Attention into Vision Transformer (ViT) yields our novel Group-Decoupled ViT (GD-ViT) model. Moreover, to further attenuate the statistical correlation between sensitive and target attributes, we propose a Mask-Guided Correlation Suppression learning strategy. Specifically, in Stage 1, it first leverages a min-max dual-loss optimization strategy to train GD-ViT in capturing key regions related to sensitive attributes yet irrelevant to target attributes. Then, in Stage 2, it trains another GD-ViT by masking sensitive regions identified in Stage 1, fusing the masked output (as intermediate input) with the model’s intermediate outputs. This weakens regions associated with sensitive attributes while enhancing others, suppressing the learning of key features related to sensitive attributes. Consequently, it encourages the model to focus more on intrinsic target attribute regions and balances the learning process between the sensitive attribute and the target attribute. Extensive experiments demonstrate that our method achieves superior performance across three benchmark datasets for fair facial attribute recognition.
Huichang Huang, Kunchi Li, Si Chen 0002, Dahan Wang
AAAI4
2026 Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
abstract
Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance.
Dahan Wang, Changbao Zhu, Kai Chen 0029
AAAI1
2026 DAPE: Dynamic Non-uniform Alignment and Progressive Detail Enhancement Techniques for Improving the Performance of Efficient Visual Language Models
Mengyuan Tian, Qiyan Zhao, Yannan He, Dahan Wang
ICIC (13)4
2026 HDFNet:Hybrid-domain fusion network for medical image restoration
Liqun Lin, Shunzhou Wang, Si Chen 0002, Chao Zeng 0005, Nanfeng Jiang, Dahan Wang
Expert Syst. Appl.7
2026 Diverse feature generation for zero-shot Chinese character recognition
Song-Liang Pan, Kunchi Li, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
Expert Syst. Appl.3
2026 MoKA-HP: Motion-aware KAdaptation with historical prompts for efficient and robust RGB-T tracking
Zhixi Wu, Si Chen 0002, Dahan Wang, Shunzhi Zhu
Neurocomputing3
2026 Learning relationship-guided vision-language transformer for facial attribute recognition
Si Chen 0002, Mingxuan Lei, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu
Pattern Recognit.3
2026 A comprehensive survey of oracle character recognition: Challenges, datasets, methodology, and beyond
Xueke Chi, Qiufeng Wang 0001, Kaizhu Huang, Dahan Wang, Yongge Liu
Pattern Recognit.5
2026 AFH-Net: An adaptive feature harmonization network for document image De-warping
Xinyue Zhou, Nanfeng Jiang, Wang Man, Xu-Yao Zhang, Shunzhou Wang, Dahan Wang
Pattern Recognit.7
2026 CPFormer-Net: Correspondence Pruning Transformer With Structured Context Aggregation
abstract
Finding reliable correspondences between two-view images remains challenging, particularly under high outlier ratios. While existing transformer-based methods excel in capturing global context, they often fail to maintain robust long-range semantic dependencies across input correspondences. To address this, we propose a novel correspondence pruning transformer network, called CPFormer-Net, which enhances long-range dependency modeling via structured context aggregation for accurate correspondence pruning. Specifically, we propose the CPFormer module, which integrates global context with local geometric cues to ensure both structural coherence and semantic consistency. Then, we propose the Semantic Graph Network (SEGN) module with a dual-branch architecture to improve representation. One branch applies Structured Context Aggregation (SCA) block with an agent attention mechanism to explicitly model long-range dependencies, enabling it to filter channels, recalibrate features, and guide attention, while the other branch employs pooling and normalization to preserve spatial relations. Extensive experiments on indoor and outdoor benchmarks demonstrate that CPFormer-Net consistently outperforms state-of-the-art methods in inlier identification and correspondence matching.
Hanlin Guo, Weijia Lv, Zhi Shen, Dahan Wang
IEEE Signal Process. Lett.4
2026 CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement
Leyan Yang, Dahan Wang, Xiaobin Rong, Jiadong Zhao
IEEE Signal Process. Lett.2
2026 ProMoT: Progressive Prompting of Modality and Temporal Dynamics for RGB-T Tracking
abstract
RGB-T tracking benefits from the complementary nature of RGB and TIR modalities, yet their relative reliability for target localization often shifts over time. Most existing trackers fail to adapt to such modality and temporal dynamics in a unified and effective manner, resulting in target representations that are neither discriminative nor temporally consistent. In this paper, we propose ProMoT, a novel tracking framework that jointly integrates cross-modal and temporal cues into a progressive prompting process, enabling continuous retrieval of target-aware representations. Specifically, we design an adaptive target query generator (QueryGen), which selectively aggregates informative spatio-temporal cues from diverse ghost representations through the dynamic sparse ghost fusion mechanism, thereby enabling the generation of target-aware queries. To further preserve fine-grained, temporally consistent target cues, we introduce a high-order contextual prompt updater (PromptUpdater), which encodes high-order cross-modal representations from current and previous frames. These prompts establish the compact and discriminative inter-frame context to not only refine the current frame’s features but also guide target localization in future frames. All components are built upon a parameter-shared backbone for RGB and TIR inputs, forming our complete ProMoT framework. Extensive experiments on both complete and missing modality RGB-T tracking benchmarks show that ProMoT consistently achieves state-of-the-art performance while balancing efficiency.
Rui Xu 0028, Si Chen 0002, Yuzhen Niu, Yan Yan 0001, Dahan Wang
IEEE Trans. Circuits Syst. Video Technol.6
2026 HOH-Net: High-Order Hierarchical Middle-Feature Learning Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match images of the same person across visible (VIS) and infrared (IR) modalities. Existing VI-ReID methods ignore high-order structure information of features and struggle to learn a reliable common feature space due to the modality discrepancy between VIS and IR images. To alleviate the above issues, we propose a novel high-order hierarchical middle-feature learning network (HOH-Net) for VI-ReID. We introduce a high-order structure learning (HSL) module to explore the high-order relationships of short- and long-range feature nodes, for significantly mitigating model collapse and effectively obtaining discriminative features. We further develop a fine-coarse graph attention alignment (FCGA) module, which efficiently aligns multi-modality feature nodes from node-level and region-level perspectives, ensuring reliable middle-feature representations. Moreover, we exploit a hierarchical middle-feature agent learning (HMAL) loss to hierarchically reduce the modality discrepancy at each stage of the network by using the agents of middle features. The proposed HMAL loss also exchanges detailed and semantic information between low- and high-stage networks. Finally, we introduce a modality-range identity-center contrastive (MRIC) loss to minimize the distances between VIS, IR, and middle features. Extensive experiments demonstrate that the proposed HOH-Net yields state-of-the-art performance on the image-based and video-based VI-ReID datasets. The code is available at: https://github.com/Jaulaucoeng/HOS-Net.
Liuxiang Qiu, Si Chen 0002, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 CAMD: Context-Aware Masked Distillation for General Self-Supervised Facial Representation Pre-Training
abstract
Self-supervised pre-training has been shown to effectively learn transferable representations from unlabeled images in many visual tasks. However, existing self-supervised pre-training methods lack sufficient context-awareness and are difficult to obtain fine-grained facial representations, thus resulting in the weak generalization ability of the model to deal with various facial analysis tasks. To address this issue, we propose a Context-Aware Masked Distillation method, termed CAMD, to effectively learn general facial representations for fine-grained facial analysis tasks. The CAMD method first designs an innovative local-to-global masked image modeling framework to learn the contextual spatial structures and semantic relationships between local and global features, enabling effective self-supervised pre-training. In this framework, our pre-training task predicts the dense global feature representations based on the visible local feature representations after masking, so as to achieve semantic alignment across local and global views and significantly enhance spatial sensitivity. Moreover, the CAMD method leverages an attention-driven cross-view hierarchical distillation module to fully distill the features of related regions between different encoder layers of the online and target encoders. This module can learn contextual dependencies and capture discriminative fine-grained facial feature representations. Our method is evaluated on multiple downstream facial analysis tasks, including face alignment, face parsing, facial attribute recognition, facial expression recognition, and head pose estimation, all achieving state-of-the-art results and exhibiting the strong generality and effectiveness. The code is available at: https://github.com/mumumu-wss/CAMD.
Sensen Wang 0001, Si Chen 0002, Dahan Wang, Yang Hua 0001, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 MOKAN: A Multi-Omics Data Analysis Framework Using Kolmogorov-Arnold Networks
abstract
With the advancement of high-throughput sequencing technology, large amounts of multi-omics data have been accumulated. Due to the comprehensive representation of different molecular layers, multi-omics data enable a deep understanding of cancer analysis, including subtype classification, biomarker identification and so on. However, existing multi-omics integration methods either rely on complex feature engineering or fail to capture the heterogeneity of multi-source data. To integrate multi-omics data in a complementary and collaborative manner, we propose a Kolmogorov-Arnold Networks-based multi-omics integration analysis framework, termed MOKAN. In detail, MOKAN first employs a sample-weighted random sampler to reduce the difference in the number of samples involved in training. Then, to perform this data analysis, our proposed model, MOKAN, leverages KAN's architecture, utilizing its learnable edge-weight activation functions and flexible network structure. This design makes MOKAN particularly well-suited for handling multi-omics data by efficiently capturing diverse feature spaces across different omics layers. It integrates heterogeneous data types while preserving the unique contributions of each omics for comprehensive and accurate representation. Moreover, MOKAN employs KAN's decomposition property to break down high-dimensional complex data into multiple low-dimensional subspaces, each independently capturing a specific aspect of the data. These subspace representations are then systematically integrated through a one-dimensional function that overlays and combines their contributions, effectively reconstructing the global data structure. Experimental results demonstrate that our proposed model outperforms other existing methods in cancer classification tasks.
Jingguang He, Shunxin Xiao, Sujia Huang, Dahan Wang
IEEE J. Biomed. Health Informatics5
2026 Video-Level Cross-Modal Temporal-Navigation for RGBT Tracking
abstract
RGBT tracking has recently garnered significant attention due to its all-weather tracking capability. Traditional RGBT tracking methods primarily concentrate on the fusion of cross-modal spatial information. However, these methods ignore contextual relationships between consecutive video frames and lack effective interactions between modalities, easily resulting in tracking drift gradually due to the accumulation of errors. To avoid this limitation, we propose a novel Video-Level Cross-Modal Temporal-Navigation method termed VCT for robust RGBT tracking, which fully leverages the complementary spatio-temporal information across modalities to improve cross-modal tracking accuracy. The VCT employs a simple, flexible, and effective video-level dual-stream architecture that accommodates video sequences of arbitrary length, enabling RGB and TIR streams to capture and synergize spatio-temporal features across frames. To achieve temporal consistency and adaptability, we design a Cross-Modal Temporal Prompt Navigator (CM-TPN) that dynamically aggregates and compresses the historical frame context to navigate predictions of subsequent frames through temporal prompts. In addition, we introduce a Modality-Specific Mixture of Adapters (MS-MoA) to promote the spatio-temporal interaction both within and between modalities, thereby dramatically adapting to appearance changes. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the four popular RGBT tracking benchmarks.
Si Chen 0002, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001
IEEE Trans. Multim.3
2025 HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
abstract
Haiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu, Da-Han Wang, Xu-Yao Zhang, Cheng-Lin Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Haiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu 0004, Dahan Wang, Xu-Yao Zhang
ACL (1)5
2025 Federated Continual Instruction Tuning
Haiyang Guo, Fanhu Zeng, Fei Zhu 0004, Wenzhuo Liu, Dahan Wang, Jian Xu 0015, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICCV5
2025 OracleGCD: Generalized Category Discovery for Oracle Bone Scripts
Hetao Wu, Kunchi Li, Xu-Yao Zhang, Dahan Wang
ICDAR (3)5
2025 FCD-Net: Frequency and Contrastive Learning-Driven Network for Document Image Shadow Removal
Nanfeng Jiang, Dahan Wang, Yun Wu 0001
ICDAR (2)3
2025 MSFM-UNet: Multi-scan and Frequency Domain Mamba UNet for Medical Image Segmentation
Weitao Lin, Dahan Wang
ICIC (28)5
2025 Multi-hop Aware Graph Convolutional Network and Collaborative Transformer for Traffic Flow Prediction
Dahan Wang
ICIC (7)3
2025 TextMamba: Scene Text Detector with Mamba
abstract
In scene text detection, Transformer-based methods have addressed the global feature extraction limitations inherent in traditional convolution neural network-based methods. However, most directly rely on native Transformer attention layers as encoders without evaluating their cross-domain limitations and inherent shortcomings: forgetting important information or focusing on irrelevant representations when modeling long-range dependencies for text detection. The recently proposed state space model Mamba has demonstrated better long-range dependencies modeling through a linear complexity selection mechanism. Therefore, we propose a novel scene text detector based on Mamba that integrates the selection mechanism with attention layers, enhancing the encoder’s ability to extract relevant information from long sequences. We adopt the Top_k algorithm to explicitly select key information and reduce the interference of irrelevant information in Mamba modeling. Additionally, we design a dual-scale feed-forward network and an embedding pyramid enhancement module to facilitate high-dimensional hidden state interactions and multi-scale feature fusion. Our method achieves state-of-the-art or competitive performance on various benchmarks, with F-measures of 89.7%, 89.2%, and 78.5% on CTW1500, TotalText, and ICDAR19ArT, respectively. Codes will be available.
Qiyan Zhao, Dahan Wang
IJCNN3
2025 TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network
Xiaobin Rong, Dahan Wang, Qinwen Hu
INTERSPEECH2
2025 MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
abstract
Hallucinations pose a significant challenge in Large Vision Language Models (LVLMs), with misalignment between multimodal features identified as a key contributing factor. This paper reveals the negative impact of the long-term decay in Rotary Position Encoding (RoPE), used for positional modeling in LVLMs, on multimodal alignment. Concretely, under long-term decay, instruction tokens exhibit uneven perception of image tokens located at different positions within the two-dimensional space: prioritizing image tokens from the bottom-right region since in the one-dimensional sequence, these tokens are positionally closer to the instruction tokens. This biased perception leads to insufficient image-instruction interaction and suboptimal multimodal alignment. We refer to this phenomenon as ''image alignment bias.'' To enhance instruction's perception of image tokens at different spatial locations, we propose MCA-LLaVA, based on Manhattan distance, which extends the long-term decay to a two-dimensional, multi-directional spatial decay. MCA-LLaVA integrates the one-dimensional sequence order and two-dimensional spatial position of image tokens for positional modeling, mitigating hallucinations by alleviating image alignment bias. Experimental results of MCA-LLaVA across various hallucination and general benchmarks demonstrate its effectiveness and generality. The code can be accessed in https://github.com/ErikZ719/MCA-LLaVA.
Qiyan Zhao, Xiaofeng Zhang 0006, Yun Xing 0001, Xiaosong Yuan, Sinan Fan, Xuhang Chen 0002, Dahan Wang, Xu-Yao Zhang
ACM Multimedia9
2025 Dual Manifold Volume-Balanced Framework for Long-Tailed Oracle Character Recognition
Tianyu Fang, Kunchi Li, Yun Wu 0001, Dahan Wang
PRCV (7)4
2025 Dual-Branch Residual Wavelet Attention Network for Colorectal Cancer Magnifying Endoscopy Image Classification
Linrui He, Yun Wu 0001, Dahan Wang, Shunzhi Zhu, Xuyao Zhang
PRCV (14)4
2025 Low-light image enhancement with quality-oriented pseudo labels via semi-supervised contrastive learning
Nanfeng Jiang, Yiwen Cao, Xu-Yao Zhang, Dahan Wang, Chiming Wang, Shunzhi Zhu
Expert Syst. Appl.4
2025 AMST: Object tracking based on collaborative framework with adaptive multi-strategy
Rui Xu 0028, Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
Inf. Sci.4
2025 LDH-Net: Luminance-based Deep Hybrid Network for Document Image De-shadowing
Kunchi Li, Nanfeng Jiang, Yun Wu 0001, Dahan Wang
Image Vis. Comput.6
2025 ADR-Net: Attention-oriented detail recovery network for document image shadow removal
Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Yun Wu 0001, Shunzhi Zhu
Knowl. Based Syst.3
2025 Sampling Consensus by Neighborhood Interaction Information for Remote Sensing Image Matching
abstract
In remote sensing and photogrammetry, establishing reliable feature correspondences between two sets of feature points is a critical preprocessing step. In this letter, we propose a novel outlier removal method, termed sampling consensus by neighborhood interaction information (SACNI), to accurately distinguish true matches (i.e., inliers) from false matches (i.e., outliers) for remote sensing image matching. Inspired by social networks, where two individuals with close ties share numerous common connections, we propose a neighborhood interactive representation (NIR) strategy to effectively evaluate correlations among feature matches, thereby guiding the efficient sampling of outlier-free data subsets. This representation is integrated into both the initial data subset selection and optimization phases, enhancing the overall matching performance. Extensive experiments on challenging remote sensing datasets show the superiority of the proposed SACNI over several other state-of-the-art methods.
Hanlin Guo, Zhao Deng, Si Chen 0002, Dahan Wang
IEEE Geosci. Remote. Sens. Lett.7
2025 Joint radical embedding and detection for zero-shot Chinese character recognition
Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
Pattern Recognit.2
2025 Stain-adaptive self-supervised learning for histopathology image analysis
Haili Ye, Shunzhi Zhu, Dahan Wang, Xu-Yao Zhang, Heguang Huang
Pattern Recognit.4
2025 DHLA: Dynamic Hybrid Label Assignment for End-to-End Object Detection
abstract
The recent one-to-one label assignment plays a crucial role in removing the last non-differentiable component, i.e., Non-Maximum Suppression (NMS), used in the post-processing step of the one-to-many label assignment, thus building an efficient end-to-end detection system. However, due to the limited number of foreground samples, the one-to-one label assignment often suffers from insufficient representation learning, and its performance is inferior to that of traditional detectors trained using the one-to-many label assignment. To solve these problems, we introduce a novel Dynamic Hybrid Label Assignment (DHLA) method, including a Hybrid Sample Selection (HSS) strategy and a Stage-aware Soft-label Adjustment (SSA) mechanism. In order to enhance the ability of representation learning of the one-to-one label assignment, the HSS strategy subtly integrates the one-to-many and the one-to-one label assignment rules to form a simple and effective hybrid assignment rule, where high-quality samples are selected for training according to an effective task consistency metric. Moreover, the SSA mechanism dynamically adjusts the contributions of different foreground samples at different training stages, thus effectively achieving the transition from one-to-many to one-to-one label assignment. In addition, we leverage a ranking loss function to widen the score gaps between the highest scoring position and surrounding areas for effectively removing duplicate bounding boxes. As a result, our method not only learns robust feature representations during training but also performs efficient end-to-end detection during inference. Extensive experiments demonstrate our method achieves competitive performance compared to state-of-the-art detectors on the challenging COCO and CrowdHuman datasets.
Zhi-Liang Hu, Si Chen 0002, Yang Hua 0001, Dahan Wang, Shunzhi Zhu, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Hierarchical Attention-Enhanced Correlation Refinement for Robust Visual Tracking
abstract
In recent years, visual tracking has witnessed remarkable advancements with the exploration of feature extraction and correlation modeling techniques. However, inadequate robustness of either the backbone network or the correlation operation continues to plague existing trackers, leading to frustrating drift when confronted with similar distractors or cluttered backgrounds. To address this problem, we propose a hierarchical attention-enhanced correlation refinement network (HarNet) for achieving robust visual tracking. Specifically, a gated dual-view attention (GDA) module is first designed to aggregate the intra-layer attention and the inter-layer self-attention based on a fusion gate, so as to enhance hierarchical feature representations of the template. Meanwhile, a target-aware attention (TA) module introduces the template information to the inter-layer self-attention, which can highlight the target information in the search region. Moreover, a graph guided correlation (GGC) module leverages the pixel-to-local and pixel-to-global correlations to fully exploit both local-and global-spatial information between the template and the search region, and then uses the graph convolutional network (GCN) to further learn the node relationships of the correlation map for more finegrained correlations. Thus, with the above three elaborately designed modules, the HarNet is beneficial for the enhancement of feature representation and the precise localization of the target. Extensive experiments on popular visual tracking datasets (including OTB100, VOT2016, VOT2018, VOT2019, UAV123, UAV20L, GOT-10k, and LaSOT) demonstrate the superiority of our proposed method against several state-of-the-art tracking methods.
Si Chen 0002, Rui Xu 0028, Yan Yan 0001, Yang Hua 0001, Dahan Wang, Shunzhi Zhu
IEEE Trans. Intell. Transp. Syst.5
2025 Hierarchical Token-Aware Cross-Modality Reconstruction for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to query the same pedestrian's visible (infrared) images in the gallery set from the infrared (visible) images. VI-ReID not only needs to deal with the challenging factors like pose variation and occlusion, but also requires handling the large modality discrepancy. Previous methods mainly focus on learning single-scale modality-shared features and do not effectively explore the multi-scale features of two modalities from both short-range and long-range perspectives. In order to solve these problems, this paper proposes a novel Hierarchical Token-Aware Cross-Modality Reconstruction (HTCR) network to significantly mitigate the modality discrepancy for effective VI-ReID. The HTCR network consists of two main components, i.e., Hierarchical Token-aware Fusion (HTF) and Cross-modality Feature Reconstruction (CFR). The HTF module first bidirectionally exchanges the short-range and long-range multi-scale modality-shared features with a few learnable tokens to achieve discriminative pedestrian features by making full use of the advantages of both Convolutional Neural Network (CNN) and Transformer. Moreover, the CFR module reconstructs global and local pedestrian features of one modality by using the token sequence of the other modality with multi-scale cues to further explore the relationship between the two distinct modalities and alleviate the modality discrepancy. In addition, the Modality-shared feature Reconstruction (MR) loss is leveraged to reduce the noises between the reconstructed and the target features. Experimental results indicate that the proposed HTCR can significantly improve the VI-ReID performance and outperform the state-of-the-art methods on the cross-modality SYSU-MM01, RegDB, and LLCM datasets.
Si Chen 0002, Liuxiang Qiu, Dahan Wang, Wentao Zhu 0002, Yang Hua 0001, Yan Yan 0001
IEEE Trans. Multim.3
2024 High-Order Structure Based Middle-Feature Learning for Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same persons captured by visible (VIS) and infrared (IR) cameras. Existing VI-ReID methods ignore high-order structure information of features while being relatively difficult to learn a reasonable common feature space due to the large modality discrepancy between VIS and IR images. To address the above problems, we propose a novel high-order structure based middle-feature learning network (HOS-Net) for effective VI-ReID. Specifically, we first leverage a short- and long-range feature extraction (SLE) module to effectively exploit both short-range and long-range features. Then, we propose a high-order structure learning (HSL) module to successfully model the high-order relationship across different local features of each person image based on a whitened hypergraph network. This greatly alleviates model collapse and enhances feature representations. Finally, we develop a common feature space learning (CFL) module to learn a discriminative and reasonable common feature space based on middle features generated by aligning features from different modalities and ranges. In particular, a modality-range identity-center contrastive (MRIC) loss is proposed to reduce the distances between the VIS, IR, and middle features, smoothing the training process. Extensive experiments on the SYSU-MM01, RegDB, and LLCM datasets show that our HOS-Net achieves superior state-of-the-art performance. Our code is available at https://github.com/Jaulaucoeng/HOS-Net.
Liuxiang Qiu, Si Chen 0002, Yan Yan 0001, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu
AAAI5
2024 OCR4HSV: A Multi-task Learning Approach for Handwritten Signature Verification
Chao-Qun Lin, Dahan Wang, Yanfei Su, De-Wu Ge, Xu-Yao Zhang
ICPR (31)2
2024 SANS: Spatial-Aware Neural Solver for Plane Geometry Problem
Shun-Xin Xiao, Zirong Chen, Dahan Wang, Xu-Yao Zhang
ICPR (31)5
2024 Learning Explicit Radical Representations for Zero-Shot Chinese Character Recognition
Song-Liang Pan, Dahan Wang, Nanfeng Jiang, Xu-Yao Zhang, Shunzhi Zhu
ICPR (31)2
2024 Document Image Shadow Removal via Frequency Information-Oriented Network
Xinyue Zhou, Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Guantin Li, Wang Man, Yun Wu 0001
ICPR (31)4
2024 Adjustable Gating Prompt Transformer for Facial Attribute Recognition with Limited Labeled Data
Qinxian Ye, Si Chen 0002, Dahan Wang, Nanfeng Jiang, Yanfei Su, Yan Yan 0001
ICPR (28)3
2024 DocHFormer: Document Image Dewarping via Harmonized Modeling of Hierarchical Priors
Xinyue Zhou, Guanting Li, Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
ICPR (31)4
2024 MOA-Net: Multilevel Object Aware Network for Remote Sensing Image Semantic Segmentation
abstract
Remote sensing image semantic segmentation is an essential aspect of the intelligent analysis of remote sensing, extensively applied in urban planning, economic assessment, and disaster monitoring. However, the expansive field of view and intricate backgrounds in remote sensing images cause numerous objects of varying sizes and categories to coexist. This results in incomplete object segmentation and poses challenges in restoring the spatial distribution of objects. In this paper, we present a Multilevel Object Aware Network (MOA-Net), which addresses the challenges of remote sensing image semantic segmentation. To be specific, this method is designed from three perspectives. Firstly, we establish a Progressive Multiscale Global-Local Decoder (PMGLD), integrating global-local context information of objects at varying scales through a progressive convolution strategy. Secondly, the Orientation-Aware Attention Mechanism (OAAM) provides orientation information and guides the restoration of interclass 2-D spatial relationships. Finally, obtaining fine-grained features through CNN Stem improves edge segmentation. Experimental outcomes on the Potsdam and Vaihingen datasets indicate that our method of performance and efficiency surpass those of existing methods.
Yun Wu 0001, Dahan Wang
IJCNN4
2024 RCFormer : Interactive Image Segmentation via Reconstructing Click Vision Transformers
abstract
Click-based interactive image segmentation intends to segment an object from the background under user click guidance. Recently, Vision Transformer has made significant strides in interactive image segmentation. However, the previous studies 1) overlook the importance of different clicks in terms of their contribution to the segmentation results; and 2) suffer from inconsistency across different feature scales in the multi-scale structure. In this paper, we propose a new interactive segmentation framework, named RCFormer, with two novel components: reconstruct click patch embedding (RCPE) for encoding the importance of clicks, and multi-scale adaptive fusion (MSAF) for the adaptive fusion of feature maps across different scales. RCPE enhances the effectiveness of click interactions by spatially distinguishing the importance of clicks. MSAF adaptively fuses useful spatial information and filters the redundant feature at multi-scales. The experiments on several benchmarks show that our proposed approach achieves state-of-the-art performance. Notably, our method achieves 2.31 NoC@90 on the Berkeley dataset, improving by 8.6% over the previous best results.
PanPan Chen, Dahan Wang, Yun Wu 0001, Xu-Yao Zhang, Shunzhi Zhu
IJCNN2
2024 Enhancing Lightweight Remote Sensing Semantic Segmentation via Weak Consistency Regularization
abstract
Remote sensing image semantic segmentation has widespread applications in urban planning and land monitoring. In recent years, U-Net and its variant networks have almost dominated the research in the field of semantic segmentation. However, many models pay less attention to computational efficiency, rendering them ineffective in scenarios with computational resource and timeliness constraints, such as autonomous driving and disaster monitoring. To address this issue, we propose the USA-Net (UNet-like with Shifted Axial), a lightweight hybrid model based on convolution and MLP (Multi-Layer Perceptron). Specifically, we design the ST Block (Shift Tokenized Block), which introduces local features into global operations in MLP through spatial shift, and then use ELCM (Efficient Large-kernel Convolution Module) to enlarge the model’s receptive field and learn the shape features of objects. Additionally, we propose a new semi-supervised learning framework to further improve the model’s generalization performance. On the ISPRS Vaihingen and ISPRS Potsdam datasets, USA-Net significantly outperforms most state-of-the-art methods in terms of segmentation accuracy and efficiency.
Zirong Chen, Shunxin Xiao, Wang Man, Dahan Wang, Shunzhi Zhu
IJCNN4
2024 Character Relationship Refinement Network for Handwritten Mathematical Expression Recognition
abstract
Most current Handwritten Mathematical Expression Recognition (HMER) methods employ an attention-based encoder-decoder framework, which generates LaTeX sequences from the given images, following the paradigm of predicting "one-by-one". However, this paradigm may have some challenges: 1) without considering the connectivity between characters, the prior information in the prediction process will be ignored inadvertently, especially implicit information, such as " " and " ˆ ". 2) Some characters of high similarities, such as "6/b" and "o/O", will have negative effects on prediction results. To solve these issues, we propose a simple but effective Character Relationship Refinement Network (CRRN), which consists of Joint Character Learning (JCL) and Character Refinement Mask (CRM). Specifically, JCL calculates the relationship probability between characters and uses them to improve prediction accuracy. CRM takes the character confidence coefficient in a coarse-to-fine way that can reassign the weights of all characters to improve model discriminability on easily confused characters. With the collaboration of both modules, our proposed CRRN can outperform the state-of-the-art on popular datasets.
LiWei Jiang, Nanfeng Jiang, Yun Wu 0001, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
IJCNN4
2024 Bridge the Gap of Semantic Context: A Boundary-Guided Context Fusion UNet for Medical Image Segmentation
Dahan Wang, Shunzhi Zhu
PRCV (15)3
2024 Local neighbor propagation on graphs for mismatch removal
Hanlin Guo, Guobao Xiao, Lumei Su, Jiaxing Zhou, Dahan Wang
Inf. Sci.5
2024 GCAT: graph calibration attention transformer for robust object tracking
Si Chen 0002, Xinxin Hu, Dahan Wang, Yan Yan 0001, Shunzhi Zhu
Neural Comput. Appl.3
2024 HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-ID) aims to match the same pedestrian of video sequences across non-overlapping cameras. Video re-ID methods generally adopt frame-level feature extraction for different video frames, but they still lack effective spatio-temporal interaction, easily leading to the multi-frame misalignment problem. In this paper, we propose a Hierarchical Attention-aware Spatio-temporal Interaction (HASI) network, including an Attention-aware Temporal Interaction (ATI) module and a Hierarchical Local-spatial Enhancement (HLE) module for video-based person re-ID. In order to avoid the spatial misalignment between video frames, the ATI module employs multiple Frame-to-Frame Temporal Interaction (2FTI) blocks with the Multi-head Inter-frame Alignment Attention (MIAA) to make the current frame iteratively interact with each rest frame of a video in a positive single-cycle manner, rather than only interacting with the adjacent frame or directly building the relationship of all frames at once. This module can not only obtain the long-range non-adjacent temporal information, but also learn the pairwise frame-to-frame relationships. Moreover, the HLE module is designed to enhance the local fine-grained features from multiple Transformer layers, whilst delivering low-level information to further enrich middle-level and high-level semantic knowledge. Thus, our method can learn multi-perspective pedestrian information, including inter-frame long-range interaction information and intra-frame multi-layer global and local information. Extensive experiments demonstrate the superiority of the proposed HASI method compared with the state-of-the-art methods on the three challenging video-based re-ID datasets, i.e., MARS, iLIDS-VID, and PRID-2011.
Si Chen 0002, Hui Da, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Second-Order Proximity Guided Sampling Consensus for Robust Model Fitting
abstract
Robust model fitting plays a critical role in artificial intelligence and computer vision, with its performance primarily depends on the utilization of sampling algorithms. However, existing sampling algorithms become less effective when initial correspondences between two images are corrupted by a large number of outliers, especially in the presence of multi-structure data. In this paper, we propose a novel sampling algorithm (called SPGSC) for robust model fitting, where minimal subsets are sampled with the guidance of the second-order proximity measure, which involves global geometric relationships instead of local consistency relationships. Specifically, we first propose a second-order proximity measure to facilitate graph construction, which helps detect a potential inlier from input data as the first datum (i.e., the seed datum) of a minimal subset. After that, we propose a second-order proximity based initial minimal subset generation strategy, which is able to choose a certain number of minimal subsets by the seed data for efficiently producing significant model hypotheses. Furthermore, to achieve better fitting performance, we propose a maximum spanning tree based refinement (MSTR) strategy, which is used to refine the previous sampled minimal subsets and improve the effectiveness and efficiency of the sampling process. Experimental results on three vision tasks (i.e., two-view based motion segmentation, affine matrix based segmentation, and 3D motion segmentation) show the superiority of the proposed SPGSC in comparison with other state-of-the-art algorithms.
Hanlin Guo, Guobao Xiao, Lumei Su, Tianyou Li, Dahan Wang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.5
2024 Multi-Branch Enhanced Discriminative Network for Vehicle Re-Identification
abstract
Vehicle re-identification (ReID) is the task of identifying the same vehicle across numerous cameras. This is a complex classification task, and the fine-grained information and strong discrimination features have proven to be effective in handling the re-identification classification task. However, most existing methods focuses on extracting local area features in combination with global features, while exploring subtle distinguishing features, which is a difficult task, remains an open problem and unsolved. In this paper, we propose a multi-branch enhanced discriminative network (MED) to better extract subtle distinguishing features that have high discriminative power to improve the ReID performance. In the proposed MED method, each feature map obtained by convolutional neural network (CNN) is divided into 4 spatial sub-maps, on each of which, the vertical and the horizontal branches are used to extract the subtle distinguishing features intrinsically contained in sub-areas. The vertical and the horizontal branches are combined with the global branch to perform the ReID task. Moreover, our proposed method is capable of extracting rich fine-grained features without the need of extra manual annotation while maintaining a simple design structure. We conducted extensive experiments on the vehicle ReID datasets (VehicleID and VeRi-776), showing that the proposed MED method outperforms most existing methods. Further, we directly apply the MED method to the pedestrian ReID problem on the Market-1501, DUKEMTMC, and MSMT17 datasets, achieving the state-of-the-art (SOTA) performance as well. This demonstrates that the proposed method has good generality and can be flexibly applied to the ReID tasks.
Jiawei Lian, Dahan Wang, Yun Wu 0001, Shunzhi Zhu
IEEE Trans. Intell. Transp. Syst.2
2023 Multi-Zone Transformer Based on Self-Distillation for Facial Attribute Recognition
abstract
Recently, transformers have shown great promising performance in various computer vision tasks. However, the current transformer based methods ignore the information exchanges between transformer blocks, and they have not been applied in the facial attribute recognition task. In this paper, we propose a multi-zone transformer based on self-distillation for FAR, termed MZTS, to predict the facial attributes. A multi-zone transformer encoder is firstly presented to achieve the interactions of the different transformer encoder blocks, thus avoiding forgetting the effective information between the transformer encoder block groups during the iteration process. Furthermore, we introduce a new self-distillation mechanism based on class tokens, which distills the class tokens obtained from the last transformer encoder block group to the other shallow groups by interacting with the significant information between the different transformer blocks through attention. Extensive experiments on the challenging CelebA and LFWA datasets have demonstrated the excellent performance of the proposed method for FAR.
Si Chen 0002, Xueyan Zhu, Dahan Wang, Shunzhi Zhu, Yun Wu 0001
FG3
2023 A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids
abstract
This paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing module. Without using the head rotation information and the enrollment speech, the system can reach an average hearing aid speech perception index (HASPI) score of 0.696 and hearing aid speech quality index (HASQI) score of 0.320 on the official development set. The corresponding scores are 0.729 and 0.316 respectively on the Eval1 set for the challenge ranking.
Zhongshu Hou, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen 0029
ICASSP7
2023 A Shallow Graph Neural Network with Innovative Node Updating for Online Handwritten Stroke Classification
Yan-Rong Wang, Dahan Wang, Xiao-Long Yun, Shunzhi Zhu
ICDAR (4)2
2023 UAM-Net: An Attention-Based Multi-level Feature Fusion UNet for Remote Sensing Image Segmentation
Yiwen Cao, Nanfeng Jiang, Dahan Wang, Yun Wu 0001, Shunzhi Zhu
PRCV (4)3
2023 MCKIE: Multi-class Key Information Extraction from Complex Documents Based on Graph Convolutional Network
Zhicai Huang, Shunxin Xiao, Dahan Wang, Shunzhi Zhu
PRCV (7)3
2023 Pseudo Labels Refinement with Stable Cluster Reconstruction for Unsupervised Re-identification
Jiawei Lian, Dahan Wang, Yun Wu 0001, Shunzhi Zhu, Dewu Ge
PRCV (4)4
2023 GridIIS: Grid Based Interactive Image Segmentation
Pengqi Zhu, Dahan Wang, Shunzhi Zhu
PRCV (11)2
2023 SiamCCF: Siamese visual tracking via cross-layer calibration fusion
abstract
Abstract Siamese networks have attracted wide attention in visual tracking due to their competitive accuracy and speed. However, the existing Siamese trackers usually leverage a fixed linear aggregation of feature maps, which does not effectively fuse the different layers of features with attention. Besides, most of Siamese trackers calculate the similarity between the template and the search region through a cross‐correlation operation between the features of the last blocks from the two branches, which might introduce the redundant noise information. In order to solve these problems, this study proposes a novel Siamese visual tracking method via cross‐layer calibration fusion, termed SiamCCF. An attention‐based feature fusion module is employed using local attention and non‐local attention to fuse the features from the deep and shallow layers, so as to capture both local details and high‐level semantic information. Moreover, a cross‐layer calibration module can use the fused features to calibrate the features of the last network blocks and build the cross‐layer long‐range spatial and inter‐channel dependencies around each spatial location. Extensive experiments demonstrate that the proposed method has achieved competitive tracking performance compared with state‐of‐the‐art trackers on challenging benchmarks, including OTB100, OTB2013, UAV123, UAV20L, and LaSOT.
Si Chen 0002, Shunzhi Zhu, Huarong Xu, Dahan Wang
IET Comput. Vis.6
2023 Learning an attention-aware parallel sharing network for facial attribute recognition
Si Chen 0002, Xinyu Lai, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
J. Vis. Commun. Image Represent.4
2023 MTNet: Mutual tri-training network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Liuxiang Qiu, Zimin Tian, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
J. Vis. Commun. Image Represent.5
2023 Self-information of radicals: A new clue for zero-shot Chinese character recognition
Dahan Wang, Xia Du, Huayi Yin, Xu-Yao Zhang, Shunzhi Zhu
Pattern Recognit.2
2023 Identity-Aware Contrastive Knowledge Distillation for Facial Attribute Recognition
abstract
Facial attribute recognition (FAR) is an important and yet challenging multi-label learning task in computer vision. Existing FAR methods have achieved promising performance with the development of deep learning. However, they usually suffer from prohibitive computational and memory costs. In this paper, we propose an identity-aware contrastive knowledge distillation method, termed ICKD, to compress the FAR model. A nonlinear weight-sharing mapping (NWSM) mechanism is firstly designed to avoid the difficulty of directly matching features of the teacher and student networks due to the lower representation ability of the student network. Furthermore, an identity-aware contrastive distillation (ICD) loss is employed to guide the student network to effectively learn the mutual relations between samples with multiple attributes. In addition, an adjustable ladder distillation (ALD) loss is developed to automatically adjust the importance of different distillation points with the progress of training. Extensive experiments demonstrate that our method can significantly improve the performance of student networks and outperforms the existing FAR methods on the public challenging datasets.
Si Chen 0002, Xueyan Zhu, Yan Yan 0001, Shunzhi Zhu, Shaozi Li, Dahan Wang
IEEE Trans. Circuits Syst. Video Technol.6
2022 Graph Attention Transformer Network for Robust Visual Tracking
Si Chen 0002, Dahan Wang, Shunzhi Zhu
ICONIP (4)4
2022 Critical Radical Analysis Network for Chinese Character Recognition
abstract
Zero-shot learning is a challenging problem in many tasks due to the lack of training samples of the unseen classes. The radical-based zero-shot Chinese character recognition methods treat Chinese characters as a combination of radicals and structures, and recognize Chinese characters by identifying the radicals and structures contained in them. Current approaches generally treat the contribution of all radicals to Chinese character recognition as the same, and the recognition results rely on the network’s ability to recognize radicals and their corresponding position information, ignoring the potential value of radicals themselves in eliminating the uncertainty of Chinese characters. In this paper, we model the problem of radical-based Chinese character recognition as an uncertainty elimination problem and propose a Critical Radical Analysis Network (CRAN) to explore the Ideographic Description Sequence (IDS) information for zero-shot Chinese character recognition. Specifically, we propose a novel method to compute the critical values of radicals based on information theory using the predefined Chinese character IDS dictionary. In recognition, we use an iterative approach to translate the predicted radical sequence to target Chinese characters. That is, the radicals of the predicted sequence are sorted in descending order of the critical value, and then the radicals are continuously selected in this order as the information obtained to eliminate the uncertainty of the Chinese character until the character is recognized. We conduct experiments on the CTW, CASIA-AHCDB, and CASIA-HWDB datasets. The experimental results show that the proposed method improves the ability of recognizing unseen Chinese characters, demonstrating the effectiveness of the proposed method.
Huayi Yin, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
ICPR3
2022 Exploiting Robust Memory Features for Unsupervised Reidentification
Jiawei Lian, Dahan Wang, Xia Du, Yun Wu 0001, Shunzhi Zhu
PRCV (2)2
2022 Semantic-Aware Non-local Network for Handwritten Mathematical Expression Recognition
Xiang-Hao Liu, Dahan Wang, Xia Du, Shunzhi Zhu
PRCV (3)2
2022 Joint Pixel-Level and Feature-Level Unsupervised Domain Adaptation for Surveillance Face Recognition
Huangkai Zhu, Huayi Yin, Du Xia, Dahan Wang, Xianghao Liu, Shunzhi Zhu
PRCV (3)4
2022 Learning meta-adversarial features via multi-stage adaptation network for robust visual object tracking
Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu
Neurocomputing5
2021 Detecting urban hot regions by using massive geo-tagged image data
Dahan Wang, Shunzhi Zhu
Neurocomputing1
2021 Few-labeled visual recognition for self-driving using multi-view visual-semantic representation
Dahan Wang, Shunzhi Zhu
Neurocomputing1
2021 Discriminative semantic region selection for fine-grained recognition
Chunjie Zhang 0001, Dahan Wang, Hai-Sheng Li 0002
J. Vis. Commun. Image Represent.2
2021 Exploring the Prediction Consistency of Multiple Views for Transductive Visual Recognition
abstract
Although great process has been achieved to accurately classify images, many methods only use labeled images while ignoring the large quantity of unlabeled images. To make use of unlabeled images, in this letter, we propose a novel transductive visual recognition method using the prediction consistency of multiple views (T-PCMV). Both labeled and unlabeled images are used in a unified framework. The predictions of unlabeled images are learned by linearly combining the discriminative information of multiple views. We ensure the smooth constraint that visually similar images should be predicted with similar labels. To learn the classifier, we jointly minimize the classification loss and the discrepancy of predicted labels. To evaluate the usefulness of the proposed method, we conduct transductive visual recognition experiments on four image datasets. Experimental results well demonstrate the effectiveness of the proposed T-PCMV method.
Chunjie Zhang 0001, Dahan Wang
IEEE Signal Process. Lett.2
2020 ICFHR 2020 Competition on Offline Recognition and Spotting of Handwritten Mathematical Expressions - OffRaSHME
abstract
This paper presents the competition on Offline Recognition and Spotting of Handwritten Mathematical Expressions (OffRaSHME) held at the 17th International Conference on Frontiers in Handwriting Recognition (ICFHR 2020). Handwritten Mathematical Expression Recognition (HMER) has wide potential applications and is capturing increasing attention in recent years. Previous HMER competitions mainly focused on online datasets or offline datasets that are converted from online data. In this competition, we have collected a dataset of offline handwritten mathematical expressions by scanning papers that contain expressions. Moreover, we labeled the offline dataset at symbol level, i.e., the bounding boxes of each symbol are also provided, to facilitate the research of HMER. At last, 19,749 offline handwritten mathematical expressions are collected for training, and 2,000 ones are provided for evaluating the participating systems. In the competition, 7 teams submitted 8 systems for the task of offline HMER, among which 5 systems only use the provided datasets without any extra data while 3 systems use extra data. The winner team achieved a recognition accuracy of 79.85% (without extra data) and 81.85% (with extra data) on the offline formula recognition task.
Dahan Wang, Jin-Wen Wu, Yu-Pei Yan, Zhicai Huang, Gui-Yun Chen, Cheng-Lin Liu 0001
ICFHR1
2020 Attention Combination of Sequence Models for Handwritten Chinese Text Recognition
abstract
Handwritten Chinese text recognition (HCTR) methods can generally be divided into two categories: explicit segmentation based methods and implicit segmentation based methods. The explicit segmentation based approaches are superior in the accuracy rate of character recognition, while the implicit segmentation based approaches can provide more fluent recognition result due to the end-to-end recognition mode of it. In this paper, we propose to combine this two kinds of approaches using convolutional combination strategy. The proposed system includes multiple encoders and a decoder in the form of full convolution. It combines the recognition results of explicit segmentation based HCTR methods and implicit segmentation based HCTR methods, and generates the final recognition output. A novel attention combination mechanism is proposed to solve this one-to-many attention problem. Experiments on the ICDAR13 dataset show that the proposed method improves the accurate rate by 4.96%. The experimental results demonstrate the effectiveness of proposed method.
Dahan Wang
ICFHR3
2020 Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text Classification
abstract
Extreme multi-label text classification (XMTC) is a task for tagging a given text with the most relevant labels from an extremely large label set. We propose a novel deep learning method called APLC-XLNet. Our approach fine-tunes the recently released generalized autoregressive pretrained model (XLNet) to learn a dense representation for the input text. We propose Adaptive Probabilistic Label Clusters (APLC) to approximate the cross entropy loss by exploiting the unbalanced label distribution to form clusters that explicitly reduce the computational time. Our experiments, carried out on five benchmark datasets, show that our approach has achieved new state-of-the-art results on four benchmark datasets. Our source code is available publicly at https://github.com/huiyegit/APLC_XLNet.
Zhiyu Chen 0001, Dahan Wang, Brian D. Davison 0001
ICML3
2020 2D License Plate Recognition based on Automatic Perspective Rectification
abstract
License plate recognition (LPR) remains a challenging task in face of some difficulties such as image deformation and multi-line character distribution. Text rectification that is crucial to eliminate the effects of image deformation has attracted increasing attentions in scene text recognition. However, current text rectification methods are not designed specifically for LPR, which did not take the features of plate deformation into account. Considering the fact that a license plate (LP) can only generate perspective distortion in the image due to its rigid feature, in this paper we propose a novel perspective rectification network (PRN) to automatically estimate the perspective transformation and rectify the distorted LP accordingly. For recognition, we propose a location-aware 2D attention based recognition network that is capable of recognizing both single-line and double-line plates with perspective deformation. The rectification network and recognition network are connected for end-to-end training. Experiments on common datasets show that the proposed method achieves the state-of-the-art performance, demonstrating the effectiveness of the proposed approach.
Zhaohong Guo, Dahan Wang, Yu Shi 0003
ICPR3
2020 Inferring region significance by using multi-source spatial data
Shunzhi Zhu, Dahan Wang, Danhuai Guo
Neural Comput. Appl.2
2020 Generative adversarial networks with decoder-encoder output noises
Guoqiang Zhong 0001, Youzhao Yang, Dahan Wang, Kaizhu Huang
Neural Networks5
2019 CASIA-AHCDB: A Large-Scale Chinese Ancient Handwritten Characters Database
abstract
This paper introduces a Chinese Ancient Handwritten Characters Database (CASIA-AHCDB) for character recognition research. The database was built by annotating 11,937 pages of Chinese ancient handwritten documents. It consists of more than 2.2 million annotated handwritten character samples of 10,350 categories. According to the source of these documents, the database is divided into two datasets of different styles: Complete Library in Four Sections (AHCDB-style1) and Ancient Buddhist Scriptures (AHCDB-style2). Each dataset can be divided into three parts based on its applications. The first part, called basic category set, contains samples of common categories in two datasets, and is suitable for basic character recognition task. The second part, called enhanced category set, is mainly used for open-set character recognition task based on the basic character recognition. The third part, called the reserved category set, can be used in many pattern recognition tasks in the future. Based on the large category set, the various writing styles and the imbalanced sample number per category, CASIA-AHCDB can also be used for various classification and learning tasks such as transfer learning, few-shot learning. We performed experiments of basic character recognition on the basic category set, and report the results for benchmark. More techniques can be evaluated on this challenging database in the future.
Dahan Wang, Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001
ICDAR3
2016 Rapid hypothesis generation by combining residual sorting with local constraints
Taotao Lai, Hanzi Wang, Yan Yan 0001, Dahan Wang, Guobao Xiao
Multim. Tools Appl.4
2015 Scene Character and Text Recognition: The State-of-the-Art
Chongmu Chen, Dahan Wang, Hanzi Wang
ICIG (3)2
2014 Combining preference analysis with local constraints for rapid hypothesis generation
abstract
Hypothesis generation is crucial to many robust model fitting methods. In this paper, we propose an effective hypothesis generation method by adopting conditional sampling with local constraints. We choose data to generate hypotheses according to sampling weights, which are computed according to ordered residual indices. To sample a minimal subset, we randomly choose a seed datum, compute sampling weights of all data with regard to the seed datum, search the neighborhood set of the seed datum by using the sampling weights, and then sample the remaining data of the minimal subset from the neighborhood set. It has two advantages to consider the neighboring information in guided sampling: It raises the probability of generating all-inlier minimal subsets and it reduces the computational loads in hypotheses generation. The proposed method shows good performance in fundamental matrix estimation using real image pairs.
Taotao Lai, Dahan Wang, Guobao Xiao, Hanzi Wang
ICARCV2
2014 Scene text recognition using sparse coding based features
abstract
In this paper, we propose an effective scene text recognition method using sparse coding based features, called Histograms of Sparse Codes (HSC) features. For character detection, we use the HSC features instead of using the Histograms of Oriented Gradients (HOG) features. HSC features are extracted by computing sparse codes with dictionaries, which are learned from data using K-SVD, and aggregating perpixel sparse codes to form local histograms. For word recognition, we integrate multiple cues including character detection scores and geometric contexts in an objective function. The final recognition result is obtained by searching for the word which corresponds to the maximum value of the objective function. The parameters in the objective function are learned using the Minimum Classification Error (MCE) training method. Experiments on the ICDAR2003 and SVT datasets demonstrate that the HSC-based scene text recognition method outperforms the HOG-based method significantly and achieves the state-of-the-art performance.
Dahan Wang, Hanzi Wang
ICIP2
2014 Learning confidence transformation for handwritten Chinese text recognition
Dahan Wang, Cheng-Lin Liu 0001
Int. J. Document Anal. Recognit.1
2014 Character confidence based on N-best list for keyword spotting in online Chinese handwritten documents
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001
Pattern Recognit.2
2013 Learning-Based Candidate Segmentation Scoring for Real-Time Recognition of Online Overlaid Chinese Handwriting
abstract
In overlaid handwriting, multiple characters are written sequentially in the same area. This needs special consideration for segmenting the stroke sequence into characters. We propose a learning-based model for scoring the candidate stroke cuts and segments for online overlaid Chinese handwriting recognition. Based on stroke cut classification using support vector machine (SVM), strokes are grouped into segments, and consecutive segments are concatenated into candidate characters. The likeliness of candidate characters (unary geometry) and the compatibility between adjacent characters (binary geometry) are measured by combining the stroke cut score and the between-segment geometric score, and are integrated with the character classification score and linguistic context for character string recognition. Experiments on a large database of online Chinese handwriting demonstrate the effectiveness of the proposed method.
Yan-Fei Lv, Linlin Huang 0001, Dahan Wang, Cheng-Lin Liu 0001
ICDAR3
2013 Keyword Spotting from Online Chinese Handwritten Documents using One-versus-All Character Classification Model
abstract
In this paper, we propose a method for text-query-based keyword spotting from online Chinese handwritten documents using character classification model. The similarity between the query word and handwriting is obtained by combining the character classification scores. The classifier is trained by one-versus-all strategy so that it gives high similarity to the target class and low scores to the others. Using character classification-based word similarity also helps overcome the out-of-vocabulary (OOV) problem. We use a character-synchronous dynamic search algorithm to efficiently spot the query word in large database. The retrieval performance is further improved by using competing character confusion and writer-adaptive thresholds. Our experimental results on a large handwriting database CASIA-OLHWDB justify the superiority of one-versus-all trained classifiers and the benefits of confidence transformation, character confusion and adaptive thresholds. Particularly, a one-versus-all trained prototype classifier performs as well as a linear support vector machine (SVM) classifier, but consumes much less storage of index file. The experimental comparison with keyword spotting based on handwritten text recognition also demonstrates the effectiveness of the proposed method.
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001, Horst Bunke
Int. J. Pattern Recognit. Artif. Intell.2
2013 Handwritten Chinese/Japanese Text Recognition Using Semi-Markov Conditional Random Fields
abstract
This paper proposes a method for handwritten Chinese/Japanese text (character string) recognition based on semi-Markov conditional random fields (semi-CRFs). The high-order semi-CRF model is defined on a lattice containing all possible segmentation-recognition hypotheses of a string to elegantly fuse the scores of candidate character recognition and the compatibilities of geometric and linguistic contexts by representing them in the feature functions. Based on given models of character recognition and compatibilities, the fusion parameters are optimized by minimizing the negative log-likelihood loss with a margin term on a training string sample set. A forward-backward lattice pruning algorithm is proposed to reduce the computation in training when trigram language models are used, and beam search techniques are investigated to accelerate the decoding speed. We evaluate the performance of the proposed method on unconstrained online handwritten text lines of three databases. On the test sets of databases CASIA-OLHWDB (Chinese) and TUAT Kondate (Japanese), the character level correct rates are 95.20 and 95.44 percent, and the accurate rates are 94.54 and 94.55 percent, respectively. On the test set (online handwritten texts) of ICDAR 2011 Chinese handwriting recognition competition, the proposed method outperforms the best system in competition.
Dahan Wang, Feng Tian 0001, Cheng-Lin Liu 0001, Masaki Nakagawa
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Online and offline handwritten Chinese character recognition: Benchmarking on new databases
Cheng-Lin Liu 0001, Dahan Wang, Qiufeng Wang 0001
Pattern Recognit.3
2012 String-level learning of confidence transformation for Chinese handwritten text recognition
Dahan Wang, Cheng-Lin Liu 0001
ICPR1
2012 A confidence-based method for keyword spotting in online Chinese handwritten documents
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001
ICPR2
2012 An approach for real-time recognition of online Chinese handwritten sentences
Dahan Wang, Cheng-Lin Liu 0001
Pattern Recognit.1
2011 CASIA Online and Offline Chinese Handwriting Databases
abstract
This paper introduces a pair of online and offline Chinese handwriting databases, containing samples of isolated characters and handwritten texts. The samples were produced by 1,020 writers using Anoto pen on papers for obtaining both online trajectory data and offline images. Both the online samples and offline samples are divided into six datasets, three for isolated characters (DB1.0-C1.2) and three for handwritten texts (DB2.0-C2.2). The (either online or offline) datasets of isolated characters contain about 3.9 million samples of 7,356 classes (7,185 Chinese characters and 171 symbols), and the datasets of handwritten texts contain about 5,090 pages and 1.35 million character samples. Each dataset is segmented and annotated at character level, and is partitioned into standard training and test subsets. The online and offline databases can be used for the research of various handwritten document analysis tasks.
Cheng-Lin Liu 0001, Dahan Wang, Qiufeng Wang 0001
ICDAR3
2011 ICDAR 2011 Chinese Handwriting Recognition Competition
abstract
In the Chinese handwriting recognition competition organized with the ICDAR 2011, four tasks were evaluated: offline and online isolated character recognition, offline and online handwritten text recognition. To enable the training of recognition systems, we announced the large databases CASIA-HWDB/OLHWDB. The submitted systems were evaluated on un-open datasets to report character-level correct rates. In total, we received 25 systems submitted by eight groups. On the test datasets, the best results (correct rates) are 92.18% for offline character recognition, 95.77% for online character recognition, 77.26% for offline text recognition, and 94.33% for online text recognition, respectively. In addition to the evaluation results, we provide short descriptions of the recognition methods and have brief discussions.
Cheng-Lin Liu 0001, Qiufeng Wang 0001, Dahan Wang
ICDAR4
2011 Dynamic Text Line Segmentation for Real-Time Recognition of Chinese Handwritten Sentences
abstract
Real-time recognition of handwritten sentences enables fast text input but the dynamic nature of writing makes reliable text line segmentation difficult. This paper proposes a method for real-time dynamic text line segmentation of online Chinese handwriting. The core of the method is a statistical classifier for modeling the geometric relationship between an ongoing stroke and the previous text lines, to assign the stroke into a previous line or form a new line. The method can deal with delayed strokes and therefore enables robust real-time recognition. We evaluated the segmentation performance on a dataset of online Chinese handwriting by simulating the real-time writing and recognition process. The experimental results demonstrate the effectiveness and robustness of the proposed method.
Dahan Wang, Cheng-Lin Liu 0001
ICDAR1
2011 Transcript Mapping for Handwritten Text Lines Using Conditional Random Fields
abstract
This paper presents a conditional random field (CRF) model for aligning online handwritten Chinese/Japanese text lines (character strings) with the corresponding transcripts. The CRF model is defined on a lattice which contains all possible segmentation hypotheses. The feature functions characterize the shape and context dependences of characters, including the scores of character recognition and the geometric compatibilities between characters. The combining parameters are optimized by energy minimization. Experimental results on two online databases: CASIA-OLHWDB and TUAT Kondate demonstrate the effectiveness of the proposed method.
Dahan Wang, Qiufeng Wang 0001, Masaki Nakagawa, Cheng-Lin Liu 0001
ICDAR3
2010 Keyword Spotting from Online Chinese Handwritten Documents Using One-vs-All Trained Character Classifier
abstract
This paper presents a text query-based method for keyword spotting from online Chinese handwritten documents. The similarity between a text word and handwriting is obtained by combining the character similiarity scores given by a character classifier. To overcome the ambiguity of character segmentation, multiple candidates of character patterns are generated by over-segmentation, and sequences of candidate characters are matched with the query word in beam search. The character classifier is trained by one-vs-all strategy so that it gives high similarity to the target class and low scores to the others. Particularly, we use a one-vs-all trained prototype classifier and a support vector machine (SVM) classifier for similarity scoring. The method yielded promising performance in experiments on a database containing 550 pages of 110 writers. For words of four characters, the recall, precision and F measure are 87.25%, 94.84% and 90.88%, respectively.
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001
ICFHR2
2010 Error Reduction by Confusing Characters Discrimination for Online Handwritten Japanese Character Recognition
abstract
To reduce the classification errors of online handwritten Japanese character recognition, we propose a method for confusing characters discrimination with little additional costs. After building confusing sets by cross validation using a baseline quadratic classifier, a logistic regression (LR) classifier is trained to discriminate the characters in each set. The LR classifier uses subspace features selected from existing vectors of the baseline classifier, thus has no extra parameters except the weights, which consumes a small storage space compared to the baseline classifier. In experiments on the TUAT HANDS databases with the modified quadratic discriminant function (MQDF) as baseline classifier, the proposed method has largely reduced the confusion caused by non-Kanji characters.
Dahan Wang, Masaki Nakagawa, Cheng-Lin Liu 0001
ICFHR2
2009 CASIA-OLHWDB1: A Database of Online Handwritten Chinese Characters
abstract
This paper describes a publicly available database, CASIA-OLHWDB1, for research on online handwritten Chinese character recognition. This database is the first of our series of online/offline handwritten characters and texts, collected using Anoto pen on paper. It contains unconstrained handwritten characters of 4,037 categories (3,866 Chinese characters and 171 symbols) produced by 420 persons, and 1,694,741 samples in total. It can be used for design and evaluation of character recognition algorithms and classifier design for handwritten text recognition systems. We have partitioned the samples into three grades and into training and test sets. Preliminary experiments on the database using a state-of-the-art recognizer justify the challenge of recognition.
Dahan Wang, Cheng-Lin Liu 0001, Jin-Lun Yu
ICDAR1
2009 A robust approach to text line grouping in online handwritten Japanese documents
Dahan Wang, Cheng-Lin Liu 0001
Pattern Recognit.2
2008 Grouping Text Lines in Online Handwritten Japanese Documents by Combining Temporal and Spatial Information
abstract
We present an effective approach for grouping text lines in online handwritten Japanese documents by combining temporal and spatial information. Initially, strokes are grouped into text line strings according to off-stroke distances. Each text line string is segmented into text lines by dynamic programming (DP) optimizing a cost function trained by the minimum classification error (MCE) method. Over-segmented text lines are then merged with a support vector machine (SVM) classifier for making merge/non-merge decisions, and last, a spatial merge module corrects the segmentation errors caused by delayed strokes. In experiments on the TUAT Kondate database, the proposed approach achieves the Entity Detection Metric (EDM) rate of 0.8816, the Edit-Distance Rate (EDR) of 0.1234, which demonstrates the superiority of our approach.
Dahan Wang, Cheng-Lin Liu 0001
Document Analysis Systems2