Shuangping Huang

dblp:26/7950 · DBLP profile ↗
← Back
51ranked-venue papers
8as first author
39since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 4 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2027 Dual-granularity image-text alignment for zero-shot composed image retrieval
Wenjie Peng, Shuangping Huang, Yunqing Hu, Tianshui Chen
Expert Syst. Appl.2
2026 Knowledge-embedded graph representation learning for document-level relation extraction
Jinglin Liang 0001, Yutao Qin, Shuangping Huang, Yunqing Hu, Xinwu Liu, Tianshui Chen
Expert Syst. Appl.3
2026 DC-RRG: Diagnosis-centered cascaded radiology report generation
Zinan Hong, Zijian Zhou 0002, Miaojing Shi, Jin-Gang Yu, Jingping Yun, Shuangping Huang
Expert Syst. Appl.7
2026 An alignment-error-free framework for end-to-end table recognition
Fan Yang 0082, Ling Deng, Zhiyong Gan, Shuangping Huang, Tianshui Chen
Expert Syst. Appl.4
2026 Proxy-AN loss for deep metric learning
Wenjie Peng, Quhui Ke, Jinglin Liang 0001, Shuangping Huang, Tianshui Chen
Neural Networks4
2026 The unambiguous structure representation of tabular data for recognition
Fan Yang 0082, Junwen Tan, Tianshui Chen, Shuangping Huang, Yunqing Hu
Neural Networks4
2026 ATCMD-Bench: Agentic Traditional Chinese Medicine diagnosis benchmark for Large Language Models via multi-agent simulation
Junxiang Lin, Gang Dai 0002, Wenjie Peng, Shuangping Huang, Yingrong Lao, Huamin Zhang, Tianshui Chen
Pattern Recognit.4
2026 A text-only weakly supervised learning framework for text spotting via text-to-polygon generator
Zhiyong Gan, Ling Deng, Shuaicheng Niu, Zhenghua Peng, Shuangping Huang
Pattern Recognit.7
2026 Improving Pseudo-Labeling by Dynamic Confidence Calibration for Semi-Supervised Sequence Recognition
abstract
Sequence recognition models under a fully supervised paradigm require large-scale training data, incurring substantial annotation costs. Pseudo-labeling is one of the most effective techniques in semi-supervised learning, which leverages predicted confidence to filter pseudo-labels on unlabeled data for model training. However, recent studies indicate that the performance of semi-supervised learning is compromised by overconfident models, as the predicted unreliable confidences will filter noisy samples into training. In this work, we discover that the overconfidence in sequence recognition models is influenced by the linguistic properties of a sequence, where the tail character classes are prone to be mispredicted as the head ones that frequently appear in the language with high confidence. And this overconfidence continuously intensifies throughout the semi-supervised training process. To address this limitation, we propose a Dynamic Sequential Class-Aware Smoothing (DSCS) method that calibrates the overconfidence of the head class to alleviate the inaccurate pseudo-labeling caused by overconfident misprediction to improve the quality of pseudo-labels. Specifically, we design a sequential class-aware smoothing module that incorporates token class frequency information to regularize the model and prevent it from becoming overconfident toward the head class. Meanwhile, to address the overconfidence problem intensifying throughout the semi-supervised learning processes, we introduce a dynamic regularization module to adjust the calibration strength dynamically for the coordination between the calibration and semi-supervised learning processes. Extensive experiments demonstrate the effectiveness and generality of our method, which significantly reduces annotation efforts while maintaining competitive recognition performance.
Keke Xu, Zhenghua Peng, Shuangping Huang, Yunqing Hu, Wenjie Peng
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Monocular and Generalizable Gaussian Talking Head Animation
abstract
In this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without personalized re-training. Compared with previous 3D Gaussian Splatting (3DGS) methods that requires elusive multi-view datasets or tedious personalized learning/inference, MGGtalk enables more practical and broader applications. However, in the absence of multi-view and personalized training data, the incompleteness of geometric and appearance information poses a significant challenge. To address these challenges, MGGTalk explores depth information to enhance geometric and facial symmetry characteristics to supplement both geometric and appearance features. Initially, based on the pixel-wise geometric information obtained from depth estimation, we incorporate symmetry operations and point cloud filtering techniques to ensure a complete and precise position parameter for 3DGS. Subsequently, we adopt a two-stage strategy with symmetric priors for predicting the remaining 3DGS parameters. We begin by predicting Gaussian parameters for the visible facial regions of the source image. These parameters are subsequently utilized to improve the prediction of Gaussian parameters for the non-visible regions. Extensive experiments demonstrate that MGGTalk surpasses previous state-of-the-art methods, achieving superior performance across various metrics. Project page: https://scut-mmpr.github.io/MGGTalk-Homepage/.
Shengjie Gong, Jiapeng Tang, Dongming Hu, Shuangping Huang, Tianshui Chen, Zhuoman Liu
CVPR5
2025 MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
abstract
Autonomous driving visual question answering (AD-VQA) aims to answer questions related to perception, prediction, and planning based on given driving scene images, heavily relying on the model’s spatial understanding capabilities. Prior works typically express spatial information through textual representations of coordinates, resulting in semantic gaps between visual coordinate representations and textual descriptions. This oversight hinders the accurate transmission of spatial information and increases the expressive burden. To address this, we propose a novel Marker-based Prompt learning framework (MPDrive), which represents spatial coordinates by concise visual markers, ensuring linguistic expressive consistency and enhancing the accuracy of both visual perception and spatial expression in AD-VQA. Specifically, we create marker images by employing a detection expert to overlay object regions with numerical labels, converting complex textual coordinate generation into straightforward text-based visual marker predictions. Moreover, we fuse original and marker images as scene-level features and integrate them with detection priors to derive instance-level features. By combining these features, we construct dual-granularity visual prompts that stimulate the LLM’s spatial perception capabilities. Extensive experiments on the DriveLM and CODA-LM datasets show that MPDrive achieves state-of-the-art performance, particularly in cases requiring sophisticated spatial understanding.
Wenjie Peng, Zijian Zhou 0002, Miaojing Shi, Shuangping Huang
CVPR7
2025 Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation
Gang Dai 0002, Yifan Zhang 0004, Yutao Qin, Qiangya Guo, Shuangping Huang, Shuicheng Yan
ICCV5
2025 Deep Unfolding for Task-Decomposed Image Restoration Under Diverse Degradations
Jiaxuan Cheng, Lingyu Liang, Xinchao Li, Guoxi Sun, Shuangping Huang
ICIC (3)6
2025 COS-SLAM: Coordinate Attention Semantic SLAM with Pixel-to-Line Transformer
Handong Shen, Lingyu Liang, Xiaohao Liu, Xinchao Li, Guoxi Sun, Shuangping Huang
ICIG (2)7
2025 ReplayCAD: Generative Diffusion Replay for Continual Anomaly Detection
abstract
Continual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrophic forgetting and segmentation of small anomalous regions. Existing CAD methods store image distributions or patch features to mitigate catastrophic forgetting, but they fail to preserve pixel-level detailed features for accurate segmentation. To overcome this limitation, we propose ReplayCAD, a novel diffusion-driven generative replay framework that replay high-quality historical data, thus effectively preserving pixel-level detailed features. Specifically, we compress historical data by searching for a class semantic embedding in the conditional space of the pre-trained diffusion model, which can guide the model to replay data with fine-grained pixel details, thus improving the segmentation performance. However, relying solely on semantic features results in limited spatial diversity. Hence, we further use spatial features to guide data compression, achieving precise control of sample space, thereby generating more diverse data. Our method achieves state-of-the-art performance in both classification and segmentation, with notable improvements in segmentation: 11.5% on VisA and 8.1% on MVTec. Our source code is available at https://github.com/HULEI7/ReplayCAD.
Lei Hu 0012, Zhiyong Gan, Ling Deng, Jinglin Liang 0001, Lingyu Liang, Shuangping Huang, Tianshui Chen
IJCAI6
2025 Optimal Feature Embedding for Document Large Visual Language Model
abstract
Document Large Vision Language Models excel in document-centric tasks and have become a key focus of research. Existing frameworks embed features from a lightweight, document-specific encoder into the first layer of a general-purpose Vision Language Model (VLM). However, this introduces a feature mismatch problem. VLMs typically consist of many stacked layers, with the feature hierarchy becoming increasingly abstract at higher layers. Specifically, the first-layer feature in a VLM is token-level, whereas the feature from the encoder is task-level, resulting in a mismatch. Consequently, it is crucial to identify an optimal layer within the VLM for embedding the encoder's features. Inspired by physics, we reformulate the search for the optimal embedding as a problem of finding the shortest time curve. Leveraging the properties of the shortest time curve, we theoretically derive a task-agnostic proxy score that requires only partial training and propose our searching framework, Brac4VLM. Our theoretical derivation shows that Brac4VLM reduces search time by 97.8% compared to brute-force methods. Experimental results further demonstrate that Brac4VLM identifies embedding points that closely align with the true optima. Moreover, the DocVLM with the optimal embedding position identified achieves state-of-the-art performance across various document-centric tasks. Codes: https://github.com/MaxKinny/Brac4VLM.
Fan Yang 0082, Ling Deng, Zhiyong Gan, Qisheng He, Yuanbo Fang, Xiangmin Xu 0001, Shuangping Huang, Tianshui Chen
ACM Multimedia7
2025 Order-Level Attention Similarity Across Language Models: A Latent Commonality
abstract
In this paper, we explore an important yet previously neglected question: Do context aggregation patterns across Language Models (LMs) share commonalities? While some works have investigated context aggregation or attention weights in LMs, they typically focus on individual models or attention heads, lacking a systematic analysis across multiple LMs to explore their commonalities. In contrast, we focus on the commonalities among LMs, which can deepen our understanding of LMs and even facilitate cross-model knowledge transfer. In this work, we introduce the Order-Level Attention (OLA) derived from the order-wise decomposition of Attention Rollout and reveal that the OLA at the same order across LMs exhibits significant similarities. Furthermore, we discover an implicit mapping between OLA and syntactic knowledge. Based on these two findings, we propose the Transferable OLA Adapter (TOA), a training-free cross-LM adapter transfer method. Specifically, we treat the OLA as a unified syntactic feature representation and train an adapter that takes OLA as input. Due to the similarities in OLA across LMs, the adapter generalizes to unseen LMs without requiring any parameter updates. Extensive experiments demonstrate that TOA's cross-LM generalization effectively enhances the performance of unseen LMs. Code is available at \url{https://github.com/jinglin-liang/OLAS}.
Jinglin Liang 0001, Jin Zhong 0001, Shuangping Huang, Yunqing Hu, Lixin Fan, Hanlin Gu
NeurIPS3
2025 Spatio-temporal collaborative multiple-stream transformer network for liver lesion classification on multiple-sequence magnetic resonance imaging
Shuangping Huang, Zinan Hong, Bianzhe Wu, Jinglin Liang 0001, Qinghua Huang
Eng. Appl. Artif. Intell.1
2025 Globally Correlation-Aware Hard Negative Generation
Wenjie Peng, Hongxiang Huang, Tianshui Chen, Quhui Ke, Gang Dai 0002, Shuangping Huang
Int. J. Comput. Vis.6
2025 Heterogeneous Correlation Aware Regularization for Sequential Confidence Calibration
abstract
Despite notable advancements across various tasks, deep sequence recognition models are shown to grapple with the dilemma of over-confidence, leading to unreliable predicted confidence, necessitating the need for calibration. Current efforts predominantly focus on classification model calibration, leaving the sequence recognition model calibration analysis underexplored and challenging. In this work, we discover that the primary reason for over-confidence in sequence recognition models stems from the one-hot encoding target sequence training paradigm and identify two distinct manifestations of over-confidence: perception and semantic context over-confidence. To address these challenges, we propose a heterogeneous correlation aware sequence regularization (HCSR) method that adaptively incorporates correlated sequences into training alongside the target sequence as additional supervision to regularize the probability of the target sequence from arbitrarily escalating. Specifically, a correlated sequence mining (CSM) model is designed, capable of efficiently mining heterogeneous correlated sequences, which can be flexibly customized to search for specific types of correlated sequences in demand to facilitate the calibration of corresponding types of over-confidence in the calibrating model, thereby achieving fine-grained calibration. Meanwhile, an adaptive calibration module is introduced to adaptively coordinate the optimization weights between the target sequence and correlated sequences, enabling the co-calibration among different samples. Comprehensive experiments conducted on several widely employed sequence recognition tasks demonstrate that the proposed method outperforms the current competing methods by a substantial margin.
Zhenghua Peng, Tianshui Chen, Shuangping Huang, Yunqing Hu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Neural Scene Designer: Self-Styled Semantic Image Manipulation
abstract
Maintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing methods primarily focus on the semantic control of generated content, often neglecting the critical task of preserving this consistency. In this work, we introduce the Neural Scene Designer (NSD), a novel framework that enables photo-realistic manipulation of user-specified scene regions while ensuring both semantic alignment with user intent and stylistic consistency with the surrounding environment. NSD leverages an advanced diffusion model, incorporating two parallel cross-attention mechanisms that separately process text and style information to achieve the dual objectives of semantic control and style consistency. To capture fine-grained style representations, we propose the Progressive Self-style Representational Learning (PSRL) module. This module is predicated on the intuitive premise that different regions within a single image share a consistent style, whereas regions from different images exhibit distinct styles. The PSRL module employs a style contrastive loss that encourages high similarity between representations from the same image while enforcing dissimilarity between those from different images. Furthermore, to address the lack of standardized evaluation protocols for this task, we establish a comprehensive benchmark. This benchmark includes competing algorithms, dedicated style-related metrics, and diverse datasets and settings to facilitate fair comparisons. Extensive experiments conducted on our benchmark demonstrate the effectiveness of the proposed framework.
Jianman Lin, Tianshui Chen, Chunmei Qing, Zhijing Yang, Shuangping Huang, Yuheng Ren, Liang Lin 0004
IEEE Trans. Image Process.5
2025 Enhancing Lip Dynamic Authenticity: Learning 3D Temporal Representations for Talking Head Synthesis
abstract
Audio-driven talking head synthesis aims to generate lifelike facial animations synchronized with audio. Current approaches primarily focus on lip motion information in 2D visual space for lip-audio synchronization and expressive lip dynamic, often neglecting 3D geometric motion representations of the lips that can more accurately capture lip movements in real-world scenarios. This oversight can result in suboptimal lip dynamic authenticity. In this work, we introduce a novel 3D Temporal Representation Learning (3D-TRL) algorithm that models 3D lip temporal information as latent representations and utilizes these representations as additional supervision to enhance dynamic authenticity. To achieve this, we leverage the geometric mesh constructed from the 3D Morphable Model (3DMM) as our 3D information of the lip and explore two self-supervised strategies to learn temporal representation in 3D geometric space. First, we propose a Reconstruction-Oriented 3D-TRL algorithm that reconstructs the input to obtain motion tokens in hidden space, encapsulating content while capturing richer contextual representations of the sequence. Second, we develop a Contrastive-Based 3D-TRL algorithm that utilizes contrastive learning to extract hidden 3D motion representations. This algorithm employs data augmentation strategies appropriate specifically for the 3D temporal sequences of the lips. Extensive experiments demonstrate that our approach, as a versatile and adaptable supervisory, can be integrated into various state-of-the-art network frameworks, leading to substantial enhancements in lip dynamic authenticity.
Yining Huang, Tianshui Chen, Shuangping Huang
ACM Trans. Multim. Comput. Commun. Appl.5
2024 One-DM: One-Shot Diffusion Mimicker for Handwritten Text Generation
Gang Dai 0002, Yifan Zhang 0004, Quhui Ke, Qiangya Guo, Shuangping Huang
ECCV (58)5
2024 Diffusion-Driven Data Replay: A Novel Approach to Combat Forgetting in Federated Class Continual Learning
Jinglin Liang 0001, Jin Zhong 0001, Hanlin Gu, Zhongqi Lu, Xingxing Tang, Gang Dai 0002, Shuangping Huang, Lixin Fan, Qiang Yang 0001
ECCV (30)7
2024 KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding
Shengjie Gong, Jiapeng Tang, Lingyu Liang, Yining Huang, Shuangping Huang
ECCV (56)7
2024 Document Image Dewarping Guided by 3D Geometry and Layout Priors
abstract
Document image dewarping aims to reconstruct the flat document image from distorted inputs. Previous methods often use geometric or text-line priors to guide the dewarping process. However, document images contain diverse contents, including figures, tables, or paragraph structures, image dewarping without considering the layout structure may fail to obtain global optimization. This paper proposes an encoder-decoder neural network, called DocTLNet, to achieve document image dewarping, which uses both 3D geometry and layout as constraints to refine the content details. To further enhance the layout details, a layout-aug loss is also proposed to explicitly guides the network to handle the distorted layout boundaries. Qualitative and quantitative experiments were conducted on DocUNet benchmark, and the results indicate that our DocTLNet is superior to related methods.
Lingyu Liang, Shuangping Huang
ICME3
2024 Enhancing Table Structure Recognition via Bounding Box Guidance
Lei Hu 0012, Shuangping Huang
ICPR (20)2
2024 Handwriting Trajectory Recovery Via Trajectory Transformer With Global Radical Context-Aware Module
Junxiang Lin, Zhounan Chen, Lingyu Liang, Wenjie Peng, Shuangping Huang
ICPR (20)5
2024 An Industrial Scene Text Detection with Spectral Domain Enhancement and Graph Fourier Mapping
abstract
Text detection is a task of great significance in different scenarios, which has wide applications for downstream tasks such as text recognition and text retrieval. Varieties of text detection methods have been proposed to solve this problem in natural scenes and have achieved good results. However, these methods cannot get satisfied performance in industrial scenes for various interferences caused by the industrial environment like the image noise, background material texture and low contrast. To deal with these challenges, we first propose a contour modeling algorithm based on graph Fourier transform mapping to represent arbitrary shaped text contours. A refined fast Fourier convolutional network module with ability of spectral-domain sensing is also intruduced to enhance text feature and suppress interference. Based on these two components, we construct a novel network to achieve accurate industrial scene text detection. Quantitative evaluations are conducted on benchmark datasets MPSC and IcText, experimental results show that our method obtains the state-of-the-art detection accuracy.
Wocheng Xiao, Lingyu Liang, Shuangping Huang
SMC3
2024 GNF-Net: An Adaptive Mesh Denoising Method with GCN-Based Guided Normal Filtering
abstract
Mesh Denoising has become a popular area of research, many traditional and learning-based methods have been proposed to remove noise from meshes. However, most approaches only focus on denoising meshes with low levels of noise. When the noise is high-frequency, it can be difficult to recover the original shape. In this paper, we present a mesh denoising approach that utilizes graph convolution representations to enhance the understanding of the mesh characteristics. It integrates precisely designed graphs to explore the inherent compositional structure of the mesh. When analyzing meshes under the influence of various noises, we extract information about the original features of the mesh by capturing spatial geometric features through graph convolution calculations. Our method is based on Guided Normal Filtering (GNF) to design a Graphical Representation Module (GRM) and a GCN-Based Normal Prediction Module (NPM). It can adaptively obtain the optimal guided normal vectors for noisy meshes. We have compared and analyzed the various methods to produce state-of-the-art results.
Lingyu Liang, Yutian Yang, Shuangping Huang
SMC4
2024 Contrastive representation enhancement and learning for handwritten mathematical expression recognition
Zihao Lin 0008, Gang Dai 0002, Tianshui Chen, Shuangping Huang, Jianmin Lin
Pattern Recognit. Lett.5
2023 Disentangling Writer and Character Styles for Handwriting Generation
abstract
Training machines to synthesize diverse handwritings is an intriguing task. Recently, RNN-based methods have been proposed to generate stylized online Chinese characters. However, these methods mainly focus on capturing a person's overall writing style, neglecting subtle style inconsistencies between characters written by the same person. For example, while a person's handwriting typically exhibits general uniformity (e.g., glyph slant and aspect ratios), there are still small style variations in finer details (e.g., stroke length and curvature) of characters. In light of this, we propose to disentangle the style representations at both writer and character levels from individual handwritings to synthesize realistic stylized online handwritten characters. Specifically, we present the style-disentangled Transformer (SDT), which employs two complementary contrastive objectives to extract the style commonalities of reference samples and capture the detailed style patterns of each sample, respectively. Extensive experiments on various language scripts demonstrate the effectiveness of SDT. Notably, our empirical findings reveal that the two learned style representations provide information at different frequency magnitudes, underscoring the importance of separate style extraction. Our source code is public at: https://github.com/dailenson/SDT.
Gang Dai 0002, Yifan Zhang 0004, Zhu Liang Yu, Zhuoman Liu, Shuangping Huang
CVPR7
2023 Perception and Semantic Aware Regularization for Sequential Confidence Calibration
abstract
Deep sequence recognition (DSR) models receive increasing attention due to their superior application to various applications. Most DSR models use merely the target sequences as supervision without considering other related sequences, leading to over-confidence in their predictions. The DSR models trained with label smoothing regularize labels by equally and independently smoothing each token, reallocating a small value to other tokens for mitigating overconfidence. However, they do not consider tokens/sequences correlations that may provide more effective information to regularize training and thus lead to sub-optimal performance. In this work, we find tokens/sequences with high perception and semantic correlations with the target ones contain more correlated and effective information and thus facilitate more effective regularization. To this end, we propose a Perception and Semantic aware Sequence Regularization framework, which explore perceptively and semantically correlated tokens/sequences as regularization. Specifically, we introduce a semantic context-free recognition and a language model to acquire similar sequences with high perceptive similarities and semantic correlation, respectively. Moreover, over-confidence degree varies across samples according to their difficulties. Thus, we further design an adaptive calibration intensity module to compute a difficulty score for each samples to obtain finer-grained regularization. Extensive experiments on canonical sequence recognition tasks, including scene text and speech recognition, demonstrate that our method sets novel state-of-the-art results. Code is available at https://github.com/husterpzh/PSSR.
Zhenghua Peng, Tianshui Chen, Keke Xu, Shuangping Huang
CVPR5
2023 Nash equilibria of two-round auctions
abstract
In a two-round auction, a subset of bidders is selected (probabilistically), according to their bids in the first round, for the second round, where they can increase their bids. We formalize the two-round auction model, restricting the second round to a dominant strategy incentive compatible (DSIC) auction for the selected bidders. It turns out that, however, such two-round auctions are not directly DSIC, even if the probability of each bidder being selected for the second round is monotonic to its first bid, which is surprisingly counter-intuitive. We also illustrate the necessary and sufficient conditions of two-round auctions being DSIC. Besides, we characterize the Nash equilibria for untruthful two-round auctions. One can achieve better revenue performance by setting proper probability for selecting bidders for the second round compared with single-round auctions.
Chulong Zhong, Yuyi Wang 0001, Shuangping Huang, Jin Zhong 0001
DAI4
2022 Complex Handwriting Trajectory Recovery: Evaluation Metrics and Algorithm
Zhounan Chen, Daihui Yang, Jinglin Liang 0001, Xinwu Liu, Yuyi Wang 0001, Zhenghua Peng, Shuangping Huang
ACCV (2)7
2022 Spatial Attention and Syntax Rule Enhanced Tree Decoder for Offline Handwritten Mathematical Expression Recognition
Zihao Lin 0008, Fan Yang 0082, Shuangping Huang, Xu Yang 0038, Jianmin Lin
ICFHR4
2022 AGTGAN: Unpaired Image Translation for Photographic Ancient Character Generation
abstract
The study of ancient writings has great value for archaeology and philology. Essential forms of material are photographic characters, but manual photographic character recognition is extremely time-consuming and expertise-dependent. Automatic classification is therefore greatly desired. However, the current performance is limited due to the lack of annotated data. Data generation is an inexpensive but useful solution to data scarcity. Nevertheless, the diverse glyph shapes and complex background textures of photographic ancient characters make the generation task difficult, leading to unsatisfactory results of existing methods. To this end, we propose an unsupervised generative adversarial network called AGTGAN in this paper. By explicitly modeling global and local glyph shape styles, followed by a stroke-aware texture transfer and an associate adversarial learning mechanism, our method can generate characters with diverse glyphs and realistic textures. We evaluate our method on photographic ancient character datasets, e.g., OBC306 and CSDD. Our method outperforms other state-of-the-art methods in terms of various metrics and performs much better in terms of the diversity and authenticity of generated samples. With our generated images, experiments on the largest photographic oracle bone character dataset show that our method can achieve a significant increase in classification accuracy, up to 16.34%. The source code is available at https://github.com/Hellomystery/AGTGAN.
Hongxiang Huang, Daihui Yang, Gang Dai 0002, Zhen Han 0003, Yuyi Wang 0001, Kin-Man Lam 0001, Fan Yang 0082, Shuangping Huang, Yongge Liu, Mengchao He
ACM Multimedia8
2021 A New Semi-automatic Annotation Model via Semantic Boundary Estimation for Scene Text Detection
Zhenzhou Zhuang, Zonghao Liu, Kin-Man Lam 0001, Shuangping Huang, Gang Dai 0002
ICDAR (3)4
2021 Context-Aware Selective Label Smoothing for Calibrating Sequence Recognition Model
abstract
Despite the success of deep neural network (DNN) on sequential data (i.e., scene text and speech) recognition, it suffers from the over-confidence problem mainly due to overfitting in training with the cross-entropy loss, which may make the decision-making less reliable. Confidence calibration has been recently proposed as one effective solution to this problem. Nevertheless, the majority of existing confidence calibration methods aims at non-sequential data, which is limited if directly applied to sequential data since the intrinsic contextual dependency in sequences or the class-specific statistical prior is seldom exploited. To the end, we propose a Context-Aware Selective Label Smoothing (CASLS) method for calibrating sequential data. The proposed CASLS fully leverages the contextual dependency in sequences to construct confusion matrices of contextual prediction statistics over different classes. Class-specific error rates are then used to adjust the weights of smoothing strength in order to achieve adaptive calibration. Experimental results on sequence recognition tasks, including scene text recognition and speech recognition, demonstrate that our method can achieve the state-of-the-art performance.
Shuangping Huang, Yu Luo 0007, Zhenzhou Zhuang, Jin-Gang Yu, Mengchao He, Yongpan Wang
ACM Multimedia1
2020 Two-dimensional multi-scale perceptive context for scene text recognition
Daihui Yang, Shuangping Huang, Kin-Man Lam 0001, Zhenzhou Zhuang
Neurocomputing3
2019 OBC306: A Large-Scale Oracle Bone Character Recognition Dataset
abstract
The oracle bone script from ancient China is among the world's most famous ancient writing systems. Identifying and deciphering oracle bone scripts is one of the most important topics in oracle bone study and requires a deep familiarity with the culture of ancient China. This task remains very challenging for two reasons. The first is that it is executed mainly by humans and requires a high level of experience, aptitude, and commitment. The second is due to the scarcity of domain-specific data, which hinders the advancement of automatic recognition research. A collection of well-labeled oracle-bone data is necessary to bridge the oracle bone and information processing fields; however, such a dataset has not yet been presented. Hence, in this paper, we construct a new large-scale dataset of oracle bone characters called OBC306. We also present the standard deep convolutional neural network-based evaluation for this dataset to serve as a benchmark. Through statistical and visual analyses, we describe the inherent difficulties of oracle bone recognition and propose future challenges for and extensions of oracle bone study using information processing. This dataset contains more than 300,000 character-level samples cropped from oracle-bone rubbings or images. It covers 306 glyph classes and is the largest existing raw oracle-bone character set, to the best of our knowledge. It is anticipated the publication of this dataset will facilitate the development of oracle bone research and lead to optimal algorithmic solutions.
Shuangping Huang, Haobin Wang, Yongge Liu, Xiaosong Shi
ICDAR1
2019 A New Parallel Detection-Recognition Approach for End-to-End Scene Text Extraction
abstract
In this work, we present a new conceptually simple and flexible network for accurate scene text extraction, which handle text detection and text recognition concurrently. Apart from solving feature sharing by designing a unified network, we implement end-to-end training of the whole system by a set of new optimization strategies. More importantly, this method highlights its novel parallel detection-recognition structure, which constructs a loose connection between both detection and recognition. This loose connection is embodied in the definition of overall loss and derivative back propagation for the model parameter updating, which automatically balances the contribution of the two branches to the system performance. It is different from the existing end-to-end methods where two subtasks are connected serially and thus yielding heavy dependence of the predecessor text detection task on the follow-up text recognition task and sensitivity of recognition to detection noise. In addition, a simple Mask-Rectifier mechanism is applied to easily adapt our system to incidental text recognition with arbitrary orientation and shape. Experiment results on Incidental Scene Text ICDAR2015 dataset surpass the current state-of-the-art FOTS method, as suggests the effectiveness of the proposed approach.
Zijian Zhou 0002, Zhizhong Su, Shuangping Huang
ICDAR4
2018 Focus on Scene Text Using Deep Reinforcement Learning
abstract
Scene text detection has been attracting increasing interests in recent years and a rich body of approaches has been proposed. These previous works of detecting scene text have been dominated by region proposals based approaches, which always generate too many text candidates relative to the number of ground truth bounding boxes. Only a few of those candidates are output as true predictions, and most of the other is fruitlessly involved in regression or classification predictions that consume a great amount of time and storage. Thus emerges the problem of low efficiency of generating text candidates. To address the issue, we propose a method for focusing on scene text gradually guided by an active model. The model allows an agent to take the whole image as the only region proposal in each episode when locating text and therefore significantly reduces the region proposals needed. The agent is trained by deep reinforcement strategy to learn how to estimate future returns of given states and sequentially make decisions to find scene text. Considering the characteristics of scene text, we additionally propose a flexible action scheme and a new reward scheme together with lazy punishment. The experiments on the ICDAR 2013 dataset shows that the proposed method achieve a promising performance while using region proposals as few as the ground truth bounding boxes.
Haobin Wang, Shuangping Huang
ICPR2
2018 DropRegion training of inception font network for high-performance Chinese font recognition
Shuangping Huang, Zhuoyao Zhong, Shuye Zhang, Haobin Wang
Pattern Recognit.1
2017 DeepText: A new approach for text proposal generation and text detection in natural images
abstract
In this paper, we develop a new approach called DeepText for text region proposal generation and text detection in natural images via a fully convolutional neural network (CNN). First, we propose the novel inception region proposal network (Inception-RPN), which slides an inception network with multi-scale windows over the top of convolutional feature maps and associates a set of text characteristic prior bounding boxes with each sliding position to generate high recall word region proposals. Next, we present a powerful text detection network that embeds ambiguous text category (ATC) information and multi-level region-of-interest pooling (MLRP) for text and non-text classification and accurate localization refinement. Our approach achieves an F-measure of 0.83 and 0.85 on the ICDAR 2011 and 2013 robust text detection benchmarks, outperforming previous state-of-the-art results.
Zhuoyao Zhong, Shuangping Huang
ICASSP3
2017 Robust shared feature learning for script and handwritten/machine-printed identification
Ziyong Feng, Zhaoyang Yang, Shuangping Huang, Jun Sun 0004
Pattern Recognit. Lett.4
2015 DLANet: A manifold-learning-based discriminative feature learning network for scene classification
Ziyong Feng, Dapeng Tao, Shuangping Huang
Neurocomputing4
2015 Online primal-dual learning for a data-dependent multi-kernel combination model with multiclass visual categorization applications
Shuangping Huang, Kunnan Xue
Inf. Sci.1
2014 Online heterogeneous feature fusion machines for visual recognition
Shuangping Huang, Xiaoxin Wei
Neurocomputing1
2010 Enhanced visual categorization performances by incorporation of simple features into bim features
abstract
Recent studies have demonstrated that Biologically Inspired Model features (BIM) are effective for object or scene categorization. However, according to BIM's forming mechanism, it may not hit the typical pattern due to its blind feature selection. In order to provide more informative pattern information for different visual object classes, a large number of prototypes have to be used in describing images. This leads to huge redundancy which may decrease categorization accuracy. Thus, improving BIM feature's performance by just increasing the number of prototypes is not adequate. In this paper, we propose an integrated approach to address this problem. In our approach, some simple non-biological features such as color histogram and Edge Orientation Histogram (EOH) are incorporated into BIM for discriminative image representation. Experimental results have shown that combination of BIM and simple features can improve visual categorization performances significantly.
Shuangping Huang
ICIP1
2009 A PLSA-Based Semantic Bag Generator with Application to Natural Scene Classification under Multi-instance Multi-label Learning Framework
abstract
Classifying natural scenes into semantic categories has always been a challenging task. So far, many works in this field are primarily intended for single label classification, where each scene example is represented as a single instance vector. The multi-instance multi-label (MIML) learning framework proposed by Z.H. Zhou et al. provides a new solution to the problem of scene classification in a different way. In this paper, we propose a novel scene classification method based on pLSA-based semantic bag generator and MIML learning framework. Under the framework of MIML learning, we introduce the mechanism that transfers an image into a set of instances through the pLSA-based bag generator. Experiments show that our approach achieves better classification performance comparing with the previous work.
Shuangping Huang
ICIG1