Tonghua Su

dblp:41/1756 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0002-8869-1664ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 GUSLO: General and Unified Structured Light Optimization
abstract
Structured light (SL) 3D reconstruction captures the precise surface shape of objects, providing high-accuracy 3D data essential for industrial inspection and cultural heritage digitization. However, existing methods suffer from two key limitations: reliance on scene-specific calibration with manual parameter tuning, and optimization frameworks tailored to specific SL patterns, limiting their generalizability across varied scenarios. We propose General and Unified Structured Light Optimization (GUSLO), a novel framework addressing these issues through two coordinated innovations: (1) single-shot calibration via 2D triangulation-based interpolation that converts sparse matches into dense correspondence fields, and (2) artifact-aware photometric adaptation via explicit transfer functions, balancing generalization and color fidelity. We conduct diverse experiments covering binary, speckle, and color-coded settings. Results show that GUSLO consistently improves accuracy and cross-encoding robustness over conventional methods in challenging industrial and cultural scenarios.
Tinglei Wan, Zhongjie Wang 0003, Tonghua Su
AAAI3
2026 DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation
Yongjia Ma, Lei Fan 0007, Donglin Di, Tonghua Su
Comput. Vis. Image Underst.5
2026 OSTE: Omni-Scene Text Editing with Latent Decoupling
Tonghua Su, Fuxiang Yang, Lei Fan 0007, Donglin Di, Zhongjie Wang 0003, Xiangqian Wu 0002
Comput. Vis. Image Underst.1
2026 Adapt, project and fuse: Parameter-efficient framework for vision large language models
Yuting Bai, Tonghua Su, Zixing Bai
Neurocomputing2
2026 DUKAE: DUal-level Knowledge Accumulation and Ensemble for pre-trained model-based continual learning
Tonghua Su, Xu-Yao Zhang, Qixing Xu, Zhongjie Wang 0003
Pattern Recognit.2
2026 Learning priority-aware controllable poster layout generation
Fuxiang Yang, Wendi Hou, Lei Fan 0007, Tonghua Su, Lingxiao He, Chengzhou Li, Meng Wang 0001, Qianlong Xie, Donglin Di, Xun Yang 0001
Pattern Recognit.4
2026 Noise-aware cross attention for image manipulation localization
abstract
• A Gated Noise Extractor that dynamically captures noise features from multiple strategies. • Dual-granularity contrastive learning for more discriminative noise extraction. • Noise-domain guided fusion module t • reduce interference from irrelevant in- formation. • An efficient model with low parameter count and computational complexity. Modern image manipulation techniques have achieved visual realism that often deceives the human eye and semantic-based detectors. However, manipulation operations typically disturb the intrinsic statistical properties of images. Unlike high-level semantic content, which remains visually consistent, such disturbances manifest as anomalies in noise characteristics, including inconsistencies in sensor pattern noise, distinct high-frequency residuals, and unnatural frequency-domain artifacts introduced by resampling or synthesis. These subtle forensic cues provide more reliable evidence for manipulation localization but are often suppressed by standard RGB-domain feature extractors. Existing IML methods often rely on a single noise feature extraction strategy or treat all tampering techniques uniformly, leading to two major limitations, incomplete noise characterization and insufficient tampering-type awareness . We propose a Noise-aware Contrastive localization Network (NC-Net), which introduces two key modules. Firstly, a Gated Noise Extractor that captures mixed noise-domain patterns using a gated network combining features derived from BayarConv and Discrete Wavelet Transform (DWT) operations. This extractor is further enhanced by a dual-granularity contrastive learning strategy, which models distributional discrepancies both within images (between manipulated and authentic regions) and across images (among different manipulation types). Secondly, a Multi-Scale Fusion Module that adaptively integrates noise-domain and RGB-domain semantic features via a cross-domain attention mechanism and a top-down feature pyramid. A lightweight decoder then produces the final localization map with high precision. NC-Net enables end-to-end joint optimization of the noise extraction and RGB branches, achieving state-of-the-art performance with competitive computational overhead. Extensive experiments demonstrate its superiority over existing methods. Source code is available at https://github.com/HIT-liar/NC-Net .
Hongshi Zhang, Tonghua Su, Fuxiang Yang, Donglin Di, Yang Song 0001, Lei Fan 0007
Pattern Recognit.2
2025 DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
Donglin Di, Wenzhang Sun, Yongjia Ma, Hao Li 0030, Wei Chen 0089, Lei Fan 0007, Tonghua Su, Xun Yang 0001
ICCV8
2025 EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
abstract
Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models often focus on more common categories. In large-scale fine-grained image generation, issues of semantic information entanglement and insufficient detail in the generated images still persist. This paper attempts to introduce a concept of a "tiered embedder" in fine-grained image generation, which integrates semantic information from both super and child classes, allowing the diffusion model to better incorporate semantic information and address the issue of semantic entanglement. To address the issue of insufficient detail in fine-grained images, we introduce the concept of super-resolution during the perceptual information generation stage, enhancing the detailed features of fine-grained images through enhancement and degradation models. Furthermore, we propose an efficient ProAttention mechanism that can be effectively implemented in the diffusion model. We evaluate our method through extensive experiments on public benchmarks, demonstrating that our approach outperforms other state-of-the-art fine-tuning methods in terms of performance.
Donglin Di, Tonghua Su, Lei Fan 0007
ICME3
2025 Global-Local Aware Scene Text Editing
abstract
Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer from two major challenges: inconsistency and length-insensitivity. They often fail to maintain coherence between the edited local patch and the surrounding area, and they struggle to handle significant differences in text length before and after editing. To tackle these challenges, we propose an end-to-end framework called Global-Local Aware Scene Text Editing (GLASTE), which simultaneously incorporates high-level global contextual information along with delicate local features. Specifically, we design a global-local combination structure, joint global and local losses, and enhance text image features to ensure consistency in text style within local patches while maintaining harmony between local and global areas. Additionally, we express the text style as a vector independent of the image size, which can be transferred to target text images of various sizes. We use an affine fusion to fill target text images into the editing patch while maintaining their aspect ratio unchanged. Extensive experiments on real-world datasets validate that our GLASTE model outperforms previous methods in both quantitative metrics and qualitative results and effectively mitigates the two challenges.
Fuxiang Yang, Tonghua Su, Donglin Di, Xiangqian Wu 0002, Zhongjie Wang 0003, Lei Fan 0007
ICME2
2025 Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
abstract
Multimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. Furthermore, the disparity in data granularity and dimensionality between pathology and genomics leads to a significant modality imbalance. The high spatial resolution inherent in pathology data renders it a dominant role while overshadowing genomics in multimodal integration. In this paper, we propose a multimodal survival prediction framework that incorporates hypergraph learning to effectively capture both contextual and hierarchical details from pathology images. Moreover, it employs a modality rebalance mechanism and an interactive alignment fusion strategy to dynamically reweight the contributions of the two modalities, thereby mitigating the pathology-genomics imbalance. Quantitative and qualitative experiments are conducted on five TCGA datasets, demonstrating that our model outperforms advanced methods by over 3.4% in C-Index performance. Code: https://github.com/MCPathology/MRePath.
Mingcheng Qu, Donglin Di, Tonghua Su, Yue Gao 0002, Yang Song 0001, Lei Fan 0007
IJCAI4
2025 Spatially Gene Expression Prediction Using Dual-Scale Contrastive Learning
Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007
MICCAI (15)5
2025 Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
Mingcheng Qu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007
MICCAI (10)5
2025 Continual Learning With Knowledge Distillation: A Survey
abstract
The foremost challenge in continual learning is to mitigate catastrophic forgetting, allowing a model to retain knowledge of previous tasks while learning new tasks. Knowledge distillation (KD), a form of regularization, has gained significant attention for its ability to maintain a model's performance on previous tasks by mimicking the outputs of earlier models during the learning of new tasks, thus reducing forgetting. This article offers a comprehensive survey of continual learning methods employing KD within the realm of image classification. We provide a detailed analysis of how KD is utilized in continual learning methods, categorizing its application into three distinct paradigms. Besides, we classify these methods based on the type of knowledge source used and thoroughly examine how KD consolidates memory in continual learning from the perspective of loss functions. In addition, we have conducted extensive experiments on CIFAR-100, TinyImageNet, and ImageNet-100 across ten KD-integrated continual learning methods to analyze the role of KD in continual learning, and we have further discussed its effectiveness in other continual learning tasks. Our extensive experimental evidence demonstrates that KD plays a crucial role in mitigating forgetting in continual learning and substantiates that, when used with data replay, classification bias adversely affects the effectiveness of KD, whereas employing a separated softmax loss can significantly enhance its efficacy.
Tonghua Su, Xu-Yao Zhang, Zhongjie Wang 0003
IEEE Trans. Neural Networks Learn. Syst.2
2024 Boundary-Guided Learning for Gene Expression Prediction in Spatial Transcriptomics
abstract
Spatial transcriptomics (ST) has emerged as an advanced technology that provides spatial context to gene expression. Recently, deep learning-based methods have shown the capability to predict gene expression from WSI data using ST data. Existing approaches typically extract features from images and the neighboring regions using pretrained models, and then develop methods to fuse this information to generate the final output. However, these methods often fail to account for the cellular structure similarity, cellular density and the interactions within the microenvironment.In this paper, we propose a framework named BG-TRIPLEX, which leverages boundary information extracted from pathological images as guiding features to enhance gene expression prediction from WSIs. Specifically, our model consists of three branches: the spot, in-context and global branches. In the spot and in-context branches, boundary information, including edge and nuclei characteristics, is extracted using pretrained models. These boundary features guide the learning of cellular morphology and the characteristics of microenvironment through Multi-Head Cross-Attention. Finally, these features are integrated with global features to predict the final output.Extensive experiments were conducted on three public ST datasets. The results demonstrate that our BG-TRIPLEX consistently outperforms existing methods in terms of Pearson Correlation Coefficient (PCC). This method highlights the crucial role of boundary features in understanding the complex interactions between WSI and gene expression, offering a promising direction for future research. Codes are available at: https://github.com/WcloudC0416/BG-TRIPLEX
Mingcheng Qu, Yuncong Wu, Donglin Di, Anyang Su, Tonghua Su, Yang Song 0001, Lei Fan 0007
BIBM5
2024 Divide-Aggregate Heterogeneous Hypergraph for large-scale user intention detection
Mingcheng Qu, Xianyang Song, Donglin Di, Tonghua Su
Knowl. Based Syst.4
2024 BehaviorNet: A Fine-grained Behavior-aware Network for Dynamic Link Prediction
abstract
Dynamic link prediction has become a trending research subject because of its wide applications in the web, sociology, transportation, and bioinformatics. Currently, the prevailing approach for dynamic link prediction is based on graph neural networks, in which graph representation learning is the key to perform dynamic link prediction tasks. However, there are still great challenges because the structure of graphs evolves over time. A common approach is to represent a dynamic graph as a collection of discrete snapshots, in which information over a period is aggregated through summation or averaging. This way results in some fine-grained time-related information loss, which further leads to a certain degree of performance degradation. We conjecture that such fine-grained information is vital because it implies specific behavior patterns of nodes and edges in a snapshot. To verify this conjecture, we propose a novel fine-grained behavior-aware network (BehaviorNet) for dynamic network link prediction. Specifically, BehaviorNet adapts a transformer-based graph convolution network to capture the latent structural representations of nodes by adding edge behaviors as an additional attribute of edges. GRU is applied to learn the temporal features of given snapshots of a dynamic network by utilizing node behaviors as auxiliary information. Extensive experiments are conducted on several real-world dynamic graph datasets, and the results show significant performance gains for BehaviorNet over several state-of-the-art (SOTA) discrete dynamic link prediction baselines. Ablation study validates the effectiveness of modeling fine-grained edge and node behaviors.
Zhiying Tu, Tonghua Su, Xianzhi Wang 0001, Xiaofei Xu 0001, Zhongjie Wang 0003
ACM Trans. Web3
2023 Multimodal Scoring Model for Handwritten Chinese Essay
Tonghua Su, Hongming You, Zhongjie Wang 0003
ICDAR (1)1
2023 Self-Supervised Cross-Language Scene Text Editing
abstract
We propose and formulate the task of cross-language scene text editing, modifying the text content of a scene image into new text in another language, while preserving the scene text style and background texture. The key challenges of this task lie in the difficulty in distinguishing text and background, great distribution differences among languages, and the lack of fine-labeled real-world data. To tackle these problems, we propose a novel network named Cross-LAnguage Scene Text Editing (CLASTE), which is capable of separating the foreground text and background, as well as further decomposing the content and style of the foreground text. Our model can be trained in a self-supervised training manner on the unlabeled and multi-language data in real-world scenarios, where the source images serve as both input and ground truth. Experimental results on the Chinese-English cross-language dataset show that our proposed model can generate realistic text images, specifically, modifying English to Chinese and vice versa. Furthermore, our method is universal and can be extended to other languages such as Arabic, Korean, Japanese, Hindi, Bengali, and so on.
Fuxiang Yang, Tonghua Su, Donglin Di, Zhongjie Wang 0003
ACM Multimedia2
2023 Dual attentional transformer for video visual relation prediction
Mingcheng Qu, Ganlin Deng, Donglin Di, Jianxun Cui, Tonghua Su
Neurocomputing5
2022 FPRNet: End-to-End Full-Page Recognition Model for Handwritten Chinese Essay
Tonghua Su, Hongming You, Shuchen Liu, Zhongjie Wang 0003
ICFHR1
2022 A Deep Learning based Personalized QoE/QoS Correlation Model for Composite Services
abstract
Classical services computing tasks such as service design and service recommendation need to comprehensively consider objective Quality of Services (QoS) and subjective Quality of Experiences (QoE) of users. There are close relationships between QoS and QoE, and how to construct an accurate QoS/QoE correlation model has been a hot topic in academia for years. Particularly, it is a challenge to construct such a model for complex composite services that are composed of services and their corresponding providers from multiple domains. This is because the number of QoS parameters is huge while the number of QoE parameters is comparatively smaller, and consequently, to reasonably encode the imbalanced QoS and QoE parameters of composite services becomes challenging. In addition, different users have different concerns and personalized experiences on the same service, and the QoS/QoE correlation model should be personalized, too; however, traditional end-to-end models which simply use QoS as input and QoE as output ignore such personalized preferences of different users, thus the model accuracy is not high enough. Based on the transformer pre-trained language model, this paper mines users’ fine-grained concerns and their sentiment polarity from comments. Then, personalized preferences of users are encoded with CNN, QoS of composite services are encoded with multi-layer Bi-LSTM, and the QoS/QoE correlation is established based on the attention mechanism. In the experiments, our model achieves the highest on the accuracy of sentiment polarity prediction of user concerns, and the QoS and QoE encoded by the proposed model can accurately express the differentiated preferences of different users in a concrete composite service scenario. Potential downstream applications of the proposed QoS/QoE correlation model are comprehensively discussed.
Min Li 0051, Hanchuan Xu, Zhiying Tu, Tonghua Su, Xiaofei Xu 0001, Zhongjie Wang 0003
ICWS4
2022 Intention model based multi-round dialogue strategies for conversational AI bots
Junrui Tian, Zhiying Tu, Tonghua Su, Xiaofei Xu 0001, Zhongjie Wang 0003
Appl. Intell.4
2022 LTP: A New Active Learning Strategy for CRF-Based Named Entity Recognition
Zhiying Tu, Tonghua Su, Xiaofei Xu 0001, Zhongjie Wang 0003
Neural Process. Lett.4
2021 RTNet: An End-to-End Method for Handwritten Text Image Translation
Tonghua Su, Shuchen Liu
ICDAR (2)1
2021 SRaSLR: A Novel Social Relation Aware Service Label Recommendation Model
abstract
With the rapid development of new technologies such as cloud, edge and mobile computing, the number and diversity of available services are dramatically exploding and services have become increasingly important to people's daily work and life. As a consequence, using service label recommendation techniques to automatically categorize services plays a crucial role in many service computing tasks, such as service discovery, service composition, and service organization. There have been many service label recommendation studies that have achieved remarkable performance. However, these studies mainly focus on using the text information in service profiles to recommend labels for services while overlooking those social relations that widely exist among services. We argue that such social relations can help to obtain more precise recommendation results. In this paper, we propose a novel Social Relation aware Service Label Recommendation model called SRaSLR, which combines text information in service profiles and social network relations among services. A deep learning based model is constructed based on feature fusion of the two perspectives. We conduct extensive experiments on the real-world Programmable Web dataset, and the experiment results show that SRaSLR yields better performance than existing methods. Additionally, we discuss how service social network affects service label recommendation performance based on the experiment results.
Yeqi Zhu, Zhiying Tu, Tonghua Su, Zhongjie Wang 0001
ICWS4
2020 Empirical Study on the Skill Market of Virtual Personal Assistants (VPA)
abstract
Ever since smart speakers became popular, the functions that can help users complete a series of tasks through voice interaction are called “skills”. The market integrates all “skills” is called “skill market”. There is a serious imbalance in the distribution of hot spots and user concerns in the skill market, and the research on the distribution of user needs satisfied by skills and points of interest(POI) that users pay attention to is insufficient. User needs and POIs are contained in unstructured data, in order to analyze the distribution of user needs and POIs from unstructured data, this paper conducted an empirical study that used the BERT multi-label classification model to extract the user needs that meets the Maslow's hierarchy of needs from the skill description, and used RAKE algorithm to extract user POIs from user reviews and used knowledge graph to extract the relationships between POIs. Using the analysis results of the extracted data, the paper gives suggestions related to the development direction and POIs that should pay attention to in development for skill developers.
Tonghua Su, Zhiying Tu, Zhongjie Wang 0003
ICSS2
2019 HITHCD-2018: Handwritten Chinese Character Database of 21K-Category
abstract
Current state of handwritten Chinese character recognition (HCCR) conducted on well-confined character set, far from meeting industrial requirements. The paper describes the creation of a large-scale handwritten Chinese character database. Constructing the database is an effort to scale up Chinese handwritten character classification task to cover the full list of GBK character set specification. It consists of 21-thousand Chinese character categories and 20-million character images, larger than previous databases both in scale and diversity. We present solutions to the challenges of collecting and annotating such large-scale handwritten character samples. We elaborately design the sampling strategy, extract salient signals in a systematic way, annotate the tremendous characters through three distinct stages. Experiments are conducted the generalization to other handwritten character databases and our database demonstrates great values. Surely, its scale opens unprecedented opportunities both in evaluation of character recognition algorithms and in developing new techniques.
Tonghua Su, Lijuan Yu
ICDAR1
2017 Propagation Based Prototype Prediction
abstract
The prediction phase is used to interact with end users, so its response speed is critical for a good user experience to large category recognition tasks. This paper presents a novel and fast algorithm for prototype prediction which may solve the current computing challenges in character input applications on smart terminals. We construct a social network for prototypes and their pair-wise connections. Such "prototype network" falls into "scale-free network", emerging "small-world effect". Under reasonable conditions, there exists a small geodesic path between each node pair. This feature guarantees us to search "better" nodes following the directed edges. Unfortunately, the naive search strategy results in a computing complexity of exponential order. To convert the problem to a manageable scale, we enhance the generic breadth-first search with greedy selectivity. As a result, our method just consider a small candidate set and further propagate from those seeds. Thorough analysis both on network structure and algorithmic propagation patterns is conducted and advantages in efficiency and practicality are revealed. Finally, we evaluated the proposed algorithm on a large-scale, large-category handwritten Chinese character recognition task. Experimental results show that the proposed algorithm can be tuned with either faster prediction speed or higher prediction accuracy.
Tonghua Su, Lijun Yu
ICDAR1
2017 GMU: A Novel RNN Neuron and Its Application to Handwriting Recognition
abstract
Recurrent neural networks (RNNs) have been widely used in many sequential labeling fields. Decades of research fruits show that artificial neuron as the building blocks plays great role in its success. Different RNN neurons are proposed, such as long-short term memory (LSTM) and gated recurrent unit (GRU), and used in most applications let alone character recognition, to encode the long-term contextual dependencies. Inspired by both LSTM and GRU, a new structure named gated memory unit (GMU) is presented which carries forward their merits. GMU preserves the constant error carousels (CEC) which is devoted to enhance a smooth information flow. GMU also lends both the cell structure of LSTM and the interpolation gates of GRU. The proposed neuron is evaluated on both online English handwriting recognition and online Chinese handwriting recognition tasks in terms of parameter volumes, convergence and accuracy. The results show that GMU is of potential choice in handwriting recognition tasks.
Tonghua Su, Lijun Yu
ICDAR2
2016 Deep LSTM Networks for Online Chinese Handwriting Recognition
abstract
Currently two heavy burdens are borne in online Chinese handwriting recognition: a large-scale training data needs to be annotated with the boundaries of each character and effective features should be handcrafted by domain experts. To relieve such issues, the paper presents a novel end-to-end recognition method based on recurrent neural networks. A mixture architecture of deep bidirectional Long Short-Term Memory (LSTM) layers and feed forward subsampling layers is used to encode the long contextual history trajectories. The Connectionist Temporal Classification (CTC) objective function makes it possible to train the model without providing alignment information between input trajectories and output strings. During decoding, a modified CTC beam search algorithm is devised to integrate the linguistic constraints wisely. Our method is evaluated both on test set and competition set of CASIA-OLHWDB 2. x. Comparing with state-of-the-art methods, over 30% relative error reductions are observed on test set in terms of both correct rate and accurate rate. Even to the more challenging competition set, better results can be achieved by our method if the out-of-vocabulary problem can be ignored.
Tonghua Su
ICFHR2
2016 Novel character segmentation method for overlapped Chinese handwriting recognition based on LSTM neural networks
abstract
Overlapped handwriting recognition is widely used to input text in smart devices since it allows to write continuous characters on an size-restricted screens. How to segment the stroke sequences into characters is a crucial step before recognition. It is currently formulated as a two-class classification problem merely evaluating on the relationships between a pair of adjacent strokes. To facilitate the long contextual dependency, the paper novelly presents the problem as a sequential classification problem. Firstly each adjacent stroke pair is expressed as a feature vector. Secondly a LSTM model is learned to encode the long contextual history information from massive data. Finally the model is propagated forward to predict the labels once new samples are fed. Experiments are conducted on a public online Chinese handwriting database. The results show that the proposed method outperforms the traditional ones with about 10 percent improvement in terms of both specificity and precision.
Tonghua Su, Shukai Jia
ICPR1
2013 Exploring MPE/MWE Training for Chinese Handwriting Recognition
abstract
The HMM-based segmentation-free strategy for Chinese handwriting recognition has the merit that the model parameters can be trained with text line samples without annotation of character boundaries. However, the recognition performance has been limited to the general maximum likelihood estimation framework. In this paper, we investigate the discriminative training framework based on MPE/MWE criteria in the context of Chinese handwriting recognition for the first time. It optimizes a objective function that is a smooth measure of recognition error. Then EBW procedure is used to solve such criteria. Some key issues for robust MPE/MWE training are explored. We reveal that MPE/MWE requires more training samples, however, Chinese handwriting recognition poses severe data sparsity problem. We explore the sample synthesizing to help the training process. Experiments are conducted on Chinese handwriting database and the effectiveness of MPE/MWE training is manifested. In particular, at least 28% error reduction of recognition rates is observed in MPE/MWE training with 50 copies of synthetic sample when big ram is used to approximate the language model.
Tonghua Su, Peijun Ma, Shengchun Deng
ICDAR1
2008 Transformation-based hierarchical decision rules using genetic algorithms and its application to handwriting recognition domain
abstract
This paper describes a new approach based on Transformation-Based Learning for extracting hierarchical decision rules. Genetic algorithms are adapted to establish the context environment for transformation operation and the transformation operation can lengthen the life cycle of “good” candidate rules. The experiments are conducted on iris, wine and glass datasets with a 10-fold cross validation setup. The results show that transformation operation can improve the precision of the classifier with a smaller number of rules and generations than hierarchical decision rules. The approach also works well in touching block extraction of Chinese handwritten text.
Tonghua Su, Tianwen Zhang, Hujie Huang, Guixiang Xue
IEEE Congress on Evolutionary Computation1
2008 Task scheduling by Mean Field Annealing algorithm in grid computing
abstract
Desirable goals for grid task scheduling algorithms would shorten average delay, maximize system utilization and fulfill user constraints. In this work, an agent-based grid management infrastructure coupled with mean field annealing (MFA) scheduling algorithm has been proposed. An agent in grid utilizes a neural network algorithm to manage and schedule tasks. The Hopfield neural network is good at finding optimal solution with multi-constraints and can be fast to converge to the result. However, it is often trapped in a local minimum. Stochastic simulated annealing algorithm has an advantage in finding the optimal solution and escaping from the local minimum. Both significant characteristics of Hopfield neural network structure and stochastic simulated annealing algorithm are combined together to yield a mean field annealing scheme. A modified cooling procedure to accelerate reaching equilibrium for normalized mean field annealing has been applied to this scheme. The simulation results show that the scheduling algorithm of MFA works effectively.
Guixiang Xue, Maode Ma, Tonghua Su, Tianwen Zhang
IEEE Congress on Evolutionary Computation4