Yonghong Song

dblp:48/544 · DBLP profile ↗
← Back
76ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0001-5978-3781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 11 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 4 since 2021Systems, architecture and hardware · 8 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Where & What Anomaly? A Framework for Pose-Agnostic Anomaly Detection and Zero-Shot Semantic Classification
abstract
Industrial anomaly detection is commonly formulated as a binary classification task, providing limited semantic understanding for real-world inspection. In this work, we propose a zero-shot, pose-agnostic framework for fine-grained anomaly detection and semantic classification, capable of answering both where the anomaly is and what it is without defect-specific training data. Our approach leverages a pose-agnostic 3D Gaussian Splatting reference with a SAM-guided ROI Predictor for robust alignment under varying viewpoints. To overcome the global–local granularity mismatch of vision–language models, we introduce Visual–Language Differential Alignment, which aligns visual and textual residuals in the CLIP feature space to capture subtle defect differences. Moreover, a lightweight Multimodal Large Language Model dynamically generates ROI-conditioned, anomaly-aware hypotheses, and geometric priors from Depth Anything help disambiguate texture artifacts from structural defects. On the LEGO benchmark, our method improves pose-agnostic anomaly detection by 2.4% in image-level AUROC and achieves 74.7% zero-shot fine-grained defect classification accuracy, demonstrating consistent gains over both geometry-based and large-scale vision–language baselines.
Zhaokun Huang, Yonghong Song, Heyao Zhang
ICMR2
2026 Class-Missing Semi-supervised document key information extraction via synergistic refinement estimation
abstract
Current methods for document key information extraction (DKIE) rely heavily on labeled data with high annotation costs. To mitigate this issue, the semi-supervised learning (SSL) paradigm, which utilizes unlabeled document samples, has gained broad attention in DKIE. However, existing SSL methods require labeled and unlabeled data to share an identical label space, which is impractical in many DKIE tasks (i.e., some unlabeled samples do not belong to any known classes in the labeled set). In this paper, we formulate this problem as Class-Missing Semi-supervised (CMSS) DKIE. In DKIE, unknown classes usually belong to minority and fine-grained categories, intensifying the misconnections between known and unknown classes and making CMSS more challenging. To address this issue, we propose Synergistic Refinement Estimation (SRE), a progressive prototype estimation scheme that alleviates the unknown classes bias to the majority known classes on long-tailed unlabeled data. Furthermore, dynamic threshold hash rectification and structural calibration mechanisms are proposed to correct connections between fine-grained classes. Extensive experimental results demonstrate that SRE surpasses existing state-of-the-art methods on several DKIE benchmarks. Code is available at https://github.com/anonymoulink/SRE_DKIE .
Yonghong Song, Boyu Wang 0004, Yankai Cao, Jiayang Ren, Chaojie Ji, Qi Zhang 0096, Qiangqiang Mao
Inf. Process. Manag.2
2026 MoE-Morph: Lightweight Pyramid Model With Heterogeneous Mixture of Experts for Deformable Medical Image Registration
abstract
Deformable image registration aims to achieve nonlinear alignment of image spaces by estimating dense displacement fields. It is widely used in clinical tasks such as surgical planning, assisted diagnosis, and surgical navigation. While efficient, deep learning registration methods often struggle with large, complex displacements. Pyramid-based approaches address this with a coarse-to-fine strategy, but their single-feature processing can lead to error accumulation. In this paper, we introduce a dense Mixture of Experts (MoE) pyramid registration model, using routing schemes and multiple heterogeneous experts to increase the width and flexibility of feature processing within a single layer. The collaboration among heterogeneous experts enables the model to retain more precise details and maintain greater feature freedom when dealing with complex displacements. We use only deformation fields as the information transmission paradigm between different levels, with deformation field interactions between layers, which encourages the model to focus on the feature location matching process and perform registration in the correct direction. We do not utilize any complex mechanisms such as attention or ViT, keeping the model at its simplest form. The powerful deformable capability allows the model to perform volume registration directly and accurately without the need for affine registration. Experimental results show that the model achieves outstanding performance across four public datasets, including brain registration, lung registration, and abdominal multi-modal registration. The code will be published at https://github.com/Darlinglinlinlin/MOE_Morph.
Hao Lin 0009, Yonghong Song, You Su
IEEE Trans. Medical Imaging2
2025 A Multi-Degradation Dataset and a Universal Method for Image Deraining in All-Time Driving Scenes
abstract
Improving the visual quality of rainy driving scenes poses a significant challenge, as the rain streaks in the distance and the raindrops attached to nearby surfaces exhibit different characteristics under varying lighting conditions, both during the daytime and nighttime. We note that existing image deraining approaches are trained independently for specific types of rain degradation, which limits the model’s ability to adapt to dynamic driving scenes. In this paper, we introduce a new task: all-time rainy driving scene reconstruction, which aims to simultaneously address both daytime and nighttime rain degradation using a universal mix-trained model. Firstly, we construct a high-quality benchmark dataset termed RainDrive-10K, which contains four patterns: daytime rain streak, daytime raindrop, nighttime rain streak and nighttime raindrop. Furthermore, we also develop an effective Mamba-based baseline de-raining model, which employs a multi-patch progressive learning strategy to better help image restoration. Unlike existing Mamba-based methods that use fixed-scale scanning for feature extraction, we design a new multi-patch hierarchical scanning block that improves the model’s robustness to diverse rain appearances. Extensive experiments demonstrate the effectiveness of our proposed model, and show that it achieves favorable performance against state-of-the-art ones. The dataset is available at https://github.com/ZXXaaaa/MP-RainMamba.
Yonghong Song, Xinyue Su, Xiaoyun Yang
ECAI2
2025 Towards Debiased Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) aims at classifying unlabeled training data coming from old and novel classes by leveraging the information of partially labeled old classes. In this paper, we reveal that existing methods often suffer from competition between new and old classes, where the focus on learning new classes often results in a notable performance degradation on the old classes. Moreover, we delve into the reason behind this problem: the GCD classifier can be overconfident and biased towards the new class. With this insight, we propose Debiased GCD (DeGCD), a simple but effective approach that mitigates the bias caused by the overconfidence from new categories by a debiased head. Specifically, we first propose semantic calibration loss that aids the GCD classifier in debiasing by enforcing neighborhood prediction consistency with the latent representation of the debiased head. Furthermore, a debiased contrastive objective is proposed to refine the similarity matrix from the GCD classifier and the debiased classifier, suppressing the overconfidence in new classes in unlabeled data. In addition, an alignment constraint loss is designed to prevent damaging the distribution of the old categories caused by overconfidence in the new categories. Experiments on various datasets shows DeGCD achieves state-of-the-art performance and maintains a good balance between new and old classes. In addition, this method can be seamlessly adapted to other GCD methods, not only to achieve further performance gains but also to effectively balance the performance of the new class with that of the old class.
Yonghong Song
IJCAI2
2025 CoDifFu: Diffusion-Based Collaborative Perception with Efficient Heterogeneous Feature Fusion
abstract
Multi-Agent collaborative perception is currently experiencing a surge in attention as a novel approach to addressing autonomous driving challenges. Despite advances in previous efforts, challenges remain due to various dilemmas in the perception process, such as imperfect localization and collaboration heterogeneity. To tackle these issues, we propose CoDifFu, a novel diffusion-based collaborative perception framework that enhances robustness against localization uncertainty and improves efficiency in heterogeneous feature fusion. A diffusion-based detection head progressively denoises object centers through a learnable reverse process. During training, the center coordinates of objects diffuse from the ground truth to the Gaussian distribution, then the network learns to reverse the diffusion process. In the inference, the model progressively refines a set of random centers of boxes to align with the ground truth centers. Moreover, we devised a confidence-guided multi-agent communication module(CMC), utilizing the confidence map as guidance to effectively achieve complementary feature fusion of multi-agent’s features and alleviates collaboration heterogeneity. To thoroughly evaluate CoDifFu, we consider 3D object detection in both real-world and simulation scenarios. Extensive experiments demonstrate the superiority of CoDifFu and the effectiveness of all its vital components. The code will be released.
ZeYu Meng, Yonghong Song, Yuanlin Zhang 0001, ZeNan Bai, Jiayi Duan
IROS2
2025 Con2Diff: Controllable Condition Diffusion Model for Unsupervised Anomaly Detection
abstract
Recent advances in reconstruction-based anomaly detection have demonstrated strong performance, with diffusion models increasingly applied due to their superior image reconstruction capabilities. However, these methods often suffer from a lack of guidance, resulting in uncontrollable reconstructions and reduced quality, which impacts anomaly detection and localization. To address these issues, we propose a Controllable Condition Diffusion Model (Con2Diff). First, we introduce a Controllable Condition Guidance mechanism (CCG) that incorporates target image guidance during denoising, enhancing controllability and improving reconstruction quality while preserving original image details. Second, we propose a Category Domain Enhancement strategy (CDE) to fine-tune the feature extractor, reducing the gap between pre-trained and industrial image features and improving feature comparison accuracy. Finally, we design a Dual-Comparison Paradigm (DCP) that integrates pixel-level and feature-level anomaly scores, addressing the limitations of single-method approaches and boosting anomaly score precision. Experiments on the MVTec AD and VisA datasets achieve state-of-the-art performance, with image AUROC scores of 99.8% and 99.0%, respectively.
Yonghong Song
ICMR2
2025 SEG-Doc: A simple yet efficient graph neural network framework for document key information extraction
Yangyang Hui, Tianyi Zhou 0006, Yonghong Song
Neurocomputing5
2025 CRMSP: A semi-supervised approach for key information extraction with Class-Rebalancing and Merged Semantic Pseudo-Labeling
Qi Zhang 0096, Yonghong Song, Yangyang Hui
Neurocomputing2
2025 Learning the Beneficial, Forgetting the Harmful: High generalization reinforcement learning with in evolving representations
Yonghong Song, Guojiao Lin, Jiayi Duan, Shuaitao Li
Neurocomputing2
2025 Semi-Supervised Change Detection With Boundary Refinement Teacher
abstract
High-quality pseudo-labels are critical for guiding learning in semi-supervised change detection (SSCD). Recently, many SSCD methods based on consistency regularization (CR) have achieved advanced performance. These methods typically generate pseudo-labels by setting a fixed threshold. However, this strategy struggles to generate pseudo-labels with high-quality boundary details. To this end, we propose a novel SSCD method boundary refinement teacher (BRT) to enhance the boundary quality of pseudo-labels. A bi-temporal image boundary refinement (BIBR) module is designed to uncover boundary details in unlabeled images at first. BIBR explores the boundary characteristics of pseudo-label by extracting the boundary blocks from its change map and re-delineates the boundary decision points with their magnified views. Then, a stable teacher parameter update (STPU) module is devised to sustain the semi-supervised learning state steady, avoiding frequent updates to the teacher model parameters. These stable updates to teacher model parameters provide continuous and high-quality guidance, preventing pseudo-label fluctuations from disrupting the student model’s acquisition of new knowledge. Extensive experiments are conducted on three commonly used change detection datasets, encompassing buildings and multiple categories, covering SSCD settings, boundary metrics, and detailed ablation studies. Our results show that simply enhancing the boundary quality of the pseudo-labels allows BRT to consistently deliver state-of-the-art (SOTA) performance in SSCD. The code is available at: https://github.com/yoghurts-sy/BRT.
You Su, Yonghong Song, Xiaomeng Wu, Jingqi Chen, Zehan Wen
IEEE Trans. Geosci. Remote. Sens.2
2024 Video Anomaly Detection Via Self-Supervised Learning With Frame Interval and Rotation Prediction
abstract
Video Anomaly Detection (VAD) presents a substantial challenge in the field of computer vision. The emergence of self-supervised learning has played a crucial role in tackling this challenge, with the design of self-supervised pretext tasks proving to be an exceptionally effective method. However, how to imbue the devised pretext task with a more comprehensive understanding of video content is still a problem worth exploring. In this paper, we contribute by introducing two object-level self-supervised tasks tailored for mining temporal and spatial information in videos, subsequently applied to VAD. The self-supervised pretext tasks we have formulated are as follows: (i) Frame Rotation Prediction and (ii) Frame Interval Prediction. These two pretext tasks Focus on handling abnormality of appearance and action respectively. Our approach follows an end-to-end methodology that does not depend on pre-trained models, skeleton data, or optical flow information. In our experimental evaluations, our model has demonstrated superior performance, outperforming established competitors on two publicly available benchmarks. In particular, we achieved a micro-AUROC of 86.9 on the ShanghaiTech dataset.
Ke Jia, Yonghong Song, Xiaomeng Wu, You Su
ICME2
2024 Effective Representation Learning is More Effective in Reinforcement Learning than You Think
abstract
In reinforcement learning (RL), learning directly from pixels, is commonly known as vision-based RL. Effective state representations are crucial for high performance in vision-based RL. However, in order to learn effective state representations, most current vision-based RL methods based on contrastive unsupervised learning use auxiliary tasks similar to those in computer vision, which does not guarantee the effective information interaction between representation learning and RL. To learn more efficient states, we propose a simple and effective vision-based RL method. It leverages the representations acquired through contrastive learning by the Teacher Encoder and the Student Encoder to collaboratively estimate the Q-function. This cooperative process utilizes the TD error to steer updates to the Teacher Encoder, thereby ensuring effective information exchange between representation learning and RL. We refer to this approach as Reinforcement Learning with Teacher-Student Collaboration (RLTSC). RLTSC incorporates recent advancements in contrastive unsupervised learning, endowing it with potent representation learning capabilities. It provides a robust estimate of the Q-function with minimal variance and effectively guides the Teacher Ecoder to update and acquire a more efficient representation. RLTSC substantially enhances data efficiency in vision-based RL, surpassing state-of-the-art methods on various continuous and discrete control benchmarks. Remarkably, RLTSC even outperforms RL methods based on physical state features in terms of data efficiency for continuous control benchmarks. This may enlighten us: effective representation learning is more effective in reinforcement learning than you think!
Yonghong Song
ICRA2
2024 Image Caption Method from Coarse to Fine Based On Dual Encoder-Decoder Framework
abstract
Encoders are widely used in the field of image caption, but the statements generated by the current image caption method may miss the target and the generated description statements are not appropriate enough for the image content. In order to solve the above problems, we propose a coarse-fine image caption method based on dual encoder-decoder framework, which provides a mechanism for discovering and correcting omissions and enables the model to generate a complete image description. Firstly, an image feature extractor based on global and local information is designed, which can extract global information and local information of image and obtain more abundant image representation. Secondly, a dual encoder-decoder framework is designed, which consists of a coarse-grained encoder-decoder and a fine-grained encoder-decoder. Coarse-grained encoder-decoder requires only the original image features as input, which is processed by transformer to produce a coarse text description. In addition, an image feature auto-enhancement module is proposed to detect missing objects in coarse text and enhance their feature expression. Finally, the fine-grained encoder-decoder uses both the image feature and the coarse text caption as input, and generates the final fine-grained caption after multi-modal information fusion. Experimental results on MSCOCO datasets show that our proposed method outperforms previous image caption methods and achieves a performance of 39.7 BLEU-4 score and 121.6 CIDEr-A score.
Zefeng Li, Yuehu Liu, Yonghong Song
IJCNN3
2024 Semi-supervised Crowd Counting Based on Hard Pseudo-labels
abstract
The problem of crowd counting is currently an important topic of research in the field of computer vision, and the accuracy of crowd counting has been improving in recent years with the proposal and development of the method of predicting crowd density maps using deep neural networks. However, there are still many problems with existing crowd counting methods, such as the effect of occlusion between people and occlusion between people and obstacles on the counting results, and the use of crowd density maps as soft pseudo-labels is not applicable to the training of unlabeled data. To address these issues, in this paper, we propose a simple and effective crowd counting network based on the Mean-Teacher semi-supervised model using hard pseudo-labels for training on unlabeled data and using the Prediction Point Extraction Algorithm - Hungary Algorithm (PH) module to supervise labeled data. Specifically, we use a more advanced focal inverse distance transform map instead of crowd density map as soft pseudo-labels and extract hard pseudo-labels from Focal Inverse Distance Transform Map (FIDTM) using the Local Coordinate Extraction (LCE) algorithm, which provides stronger supervision in the training of unlabeled data; whereas, for labeled data, we use the PH module to classify the points in the image into foreground and background, and use cross-entropy loss to supervise, which converts the crowd counting problem into a classification problem. We have conducted extensive experiments on three commonly used crowd counting datasets, which show that our method achieves excellent performance in semi-supervised crowd counting and is close to the counting accuracy of fully supervised methods.
Yonghong Song, Tong Geng
IJCNN2
2024 PNSP: Overcoming catastrophic forgetting using Primary Null Space Projection in continual learning
Dailiang Zhou, Yonghong Song
Pattern Recognit. Lett.2
2024 DCMAI: A Dynamical Cross-Modal Alignment Interaction Framework for Document Key Information Extraction
abstract
Document key information extraction (DKIE) is a crucial topic that aims at automatically comprehending documents with complex formats and layouts (invoices, business insurance, etc.). While pre-trained approaches have shown high performance on many DKIE tasks, they suffer from three major challenges. First of all, these approaches ignore the ambiguity resulting from similar text representations before cross-modal interaction. Secondly, they do not consider cross-modal representation alignment before cross-modal interaction. Finally, self-attention layers in cross-modal interaction incur significant computing costs, making it hard to perform joint representation learning from all negative samples. To address these issues, we propose a Dynamical Cross-Modal Alignment Interaction framework (DCMAI). To be more specific, 1) A prior knowledge-guided module is designed to adaptively mine fine-grained visual information to disambiguate similar text representations. 2) A crossover alignment loss is formulated to align cross-modal representations before cross-modal interaction. 3) A hierarchical interaction sampling scheme is introduced to obtain a small but efficient subset of cross-modal negative samples, and a contrastive loss is employed to improve joint representation learning. Comprehensive experiments show that the proposed DCMAI achieves state-of-the-art performance than competitive baselines on several public downstream benchmarks. Code will be open to the public.
Yonghong Song, Yongbiao Deng, Kangkang Xie, Mingjie Xu, Haijun Ren
IEEE Trans. Circuits Syst. Video Technol.2
2024 Traffic Object Detection for Autonomous Driving Fusing LiDAR and Pseudo 4D-Radar Under Bird's-Eye-View
abstract
To ensure safe and efficient intelligent transportation systems (ITS), autonomous driving systems must have excellent abilities about object detection and environmental perception. Fusing data from multiple sensors can overcome inherent limitations of single-sensor perception in 3D object detection for autonomous driving. LiDAR is great at pinpointing objects but doesn’t capture their velocity. Radar, on the other hand, accurately measures velocity but doesn’t provide height details. Fusing Radar and LiDAR can extend the detection range and improve the detection performance for dynamic objects. Nevertheless, direct integration of two sensors for performance improvement is hindered by different data characteristics and noise distributions. To address this, we propose a novel fusion framework termed LiDAR and Pseudo 4D-Radar fusion under Bird’s-Eye-View, dubbed L4R-BEVFusion, to overcome the challenge of multi-modal fusion. Specifically, Radar is a sensor that lacks object height information, we firstly design a pseudo 4D-Radar generation process that includes the sparse to dense(S2P) module and the height completion(RHC) module to transform the original 3D-Radar feature map into pseudo 4D-Radar feature map that is more dense and has enriched height information. Secondly, the fusion framework encodes the LiDAR features and the pseudo 4D-Radar features into the same Bird’s-Eye-View(BEV) through cross-guided BEV encoder(CGBE) module. Extensive experiments show that L4R-BEVFusion achieves state-of-the-art performance (71.3% NDS and 66.7% mAP) for detecting dynamic objects(only use LiDAR and Radar) on the NuScenes dataset.
ZeYu Meng, Yonghong Song, Yuanlin Zhang 0001, YueXing Nan, ZeNan Bai
IEEE Trans. Intell. Transp. Syst.2
2024 PLBR: A Semi-Supervised Document Key Information Extraction via Pseudo-Labeling Bias Rectification
abstract
Document key information extraction (DKIE) methods often require a large number of labeled samples, imposing substantial annotation costs in practical scenarios. Fortunately, pseudo-labeling based semi-supervised learning (PSSL) algorithms provide an effective paradigm to alleviate the reliance on labeled data by leveraging unlabeled data. However, the main challenges for PSSL in DKIE tasks: 1) context dependency of DKIE results in incorrect pseudo-labels. 2) high intra-class variance and low inter-class variation on DKIE. To this end, this paper proposes a similarity matrix Pseudo-Label Bias Rectification (PLBR) semi-supervised method for DKIE tasks, which improves the quality of pseudo-labels on DKIE benchmarks with rare labels. More specifically, the Similarity Matrix Bias Rectification (SMBR) module is proposed to improve the quality of pseudo-labels, which utilizes the contextual information of DKIE data through the analysis of similarity between labeled and unlabeled data. Moreover, a dual branch adaptive alignment (DBAA) mechanism is designed to adaptively align intra-class variance and alleviate inter-class variation on DKIE benchmarks, which is composed of two adaptive alignment ways. One is the intra-class alignment branch, which is designed to adaptively align intra-class variance. The other one is the inter-class alignment branch, which is developed to adaptively alleviate inter-class variance changes on the representation level. Extensive experiment results on two benchmarks demonstrate that PLBR achieves state-of-the-art performance and its performance surpasses the previous SOTA by$2.11\% \sim 2.53\%$,$2.09\% \sim 2.49\%$F1-score on FUNSD and CORD with rare labeled samples, respectively. Code will be open to the public.
Yonghong Song, Boyu Wang 0004, Jiaohao Liu, Qi Zhang 0096
IEEE Trans. Knowl. Data Eng.2
2023 NC-WAMKD: Neighborhood Correction Weight-Adaptive Multi-Teacher Knowledge Distillation for Graph-Based Semi-Supervised Node Classification
abstract
Multi-teacher knowledge distillation can improve the performance of student networks in semi-supervised node classification tasks, but existing works ignore the importance of different teachers, using average of multiple teachers as final prediction. In addition, they rely on a large amount of labeled data, which is inconsistent with semi-supervised learning requirements. To solve these limitations, we propose a Neighborhood Correction Weight-Adaptive Multi-teacher Knowledge Distillation (NC-WAMKD) framework, which involves the knowledge distillation strategy WAMKD and the student model with Neighborhood Correction Label Propagation and Feature Transformation (NCLF). Specifically, WAMKD is designed to adaptively assign weights for multiple teachers to avoid misleading student by low-quality teachers. NCLF relies on neighborhood correction label propagation and feature transformation to validly predict unlabeled nodes. Experiments on semi-supervised node classification tasks demonstrate the effectiveness of the proposed framework. Code is available: https://github.com/para-99/NC-WAMKD.
Yonghong Song
ICASSP3
2023 Abnormal-Aware Loss and Full Distillation for Unsupervised Anomaly Detection Based on Knowledge Distillation
abstract
Knowledge Distillation based Unsupervised Anomaly Detection (KD-UAD), which aims to detect anomalies according to the differences between features of student network and teacher network on the sample, has been widely concerned by researchers. However, the strong learning ability of student networks may lead to small differences on abnormal samples, making it impossible to distinguish between anomalies and non-anomalies. To alleviate the above problem, we propose an improved KD-UAD approach to enhance the network’s ability to perceive anomalies. Firstly, we propose abnormal-aware loss (AAL), which allows student network to gain anomaly repair capabilities. AAL enhances the differences between the features extracted by the student and teacher network on abnormal samples. Secondly, we design a full distillation loss (FDL) to enhance distillation strength. FDL allows the student network to learn the feature distribution more comprehensively. Experimental results show that our method outperforms the current state-of-the-art methods on the MVTec AD and BTAD datasets.
Yonghong Song
ICIP2
2023 SGLP-Net: Sparse Graph Label Propagation Network for Weakly-Supervised Temporal Action Localization
Xiaoyao Wu, Yonghong Song
ICONIP (5)2
2023 Efficient Temporal Action Localization with Temporal Attention and Gaussian Weight
abstract
The task of temporal action localization is to recognize the action categories and meanwhile detect the start and end time of each action instance. In this paper, we propose a temporal attention and gaussian weighted anchor-free method, named TG-TAL, for temporal action localization. Rather than using anchors, our method regresses action instances directly with video frames as samples. To better address the variable length of action instance, we introduce a multi-level prediction framework with temporal attention. An additional gaussian weight branch is also defined to enhance the classification performance on low-quality temporal segments. Extensive experiments demonstrate that our method is effective on various datasets. In particular, on THUMOS14, our method outperforms one-stage temporal action localization methods and establishes a new state-of-the-art performance with an mAP(%) of 41.9 at tIoU threshold 0.5. Our method also works with two-stages methods and proposal postprocessing methods. Combined with PGCN, our method surpasses the state-of-the-art methods at tIoU threshold 0.7 and achieves a new state-of-the-art performance of 24.1 in terms of mAP(%) on THUMOS14.
Mengbo Sun, Yonghong Song, Hongda Wang
IJCNN2
2023 Handwritten Chinese Character Generation via Embedding, Decomposition and Discrimination
abstract
Handwritten Chinese characters generation (HCCG) aims to automatically generate handwritten Chinese character images with various styles, which is very important for training a high performance handwritten Chinese character recognition model. Most existing HCCG methods take Chinese characters as a whole and ignore their structures and details, resulting in the lack of fine-grained representations and monotonous styles, which further affects the generalization ability of handwritten Chinese character recognition model. In this paper, we propose a novel HCCG method, named Embedding, Decomposition and Discrimination Network (EDDNet), which decomposes a Chinese character into a component sequence to represent its internal structures and radicals, so as to achieve the diversity of writing styles through fine-grained representations of handwritten Chinese characters. First, we propose a local style embedding module (LSEM) to inject target styles into content features. Then we decompose characters into sequences of components and propose a fine-grained content discriminator (FCD) to maintain the content integrity of generated images. Extensive experiments demonstrate that EDDNet can generate handwritten Chinese character images with more perfect details and realistic writing styles. Moreover, it is easy to extend to unseen styles and unseen characters.
Yonghong Song
IJCNN2
2023 ECFNet: A Siamese Network With Fewer FPs and Fewer FNs for Change Detection of Remote-Sensing Images
abstract
High-resolution remote-sensing image change detection plays an important role in areas such as land and resources investigation, natural disaster prediction, and military strategy research. Current change detection methods often focus on extracting more discriminative features, while ignoring the information loss and imbalance problems in the process of feature fusion, which results in weakness in small change objects and edge pixels of change objects. In this letter, a simple but efficient network architecture, extraction, comparison and fusion network (ECFNet), for change detection in remote-sensing images is proposed. By constraining the number of feature channels in the fusion process by the feature comparison module (FCM), ECFNet can better utilize the fine-grained information in the multiscale feature map for result prediction, which not only improves the detection performance of small objects, but also reduces the false detection around the edge pixels of change objects. Experiments on the deeply supervised image fusion network for change detection (DSIFN-CD) test set show that ECFNet achieves state-of-the-art results with a small amount of computation.
Siyuan Zhu, Yonghong Song, Yu Zhang 0247, Yuanlin Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 Interaction in Transformer for Change Detection in VHR Remote Sensing Images
abstract
With the development of deep learning, very high-resolution remote sensing image change detection (VHRCD) methods are becoming more popular. However, the most existing change detection methods are not good at processing edge details and small target detection. To this end, in this paper, InterFormer, a bidirectional interactive framework based on transformer is proposed to find small changes and extract more accurate edge information of the change area. First, we designed an asymmetric Interaction Attention Module (IAM) to identify the edge details for the bi-temporal image. The IAM fully leverages the benefits of self-attention, performing feature fusion during feature extraction. This approach improves edge feature extraction capability and reduces the number of parameters, compared to other transformer methods. Second, we designed a global attention based feature fusion module called GFFM to enhance the detection performance of small targets. The GFFM further improves small target detection ability by augmenting the network’s selectivity to spatial information during feature fusion. The method is applicable to scenarios involving small changes and possesses enhanced edge detection capabilities. Our method outperforms state-of-the-art counterparts on three public benchmarks and has fewer parameters.
ZiJian Chen, Yonghong Song, GuoFu Li
IEEE Trans. Geosci. Remote. Sens.2
2023 Spatiotemporal Analysis of Static and Dynamic Traffic Elements From Road Scenes
abstract
Spatiotemporal analysis of road scenes is a hot research topic in the communities of computer vision and intelligent transportation systems. In this paper, we propose a new framework for spatiotemporal analysis of static and dynamic traffic elements from road scenes. In the first stage, a bottom-up analysis method for static traffic elements is proposed based on a hierarchical spatiotemporal model using hidden conditional random fields (HCRF). The bottom-level features are extracted from sub-regions in the hierarchical model, and the local and global features of the image sequence are then fully combined for spatial and temporal layers. In the second stage, a lightweight multi-stream 3DCNN network is developed for the behavior classification of dynamic traffic elements, which is composed of three parts. Firstly, a SELayer-3DCNN is designed to extract the appearance, motion and edge information from the image sequences. Secondly, the channel attention fusion strategy (CAF) is introduced to enhance the feature fusion ability. Finally, the 3D-RFB module is incorporated to expand the receptive field of the convolution kernel. The experimental results well demonstrate the effectiveness of the proposed framework.
Yaochen Li, Haochuan Hou, Zikun Dong, Yujie Zang, Yonghong Song
IEEE Trans. Intell. Transp. Syst.6
2022 SA-CenterNet: Scale Adaptive CenterNet for X-RAY Luggage Image Detection
abstract
Using deep-learing based object detection technology to detect contraband is to locate and classify it in X-RAY luggage images. It can circumvent the problem of missed and false detection caused by traditional manual security inspection. This paper proposes a novel anchor-free detector named Scale Adaptive CenterNet (SA-CenterNet), which models the contraband as its center point at multiple scales. Our method uses keypoint estimation to find center points and regresses to object’s sizes to obtain bounding boxes. SA-CenterNet introduces a scale adaptive module which uses multi-level feature maps to predict occluded contraband with the different sizes and the same center point. In additIon, we propose a feature enhancement module (FEM) to better extract the feature of small and obscure contraband. We use lightweight backbone and group convolution in detection head to reduce the amount of calculation and speed up the inference speed. Experiments demonstrate that our method achieves 95.2%AP50at 105 FPS on our dataset named UNICOMP, which surpasses the state-of-the-art methods. Our method can robustly detect contraband passing the security inspection machine in real time, saving manpower and material resources.
Yonghong Song, Yuanlin Zhang 0001, Siyuan Zhu, Hengyu Zhu
ICPR2
2022 Adaptive Sparse Self-attention for Object Detection
abstract
Object detection is a fundamental task for computer vision. The majority of prior methods employ global contextual information to enhance features representation. However, we argue that prior methods suffer from the superfluous features. To address the above mentioned problems, we explore the sparsity on object detection tasks in two dimensions. Specifically, a sparse spatial attention is proposed to capture global sparse long-range relationship into local features adaptively, where a learnable channel-wise mask is obtained to reduce the superfluous channels. Meanwhile, a sparse channel attention module is used to enhance the representation of each channel by introducing sparse global semantics. Experiments demonstrate that our proposed method outperforms comparative methods on the commonly used benchmark dataset, i.e., MS-COCO. The ablative experiments show that the sparsity the effectiveness in feature extraction and bounding boxes selection.
Mingjie Xu, Yonghong Song, Kangkang Xie, Jiaxi Mu
IJCNN2
2022 SiamPolar: Semi-supervised realtime video object segmentation with polar representation
Yaochen Li, Yuhui Hong, Yonghong Song
Neurocomputing3
2021 Feature Separation GAN for Cross View Gait Recognition
Chongdong Huang, Yonghong Song, Yuanlin Zhang 0001
ICIG (1)2
2021 Multi-view Gait Recognition by Inception-Encoder and CL-GEI
Chongdong Huang, Yonghong Song
ICIG (2)2
2020 Gait Recognition Based on 3D Skeleton Data and Graph Convolutional Network
abstract
Gait recognition is a hot topic in the field of biometrics because of its unique advantages such as non-contact and long distance. The appearance-based gait recognition methods usually extract features from the silhouettes of human body, which are easy to be affected by factors such as clothing and carrying objects. Although the model-based methods can effectively reduce the influence of appearance factors, it has high computational complexity. Therefore, this paper proposes a gait recognition method based on the 3D skeleton data and graph convolutional network. The 3D skeleton data is robust to the change of view. In this paper, we extract 3D joint feature and 3D bone feature based on 3D skeleton data, design a dual graph convolutional network to extract corresponding gait features and fuse them at feature level. At the same time, we use a multi-loss strategy to combine center loss and softmax loss to optimize the network. Our method is evaluated on the dataset CASIA B. The experimental results show that the proposed method can achieve state-of-the-art performance, and it can effectively reduce the influence of view, clothing and other factors.
Mengge Mao, Yonghong Song
IJCB2
2020 You Ought to Look Around: Precise, Large Span Action Detection
abstract
For the action localization task, pre-defined action anchors are the cornerstone of mainstream techniques. State-of-the-art models mostly rely on a dense segmenting scheme, where anchors are sampled uniformly over the temporal domain with a predefined set of scales. However, it is not sufficient because action duration varies greatly. Therefore, it is necessary for the anchors or proposals to have a variable receptive field. In this paper, we propose a method called YOLA (You Ought to Look Around) which includes three parts: 1) a robust backbone SPN-I3D for extracting spatio-temporal features. In this part, we employ a stronger backbone I3D with SPN (Segment Pyramid Network) instead of C3D to obtain multi-scale features; 2) a simple but useful feature fusion module named LFE (Local Feature Extraction). Compared with the fully connected layer and global average pooling, our LFE model is more advantageous for network to fit and fuse features. 3) a new feature segment aligning method called TPGC (Two Pathway Graph Convolution), which allows one proposal to leverage semantic features of adjacent proposals to update its content and make sure the proposals have a variable receptive field. YOLA add only a small overhead to the baseline network, and is easy to train in an end-to-end manner, running at a speed of 1097 fps. YOLA achieves a mAP of 58.3%, outperforming all existing models including both RGB-based and two stream on THUMOS'14, and achieves competitive results on ActivityNet 1.3.
Ge Pan, Fan Yu 0004, Yonghong Song, Yuanlin Zhang 0001
ICPR4
2020 SCA Net: Sparse Channel Attention Module for Action Recognition
abstract
Channel attention has shown its great performance recently when it was incorporated into deep convolutional neural networks. However, existing methods usually require extensive computing resources due to their involuted structure, which further increase the computational burden of 3D CNNs. In this paper, a lightweight sparse channel attention (SCA) module implemented by efficient group convolution is proposed, which adopts the idea of sparse channel connection and involves much fewer parameters but brings clear performance gain. Meanwhile, to solve the lack of local channel interaction brought by group convolution, a dominant function called Aggregate-Shuffle-Diverge (ASD) is leveraged to enhance information flow over each group with no additional parameters. We also adjust the existing mainstream 3D CNNs by employing 3D convolution factorization, so as to further reduce the parameters. Our SCA module can be flexibly incorporated into most existing 3D CNNs, all of which can achieve a perfect trade-off between performance and complexity on action recognition task with factorized I3D or 3D ResNext backbone networks. The experimental results also indicate that the resulting network, namely, SCA Net can achieve outstanding performance on UCF-101 and HMDB-51 datasets.
Yonghong Song, Yuanlin Zhang 0001
ICPR2
2020 Enhanced Darknet53 Combine MLFPN Based Real-Time Defect Detection in Steel Surface
Xiao Yi, Yonghong Song, Yuanlin Zhang 0001
PRCV (1)2
2019 Enhanced EAST: Improving Network's Feature Extraction Ability and Text Complete Shape Perception
abstract
EAST [1] is a popular end to end text detector, however, it performs deficiently with long or large texts, because of limited network's feature extraction capacity and narrow receptive fields. Based on EAST, we propose an approach named Enhanced EAST. Firstly, we offer low-level feature layers more semantic information by introducing information from high levels to low levels, which reduces the information gap between different layers. Meanwhile, we utilize a two-stream large kernel convolution to increase receptive fields with reasonable computational cost, therefore, improving the network's features detection and fusion ability. In addition, we also optimize the label generation of training data and design a weighted mask for each text, which can guide the training process to enhance the network's complete shape perception of texts, thus impelling the predicted text boxes locate more accurately. In the end, we perform data equalization and augmentation in the experiments and experiment results on ICDAR 2015, MSRA-TD500 and ICDAR 2017 MLT datasets demonstrate the proposed algorithm achieves a state-of-art performance in multi-oriented scene text detection.
Yonghong Song, Yuanlin Zhang 0001
ICDAR2
2019 Spatial Mask ConvLSTM Network and Intra-Class Joint Training Method for Human Action Recognition in Video
abstract
For action recognition, attention model is widely used, but most of them lack consideration of the relationship of spatial and temporal information. We thus propose a Spatial Mask ConvLSTM Network (SM_ConvLSTM-Net) to determine the attention score of each pixel position. SM_ConvLSTM-Net is used to combine the information of space and time for getting more precise spatial mask, which has a long receptive field in time domain. Furthermore, to combine the connection of different samples from same category, a novel training method called intra-class joint training method is proposed to make network extract the common characteristics related to actions of the same class in different background. Extensive experiments illustrate the effectiveness of our method and our method significantly outperforms the baseline C3D network on UCF101 and HMDB51. Moreover, our approach achieves the best performance on UCF101 and a compared result on HMDB51 in comparison to some state-of-the-art approaches with RGB input.
Jingjun Chen, Yonghong Song, Yuanlin Zhang 0001
ICME2
2019 Graph Convolutional LSTM Model for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has made impressive progress these years. Yet few methods consider spatial configuration of joints and temporal correlation meanwhile as a unity. To model action sequences in a way which regard both two dimensions, a Graph Convolutional Long Short Term Memory Networks (GC-LSTM) model is proposed in this paper, which automatically learns spatiotemporal features to model the action. Our model introduces the GCN operation into conventional RNN unit including graph convolution at each time step for input-to-state and state-to-state transition. Plenty of experiment analyses show that the proposed GC-LSTM model strives (1) to focus more on discriminative parts at discriminative frames and (2) to be insensitive to the redundant parts which are irrelevant for recognition. Moreover, several methods are compared with ours on two publicly available datasets and experimental results demonstrate that our model achieves the state-of-the-art performance.
Yonghong Song, Yuanlin Zhang 0001
ICME2
2019 Multi-view gait recognition using NMF and 2DLDA
Yonghong Song, Yuanlin Zhang 0001
Multim. Tools Appl.2
2018 A Bidirectional Information Aggregation Architecture for Scene Text Detection
abstract
TextBoxes[1] is one of the most advanced text detection method in both aspects of accuracy and efficiency, but it is still not very sensitive to the small text in natural scenes and often can not localize text regions precisely. To tackle these problems, we first present a Bidirectional Information Aggregation (BIA) architecture by effectively aggregating multi-scale feature maps to enhance local details and strengthen context information, making the detector not only work reliably on multi-scale text, especially the small text, but also predict more precise boxes for texts. This architecture also results in a single classifier network, which allows our model to be trained much faster and easily with better generalization power. Then, we propose to use multiple symmetrical feature maps for feature extraction in the test stages for further improving the performance on the small text. To further promote precise predicting boxes, we present a statistical grouping method that operates on the training set bounding boxes to generate aspect ratios for default boxes. Finally, our model not only outperforms the TextBoxes without much time overhead, but also provides promising performance compared to the recent state-of-theart methods on the ICDAR 2011 and 2013 database.
Yonghong Song, Yuanlin Zhang 0001
DAS2
2018 Which Part is Better: Multi-Part Competition Network for person Re-Identification
abstract
Person re-identification is a challenging task due to the background clutters, occlusion and illumination variations. In addition, the pedestrian misalignment always exists in some automatic-detection datasets. In this paper, we propose a Multi-Part Competition Network (MPCN) consisting of Multi-Part Network (MPN) and Part Competition Network (PCN), which aims to solve the misalignment problem caused by the detector errors and human pose variations. First, we construct original body parts and enlarged body parts using human pose estimation algorithm. These two kinds of body parts not only alleviate the misalignment from background and varying human pose but also solve the missing details and imprecise body parts introduced by human pose estimator. Then, we use MPN to acquire global features and two different body parts features. The components of MPN, a global branch and two part branches, are combined by ROI pooling layer. Finally, we apply PCN to achieve a tradeoff between the original body parts and the enlarged body parts and acquire discriminative part features from these two different body parts. Extensive evaluations on three widely used re-id datasets, Market-1501, CUHK03, VIPeR demonstrate that our proposed network have a competitive result compared to the state-of-the-art methods.
Yonghong Song, Yuanlin Zhang 0001
ICPR2
2018 Face Aging with Improved Invertible Conditional GANs
abstract
Due to the continuous development of GAN, vivid faces can be generated, and the use of GAN for face aging becomes a novel trend. However, many existing works for face aging require tedious pre-processing of datasets. This brings a lot of computational burden and limits the application of face aging. In order to solve these problems, a face aging network is constructed using IcGAN without any data pre-processing which map a face image into personality and age vector spaces through encoders Z and Y. Different from the previous work, we make an emphasis on the preservation of both personalized and aging features. Thus, the minimize absolute reconstructing loss is proposed to optimize vector z, which can remain the personality characteristics, meanwhile preserving the pose, hairstyle and background of the input face. Additionally, we introduce a novel age vector optimization approach by classifying reconstruction loss and introduce the parameter λ which is well-balanced between large age features and subtle texture features. The experimental results demonstrate our proposed AlGAN provides better aging faces over other state-of-the-art age progression methods.
Yonghong Song, Yuanlin Zhang 0001
ICPR2
2018 A coarse-to-fine scene text detection method based on Skeleton-cut detector and Binary-Tree-Search based rectification
Yonghong Song, Yuanlin Zhang 0001
Pattern Recognit. Lett.2
2017 Gesture Recognition Using Enhanced Depth Motion Map and Static Pose Map
abstract
In this paper, we propose a gesture recognition method using Enhanced Depth Motion Map (eDMM) and Static Pose Map (SPM) from depth videos. Firstly, the eDMM is proposed to describe motions in gesture videos, which is more robust to noise than DMM. The SPM is constructed to describe static postures of a gesture, which provides complementary information for the eDMM. Then a 2-CNN architecture is used to extract features from eDMM and SPM. The extracted features are fused to form gesture feature which contains both dynamic movement information and static pose information. Finally an ANN is trained for gesture recognition. The proposed method is evaluated on the Chalearn IsoGD dataset and the NATOPS dataset. The experimental results show that the proposed method achieves higher recognition rate than the baseline method of the Chalearn IsoGD dataset and is competitive with the-state-of-art of the NATOPS dataset.
Shenghua Wei, Yonghong Song, Yuanlin Zhang 0001
FG3
2017 Information Gain Product Quantization for Image Retrieval
Jingjia Chen, Yonghong Song, Yuanlin Zhang 0001
ICIG (2)2
2017 Scene text detection based on skeleton-cut detector
abstract
As the structural information of an object can be well descripted by edge pixels, we observe that the greatest challenge for locating text edges on scene image is how to handle the edge-adhesion problem. In this paper we propose the Skeleton-cut Text Detector, which take text-specific edge cues such as a novel presentation skeleton into account to hunt text efficiently with improved recall rate. To address edge-adhesion problem, skeleton-junctions detection and elimination are performed first to cut candidate text out of the edge map. Then the candidates are verified through a two-stage classifier based on properties like concentration ratio. Finally iteratively local refinement (IRL) is applied to enhance the overlap of proposals. Experimental results on public benchmarks, ICDAR 2013 and MSRA, demonstrate that our algorithm achieves state-of-the-art performance. Moreover in severe scenarios, our proposed method shows stronger adaptability to texts by exploiting skeleton compared to conventional presentations like MSERs.
Yonghong Song, Yuanlin Zhang 0001
ICIP2
2017 Human skeleton tree recurrent neural network with joint relative motion feature for skeleton based action recognition
abstract
Recently, the recurrent neural network(RNN) has been widely used for skeleton based action recognition because of its ability to model long-term temporal dependencies automatically. However, current methods cannot accurately describe the characteristics of actions, because they only consider joint positions rather than high order features like relative motion to different joints and ignore the impact of human physical structure. In this paper, a novel high order joint relative motion feature(JRMF) and a novel human skeleton tree RNN network(HST-RNN) are proposed. Human skeleton joints structure can be represented by a tree. The JRMF for each skeleton joint consists of the relative position, velocity and acceleration to this joint of all its descendant joints. It describes the instantaneous status of the skeleton joint better than joint positions. The HST-RNN network is constructed with the same tree structure as the human skeleton joints. Each node of the tree is a Gated Recurrent Unit(GRU) and represents a skeleton joint. The outputs of its child nodes and the corresponding JRMF are concatenated and fed into each GRU. The network combines low-level features and extracts high level features from the leaf nodes to the root node in a hierarchical way according to the human physical structure. The experimental results demonstrates that the proposed HST-RNN with JRMF achieves the state-of-art performance on challenging datasets like MSR-Action3D, UT-Kinect and UTD-MHAD.
Shenghua Wei, Yonghong Song, Yuanlin Zhang 0001
ICIP2
2017 Fast document image comparison in multilingual corpus without OCR
Yuping Lin, Yingyu Li, Yonghong Song
Multim. Syst.3
2017 Multilingual corpus construction based on printed and handwritten character separation
Yuping Lin, Yonghong Song, Yingyu Li
Multim. Tools Appl.2
2017 A hierarchical recursive method for text detection in natural scene images
Yonghong Song, Yuanlin Zhang 0001, Jingmin Xin
Multim. Tools Appl.2
2016 Hand gesture recognition using view projection from point cloud
abstract
In this paper we propose a multi-view method to recognize hand gestures using point cloud. The main idea of this paper is to project point cloud into view images and hand gestures are described by extracting and fusing features in view images. The conversion of feature space increases the inner-class similarity and meanwhile reduces the inter-class similarity. The features of view images are extracted in parallel so the scale of each feature extractor can be reduced to converge easily. In our method we perform a refined hand segmentation to segment hand form background firstly. Then the segmented hand point cloud is projected into different view planes to form view images. Next we use convolutional neural networks as feature extractors to extract features of view images. The extracted view image features are fused to form the features of hand gestures. Finally a SVM is trained for hand gesture recognition. The experimental results show that our multiview method achieves higher recognition rate and more robust to the challenging rotation changes especially out-plane rotations.
Chaoyu Liang, Yonghong Song, Yuanlin Zhang 0001
ICIP2
2016 Scene text detection based on multi-scale SWT and edge filtering
abstract
This paper presents a text detection method based on multi-scale Stroke Width Transform (SWT). First, an image pyramid is built and SWT is performed on each level of the pyramid. Second, edge components are filtered using two novel features, stroke pair ratio (SPR) and edge density of a connected component (EDC). Next, the remaining edge components on each level are grouped into text lines. And these lines are projected back onto a single image and merged. Finally, candidate text lines are verified by integrating block level features and line level features. The multi-scale mechanism makes it possible to detect text defected by reflection or blurring. And the two features are proved to be both effective and efficient in filtering non-text edges. Moreover, experimental results on the ICDAR Robust Reading Competition datasets show that the proposed text detection method provides promising performance.
Yuanyuan Feng, Yonghong Song, Yuanlin Zhang 0001
ICPR2
2016 Scene text localization using edge analysis and feature pool
Yonghong Song, Yuanlin Zhang 0001
Neurocomputing2
2015 Extraction of Virtual Baselines from Distorted Document Images Using Curvilinear Projection
abstract
The baselines of a document page are a set of virtual horizontal and parallel lines, to which the printed contents of document, e.g., text lines, tables or inserted photos, are aligned. Accurate baseline extraction is of great importance in the geometric correction of curved document images. In this paper, we propose an efficient method for accurate extraction of these virtual visual cues from a curved document image. Our method comes from two basic observations that the baselines of documents do not intersect with each other and that within a narrow strip, the baselines can be well approximated by linear segments. Based upon these observations, we propose a curvilinear projection based method and model the estimation of curved baselines as a constrained sequential optimization problem. A dynamic programming algorithm is then developed to efficiently solve the problem. The proposed method can extract the complete baselines through each pixel of document images in a high accuracy. It is also scripts insensitive and highly robust to image noises, non-textual objects, image resolutions and image quality degradation like blurring and non-uniform illumination. Extensive experiments on a number of captured document images demonstrate the effectiveness of the proposed method.
Gaofeng Meng, Zuming Huang, Yonghong Song, Shiming Xiang, Chunhong Pan
ICCV3
2015 Text detection and recognition in natural scene with edge analysis
abstract
Text plays an important role in daily life because of its rich information, thus automatic text detection in natural scenes has many attractive applications. However, detecting and recognising such text is always a challenging problem. In this study, the authors propose a method which extends the widely‐used stroke width transform by two steps of edge analysis, namely candidate edge recombination and edge classification. A new method that recognises text through candidate edge recombination and candidate edge recognition is also proposed. In the step of candidate edge recombination, they use the idea of over‐segmentation and region merging. To separate text edge from background, the edge of the input image is first divided into small segments. Then, neighbour edge segments are merged, if they have similar stroke width and colour. Through this step, each character is described by one candidate boundary. In the step of boundary classification, candidate boundaries are aggregated into text chains, followed by chain classification using character‐based and chain‐based features. To recognise text, the grey image is extracted based on the location of each candidate edge after the step of candidate edge recombination. Then, histogram of gradient features and a classifier are used to recognise each character. To evaluate the effectiveness of their method, the algorithm is run on the ICDAR competition dataset and Street View Text database. The experimental results show that the proposed method provides promising performance in comparison with the existing methods.
Yonghong Song, Quan Meng, Yuanlin Zhang 0001
IET Comput. Vis.2
2015 Natural scene text detection with multi-layer segmentation and higher order conditional random field based analysis
Yonghong Song, Yuanlin Zhang 0001, Jingmin Xin
Pattern Recognit. Lett.2
2014 Real Time Fingertip Detection with Kinect Depth Image Sequences
abstract
Gesture recognition has been a research focus with the popularity of depth sensing device. In this paper, we propose a new fingertip detection method based on a novel definition of fingers. This method consists of two steps. Firstly, finger bases are detected and estimated as prior information. Secondly, the finger regions and fingertips are located. In the second module, the point cloud of hand is represented as a graph to obtain all geodesic paths originated from palm center. If one path travels through a finger base, then its terminal point is defined as a finger point. The fingertips are determined within these finger points by utilizing geodesic distances. To our knowledge, such definition has never been applied in gesture recognition before and its performance surpasses the definition of geodesic maxima. Experimental results demonstrates the effectiveness of our method even when hand is not parallel to camera. Compared with state-of-the-art approach, our method shows much less error.
Yonghong Song, Yuanlin Zhang 0001
ICPR2
2013 A Novel Multi-oriented Chinese Text Extraction Approach from Videos
abstract
Video texts contain useful high level information which contributes to video indexing and retrieval. This paper proposed a novel video text extraction method for Chinese text of any orientation. Firstly, candidate text regions are detected by a wavelet based algorithm. Secondly, horizontal, vertical, slant, curve or arc text lines are merged by color and space relationship based on arrangement structures of multi-oriented text lines in these candidate regions. And then character segmentation will automatically choose best fit strategies based on structure analysis, due to the complex structures (single, up-down, left-right, encircling) of Chinese characters. Finally, a SVM classifier eliminates false positives. The experimental results show the proposed approach is robust for Chinese texts of any orientation in videos.
Yonghong Song, Yuanlin Zhang 0001, Quan Meng
ICDAR2
2013 Natural Scene Text Detection with Multi-channel Connected Component Segmentation
abstract
Text detection attracts more and more attention these years. But natural scene text detection is still a challenge problem due to the variations of text and the complexity of the background. In this paper an efficient text detection method with multi-channel connected component segmentation is proposed. First, connected component segmentation is done using Markov Random Field with local contrasts, colors and gradients of RGB channels. Three segmentation images are obtained corresponding to the three channels. Then, non-text connected components in the three segmentation images are removed. Finally, the remaining text components in the three segmentation images are merged and then grouped into words. Experiments on the ICDAR 2003 dataset and the ICDAR2011 dataset demonstrate that this method compares favorably with the state-of-the-art methods.
Yonghong Song, Yuanlin Zhang 0001
ICDAR2
2013 Text detection in natural scene with edge analysis
abstract
Text plays an important role in daily life due to its rich information, thus automatic text detection in natural scenes has many attractive applications. However, detecting such text is a challenge problem, because of the variations of scale, font, color, lighting and shadow. In this paper, we propose a method that detects text in natural scene through two steps of edge analysis, namely candidate edge combination and edge classification. In the step of candidate edge combination, the edge of input image is divided into small segments firstly. Then neighbor edge segments are merged, when they have similar stroke width and color. Through this step, each character is described by one edge segment set. Because sole letter rarely appears in natural scene, in the step of edge classification, candidate edges are aggregated into text chain, following with chain classification based on character-based and chain-based features. In order to evaluate the effectiveness of our method, we run our algorithm on the ICDAR 2011 public database and Street View Text database. The experimental results show that the proposed method provides promising performance in comparison with existing methods.
Quan Meng, Yonghong Song, Yuanlin Zhang 0001
ICIP2
2012 A Phase-Based Approach for Caption Detection in Videos
Shu Wen, Yonghong Song, Yuanlin Zhang 0001
ACCV (2)2
2012 A Shadow Repair Approach for Kinect Depth Maps
Yonghong Song, Yuanlin Zhang 0001, Shu Wen
ACCV (4)2
2012 Text Detection in Natural Scenes with Salient Region
abstract
In this paper, we present a novel approach to detect text in natural scenes. This approach is a type of bionic method, which imitates how human beings detect text exactly and robustly. Practically, human beings follow two steps to detect text: the first step is to find salient regions in a scene and the second step is to determine whether these salient regions are text or not. Therefore, two similar steps namely salient regions computation and text localization are used in our method. In the step of salient regions computation, a set of salient features including multi-sacle contrast, modified center-surround histogram, color spatial distribution and similarity of stroke width are used to describe an image, following with computation of salient regions based on the combination of Conditional Random Fields model and above features. Because sole letter rarely appear, in the step of text localization, salient regions are segmented and the connected components are grouped into text strings based on their features such as spatial relationships, color difference and stroke width. As an elementary unit, the text string is refined by connected component analysis. We tested the effectiveness of our method on the ICDAR 2003 database. The experimental results show that the proposed method provides promising performance in comparison with existing methods.
Quan Meng, Yonghong Song
Document Analysis Systems2
2011 A Handwritten Character Extraction Algorithm for Multi-language Document Image
abstract
In this paper, we propose a novel method for extracting handwritten characters from multi-language document images, which may contain various types of characters, e.g. Chinese, English, Japanese or their mixture. Firstly, text patches in document image are segmented based on connected component analysis. Rules for merging connected components are chosen according to the results of language identification. Then features are extracted for each basic analysis unit-text patch. Genetic algorithm is applied for feature fusion and patch type classification. Finally, a Markov Random Field model is utilized as a post-processing step to further correct the misclassification of text patch type by considering the document context. Experimental results show that the proposed algorithm can apparently improve the performance of handwritten character extraction.
Yonghong Song, Guilin Xiao, Yuanlin Zhang 0001, Lei Yang 0063, Liuliu Zhao
ICDAR1
2008 Shading Extraction and Correction for Scanned Book Images
abstract
When one scans document pages from a bound book, shading artifacts are commonly occurred in the book spine area. In this letter, we propose a general-purpose method for image shading correction based on an assumption that the reflectance function of the page surface is piecewise constant and the illumination function is smooth. The proposed method is able to completely correct more general types of shading artifacts which are nonuniformly distributed along the book spine. Comparison experiments on a synthetic and a variety of real scanned book images demonstrate the feasibility and effectiveness of the proposed method.
Gaofeng Meng, Nanning Zheng 0001, Shaoyi Du, Yonghong Song, Yuanlin Zhang 0001
IEEE Signal Process. Lett.4
2007 Document Images Retrieval Based on Multiple Features Combination
abstract
Retrieving the relevant document images from a great number of digitized pages with different kinds of artificial variations and documents quality deteriorations caused by scanning and printing is a meaningful and challenging problem. We attempt to deal with this problem by combining up multiple different kinds of document features in a hybrid way. Firstly, two new kinds of document image features based on the projection histograms and crossings number histograms of an image are proposed. Secondly, the proposed two features, together with density distribution feature and local binary pattern feature, are combined in a multistage structure to develop a novel document image retrieval system. Experimental results show that the proposed novel system is very efficient and robust for retrieving different kinds of document images, even if some of them are severely degraded.
Gaofeng Meng, Nanning Zheng 0001, Yonghong Song, Yuanlin Zhang 0001
ICDAR3
2007 Circular Noises Removal from Scanned Document Images
abstract
Defects inspection and correction is an important topic in the fields of scanned documents preprocessing. In this paper, a very fast and robust algorithm is proposed for locating and removing a special kind of circular noises caused by scanning documents with punched holes. Firstly, original image is reduced according to an elaborately selected ratio. Punched holes after reduction will leave some distinctive small regions. By examining such small regions, holes noises can be fast detected and located. To diminish false detections, Hough transformation is applied to the roughly located regions to further confirm the located holes. Finally, circular noise is eliminated by fitting a bi-linear blending Coons surface which interpolates along the four edges of noisy region. Experiments on a variety of scanned documents with punched holes demonstrate the feasibility and efficiency of the proposed algorithm.
Gaofeng Meng, Nanning Zheng 0001, Yuanlin Zhang 0001, Yonghong Song
ICDAR4
2004 Processor Aware Anticipatory Prefetching in Loops
abstract
As microprocessor speeds increase, a large fraction of the execution time is often lost to cache miss penalties. This loss can be particularly severe in processors such as the UltraSPARC-IIICu which have in-order execution and block on cache misses. Such processors rely greatly on the compiler to reduce stalls and achieve high performance. This paper describes a compiler technique for software prefetching that is aware of the specific prefetch behaviors of the target processor. The implementation targets loops containing control-flow and strided or irregular memory access patterns. A two phase locality analysis, capable of handling complex subscript expressions, is used for enhanced identification of prefetch candidates. Prefetch instructions are scheduled with careful consideration of the prefetch behaviors in the target system. Compared to a previous implementation, our technique produced performance improvements of 9% on the geometric mean, and up to 44% on individual tests, in Sun’s first UltraSPARC-IIICu based SPEC CPU2000 submission [5] and has been used in all later submissions to date.
Spiros Kalogeropulos, Mahadevan Rajagopalan, Vikram Rao, Yonghong Song, Partha Tirumalai
HPCA4
2004 Applying Array Contraction to a Sequence of DOALL Loops
abstract
Efficient program execution on multiprocessor computers requires both sufficient parallelism and good data locality. Recent research found that, using a combination of loop shifting, loop fusion, and array contraction, one can reduce the memory required to execute a sequence of serial loops, thereby to improve the cache locality. This paper studies how to extend such a memory-reduction scheme to a sequence of DOALL loops, which are executed in parallel on multiprocessors. Two methods are proposed to overcome difficulties caused by loop-carried dependences. Data copy-in is performed to remove anti-dependences between different parallel threads, and computation duplication is performed to remove flow dependences. Experiments performed on a number of benchmark programs show that the proposed technique improves both cache locality and parallel execution speed for the DOALL loops. The scheme achieves an average speedup of 1.41 for 17 programs on a 4-processor SUN machine.
Yonghong Song, Zhiyuan Li 0001
ICPP1
2004 Improving Data Locality by Array Contraction
abstract
Array contraction is a program transformation which reduces array size while preserving the correct output. In this paper, we present an aggressive array-contraction technique and study its impact on memory system performance. This technique, called controlled SFC, combines loop shifting and controlled loop fusion to maximize opportunities for array contraction within a given loop nesting. A controlled fusion scheme is used to prevent overfusing loops and to avoid excessive pressure on the cache and the registers. Reducing the array size increases data reuse because of the increased average number of memory operations on the same memory addresses. Furthermore, if the data size of a loop nest fits in the cache after array contraction, then repeated references to the same variable in the loop nest generate cache hits, assuming set conflicts are eliminated successfully.
Yonghong Song, Cheng Wang 0019, Zhiyuan Li 0001
IEEE Trans. Computers1
2004 Automatic tiling of iterative stencil loops
abstract
Iterative stencil loops are used in scientific programs to implement relaxation methods for numerical simulation and signal processing. Such loops iteratively modify the same array elements over different time steps, which presents opportunities for the compiler to improve the temporal data locality through loop tiling. This article presents a compiler framework for automatic tiling of iterative stencil loops, with the objective of improving the cache performance. The article first presents a technique which allows loop tiling to satisfy data dependences in spite of the difficulty created by imperfectly nested inner loops. It does so by skewing the inner loops over the time steps and by applying a uniform skew factor to all loops at the same nesting level. Based on a memory cost analysis, the article shows that the skew factor must be minimized at every loop level in order to minimize cache misses. A graph-theoretical algorithm, which takes polynomial time, is presented to determine the minimum skew factor. Furthermore, the memory-cost analysis derives the tile size which minimizes capacity misses. Given the tile size, an efficient and general array-padding scheme is applied to remove conflict misses. Experiments were conducted on 16 test programs and preliminary results showed an average speedup of 1.58 and a maximum speedup of 5.06 across those test programs.
Zhiyuan Li 0001, Yonghong Song
ACM Trans. Program. Lang. Syst.2
2001 Data locality enhancement by memory reduction
abstract
In this paper, we propose memory reduction as a new approach to data locality enhancement. Under this approach, we use the compiler to reduce the size of the data repeatedly referenced in a collection of nested loops. Between their reuses, the data will more likely remain in higher-speed memory devices, such as the cache. Specifically, we present an optimal algorithm to combine loop shifting, loop fusion and array contraction to reduce the temporary array storage required to execute a collection of loops. When applied to 20 benchmark programs, our technique reduces the memory requirement, counting both the data and the code, by 51% on average. The transformed programs gain a speedup of 1.40 on average, due to the reduced footprint and, consequently, the improved data locality.
Yonghong Song, Cheng Wang 0019, Zhiyuan Li 0001
ICS1
2000 Unroll-and-jam for imperfectly-nested loops in DSP applications
abstract
Unroll-and-jam, combined with scalar replacement, is a wellknown technique to balance memory operations and computations within a loop body, hence improving instructionlevel parallelism.Previous work on unroll-and-jam applies to perfectly-nested loops only.However, most loop nests in DSP applications are imperfectly-nested.In this paper, we p r e s e n t a framework to unroll-and-jam imperfectlynested loops.We d e v elop a graph-based algorithm to determine the maximum legal unroll factor.A simple heuristic is applied to compute a pro table unroll factor.Compared with a straightforward approach, which applies strip-mining, loop distribution and loop unrolling in order, our scheme is more eÆcient and allows a potentially larger legal unroll factor.To study the eectiveness of our technique, we hand-applied unroll-and-jam and scalar replacement t o s e veral typical DSP benchmarks.The results demonstrate the importance of unroll-and-jamming imperfectly-nested loops in performance improvement.Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page.To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior specific
Yonghong Song
CASES1
1999 New Tiling Techniques to Improve Cache Temporal Locality
abstract
Tiling is a well-known loop transformation to improve temporal locality of nested loops. Current compiler algorithms for tiling are limited to loops which are perfectly nested or can be transformed, in trivial ways, into a perfect nest. This paper presents a number of program transformations to enable tiling for a class of nontrivial imperfectly-nested loops such that cache locality is improved. We define a program model for such loops and develop compiler algorithms for their tiling. We propose to adopt odd-even variable duplication to break anti- and output dependences without unduly increasing the working-set size, and to adopt speculative execution to enable tiling of loops which may terminate prematurely due to, e.g. convergence tests in iterative algorithms. We have implemented these techniques in a research compiler, Panorama. Initial experiments with several benchmark programs are performed on SGI workstations based on MIPS R5K and R10K processors. Overall, the transformed programs run faster by 9% to 164%.
Yonghong Song, Zhiyuan Li 0001
PLDI1
1998 High-Level Information - An Approach for Integrating Front-End and Back-End Compilers
abstract
We propose a new universal High-Level Information (HLI) format to effectively integrate front-end and back-end compilers by passing front-end information to the back-end compiler. Importing this information into an existing back-end leverages the state-of-the-art analysis and transformation capabilities of existing front-end compilers to allow the back-end greater optimization potential than it has when relying on only locally-extracted information. A version of the HLI has been implemented in the SUIF parallelizing compiler and the GCC back-end compiler. Experimental results with the SPEC benchmarks show that HLI can provide GCC with substantially more accurate data dependence information than it can obtain on its own. Our results show that the number of dependence edges in GCC can be reduced by an average of 48% for the integer benchmark programs and an average of 54% for the floating-point benchmark programs studied, which provides greater flexibility to GCC's code scheduling pass. Even with the scheduling optimization limited to basic blocks, the use of HLI produces moderate speedups compared to using only GCC's dependence tests when the optimized programs are executed on MIPS R4600 and R10000 processors.
Sangyeun Cho, Jenn-Yuan Tsai, Yonghong Song, Bixia Zheng, Stephen J. Schwinn, Zhiyuan Li 0001, David J. Lilja, Pen-Chung Yew
ICPP3