Weiqiang Wang 0001

dblp:14/6202-1 · DBLP profile ↗
← Back
113ranked-venue papers
2as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 78 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 34 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 FAST: Failure-Aware Asynchronous Search with Early Termination for Physical Design
Sihang Lei, Xueyan Zhao, Yihang Qiu, Biwei Xie, Weiqiang Wang 0001
ACM Great Lakes Symposium on VLSI5
2026 Transformer-Based 3-D Hand Pose Estimation via Bidirectional Multiscale Fusion and Learnable Anchor Guidance
abstract
Accurate 3D hand pose estimation faces inherent challenges, including self-occlusions, joint similarities, and high degrees of freedom. Although most existing CNN-based or Transformer-based methods leverage global contexts, they often fail to capture fine-grained local details and robustly handle occluded joints. To address these limitations, we propose a novel Transformer framework for depth-based 3D hand pose estimation, which incorporates two key designs: the Bidirectional Multiscale Fusion and Learnable Anchor Guidance. Firstly, we propose bidirectional multiscale fusion that sequentially propa-gates features from different encoder levels in both top-down and bottom-up directions, followed by a final aggregation of all scale features. Such an intricate feature interaction eventually enables effective joint modeling of fine-grained local details (e.g., fingertip positions) and high-level semantic context information (e.g., palm orientation). Secondly, we introduce the learnable anchor query as prior guidance to dynamically guide the decoder to localize ambiguous joints better. The learnable anchors are derived from joint-specific attention maps under 3D ground-truth supervision and then are concatenated with static grid anchors to form hybrid anchors, which effectively enable more precise 3D hand pose estimation, especially for occlusions. To further demonstrate the practicality of our framework for IoT-oriented deployment, we conduct edge-device experiments to validate its deployment feasibility. Experiments on benchmark datasets (including NYU, ICVL, MSRA, and DexYCB) demonstrate the superiority of our method over previous state-of-the-art approaches.
Ji Gan, Weiqiang Wang 0001, Feng Gao 0005, Jiaxu Leng, Haosheng Chen 0001, Xinbo Gao 0001
IEEE Internet Things J.3
2026 AiEDA: An Open-Source AI-Aided Design Library for Design-to-Vector
abstract
Recent research has demonstrated that artificial intelligence (AI) can assist electronic design automation (EDA) in improving both the quality and efficiency of chip design. But current AI for EDA (AI-EDA) infrastructures remain fragmented, lacking comprehensive solutions for the entire data pipeline from design execution to AI integration. Key challenges include fragmented flow engines that generate raw data, heterogeneous file formats for data exchange, non-standardized data extraction methods, and poorly organized data storage. This work introduces a unified open-source library for EDA (AiEDA) that addresses these issues. AiEDA integrates multiple design-to-vector data representation techniques that transform diverse chip design data into universal multi-level vector representations, establishing an AI-aided design (AAD) paradigm optimized for AI-EDA workflows. AiEDA provides complete physical design flows with programmatic data extraction and standardized Python interfaces that bridge EDA datasets and AI frameworks. Leveraging the AiEDA library, we generate iDATA, a 600GB dataset of structured data derived from 50 real chip designs (28nm), and validate its effectiveness through five representative AAD tasks spanning prediction, generation, and optimization. The code of AiEDA is publicly available at https://github.com/OSCC-Project/AiEDA, providing a foundation for future AI-EDA research.
Yihang Qiu, Zengrong Huang, Simin Tao, Hongda Zhang, Xinhua Lai, Weiqiang Wang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2026 Talk With Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
abstract
Air-writing has emerged as a promising communication modality for AR/VR and metaverse environments, enabling quiet, non-contact text input by translating finger movements into natural language. However, existing approaches typically project in-air writing onto a virtual 2D plane and assume characters are formed with a single continuous stroke–an oversimplification that neglects the rich 3D structure inherent in natural handwriting. In this work, we challenge the “single-stroke 2D” paradigm and explore the role of depth cues in enhancing air-writing recognition. To this end, we present DAAWBench, the first large-scale Depth-Aware Air-Writing dataset, featuring 8.8 million RGB-D frames annotated with 3,755 Chinese characters from the GB2312-80 Level-1 set. Our analysis reveals consistent depth variations at stroke boundaries, indicating that stroke segmentation and character recognition can benefit from depth modeling. Based on these insights, we propose DARec, a novel 3D trajectory-based recognition model that effectively leverages depth-aware priors. Extensive experiments across in-domain and out-of-domain settings, including evaluations with vision-language models (e.g., GPT-4o, Qwen-VL) and human baselines, show that DARec significantly outperforms 2D-only counterparts, achieving 87.73% accuracy versus 9.05%. Our findings demonstrate the critical importance of depth modeling in human-computer co-creative interfaces, and we will publicly release our dataset and code at https://github.com/wmeiqi/DAAWBench.
Meiqi Wu, Yuzhong Zhao, Xuchen Li 0001, Yuanqiang Cai, Jiahong Wu 0005, Weiqiang Wang 0001, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.7
2025 Adaptive Feature Aggregation for In-Air Handwritten Trajectory
abstract
As a novel human-computer interaction modality, in-air handwriting character recognition has drawn the attention of researchers. Nevertheless, existing camera-based in-air handwriting character recognition algorithms, such as those employing average pooling or CTC decoding, overlook the varying importance of features, leading to relatively low recognition accuracy. In this study, we present an adaptive feature aggregation approach. Our method can allocate weights based on the significance of features in the time dimension and perform aggregation in accordance with these weights. Multiple experiments demonstrate that our module surpasses existing methods in terms of performance. We have achieved remarkable enhancements in recognition accuracy on the AWCV-100K-UCAS2024, and WiTA datasets. Our project is open-sourced at: https://github.com/Mayo001/AFA.
Zeyu Qiu, Weiqiang Wang 0001
ICASSP2
2025 Keeping Your Eyes on the Fingertip: A Two-Stage In-Air Handwritten Recognition Method
abstract
As human-computer interaction methods diversify, in-air handwritten recognition has gained growing research interest. However, existing camera-based in-air handwritten recognition models typically overlook a crucial fact: the background information in the video, despite its large presence, is irrelevant to the recognition task. In this study, we propose a two-stage model that first extracts spatial features around the writer’s index fingertip in each video frame by a fingertip feature extraction module, then encodes these features along the temporal dimension through a text recognition module, and ultimately performs CTC decoding for prediction. Our method enables the text recognition module to focus on features near the writer’s index fingertip in the video, allowing the module to concentrate on the recognition task without being distracted by background features. Multiple experiments demonstrate superior recognition accuracy compared to state-of-the-art methods, while utilizing only about half the parameters. Our project is open-sourced at https://github.com/Mayo001/FETR.
Zeyu Qiu, Weiqiang Wang 0001
ICASSP2
2025 Decoding BatchNorm statistics via anchors pool for data-free models based on continual learning
Xiaobin Li 0006, Weiqiang Wang 0001, Guangluan Xu
Neural Comput. Appl.2
2024 MAS-NET: Mixed-Feature Attention Siamese Network for Change Detection on Remote Sensing Images
abstract
Change detection plays a crucial role in remote sensing tasks. However, current deep learning-based change detection methods suffer from issues such as misclassified pixels and unclear segmentation result on edges. To address these challenges, we propose a novel approach called Mixed-feature Attention Siamese Network (MAS-Net). MAS-Net adopts an encoder-decoder structure, where the encoder part effectively fuses image features from different time points, thereby preserving more cross-time information in the feature map. In the decoder part, we introduce the Feature Global Attention Module (FGAM) to leverage the global attention mechanism for extracting deep semantic information from the fused feature map. By incorporating these proposed modules and strategies, MAS-Net achieves fewer misclassified pixels and clearer edges in the resulting change detection maps. Experimental evaluations on the LEVIR-CD and CDD datasets demonstrate that MAS-Net outperforms state-of-the-art models by 0.15% (91.88% vs. 91.73%) and 1.3% (97.5% vs. 96.2%) in terms of F1-Score, respectively, thus establishing a solid baseline for change detection.
Xingyu Ding, Weiqiang Wang 0001
ICASSP2
2024 VS-LLM: Visual-Semantic Depression Assessment Based on LLM for Drawing Projection Test
Meiqi Wu, Yaxuan Kang, Xuchen Li 0001, Xiaotang Chen, Yunfeng Kang, Weiqiang Wang 0001, Kaiqi Huang
PRCV (9)7
2024 Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
abstract
Air-writing is a challenging task that combines the fields of computer vision and natural language processing, offering an intuitive and natural approach for human-computer interaction. However, current air-writing solutions face two primary challenges: (1) their dependency on complex sensors (e.g., Radar, EEGs and others) for capturing precise handwritten trajectories, and (2) the absence of a video-based air-writing dataset that covers a comprehensive vocabulary range. These limitations impede their practicality in various real-world scenarios, including the use on devices like iPhones and laptops. To tackle these challenges, we present the groundbreaking air-writing Chinese character video dataset (AWCV-100K), serving as a pioneering benchmark for video-based air-writing. This dataset captures handwritten trajectories in various real-world scenarios using commonly accessible RGB cameras, eliminating the need for complex sensors. AWCV-100K includes 8.8 million video frames, encompassing the complete set of 3,755 characters from the GB2312-80 level-1 set (GB1). Furthermore, we introduce our baseline approach, the video-based character recognizer (VCRec). VCRec adeptly extracts fingertip features from sparse visual cues and employs a spatio-temporal sequence module for analysis. Experimental results showcase the superior performance of VCRec compared to existing models in recognizing air-written characters, both quantitatively and qualitatively. This breakthrough paves the way for enhanced human-computer interaction in real-world contexts. Moreover, our approach leverages affordable RGB cameras, enabling its applicability in a diverse range of scenarios. The code and data examples will be made public at https://github.com/wmeiqi/AWCV.
Meiqi Wu, Kaiqi Huang, Yuanqiang Cai, Yuzhong Zhao, Weiqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 Change Detection for Remote Sensing Images based on Semantic Prototypes and Contrastive Learning
abstract
Deep networks have achieved remarkable success in addressing the problem of change detection for bi-temporal remote sensing images. While most networks focus on learning discriminating features using diverse backbone network architectures, few of them prioritize improving the performance of classification networks. In this paper, we propose a method called SPCL-BIT, which enhances the BIT method [1] by replacing its Softmax classifier with a more discriminative classifier based on semantic prototypes and contrastive learning. In the proposed classifier, we introduce two semantic prototypes to represent the change class and the unchanged class. A semantic encoder is then employed to model the relationship between these two prototypes. To obtain effective prototype representations with good inter-class separability and intra-class cohesiveness, we adopt a contrastive loss. The learned classifier, named SPCL-CD, based on semantic prototypes and contrastive learning, can be integrated with the backbone networks of most existing change detection models to enhance their performance. Comprehensive experimental results demonstrate the effectiveness of the SPCL-BIT network, and the proposed classification network, SPCL-CD, shows varying degrees of improvement when combined with the backbones of state-of-the-art networks in change detection.
Guiqin Zhao, Weiqiang Wang 0001
ICIP2
2023 Explore Faster Localization Learning For Scene Text Detection
abstract
Generally, pre-training and long-time training computation are necessary for obtaining a good-performance text detector based on deep networks. In this paper, we present a new scene text detection network (called FANet) with a Fast convergence speed and Accurate text localization. The proposed FANet is an end-to-end text detector based on transformer feature learning and normalized Fourier descriptor modeling, where the Fourier Descriptor Proposal Network and Iterative Text Decoding Network are designed to efficiently and accurately identify text proposals. Additionally, a Dense Matching Strategy and a well-designed loss function are also proposed for optimizing the network performance. Extensive experiments are carried out to demonstrate that the proposed FANet can achieve the SOTA performance with fewer training epochs and no pretraining. When we introduce additional data for pre-training, the proposed FANet can achieve SOTA performance on MSRA-TD500, CTW1500, and TotalText. The ablation experiments also verify the effectiveness of our contributions. Code is available at https://github.com/callsys/FANet.
Yuzhong Zhao, Yuanqiang Cai, Weijia Wu 0001, Weiqiang Wang 0001
ICME4
2023 FlowText: Synthesizing Realistic Scene Text Video with Optical Flow Estimation
abstract
Current video text spotting methods can achieve preferable performance, powered with sufficient labeled training data. However, labeling data manually is time-consuming and labor-intensive. To overcome this, using low-cost synthetic data is a promising alternative. This paper introduces a novel video text synthesis technique called FlowText, which utilizes optical flow estimation to synthesize a large amount of text video data at a low cost for training robust video text spotters. Unlike existing methods that focus on image-level synthesis, FlowText concentrates on synthesizing temporal information of text instances across consecutive frames using optical flow. This temporal information is crucial for accurately tracking and spotting text in video sequences, including text movement, distortion, appearance, disappearance, shelter, and blur. Experiments show that combining general detectors like TransDETR with the proposed FlowText produces remarkable results on various datasets, such as ICDAR2015video and ICDAR2013video. Code is available at https://github.com/callsys/FlowText.
Yuzhong Zhao, Weijia Wu 0001, Zhuang Li 0002, Weiqiang Wang 0001
ICME5
2023 Characters as graphs: Interpretable handwritten Chinese character recognition via Pyramid Graph Transformer
Ji Gan, Yuyan Chen, Bo Hu 0008, Jiaxu Leng, Weiqiang Wang 0001, Xinbo Gao 0001
Pattern Recognit.5
2023 Detect Arbitrary-Shaped Text via Adaptive Thresholding and Localization Quality Estimation
abstract
The two-stage scene text detection algorithms based on Mask R-CNN have achieved good performances on multiple challenging benchmarks. However, their effectiveness is degraded due to artificially setting constant thresholds and low localization quality of candidate boxes. In this paper, we present a novel scene text detection method based on Mask R-CNN and the proposed method, named LOAD, proposes adaptive threshold module and localization quality estimation module to address the above two problems. We propose two kinds of adaptive thresholds which are used for the filtering of candidate boxes and the binarization of pixels respectively. We introduce the self-attention mechanism to obtain the global information for generating the adaptive thresholds. Besides, we introduce the localization quality estimation into our model to obtain more accurate candidate boxes for subsequent segmentation. Comparative experiments are conducted on five benchmarks(ICDAR 2015, ICDAR 2017, MSRA-TD500, Total-Text and CTW1500), and the results demonstrate that the proposed method achieves the state-of-the-art performance with an F-measure of 91.0%, 78.7%, 87.4%, 90.6% and 86.0%. We also provide adequate ablation experiments to demonstrate the effectiveness of the proposed components.
Peirui Cheng, Yuzhong Zhao, Weiqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Boosting Semantic Segmentation of Aerial Images via Decoupled and Multilevel Compaction and Dispersion
abstract
Semantic segmentation is a valuable task in practical applications for aerial images. Nevertheless, the segmentation performance is unsatisfactory due to aerial images’ huge intra-class variance and inter-class similarity. To solve this problem, we propose an approach to increase the distinction between classes and compact the features of the same class. Specifically, since a single aerial image contains only a small number of categories, which is fatal for previous contrastive learning, we discard InfoNCE loss in contrastive learning and use the simple Mean Square Error (MSE) loss that does not require negative samples to decouple the dispersion and compaction operations. Besides, we set up more representative prototypes for classes and extend the prototypes to the whole dataset level, which we call image- and dataset-level prototypes. Based on the calculated prototypes, we propose Multi-level intra-class Feature Compaction (MFC) and Multi-level inter-class Feature Dispersion (MFD) to compact the features of the same class and disperse the features of different classes in the latent feature space. More importantly, some measures are proposed to ensure the two do not conflict. MFC and MFD can be applied to any existing segmentation network to improve performance significantly without increasing computational complexity during inference. Moreover, we feed the calculated multi-level prototypes directly into the classifier, thus keeping the feature extraction and classifier consistent. Results on four challenging datasets, Deepglobe, iSAID, Potsdam, and Vaihingen, demonstrate the significant effect of our method, and sufficient ablation studies verify the role of each module.
Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 GopGAN: Gradients Orthogonal Projection Generative Adversarial Network With Continual Learning
abstract
The generative adversarial networks (GANs) in continual learning suffer from catastrophic forgetting. In continual learning, GANs tend to forget about previous generation tasks and only remember the tasks they just learned. In this article, we present a novel conditional GAN, called the gradients orthogonal projection GAN (GopGAN), which updates the weights in the orthogonal subspace of the space spanned by the representations of training examples, and we also mathematically demonstrate its ability to retain the old knowledge about learned tasks in learning a new task. Furthermore, the orthogonal projection matrix for modulating gradients is mathematically derived and its iterative calculation algorithm for continual learning is given so that training examples for learned tasks do not need to be stored when learning a new task. In addition, a task-dependent latent vector construction is presented and the constructed conditional latent vectors are used as the inputs of generator in GopGAN to avoid the disappearance of orthogonal subspace of learned tasks. Extensive experiments on MNIST, EMNIST, SVHN, CIFAR10, and ImageNet-200 generation tasks show that the proposed GopGAN can effectively cope with the issue of catastrophic forgetting and stably retain learned knowledge.
Xiaobin Li 0006, Weiqiang Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 HiGAN+: Handwriting Imitation GAN with Disentangled Representations
abstract
Humans remain far better than machines at learning, where humans require fewer examples to learn new concepts and can use those concepts in richer ways. Take handwriting as an example, after learning from very limited handwriting scripts, a person can easily imagine what the handwritten texts would like with other arbitrary textual contents (even for unseen words or texts). Moreover, humans can also hallucinate to imitate calligraphic styles from just a single reference handwriting sample (that even have never seen before). Humans can do such hallucinations, perhaps because they can learn to disentangle the textual contents and calligraphic styles from handwriting images. Inspired by this, we propose a novel handwriting imitation generative adversarial network (HiGAN+) for realistic handwritten text synthesis based on disentangled representations. The proposed HiGAN+ can achieve a precise one-shot handwriting style transfer by introducing the writer-specific auxiliary loss and contextual loss, and it also attains a good global & local consistency by refining local details of synthetic handwriting images. Extensive experiments, including human evaluations, on the benchmark dataset validate our superiority in terms of visual quality, scalability, compactness, and style transferability compared with the state-of-the-art GANs for handwritten text synthesis.
Ji Gan, Weiqiang Wang 0001, Jiaxu Leng, Xinbo Gao 0001
ACM Trans. Graph.2
2022 GraFormer: Graph-oriented Transformer for 3D Pose Estimation
abstract
In 2D-to-3D pose estimation, it is important to exploit the spatial constraints of 2D joints, but it is not yet well modeled. To better model the relation of joints for 3D pose estimation, we propose an effective but simple net-work, called Graliormer11Codes:https://girhub.com/zhoawexi/GraFormer, where a novel transformer architecture is designed via embedding graph convolution layers after multi-head attention block. The proposed GraFormer is built by repeatedly stacking the GraAttention block and the ChebGConv block. The proposed GraAttention block is a new transformer block designed for processing graph-structured data, which is able to learn better features through capturing global information from all the nodes as well as the explicit adjacency structure of nodes. To model the implicit high-order connection relations among non-neighboring nodes, the ChebGConv block is introduced to exchange information between non-neighboring nodes and attain a larger receptive field. We have empirically shown the superiority of GraFormer through extensive experiments on popular public datasets. Specifically, GraFormer outperforms the state-of-the-art GraghSH [38] on the Hu-man3.6M dataset yet only contains 18% parameters of it.
Weixi Zhao, Weiqiang Wang 0001, Yunjie Tian
CVPR2
2022 MBNet: A Multi-Resolution Branch Network for Semantic Segmentation Of Ultra-High Resolution Images
abstract
Semantic segmentation of ultra-high resolution images is more challenging than ordinary images since high-resolution images need to be cropped into patches in training due to GPU memory limitation. To solve this problem, we design a multibranch structure to deal with multi-resolution inputs, called Multi-resolution Branch Network (MBNet). MBNet takes patches of various instead of only one resolution as inputs, so it can make the extracted features pure and different from each other so as to cover the complex scenes with tremendous variation. Moreover, to make full use of the multi-branch structure, we design a zoom module. Zoom module abandons the previous L2-norm before feature concatenation but combines different features according to the learned attention, which fully releases the advantages of multi-resolution. Results on two benchmark datasets show that our method improves significantly over the previous state-of-the-art methods.
Lianlei Shan, Weiqiang Wang 0001
ICASSP2
2022 Direct regression scene text detection with accuracy scoring
Peirui Cheng, Yuzhong Zhao, Yuanqiang Cai, Weiqiang Wang 0001
Neurocomputing4
2022 DBRANet: Road Extraction by Dual-Branch Encoder and Regional Attention Decoder
abstract
Although widely exploited in recent decades, road extraction is still a very significant and challenging research in the field of remote sensing image processing due to the complex background and road distribution. Among the existing CNN-based methods, U-shape architectures composed of encoders and decoders have shown their effectiveness. In this letter, we propose an improved encoder–decoder method, named DBRANet, for extracting roads from remote sensing images. In the encoding phase, we present a dual-branch network module (DBNM) to construct more effective features, thus improving the fusion feature maps of different scales. One branch utilizes the residual block, and the other branch utilizes the refined asymmetric block, which effectively increases the feature extraction capability of the backbone. In the decoding phase, considering the sinuous shape and the unbalanced distribution of roads in remote sensing images, we design a novel attention module, named the regional attention network module (RANM), to automatically learn the importance of each channel according to the regional information. Extensive experiments on several public remote sensing road data sets show that our DBRANet achieves higher segmentation [$F1$score and Intersection over Union (IoU)] and connectivity [average path length similarity (APLS)] accuracy, which verifies the effectiveness of our approach.
Sibao Chen 0001, Yu-Xin Ji, Jin Tang 0001, Bin Luo 0001, Weiqiang Wang 0001, Ke Lu 0002
IEEE Geosci. Remote. Sens. Lett.5
2022 DenseNet-Based Land Cover Classification Network With Deep Fusion
abstract
Recently, fully convolutional network (FCN)-based (Longet al., 2015) networks have made impressive success in semantic segmentation, and these approaches achieve satisfactory results in natural images. However, in the field of high-resolution remote sensing image segmentation, the accuracy has a considerable huge gap compared with that of natural images. Through the development process of semantic segmentation, we found that the key to accurate segmentation is the context. Effective networks can always obtain large contexts, which means that context is the key to one successful segmentation network. For high-resolution remote sensing images, their elements always extend to large scope and they have no clear or regular boundaries. As a result, it needs more context to correctly classify each pixel. However, the networks designed for natural images obviously do not meet this requirement, and thus, they achieve poor segmentation results for high-resolution images. Therefore, we do some targeted improvements. Based on one powerful backbone, we add two new fusions called unit fusion and cross-level fusion, respectively. Unit fusion makes the connection from the encoder part to the decoder part not only occur in the final output of each dense block but also in the middle feature layers inside one dense block. These added fusions make feature fusion in the same level more complete, which is of great significance for complex and zigzag boundary areas. As a complement for unit fusion, cross-level fusion aims to enhance the fusion of different dense blocks. Specifically, cross-level fusion learns from the internal structure of the dense block and applies the design to the whole network level. It can incorporate nonadjacent features and rapidly increase the receptive field and context, which is very effective for the segmentation of targets with very large sizes. Experiments on Deepglobe (Demiret al., 2018) show significant improvements in our work.
Lianlei Shan, Weiqiang Wang 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Class-Incremental Learning for Semantic Segmentation in Aerial Imagery via Distillation in All Aspects
abstract
Incremental learning using neural networks achieves great success in semantic segmentation but still suffers from catastrophic forgetting. In this article, we propose an effective class-incremental segmentation method without storing old data. To alleviate the issue of forgetting, we present two important modules, i.e., the deep feature distillation (DFD) module and the label mixed (LM) module. The DFD module is established to learn a good feature representation of old classes by distilling a new compact feature representation from different layers of networks. The proposed LM module first identifies the examples (pixels) of old classes with high confidences utilizing the output of old models, and then, they are combined with examples of new classes to supervise the training of new models, which can achieve a good balance between learning new classing and avoiding forgetting old ones. Our ablation studies show that the DFD module and the LM module can make the learning network obtain 6.2% and 15% performance gains [mean Intersection over Union (mIOU)], respectively. Furthermore, by introducing the supervision of output distillation loss, we compare our method with several state-of-the-art methods in the extensive experiments, and the experimental results all show that our method is significantly superior to them on the dataset of aerial images.
Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Class-Incremental Semantic Segmentation of Aerial Images via Pixel-Level Feature Generation and Task-Wise Distillation
abstract
Deep neural networks achieve significant progress in semantic segmentation but still suffer from the catastrophic forgetting problem, i.e., networks will forget old classes as they learn new ones. In this article, we propose an effective class-incremental segmentation framework without storing old data. Specifically, to alleviate the issue of catastrophic forgetting, we present two important modules, i.e., the Pixel-level Feature Generation (PFG) module, and the Task-wise Knowledge distillation (TKD) module. The PFG module is designed to constantly generate any number of features of the old classes to keep the old memory. The PFG module is the first attempt to use the generative method in class-incremental segmentation of aerial images, and it abandons the previous image generation approach but to generate pixel-level features, which is more suitable for the segmentation task. Meanwhile, the proposed TKD module is specially designed for class incremental tasks, and it only compares classes in the same learning step (task), thus avoiding the squeezing of new classes to old classes when the output is normalized (softmax), making distillation more effective. Sufficient experiments show that our method is remarkably effective and achieves more than 4.5% gains compared with state-of-the-art methods, and more than 13% compared with baselines, on all learning conditions. The ablation studies show that the PFG module and the TKD module are both indispensables. Besides, the proposed framework can be well combined with any existing class incremental learning method to achieve better performance.
Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Rethinking Object Detection in Retail Stores
abstract
The conventional standard for object detection uses a bounding box to represent each individual object instance. However, it is not practical in the industry-relevant applications in the context of warehouses due to severe occlusions among groups of instances of the same categories. In this paper, we propose a new task, i.e., simultaneously object localization and counting, abbreviated as Locount, which requires algorithms to localize groups of objects of interest with the number of instances. However, there does not exist a dataset or benchmark designed for such a task. To this end, we collect a large-scale object localization and counting dataset with rich annotations in retail stores, which consists of 50,394 images with more than 1.9 million object instances in 140 categories. Together with this dataset, we provide a new evaluation protocol and divide the training and testing subsets to fairly evaluate the performance of algorithms for Locount, developing a new benchmark for the Locount task. Moreover, we present a cascaded localization and counting network as a strong baseline, which gradually classifies and regresses the bounding boxes of objects with the predicted numbers of instances enclosed in the bounding boxes, trained in an end-to-end manner. Extensive experiments are conducted on the proposed dataset to demonstrate its significance and the analysis is provided to indicate future directions. Dataset is available at https://isrc.iscas.ac.cn/gitlab/research/locount-dataset.
Yuanqiang Cai, Longyin Wen, Libo Zhang 0001, Dawei Du, Weiqiang Wang 0001
AAAI5
2021 HiGAN: Handwriting Imitation Conditioned on Arbitrary-Length Texts and Disentangled Styles
abstract
Given limited handwriting scripts, humans can easily visualize (or imagine) what the handwritten words/texts would look like with other arbitrary textual contents. Moreover, a person also is able to imitate the handwriting styles of provided reference samples. Humans can do such hallucinations, perhaps because they can learn to disentangle the calligraphic styles and textual contents from given handwriting scripts. However, computers cannot study to do such flexible handwriting imitation with existing techniques. In this paper, we propose a novel handwriting imitation generative adversarial network (HiGAN) to mimic such hallucinations. Specifically, HiGAN can generate variable-length handwritten words/texts conditioned on arbitrary textual contents, which are unconstrained to any predefined corpus or out-of-vocabulary words. Moreover, HiGAN can flexibly control the handwriting styles of synthetic images by disentangling calligraphic styles from the reference samples. Experiments on handwriting benchmarks validate our superiority in terms of visual quality and scalability when comparing to the state-of-the-art methods for handwritten word/text synthesis. The code and pre-trained models can be found at https://github.com/ganji15/HiGAN.
Ji Gan, Weiqiang Wang 0001
AAAI2
2021 Fusing Multitask Models by Recursive Least Squares
abstract
It is easy to obtain multi-tasking models from open source platforms or various organizations. However, using these models at the same time will bring a great burden on storage and reduce computing efficiency. In this paper, we propose a transformation-based multi-task fusion method, called transformation fusion(TF), which is implemented by recursive least squares. The recursive transformation fusion not only reduces the storage burden brought by models fusion but also avoids computing the inverse matrix of high-dimensional matrix. Our multi-model fusion method can also be applied to many mainstream tasks, such as multi-task learning and offline distributed learning. Our fusion method can be applied to data-based fusion tasks as well as data-free fusion tasks. Extensive experiments demonstrate the effectiveness of our fusion method.
Xiaobin Li 0006, Lianlei Shan, Weiqiang Wang 0001
ICASSP3
2021 Decouple the High-Frequency and Low-Frequency Information of Images for Semantic Segmentation
abstract
As a special kind of signal processing technology, image processing has been developed rapidly after the appearance of convolutional neural network (CNN). At present, the semantic segmentation methods are all based on CNN and ignore the advantages of traditional image processing technology. We combine the two and make them promote each other. The high frequency component of the image represents the edge part and the low frequency represents the body part. Based on this assumption, we use Fourier transform to obtain the high and low frequency component from images. Then, a multi-branch parallel network structure is designed, and the high and low frequency components are sent into two branches respectively to obtain the body and edge features. Finally, the two features are fused together through one deep feature fusion to obtain the final output. In this way, the edge information and body information are decoupled on original images and extracted separately, which not only ensures the consistency of the internal information within objects, but also strengthens the supervision of the edge part which is the most error prone area in semantic segmentation. The results on Cityscapes and KITTI fully demonstrate the effectiveness of our work.
Lianlei Shan, Xiaobin Li 0006, Weiqiang Wang 0001
ICASSP3
2021 Scale-Residual Learning Network for Scene Text Detection
abstract
Detecting incidentally captured text in the wild remains an open problem due to challenging factors including unconstrained scenarios and large scale variation. In this paper, we establish a large-scale scene text detection dataset (LS-Text), containing 36, 000 images and 270, 783 text instances with various scales and complex scenarios, to promote the research of text detection. We propose a Scale-residual Learning Network (SLN) to deal with the scale variation problem in a progressive optimization manner. Specifically, we integrate both learnable feature concatenation and feature up-sampling operator. It can effectively eliminate the residuals between the outputs of SLN and ground-truth text instances by processing both the Feature Fusion Residuals (FFR) and the Scale Transformation Residuals (STR), simultaneously. By stacking multi-scale feature maps in a deep-to-shallow manner, SLN continuously optimizes feature representation by accumulating strong semantic information and rich texture details in a scale-residual learning way. Extensive experimental results on five challenging datasets demonstrate the state-of-the-art performance of the proposed SLN model, and the challenging aspects related to real-world scenarios of the proposed LS-Text dataset. Both the source code of SLN and the LS-Text dataset are available athttps://github.com/SLN-Text-Detection.
Yuanqiang Cai, Chang Liu 0047, Peirui Cheng, Dawei Du, Libo Zhang 0001, Weiqiang Wang 0001, Qixiang Ye
IEEE Trans. Circuits Syst. Video Technol.6
2020 Energy Minimum Regularization in Continual Learning
abstract
How to give agents the ability of continuous learning like human and animals is still a challenge. In the regularized continual learning method OWM, the constraint of the model on the energy compression of the learned task is ignored, which results in the poor performance of the method on the dataset with a large number of learning tasks. In this paper, we propose an energy minimization regularization(EMR) method to constrain the energy of learned tasks, providing enough learning space for the following tasks that are not learned, and increasing the capacity of the model to the number of learning tasks. A large number of experiments show that our method can effectively increase the capacity of the model and reduce the sensitivity of the model to the number of tasks and the size of the network.
Xiaobin Li 0006, Lianlei Shan, Minglong Li, Weiqiang Wang 0001
ICPR4
2020 Global-Local Attention Network for Semantic Segmentation in Aerial Images
abstract
Errors in semantic segmentation could be classified into two types: the large area misclassification and inaccurate local boundaries. Previously attention-based methods typically capture rich global contextual information, which benefits the large area classification but cannot address the local errors of boundaries. In this paper, we propose a Global-Local Attention Network (GLANet) which can simultaneously consider the global context and local details. Specifically, our GLANet consists of two branches: (1) the global attention branch and (2) local attention branch. Furthermore, three different modules are embedded in GLANet for respectively modelling the semantic interdependencies in spatial, channel and boundary dimension. Lastly, we merge the outputs of different branches to enhance the feature representation further, resulting in more precise segmentation. Overall, the proposed method achieves the competitive segmentation accuracy on two public aerial image datasets, bringing significant improvements over the existing baselines.
Minglong Li, Lianlei Shan, Xiaobin Li 0006, Dengji Zhou, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001
ICPR6
2020 UHRSNet: A Semantic Segmentation Network Specifically for Ultra-High-Resolution Images
abstract
Semantic segmentation is a basic task in computer vision, but only limited attention has been devoted to the ultra-high-resolution (UHR) image segmentation. Since UHR images occupy too much memory, they cannot be directly put into GPU for training. Previous methods are cropping images to small patches or downsampling the whole images. Cropping and downsampling cause the loss of contexts and details, which is essential for segmentation accuracy. To solve this problem, we improve and simplify the local and global feature fusion method in previous works. Local features are extracted from patches and global features are from downsampled images. Meanwhile, we propose one new fusion called local feature fusion for the first time, which can make patches get information from surrounding patches. We call the network with these two fusions ultra-high-resolution segmentation network (UHRSNet). These two fusions can effectively and efficiently solve the problem caused by cropping and downsampling. Experiments show a remarkable improvement on Deepglobe dataset [1].
Lianlei Shan, Minglong Li, Xiaobin Li 0006, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001, Weiqiang Wang 0001
ICPR8
2020 Guided Attention Network for Object Detection and Counting on Drones
abstract
Object detection and counting are related but challenging problems, especially for drone based scenes with small objects and cluttered background. In this paper, we propose a new Guided Attention network (GAnet) to deal with both object detection and counting tasks based on the feature pyramid. Different from the previous methods relying on unsupervised attention modules, we fuse different scales of feature maps by using the proposed weakly-supervised Background Attention (BA) between the background and objects for more semantic feature representation. Then, the Foreground Attention (FA) module is developed to consider both global and local appearance of the object to facilitate accurate localization. Moreover, the new data argumentation strategy is designed to train a robust model in the drone based scenes with various illumination conditions. Extensive experiments on three challenging benchmarks (i.e., UAVDT, CARPK and PUCPR+) show the state-of-the-art detection and counting performance of the proposed method compared with existing methods. Code can be found at https://isrc.iscas.ac.cn/gitlab/research/ganet.
Yuanqiang Cai, Dawei Du, Libo Zhang 0001, Longyin Wen, Weiqiang Wang 0001, Siwei Lyu
ACM Multimedia5
2020 Robustly detect different types of text in videos
Yuanqiang Cai, Weiqiang Wang 0001
Neural Comput. Appl.2
2020 SPN: short path network for scene text detection
Yuanqiang Cai, Weiqiang Wang 0001, Haiqing Ren, Ke Lu 0002
Neural Comput. Appl.2
2020 IOS-Net: An inside-to-outside supervision network for scale robust text detection in the wild
Yuanqiang Cai, Weiqiang Wang 0001, Qixiang Ye
Pattern Recognit.2
2020 In-air handwritten Chinese text recognition with temporal convolutional recurrent network
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002
Pattern Recognit.2
2020 Reinterpreting CTC training as iterative fitting
Hongzhu Li, Weiqiang Wang 0001
Pattern Recognit.2
2020 Learning discriminative features via weights-biased softmax loss
Xiaobin Li 0006, Weiqiang Wang 0001
Pattern Recognit.2
2020 Compressing the CNN architecture for in-air handwritten Chinese character recognition
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002
Pattern Recognit. Lett.2
2020 A Direct Regression Scene Text Detector With Position-Sensitive Segmentation
abstract
Direct regression methods have demonstrated their success on various multi-oriented benchmarks for scene text detection due to the high recall rate for small targets and the direct regression for text boxes. However, too many false positive candidates and inaccurate position regression still limit the performance of these methods. In this paper, we propose an end-to-end method by introducing position-sensitive segmentation into the direct regression method to overcome these shortcomings. We generate the ground truth of position-sensitive segmentation maps based on the information of text boxes so that the position-sensitive segmentation module can be trained synchronously with the direct regression module. Besides, more information about the relative position of text is provided for the network through the training of position-sensitive segmentation maps, which improves the expressiveness of the network. We also introduce spatial pyramid of position-sensitive segmentation into the proposed method considering the huge differences in sizes and aspect ratios of scene texts and we propose position-sensitive COI(Corner area of Interest) pooling into the proposed method to speed up the inference. Experiments on datasets ICDAR2015, MLT-17 and COCO-Text demonstrate that the proposed method has a comparable performance with state-of-the-art methods while it is more efficient. We also provide abundant ablation experiments to demonstrate the effectiveness of these improvements in our proposed method.
Peirui Cheng, Yuanqiang Cai, Weiqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 ACPNet: Anchor-Center Based Person Network for Human Pose Estimation and Instance Segmentation
abstract
We present an effective approach to tackle the multi-person pose estimation and person instance segmentation jointly. The approach based on Mask R-CNN uses a set of well-designed labels, called anchor-center based label, to learn keypoints localization in complex and crowded multi-person scenes. Combining the annotation of bounding boxes, we adopt a regression method to predict heatmaps and 2D-offset vector fields based on anchor-center for each keypoint type. Instead of using a person-detector, we use the person proposal network to predict anchors' locations in images. To generate high-quality segmentation mask, we use the ResNet-FPN with the deformable convolutions to model geometric transformations for non-rigid objects. Our method can efficiently and effectively deal with human pose estimation and instance segmentation tasks in a flexible end-to-end manner. Without bells and whistles, our method achieves a comparable result on the COCO keypoints task and the state-of-the-art accuracy on the COCO person instance segmentation task.
Weiqiang Wang 0001
ICME2
2019 Multi-scale Scene Text Detection via Resolution Transform
abstract
Scene text detection is a challenging task because there are many small text targets in the natural scene and the size of scene text varies greatly. For the current popular scene text detection methods, such as EAST and Textboxes, it is difficult to detect small text and text with large differences in size well at the same time. To solve the problem, we propose a novel multi-scale scene text detection method based on EAST. The proposed method extracts high-resolution feature maps at multiple scales via resolution transform and detects text on these feature maps. Through the resolution transform module, the proposed method can detect both kinds of text well. Besides, we use aggregated feature pyramid module to efficiently pass both low-level and high-level information to feature maps at each scale. Experiments on datasets ICDAR2015 and COCO-Text demonstrate that the proposed method has a comparable performance with state-of-the-art methods and it is more efficient. For ICDAR2015 and COCO-Text datasets, the proposed method achieves an F-score of 0.84 and 0.42 respectively.
Peirui Cheng, Weiqiang Wang 0001, Yuanqiang Cai
ICME2
2019 A new perspective: Recognizing online handwritten Chinese characters via 1-dimensional CNN
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002
Inf. Sci.2
2019 N-FTRN: Neighborhoods based fully convolutional network for Chinese text line recognition
Hongzhu Li, Weiqiang Wang 0001, Ke Lu 0002
Multim. Tools Appl.2
2019 In-air handwritten English word recognition using attention recurrent translator
Ji Gan, Weiqiang Wang 0001
Neural Comput. Appl.2
2019 Recognizing online handwritten Chinese characters using RNNs with new computing architectures
Haiqing Ren, Weiqiang Wang 0001
Pattern Recognit.2
2019 A new hybrid-parameter recurrent neural network for online handwritten chinese character recognition
Haiqing Ren, Weiqiang Wang 0001, Xiwen Qu, Yuanqiang Cai
Pattern Recognit. Lett.2
2018 A Unified CNN-RNN Approach for in-Air Handwritten English Word Recognition
abstract
As a new human-computer interaction application, in-air handwriting allows the user to write in the air in a natural way. In this paper, we propose a unified CNN-RNN approach for in-air handwritten English word recognition (IAHEWR), which integrates the advantages of both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). Specifically, the proposed approach follows an encoder-decoder framework, where the encoder is a deep CNN for efficiently processing the input temporal-sequential features, and the decoder is a RNN for accurately generating the target character sequence. We evaluate the proposed approach on an in-air handwritten English word dataset IAHEW-UCAS2016, and the experimental results demonstrate that the proposed approach achieves the comparable recognition accuracy and much higher computation efficiency when compared with the state-of-the-art approach for IAHEWR.
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002
ICME2
2018 Entity Competition Network for Video Classification
abstract
Feature aggregation focuses on the problem of how to fuse the extracted frame-level features into video-level features, which is crucial to the video classification task. The common aggregation networks directly deploy a pooling strategy over time to generate video-level feature. Such approaches treat inputs equally, therefore the spatio-temporal information cannot be fully utilized for varied inputs. To tackle this issue, we introduce a learnable non-linear network called Entity Competition Network (ECN) which harnesses the relationships among entities from adjacent frames. By doing this, we can reassess the importance of each entity and distinguish the entities related to target event from redundant visual information, thus better video-level features can be generated. Our approach achieves the state-of-the-art results on the action recognition datasets of HMDB51 and UCF101. Improvement is also obtained with RGB modality on the untrimmed video classification dataset ActivityNet 200.
Kang Shi, Weiqiang Wang 0001, Changsheng Xu
ICME2
2018 Pre-trained VGGNet Architecture for Remote-Sensing Image Scene Classification
abstract
The visual geometry group network (VGGNet) is used widely for image classification and has proven to be very effective method. Most existing approaches use features of just one type, and traditional fusion methods generally use multiple manually created features. However, to get the benefits of multilayer features remain a significant challenge in the remote-sensing domain. To address this challenge, we present a simple yet powerful framework based on canonical correlation analysis and 4-layer SVM classifier. Specifically, the pretrained VGGNet is employed as a deep feature extractor to extract mid-level and deep features for remote-sensing scene images. We then choose two convolutional (mid-level) and two fully-connected layers produced by VGGNet in which each layer is treated as a separated feature descriptor. Next, canonical correlation analysis (CCA) is used as a feature fusion strategy to refine the extracted features, and to fuse them with more discriminative power. Finally, the support vector machine (SVM) classifier is used to construct the 4-layer representation of the scenes images. Experimenting on a UC Merced and WHU-RS datasets, demonstrate that the proposed approach, even without data augmentation, fine tuning or coding strategy, has a superior performance than state-of-the-art methods used now.
Weiqiang Wang 0001, Shahbaz Pervez Chattha, Sajid Ali 0002
ICPR2
2018 A Multi-Oriented Scene Text Detector with Position-Sensitive Segmentation
abstract
Scene text detection has been studied for a long time and lots of approaches have achieved promising performances. Most approaches regard text as a specific object and utilize the popular frameworks of object detection to detect scene text. However, scene text is different from general objects in terms of orientations, sizes and aspect ratios. In this paper, we present an end-to-end multi-oriented scene text detection approach, which combines the object detection framework with the position-sensitive segmentation. For a given image, features are extracted through a fully convolutional network. Then they are input into text detection branch and position-sensitive segmentation branch simultaneously, where text detection branch is used for generating candidates and position-sensitive segmentation branch is used for generating segmentation maps. Finally the candidates generated by text detection branch are projected onto the position-sensitive segmentation maps for filtering. The proposed approach utilizes the merits of position-sensitive segmentation to improve the expressiveness of the proposed network. Additionally, the approach uses position-sensitive segmentation maps to further filter the candidates so as to highly improve the precision rate. Experiments on datasets ICDAR2015 and COCO-Text demonstrate that the proposed method outperforms previous state-of-the-art methods. For ICDAR2015 dataset, the proposed method achieves an F-score of 0.83 and a precision rate of 0.87.
Peirui Cheng, Weiqiang Wang 0001
ICMR2
2018 Fast exact fingerprint indexing based on Compact Binary Minutia Cylinder Codes
Chaochao Bai, Weiqiang Wang 0001, Tong Zhao 0004, Mingqiang Li
Neurocomputing2
2018 Correlation filter tracker based on sparse regularization
Zhangjian Ji, Weiqiang Wang 0001
J. Vis. Commun. Image Represent.2
2018 Deep learning compact binary codes for fingerprint indexing
abstract
With the rapid growth in fingerprint databases, it has become necessary to develop excellent fingerprint indexing to achieve efficiency and accuracy. Fingerprint indexing has been widely studied with real-valued features, but few studies focus on binary feature representation, which is more suitable to identify fingerprints efficiently in large-scale fingerprint databases. In this study, we propose a deep compact binary minutia cylinder code (DCBMCC) as an effective and discriminative feature representation for fingerprint indexing. Specifically, the minutia cylinder code (MCC), as the state-of-the-art fingerprint representation, is analyzed and its shortcomings are revealed. Accordingly, we propose a novel fingerprint indexing method based on deep neural networks to learn DCBMCC. Our novel network restricts the penultimate layer to directly output binary codes. Moreover, we incorporate independence, balance, quantization-loss-minimum, and similarity-preservation properties in this learning process. Eventually, a multi-index hashing (MIH) based fingerprint indexing scheme further speeds up the exact search in the Hamming space by building multiple hash tables on binary code substrings. Furthermore, numerous experiments on public databases show that the proposed approach is an outstanding fingerprint indexing method since it has an extremely small error rate with a very low penetration rate.
Chaochao Bai, Weiqiang Wang 0001, Tong Zhao 0004, Mingqiang Li
Frontiers Inf. Technol. Electron. Eng.2
2018 Spatiotemporal text localization for videos
Yuanqiang Cai, Weiqiang Wang 0001, Shao Huang, Ke Lu 0002
Multim. Tools Appl.2
2018 In-air handwritten Chinese character recognition with locality-sensitive sparse representation toward optimized prototype classifier
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou
Pattern Recognit.2
2018 Data augmentation and directional feature maps extraction for in-air handwritten Chinese character recognition based on convolutional neural network
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou
Pattern Recognit. Lett.2
2018 Egocentric Temporal Action Proposals
abstract
We present an approach to localize generic actions in egocentric videos, called temporal action proposals (TAPs), for accelerating the action recognition step. An egocentric TAP refers to a sequence of frames that may contain a generic action performed by the wearer of a head-mounted camera, e.g., taking a knife, spreading jam, pouring milk, or cutting carrots. Inspired by object proposals, this paper aims at generating a small number of TAPs, thereby replacing the popular sliding window strategy, for localizing all action events in the input video. To this end, we first propose to temporally segment the input video into action atoms, which are the smallest units that may contain an action. We then apply a hierarchical clustering algorithm with several egocentric cues to generate TAPs. Finally, we propose two actionness networks to score the likelihood of each TAP containing an action. The top ranked candidates are returned as output TAPs. Experimental results show that the proposed TAP detection framework performs significantly better than relevant approaches for egocentric action detection.
Shao Huang, Weiqiang Wang 0001, Shengfeng He, Rynson W. H. Lau
IEEE Trans. Image Process.2
2018 Egocentric Hand Detection Via Dynamic Region Growing
abstract
Egocentric videos, which mainly record the activities carried out by the users of wearable cameras, have drawn much research attention in recent years. Due to its lengthy content, a large number of ego-related applications have been developed to abstract the captured videos. As the users are accustomed to interacting with the target objects using their own hands, while their hands usually appear within their visual fields during the interaction, an egocentric hand detection step is involved in tasks like gesture recognition, action recognition, and social interaction understanding. In this work, we propose a dynamic region-growing approach for hand region detection in egocentric videos, by jointly considering hand-related motion and egocentric cues. We first determine seed regions that most likely belong to the hand, by analyzing the motion patterns across successive frames. The hand regions can then be located by extending from the seed regions, according to the scores computed for the adjacent superpixels. These scores are derived from four egocentric cues: contrast, location, position consistency, and appearance continuity. We discuss how to apply the proposed method in real-life scenarios, where multiple hands irregularly appear and disappear from the videos. Experimental results on public datasets show that the proposed method achieves superior performance compared with the state-of-the-art methods, especially in complicated scenarios.
Shao Huang, Weiqiang Wang 0001, Shengfeng He, Rynson W. H. Lau
ACM Trans. Multim. Comput. Commun. Appl.2
2017 Scene text detection based on pruning strategy of MSER-trees and Linkage-trees
abstract
The extraction and recognition of scene text in images is an important way to understand the semantic information in image. By now, scene text detection is still a challenging problem. In this paper, we present a scene text localization method based on the pruning of Maximally Stable Extremal Region (MSER) tree and Linkage-tree. Concretely, the MSER-tree is first constructed and overlap MSERs are removed by the nonmaximum suppression strategy. Further, the linkage-trees are constructed and pruned based on the features considering the nodes themselves and their siblings. Finally, the non-text lines are eliminated based on to the CNN-based features and intensity contrast. Experimental results on two challenging datasets demonstrate the effectiveness of the proposed method.
Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou
ICME2
2017 An end-to-end recognizer for in-air handwritten Chinese characters based on a new recurrent neural networks
abstract
In-air handwriting is becoming a new human-computer interaction way. It is a challenging task to accurately recognizing in-air handwritten Chinese characters. In this paper, we present an end-to-end recognizer for in-air handwritten Chinese characters by using recurrent neural networks (RNN). Compared with the existing methods, the proposed RNN based methods does not need to explicitly extract features and directly take a sequence of dot locations as input. We have made two aspects of modifications on traditional RNN for improving the recognition accuracy. Concretely, the sum-pooling is performed on the states of each hidden layers, and a faster convergence in training can be obtained. Additionally, an assistant objective function is introduced into the conventional loss function, which brings a slight increase of performance. To evaluate the performance of the proposed method, the experiments are carried out on the IAHCC-UCAS2016 datasets to compare ours with other state-of-art methods. The experimental results show that the proposed RNN model has a fairly high recognition accuracy for in-air handwritten Chinese characters.
Haiqing Ren, Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou, Qiuchen Yuan
ICME2
2017 Exposure fusion via sparse representation and shiftable complex directional pyramid transform
Weiqiang Wang 0001, Guangmei Xu, Ruizhe Zhang 0006, Jingzun Zhang
Multim. Tools Appl.2
2017 Saliency-guided Pairwise Matching
Shao Huang, Weiqiang Wang 0001
Pattern Recognit. Lett.2
2017 Stereo Object Proposals
abstract
Object proposal detection is an effective way of accelerating object recognition. Existing proposal methods are mostly based on detecting object boundaries, which may not be effective for cluttered backgrounds. In this paper, we leverage stereopsis as a robust and effective solution for generating object proposals. We first obtain a set of candidate bounding boxes through adaptive transformation, which fits the bounding boxes tightly to object boundaries detected by rough depth and color information. A two-level hierarchy composed of proposal and cluster levels is then constructed to estimate object locations in an efficient and accurate manner. Three stereo-based cues "exactness," "focus," and "distribution" are proposed for objectness estimation. Two-level hierarchical ranking is proposed to accurately obtain ranked object proposals. A stereo data set with 400 labeled stereo image pairs is constructed to evaluate the performance of the proposed method in both indoor and outdoor scenes. Extensive experimental evaluations show that the proposed stereo-based approach achieves a better performance than the state of the arts with either a small or a large number of object proposals. As stereopsis can be a complement to the color information, the proposed method can be integrated with existing proposal methods to obtain superior results.
Shao Huang, Weiqiang Wang 0001, Shengfeng He, Rynson W. H. Lau
IEEE Trans. Image Process.2
2016 MQDF with a novel covariance matrix estimation and discriminant LSRC, which is better for in-air handwritten Chinese character recognition
abstract
With the advance of 3-dimensional sensing devices, the in-air handwriting, as a more natural way for human and computer interaction, is being developed by the UCAS-CVMT Lab. Compared with the conventional handwritten Chinese characters generated by touching, it is more challenging to accurately recognize them due to unconstrained one-stroke writing style. This paper presents two recognizers to address this problem. One is built on the modified quadratic discriminant function (MQDF), where a new kernel method is used to estimate the covariance matrix. Additionally, we present a new classifier built on the locality-sensitive sparse representation based classifier (LSRC). It introduces the discriminant information to reinforce the recognition ability, called discriminant LSRC. We have applied them on the in-air handwritten Chinese character dataset IAHCC-UCAS2015. The experimental results show that the first one has the higher accuracy but huge storage cost, while the DLSRC recognizer has the comparable accuracy and very compact storage space.
Weiqiang Wang 0001, Ke Lu 0002
ICIP2
2016 High-order directional features and sparse representation based classification for in-air handwritten Chinese character recognition
abstract
The in-air handwriting is a natural and promising human-computer interaction way. Compared with handwritten Chinese characters on touch screen, the in-air handwritten Chinese characters have their unique characteristics, e.g., each character is always written in a single stroke. In this paper, we propose a high-order directional feature for recognizing in-air handwritten Chinese characters. The proposed highorder features characterize the rate of direction change and the change rate of the rate of direction change, and can be easily be combined with the 8-directional features to get an enhanced version. Additionally, we exploit the locality-sensitive dictionary learning and sparse representation based classifier (LSRC) to recognize in-air handwritten Chinese characters. Since few work has applied the SRC in on-line handwritten Chinese character recognition (OHCCR), we evaluate the proposed system on both an in-air handwritten Chinese character dataset, the IAHCC-UCAS2015 dataset, and a handwritten Chinese character dataset, the SCUT-COUCH2009 database. The experimental results show that the proposed high-order directional features, when combined with the firstorder directional features (8-directional features), can improve the recognition accuracy, and they are more suitable for IAHCCR than traditional OHCCR. Additionally, the results also demonstrate the LSRC is a good choice for IAHCCR.
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Zhangjian Ji
ICME2
2016 Egocentric hand detection via region growth
abstract
Wearable cameras used to record daily life are attracting researchers' attention, and a large number of ego-related applications have been developed in recent years. Hand detection is one of the key steps for the tasks like gesture recognition, action recognition and understanding hand-based interaction in egocentric videos, since humans are accustomed to interacting with objects using their hands. In this work a novel region growth approach is proposed for egocentric hand detection. The proposed method first identifies seed regions most likely containing the parts of hand based on the distribution of hand-related matches. Then hand regions are gradually located by extending from the seed regions. Finally a whole hand is obtained according to the scores of adjacent superpixels, which are evaluated based on four egocentric cues: contrast, distance, previous hand position and skin appearance. The experimental results on two publicly available datasets demonstrate this work achieves satisfactory performance.
Shao Huang, Weiqiang Wang 0001, Ke Lu 0002
ICPR2
2016 In-air handwritten Chinese character recognition using discriminative projection based on locality-sensitive sparse representation
abstract
Dimensionality reduction methods have been shown to be effective for handwritten Chinese character recognition. In this paper, we propose discriminative projection based on locality-sensitive sparse representation (DPLSR) for in-air handwritten Chinese character recognition. DPLSR based on the locality-sensitive sparse representation based classifier (LSRC), which can provide closed-form solutions and maintain the data locality constraint during the sparse coding stage. In contrast to sparse representation classifier steered discriminative projection (SRC-DP), which did not consider global structure of data and use all training samples as dictionary atoms, DPLSR is able to use fewer atoms and spend less training time to achieve better performance. Experiments are conducted on the IAHCC-UCAS2016 dataset built by us, experimental results demonstrate the effectiveness of proposed method.
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002
ICPR2
2016 Detecting Salient Objects via Color and Texture Compactness Hypotheses
abstract
In recent years, the object-level saliency detection has attracted much research attention, due to its usefulness in many high-level tasks. Existing methods are mostly based on the contrast hypothesis, which regards the regions with high contrast in a certain context as salient objects. Although the contrast hypothesis is effective in many scenarios, it cannot handle some difficult cases. As a remedy to address the weakness of contrast hypothesis, we propose a novel compactness hypothesis, which assumes salient regions are more compact than background from the perspectives of both color layout and texture layout. Based on the compactness hypotheses, we implement an effective object-level saliency detection method. In the proposed method, we first construct a weak saliency map based on the compact hypotheses, then collect samples from the weak saliency map to train a dedicated classifier. This classifier is applied on each individual pixel of the input image to produce a confidence score. Finally, the confidence scores are used to form a saliency map. This process is carried out at different scales, and the corresponding results are integrated into the formation of the final saliency map. The proposed approach is evaluated on eight benchmark data sets, where it delivers the competitive performance compared with the state-of-the-art methods.
Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002
IEEE Trans. Image Process.2
2015 An Efficient Indexing Scheme Based on K-Plet Representation for Fingerprint Database
Chaochao Bai, Tong Zhao 0004, Weiqiang Wang 0001
ICIC (1)3
2015 Recognition of In-air Handwritten Chinese Character Based on Leap Motion Controller
Ning Xu 0008, Weiqiang Wang 0001, Xiwen Qu
ICIG (3)2
2015 A novel binarization approach for text in images
abstract
Accurate recognition of scene text and overlaid text is still a challenging issue due to degradation and complex background, and text binarization is crucial for recognition accuracy. This paper presents an effective method to extract characters in images and video frames. Our method assumes that background pixels possess good spatial connectivity and high appearance similarity to boundary pixels in a cropped text string image. It first computes the confidence of pixels as text. Then the confidence map is exploited to partition text regions into characters. Further, each character region is clustered into different layers and background components are removed to generate candidate binarization results. The final result is obtained based on the scores of each layer. Our method is validated by better recognition rates and segmentation accuracy on the ICADR03 dataset and a big dataset of overlaid text.
Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002
ICIP2
2015 Retrieving images combining saliency detection with IRM
abstract
The research from cognitive psychology and neurobiology suggests that people have a strong ability to perceive objects before identifying them, and the human vision system (HVS) processes some salient regions of an image while leaving others nearly unprocessed. Based on this observation, we assume people prefer to pay more attention to salient regions when judging whether images are matched. A novel saliency detection based on reconstruction residual is proposed to calculate saliency maps and generate salient regions, which helps decrease the interference from irrelevant regions in image retrieval. The introduction of integrated region matching (IRM) can better characterize the spatial constraints among salient regions to conduct similarity measurement. The experimental results demonstrate that the proposed saliency detection has a good performance. At the same time, the combination of salient regions and IRM can help improve the state-of-the-art retrieval algorithms.
Shao Huang, Weiqiang Wang 0001
ICIP2
2015 Robustly tracking objects via multi-task kernel dynamic sparse model
abstract
Recently, sparse representation has been successfully applied by some generative tracking methods. However, few methods consider the correlation between the representations of each particle in time domain and space domain. Additionally, most methods use the raw pixels as templates which can not well adapt to the sophisticated object changes. To solve these problems, we consider the sparse representation in kernel space and propose a multi-task kernel dynamic sparse tracking algorithm (MTKDST). As compared to previous methods, our method exploits the dependencies between particles in the space domain and the correlation of particle representation in the time domain to improve the tracking performance. Furthermore, we also adopt the multikernel fusion mechanism to utilize multiple complementary visual features (e.g., spatial color histogram and spatial gradient-orientation histogram) to enhance the robustness of the proposed method. The comprehensive experiments on several challenging image sequences demonstrate that the proposed method outperforms the state-of-the-art approaches in tracking accuracy.
Zhangjian Ji, Weiqiang Wang 0001, Ke Lu 0002
ICIP2
2015 In-air handwritten Chinese character recognition using multi-stage classifier based on adaptive discriminative locality alignment
abstract
The in-air handwriting is a natural and useful way for human-computer interaction. Yet, to our knowledge, few work has been done for the in-air handwritten Chinese character recognition (IAHCCR). In this paper, we present a multi-stage recognizer to address the problem of IAHCCR. The proposed methods can also deal with the classical handwritten Chinese character recognition (HCCR). We find that the discriminative locality alignment (DLA) technique in HCCR heavily depends on the choice of parameters in practice. To overcome the disadvantage, we present an adaptive discriminative locality alignment (ADLA), which does not involve the parameter optimization process. At the same time, a new static similar characters collection technique is proposed. We evaluate the proposed methods on the IAHCC-UCAS2014 dataset, an in-air handwritten Chinese character dataset constructed by us, as well as the SCUT-COUCH2009 database, a HCCR dataset. The experimental results demonstrate the effectiveness of the proposed methods on two different kinds of dataset.
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Ning Xu 0008
ICIP2
2015 Detecting Salient Objects via Spatial and Appearance Compactness Hypotheses
abstract
Object-level saliency detection has been attracting a lot of attention, due to its potential enhancement in many high-level vision tasks. Many previous methods are based on the contrast hypothesis which regards the regions with high contrast in a certain context as salient. Although the contrast hypothesis is valid in many cases, it cannot handle some difficult cases. To make up for the weakness of contrast hypothesis, we propose a novel compactness hypothesis which assumes salient regions are more compact than background spatially and in appearance. Based on compactness hypotheses, we implement an effective object-level saliency detection method, which is demonstrated to be effective even in difficult cases. In addition, we present an adaptive multiple saliency maps fusion framework which can automatically select saliency maps of high quality according to three quality assessment rules. We evaluate the proposed method on four benchmark datasets and the comparable performance as the state-of-the-art methods has been achieved.
Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002
ACM Multimedia2
2015 Object tracking based on local dynamic sparse model
Zhangjian Ji, Weiqiang Wang 0001
J. Vis. Commun. Image Represent.2
2015 An efficient multi-threshold AdaBoost approach to detecting faces in images
Weiqiang Wang 0001, Ke Lu 0002
Multim. Tools Appl.2
2015 An effective graph-cut scene text localization with embedded text segmentation
Weiqiang Wang 0001
Multim. Tools Appl.2
2014 Retrieving images using saliency detection and graph matching
abstract
The need for fast retrieving images has recently increased tremendously in many application areas (biomedicine, military, commerce, education, etc.). In this work, we exploit the saliency detection to select a group of salient regions and utilize an undirected graph to model the dependency among these salient regions, so that the similarity of images can be measured by calculating the similarity of the corresponding graphs. Identification of salient pixels can decrease interferences from irrelevant information, and make the image representation more effective. The introduction of the graph model can better characterize the spatial constraints among salient regions. The comparison experiments are carried out on the three representative datasets publicly available (Holidays, UKB, and Oxford 5k), and the experimental results show that the integration of the proposed method and the SIFT-like local descriptors can better improve the existing state-of-the-art retrieval accuracy.
Shao Huang, Weiqiang Wang 0001
ICIP2
2014 Robust object tracking via multi-task dynamic sparse model
abstract
Recently, sparse representation has been widely applied to some generative tracking methods, which learn the representation of each particle independently and do not consider the correlation between the representation of each particle in the time domain. In this paper, we formulate the object tracking in a particle filter framework as a multi-task dynamic sparse learning problem, which we denote as Multi-Task Dynamic Sparse Tracking(MTDST). By exploring the popular sparsity-inducing ℓ1, 2mixed norms, we regularize the representation problem to enforce joint sparsity and learn the particle representations together. Meanwhile, we also introduce the innovation sparse term in the tracking model. As compared to previous methods, our method mines the independencies between particles and the correlation of particle representation in the time domain, which improves the tracking performance. In addition, because the loft least square is robust to the outliers, we adopt the loft least square to replace the least square to calculate the likelihood probability. In the updating scheme, we eliminate the influences of occlusion pixels when updating the templates. The comprehensive experiments on the several challenging image sequences demonstrate that the proposed method consistently outperforms the existing state-of-the-art methods.
Zhangjian Ji, Weiqiang Wang 0001
ICIP2
2014 Robust object tracking via incremental subspace dynamic sparse model
abstract
Sparse representation has been widely applied to some generative tracking methods. However, these methods do not consider the correlation between sparse representation coefficients in the time domain. In this paper, we propose a novel incremental subspace dynamic sparse tracking (ISDST) model with the error term of Gaussian-Laplacian distribution, which fully considers the correlation of object representations between consecutive frames by compressive sensing, and can effectively handle the occlusion in scenes. Next, the outlier entries, especially caused by the occlusion, have some group effect, so we adopt the spatial structured sparse via l1, 2mixed norms instead of the original l1sparse items. In addition, since the occlusion changes is very little between consecutive frames, we maintain an occlusion mask and eliminate the influence of occlusion pixels in the process of calculating the likelihood probability. Extensive experiments on challenging sequences demonstrate that our method consistently outperforms existing state-of-the-art methods.
Zhangjian Ji, Weiqiang Wang 0001, Ning Xu 0008
ICME2
2014 A Filtration Strategy Based on FD-CRF for Image Matching
abstract
The need for fast retrieving images has recently increased tremendously in many application areas. SIFT-like local descriptor-based matching is widely adopted and has achieved state-of-the-art performance. However, it becomes inefficient when computational and storage resources are limited. Besides, local descriptor-based methods may suffer difficulties when an image pair contains multiple similar local regions. In this work, we propose a novel and effective filtration strategy based on the Conditional Random Field (CRF) model to enhance image retrieval. As the CRF model is used to depict the dependencies of adjacent components, we regard the essential components of an image as the basic structure of CRF. The novel Fourier Descriptor CRF (FD-CRF) method is first proposed to utilize the advantages of CRF and global shape features, then the filtration strategy is adopted to integrate FD-CRF and SIFT-like descriptors for better retrieval results. The experiments demonstrate that our method is practical and outperforms state-of-the-art methods in matching accuracy.
Shao Huang, Weiqiang Wang 0001
ICPR2
2014 Video Text Extraction Using the Fusion of Color Gradient and Log-Gabor Filter
abstract
Video text which contains rich semantic information can be utilized for video indexing and summarization. However, compared with scanned documents, text recogniton for video text is still a challenging problem due to complex background. Segmenting text line into single characters before text extraction can achieve higher recognition accuracy, since background of single character is less complex compared with whole text line. Therefore, we first perform character segmentation, which can accurately locate the character gap in the text line. More specifically, we get a fusion map which fuses the results of color gradient and log-gabor filter. Then, candidate segmentation points are obtained by vertical projection analysis of the fusion map. We get segmentation points by finding minimum projection value of candidate points in a limited range. Finally, we get the binary image of the single character image by applying K-means clustering and combine their results to form binary image of the whole text line. The binary image is further refined by inward filling and the fusion map. The experimental results on a large amount of data show that the proposed method can contribute to better binarization result which leads to a higher character recognition rate of OCR engine.
Zhike Zhang, Weiqiang Wang 0001, Ke Lu 0002
ICPR2
2014 Detect foreground objects via adaptive fusing model in a hybrid feature space
Zhangjian Ji, Weiqiang Wang 0001
Pattern Recognit.2
2013 Extract foreground objects based on sparse model of spatiotemporal spectrum
abstract
In this paper, we present a novel foreground object detection method based on the sparse model of the spectrum of spatiotemporal DCT domain, which is robust for high dynamic scenes. First, we adopt the three-dimensional Discrete Cosine Transform (DCT) to calculate the spatiotemporal spectrum representation of the current frame. Then, identification of foreground pixels is formulated as the analysis of the sparse solution of an optimization problem, where foreground pixels correspond to an outlier of the sparse model. Finally, the background updating method is presented to adaptively update the dictionary of sparse model corresponding to background representation. The experimental results on four challenging video sequences show that the proposed method is more robust to high dynamic changes of scenes compared with four representative methods.
Zhangjian Ji, Weiqiang Wang 0001, Ke Lu 0002
ICIP2
2013 Object-level saliency detection based on spatial compactness assumption
abstract
Object-level saliency detection is an important aspect of visual saliency. Most existing methods build on the contrast assumption. It tends to highlight the saliency of the regions with high contrast in a certain context, but it does not work well in some scenarios. In this paper, we propose a novel spatial compactness assumption which considers that salient regions are spatially more compact than background regions. Based on it, we present two object-level saliency detection methods: the patch-based method and the region-based method. In the experiments, both methods are compared with nine state-of-the-art methods on a public dataset and the best performances are obtained. The experimental results show that the spatial compactness assumption is valid and the proposed methods can uniformly highlight salient objects, even for large ones.
Weiqiang Wang 0001
ICIP2
2013 Foreground Detection Utilizing Structured Sparse Model via l1, 2 Mixed Norms
abstract
Foreground object detection is a crucial technique of intelligent surveillance systems, and it is still a challenging problem in complex scenes with illumination variations and dynamic backgrounds. Intuitively, the foreground object pixels are often not sparsely distributed but tend to be clustered. Motivated by this hypothesis, we present a new structured sparse model to extract foreground objects, which introduces the spatial neighborhood information into a unified optimization framework by l1,2 mixed norms. Simultaneously, we also give the solving method of the proposed model in details. Moreover, we apply the model to the sparse signal recovery and background subtraction in videos. In the experiments, better performance is obtained over previous methods. The experimental results validate the hypothesis and the effectiveness of the proposed method.
Zhangjian Ji, Weiqiang Wang 0001, Ke Lu 0002
SMC2
2013 A Novel Approach for Binarization of Overlay Text
abstract
In this paper, we presents a new binarization approach to extract text pixels from complex background in video frames. The binarization computation is a crucial step for, video text recognition, which can greatly increase the recognition, accuracy of an OCR software. The proposed approach consists, of four phases. First, the text polarity is determined, i.e. light text with dark background or dark text with light background., Then the pixels in the given image are clustered into K clusters, using the K-means algorithm in the RGB color space and the, text cluster is selected based on the text polarity. Further, the, MRF Model is exploited to get the binarization result. Finally, the, result is further refined by the Log-Gabor filter. The Experimental, results on a large dataset show that the significant gains have been, obtained according to the segmentation performance on the pixel, level as well as the OCR accuracy.
Zhike Zhang, Weiqiang Wang 0001
SMC2
2012 Effectively localize text in natural scene images
Ke Lu 0002, Weiqiang Wang 0001
ICPR3
2012 A robust and efficient shot boundary detection approach based on fisher criterion
abstract
In this paper, we present a robust and efficient approach which is capable to simultaneously detect various shot boundaries in a unified way. The proposed approach first detects general shot boundaries based on the idea of Fisher criterion, and then classifies them into two categories, cut and gradual transition (GT), by an SVM classifier. Further computation is performed for the GT shot boundaries to expand rough boundary locations between two frames into the transition interval consisting of all the transitional frames. Finally, a postprocessing module is employed to merge overlapped transitions. The evaluation experiments show that the proposed approach has the impressive performance in both efficiency and accuracy, and outperforms the best results of all the participants of TRECVID 2006.
Weiqiang Wang 0001
ACM Multimedia2
2012 Robustly Extracting Captions in Videos Based on Stroke-Like Edges and Spatio-Temporal Analysis
abstract
This paper presents an effective and efficient approach to extracting captions from videos. The robustness of our system comes from two aspects of contributions. First, we propose a novel stroke-like edge detection method based on contours, which can effectively remove the interference of non-stroke edges in complex background so as to make the detection and localization of captions much more accurate. Second, our approach highlights the importance of temporal feature, i.e., inter-frame feature, in the task of caption extraction (detection, localization, segmentation). Instead of regarding each video frame as an independent image, through fully utilizing the temporal feature of video together with spatial analysis in the computation of caption localization, segmentation and post-processing, we demonstrate that the use of inter-frame information can effectively improve the accuracy of caption localization and caption segmentation. In the comprehensive our evaluation experiments, the experimental results on two representative datasets have shown the robustness and efficiency of our approach.
Weiqiang Wang 0001
IEEE Trans. Multim.2
2010 Extracting Captions in Complex Background from Videos
abstract
Captions in videos play a significant role for automatically understanding and indexing video content, since much semantic information is associated with them. This paper presents an effective approach to extracting captions from videos, in which multiple different categories of features (edge, color, stroke etc.) are utilized, and the spatio-temporal characteristics of captions are considered. First, our method exploits the distribution of gradient directions to decompose a video into a sequence of clips temporally, so that each clip contains a caption at most, which makes the successive extraction computation more efficient and accurate. For each clip, the edge and corner information are then utilized to locate text regions. Further, text pixels are extracted based on the assumption that text pixels in text regions always have homogeneous color, and their quantity dominates the region relative to non-text pixels with different colors. Finally, the segmentation results are further refined. The encouraging experimental results on 2565 characters have preliminarily validated our approach.
Weiqiang Wang 0001, Tingshao Zhu
ICPR2
2010 Extracting captions from videos using temporal feature
abstract
Captions in videos provide much useful semantic information for indexing and retrieving video contents. In this paper, we present an effective approach to extracting captions from videos. Its novelty comes from exploiting the temporal information in both localization and segmentation of captions. Since some simple features such as edges, corners and color are utilized, our approach is efficient. It involves four steps. First, we exploit the distribution of corners to spatially detect and locate the caption in a frame. Then the temporal localization for different captions in a video is performed by identifying the change of stroke directions. After that, we segment the caption pixels in a clip with a same caption based on the consistency and dominant distribution of caption color. Finally, the segmentation results are further refined. The experimental results on two representative movies have preliminarily verified the validity of our approach.
Weiqiang Wang 0001
ACM Multimedia2
2009 A hybrid text segmentation approach
abstract
In this paper, we present a hybrid text segmentation approach for embedded text in images, aiming to combining the advantages of the difference-based methods and the similarity-based methods together. First a new stroke edge filter is applied to obtain stroke edge map. Then a two-threshold method based on the improved Niblack thresholding technique is utilized to identify stroke edges. Those pixels between the edge pairs above the high threshold are collected to estimate the representative of stroke color, so that stroke pixels are further extracted by computing the color similarity. Finally some heuristic rules are devised to integrate stroke edge and stroke region information to obtain better segmentation results. The experimental results show that our approach can effectively segment text from background.
Weiqiang Wang 0001, Qingming Huang, Wen Gao 0001, Laiyun Qing
ICME2
2008 Fast and effective text detection
abstract
Text in images and videos is a significant cue for visual content understanding and retrieval. In this paper, we present a fast and effective approach to locate text lines even under complex background. First, our algorithm uses the stroke filter to calculate the stroke maps in horizontal, vertical, left-diagonal, right-diagonal directions. Then a 24- dimensional feature is extracted for each sliding window and a SVM is used to obtain rough text regions. The rough text regions are further refined through a group of rules. And candidate text lines were localized more accurately through projection profile of the refined text regions. Finally another SVM classifier based on a 6-dimensional feature is used to verify the candidate text lines. The experimental results on challenging databases show that this approach can fast and effectively detect and localize text lines.
Weiqiang Wang 0001, Shuqiang Jiang, Qingming Huang, Wen Gao 0001
ICIP2
2008 Effective scene matching with local feature representatives
abstract
Scene matching measures the similarity of scenes in photos and is of central importance in applications where we have to properly organize large amount of digital photos by scene categories. In this paper, we present a novel scene matching method using local features representatives. For a given image, its scene is compactly represented as a set of cluster centers, called local feature representatives, where the clusters are obtained using the affinity propagation (AP) algorithm to aggregate local features according to their spatial closeness and appearance similarity. The similarity of scenes in two images is then measured by a modified Earth Mover Distance (EMD) between their corresponding sets of local feature representatives. Empirical experiments on real world photos shows that our method is comparable to the state-of-the-arts.
Shugao Ma, Weiqiang Wang 0001, Qingming Huang, Shuqiang Jiang, Wen Gao 0001
ICPR2
2008 Matching images more efficiently with local descriptors
abstract
Image matching is a fundamental task for many applications of computer vision. Today it is very popular to represent two matched images as two bags of local descriptors, and the classic RANSAC based matching procedure is always exploited in the task. In this paper, we present a much efficient image matching approach based on sets of any local descriptors. A block-to-block strategy is devised to speed up the establishment of local correspondences. Additionally, the weighted RANSAC (w-RANSAC) technique is proposed to make the search of optimal global models converge faster. Comparative experiments with the RANSAC based paradigm show our approach can not only generate more accurate correspondences, but also double the matching speed.
Dong Zhang 0001, Weiqiang Wang 0001, Qingming Huang, Shuqiang Jiang, Wen Gao 0001
ICPR2
2008 Modeling Background and Segmenting Moving Objects from Compressed Video
abstract
Modeling background and segmenting moving objects are significant techniques for video surveillance and other video processing applications. Most existing methods of modeling background and segmenting moving objects mainly operate in the spatial domain at pixel level. In this paper, we present three new algorithms (running average, median, mixture of Gaussians) modeling background directly from compressed video, and a two-stage segmentation approach based on the proposed background models. The proposed methods utilize discrete cosine transform (DCT) coefficients (including ac coefficients) at block level to represent background, and adapt the background by updating DCT coefficients. The proposed segmentation approach can extract foreground objects with pixel accuracy through a two-stage process. First a new background subtraction technique in the DCT domain is exploited to identify the block regions fully or partially occupied by foreground objects, and then pixels from these foreground blocks are further classified in the spatial domain. The experimental results show the proposed background modeling algorithms can achieve comparable accuracy to their counterparts in the spatial domain, and the associated segmentation scheme can visually generate good segmentation results with efficient computation. For instance, the computational cost of the proposed median and MoG algorithms are only 40.4% and 20.6% of their counterparts in the spatial domain for background construction.
Weiqiang Wang 0001, Jie Yang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2007 Object Recognition Based on Dependent Pachinko Allocation Model
abstract
Recently the "bag of words" model becomes popular in the approaches to object recognition. These approaches model an image as a collection of local patches called "visual words", and recognize objects in the image through inferring latent topics associated with the set of visual words. In this paper, we apply an extension version of Pachinko allocation model (PAM) to object recognition. Our PAM based approach models the correlation-ship of latent topics explicitly in a hierarchical structure. To relax the independent assumption for visual words and refine the topic inferring, we incorporate the prior knowledge of cooccurrence dependence among visual words into PAM. Highly competitive recognition results on both Caltech4 and Caltech101 datasets show the proposed approach is more expressive and discriminative than most existing methods of object recognition.
Yuanning Li, Weiqiang Wang 0001, Wen Gao 0001
ICIP (5)2
2007 A Fast Approach for Natural Image Matting using Structure Information
abstract
Natural image matting plays an important role in image and video editing. It has been addressed hotly because it is inherently an under-constrained problem - we must estimate both the foreground and background colors (we call them component colors) at each pixel to calculate its opacity, according to the only known observed color. Prior assumption such as statistics and smoothness are utilized for estimation, but these methods either are simple to handle situations such as complex background and large interaction region or have high computational complexity. This paper proposes a fast technique to estimate the component colors based on structure information. Our approach exploits a simple convolution operation to detect structure information in images. It then uses two kinds of estimation methods to propagate colors based on the structure types. Experimental results show that our method is fast and efficient to handle objects with strong structures and large part of interaction region with the background.
Qianhui Ning, Weiqiang Wang 0001, Caifeng Zhu, Laiyun Qing, Qingming Huang
ICME2
2007 An Effective Local Invariant Descriptor Combining Luminance and Color Information
abstract
Extraction of stable local invariant features is very important in many computer vision applications, such as image matching, object recognition and image retrieval. Most existing local invariant features mainly characterize luminance information, and neglect color information. In this paper, we present a new local invariant descriptor characterizing both of them, which combines three photometric invariant color descriptors with the famous SIFT descriptor. To reduce the dimension of the combined high-dimensional invariant feature the principal component analysis (PCA) is used. Our experiments show the proposed local descriptor through combining luminance and color information outperforms the descriptors that only utilize a single category of information, and combining the three color feature representations is more effective than only one.
Dong Zhang 0001, Weiqiang Wang 0001, Wen Gao 0001, Shuqiang Jiang
ICME2
2006 Effective and efficient object-based image retrieval using visual phrases
abstract
In this paper, we draw an analogy between image retrieval and text retrieval and propose a visual phrase-based approach to retrieve images containing desired objects. The visual phrase is defined as a pair of adjacent local image patches and is constructed using data mining. We devise methods on how to construct visual phrases from images and how to encode the visual phrase for indexing and retrieval. Our experiments demonstrate that visual phrase-based retrieval approach can be very efficient and can be 20% more effective than its visual word-based counterpart.
Weiqiang Wang 0001, Wen Gao 0001
ACM Multimedia2
2005 Local invariant descriptor for image matching
abstract
Image matching is a fundamental task of many computer vision problems. In this paper we present a novel approach for matching two images in the presence of image rotation, scale, and illumination changes. The proposed approach is based on local invariant features. A two-step process detects local invariant regions. Characteristic circles associated with these regions illustrate the position and radius of the regions. Then, the regions are represented by a new image descriptor. To test the new descriptor, we evaluate it in image matching and retrieval experiments. The experimental results show that using our descriptors results in effective and faster matching.
Wei Zeng 0006, Wen Gao 0001, Weiqiang Wang 0001
ICASSP (2)4
2005 Research on the Discrimination of Pornographic and Bikini Images
abstract
Our research extends the general technologies detecting pornographic images to prevent the benign images whose content is approximate with the pornographic ones from being screened. This paper presents a multiple-step method to distinguish the benign bikini photos from pornographic photos. The proposed approach utilizes the information about the body shape and face to determine the feature region where the specific patterns of the skin distribution of bikini model's body are represented, and uses a neural network to learn these, patterns to recognize the bikini photos. The experimental results show that more than 80% normal bikini photos are recognized from the pornographic images by the method, and at the same time few pornographic images are mistaken as bikini photos. The results indicate it is possible to distinguish the bikini photos from pornographic ones.
Weiqiang Wang 0001, Wen Gao 0001
ISM2
2004 Skin-color detection based on adaptive thresholds
abstract
In this paper a new skin detection method based on adaptive thresholds is proposed. As compared with the fixed threshold histogram widely method used, ours can find optimal thresholds to the different complex backgrounds. Four clues are summarized from the skin probability distribution histogram (SPDH) to help search candidates of optimum thresholds, and an ANN classifier is trained to select the final optimum threshold. A novel image relation operation is also proposed to eliminate the confusing backgrounds. The method is fast thus appropriate for real-time applications since no iterative operation is involved. Experimental results show that the proposed method can achieve better performance than the fixed threshold histogram method.
Ming-Ji Zhang, Weiqiang Wang 0001, Wen Gao 0001
ICIG2
2004 Shape-based adult images detection
abstract
This paper reports an investigation on adult images detection based on the shape features of skin regions. In order to accurately detect skin regions, we propose a skin detection method using multi-Bayes classifiers in the paper. Based on skin color detection results, shape features are extracted and fed into a boosted classifier to decide whether or not the skin regions represent a nude. We evaluate adult image detection performance using different boosted classifiers and different shape descriptors. Experimental results show that classification using boosted C4.5 classifier and combination of different shape descriptors outperforms other classification schemes.
Wei Zeng 0006, Wen Gao 0001, Weiqiang Wang 0001
ICIG4
2004 A novel compressed domain shot segmentation algorithm on H.264/AVC
abstract
This paper presents a novel shot segmentation algorithm on the H.264/AVC video, which operates in the compressed domain. First, the algorithm exploits the intra prediction mode histogram to locate those potential GOPs, where shot transitions occur with great probability. Secondly, to further find shot boundaries at the frame level, we count the number of macroblocks with different inter prediction modes as the features and exploit HMMs to automatically model different cases in which shot transitions can occur among I, P and B frames. Since H.264/AVC provides more motion compensation modes, using HMMs can avoid the tediousness of manually tuning multiple thresholds simultaneously. The experimental results show that the algorithm is efficient and robust and it can not only locate cuts, but also work for gradual shot transitions.
Yang Liu 0006, Weiqiang Wang 0001, Wen Gao 0001, Wei Zeng 0006
ICIP2
2003 Illumination Invariant Shot Boundary Detection
Laiyun Qing, Weiqiang Wang 0001, Wen Gao 0001
IDEAL2
2003 Objectionable Image Recognition System in Compression Domain
Qixiang Ye, Wen Gao 0001, Wei Zeng 0006, Weiqiang Wang 0001, Yang Liu 0006
IDEAL5
2000 News Content Highlight via Fast Caption Text Detection on Compressed Video
Weiqiang Wang 0001, Wen Gao 0001, Jintao Li 0001, Shouxun Lin
IDEAL1