Shuwu Zhang

dblp:03/2306 · DBLP profile ↗
← Back
63ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-6013-6351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 36 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Dual-Branch Asymmetric Discrepancy Learning Based on Fake Image Pattern-Coexistence for AI-Generated Image Detection
abstract
With the rapid advancement of generative models, high-fidelity AI-generated images have become increasingly indistinguishable from real images, posing significant challenges to traditional detection methods that rely on explicit artifacts or uniform feature learning. We hypothesize that detection ambiguity originates from pattern coexistence: synthetic images simultaneously embed (a) authentic patterns inherited from real-image distributions and (b) synthetic patterns induced by generative architectures, whereas real images maintain consistent patterns. We validate this hypothesis through SHAP-based quantitative analysis, demonstrating that synthetic images inherently exhibit a dual distribution—simultaneously containing authentic patterns and synthetic traces—while real images show a unimodal distribution. Building on this insight, this paper proposes a Dual-Branch Asymmetric Discrepancy Learning (DADL) framework. The DADL leverages multi-scale feature extraction and Asymmetric Feature Discrepancy Loss to capture and amplify such pattern differences across multiple scales. Extensive experiments on three benchmarks (AIGCDetectBenchmark, GenImage, and Chameleon) show that DADL achieves state-of-the-art performance, with particular strengths in detecting high-fidelity synthetic images from diffusion models (e.g., Midjourney, SDv1.4, SDv1.5) and enhancing generalization across diverse generative paradigms. This study not only offers an effective approach for AIGI detection but also sheds light on the intrinsic properties of synthetic images, providing a new perspective for advancing AIGI forensics.
Chunli Song, Peiyang Wang, Guixuan Zhang, Shuwu Zhang
AAAI7
2026 Keypoint-enhanced image watermarking with spatial-frequency mapping and perceptual optimization
Fei Ge, Jie Liu 0028, Guixuan Zhang, Shuwu Zhang, Hu Guan
Inf. Sci.7
2025 Speech2Face3D: A Two-Stage Transfer-Learning Framework for Speech-Driven 3D Facial Animation
abstract
ABSTRACT High‐fidelity, speech‐driven 3D facial animation is crucial for immersive applications and virtual avatars. Nevertheless, advancement is impeded by two principal challenges: (1) a lack of high‐quality 3D data, and (2) inadequate modelling of the multi‐scale characteristics of speech signals. In this paper, we present Speech2Face3D, a novel two‐stage transfer‐learning framework that pretrains on large‐scale pseudo‐3D facial data derived from 2D videos and subsequently finetunes on smaller yet high‐fidelity 3D datasets. This design leverages the richness of easily accessible 2D resources while mitigating reconstruction noise through a simple temporal smoothing step. Our approach further introduces a Multi‐Scale Hierarchical Audio Encoder to capture subtle phoneme transitions, mid‐range prosody, and longer‐range emotional cues. Extensive experiments on public 3D benchmarks demonstrate that our method achieves state‐of‐the‐art performance on lip synchronization, expression fidelity, and temporal coherence metrics. Qualitative user evaluations validate these quantitative improvements. Speech2Face3D is a robust and scalable framework for utilizing extensive 2D data to generate precise and realistic 3D facial animations only based on speech.
Liming Pang, Guixuan Zhang, Shuwu Zhang
IET Image Process.5
2025 Weakening the Dominant Role of Text: CMOSI Dataset and Multimodal Semantic Enhancement Network
abstract
Multimodal sentiment analysis (MSA) is important for quickly and accurately understanding people's attitudes and opinions about an event. However, existing sentiment analysis methods suffer from the dominant contribution of text modality in the dataset; this is called text dominance. In this context, we emphasize that weakening the dominant role of text modality is important for MSA tasks. To solve the above two problems, from the perspective of datasets, we first propose the Chinese multimodal opinion-level sentiment intensity (CMOSI) dataset. Three different versions of the dataset were constructed: manually proofreading subtitles, generating subtitles using machine speech transcription, and generating subtitles using human cross-language translation. The latter two versions radically weaken the dominant role of the textual model. We randomly collected 144 real videos from the Bilibili video site and manually edited 2557 clips containing emotions from them. From the perspective of network modeling, we propose a multimodal semantic enhancement network (MSEN) based on a multiheaded attention mechanism by taking advantage of the multiple versions of the CMOSI dataset. Experiments with our proposed CMOSI show that the network performs best with the text-unweakened version of the dataset. The loss of performance is minimal on both versions of the text-weakened dataset, indicating that our network can fully exploit the latent semantics in nontext patterns. In addition, we conducted model generalization experiments with MSEN on MOSI, MOSEI, and CH-SIMS datasets, and the results show that our approach is also very competitive and has good cross-language robustness.
Ming Yan 0005, Guangzhe Zhao, Guixuan Zhang, Shuwu Zhang
IEEE Trans. Neural Networks Learn. Syst.6
2024 A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
Zeyu Zhao 0005, Nan Gao 0001, Guixuan Zhang, Jie Liu 0028, Shuwu Zhang
MMAsia6
2024 Degradation regression with uncertainty for blind super-resolution
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang
Neurocomputing6
2023 Learning Video Localization on Segment-Level Video Copy Detection with Transformer
Chi Zhang 0093, Jie Liu 0028, Shuwu Zhang
ICANN (7)3
2023 Adversarial Audio Watermarking: Embedding Watermark into Deep Feature
abstract
Audio watermarking is a promising technology for copyright protection, yet traditional methods are limited that must be combined with auxiliary techniques against attacks. This article proposes a new audio watermarking method that embeds watermarks through a trained neural network. It adds small imperceptible perturbations to the original audio so that its deep features point to specific watermark features. Data augmentation and error correcting coding are employed to guarantee its practicable robustness. This method is robust against many attacks without auxiliary techniques and shows better performance than other deep learning-based methods.
Shiqiang Wu, Jie Liu 0028, Hu Guan, Shuwu Zhang
ICME5
2023 Robust Texture-Aware Local Adaptive Image Watermarking With Perceptual Guarantee
abstract
Watermarking involves embedding a watermark in an image and later extracting it to prove the image’s copyright. In most cases, a complete image contains both smooth and textured regions. As a rule of thumb, the visual quality of an image with a watermark embedded in its textured regions is better than that of the same image with a watermark in smooth regions. This paper, by taking advantage of the fact, proposes a texture-aware local adaptive watermarking algorithm to maximize the watermark’s robustness while maintaining its imperceptibility. To identify textured regions in an image, we introduce the texture value, an efficient and proper metric of the richness of image texture. It combines the texture correlation of the AC coefficients, the luminance masking of the DC coefficient, and the distribution of image texture. A watermark is embedded adaptively into multiple non-overlapping textured regions of an image under the specified SSIM condition. Its adaptiveness comes from a novel texture-aware adaptive parameter model derived by multivariate regression analysis. Correct extraction of watermarks from multiple textured regions can be done by the cooperation of embedding and extraction strategies, with the assistance of RS-based watermark coding model. They allow for greater robustness, faster extraction, and adjustable watermark capacity. The simulation experiments on 100 images demonstrate that our proposed algorithm outperforms state-of-the-art algorithms with respect to imperceptibility, robustness, and adaptability.
Hu Guan, Jie Liu 0028, Shuwu Zhang, Baoning Niu, Guixuan Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2022 Heterogeneous Avatar Synthesis Based on Disentanglement of Topology and Rendering
Guixuan Zhang, Shuwu Zhang
ACCV (4)4
2022 From general to specific: Online updating for blind super-resolution
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang
Pattern Recognit.6
2021 Approaching the Limit of Image Rescaling via Flow Guidance
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang
BMVC6
2021 Learning to predict more accurate text instances for scene text detection
Jie Liu 0028, Guixuan Zhang, Yang Zheng 0002, Shuwu Zhang
Neurocomputing6
2021 Detection of Fake Reviews Using Group Model
Yuejun Li, Fangxin Wang 0002, Shuwu Zhang, Xiaofei Niu
Mob. Networks Appl.3
2020 IBN-STR: A Robust Text Recognizer for Irregular Text in Natural Scenes
abstract
Although text recognition methods based on deep neural networks have promising performance, there are still challenges due to the variety of text styles, perspective distortion, text with large curvature, and so on. To obtain a robust text recognizer, we have improved the performance from two aspects: data aspect and feature representation aspect. In terms of data, we transform the input images into S-shape distorted images in order to increase the diversity of training data. Besides, we explore the effects of different training data. In terms of feature representation, the combination of instance normalization and batch normalization improves the model's capacity and generalization ability. This paper proposes a robust scene text recognizer IBN-STR, which is an attention-based model. Through extensive experiments, the model analysis and comparison have been carried out from the aspects of data and feature representation, and the effectiveness of IBN-STR on both regular and irregular text instances has been verified. Furthermore, IBN-STR is an end-to-end recognition system that can achieve state-of-the-art performance.
Jie Liu 0028, Guixuan Zhang, Shuwu Zhang
ICPR4
2020 Single shot multi-oriented text detection based on local and non-local features
Jie Liu 0028, Shuwu Zhang, Guixuan Zhang, Yang Zheng 0002
Int. J. Document Anal. Recognit.3
2019 Enhancing Image Watermarking With Adaptive Embedding Parameter and PSNR Guarantee
abstract
Watermarking plays an important role in identifying the copyright of an image and related issues. The state-of-the-art watermark embedding schemes, spread spectrum and quantization, suffer from host signal interference (HSI) and scaling attacks, respectively. Both of them use a fixed embedding parameter, which is difficult to take both robustness and imperceptibility into account for all images. This paper solves the problems by proposing two novel blind watermarking schemes: a spread spectrum scheme with adaptive embedding strength (SSAES) and a differential quantization scheme with adaptive quantization threshold (DQAQT). Their adaptiveness comes from the proposed adaptive embedding strategy (AEP), which maximizes the embedding strength or quantization threshold by guaranteeing the peak signal-to-noise ratio (PSNR) of the host image after embedding the watermark, and strikes the balance between robustness and imperceptibility. SSAES is HSI free by factoring in the priori knowledge about HSI. In DQAQT, an effective quantization mode is proposed to resist scaling attacks by utilizing the difference between two selected DCT coefficients with high stability. Both SSAES and DQAQT can be easily applied to other watermarking frameworks. We introduce a notion called error threshold to theoretically analyze the performance of our proposed methods in details. The experimental results consistently demonstrate that SSAES and DQAQT outperform the state-of-the-art methods in terms of imperceptibility, robustness, computational cost, and adaptability.
Baoning Niu, Hu Guan, Shuwu Zhang
IEEE Trans. Multim.4
2018 Aspect-Level Sentiment Classification with Conv-Attention Mechanism
Jie Liu 0028, Guixuan Zhang, Shuwu Zhang
ICONIP (4)4
2017 Region based image retrieval with query-adaptive feature fusion
abstract
Recently, image representation based on convolutional neural network (CNN) becomes more popular than SIFT based feature, such as Fisher vector (FV). However, which of the two works better for image retrieval is not entirely clear yet. In this paper, we propose to fuse CNN and FV to incorporate the advantages of both features for image retrieval. We extract CNN feature and FV from multi-scale regions, which makes the representation more robust to image noise. Then a query-adaptive feature fusion method is proposed, which is used jointly with 2-D inverted index under the framework of bag-of-words. Moreover, we make an evaluation of different CNN feature extraction methods for the region based method. Extensive experiments on four benchmark datasets demonstrate the effectiveness of our method with efficiency in both time cost and memory usage.
Guixuan Zhang, Shuwu Zhang, Hu Guan, Fangxin Wang 0002
ICIP2
2017 SIFT Matching with CNN Evidences for Particular Object Retrieval
Guixuan Zhang, Shuwu Zhang, Wanchun Wu
Neurocomputing3
2017 A cascaded method for text detection in natural scene images
Yang Zheng 0002, Qing Li 0015, Jie Liu 0028, Heping Liu, Shuwu Zhang
Neurocomputing6
2016 Region matching and similarity enhancing for image retrieval
abstract
Many image retrieval systems adopt the bag-of-words model and rely on matching of local descriptors. However, these descriptors of keypoints, such as SIFT, may lead to false matches, since they do not consider the contextual information of the keypoints. In this paper, we incorporate the cues of meaningful regions where local descriptors are extracted. We describe a matching region estimation (MRE) method to find appropriate matching regions for local descriptor matching pairs. Then the region matching quality is evaluated and the true matched regions will enhance the similarity of local descriptors. Consequently, the image retrieval accuracy can be improved. Extensive experiments on benchmark datasets show the effectiveness of our method and our result compares favorably with the state-of-the-art.
Guixuan Zhang, Shuwu Zhang, Hu Guan, Qin-Zhen Guo
ICASSP3
2016 Scene text detection with extremal region based cascaded filtering
abstract
In this paper, we present a robust Extremal Region (ER) based scene text detection system. To eliminate the vast non-text components generated by ER operator, a three-stage cascaded filter is proposed. In the first stage, a powerful character classifier enhanced by recursive local search is introduced to separate text components from noises. Then, an efficient heuristic pruning method is designed to further clean overlapped duplicate characters. Finally, after text line construction, a cascaded text line classification model integrating word entropy and sliding window based CNN is proposed to remove false text lines. Experiments on benchmarks show that our method achieves state-of-the-art performance.
Jie Liu 0028, Shuwu Zhang, Yang Zheng 0002
ICIP3
2016 Adaptive bit allocation product quantization
Qin-Zhen Guo, Shuwu Zhang, Guixuan Zhang
Neurocomputing3
2015 Transmitting informative components of fisher codes for mobile visual search
abstract
Existing techniques usually adopt compact descriptors such as Fisher vector for mobile visual search, since compact descriptors are memory-efficient and suitable for fast transmission. In common Fisher vector methods, in order to make the size of image representations small enough for efficient transmission, only a small number of visual words are used. However, this choice usually sacrifices the search accuracy. In this paper, a Soft-Assignment Adjusting approach is proposed to just select informative components of descriptors for query. With this method, we can adopt more visual words to improve accuracy, while the memory usage is still low. Furthermore, efficient bitrate scalable codes are proposed in order to accommodate the network bandwidth variation. Experiments performed on benchmark datasets show that our proposed approach outperforms the state-of-the-art methods for mobile visual search.
Guixuan Zhang, Shuwu Zhang, Qin-Zhen Guo
ICASSP3
2015 Adaptive bit allocation hashing for approximate nearest neighbor search
Qin-Zhen Guo, Shuwu Zhang
Neurocomputing3
2014 Real-Time Event Detection Based on Geo Extraction and Temporal Analysis
Shuwu Zhang, Wei Liang 0009, Zhe Tu
ADMA2
2014 Individualized matching based on logo density for scalable logo recognition
abstract
Although many systems based on global or local descriptors have shown promising results for logo recognition, they have handled all logos with the same structure and not considered their diversities. Therefore, with the logo scale increasing, the general way cannot recognize each logo perfectly. To overcome this limitation, we propose a novel strategy to match query and each logo individually using these features. First, a new conception named logo density is introduced as important semantic information for logos. Second, matching density is given according to the logo density and by utilizing it in logistic function an individualized matching strategy is developed to obtain accurate similarity for query and a logo. Finally, we present a fast recognition algorithm based upon bag-of-words model to realize scalable logo recognition. Our method is evaluated on two challenging datasets (our 10,000-class logo dataset and FlickrLogos-27). Experiments demonstrate its superior performance comparing to previous methods.
Shuwu Zhang, Wei Liang 0009, Qin-Zhen Guo
ICASSP2
2014 Multi-feature hierarchical topic models for human behavior recognition
Heping Li, Shuwu Zhang
Sci. China Inf. Sci.3
2013 Adaptive bit allocation hashing for approximate nearest neighbor search
abstract
Using hashing algorithms to learn binary codes representation of data for fast approximate nearest neighbor (ANN) search has attracted more and more attentions. Most existing hashing methods employ various hash functions to encode data. The resulting binary codes can be obtained by concatenating bits produced by those hash functions. These methods usually have two main steps: projection and thresholding. One problem of these methods is that every dimension of the projected data is regarded as the same importance and represented by one bit, which may result in ineffective codes. We introduce an adaptive bit allocation hashing (ABAH) method to encode data for ANN search. The basic idea is, according to the dispersion of every dimension after projection we use different number of bits to encode them. ABAH can effectively preserve the neighborhood structure in the original data space. Extensive experiments show that ABAH significantly outperforms three state-of-the-art methods.
Qin-Zhen Guo, Shuwu Zhang, Fangyuan Wang 0003
ICME3
2013 Chinese Short Text Classification Based on Domain Knowledge
Chengyong Liu, Wei Liang 0009, Shuwu Zhang
IJCNLP5
2013 Large Scale Image Retrieval with Practical Spatial Weighting for Bag-of-Visual-Words
Fangyuan Wang 0003, Heping Li, Shuwu Zhang
MMM (1)4
2012 Spatial connected component pre-locating algorithm for rapid logo detection
abstract
This paper introduces a novel pre-locating algorithm for rapid logo detection in unconstrained color images. This work is distinguished by two major contributions. The first is a new method of representation for logo called “spatial connected component descriptor” (SCCD) containing connected component (CC) prediction model and effective-CC pixel distribution histogram. The former represents combinations between CCs based on color and spatial relationships of CCs. While the latter describes the pixel distribution information of effective CCs. The two parts capture the layout of logos from different points. The second is a logo pre-locating algorithm by the means of SCCD to search for logo prediction regions in test images, on which some content-based features are used for logo matching. Experimental results illustrate that our pre-locating algorithm speeds up logo detection to a great extent and shows precise location compared to previous systems.
Shuwu Zhang, Wei Liang 0009
ICASSP2
2011 Hierarchical Latent Dirichlet Allocation models for realistic action recognition
abstract
It has always been very difficult to recognize realistic actions from unconstrained videos because there are tremendous variations from camera motion, background clutter, object appearance and so on. In this paper, a Single-Feature Hierarchical Latent Dirichlet Allocation model called SF-HLDA by extending Latent Dirichlet Allocation to the hierarchical one is first proposed for realistic action recognition. And then, by extending SF-HLDA, we present another model called Multi-Feature Hierarchical Latent Dirichlet Allocation model MF-HLDA which can effectively fuse several different features into one model for recognizing the realistic actions. Experiments demonstrate the effectiveness of our proposed models.
Heping Li, Jie Liu 0028, Shuwu Zhang
ICASSP3
2011 A hierarchical generative model for Generic Audio Document Categorization
abstract
In this paper, we call the pattern classification problem that consists in assigning a category label to a long audio signal based on its semantic content as Generic Audio Document Categorization (GADC). A novel generative model is proposed to describe the generic audio document categories and solve the GADC problem. This model is a four-level hierarchical model in which two latent variables "audio topic" and "audio word" are introduced in addition to the two observed variables category and audio feature. We present an iterative learning algorithm including two Expectation-Maximization (EM) cycles to estimate the model parameters and give a discriminative document weighting procedure to make the model more discriminative. Subsequently, the distribution of "audio topic" in the well-trained model is utilized to represent each generic audio document category. This is same with some bag-of-word methods. However, our method is advanced since it does not require quantizing the continuous audio features to a vocabulary of "audio words". Finally, experiment results show the effectiveness of our approach.
Shuwu Zhang
ICASSP2
2011 A Novel Italic Detection and Rectification Method for Chinese Advertising Images
abstract
The italic detection and slant rectification is a key step of optical character recognition (OCR). In this paper, a novel method is proposed to detect and rectify italic characters in Chinese advertising images. Based on observations on structures of many characters, the centroid angle is proposed and a statistical study on it is presented. According to the statistical results, the centroid angle of a Chinese character approximately obeys a Gaussian distribution with its slant angle. Moreover, a Markov Random Field (MRF) model, considering the font-face similarity of neighboring characters and the strong correlation between the centroid angle and the slant angle of a character, is then presented to estimate the slant angle of a character. The italic characters can be detected and rectified by the estimated angle. The experimental results demonstrate the proposed method is effective and applicable.
Jie Liu 0028, Heping Li, Shuwu Zhang, Wei Liang 0009
ICDAR3
2011 A Chinese Character Localization Method Based on Intergrating Structure and CC-Clustering for Advertising Images
abstract
In this paper, a novel Chinese character localization method is proposed for texts in advertising images. To deal with the texts with gradient color, a color clustering method based on edge is introduced to separate the color image into homogeneous color layers. To solve the problem of locating characters varied in size, style and arranged in irregular direction, a novel character localization method is proposed, which integrates structure and CC-clustering to locate characters according to reliable features of characters. Finally, a new noise removal method based on stroke width histogram is employed to remove all non-characters connected components, and then all characters are located. The experimental results show that the proposed method can effectively locate characters in advertising images.
Jie Liu 0028, Shuwu Zhang, Heping Li, Wei Liang 0009
ICDAR2
2011 Automatic behavior model selection by iterative learning and abnormality recognition
abstract
Automatic behavior recognition is one important task of community security and surveillance system. In this paper, a novel method is proposed for automatic selection of behavior models by iterative learning and abnormality recognition. The method is mainly composed of the following two steps: (1) The models of normal behaviors are automatically selected and trained by combining Dynamic Time Warping based spectral clustering and iterative learning; (2) Maximum A Posteriori adaptation technique is used to estimate the parameters of abnormal behavior models from those of normal behavior models. Compared with the related works in the literature, our method has three advantages: (1) automatic selection of the class number of normal behaviors from large unlabeled video data according to the process of iterative learning, (2) semi-supervised learning of abnormal behavior models, and (3) avoidance of the running risk of over-fitting during learning the Hidden Markov Models of behaviors in case of sparse data. Experiments demonstrate the effectiveness of our proposed method.
Heping Li, Jie Liu 0028, Shuwu Zhang
ISI3
2011 Global and Local Features based topic model for scene recognition
abstract
This paper presents a novel Global and Local Features based Latent Dirichlet Allocation model for scene recognition. The proposed model follows the bag-of-word framework like the Latent Dirichlet Allocation model. The traditional Latent Dirichlet Allocation model for scene recognition only uses the orderless bag of features called global features without considering spatial constraints on these features. Different from this model, our proposed model can combine both global features and local region features for improving the recognition performance. In our method, local region features are gotten by adding a simple spatial constraint on the orderless bag of features. Experiments on three scene datasets demonstrate the effectiveness of our proposed model.
Heping Li, Fangyuan Wang 0003, Shuwu Zhang
SMC3
2010 A local appearance contextual descriptor for object matching
abstract
We present a novel approach to measuring similarity between objects based on matching local “appearance contextual descriptor”. The descriptor has two components: Histogram of Oriented Gradient feature representing local patch appearance and the contextual descriptor capturing not only the spatial distribution of the non-reference patches relative to the reference patch but also the appearance similarities between the reference patch and the non-reference patches in the region. Corresponding patches within two similar objects will have similar contextual descriptors, though the patch appearances may have some difference. We treat recognition in a nearest-neighbor classification framework and match object in regions with no prior learning. We compare our method to commonly used methods and demonstrate its applicability to object detection and recognition.
Xiaozhen Xia, Shuwu Zhang, Wei Liang 0009
ICASSP2
2010 Similarity-based image classification via kernelized sparse representation
abstract
We consider the image classification problem based on the similarities between images. The choice of the similarity is related to the particular applications, and it could be based on color, texture, bag-of-features, or even more complex kernels. As long as the pair-wise similarity matrix is transformed into a positive semidefinite one, the similarities of images could be treated as kernels. This transformation makes it possible for kernel methods to solve the similarity-based image classification problem. In this paper, we propose a novel kernelized classification framework based on sparse representation. This new framework casts the classification as finding a sparse linear representation of test image with respect to training images. Unlike the former works, we do this sparse coding procedure through a proposed kernelized orthogonal matching pursuit algorithm, which is performed in inner product space rather than Euclidean space. Through a proper choice of the similarity function, the proposed approach can be applied to diverse image classification problems. Comparative experiments between the proposed method and other existing methods, on two real datasets (Caltech-101 and Face Rec) show that our method performed better.
Heping Li, Wei Liang 0009, Shuwu Zhang
ICIP4
2010 A Fast Image Inpainting Method Based on Hybrid Similarity-Distance
abstract
A fast image in painting method based on hybrid similarity-distance is proposed in this paper. In Criminisi et al.'s work, similarity distance are not reliable enough in many cases and the algorithm performs inefficiently. To solve these problems, we propose a new searching strategy to accelerate the algorithm. In addition, we modify the confidence-updating rule to make more reasonable the distributions of the confidences in source region. Besides, taking account of the stationarity of texture and the reliability of the source regions, we present a hybrid similarity-distance, which combines the distance in color space with the distance in spatial space by weight coefficients related to the confidence value. A more reasonable patch will be found out by this hybrid similarity-distance. The experiments verify that the proposed method yields qualitative improvements compared to Criminisi et al.'s work.
Jie Liu 0028, Shuwu Zhang, Wuyi Yang, Heping Li
ICPR2
2009 Part-Based Object Detection Using Cascades of Boosted Classifiers
Xiaozhen Xia, Wuyi Yang, Heping Li, Shuwu Zhang
ACCV (2)4
2009 A novel approach to musical genre classification using probabilistic latent semantic analysis model
abstract
A novel approach based on the probabilistic latent semantic analysis model (pLSA) for automatic musical genre classification is proposed in this paper. Unlike traditional usage, the pLSA is used to model musical genre instead of single music signal in the proposed approach. First, an unsupervised clustering algorithm is utilized to group temporal segments in music signals into several natural clusters. By this means, each music signal is decomposed into a bag of ldquoaudio wordsrdquo. Subsequently, the pLSA model of each musical genre is trained through a new iterative training procedure and well-known EM algorithm. This training procedure can iteratively update the pLSA model parameters by discriminatively computing weight of each training music signal and evidently improve the model's discriminative performance. Finally, these models can be used to classify new unseen music signals. Experiments on two commonly utilized databases show that our pLSA based approach can give promising results and the iterative learning procedure is effective.
Shuwu Zhang, Heping Li, Wei Liang 0009, Haibo Zheng
ICME2
2008 A Graph Based Subspace Semi-supervised Learning Framework for Dimensionality Reduction
Wuyi Yang, Shuwu Zhang, Wei Liang 0009
ECCV (2)2
2006 A Question Answering System on Special Domain and the Implementation of Speech Interface
Fuji Ren, Shingo Kuroiwa, Shuwu Zhang
CICLing4
2006 A Mongolian Speech Recognition System Based on HMM
Guanglai Gao, Biligetu, Nabuqing, Shuwu Zhang
ICIC (2)4
2006 Fast SVM training based on the choice of effective samples for audio classification
Shilei Zhang, Hongchen Jiang, Shuwu Zhang, Bo Xu 0002
INTERSPEECH3
2006 A quality measure method using Gaussian mixture models and divergence measure for speaker identification
Rong Zheng 0005, Shuwu Zhang, Bo Xu 0002
INTERSPEECH2
2005 Optimal model order selection based on regression tree in speaker identification
Shilei Zhang, Junmei Bai, Shuwu Zhang, Bo Xu 0002
INTERSPEECH3
2004 Chinese-English bilingual phone modeling for cross-language speech recognition
abstract
In this paper, three different approaches to Chinese-English bilingual phone modeling are investigated and compared. The first approach is to simply combine Chinese and English phone inventories together without phone sharing across the languages. The second one is to map language-dependent phones to the inventory of the International Phonetic Association (IPA) based on phonetic knowledge to construct the bilingual phone inventory. The third one is to merge the language-dependent phone models by an hierarchical phone clustering algorithm to get a compact bilingual inventory. In the third approach, two distance measures are used to perform the bottom-up clustering. One is the Bhattacharyya distance. The other is the acoustic likelihood distance. Experimental results show that the phone clustering approach outperforms the IPA-based phone mapping approach, and it can also achieve comparable performance to the simple combination of language-dependent phone inventories with fewer model parameters, especially when using acoustic likelihood distance measurement.
Shengmin Yu, Shuwu Zhang, Bo Xu 0002
ICASSP (1)2
2004 A novel target-driven generalized JMAP adaptation algorithm
abstract
Adapting the parameters of a statistical speaker independent continuous speech recognizer to the speaker can significantly improve the recognition performance and robustness of the system. In this paper, we propose a novel target-driven speaker adaptation method, Generalized Joint Maximum a Posteriori (GJMAP), which extends and improves the previous successful method JMAP. GJMAP partitions the HMM parameters with respect to the adaptation data, using the priori phonetic knowledge. The generation of regression class trees is dynamically constructed on the target-driven principle in order to obtain the maximum increase of the auxiliary function. An off-line adaptation experiment on large vocabulary continuous speech recognition is carried out. The experimental results show GJMAP has more advantages than the conventional methods.
Zhaobing Han, Shuwu Zhang, Bo Xu 0002
INTERSPEECH2
2004 Combining agglomerative and tree-based state clustering for high accuracy acoustic modeling
Zhaobing Han, Shuwu Zhang, Bo Xu 0002
INTERSPEECH2
2004 Multi-layer structure MLLR adaptation algorithm with subspace regression classes and tying
Xiangyu Mu, Shuwu Zhang, Bo Xu 0002
INTERSPEECH2
2003 Comparison and study of some variants of partially tied covariance modeling
abstract
Some practical implementation issues on partially tied covariance (PTC) modeling are discussed. First, from the view of model complexity and computational load, a comparison is made for some variants of PTC. From the analysis, two representatives, STC and Ortho-STC are compared in detail. Second, based on these variants, two techniques are studied. One technique is joint optimization of both transformation and HMM parameters, which will exploit the potential of PTC. The other technique is model selection by hierarchical tree via Bayesian information criterion (BIC), which will decide the number and structure of transformation classes thus to assure the generalization capacity. Experiment results showed that STC always outperforms Ortho-STC due to the effect of parameter tying and by the application of above two techniques the system performance can be much improved.
Peng Ding 0003, Shuwu Zhang, Bo Xu 0002
ICASSP (1)2
2003 A vector statistical piecewise polynomial approximation algorithm for environment compensation in telephone LVCSR
abstract
A vector statistical piecewise polynomial (VPP) approximation algorithm is proposed for environment compensation in speech signals that are degraded by both additive and convolutive noise. By investigating the model of the telephone environment, we address a piecewise polynomial, namely two linear polynomials and a quadratic polynomial, to approximate the environment function precisely. The VPP is applied either to stationary noise, or to non-stationary noise. In the first case, batch EM is used in the log-spectral domain; in the second case, recursive EM with iterative stochastic approximation is developed in the cepstral domain. Both approaches are based on the minimum mean squared error (MMSE) sense. Experimental results are presented on the application of this approach in improving the performance of Mandarin large vocabulary continuous speech recognition (LVCSR) in background noise and different transmission channels (such as fixed telephone line and GSM). The method can reduce the average character error rate (CER) by about 18%.
Zhaobing Han, Shuwu Zhang, Huayun Zhang, Bo Xu 0002
ICASSP (2)2
2003 Discriminative optimization of large vocabulary Mandarin conversational speech recognition system
Peng Ding 0003, Zhenbiao Chen, Shuwu Zhang, Bo Xu 0002
INTERSPEECH4
2003 Statistical speech-to-speech translation with multilingual speech recognition and bilingual-chunk parsing
Bo Xu 0002, Shuwu Zhang, Chengqing Zong
INTERSPEECH2
2001 A hybrid approach to enhance task portability of acoustic models in Chinese speech recognition
abstract
This paper presents our approach to enhance the portability of acoustic models by mitigating the phonetic mismatch arising from a new testing task which is rather different from the training data. The approach is a hybrid one which combines knowledge-based context categorization to generate a context rich set of subword units, and data-driven-based acoustic model clustering on the level of context category. Compared with the conventional approach of only phonetic decision tree based model clustering and unseen model generation, the new approach improved greatly the desired subword coverage for the new testing domain, and achieved an error rate reduction by 10.8% for Chinese character accuracy in the recognition experiments. Together with the effect of the newly adopted basic units of 9 glottal stops, we achieved a total 23.5% error rate reduction in the testing compared to the baseline system.
Jinsong Zhang 0001, Shuwu Zhang, Yoshinori Sagisaka, Satoshi Nakamura 0001
INTERSPEECH2
2000 An embedded knowledge integration for hybrid language modelling
Shuwu Zhang, Hirofumi Yamamoto, Yoshinori Sagisaka
INTERSPEECH1
1999 Improving n-gram modeling using distance-related unit association maximum entropy language modeling
abstract
In this paper, a distance-related unit association maximum entropy (DUAME) language modeling is proposed. This approach can model an event (unit subsequence) using the co-occurrence of full distance unit association (UA) features so that it is able to pursue a functional approximation to higher order N-gram with significantly less memory requirement. A smoothing strategy related to this modeling will also be discussed. Preliminary experimental results have shown that DUAME modeling is comparable to conventional N-gram modeling in perplexity with significantly small number of parameters.
Shuwu Zhang, Harald Singer, Dekai Wu, Yoshinori Sagisaka
EUROSPEECH1
1997 An integrated language modeling with n-gram model and WA model for speech recognition
abstract
As to traditional n-gram model, smaller n value is an inherent defect for estimating language probabilities in speech recognition, simply because that estimation could not be executed over farther word association but by means of short sequential word correlated information. This has an strong effect on the performance of speech recognition. This paper introduces an integrated language modeling with ngram model and word association model (abbreviated as WA model). This model integrated two kind of joint probabilities, traditional n-gram probability and word association probability, to estimate actual output probability. WA model are based on a combined probability estimation of orderly word association without distant and strict sequential limitation. In addition, two kinds of local linguistic constraints have also been incorporated into n-gram estimation for smoothing date sparse and adjusting special language unit score locally. A substantial improvement for the performance of Chinese phonetic-to-text transcription in speech recognition has been obtained.
Shuwu Zhang, Taiyi Huang 0001
EUROSPEECH1
1996 Speaker-independent dictation of Chinese speech with 32k vocabulary
abstract
While early machines adopted isolated syllable as input units and needed boring enrollment, our research focus on the speaker-independent, word-based dictation.A deliberately designed 120-speaker database was built for training ; inter-syllable context ,tonal and endpoint dependent acoustic model are applied with promising MFCC feature; Two-pass acoustic matching accelerates the recognition making fully advantage of the monosyllabic structure of Chinese speech; A complete word bigram and trigram serve as language processing module.With all efforts, the system reaches 90% character accuracy performing in almost real-time on Pentium PC without DSP help.
Bo Xu 0002, Shuwu Zhang, Fei Qu, Taiyi Huang 0001
ICSLP3