VLDB 2026 Research / reviewers in the wild / expert
Zhenjiang Miao
dblp:98/6067
· DBLP profile ↗
117ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0001-8032-5769ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 88 · 13 since 2021Artificial intelligence and machine learning · 34 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debiased Multi-modal Vision-Language Network for Action Quality Assessment
Ruizhao Zhai, Wanru Xu, Zhenjiang Miao, Qinghao Kong, Ruiying Yao |
ICIC (16) | 3 |
| 2026 | Reinforced Multi-expert Ensemble Strategy for Weakly Supervised Video Anomaly Detection
Wanru Xu, Zhenjiang Miao, Ruiying Yao |
ICPR (7) | 3 |
| 2026 | Causal learning with uncertainty-aware transformer for vision-and-language navigation
Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Wangsheng He |
Neurocomputing | 3 |
| 2026 | Natural Cognizing Video: A Decoupling and Integration Network for General Event Boundary CaptioningabstractVideo captioning is a prominent and challenging research area. Previous studies have focused predominantly on describing entire video segments, often overlooking the more significant status changes within these segments. We propose a novel decoupling and integration network for general event boundary captioning (DIN-GEBC), which is applied to the Kinetics-GEBC dataset, a video dataset with fine-grained status descriptions. DIN-GEBC focuses on generating three types of captions for each video segment boundary: a caption for the dominant subject (subject caption), a caption describing the event prior to the status moment (status-before), and a caption describing the event after the status moment (status-after). To address different descriptive focuses and characteristics, DIN-GEBC is proposed for decoupling and integrating both tasks and features. For task decoupling, DIN-GEBC is designed with a dual branch structure, in which the generation of the subject caption is addressed by the dominant subject branch with a Video Q-former encoder and the generation of the status-before and status-after captions is addressed by the event branch with a Reinventing RNNs for the Transformer Era (RWKV) encoder. For task integration, DIN-GEBC enables the dominant subject branch to guide the event branch in producing detailed information regarding the subject experiencing the change. Feature disentanglement, in which the common features are used by the dominant subject branch to capture the unchanging information, is also performed, and the difference features are applied to the event branch to capture the changing information. The experimental results show that our model outperforms existing models on the Kinetics-GEBC dataset even with fewer parameters. The code is available athttps://github.com/Mrliu001219/DIN-GEBC. Wanru Xu, Zhenjiang Miao, Wangsheng He, Ruiying Yao |
IEEE Trans. Multim. | 3 |
| 2025 | Dual-Branch Diffusion Model for JPEG Artifact Correction
Wenhao Yu 0015, Wanru Xu, Zhenjiang Miao |
ICIC (1) | 4 |
| 2025 | InstructStep: Fine-Grained Localization of Step Content and Relation in Instructional VideoabstractExisting methods for video answer localization (VAL) in instructional video focus predominantly on coarse-grained themes, failing to address detailed step content and inter-step relation crucial for effective comprehension. Current datasets, such as MedVidQA, primarily capture video content but lack annotations for step structure and inter-step relation. To address this gap, we introduce InstructStep, a newly proposed VAL task, specifically designed for Instructional Video Step Content and Relation Localization. It extends original VAL task to step-centric content and relation. Accordingly, we create a InstructStep Dataset with fine-grained step content and relation QA pairs. To tackle the challenges of this task, we propose a Step-Centric Multi-Level Knowledge Distillation (SC-MLKD) approach that: (1) A two-stage training strategy that generates step-specific summaries in the first stage and introduces a step branch in the second stage to learn step relations. This is applied only during training, ensuring no additional inference time. (2) Multi-level knowledge distillation, including feature, relation, and response distillation, across visual, text and step branches to capture fine-grained and step-centric features. Comprehensive experiments demonstrate the efficiency of SC-MLKD, with notable gains of up to 5.83% in step content and up to 5.9% in step relation. The dataset have been made publicly available on https://github.com/hewangsh/InstructStep Wangsheng He, Wanru Xu, Zhenjiang Miao |
ACM Multimedia | 4 |
| 2025 | Causal Debiasing Network for Action Quality Assessment
Ruizhao Zhai, Wanru Xu, Zhenjiang Miao, Qinghao Kong |
PRCV (7) | 3 |
| 2025 | Knowledge-based and Data-driven Fusion for Unsupervised Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) has extensive applications in fields such as intelligent surveillance and autonomous driving. In the field of unsupervised VAD, data-driven methods based on pseudo-label generation have their advantages. They can utilize the characteristics of the data itself to generate pseudo-labels for model learning. However, the unsupervised setting leads to a lack of supervision information, resulting in a low confidence level of the supervision signal. On the other hand, prior knowledge can effectively supplement some supervision information. Nevertheless, prior knowledge is usually general and does not take into account the specific information of each sample. To address these limitations, this paper proposes an unsupervised video anomaly detection method that combines prior knowledge and data. This method takes unlabeled videos as input and learns to predict frame-level anomaly scores. The algorithm consists of a prior knowledge module and a data-driven module. The prior knowledge module calculates the degree of anomaly through normal propagation based on prior knowledge independent of the data. The data-driven module estimates the degree of anomaly through two branches, appearance and motion, and jointly generates pseudo-labels. Experiments conducted on the UCF-Crime and ShanghaiTech datasets, using frame-level AUC as the evaluation metric, show that the proposed method achieves state-of-the-art performance in the unsupervised category, outperforming existing one-class classification methods. Qinghao Kong, Wanru Xu, Zhenjiang Miao, Ruizhao Zhai, Wenhao Yu 0015 |
SMC | 3 |
| 2025 | CroCaps: A CLIP-assisted cross-domain video captioner
Wanru Xu, Yenan Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
Expert Syst. Appl. | 3 |
| 2025 | Counterfactual contrastive learning for weakly supervised temporal sentence grounding
Yenan Xu, Wanru Xu, Zhenjiang Miao |
Neurocomputing | 3 |
| 2025 | RESTHT: relation-enhanced spatial-temporal hierarchical transformer for video captioning
Lihuan Zheng, Wanru Xu, Zhenjiang Miao, Xinxiu Qiu, Shanshan Gong |
Vis. Comput. | 3 |
| 2024 | Hypergraph Self-Attention and Channel Topology Specialization Network for Automatic Generation of Labanotation
Wanru Xu, Zhenjiang Miao |
ICPR (15) | 3 |
| 2024 | Probabilistic Distillation Transformer: Modelling Uncertainties for Visual Abductive ReasoningabstractVisual abduction reasoning aims to find the most plausible explanation for incomplete observations, and suffers from inherent uncertainties and ambiguities, which mainly stem from the latent causal relations, incomplete observations, and the reasoning itself. To address this, we propose a probabilistic model named Uncertainty-Guided Probabilistic Distillation Transformer (UPD-Trans) to model uncertainties for Visual Abductive Reasoning. In order to better discover the correct cause-effect chain, we model all the potential causal relations into a unified reasoning framework, thus both the direct relations and latent relations are considered. In order to reduce the effect of the stochasticity and uncertainty for reasoning: 1) we extend the deterministic Transformer to a probabilistic Transformer by considering those uncertain factors as Gaussian random variables and explicitly modeling their distribution; 2) we introduce a distillation mechanism between the posterior branch with complete observations and the prior branch with incomplete observations to transfer posterior knowledge. Evaluation results on the benchmark datasets, consistently demonstrate the commendable performance of our UPD-Trans, with significant improvements after latent relation modeling and uncertainty modeling. Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
ACM Multimedia | 2 |
| 2023 | LabanFormer: Multi-scale graph attention network and transformer with gated recurrent positional encoding for labanotation generation
Zhenjiang Miao, Yuanyao Lu |
Neurocomputing | 2 |
| 2022 | Sequential Gesture Learning for Continuous Labanotation Generation Based on the Fusion of Graph Neural NetworksabstractLabanotation is a symbolic recording system for human movements, and also a powerful tool for protecting and spreading folk dances and other performing arts. State-of-the-art automatic Labanotation uses end-to-end methods with sequence-based skeleton representation, which cannot capture the relationship between joints and bones in the skeleton for accurate descriptions of continuous lower limb movements such as dance steps. In this paper, we propose a novel double-stream fusion method of directed graph neural networks (DGNN), combined with connectionist temporal classification (CTC), namely DFGNN-CTC, for sequential fine-grained motion recognition, such as the Labanotation generation of unsegmented dance movement. First, we extract double-stream directed graph feature, employing an orientation-normalized directed acyclic graph (ON-DAG) and an orientation-normalized temporal directed acyclic graph (ON-TDAG), to jointly model spatiotemporal properties of movement recorded in motion capture data. Then, we design a CTC-based fusion-pooling module to fuse the spatial and temporal streams encoded by two DGNNs. It concatenates and fuses the two streams to generate discriminative descriptions of each time step, and concentrates them to make per-time-step predictions of Laban gesture type, from which the CTC searches the optimal Laban symbol sequence, corresponding to elemental motions composing the movement. In this way, the new method enables much finer discrimination for similar Laban gestures with subtle differences in spatial and temporal properties through joint contextual spatiotemporal modeling so that it achieves much superior performance in continuous Labanotation generation to existing methods, which only have single-stream analysis either spatially or temporally. The experiments on two Labanotation-labelled motion capture datasets demonstrate the effectiveness of the components in the proposed method and its superiority comparing with the state-of-the-art methods, especially for lower limb movements. Ningwei Xie, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu, Min Li 0025, Jiaji Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Bridging Video and Text: A Two-Step Polishing Transformer for Video CaptioningabstractVideo captioning is a joint task of computer vision and natural language processing, which aims to describe the video content using several natural language sentences. Nowadays, most methods cast this task as a mapping problem, which learns a mapping from visual features to natural language and generates captions directly from videos. However, the underlying challenge of video captioning,i.e., sequence to sequence mapping across the different domains, is still not well handled. To address these problems, we introduce the polishing mechanism in an attempt to mimic human polishing process and propose a generate-and-polish framework for video captioning. In this paper, we propose a two-step transformer based polishing network (TSTPN) consisting of two sub-modules: the generation-module is to generate the caption candidate and the polishing-module is to gradually refine the generated candidate. Specifically, the candidate provides a global information of the visual contents in a semantically-meaningful order, where it is firstly considered as a semantic intersnubber to bridge the semantic gap between the text and video, with the cross-modal attention mechanism for better cross-modal modeling; and it secondly provides a global planning ability to maintain the semantic consistency and fluency of the whole sentence for better sequence mapping. In experiments, we present adequate evaluations to show that the proposed TSTPN achieves the comparable and even better performance than the state-of-the-art methods on the benchmark datasets. Wanru Xu, Zhenjiang Miao, Jian Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Rhythm-Aware Sequence-to-Sequence Learning for Labanotation Generation With Gesture-Sensitive Graph Convolutional EncodingabstractLabanotation is a professional dance notation system widely used in dance education and choreography preservation. Automatically generatingLabanotation dance scores from motion capture data can save a huge amount of manual time and effort. Recently, the sequence-to-sequence (seq2seq) model is applied to the automatic Labanotation generation. This model is based on an encoder-decoder structure, which encodes the input motion sequence to a fixed-length vector and then decodes it to generate the target sequence. However, the encoding of spatial skeleton structure of motion data is not considered in the existing work. Besides, it is challenging to align between the input motion data and the output Laban symbol sequences due to the severe imbalance of sequence lengths. Therefore, in this paper, we present a new seq2seq model for more effective Labanotation generation. In the encoder, we propose a new gesture-sensitive graph convolutional network with learned adaptive joint weights and non-physical connections to learn both spatial and temporal patterns from motion data sequences. In the decoder, we exploit motion rhythm information and propose a novel rhythm-aware attention mechanism to learn a good alignment between motion sequences and Laban symbol sequences, so that we can focus on relevant parts of the input motion sequence without searching in the whole sequence when predicting a target Laban symbol. Extensive experiments on two real-world datasets show that the proposed method achieves a better performance compared with the state-of-the-art approaches on the task of automatic Labanotation generation. Min Li 0025, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu, Cong Ma 0004, Ningwei Xie |
IEEE Trans. Multim. | 2 |
| 2022 | Deep Domain Adaptation Based Multi-Spectral Salient Object DetectionabstractSalient Object Detection (SOD) plays an important role in many image-related multimedia applications. Although there are many existing research works about the salient object detection in traditional RGB (visible-light spectrum) images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we mainly model this research problem as a deep learning based domain adaptation from the traditional RGB image data (source domain) to the multi-spectral data (target domain), and an adversarial deep domain adaptation model is proposed. We first collect and will publicize a large multi-spectral dataset, RGBN-SOD dataset, including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. Intensive experimental results show the effectiveness and accuracy of the proposed deep domain adaptation for the multi-spectral SOD. Besides, due to the absence of research on the field of multi-spectral co-saliency detection, we also collect 200 synchronized RGB and NIR image pairs in addition to explore the multi-spectral co-saliency detection. Shaoyue Song, Zhenjiang Miao, Hongkai Yu, Jianwu Fang, Cong Ma 0004, Song Wang 0002 |
IEEE Trans. Multim. | 2 |
| 2021 | An Attention-Seq2Seq Model Based on CRNN Encoding for Automatic Labanotation Generation from Motion Capture DataabstractLabanotation is an important notation system widely used for recording dances. Numerous methods have been proposed for automatic Labanotation generation from motion capture data. Recently, the sequence-to-sequence (seq2seq) model is proposed. However, the encoder of the model only encodes the temporal information of motion data, lacking the encoding for spatial information. And it is challenging for the decoder to align input and output sequences due to the imbalance of the sequence lengths. In this paper, we propose an attention-seq2seq model based on Convolutional Recurrent Neural Network (CRNN). The proposed model employs an encoder based on CRNN to learn the spatial-temporal information of motion data and applies an attention mechanism to align each target Laban symbol with relevant parts of the input motion data in decoding. Experiments show that the proposed method performs favorably against state-of-the-art algorithms in the automatic Labanotation generation task. Min Li 0025, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu |
ICASSP | 2 |
| 2021 | Video Anomaly Detection Using Dual Discriminator Based Generative Adversarial NetworkabstractVideo anomaly detection is of great significance due to its wide applications in video surveillance. Recently, there is a trend of using a video prediction framework to tackle this problem. This kind of methods detects anomalies according to the difference between a predicted frame and its ground truth. However, existing prediction methods lack the consideration of multi-scale temporal constraints on generating future frames. Therefore, based on the prediction framework, this paper proposes a temporal enhanced anomaly detection approach, which designs a generative adversarial network with dual discriminator (frame discriminator and sequence discriminator) to predict future frames. In order to obtain more realistic predictions, other than commonly used spatial constraints and the adversarial penalty from the frame discriminator, we also consider both short-range and long-range motions to impose constraints. Specifically, for short-range motion modeling, we utilize the optical flow loss to ensure temporal continuity over two adjacent frames, while for long-range motion modeling, we design a sequence discriminator to identify fake contained sequences from real sequences, making frames more consistent with their previous consecutive frames. Experiments on three datasets, UCSD Ped1, UCSD Ped2 and Avenue, demonstrate the effectiveness of our method in terms of various evaluation criteria for video anomaly detection. Zhenjiang Miao, Wanru Xu, Jiaji Wang, Qiang Zhang 0030, Shaoyue Song |
ICMLA | 2 |
| 2021 | A CRNN-based attention-seq2seq model with fusion feature for automatic Labanotation generation1
Zhenjiang Miao, Wanru Xu |
Neurocomputing | 2 |
| 2021 | Deep Reinforcement Polishing Network for Video CaptioningabstractThe video captioning task aims to describe video content using several natural-language sentences. Although one-step encoder-decoder models have achieved promising progress, the generations always involve many errors, which are mainly caused by the large semantic gap between the visual domain and the language domain and by the difficulty in long-sequence generation. The underlying challenge of video captioning, i.e., sequence-to-sequence mapping across different domains, is still not well handled. Inspired by the proofreading procedure of human beings, the generated caption can be gradually polished to improve its quality. In this paper, we propose a deep reinforcement polishing network (DRPN) to refine the caption candidates, which consists of a word-denoising network (WDN) to revise word errors and a grammar-checking network (GCN) to revise grammar errors. On the one hand, the long-term reward in deep reinforcement learning benefits the long-sequence generation, which takes the global quality of caption sentences into account. On the other hand, the caption candidate can be considered a bridge between visual and language domains, where the semantic gap is gradually reduced with better candidates generated by repeated revisions. In experiments, we present adequate evaluations to show that the proposed DRPN achieves comparable and even better performance than the state-of-the-art methods. Furthermore, the DRPN is model-irrelevant and can be integrated into any video captioning models to refine their generated caption sentences. Wanru Xu, Jian Yu 0001, Zhenjiang Miao |
IEEE Trans. Multim. | 3 |
| 2020 | Multi-Spectral Salient Object Detection by Adversarial Domain AdaptationabstractAlthough there are many existing research works about the salient object detection (SOD) in RGB images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we first collect and will publicize a large multi-spectral dataset including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. We model this research problem as an adversarial domain adaptation from the existing RGB image dataset (source domain) to the collected multi-spectral dataset (target domain). Experimental results show the effectiveness and accuracy of the proposed adversarial domain adaptation for the multi-spectral SOD. Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Jianwu Fang, Cong Ma 0004, Song Wang 0002 |
AAAI | 3 |
| 2020 | Sequence-to-Sequence Labanotation Generation Based on Motion Capture DataabstractLabanotation is an important notation system for recording dances. Automatically generating Labanotation scores from motion capture data has attracted more interest in recent years. Current methods usually focus on individual movement segments and generate Labanotation symbols one by one. This requires segmenting the captured data sequence in advance. Manual segmentation will consume a lot of time and effort, while automatic segmentation may not be reliable enough. In this paper, we propose a sequence-to-sequence approach that can generate Labanotation scores from unsegmented motion data sequences. First, we extract effective features from motion capture data based on body skeleton analysis. Then, we train a neural network under the encoder-decoder architecture to transform the motion feature sequences to corresponding Labanotation symbols. As such, the dance score is generated. Experiments show that the proposed method performs favorably against state-of-the-art algorithms in the automatic Labanotation generation task. Min Li 0025, Zhenjiang Miao, Cong Ma 0004 |
ICASSP | 2 |
| 2020 | An Automatic Framework For Generating Labanotation Scores From Continuous Motion Capture DataabstractLabanotation is a widely used dance notation system. The topic of generating Labanotation scores from captured dance motion data has attracted more research interest in recent years. Current methods usually generate Laban symbols via recognizing motion segments that each contains a single dance movement. However, they rely on the manual segmentation of raw dance motion sequences, which can cost a lot of time and effort. In this paper, we present a fully automatic framework to generate Labanotation scores from continuous motion data. First, we split the captured dance data to motion segments based on the Laban theory of body weight support transferring. Then, based on the segmented data, we utilize a network with both 1D-convolutional and recurrent layers to recognize body movements and generate Laban symbols. As such, the Labanotation score is created. Extensive experiments show that the proposed automatic framework performs favorably against previous solutions. Min Li 0025, Zhenjiang Miao, Cong Ma 0004 |
ICME | 2 |
| 2020 | End-to-End Method For Labanotation Generation From Continuous Motion Capture DataabstractLabanotation is a standardized notation system for human motion recording and archiving. Existing methods for automatic Labanotation generation require pre-segmentation of continuous motion and only recognize single movement each time. The poor performance of pre-segmentation will subsequently lead to inaccurate movement recognition and Labanotation generation. In this paper, we proposed an end-to-end trainable method for recognizing continuous motion and generating corresponding Laban symbols without presegmentation. Firstly, we design Lie group feature to represent rotation information of joints and bones underlying in continuous motion capture data. Secondly, we employ Convolutional Recurrent Neural Network (CRNN) to jointly analyze motion sequences in spatial and temporal domain. Furthermore, we adopt Connectionist Temporal Classification (CTC) layer to translate per-frame predictions into Laban symbol sequence, which enables analysis on motion sequences with arbitrary lengths. Experiments on continuous motion capture dataset demonstrate the effectiveness of proposed method and its superiority compared with the state-of-the-art. Ningwei Xie, Zhenjiang Miao, Jiaji Wang, Qiang Zhang 0030 |
ICME | 2 |
| 2020 | Abnormal event detection in surveillance videos based on low-rank and compact coefficient dictionary learning
Zhenjiang Miao, Yi-Gang Cen, Xiao-Ping Zhang 0002, Linna Zhang, Shiming Chen 0001 |
Pattern Recognit. | 2 |
| 2020 | Spatio-Temporal Deep Q-Networks for Human Activity LocalizationabstractHuman activity localization aims to recognize category labels and detect the spatio-temporal locations of activities in video sequences. Existing activity localization methods suffer from three major limitations. First, the search space is too large for three-dimensional (3D) activity localization, which requires the generation of a large number of proposals. Second, contextual relations are often ignored in these target-centered methods. Third, locating each frame independently fails to capture the temporal dynamics of human activity. To address the above issues, we propose a unified spatio-temporal deep Q-network (ST-DQN), consisting of a temporal Q-network and a spatial Q-network, to learn an optimized search strategy. Specifically, the spatial Q-network is a novel two-branch sequence-to-sequence deep Q-network, called TBSS-DQN. The network makes a sequence of decisions to search the bounding box for each frame simultaneously and accounts for temporal dependencies between neighboring frames. Additionally, the TBSS-DQN incorporates both the target branch and context branch to exploit contextual relations. The experimental results on the UCF-Sports, UCF-101, ActivityNet, JHMDB, and sub-JHMDB datasets demonstrate that our ST-DQN achieves promising localization performance with a very small number of proposals. The results also demonstrate that exploiting contextual information and temporal dependencies contributes to accurate detection of the spatio-temporal boundary. Wanru Xu, Jian Yu 0001, Zhenjiang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Deep Reinforcement Learning for Weak Human Activity LocalizationabstractHuman activity localization aims at recognizing contents and detecting locations of activities in video sequences. With an increasing number of untrimmed video data, traditional activity localization methods always suffer from two major limitations. First, detailed annotations are needed in most existing methods, i.e., bounding-box annotations in every frame, which are both expensive and time consuming. Second, the search space is too large for 3D activity localization, which requires generating a large number of proposals. In this paper, we propose a unified deep Q-network with weak reward and weak loss (DWRLQN) to address the two problems. Certain weak knowledge and weak constraints involving the temporal dynamics of human activity are incorporated into a deep reinforcement learning framework under sparse spatial supervision, where we assume that only a portion of frames are annotated in each video sequence. Experiments on UCF-Sports, UCF-101 and sub-JHMDB demonstrate that our proposed model achieves promising performance by only utilizing a very small number of proposals. More importantly, our DWRLQN trained with partial annotations and weak information even outperforms fully supervised methods. Wanru Xu, Zhenjiang Miao, Jian Yu 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Multi-scale Feature and Spatial Relation Inference for Object Detection
Zhenjiang Miao, Jiaji Wang |
ICIG (1) | 2 |
| 2019 | Labanotation Generation Based on Bidirectional Gated Recurrent Units with Joint and Line FeaturesabstractLabanotation is an effective carrier for recording and displaying three-dimensional human movements. In the existing methods of Labanotation generation, the spatial characteristics are not fully considered. In addition, Long-term correlation of time series is not reflected. In this paper, we propose a novel method based on Bidirectional Gated Recurrent Units with Joint and Line features which can efficiently convert human movements into Labanotation. Firstly, Joint feature carries the location information; Line feature contains the direction information. These two types of features make full use of the correlation and co-occurrence between adjacent joints, thus embodying good spatial characteristics. Secondly, Bidirectional Gated Recurrent Units are applied to identify human movements. The unique gate-control structure of Bi-GRU can predict current status based on historical and future information. Thus, this method is good in timing modeling especially for long time series. The experimental results show that this method achieved an accuracy of 97.4%. It is higher than the state of the art, which demonstrating the effectiveness of our proposed method. Shanshan Hao, Zhenjiang Miao, Jiaji Wang, Wanru Xu, Qiang Zhang 0030 |
ICIP | 2 |
| 2019 | Prediction-CGAN: Human Action Prediction with Conditional Generative Adversarial NetworksabstractThe underlying challenge of human action prediction, i.e. maintaining prediction accuracy at very beginning of an action execution, is still not well handled. In this paper, we propose a Prediction Conditional Generative Adversarial Network (Prediction-CGAN) for predicting action, which shares information between completely observed and partially observed videos. Instead of generating future frames, we aim at completing visual representations of unfinished video, which can be directly utilized to predict action label no matter at any progress levels. The Prediction-CGAN incorporates the completion constraint to learn a transformation from incomplete actions to complete actions; the adversarial constraint to ensure the generation has similar discriminative power to complete representation; the label consistency constraint to encourage label consistency between each segment and its corresponding complete video; and the confidence monotonically increasing constraint to yield increasingly accurate predictions as observing more frames. Meanwhile, we introduce a novel adversarial criterion especially for prediction task, which requires the generation is more discriminative than its corresponding incomplete representation, while the generation is less discriminative than its real complete representation. In experiments, we present adequate evaluations to show that the proposed Prediction-CGAN outperforms state-of-the-art methods in action prediction. Wanru Xu, Jian Yu 0001, Zhenjiang Miao |
ACM Multimedia | 3 |
| 2019 | An easy-to-hard learning strategy for within-image co-saliency detection
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Dazhou Guo, Wei Ke 0001, Cong Ma 0004, Song Wang 0002 |
Neurocomputing | 3 |
| 2019 | Action recognition and localization with spatial and temporal contexts
Wanru Xu, Zhenjiang Miao, Jian Yu 0001 |
Neurocomputing | 2 |
| 2019 | Domain Adaptation for Convolutional Neural Networks-Based Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification plays an important role in the field of earth observation. With the rapid development of the RS techniques, a large number of RS scene images are available. As manually labeling large-scale RS scene images is both labor and time consuming, when a new unlabeled data set is obtained, how to use the existing labeled data sets to classify the new unlabeled images is an important research direction. Different RS scene image data sets may be taken from different type of sensors, and the images may vary from imaging modalities, spatial resolutions, and image scales, so the distribution discrepancy exists among different image data sets. As a result, simply applying convolutional neural networks (CNN) trained on source domain cannot accurately classify the images on target domain. Domain adaptation (DA) can be helpful to solve this problem. In this letter, we design a subspace alignment (SA) and CNN-based framework to solve the DA problem in RS scene image classification. A new SA layer is proposed and added into CNN models for DA, which could align the source and target domains in some feature subspace. Fine-tuning the modified CNN model with the added SA layer makes the CNN model adapt to the aligned feature subspace and helps to relieve the domain distribution discrepancy. The experiments conducted on two public data sets show that adding the SA layer into CNN improves the scene classification on the target domain. Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Qiang Zhang 0030, Yuewei Lin, Song Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Detecting Anomalous Trajectories via Recurrent Neural Networks
Cong Ma 0004, Zhenjiang Miao, Min Li 0025, Shaoyue Song, Ming-Hsuan Yang 0001 |
ACCV (4) | 2 |
| 2018 | A method of automatically generating Labanotation from human motion capture dataabstractThis paper presents a method of automatically generating Labanotation scores from human motion capture data. Up to now the main acquisition of Labanotation is manual recording by the professionals. Our work allows the users converting human motions to Labanotation scores efficiently. The key components of our method are the analysis of motion capture data, the segmentation of motion and the recognition of each motion fragment. In motion segmentation, we make the results aligned with the beat of Labanotation to ensure the generated symbols regular and accurate. In movement recognition, according to the different properties of human motion, we deal with the data in different suitable ways. Therefore, our recognition results are more reliable than previous works. The experiments show that our work is a useful tool for converting human dance motions into Labanotation scores. Further, considering its efficiency the method can be used to record large numbers of ethnic dances that coming to the crisis of being lost. Jiaji Wang, Zhenjiang Miao |
ICPR | 2 |
| 2017 | Learning a hierarchical spatio-temporal model for human activity recognitionabstractRecent works have shown that hierarchical models lead to significant improvement in human activity recognition, which can not only enhance descriptive capability, but also improve discriminative power. However, most existing methods exploit just one of the two advantages. In this paper, a new hierarchical spatio-temporal model (HSTM) is proposed to integrate feature learning into two-layer hierarchical classification model simultaneously. On the one hand, the two-layer model has sufficient descriptive capability. The bottom layer aims at capturing spatial relations in each frame and learning high-level representations, and the top layer utilizes these learned features to characterize temporal relations in the whole video sequence. On the other hand, the hierarchical model has strong discriminative power. Both spatial similarity and temporal similarity of activities are measured. Experimental results show that the HSTM can successfully recognize human activities with higher accuracies on one-person actions (KTH and UCF), human-human interactions (CASIA), and human-object interactional activities (Gupta). Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2017 | Optimized learning instance-based image retrieval
Yueli Li, Rongfang Bie, Chenyun Zhang, Zhenjiang Miao, Jiajing Wang, Hao Wu 0022 |
Multim. Tools Appl. | 4 |
| 2017 | Anomaly detection using sparse reconstruction in crowded scenes
Zhenjiang Miao, Yi-Gang Cen |
Multim. Tools Appl. | 2 |
| 2017 | A Saliency Prior Context Model for Real-Time Object TrackingabstractReal-time object tracking has wide applications in time-critical multimedia processing areas such as motion analysis and human-computer interaction. It remains a hard problem to balance between accuracy and speed. In this paper, we present a fast real-time context-based visual tracking algorithm with a new saliency prior context (SPC) model. Based on the probability formulation, the tracking problem is solved by sequentially maximizing the computed confidence map of target location in each video frame. To handle the various cases of feature distributions generated from different targets and their contexts, we exploit low-level features as well as fast spectral analysis for saliency to build a new prior context model. Then, based on this model and a spatial context model learned online, a confidence map is computed and the target location is estimated. In addition, under this framework, the tracking procedure can be accelerated by the fast Fourier transform. Therefore, the new method generally achieves a real-time running speed. Extensive experiments show that our tracking algorithm based on the proposed SPC model achieves real-time computation efficiency with overall best performance comparing with other state-of-the-art methods. Cong Ma 0004, Zhenjiang Miao, Xiao-Ping Zhang 0002, Min Li 0025 |
IEEE Trans. Multim. | 2 |
| 2017 | A Hierarchical Spatio-Temporal Model for Human Activity RecognitionabstractThere are two key issues in human activity recognition: spatial dependencies and temporal dependencies. Most recent methods focus on only one of them, and thus do not have sufficient descriptive power to recognize complex activity. In this paper, we propose a hierarchical spatio-temporal model (HSTM) to solve the problem by modeling spatial and temporal constraints simultaneously. The new HSTM is a two-layer hidden conditional random field (HCRF), where the bottom-layer HCRF aims at describing spatial relations in each frame and learning more discriminative representations, and the top-layer HCRF utilizes these high-level features to characterize temporal relations in the whole video sequence. The new HSTM takes advantage of the bottom layer as the building blocks for the top layer and it aggregates evidence from local to global level. A novel learning algorithm is derived to train all model parameters efficiently and its effectiveness is validated theoretically. Experimental results show that the HSTM can successfully classify human activities with higher accuracies on single-person actions (UCF) than other existing methods. More importantly, the HSTM also achieves superior performance on more practical interactions, including human-human interactional activities (UT-Interaction, BIT-Interaction, and CASIA) and human-object interactional activities (Gupta video dataset). Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2016 | Abnormal event detection based on sparse reconstruction in crowded scenesabstractIn this paper, we propose an algorithm of abnormal event detection in crowded scenes using sparse representation over the bases of normal motion feature descriptors. To construct an over-complete dictionary, we extract the histogram of maximal optical flow projection (HMOFP) feature from a set of normal training frames. Then the K-SVD dictionary training method is used to get a redundant dictionary after a process of selecting the training samples, which is better than the dictionary simply composed by the HMOFP feature of the whole training frames. In order to detect whether a frame is normal or not, we use the U-norm of the sparse reconstruction coefficients (i.e., the sparse reconstruction cost, SRC) to show the anomaly of the testing frame, which is simple but very effective. The experiment results on UMN dataset and the comparison to the state-of-the-art methods show that our algorithm is promising. Zhenjiang Miao, Yi-Gang Cen, Qinghua Liang |
ICASSP | 2 |
| 2016 | Saliency preprocessing for person re-identification imagesabstractIn this paper, we propose a preprocessing strategy for refining the pedestrian images in person re-identification(re-id). The accomplishment of the person re-id task depends on the features extracted from pedestrian appearances. The best image matches are verified based on these features as identification results. The saliency information in image scenes is often exploited for feature selection in high-level vision tasks. Inspired by the pre-attentive mechanism in human visual system, we utilize the saliency information in the re-identification data to refine the person appearance in a preprocessing step. We first obtain the eye-fixation-predicting map based on the saliency analysis of image. Then this map is used to spatially weight the image features for a better appearance. Finally, we apply these processed images to the feature extraction in a standard re-identification procedure. Experiments on the widely-used VIPeR dataset show that the proposed method improves the final performance of re-identification task. Cong Ma 0004, Zhenjiang Miao, Min Li 0025 |
ICASSP | 2 |
| 2016 | Subspace clustering with a learned dimensionality reduction projectionabstractSubspace clustering aims to separate data from a union of low dimensional linear subspaces. Many recent subspace clustering methods based on self-representation are popular and achieve state-of-art performance. Dimensionality reduction is a common preprocessing procedure before applying these clustering methods. In this paper, we present an algorithm to segment subspaces with a learned dimensionality reduction projection instead of simply using PCA (Principal Component Analysis). We propose an objective function which simultaneously learns the dimensionality reduction projection and self-representation coefficients. We integrate SMR (Smooth Representation subspace clustering) into this framework and propose SMR_LP (Smooth Representation clustering with Learned Projection). We also propose an efficient method to optimize the cost function. A well learned projection helps preserving the data structure and improves the clustering performance. Experimental results demonstrate the effectiveness of our proposed method. Qiang Zhang 0030, Zhenjiang Miao |
ICASSP | 2 |
| 2016 | Saliency prior context model for visual trackingabstractIn this paper, we present a new FFT-based visual tracking algorithm based on a new model for computing prior context distribution via a spectral saliency approach. The tracking problem is formulated under a Bayesian framework where the statistical relationships between the features of the target and its spatio-temporal context are modeled. When building the context model, the prior distribution of the possible target position is an important part worth studying. To deal with various cases of distributions based on different attributes of the target and its context, we exploit low level saliency features by spectral analysis to compute prior distribution, not limited to center-surround weights. We show by extensive experiments that the performance of the new tracking algorithm based on the new saliency prior context (SPC) model achieves real-time computation efficiency with overall best location accuracy performance compared with other state-of-the-art methods. Cong Ma 0004, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2016 | Global anomaly detection in crowded scenes based on optical flow saliencyabstractIn this paper, an algorithm of global anomaly detection in crowded scenes using the saliency in optical flow field is proposed. Before the process of extracting the histogram of maximal optical flow projection (HMOFP), the scale invariant feature transforms (SIFT) method is utilized to get the saliency map of optical flow field. On the basis of the HMOFP feature of normal frames, the online dictionary learning algorithm is used to train an optimal dictionary with proper redundancy after a process of selecting the training samples, which is better than the dictionary simply composed by the HMOFP feature of the whole training frames. In order to detect whether a frame is normal or not, we use the ℓ1-norm of the sparse reconstruction coefficients (i.e., the sparse reconstruction cost, SRC) to show the anomaly of the testing frame, which is simple but very effective. The experiment results on UMN dataset and the comparison to the state-of-the-art methods show that our algorithm is promising. Zhenjiang Miao, Yi-Gang Cen |
MMSP | 2 |
| 2016 | A novel mid-level distinctive feature learning for action recognition via diffusion map
Wanru Xu, Zhenjiang Miao |
Neurocomputing | 2 |
| 2016 | A new sampling algorithm for high-quality image matting
Hao Wu 0022, Yueli Li, Zhenjiang Miao, Runsheng Zhu, Rongfang Bie, Rui Lie |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Creative and high-quality image composition based on a new criterion
Hao Wu 0022, Yueli Li, Zhenjiang Miao, Runsheng Zhu, Rongfang Bie |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Recognition improvement through optimized spatial support methodology
Hao Wu 0022, Zhenjiang Miao, Jingyue Chen |
Multim. Tools Appl. | 2 |
| 2016 | Video restoration based on PatchMatch and reweighted low-rank matrix recovery
Bo-Hua Xu, Yi-Gang Cen, Ruizhen Zhao, Zhenjiang Miao |
Multim. Tools Appl. | 6 |
| 2015 | Research on Vehicle Type Classification Based on Spatial Pyramid Representation and BP Neural Network
Shaoyue Song, Zhenjiang Miao |
ICIG (3) | 2 |
| 2015 | Structured feature-graph model for human activity recognitionabstractRecent works have shown that extracting and learning mid-level features lead to significant improvement in human activity recognition. Most existing methods represent activities as a collection of mid-level features and their spatio-temporal relations are completely neglected. Therefore, when scene contains interactional parts or high-level semantic actions, these mid-level features are not able to capture spatial structures as well as high order temporal relationships. In this paper, the activity is represented as a string of structured feature-graphs (SFGs) which models spatial structures and temporal structures simultaneously. A novel temporal graph kernel (TGK) is also proposed to measure similarity between two string representations. Experimental results show that our approach can successfully classify human activities with much higher accuracies for both single-person actions and human-human interactions. Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2015 | Rotation invariant ellipsoid projection for domain transfer in human skin detectionabstractDifferent datasets are often regarded as different domains where different data distributions exist in many pattern recognition systems. A typical problem is the different illumination conditions that need to be adapted in human skin detection. In this paper we present a method called rotation invariant ellipsoid projection (RIEP) to handle the domain transfer problem in the feature space. It uses two ellipsoids to represent the data distributions of the source and the target domains. The source ellipsoid will be projected to the target ellipsoid to complete the domain transfer. In this process the source ellipsoid does not rotate to keep the updated data avoiding distortion. We implement this model in human skin detection and observe better performance comparing to some existing methods. Zhenjiang Miao, Qinghua Liang |
MMSP | 2 |
| 2015 | Recognition improvement through the optimisation of learning instancesabstractImage recognition based on machine learning has been widely utilised in the computer vision field. In the image recognition process, quite a few positive and negative instances are needed for effective machine learning. However, some invalid instances selected from the instance candidates, particularly for negative instances, will result in reduced image recognition accuracy and wasted resources. When making the instance selections for machine learning, if a large number of negative instances are selected that exhibit a semantic distance too close to the positive instances, the recognition accuracy will be considerably reduced. The selection of valid negative instances from the instance candidates has become a new challenge in the field of computer vision. In this paper a new method containing several different algorithms is presented, that more effectively selects valid negative instances. In this innovative process, a Wordnet subtree, semantic distance based on feature descriptors and an improved K‐nearest neighbour model were all utilised. Experiments were implemented using a large database to determine the effectiveness of this method. The results have demonstrated that this method can significantly improve the recognition results. Hao Wu 0022, Zhenjiang Miao, Jingyue Chen |
IET Comput. Vis. | 2 |
| 2015 | Defect inspection for TFT-LCD images based on the low-rank matrix reconstruction
Yi-Gang Cen, Ruizhen Zhao, Lihong Cui, Zhenjiang Miao |
Neurocomputing | 5 |
| 2015 | Image completion with multi-image based on entropy reduction
Hao Wu 0022, Zhenjiang Miao, Jingyue Chen, Cong Ma 0004 |
Neurocomputing | 2 |
| 2015 | Projection transform on spatio-temporal context for action recognition
Wanru Xu, Zhenjiang Miao, Qiang Zhang 0030 |
Multim. Tools Appl. | 2 |
| 2015 | Multi-label audio concept detection using correlated-aspect Gaussian Mixture Model
Cencen Zhong, Zhenjiang Miao |
Multim. Tools Appl. | 2 |
| 2015 | A Quasi-Dense Matching Approach and its Calibration Application with Internet PhotosabstractThis paper proposes a quasi-dense matching approach to the automatic acquisition of camera parameters, which is required for recovering 3-D information from 2-D images. An affine transformation-based optimization model and a new matching cost function are used to acquire quasi-dense correspondences with high accuracy in each pair of views. These correspondences can be effectively detected and tracked at the sub-pixel level in multiviews with our neighboring view selection strategy. A two-layer iteration algorithm is proposed to optimize 3-D quasi-dense points and camera parameters. In the inner layer, different optimization strategies based on local photometric consistency and a global objective function are employed to optimize the 3-D quasi-dense points and camera parameters, respectively. In the outer layer, quasi-dense correspondences are resampled to guide a new estimation and optimization process of the camera parameters. We demonstrate the effectiveness of our algorithm with several experiments. Yanli Wan, Zhenjiang Miao, Q. M. Jonathan Wu, Xifu Wang |
IEEE Trans. Cybern. | 2 |
| 2015 | Efficient Heuristic Methods for Multimodal Fusion and Concept Fusion in Video Concept DetectionabstractSemantic models are widely used to bridge the semantic gap between low-level features and high-level features in video concept indexing. Multimodal fusion and concept fusion are two commonly used approaches in building semantic models. In the previous work, domain adaptation is neglected in multimodal fusion, and many probability maximization based and unsupervised concept fusion methods are counterintuitive since they do not incorporate subjective human intuition. In this paper, we present a new two-stage semantic model combining the multimodal fusion and the concept fusion incorporating human heuristics. In the multimodal fusion model, we employ a new generic unsupervised method, namely, domain adaptive linear combination (DALC), to update the linear combination (LC) weights by incorporating the differences of element distributions between training and testing domains. In the concept fusion model, a novel mechanical node equilibrium (NE) model is developed by using forces to model the concept correlations to update the score of concepts represented by nodes. It is intuitive and can incorporate multiple kinds of correlations simultaneously to construct more sophisticated semantic structure. Compared to other state-of-the-art supervised and unsupervised methods, the new model can use either unsupervised or supervised factors to significantly improve the mean inferred average precision (MAP) performance on all datasets. Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2015 | Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR DescriptorabstractThe appearance gap between sketches and photo- realistic images is a fundamental challenge in sketch-based image retrieval (SBIR) systems. The existence of noisy edges on photo- realistic images is a key factor in the enlargement of the appearance gap and significantly degrades retrieval performance . To bridge the gap, we propose a framework consisting of a new line segment -based descriptor named histogram of line relationship (HLR) and a new noise impact reduction algorithm known as object boundary selection . HLR treats sketches and extracted edges of photo- realistic images as a series of piece-wise line segments and captures the relationship between them. Based on the HLR, the object boundary selection algorithm aims to reduce the impact of noisy edges by selecting the shaping edges that best correspond to the object boundaries. Multiple hypotheses are generated for descriptors by hypothetical edge selection. The selection algorithm is formulated to find the best combination of hypotheses to maximize the retrieval score; a fast method is also proposed. To reduce the distraction of false matches in the scoring process, two constraints on spatial and coherent aspects are introduced . We tested the HLR descriptor and the proposed framework on public datasets and a new image dataset of three million images, which we recently collected for SBIR evaluation purposes. We compared the proposed HLR with state-of-the-art descriptors (SHoG, GF-HOG). The experimental results show that our HLR descriptor outperforms them. Combined with the object boundary selection algorithm, our framework significantly improves SBIR performance. Jian Zhang 0002, Tony X. Han, Zhenjiang Miao |
IEEE Trans. Multim. | 4 |
| 2015 | Optimized recognition with few instances based on semantic distance
Hao Wu 0022, Zhenjiang Miao, Manna Lin |
Vis. Comput. | 2 |
| 2014 | Image matching using adapted image models and its application to content-based image retrievalabstractThe image matching is an important and necessary step in a series of processes aimed at overall content-based image retrieval (CBIR). This paper presents a new image matching approach for CBIR. It first defines the concept of image-class and the process of image matching in CBIR is formulated into a set of hypothesis tests. The retrieval is based on modeling positive and negative hypotheses and testing a query image against these two hypotheses. In order to model the two hypotheses, the paper proposes to calculate first a universal image model (UIM). The derived UIM is then used as a reference for the calculation of adapted models for each image in the database. The use of Bayesian adaptation to derive well-behaved image models from an UIM in image retrieval is a main contribution of the paper. In addition, the paper discusses an acceleration technique based on ranking the closest mixture Gaussian components of the background model and using their corresponding components in the positive classes. The experimental results on publicly available datasets show that the proposed approach improves the robust and evident performance. Zhenjiang Miao |
ICIP | 2 |
| 2014 | Modeling correlation between multi-modal continuous words for pLSA-based video classificationabstractSeeing that probabilistic Latent Semantic Analysis (pLSA) deals with discrete quantity only, pLSA with Gaussian Mixtures (GM-pLSA) extends it to continuous feature space by treating continuous feature as continuous word. However, GM-pLSA does not provide a clear way of modeling multimodal features, and also neglects the intrinsic correlation between these continuous words. In this paper, we present a graph regularized multi-modal GM-pLSA (GRMMGM-pLSA) model to incorporate such correlation between multimodal continuous words into the process of model learning. First, multiple GMMs are adopted with each depicting the distribution of continuous words from each modality; and then, a graph regularizer is introduced to capture the word correlation. In the task of video classification, GRMMGM-pLSA that takes both multi-modal visual features of sub-shots and word correlation in terms of temporal consistency between sub-shots into account is exploited to perform feature mapping. Experiments on YouTube videos show the effectiveness of our proposed model. Cencen Zhong, Zhenjiang Miao |
ICIP | 2 |
| 2014 | A new approach of conditions on δ 2s (Φ) for s-sparse recovery
Yi-Gang Cen, Ruizhen Zhao, Zhenjiang Miao, Lihong Cui |
Sci. China Inf. Sci. | 3 |
| 2014 | Graph regularized GM-pLSA and its applications to video content analysis
Cencen Zhong, Zhenjiang Miao |
Multim. Syst. | 2 |
| 2014 | Continuous human action recognition in real time
Zhenjiang Miao, Yuan Shen 0004, Wanru Xu, Dianyong Zhang |
Multim. Tools Appl. | 2 |
| 2014 | Multihuman Tracking Based on a Spatial-Temporal Appearance MatchabstractIn this paper, we focus on the improvements of appearance representation for multihuman tracking. Many previous methods extracted low-level appearance features, such as color histogram and texture, even combined with spatial information for each frame. These methods ignore the temporal distribution of features. The features of each frame may not be stable due to illumination, human pose variation, and image noise. In order to improve it, we propose a novel appearance representation called the spatial-temporal appearance model based on the statistical distribution of Gaussian mixture model (GMM). It represents the appearance of a tracklet as a whole with dynamic spatial and temporal information. The spatial information is the dynamic subregions. The temporal information is the dynamic duration time of each subregion. Each subregion is modeled as the weighted Gaussian distribution of GMM. The online expectation-maximization (online EM) algorithm is used to estimate the parameters of GMM. Then, we propose a tracklet association method using Bayesian prediction and Jensen-Shannon divergence. The Bayesian prediction is used to predict the locations of targets. The Jensen-Shannon divergence is used to compute the distance of spatial-temporal appearance distribution between two tracklets. Finally, we test our approach on four challenging datasets (TRECVID, CAVIAR, ETH, and EPFL Terrace) and achieve good results. Yuan Shen 0004, Zhenjiang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Illumination Robust Video Foreground Prediction Based on Color RecoveringabstractVideo foreground prediction is a technique to estimate the probability of each pixel being foreground in current frame based on a foreground segmentation result of its previous frame. Existing foreground prediction algorithms usually assume that the illumination conditions are constant for consecutive frames. Therefore, they cannot predict foreground accurately when the illumination condition changes sharply between video frames. In this paper, a new robust video foreground prediction algorithm is proposed based on color recovering, which is derived based on an observation that the illumination changes are locally smooth. By integrating color recovering with an optical flow estimation algorithm and an opacity propagation algorithm, the negative impact of the illumination changes could be removed. Experimental results show that the proposed algorithm can get more accurate results for videos with illumination changes compared with the existing foreground prediction algorithms. Yanli Wan, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2014 | Low-resolution face recognition: a review
Zhenjiang Miao, Q. M. Jonathan Wu, Yanli Wan |
Vis. Comput. | 2 |
| 2013 | Projection-optimal tensor local fisher discriminant analysis for image feature extractionabstractTensor-based feature extraction approaches have been proved to be effective since they can solve the undersampled problem. In this paper, we propose a novel method called projection-optimal tensor local fisher discriminant analysis (PoTLFDA), which shares the character of local fisher discriminant analysis (LFDA). A novel affinity matrix is defined to effectively reflect the relationships of points in original tensor space and embedding space. The projection matrices are optimized by alternately solving the trace ratio problem. Convergence proof of the proposed algorithm is also given in this paper. Experiment results on face databases demonstrate the effectiveness of PoTLFDA. Qiuqi Ruan, Zhenjiang Miao |
ICIP | 3 |
| 2013 | A new edge feature for head-shoulder detectionabstractIn this work, we introduce a new edge feature to improve the head-shoulder detection performance. Since Head-shoulder detection is much vulnerable to vague contour, our new edge feature is designed to extract and enhance the head-shoulder contour and suppress the other contours. The basic idea is that head-shoulder contour can be predicted by filtering edge image with edge patterns, which are generated from edge fragments through a learning process. This edge feature can significantly enhance the object contour such as human head and shoulder known as En-Contour. To evaluate the performance of the new En-Contour, we combine it with HOG+LBP [1] as HOG+LBP+En-Contour. The HOG+LBP is the state-of-the-art feature in pedestrian detection. Because the human head-shoulder detection is a special case of pedestrian detection, we also use it as our baseline. Our experiments have indicated that this new feature significantly improve the HOG+LBP. Jian Zhang 0002, Zhenjiang Miao |
ICIP | 3 |
| 2013 | A UIM/ICM based approach to content-based image retrievalabstractThis paper presents a new similarity measure and matching scheme for content-based image retrieval (CBIR), based on modeling positive and negative hypotheses and testing a query image against these two hypotheses. The paper proposes to calculate first a universal image model (UIM), which is built based on a large set of images. The derived UIM is then used as a reference for the calculation of adapted models for each image class, which is done by a Bayesian adaptation of the GMM. The image class models (ICM) are therefore based on adapted versions of the background mixture components. Querying is based on the likelihood ratio between the values of these two hypotheses. A parameter adaptation technique is also introduced based on the background hypothesis. In addition, the paper discussed an acceleration technique based on ranking the closest Gaussian components of the background model and using their corresponding components in the positive classes. The experimental results show that the proposed approach improves the robust and evident performance. Zhenjiang Miao |
ICME | 2 |
| 2012 | Temporal-Spatial Refinements for Video Concept Fusion
Zhenjiang Miao, Hai Chi |
ACCV (3) | 2 |
| 2012 | Data-specific concept correlation estimation for video annotation refinementabstractFor video annotation refinement, a reasonable concept correlation representation is crucial. In this paper, we present a data-specific concept correlation estimation procedure for this task, where the resulting correlation with respect to each data encodes both its visual and high-level characteristics. Specifically, this procedure comprises two major modules: concept correlation basis estimation and data-specific concept correlation calculation. Under the framework of sparse representation, the former introduces a set of high-level concept correlation bases to represent the concept distribution of each feature-level basis, while the latter constructs the concept correlation of a specific data by combining its feature-level sparse coefficients and correlation bases together. In the end, given this new correlation, a probability-calculation based video annotation refinement is performed on TRECVID 2006 dataset. The experiments show that such a representation capturing data-specific characteristics could achieve better performance, than the generic concept correlation applied to all data. Cencen Zhong, Zhenjiang Miao |
ICASSP | 2 |
| 2012 | Human action categories using motion descriptorsabstractIn this paper, we recognize human action based on an improved BOW model and latent topic model. We proposed an improved motion descriptor to build our bag of words, which is called the local spatial-temporal maximum value of optical flow. We force similar local features that appear in different positions on the image grid to be assigned to different visual words. This approach assigns the spatial information to each visual word. Then, we use the topic model of pLSA (probabilistic Latent Semantic Analysis) to classify. Our approach is tested on two datasets, the KTH datasets and WEIZMANN datasets. The result shows our method is effective. Zhenjiang Miao |
ICIP | 2 |
| 2012 | Similarity weighted sparse representation for classification
Qiuqi Ruan, Zhenjiang Miao |
ICPR | 3 |
| 2012 | Object segmentation in multiple views without camera calibration
Qinghua Liang, Zhenjiang Miao |
ICPR | 2 |
| 2012 | Unsupervised online learning trajectory analysis based on weighted directed graph
Yuan Shen 0004, Zhenjiang Miao |
ICPR | 2 |
| 2012 | Image Composition with Color HarmonizationabstractImage composition is a very important technique in computer generated imagery. Besides some factors such as contrast, texture and noise that affect the quality of the composition, color harmony between fore- and background is also an important factor that would affect the quality of the composition. However, in the previous image composition techniques, color harmony between fore- and background is seldom considered. In this paper, an optimization method is proposed to deal with the color harmonization problem that used in image composition. A cost function is derived from the local smoothness of the hue values, and the image is harmonized by minimizing the cost function. A new matching cost function is proposed to select the best matching harmonic schemes. Our approach overcomes several shortcomings of the existing color harmonization methods. We validate the performance of our method and demonstrate its effectiveness with a variety of experiments. Yanli Wan, Zhenjiang Miao |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | A Two-View Concept Correlation Based Video Annotation RefinementabstractRecently, concept correlation defining the relationship between concepts has been playing an important role in video annotation (or concept detection). To improve the annotation performance, this paper presents a two-view concept correlation based video annotation refinement, using data-specific spatial and temporal concept correlations. Specifically, instead of generic concept correlation within shots, the spatial view estimates a data-specific concept correlation for each shot, via introducing concept correlation bases to map low-level features to high-level concept distribution under the framework of sparse representation. On the other hand, beyond the temporal consistency of one concept, a richer temporal correlation between different concepts respectively locating in the current shot and its neighbors is utilized to adjust the detection scores. In the end, these two types of concept correlations are integrated into a probability calculation based framework to refine the initial results derived from multiple concept detectors. And the experiments conducted on TRECVID 2006-2008 datasets and comparison with existing works demonstrate its effectiveness. Cencen Zhong, Zhenjiang Miao |
IEEE Signal Process. Lett. | 2 |
| 2012 | Coupled Observation Decomposed Hidden Markov Model for Multiperson Activity RecognitionabstractMultiperson activity recognition in videos is a challenging task, due to the complexity of interactions among multiple persons. In this paper, a new statistical model, named coupled observation decomposed hidden Markov model (CODHMM), is presented to model multiperson activities in videos. A human activity that involves multiple persons is analyzed in two levels: the individual level that describes each individual's motion details and the interaction level that expresses the shared information among multiple persons. The two levels are modeled by two hidden Markov chains that are interdependent and interact with each other. The observation in each chain at each time slice is decomposed into subobservations according to the number of features and the number of persons. For each activity to be recognized, a CODHMM is built and model parameters are learnt by a generalized expectation maximization (EM) algorithm. Given an input video that contains an unknown activity, maximum likelihood algorithms are developed to classify it into one of the learnt activity categories. Experimental results show that the CODHMM can successfully classify human activities involving multiple persons with high accuracy and low computations. Zhenjiang Miao, Xiao-Ping Zhang 0002, Yuan Shen 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Video matting via opacity propagation
Zhenjiang Miao, Yanli Wan, Dianyong Zhang |
Vis. Comput. | 2 |
| 2011 | A cost function approach for multi-human trackingabstractIn this paper, we propose a cost function approach for multi-human tracking. We first build short reliable trajectories based on human appearance and position. Here, we apply topic model to represent human appearance. The appearance of each person can be considered as topic distribution. Then, we compute the cost function for each pair of reliable trajectories in the time sliding window to estimate whether two trajectories can be associated. The cost function includes four parts, appearance cost, motion direction cost, object size cost and distance cost. The experimental results show that our approach has better performance compared with state-of-the-art approaches. Yuan Shen 0004, Zhenjiang Miao |
ICIP | 2 |
| 2011 | Skin detection with illumination adaptation in single imageabstractThis paper presents a novel method of skin detection in single image with illumination adaption. In this method, both the input image and the skin model will be updated dynamically using skin color region projection and skin model fusion respectively. These two steps are based on the illumination model which is derived from the clustering of the training material. Histogram model is used as the basic detection method. The experimental result shows the proposed method improved the performance obviously. Zhenjiang Miao, Cencen Zhong |
ICME | 2 |
| 2011 | Foreground prediction for bilayer segmentation of videos
Zhenjiang Miao, Yanli Wan, Forrest Fabian Jesse |
Pattern Recognit. Lett. | 2 |
| 2010 | Real Time Human Action Recognition in a Long Video SequenceabstractIn recent years, most action recognition researches focus on isolated action analysis for short videos, but ignore the issue of continuous action recognition for a long video sequence in real time. This paper proposes a novel approach for human action recognition in a video sequence with whatever length, which, unlike previous works,requires no annotations and no pre-temporal-segmentations.Based on the bag of words representation and the probabilistic Latent Semantic Analysis (pLSA) model, there cognition process goes frame by frame and the decision updates from time to time. Experimental results show that this approach is effective to recognize both isolated actions and continuous actions no matter how long a video sequence is. This is very useful for real time applications like video surveillance. Besides, we also test our approach for real time temporal video segmentation and real time keyframe extraction. Zhenjiang Miao, Yuan Shen 0004, Heng-Da Cheng |
AVSS | 2 |
| 2010 | Reconstruction of dense point cloud from uncalibrated widebaseline imagesabstractThis paper presents a new approach to reconstruct 3D dense point cloud from uncalibrated wide-baseline images. It includes three steps: acquiring quasi-dense point correspondences, recovering structure from motion, and reconstructing the 3D dense point cloud. We present a two level propagation algorithm. The first level is implemented in the 2D image space and the second level is in the 3D scene space. We use affine iterative model to acquire accurate quasi-dense correspondences, and iterative optimization to recover the camera parameters and the 3D scene structure. It increases the robust and accuracy of self-calibration. In the second level propagation, the strategy of view selection and local photometric consistency are used to minimize the effects of occlusions and highlights etc. We demonstrate our algorithm with some high-quality reconstruction examples. Yanli Wan, Zhenjiang Miao |
ICASSP | 2 |
| 2010 | Masks based human action detection in crowded videosabstractThis paper discusses the task of human action detection in crowded videos. First, we propose a novel mask based shape matching method for action recognition. Our method does not need human detection or segmentation, and it can be used in both clean and crowed backgrounds. Next, shape and flow based features are combined due to their complementary nature. For each action, a binary sequence is used as the template for both shape and flow matching. For a testing sequence and a template sequence, dynamic time warping technique is first applied for time alignment, then shape and flow matching distances are computed between matched frames. We test our algorithm on the CMU dataset and achieve an encouraging performance. Zhenjiang Miao, Heng-Da Cheng |
ICIP | 2 |
| 2010 | Unsynchronized markerless motion capture with sharp illumination changesabstractHuman motion capture typically requires synchronization hardware and specialized cameras, which makes the system expensive and technically complex. Instead, in this paper, we propose a novel approach for markerless motion capture with multiple unsynchronized cameras. Cameras synchronization is achieved using sharp illumination changes produced while recording the videos. To estimate poses with the presence of sharp illumination changes, we also propose a pose estimation method based on a bilayer Gaussian Mixture Model (GMM) and graph-cut, which can handle sharp illumination changes. And we demonstrate our video synchronization and markerless motion capture methods by the experiments on real human motion videos. Zhenjiang Miao, Heng-Da Cheng, Dianyong Zhang |
ICIP | 2 |
| 2010 | Automatic foreground extraction for images and videosabstractIn this paper, an automatic foreground extraction algorithm for images and videos is presented. It first automatically locates the foreground object coarsely by a salience detection algorithm, and then refines the object by Weighted Kernel Density Estimation (WKDE) and graph cut algorithm. A new initial probability map construction algorithm for WKDE is also presented. We then extend our algorithm to videos by an opacity tracking algorithm. It can automatically extract the foreground with high accuracy even for the videos with non-static background. The experiments demonstrate the effectiveness of our approach. Zhenjiang Miao, Yanli Wan |
ICIP | 2 |
| 2010 | A robust and accurate self-calibration approach from unordered wide-baseline imagesabstractThis paper presents a robust and accurate self-calibration approach from unordered wide-baseline images. An optimization model based on affine transformation is introduced into propagation to acquire higher accuracy of quasi-dense correspondences. Our self-calibration algorithm is completed through two-layer iteration. In the inner layer, global objective function and local photometric consistency is used to iteratively optimize camera parameters and scene structure. In the outer layer, a resampling strategy is presented to iteratively select a group of quasi-dense correspondences to solve the camera parameters and scene structure. We demonstrate our algorithm with several high accurate calibration results. Yanli Wan, Zhenjiang Miao, Mingxing Hu |
ICIP | 2 |
| 2010 | Tone Recognition of Isolated Mandarin Syllables
Zhaoqiang Xie, Zhenjiang Miao |
ICISP | 2 |
| 2010 | Multi-person activity recognition through hierarchical and observation decomposed HMMabstractMulti-person activity recognition is a challenging task due to the complex interactions between people and the multi-dimensionality of features. This paper proposes a hierarchical and observation decomposed hidden Markov model to classify multi-person activities. In order to give detailed descriptions of people's interactions by different feature scale, states of individual persons and states of interactions between people are separated. In addition, observations are decomposed into groups of subobservations to handle high dimensionality problem of feature space. This decomposition enhances the flexibility in feature selection which enables the combination of discrete and continuous features. Besides, this model has no limitations in terms of the number of persons. Experiments are successfully conducted with encouraging results. Activities of two persons and three persons are classified with good accuracies. Zhenjiang Miao |
ICME | 2 |
| 2010 | Temporally consistent video matting based on bilayer segmentationabstractIn this paper, a quasi-automatic video matting approach which can preserve the temporal consistency of the alpha mattes is presented. “Quasi-automatic” means that it only needs a few user interactions on the first frame. A new algorithm which incorporates the Bayesian Estimation, Weighted Kernel Density Estimation (WKDE) and graph cut is presented to automatically and accurately segment each frame into bilayer. And then a GMM (Gaussian Mixture Model) based algorithm is used to transform the foreground layer to the trimap. At last, a 3D closed-form method is used to achieve the temporally consistent video matting. In order to give the user real time feedbacks when the user adds more scribbles to edit the matting result, a smallest-independent-region based re-matting algorithm is proposed. We test our algorithm with variety of video sequences, and the results demonstrate the effectiveness of our approach. Zhenjiang Miao, Yanli Wan |
ICME | 2 |
| 2010 | Action Detection in Crowded Videos Using MasksabstractIn this paper, we investigate the task of human action detection in crowded videos. Different from action analysis in clean scenes, action detection in crowded environments is difficult due to the cluttered backgrounds, high densities of people and partial occlusions. This paper proposes a method for action detection based on masks. No human segmentation or tracking technique is required. To cope with the cluttered and crowded backgrounds, shape and motion templates are built and the shape templates are used as masks for feature refining. In order to handle the partial occlusion problem, only the moving body parts in each motion are involved in action training. Experiments using our approach are conducted on the CMU dataset with encouraging results. Zhenjiang Miao |
ICPR | 2 |
| 2009 | Design and Implementation of a Vision-Based Motion Capture SystemabstractMotion Capture has been widely used in many fields such as animation and game production. This paper describes the scheme and implementation of a self-designed vision-based motion capture system. This system is set up based on a Client/Server network structure. The client software tracks every marker attached to the actor's body and sends the results to the server in real-time. The server software recovers every joint's 3D position and shows the capturing results simultaneously. It can also convert the 3D position data into BVH file data. Zhenjiang Miao |
ICIG | 2 |
| 2009 | Markerless human motion capture by Markov random field and dynamic graph cuts with color constraints
Chengkai Wan, Dianyong Zhang, Zhenjiang Miao, Baozong Yuan |
Sci. China Ser. F Inf. Sci. | 4 |
| 2009 | A moving object segmentation algorithm for static camera via active contours and GMM
Chengkai Wan, Baozong Yuan, Zhenjiang Miao |
Sci. China Ser. F Inf. Sci. | 3 |
| 2008 | Moving object detection in dynamic scenes using nonparametric local kernel histogram estimationabstractRobust detection of moving objects in complex and dynamic scenes is one of the most challenging issues in computer vision. In this paper, we present an approach to segmenting moving objects with nonparametric estimated local kernel histogram (ELKH) in dynamic scenes. By using the correlation and texture of spatially proximal pixels, local kernel histogram background model is constructed. Then probability distribution of local kernel histogram is estimated with nonparametric techniques. We employ Bhattacharyya distance to measure the similarity of local kernel histogram between estimated background model with each pixel in current frame, and decide whether the pixel belongs to moving objects or not. Our approach can reduce false detections due to disturbing noise and small motions in dynamic scenes, such as swaying branches, flickering water surface, rain. Experiments show that the proposed approach achieves promising results in real dynamic videos robustly. Baozong Yuan, Zhenjiang Miao |
ICME | 3 |
| 2008 | Automatic panorama image mosaic and ghost eliminatingabstractGhost is an important problem in the process of image mosaic. It results from the inaccurate registration or the moving objects in the image. In order to get rid of the ghost coming from inaccurate registration effectively, phase correlation technique, median flow filter and improved RANSAC are used to improve the accuracy of local registration. Bundle adjustment is used to correct the accumulated errors caused by the concatenation of pairwise homographies and the disregarded multiple constraints between images. In this paper, we put forward a new blending method which searches the blending region dynamically to eliminate the ghost caused by moving objects. Experiments demonstrate our method effective. Yanli Wan, Zhenjiang Miao |
ICME | 2 |
| 2008 | Feature-based super-resolution for face recognitionabstractIn video surveillance, face images that are captured by common cameras usually have a low resolution, which is a great obstacle to face recognition. In this paper, we show that classification accuracy drops very quickly and algebraic features of face image change greatly as resolution decreases, and we further indicate that two orthogonal matrices (TOMs) of SVD contain the leading information for recognition. Based on the phenomena, we propose a feature-based face super-resolution method using eigentransformation by 2DPCA. TOMs of SVD on face image matrices are reconstructed instead of super-resolution being performed in pixel domain. Reconstructed algebraic features are formed into high-resolution face images and also be used for face recognition directly. Experiments demonstrate that our method obtains hallucinated face images with good vision, and improves performance of low-resolution face recognition. Zhenjiang Miao |
ICME | 2 |
| 2008 | A new algorithm for static camera foreground segmentation via active coutours and GMMabstractForeground segmentation is one of the most challenging problems in computer vision. In this paper, we propose a new algorithm for static camera foreground segmentation. It combines Gaussian mixture model (GMM) and active contours method, and produces much better results than conventional background subtraction methods. It formulates foreground segmentation as an energy minimization problem and minimizes the energy function using curve evolution method. Because of the integration of GMM background model, shadow elimination term and curve evolution edge stopping term into energy function, it achieves more accurate segmentation than existing method of the same type. Promising results on real images demonstrate the potential of the presented method. Chengkai Wan, Baozong Yuan, Zhenjiang Miao |
ICPR | 3 |
| 2008 | Scale invariant face recognition using probabilistic similarity measureabstractIn video surveillance, the size of face images is very small. However, few works have been done to investigate scale invariant face recognition. Our experiments on appearance-based methods in different resolutions show that such methods as neighboring preserving embedding (NPE) preserving local structure are less effective than global ones such as linear discriminant analysis (LDA) under low-resolution. Based on the phenomena, we present a new graph embedding method FisherNPE, preserving both global and local structures on the data, and using Bayesian probabilistic similarity analysis of intensity differences between high- and low-resolution images for scale robust feature extraction. Experimental results on ORL and Yale database indicate that our method obtains good results on different resolution images. Zhenjiang Miao |
ICPR | 2 |
| 2008 | Fuzzy discriminant projections for facial expression recognitionabstractA linear projective map called fuzzy discriminant projections has been proposed in this paper. Fuzzy discriminant projection (FDP) is motivated by locality preserving projections which can optimally preserve the neighborhood structure of the data set. FDP utilizes the soft assignment method to weight pairs of samples with membership degree, and tries to find the optimal projective directions by maximizing the ratio of between-class distance against within-class distance. The resulting embedding subspace has more discriminant and robust power than that of traditional methods. Experiments on Cohn-Kanade databases show that FDP can effectively distinct the confusing facial expressions and obtain higher recognition accuracies than other subspacebased methods. Ruicong Zhi, Qiuqi Ruan, Zhenjiang Miao |
ICPR | 3 |
| 2008 | Markerless human body motion capture using Markov random field and dynamic graph cuts
Chengkai Wan, Baozong Yuan, Zhenjiang Miao |
Vis. Comput. | 3 |
| 2007 | The application of information fusion in the real-time monitoring systemabstractThis paper studies the feasibility of information analysis processing technology, which fuses speech and image together in the real-time monitoring system. It emphasizes particularly on speech analysis and fuses these two technologies in terms of scoring strategy. It also makes some improvement on MFCC feature extraction and proposes a quick MFCC algorithm. The proposed algorithm can reach the requirement of real-time system in case of the high precision. To prove it, this paper compares its algorithm with LPC and FFT. The experiment indicates that the EER of LPC is 13.9% and the EER of FFT is 11.1%, but by using the Quick MFCC the EER is only 4.2%. And compared with the traditional MFCC algorithm, the quick MFCC algorithm reduces the run time greatly while maintaining recognition accuracy of the system. Finally the rate of fusion recognition is about 97.8%, which is a good result for the real-time monitoring system. Zhenjiang Miao |
FUSION | 2 |
| 2007 | Model-Based Markerless Human Body Motion Capture using Multiple CamerasabstractCurrent markerless model-based human body motion capture methods always aim at accurate human body model and reconstruction surface contour. Unfortunately, because of the factors such as loose clothing, image noise and background segmentation errors, the efforts of these methods get very limited effects. In this paper, we propose a new algorithm for markerless model-based human body motion capture which is robust to the inaccurate human body reconstruction caused by the factors mentioned above. We extracted a volume data (voxel) representation from silhouettes in multiple video images. In the consideration of the human body model, we construct an articulated model with a potential energy which emphasize the skeleton of the human body and is not sensitive to the outline details of the human body surface. Then, we fit the human body model to the volume data in an expectation-maximization framework and recover the pose of the human body. Chengkai Wan, Baozong Yuan, Zhenjiang Miao |
ICME | 3 |
| 2007 | Background Subtraction Using Running Gaussian Average and Frame Difference
Zhenjiang Miao, Yanli Wan |
ICEC | 2 |
| 2007 | An Algorithm for Seamless Image Stitching and Its Application
Zhenjiang Miao |
ICEC | 2 |
| 2007 | Bootstrap FDA for counting positives accurately in imprecise environments
Jigang Xie, Zhengding Qiu, Zhenjiang Miao, Yanqiang Zhang |
Pattern Recognit. | 3 |
| 2006 | An OOPR-based rose variety recognition system
Zhenjiang Miao, M.-H. Gandelin, Baozong Yuan |
Eng. Appl. Artif. Intell. | 1 |
| 2006 | A new image shape analysis approach and its application to flower shape analysis
Zhenjiang Miao, M.-H. Gandelin, Baozong Yuan |
Image Vis. Comput. | 1 |
| 2000 | Zernike moment-based image shape analysis and its application
Zhenjiang Miao |
Pattern Recognit. Lett. | 1 |
| 1999 | Analysis and optimal design of continuous neural networks with applications to associative memory
Zhenjiang Miao, Baozong Yuan |
Neural Networks | 1 |