Wei Zeng 0006

dblp:80/1961-6 · DBLP profile ↗
← Back
35ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0002-2870-1178ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 5 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Towards to real world vehicle privacy protection: A new dataset and benchmark
Jiayi Lin 0010, Chengming Zou, Long Lan, Yong Luo 0002, Yue Yu 0001, Yaowei Wang 0001, Wei Zeng 0006, Yonghong Tian 0001
Pattern Recognit.7
2025 NN-Former: Rethinking Graph Structure in Neural Architecture Representation
abstract
The growing use of deep learning necessitates efficient network design and deployment, making neural predictors vital for estimating attributes such as accuracy and latency. Recently, Graph Neural Networks (GNNs) and transformers have shown promising performance in representing neural architectures. However, each of both methods has its disadvantages. GNNs lack the capabilities to represent complicated features, while transformers face poor generalization when the depth of architecture grows. To mitigate the above issues, we rethink neural architecture topology and show that sibling nodes are pivotal while overlooked in previous research. We thus propose a novel predictor leveraging the strengths of GNNs and transformers to learn the enhanced topology. We introduce a novel token mixer that considers siblings, and a new channel mixer named bidirectional graph isomorphism feed-forward network. Our approach consistently achieves promising performance in both accuracy and latency prediction, providing valuable insights for learning Directed Acyclic Graph (DAG) topology. The code is available at https://github.com/XuRuihan/NNFormer.
Ruihan Xu 0002, Haokui Zhang, Yaowei Wang 0001, Wei Zeng 0006, Shiliang Zhang
CVPR4
2025 Multi-Modal Reference Learning for Fine-Grained Text-to-Image Retrieval
abstract
Fine-grained text-to-image retrieval aims to retrieve a fine-grained target image with a given text query. Existing methods typically assume that each training image is accurately depicted by its textual descriptions. However, textual descriptions can be ambiguous and fail to depict discriminative visual details in images, leading to inaccurate representation learning. To alleviate the effects of text ambiguity, we propose a Multi-Modal Reference learning framework to learn robust representations. We first propose a multi-modal reference construction module to aggregate all visual and textual details of the same object into a comprehensive multi-modal reference. The multi-modal reference hence facilitates the subsequent representation learning and retrieval similarity computation. Specifically, a reference-guided representation learning module is proposed to use multi-modal references to learn more accurate visual and textual representations. Additionally, we introduce a reference-based refinement method that employs the object references to compute a reference-based similarity that refines the initial retrieval results. Extensive experiments are conducted on five fine-grained text-to-image retrieval datasets for different text-to-image retrieval tasks. The proposed method has achieved superior performance over state-of-the-art methods. For instance, on the text-to-person image retrieval dataset RSTPReid, our method achieves the Rank1 accuracy of 56.2%, surpassing the recent CFine by 5.6%.
Zehong Ma, Hao Chen 0061, Wei Zeng 0006, Limin Su, Shiliang Zhang
IEEE Trans. Multim.3
2024 TPTE: Text-Guided Patch Token Exploitation for Unsupervised Fine-Grained Representation Learning
abstract
Recent advances in pre-trained vision-language models have successfully boosted the performance of unsupervised image representation in many vision tasks. Most of existing works focus on learning global visual features with Transformers and neglect detailed local cues, leading to suboptimal performance in fine-grained vision tasks. In this article, we propose a text-guided patch token exploitation framework to enhance the discriminative power of unsupervised representation by exploiting more detailed local features. Our text-guided decoder extracts local features with the guidance of texts or learned prompts describing discriminative object parts. We hence introduce a local-global relation distillation loss to promote the joint optimization of local and global features. The proposed method allows to flexibly extract either global or global-local features as the image representation. It significantly outperforms previous methods in fine-grained image retrieval and base-to-new fine-grained classification tasks. For instance, our Recall@1 metric surpasses the recent unsupervised retrieval method STML by 6.0% on the SOP dataset. The code is publicly available at https://github.com/maosnhehe/TPTE .
Shunan Mao, Hao Chen 0061, Yaowei Wang 0001, Wei Zeng 0006, Shiliang Zhang
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Neural Architecture Search with Representation Mutual Information
abstract
Performance evaluation strategy is one of the most important factors that determine the effectiveness and efficiency in Neural Architecture Search (NAS). Existing strategies, such as employing standard training or performance predictor, often suffer from high computational complexity and low generality. To address this issue, we propose to rank architectures by Representation Mutual Information (RMI). Specifically, given an arbitrary architecture that has decent accuracy, architectures that have high RMI with it always yield good accuracies. As an accurate performance indicator to facilitate NAS, RMI not only generalizes well to different search spaces, but is also efficient enough to evaluate architectures using only one batch of data. Building upon RMI, we further propose a new search algorithm termed RMI-NAS, facilitating with a theorem to guarantee the global optimal of the searched architecture. In particular, RMI-NAS first randomly samples architectures from the search space, which are then effectively classified as positive or negative samples by RMI. We then use these samples to train a random forest to explore new regions, while keeping track of the distribution of positive architectures. When the sample size is sufficient, the architecture with the largest probability from the aforementioned distribution is selected, which is theoretically proved to be the optimal solution. The architectures searched by our method achieve remarkable top-1 accuracies with the magnitude times faster search process. Besides, RMI-NAS also generalizes to different datasets and search spaces. Our code has been made available at https://git.openi.org.cn/PCL_AutoML/XNAS.
Xiawu Zheng, Lei Zhang 0001, Chenglin Wu 0001, Fei Chao 0001, Jianzhuang Liu, Wei Zeng 0006, Yonghong Tian 0001, Rongrong Ji
CVPR7
2022 Attribute-Aware Feature Encoding for Object Recognition and Segmentation
abstract
Existing multi-task models for object recognition and segmentation have verified the effectiveness of joint optimization of two semantic tasks. However, learning discriminative representations with insufficient training data and redundant contextual information from the background remains challenging. Semantic attributes are designed as powerful and informative mid-level features that 1) share information across categories to model the interclass correlation and that 2) can be localized in the object region to benefit foreground extraction. This paper introduces a novel attribute-aware feature encoding (AFE) module to a multi-task network for object recognition and segmentation with the aim of improving both semantic tasks by regularizing feature encoding with auxiliary attribute learning. Intuitively, attribute learning in our method not only provides extra supervision signals to capture interclass correlation in object classification but also refines the output of object segmentation via weakly supervised attribute localization. The experimental results on two public benchmarks show that our method yields remarkable improvement in both semantic tasks and auxiliary attribute estimation over existing methods.
Shu Yang 0007, Yaowei Wang 0001, Ke Chen 0004, Wei Zeng 0006, Zesong Fei
IEEE Trans. Multim.4
2021 Contrastive Neural Architecture Search With Neural Architecture Comparators
abstract
One of the key steps in Neural Architecture Search (NAS) is to estimate the performance of candidate architectures. Existing methods either directly use the validation performance or learn a predictor to estimate the performance. However, these methods can be either computationally expensive or very inaccurate, which may severely affect the search efficiency and performance. Moreover, as it is very difficult to annotate architectures with accurate performance on specific tasks, learning a promising performance predictor is often non-trivial due to the lack of labeled data. In this paper, we argue that it may not be necessary to estimate the absolute performance for NAS. On the contrary, we may need only to understand whether an architecture is better than a baseline one. However, how to exploit this comparison information as the reward and how to well use the limited labeled data remains two great challenges. In this paper, we propose a novel Contrastive Neural Architecture Search (CTNAS) method which performs architecture search by taking the comparison results between architectures as the reward. Specifically, we design and learn a Neural Architecture Comparator (NAC) to compute the probability of candidate architectures being better than a baseline one. Moreover, we present a baseline updating scheme to improve the baseline iteratively in a curriculum learning manner. More critically, we theoretically show that learning NAC is equivalent to optimizing the ranking over architectures. Extensive experiments in three search spaces demonstrate the superiority of our CTNAS over existing methods.
Yaofo Chen, Qi Chen 0014, Minli Li, Wei Zeng 0006, Yaowei Wang 0001, Mingkui Tan
CVPR5
2021 Adaptive Class Suppression Loss for Long-Tail Object Detection
abstract
To address the problem of long-tail distribution for the large vocabulary object detection task, existing methods usually divide the whole categories into several groups and treat each group with different strategies. These methods bring the following two problems. One is the training inconsistency between adjacent categories of similar sizes, and the other is that the learned model is lack of discrimination for tail categories which are semantically similar to some of the head categories. In this paper, we devise a novel Adaptive Class Suppression Loss (ACSL) to effectively tackle the above problems and improve the detection performance of tail categories. Specifically, we introduce a statistic-free perspective to analyze the long-tail distribution, breaking the limitation of manual grouping. According to this perspective, our ACSL adjusts the suppression gradients for each sample of each class adaptively, ensuring the training consistency and boosting the discrimination for rare categories. Extensive experiments on long-tail datasets LVIS and Open Images show that the our ACSL achieves 5.18% and 5.2% improvements with ResNet50-FPN, and sets a new state of the art. Code and models are available at https://github.com/CASIA-IVA-Lab/ACSL.
Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001
CVPR4
2021 DPT: Deformable Patch-based Transformer for Visual Recognition
abstract
Transformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which might destroy the semantics of objects. To address this problem, we propose a new Deformable Patch (DePatch) module which learns to adaptively split the images into patches with different positions and scales in a data-driven way rather than using predefined fixed patches. In this way, our method can well preserve the semantics in patches. The DePatch module can work as a plug-and-play module, which can easily be incorporated into different transformers to achieve an end-to-end training. We term this DePatch-embedded transformer as Deformable Patch-based Transformer (DPT) and conduct extensive evaluations of DPT on image classification and object detection. Results show DPT can achieve 81.8% top-1 accuracy on ImageNet classification, and 43.7% box AP with RetinaNet, 44.3% with Mask R-CNN on MSCOCO object detection. Code has been made available at: https://github.com/CASIA-IVA-Lab/DPT.
Zhiyang Chen 0002, Yousong Zhu, Chaoyang Zhao, Guosheng Hu, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001
ACM Multimedia5
2020 An Asymmetric Modeling for Action Assessment
Jibin Gao, Wei-Shi Zheng 0001, Chengying Gao, Yaowei Wang 0001, Wei Zeng 0006, Jian-Huang Lai
ECCV (30)6
2020 Large Batch Optimization for Object Detection: Training COCO in 12 minutes
Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Wei Zeng 0006, Yaowei Wang 0001, Jinqiao Wang, Ming Tang 0001
ECCV (21)4
2020 Hybrid Dynamic-static Context-aware Attention Network for Action Assessment in Long Videos
abstract
The objective of action quality assessment is to score sports videos. However, most existing works focus only on video dynamic information (i.e., motion information) but ignore the specific postures that an athlete is performing in a video, which is important for action assessment in long videos. In this work, we present a novel hybrid dynAmic-static Context-aware attenTION NETwork (ACTION-NET) for action assessment in long videos. To learn more discriminative representations for videos, we not only learn the video dynamic information but also focus on the static postures of the detected athletes in specific frames, which represent the action quality at certain moments, along with the help of the proposed hybrid dynamic-static architecture. Moreover, we leverage a context-aware attention module consisting of a temporal instance-wise graph convolutional network unit and an attention unit for both streams to extract more robust stream features, where the former is for exploring the relations between instances and the latter for assigning a proper weight to each instance. Finally, we combine the features of the two streams to regress the final video score, supervised by ground-truth scores given by experts. Additionally, we have collected and annotated the new Rhythmic Gymnastics dataset, which contains videos of four different types of gymnastics routines, for evaluation of action quality assessment in long videos. Extensive experimental results validate the efficacy of our proposed method, which outperforms related approaches.
Ling-An Zeng, Fa-Ting Hong, Wei-Shi Zheng 0001, Qi-Zhi Yu, Wei Zeng 0006, Yaowei Wang 0001, Jian-Huang Lai
ACM Multimedia5
2018 SFCM: Learn a Pooling Kernel for Weakly Supervised Object Localization
abstract
The weakly supervised object localization (WSOL) is to locate the objects in an image while only image-level labels are available during the training procedure. In this work, the Selective Feature Category Mapping (SFCM) method is proposed, which introduces the Feature Category Mapping (FCM) and the widely-used selective search method to solve the WSOL task. Our FCM replaces layers after the specific layer in the state-of-the-art CNNs with a set of kernels and learns the weighted pooling for previous feature maps. It is trained with only image-level labels and then map the feature maps to their corresponding categories in the test phase. Together with selective search method, the location of each object is finally obtained. Extensive experimental evaluation on ILSVRC2012 and PASCAL VOC2007 benchmarks shows that SFCM is simple but very effective, and it is able to achieve outstanding classification performance and outperform the state-of-the-art methods in the WSOL task.
Zongxian Li, Yemin Shi 0001, Yonghong Tian 0001, Wei Zeng 0006, Yaowei Wang 0001
ICME4
2017 Learning Long-Term Dependencies for Action Recognition with a Biologically-Inspired Deep Network
abstract
Despite a lot of research efforts devoted in recent years, how to efficiently learn long-term dependencies from sequences still remains a pretty challenging task. As one of the key models for sequence learning, recurrent neural network (RNN) and its variants such as long short term memory (LSTM) and gated recurrent unit (GRU) are still not powerful enough in practice. One possible reason is that they have only feedforward connections, which is different from the biological neural system that is typically composed of both feedforward and feedback connections. To address this problem, this paper proposes a biologicallyinspired deep network, called shuttleNet. Technologically, the shuttleNet consists of several processors, each of which is a GRU while associated with multiple groups of hidden states. Unlike traditional RNNs, all processors inside shuttleNet are loop connected to mimic the brain's feedforward and feedback connections, in which they are shared across multiple pathways in the loop connection. Attention mechanism is then employed to select the best information flow pathway. Extensive experiments conducted on two benchmark datasets (i.e UCF101 and HMDB51) show that we can beat state-of-the-art methods by simply embedding shuttleNet into a CNN-RNN framework.
Yemin Shi 0001, Yonghong Tian 0001, Yaowei Wang 0001, Wei Zeng 0006, Tiejun Huang 0001
ICCV4
2017 Exploiting Multi-grain Ranking Constraints for Precisely Searching Visually-similar Vehicles
abstract
Precise search of visually-similar vehicles poses a great challenge in computer vision, which needs to find exactly the same vehicle among a massive vehicles with visually similar appearances for a given query image. In this paper, we model the relationship of vehicle images as multiple grains. Following this, we propose two approaches to alleviate the precise vehicle search problem by exploiting multi-grain ranking constraints. One is Generalized Pairwise Ranking, which generalizes the conventional pairwise from considering only binary similar/dissimilar relations to multiple relations. The other is Multi-Grain based List Ranking, which introduces permutation probability to score a permutation of a multi-grain list, and further optimizes the ranking by the likelihood loss function. We implement the two approaches with multi-attribute classification in a multi-task deep learning framework. To further facilitate the research on precise vehicle search, we also contribute two high-quality and well-annotated vehicle datasets, named VD1 and VD2, which are collected from two different cities with diverse annotated attributes. As two of the largest publicly available precise vehicle search datasets, they contain 1,097,649 and 807,260 vehicle images respectively. Experimental results show that our approaches achieve the state-of-the-art performance on both datasets.
Ke Yan 0007, Yonghong Tian 0001, Yaowei Wang 0001, Wei Zeng 0006, Tiejun Huang 0001
ICCV4
2015 Detecting abnormal behaviors in surveillance videos based on fuzzy clustering and multiple Auto-Encoders
abstract
In this paper, we present a novel framework to detect abnormal behaviors in surveillance videos by using fuzzy clustering and multiple Auto-Encoders (FMAE). As detecting abnormal behaviors is often treated as an unsupervised task, how to describe normal patterns becomes the key point. Considering there are many types of normal behaviors in the daily life, we use the fuzzy clustering technique to roughly divide the training samples into several clusters so that each cluster stands for a normal pattern. Then we deploy multiple Auto-Encoders to estimate these different types of normal behaviors from weighted samples. When testing on an unknown video, our framework can predict whether it contains abnormal behaviors or not by summarizing the reconstruction cost through each Auto-Encoder. Since there are always lots of redundancies in the surveillance video, Auto-Encoder is a pretty good tool to capture common structures of normal video sequences automatically as well as estimate normal patterns. The experimental results show that our approach achieves good performance on three public video analysis datasets and statistically outperforms the state-of-the-art approaches under some scenes.
Zhengying Chen, Yonghong Tian 0001, Wei Zeng 0006, Tiejun Huang 0001
ICME3
2015 Learning Deep Trajectory Descriptor for action recognition in videos using deep neural networks
abstract
Human action recognition is widely recognized as a challenging task due to the difficulty of effectively characterizing human action in a complex scene. Recent studies have shown that the dense-trajectory-based methods can achieve state-of-the-art recognition results on some challenging datasets. However, in these methods, each dense trajectory is often represented as a vector of coordinates, consequently losing the structural relationship between different trajectories. To address the problem, this paper proposes a novel Deep Trajectory Descriptor (DTD) for action recognition. First, we extract dense trajectories from multiple consecutive frames and then project them onto a canvas. This will result in a “trajectory texture” image which can effectively characterize the relative motion in these frames. Based on these trajectory texture images, a deep neural network (DNN) is utilized to learn a more compact and powerful representation of dense trajectories. In the action recognition system, the DTD descriptor, together with other non-trajectory features such as HOG, HOF and MBH, can provide an effective way to characterize human action from various aspects. Experimental results show that our system can statistically outperform several state-of-the-art approaches, with an average accuracy of 95:6% on KTH and an accuracy of 92.14% on UCF50.
Yemin Shi 0001, Wei Zeng 0006, Tiejun Huang 0001, Yaowei Wang 0001
ICME2
2014 Data-driven hair segmentation with isomorphic manifold inference
Shiguang Shan, Hongming Zhang 0011, Wei Zeng 0006, Xilin Chen 0001
Image Vis. Comput.4
2013 Night video enhancement using improved dark channel prior
abstract
Videos taken under low lighting condition usually have serious loss of visibility and contrast and are inconvenient for observation and analysis. To solve this problem, this paper presents a real-time night video enhancement approach. As observed that a pixel-wise inversion of a night video has quite similar appearance with the video acquired at foggy days, we use the similar idea of haze removal method to enhance the perceptual quality of the night videos. We present an improved dark channel prior model and integrate it with local smoothing and image Gaussian Pyramid operators. The experimental results demonstrate that the proposed approach can improve the perceptual quality of night videos in real-time in terms of not only enhancing details, but also effectively avoiding excessive enhancement phenomenon.
Xuesong Jiang, Hongxun Yao, Shengping Zhang, Xiusheng Lu, Wei Zeng 0006
ICIP5
2013 Effective constructing training sets for object detection
abstract
This paper addresses the problem of building up effective training sets at minimal labeling cost for object detection. This problem occurs in the situation that the part-based detector is trained on a group of positive examples with bounding box labels, but the images selected by uniform sampling do not reflect the desired training distribution and need additional labeling cost in order to obtain enough positive examples. We study the active training process in which some object windows are sampled from a pool of unlabeled candidate windows, and then their corresponding bounding annotations are queried. We derive an effective training set by selecting a group of most uncertain object windows according to the current detector. Our approach has been empirically demonstrated on the object detection task of PASCAL VOC dataset. The experiment results show that our proposed algorithm outperforms common uniform sampling within the same labeling cost.
Weining Wu, Yang Liu 0006, Wei Zeng 0006, Maozu Guo 0001, Chunyu Wang 0002
ICIP3
2013 Abnormal event detection in crowded scenes based on Structural Multi-scale Motion Interrelated Patterns
abstract
Detecting abnormal events in crowded scenes remains challenging due to the diversity of events defined by various applications. Among the many application situations, motion analysis for event representation is suited for crowded scenes. In this paper, we propose a novel abnormal event detection method via likelihood estimation of dynamic-texture motion representation, called Structural Multi-scale Motion Interrelated Patterns (SMMIP). SMMIP combines both original motion patterns and their structural spatio-temporal information, which effectively represents localized events by different resolutions of motion patterns. To model normal events, the Gaussian mixture model is trained with the observed normal events, then the likelihood estimation for testing events is computed to judge whether they are abnormal. Meanwhile, the proposed model can be learned online by updating the parameters incrementally. The proposed approach is evaluated on several publicly available datasets and outperforms several other methods proposed before, which is shown that the structural spatio-temporal information added in motion representation helps increasing the anomalies detection rate.
Dawei Du, Honggang Qi, Qingming Huang, Wei Zeng 0006, Changhua Zhang
ICME4
2012 Single and Multiple View Detection, Tracking and Video Analysis in Crowded Environments
abstract
In this paper, we present our detection, tracking and event recognition methods and the results for PETS 2012. First, ROIs (Regions of Interest) based on geometric constraints are utilized in single view detection to eliminate the negative influence of clutter environment. Then, an optimized observation model is applied to address the ID switching or tracking drifting problem in single view tracking. Third, we introduce the multi-view Bayesian network (MBN) to reduce the "phantom" phenomena which frequently happen in general multi-view detection tasks. At last, a motion-based event recognition method is proposed to handle the event recognition task. Experimental results on the PETS 2012 dataset indicate that our methods are very promising.
Teng Xu 0002, Peixi Peng, Xiaoyu Fang, Chi Su, Yaowei Wang 0001, Yonghong Tian 0001, Wei Zeng 0006, Tiejun Huang 0001
AVSS7
2011 A novel coarse-to-fine hair segmentation method
abstract
Segmenting hair regions from human images facilitates many tasks like hair synthesis and hair style trends forecast. However, hair segmentation is quite challenging due to hair/background confusion and large hair pattern diversity. To address these problems to some extent, this paper proposes a novel coarse-to-fine hair segmentation method. In our approach, firstly, the recently proposed “Active Segmentation with Fixation” (ASF) is used to coarsely define an enclosed candidate region with high-recall (but possibly low-precision) of hair pixels and exclude considerable part of the backgrounds which are easily confused with hair. Then Graph Cuts (GC) method is applied to the candidate regions to remove additional false positives by incorporating hair-specific information. Specifically, Bayesian method is employed to select some reliable hair and background regions (seeds) among the ones over-segmented by Mean Shift. SVM classifier is then learnt online from these seeds and explored to predict hair/background likelihood probability, which is subsequently fed into GC algorithm. The novelty of the proposed approach lies in three folds: 1) an elaborate design of hair segmentation framework, which utilizes ASF to reduce the candidate hair regions and adopts GC to achieve more accurate hair region contours; 2) the region-based strategy for seed selection; 3) the exploration of the discriminative method, SVM, to predict the probability of each pixel belonging to hair and background regions. Extensive experimental results demonstrate the approach outperforms recently proposed methods.
Xiujuan Chai, Hongming Zhang 0011, Hong Chang 0001, Wei Zeng 0006, Shiguang Shan
FG5
2011 Exemplar-Based Image Inpainting with Collaborative Filtering
abstract
This paper proposes a novel patch synthesis approach for exemplar-based propagation in image in painting. Currently, plural non-local exemplar patches synthesis is widely adopted to fill missing pixels. It generally provides good results, but sometimes shows poor visual quality due to dissimilarity between exemplars and targets. In this paper, a collaborative filtering approach is used to enhance the exemplar-based propagation to obtain ideal in painting results. The approach works on pixel level information, while many exemplar-based propagation algorithms focus on patch level information. Object removal and stain image recovering are carried out to evaluate the proposed approach. Experiments show that our approach provides good visual quality in object removal and high PSNR in stain image recovering.
Wei Zeng 0006, Zhenzhou Li
ICIG2
2009 Foreground Based Borderline Adjusting for Real Time Multi-camera Video Stitching
abstract
In this paper, we propose a multi-camera video stitching approach, which conducts borderline adjusting based on foreground information. In real time applications of video stitching, one problem is to deal with dynamic scenes, where foreground objects often cause broken objects like artifacts in panoramic video. To address this problem, we propose a foreground based borderline adjusting method to achieve smooth video stitching. In this method, foreground objects are extracted from different viewpoints, and the stitching borderlines are adjusted according to the foreground content in dynamic scenes. Guided by the adjusted borderlines, foreground objects are smoothly synthesized into panoramic video. Experiment results show that this approach obtains more than 81% correct rate for dynamic scenes, compared with 25%~50% correct rate on the condition of not utilizing foreground information, and achieves processing speed of 10~13 frames per second.
Hongming Zhang 0011, Wei Zeng 0006
ICIG2
2009 A novel two-tier Bayesian based method for hair segmentation
abstract
In this paper, a novel two-tier Bayesian based method is proposed for hair segmentation. In the first tier, we construct a Bayesian model by integrating hair occurrence prior probabilities (HOPP) with a generic hair color model (GHCM) to obtain some reliable hair seed pixels. These initial seeds are further propagated to their neighborhood pixels by utilizing segmentation results of mean shift, to obtain more seeds. In the second tier, all of these selected seeds are used to train a hair-specific Gaussian model, which are combined with HOPP to build the second Bayesian model for pixel classification. Mean shift results are further utilized to remove holes and spread hair regions. The experimental results illustrate the effectiveness of our approach.
Shiguang Shan, Wei Zeng 0006, Hongming Zhang 0011, Xilin Chen 0001
ICIP3
2005 Local invariant descriptor for image matching
abstract
Image matching is a fundamental task of many computer vision problems. In this paper we present a novel approach for matching two images in the presence of image rotation, scale, and illumination changes. The proposed approach is based on local invariant features. A two-step process detects local invariant regions. Characteristic circles associated with these regions illustrate the position and radius of the regions. Then, the regions are represented by a new image descriptor. To test the new descriptor, we evaluate it in image matching and retrieval experiments. The experimental results show that using our descriptors results in effective and faster matching.
Wei Zeng 0006, Wen Gao 0001, Weiqiang Wang 0001
ICASSP (2)2
2005 Adaptive relevance feedback based on Bayesian inference for image retrieval
Lijuan Duan, Wen Gao 0001, Wei Zeng 0006, Debin Zhao
Signal Process.3
2004 Accurate moving object segmentation by a hierarchical region labeling approach
abstract
This paper proposes a new algorithm to segment moving objects from color sequences accurately. The segmentation procedure is treated as a Markovian labeling process and is formulated by a hierarchical Markov random field (MRF) model. Initially, the original frame is partitioned into homogeneous regions with different granularity by the rapid watershed algorithm. Then, the foreground is detected as outliers of the estimated background motion in the initial motion classification stage. After that, the motion vector is estimated for each foreground region and is validated by an elaborate occlusion detection scheme. The initial object mask is segmented by the MRF model on the larger-scale spatial partition and is refined by the other MRF model in the small-scale partition. The hierarchical MRF models provide the fine object boundary. The proposed method is evaluated on several real-world image sequences and the experimental results shows remarkable performance.
Wei Zeng 0006, Wen Gao 0001
ICASSP (3)1
2004 Shape-based adult images detection
abstract
This paper reports an investigation on adult images detection based on the shape features of skin regions. In order to accurately detect skin regions, we propose a skin detection method using multi-Bayes classifiers in the paper. Based on skin color detection results, shape features are extracted and fed into a boosted classifier to decide whether or not the skin regions represent a nude. We evaluate adult image detection performance using different boosted classifiers and different shape descriptors. Experimental results show that classification using boosted C4.5 classifier and combination of different shape descriptors outperforms other classification schemes.
Wei Zeng 0006, Wen Gao 0001, Weiqiang Wang 0001
ICIG2
2004 A novel compressed domain shot segmentation algorithm on H.264/AVC
abstract
This paper presents a novel shot segmentation algorithm on the H.264/AVC video, which operates in the compressed domain. First, the algorithm exploits the intra prediction mode histogram to locate those potential GOPs, where shot transitions occur with great probability. Secondly, to further find shot boundaries at the frame level, we count the number of macroblocks with different inter prediction modes as the features and exploit HMMs to automatically model different cases in which shot transitions can occur among I, P and B frames. Since H.264/AVC provides more motion compensation modes, using HMMs can avoid the tediousness of manually tuning multiple thresholds simultaneously. The experimental results show that the algorithm is efficient and robust and it can not only locate cuts, but also work for gradual shot transitions.
Yang Liu 0006, Weiqiang Wang 0001, Wen Gao 0001, Wei Zeng 0006
ICIP4
2003 Color image segmentation using density-based clustering
abstract
Color image segmentation is an important but still open problem in image processing. We propose a method for this problem by integrating the spatial connectivity and color features of the pixels. Considering that an image can be regarded as a dataset in which each pixel has a spatial location and a color value, color image segmentation can be obtained by clustering these pixels into different groups of coherent spatial connectivity and color. To discover the spatial connectivity of the pixels, density-based clustering is employed, which is an effective clustering method used in data mining for discovering spatial databases. The color similarity of the pixels is measured in Munsell (HVC) color space whose perceptual uniformity ensures the color change in the segmented regions is smooth in terms of human perception. Experimental results using the proposed method demonstrate encouraging performance.
Qixiang Ye, Wen Gao 0001, Wei Zeng 0006
ICASSP (3)3
2003 Color image segmentation using density-based clustering
abstract
Color image segmentation is an important but still open problem in image processing. In this paper, we propose a method for this problem by integrating the spatial connectivity and color feature of the pixels. Considering that an image can be regarded as a dataset in which each pixel has a spatial location and a color value, color image segmentation can be obtained by clustering these pixels into different groups of coherent spatial connectivity and color. To discover the spatial connectivity of the pixels, density-based clustering is employed, which is an effective clustering method used in data mining for discovering spatial databases. Color similarity of the pixels is measured in Munsell (HVC) color space whose perceptual uniformity ensures the color change in the segmented regions is smooth in terms of human perception. Experimental results using proposed method demonstrate encouraging performance.
Qixiang Ye, Wen Gao 0001, Wei Zeng 0006
ICME3
2003 Objectionable Image Recognition System in Compression Domain
Qixiang Ye, Wen Gao 0001, Wei Zeng 0006, Weiqiang Wang 0001, Yang Liu 0006
IDEAL3
2002 Video indexing by motion activity maps
abstract
Motion based video indexing is an important and active research area in content-based video retrieval. It explores the dynamic characteristics of video content and provides techniques for video representation and retrieval. In this paper, a new motion-based approach is proposed, in which the image generated from motion activity is used to index video. The proposed approach firstly computes the accumulation measurement of motion activity on the grids of video frames along the time axis. Then, the computed measurement is quantized into several gray levels and a gray image is generated from all measurements on the grids. The generated image called motion activity map (MAM) reserves the motion information on different spatial locations of video. The intensity of MAM is corresponding to the magnitude of motion activity. All MAMs are organized into a hierarchical tree according to the video structure. Therefore, user can browse video by the MAMs tree. The proposed approach provides a hierarchical, coarse to fine view of video and thus makes interactive video retrieval more simply and intuitively.
Wei Zeng 0006, Wen Gao 0001, Debin Zhao
ICIP (1)1