Yu Zhang 0004

dblp:50/671-4 · DBLP profile ↗
← Back
45ranked-venue papers
16as first author
22since 2021 · last 2026
0000-0003-0140-7366ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 7 since 2021Computer networks · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Efficient and Effective In-context Demonstration Selection with Coreset
abstract
In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection.
Zihua Wang, Jiarui Wang 0002, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002, Xu Yang 0021, Xiu-Shen Wei, Siya Mi, Yu Zhang 0004
AAAI9
2026 Adaptively Clustering Neighbor Elements for Image-Text Generation
abstract
We propose a novel Transformer-based image-to-text generation model termed asACFthat adaptively clusters vision patches into object regions and language words into phrases to implicitly learn object-phrase alignments for better visual-text coherence. To achieve this, we design a novel self-attention layer that applies self-attention over the elements in a local cluster window instead of the whole sequence. The window size is softly decided by a clustering matrix that is calculated by the current input data and thus this process is adaptive. By stacking these revised self-attention layers to construct ACF, the small clusters in the lower layers can be grouped into a bigger cluster, e.g., vision/language. ACF clusters small objects/phrases into bigger ones. In this gradual clustering process, a parsing tree is generated which embeds the hierarchical knowledge of the input sequence. As a result, by using ACF to build the vision encoder and language decoder, the hierarchical object-phrase alignments are embedded and then transferred from vision to language domains in two popular image-to-text tasks: Image captioning and Visual Question Answering. The experiment results demonstrate the effectiveness of ACF, which outperforms most SOTA captioning and VQA models and achieves comparable scores compared with some large-scale pre-trained models.
Zihua Wang, Xu Yang 0021, Haiyang Xu 0001, Hanwang Zhang, Ming Yan 0008, Fei Huang 0002, Yu Zhang 0004
IEEE Trans. Multim.7
2025 Object-Level Correlation for Few-Shot Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment objects of novel categories in the query images given only a few annotated support samples. Existing methods primarily build the image-level correlation between the support target object and the entire query image. However, this correlation contains the hard pixel noise, \textit{i.e.}, irrelevant background objects, that is intractable to trace and suppress, leading to the overfitting of the background. To address the limitation of this correlation, we imitate the biological vision process to identify novel objects in the object-level information. Target identification in the general objects is more valid than in the entire image, especially in the low-data regime. Inspired by this, we design an Object-level Correlation Network (OCNet) by establishing the object-level correlation between the support target object and query general objects, which is mainly composed of the General Object Mining Module (GOMM) and Correlation Construction Module (CCM). Specifically, GOMM constructs the query general object feature by learning saliency and high-level similarity cues, where the general objects include the irrelevant background objects and the target foreground object. Then, CCM establishes the object-level correlation by allocating the target prototypes to match the general object feature. The generated object-level correlation can mine the query target feature and suppress the hard pixel noise for the final prediction. Extensive experiments on PASCAL-${5}^{i}$ and COCO-${20}^{i}$ show that our model achieves the state-of-the-art performance.
Chunlin Wen, Yu Zhang 0004, Hongyuan Zhu 0002, Xiu-Shen Wei, Zhiqiang Kou, Shuzhou Sun
ICCV2
2025 DVC2: Deep video cascade clustering from video structures
Zihua Wang, Siya Mi, Yu Zhang 0004
Neurocomputing3
2025 Object Adaptive Self-Supervised Dense Visual Pre-Training
abstract
Self-supervised visual pre-training models have achieved significant success without employing expensive annotations. Nevertheless, most of these models focus on iconic single-instance datasets (e.g. ImageNet), ignoring the insufficient discriminative representation for non-iconic multi-instance datasets (e.g. COCO). In this paper, we propose a novel Object Adaptive Dense Pre-training (OADP) method to learn the visual representation directly on the multi-instance datasets (e.g., PASCAL VOC and COCO) for dense prediction tasks (e.g., object detection and instance segmentation). We present a novel object-aware and learning-adaptive random view augmentation to focus the contrastive learning to enhance the discrimination of object presentations from large to small scale during different learning stages. Furthermore, the representations across different scale and resolutions are integrated so that the method can learn diverse representations. In the experiment, we evaluated OADP pre-trained on PASCAL VOC and COCO. Results show that our method has better performances than most existing state-of-the-art methods when transferring to various downstream tasks, including image classification, object detection, instance segmentation and semantic segmentation.
Yu Zhang 0004, Hongyuan Zhu 0002, Siya Mi, Xi Peng 0001, Xin Geng 0001
IEEE Trans. Image Process.1
2024 IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
abstract
As Large Language Models (LLMs) become more capable of handling increasingly complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the ability of individual test items to differentiate between high and low performers. Inspired by this theory, we propose an ID-induced prompt synthesis framework for evaluating LLMs so that the evaluation set continually updates and refines according to model abilities. Our data synthesis framework prioritizes both breadth and specificity. It can generate prompts that comprehensively evaluate the capabilities of LLMs while revealing meaningful performance differences between models, allowing for effective discrimination of their relative strengths and weaknesses across various tasks and domains. To produce high-quality data, we incorporate a self-correct mechanism into our generalization framework and develop two models to predict prompt discrimination and difficulty score to facilitate our data synthesis framework, contributing valuable tools to evaluation data synthesis research. We apply our generated data to evaluate five SOTA models. Our data achieves an average score of 51.92, accompanied by a variance of 10.06. By contrast, previous works (i.e., SELF-INSTRUCT and WizardLM) obtain an average score exceeding 67, with a variance below 3.2. The results demonstrate that the data generated by our framework is more challenging and discriminative compared to previous works. We will release a dataset of over 3,000 carefully crafted prompts to facilitate evaluation research of LLMs.
Fan Lin, Shuyi Xie, Yong Dai 0001, Wenlin Yao, Tianjiao Lang, Yu Zhang 0004
NeurIPS6
2024 Temporal segment dropout for human action video recognition
Yu Zhang 0004, Zhengjie Chen, Siya Mi, Xin Geng 0001, Min-Ling Zhang
Pattern Recognit.1
2023 Transforming Visual Scene Graphs to Image Captions
abstract
Xu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu, Qinghao Ye, Chenliang Li, Songfang Huang, Fei Huang, Zhangzikang Li, Yu Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xu Yang 0021, Jiawei Peng 0001, Zihua Wang, Haiyang Xu 0001, Qinghao Ye, Chenliang Li 0003, Songfang Huang, Fei Huang 0002, Zhangzikang Li, Yu Zhang 0004
ACL (1)10
2023 Learning Trajectory-Word Alignments for Video-Language Tasks
abstract
In a video, an object usually appears as the trajectory, i.e., it spans over a few spatial but longer temporal patches, that contains abundant spatiotemporal contexts. However, modern Video-Language BERTs (VDL-BERTs) neglect this trajectory characteristic that they usually follow image-language BERTs (IL-BERTs) to deploy the patch-to-word (P2W) attention that may over-exploit trivial spatial contexts and neglect significant temporal contexts. To amend this, we propose a novel TW-BERT to learn Trajectory-Word alignment by a newly designed trajectory-to-word (T2W) attention for solving video-language tasks. Moreover, previous VDL-BERTs usually uniformly sample a few frames into the model while different trajectories have diverse graininess, i.e., some trajectories span longer frames and some span shorter, and using a few frames will lose certain useful temporal contexts. However, simply sampling more frames will also make pre-training infeasible due to the largely increased training burdens. To alleviate the problem, during the fine-tuning stage, we insert a novel Hierarchical Frame-Selector (HFS) module into the video encoder. HFS gradually selects the suitable frames conditioned on the text context for the later cross-modal encoder to learn better trajectory-word alignments. By the proposed T2W attention and HFS, our TW-BERT achieves SOTA performances on text-to-video retrieval tasks, and comparable performances on video question-answering tasks with some VDL-BERTs trained on much more data. The code will be available in the supplementary material.
Xu Yang 0021, Zhangzikang Li, Haiyang Xu 0001, Hanwang Zhang, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Yu Zhang 0004, Fei Huang 0002, Songfang Huang
ICCV8
2023 Label distribution learning for scene text detection
Ningning Lu, Junjun Mei, Yu Zhang 0004, Xin Geng 0001
Frontiers Comput. Sci.5
2023 Asymmetric bi-encoder for image-text retrieval
Haoliang Liu, Siya Mi, Yu Zhang 0004
Multim. Syst.4
2023 Self-label correction for image classification with noisy labels
Yu Zhang 0004, Fan Lin, Siya Mi, Yali Bian
Pattern Anal. Appl.1
2023 LPCL: Localized prominence contrastive learning for self-supervised dense visual pre-training
Hongyuan Zhu 0002, Siya Mi, Yu Zhang 0004, Xin Geng 0001
Pattern Recognit.5
2023 Assisting Multimodal Named Entity Recognition by cross-modal auxiliary tasks
Zhengjie Chen, Yu Zhang 0004, Siya Mi
Pattern Recognit. Lett.2
2023 A Closer Look at Video Sampling for Sequential Action Recognition
abstract
In recent years, sequential action recognition has attracted increasingly attention as it requires long-term sequential and compositional reasoning of human actions and object interactions. Existing methods perform reasoning either by using snippets that cover very short consecutive frames or key frames sampled from segments, which take a bias process of local and global temporal information. We also find ad-hoc training and ensembling of two separate networks using existing sampling strategies can easily outperform complex state-of-the-art methods, which reveals the complementary nature of current sampling strategies. Motivated by this observation, we propose a simple yet efficient strategy named Dense Segmental Sampling (DSS) and a novel network architecture named Temporal Dense Segment Network (TDSN) to capture the complementary information from DSS. Our TDSN achieves excellent results on benchmark action recognition datasets, which not only validate the proposed strategy but also help highlight the importance along this direction for sequential video reasoning.
Yu Zhang 0004, Zhengjie Chen, Siya Mi, Hongyuan Zhu 0002, Xin Geng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Pose-guided action recognition in static images using lie-group
Siya Mi, Yu Zhang 0004
Appl. Intell.2
2022 Weakly supervised temporal action localization with proxy metric modeling
Yu Zhang 0004, Xin Geng 0001, Siya Mi, Zhihong Yang
Frontiers Comput. Sci.3
2022 Head Pose Estimation Based on Multivariate Label Distribution
abstract
Accurate ground-truth pose is essential to the training of most existing head pose estimation methods. However, in many cases, the "ground truth" pose is obtained in rather subjective ways, such as asking the subjects to stare at different markers on the wall. Thus it is better to use soft labels rather than explicit hard labels to indicate the pose of a face image. This paper proposes to associate a multivariate label distribution (MLD) to each image. An MLD covers a neighborhood around the original pose. Labeling the images with MLD can not only alleviate the problem of inaccurate pose labels, but also boost the training examples associated to each pose without actually increasing the total amount of training examples. Four algorithms are proposed to learn from MLD. Furthermore, an extension of MLD with the hierarchical structure is proposed to deal with fine-grained head pose estimation, which is named hierarchical multivariate label distribution (HMLD). Experimental results show that the MLD-based methods perform significantly better than the compared state-of-the-art head pose estimation algorithms. Moreover, the MLD-based methods appear much more robust against the label noise in the training set than the compared baseline methods.
Xin Geng 0001, Zeng-Wei Huo, Yu Zhang 0004
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Multilabel Ranking With Inconsistent Rankers
abstract
While most existing multilabel ranking methods assume the availability of a single objective label ranking for each instance in the training set, this paper deals with a more common case where only subjective inconsistent rankings from multiple rankers are associated with each instance. Two ranking methods are proposed from the perspective of instances and rankers, respectively. The first method, Instance-oriented Preference Distribution Learning (IPDL), is to learn a latent preference distribution for each instance. IPDL generates a common preference distribution that is most compatible to all the personal rankings, and then learns a mapping from the instances to the preference distributions. The second method, Ranker-oriented Preference Distribution Learning (RPDL), is proposed by leveraging interpersonal inconsistency among rankers, to learn a unified model from personal preference distribution models of all rankers. These two methods are applied to natural scene images dataset and 3D facial expression dataset BU_3DFE. Experimental results show that IPDL and RPDL can effectively incorporate the information given by the inconsistent rankers, and perform remarkably better than the compared state-of-the-art multilabel ranking algorithms.
Xin Geng 0001, RenYi Zheng, Yu Zhang 0004
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Learning time-aware features for action quality assessment
Yu Zhang 0004, Siya Mi
Pattern Recognit. Lett.1
2021 Practical age estimation using deep label distribution learning
Yu Zhang 0004, Xin Geng 0001
Frontiers Comput. Sci.2
2021 Image captioning with transformer and knowledge graph
Yu Zhang 0004, Xinyu Shi 0003, Siya Mi, Xu Yang 0021
Pattern Recognit. Lett.1
2020 Label Enhancement for Label Distribution Learning via Prior Knowledge
abstract
Label distribution learning (LDL) is a novel machine learning paradigm that gives a description degree of each label to an instance. However, most of training datasets only contain simple logical labels rather than label distributions due to the difficulty of obtaining the label distributions directly. We propose to use the prior knowledge to recover the label distributions. The process of recovering the label distributions from the logical labels is called label enhancement. In this paper, we formulate the label enhancement as a dynamic decision process. Thus, the label distribution is adjusted by a series of actions conducted by a reinforcement learning agent according to sequential state representations. The target state is defined by the prior knowledge. Experimental results show that the proposed approach outperforms the state-of-the-art methods in both age estimation and image emotion recognition.
Yongbiao Gao, Yu Zhang 0004, Xin Geng 0001
IJCAI2
2020 Label Distribution for Learning with Noisy Labels
abstract
The performances of deep neural networks (DNNs) crucially rely on the quality of labeling. In some situations, labels are easily corrupted, and therefore some labels become noisy labels. Thus, designing algorithms that deal with noisy labels is of great importance for learning robust DNNs. However, it is difficult to distinguish between clean labels and noisy labels, which becomes the bottleneck of many methods. To address the problem, this paper proposes a novel method named Label Distribution based Confidence Estimation (LDCE). LDCE estimates the confidence of the observed labels based on label distribution. Then, the boundary between clean labels and noisy labels becomes clear according to confidence scores. To verify the effectiveness of the method, LDCE is combined with the existing learning algorithm to train robust DNNs. Experiments on both synthetic and real-world datasets substantiate the superiority of the proposed algorithm against state-of-the-art methods.
Ning Xu 0009, Yu Zhang 0004, Xin Geng 0001
IJCAI3
2020 Large-scale multi-label classification using unknown streaming images
Yu Zhang 0004, Xu-Ying Liu, Siya Mi, Min-Ling Zhang
Pattern Recognit.1
2020 Simultaneous 3D hand detection and pose estimation using single depth images
Yu Zhang 0004, Siya Mi, Jianxin Wu 0001, Xin Geng 0001
Pattern Recognit. Lett.1
2019 Recurrent age estimation
Xin Geng 0001, Yu Zhang 0004, Fanyong Cheng
Pattern Recognit. Lett.3
2018 Correction to: Lie-X: Depth Image Based Articulated Object Pose Estimation, Tracking, and Action Recognition on Lie Groups
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Yu Zhang 0004, Zoe Bichler, Suresh Jesuthasan, Adam Claridge-Chang, Ajay Sriram Mathuru, Wenlong Tang, Peixin Zhu, Li Cheng 0001
Int. J. Comput. Vis.3
2017 Lie-X: Depth Image Based Articulated Object Pose Estimation, Tracking, and Action Recognition on Lie Groups
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Yu Zhang 0004, Li Cheng 0001
Int. J. Comput. Vis.3
2016 Exploit Bounding Box Annotations for Multi-Label Object Recognition
abstract
Convolutional neural networks (CNNs) have shown great performance as general feature representations for object recognition applications. However, for multi-label images that contain multiple objects from different categories, scales and locations, global CNN features are not optimal. In this paper, we incorporate local information to enhance the feature discriminative power. In particular, we first extract object proposals from each image. With each image treated as a bag and object proposals extracted from it treated as instances, we transform the multi-label recognition problem into a multi-class multi-instance learning problem. Then, in addition to extracting the typical CNN feature representation from each proposal, we propose to make use of ground-truth bounding box annotations (strong labels) to add another level of local information by using nearest-neighbor relationships of local regions to form a multi-view pipeline. The proposed multi-view multiinstance framework utilizes both weak and strong labels effectively, and more importantly it has the generalization ability to even boost the performance of unseen categories by partial strong labels from other categories. Our framework is extensively compared with state-of-the-art handcrafted feature based methods and CNN based methods on two multi-label benchmark datasets. The experimental results validate the discriminative power and the generalization ability of the proposed framework. With strong labels, our framework is able to achieve state-of-the-art results in both datasets.
Hao Yang 0033, Joey Tianyi Zhou, Yu Zhang 0004, Bin-Bin Gao, Jianxin Wu 0001, Jianfei Cai 0001
CVPR3
2016 Efficient object feature selection for action recognition
abstract
Currently most action recognition or video classification tasks highly rely on the motion features such as state-of-the-art Improved Dense Trajectory (IDT) features. Despite the huge success, IDT features lack of rich static object-level information. In this paper, we make use of the object-level features for action recognition tasks. For efficiently and effectively processing large-scale video data, we propose a two-layer feature selection framework including local object feature selection (LS) and global feature selection (GS). Both of the selection methods can improve recognition accuracy while greatly reducing the feature dimension or feature processing complexity. Experimental results show that the selected object-level features contain complimentary information to IDT features and the combination with IDT features can further improve the recognition accuracy significantly.
Tianyi Zhang 0004, Yu Zhang 0004, Jianfei Cai 0001, Alex Chichung Kot
ICASSP2
2016 Material segmentation in hyperspectral images with minimal region perimeters
abstract
We propose a supervised approach to the classification and segmentation of material regions in hyperspectral imagery. Our algorithm is a two-stage process, combining a pixelwise classification step with a segmentation step aiming to minimise the total perimeters of the resulting regions. Our algorithm is distinctive in its ability to ensure label consistency within local homogeneous areas and to generate material segments with smooth boundaries. Furthermore, we establish a new hyperspectral benchmark dataset to demonstrate the advantages of the proposed approach over several state-of-the-art methods.
Yu Zhang 0004, Cong Phuoc Huynh, Nariman Habili, King Ngi Ngan
ICIP1
2016 Learning compact binary codes from higher-order tensors via Free-Form Reshaping and Binarized Multilinear PCA
abstract
For big, high-dimensional dense features, it is important to learn compact binary codes or compress them for greater memory efficiency. This paper proposes a Binarized Multilinear PCA (BMP) method for this problem with Free-Form Reshaping (FFR) of such features to higher-order tensors, lifting the structure-modelling restriction in traditional tensor models. The reshaped tensors are transformed to a subspace using multilinear PCA. Then, we unsupervisedly select features and supervisedly binarize them with a minimum-classification-error scheme to get compact binary codes. We evaluate BMP on two scene recognition datasets against state-of-the-art algorithms. The FFR works well in experiments. With the same number of compression parameters (model size), BMP has much higher classification accuracy. To achieve the same accuracy or compression ratio, BMP has an order of magnitude smaller number of compression parameters. Thus, BMP has great potential in memory-sensitive applications such as mobile computing and big data analytics.
Haiping Lu, Jianxin Wu 0001, Yu Zhang 0004
IJCNN3
2016 Good Practices for Learning to Recognize Actions Using FV and VLAD
abstract
High dimensional representations such as Fisher vectors (FV) and vectors of locally aggregated descriptors (VLAD) have shown state-of-the-art accuracy for action recognition in videos. The high dimensionality, on the other hand, also causes computational difficulties when scaling up to large-scale video data. This paper makes three lines of contributions to learning to recognize actions using high dimensional representations. First, we reviewed several existing techniques that improve upon FV or VLAD in image classification, and performed extensive empirical evaluations to assess their applicability for action recognition. Our analyses of these empirical results show that normality and bimodality are essential to achieve high accuracy. Second, we proposed a new pooling strategy for VLAD and three simple, efficient, and effective transformations for both FV and VLAD. Both proposed methods have shown higher accuracy than the original FV/VLAD method in extensive evaluations. Third, we proposed and evaluated new feature selection and compression methods for the FV and VLAD representations. This strategy uses only 4% of the storage of the original representation, but achieves comparable or even higher accuracy. Based on these contributions, we recommend a set of good practices for action recognition in videos for practitioners in this field.
Jianxin Wu 0001, Yu Zhang 0004, Weiyao Lin
IEEE Trans. Cybern.2
2016 Action Recognition in Still Images With Minimum Annotation Efforts
abstract
We focus on the problem of still image-based human action recognition, which essentially involves making prediction by analyzing human poses and their interaction with objects in the scene. Besides image-level action labels (e.g., riding, phoning), during both training and testing stages, existing works usually require additional input of human bounding boxes to facilitate the characterization of the underlying human-object interactions. We argue that this additional input requirement might severely discourage potential applications and is not very necessary. To this end, a systematic approach was developed in this paper to address this challenging problem of minimum annotation efforts, i.e., to perform recognition in the presence of only image-level action labels in the training stage. Experimental results on three benchmark data sets demonstrate that compared with the state-of-the-art methods that have privileged access to additional human bounding-box annotations, our approach achieves comparable or even superior recognition accuracy using only action annotations in training. Interestingly, as a by-product in many cases, our approach is able to segment out the precise regions of underlying human-object interactions.
Yu Zhang 0004, Li Cheng 0001, Jianxin Wu 0001, Jianfei Cai 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.1
2016 Weakly Supervised Fine-Grained Categorization With Part-Based Image Representation
abstract
In this paper, we propose a fine-grained image categorization system with easy deployment. We do not use any object/part annotation (weakly supervised) in the training or in the testing stage, but only class labels for training images. Fine-grained image categorization aims to classify objects with only subtle distinctions (e.g., two breeds of dogs that look alike). Most existing works heavily rely on object/part detectors to build the correspondence between object parts, which require accurate object or object part annotations at least for training images. The need for expensive object annotations prevents the wide usage of these methods. Instead, we propose to generate multi-scale part proposals from object proposals, select useful part proposals, and use them to compute a global image representation for categorization. This is specially designed for the weakly supervised fine-grained categorization task, because useful parts have been shown to play a critical role in existing annotation-dependent works, but accurate part detectors are hard to acquire. With the proposed image representation, we can further detect and visualize the key (most discriminative) parts in objects of different classes. In the experiments, the proposed weakly supervised method achieves comparable or better accuracy than the state-of-the-art weakly supervised methods and most existing annotation-dependent methods on three challenging datasets. Its success suggests that it is not always necessary to learn expensive object/part detectors in fine-grained image categorization.
Yu Zhang 0004, Xiu-Shen Wei, Jianxin Wu 0001, Jianfei Cai 0001, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.1
2016 Compact Representation of High-Dimensional Feature Vectors for Large-Scale Image Recognition and Retrieval
abstract
In large-scale visual recognition and image retrieval tasks, feature vectors, such as Fisher vector (FV) or the vector of locally aggregated descriptors (VLAD), have achieved state-of-the-art results. However, the combination of the large numbers of examples and high-dimensional vectors necessitates dimensionality reduction, in order to reduce its storage and CPU costs to a reasonable range. In spite of the popularity of various feature compression methods, this paper shows that the feature (dimension) selection is a better choice for high-dimensional FV/VLAD than the feature (dimension) compression methods, e.g., product quantization. We show that strong correlation among the feature dimensions in the FV and the VLAD may not exist, which renders feature selection a natural choice. We also show that, many dimensions in FV/VLAD are noise. Throwing them away using feature selection is better than compressing them and useful dimensions altogether using feature compression methods. To choose features, we propose an efficient importance sorting algorithm considering both the supervised and unsupervised cases, for visual recognition and image retrieval, respectively. Combining with the 1-bit quantization, feature selection has achieved both higher accuracy and less computational cost than feature compression methods, such as product quantization, on the FV and the VLAD image representations.
Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001
IEEE Trans. Image Process.1
2014 Towards Good Practices for Action Video Encoding
abstract
High dimensional representations such as VLAD or FV have shown excellent accuracy in action recognition. This paper shows that a proper encoding built upon VLAD can achieve further accuracy boost with only negligible computational cost. We empirically evaluated various VLAD improvement technologies to determine good practices in VLAD-based video encoding. Furthermore, we propose an interpretation that VLAD is a maximum entropy linear feature learning process. Combining this new perspective with observed VLAD data distribution properties, we propose a simple, lightweight, but powerful bimodal encoding method. Evaluated on 3 benchmark action recognition datasets (UCF101, HMDB51 and Youtube), the bimodal encoding improves VLAD by large margins in action recognition.
Jianxin Wu 0001, Yu Zhang 0004, Weiyao Lin
CVPR2
2014 Compact Representation for Image Classification: To Choose or to Compress?
abstract
In large scale image classification, features such as Fisher vector or VLAD have achieved state-of-the-art results. However, the combination of large number of examples and high dimensional vectors necessitates dimensionality reduction, in order to reduce its storage and CPU costs to a reasonable range. In spite of the popularity of various feature compression methods, this paper argues that feature selection is a better choice than feature compression. We show that strong multicollinearity among feature dimensions may not exist, which undermines feature compression's effectiveness and renders feature selection a natural choice. We also show that many dimensions are noise and throwing them away is helpful for classification. We propose a supervised mutual information (MI) based importance sorting algorithm to choose features. Combining with 1-bit quantization, MI feature selection has achieved both higher accuracy and less computational cost than feature compression methods such as product quantization and BPBC.
Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001
CVPR1
2014 Flexible Image Similarity Computation Using Hyper-Spatial Matching
abstract
Spatial pyramid matching (SPM) has been widely used to compute the similarity of two images in computer vision and image processing. While comparing images, SPM implicitly assumes that: in two images from the same category, similar objects will appear in similar locations. However, this is not always the case. In this paper, we propose hyper-spatial matching (HSM), a more flexible image similarity computing method, to alleviate the mis-matching problem in SPM. Besides the match between corresponding regions, HSM considers the relationship of all spatial pairs in two images, which includes more meaningful match than SPM. We propose two learning strategies to learn SVM models with the proposed HSM kernel in image classification, which are hundreds of times faster than a general purpose SVM solver applied to the HSM kernel (in both training and testing). We compare HSM and SPM on several challenging benchmarks, and show that HSM is better than SPM in describing image similarity.
Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001, Weiyao Lin
IEEE Trans. Image Process.1
2012 Exclusive Visual Descriptor Quantization
Yu Zhang 0004, Jianxin Wu 0001, Weiyao Lin
ACCV (1)1
2008 Improving Videophone Transmission over Multi-Rate IEEE 802.11e Networks
abstract
In this paper, we propose an adaptive system for improving videophone transmission over EDCA. We consider that a videophone contains a constant bit rate (CBR) voice source and a rate-adaptive video source. Two issues are addressed in this research. Firstly, how to solve the AP bottleneck problem, and secondly, how to adjust video source rate to improve the network performance. For the first issue, we propose the adjustment of the transmission opportunity (TXOP) to give AP a higher priority in voice transmission in order to eliminate the AC3 transmission bottleneck at the AP. For the second issue, our principle is to guarantee the throughput of voice traffic while transmitting as much video traffic as possible. Moreover, we consider more realistic multi-rate WLANs, where multiple transmission rates are used in the PHY layer depending on the underlaying channel conditions.
Jianfei Cai 0001, Chuan Heng Foh, Yu Zhang 0004
ICC4
2008 An On-Off Queue Control Mechanism for Scalable Video Streaming over the IEEE 802.11e WLAN
abstract
In this paper, we study the issue of scalable video streaming over IEEE 802.11e EDCA WLANs. Our basic idea is to control the number of "active" nodes on the channel in order to reduce collisions under heavy traffic conditions. Specifically, we propose a distributed on-off queue control (OOQC) mechanism, which is designed to maintain high network throughput while keeping packet loss due to collision as low as possible. A low priority early drop (LPED) method is also employed to drop the packets at the queue according to packet relative priority index (RPI) provided by scalable video coding. Simulation results show that our proposed OOQC scheme significantly outperforms EDCA in received video quality.
Yu Zhang 0004, Chuan Heng Foh, Jianfei Cai 0001
ICC1
2007 Scalable Video Transmission over the IEEE 802.11e Networks Using Cross-Layer Rate Control
abstract
This work presents a novel cross-layer rate control scheme for optimizing 3D wavelet scalable video transmission over the IEEE 802.11e wireless local area networks. The proposed scheme consists of a macro and a micro rate control schemes residing at the application layer and the network sublayer respectively. The macro rate control uses bandwidth estimation to achieve optimal bit allocation with minimum distortion. The micro rate control employs an adaptive mapping of packets using video classifications. This prioritizes appropriately the video traffic to maximize the transmission protection to the important video packets. The performance is investigated by simulations showing advantages of our cross-layer design.
Chuan Heng Foh, Yu Zhang 0004, Zefeng Ni, Jianfei Cai 0001
ICC2
2007 Optimized Cross-Layer Design for Scalable Video Transmission Over the IEEE 802.11e Networks
abstract
A cross-layer design for optimizing 3-D wavelet scalable video transmission over the IEEE 802.11e networks is proposed. A thorough study on the behavior of the IEEE 802.11e protocol is conducted. Based on our findings, all timescales rate control is developed featuring a unique property of soft capacity support for multimedia delivery. The design consists of a macro timescale and a micro timescale rate control schemes residing at the application layer and the network sublayer respectively. The macro rate control uses bandwidth estimation to achieve optimal bit allocation with minimum distortion. The micro rate control employs an adaptive mapping of packets from video classifications to appropriate network priorities which preemptively drops less important video packets to maximize the transmission protection to the important video packets. The performance is investigated by simulations highlighting advantages of our cross-layer design.
Chuan Heng Foh, Yu Zhang 0004, Zefeng Ni, Jianfei Cai 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2