Jian Zhang 0002

dblp:07/314-2 · DBLP profile ↗
← Back
225ranked-venue papers
5as first author
71since 2021 · last 2026
0000-0002-7240-3541ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 157 · 3 first-author · 46 since 2021Artificial intelligence and machine learning · 81 · 1 first-author · 29 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Computer networks · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer
abstract
Single-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%).
Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002
AAAI10
2025 Fully-Geometric Cross-Attention for Point Cloud Registration
abstract
Point cloud registration approaches often fail when the overlap between point clouds is low due to noisy point correspondences. This work introduces a novel cross-attention mechanism tailored for Transformer-based architectures that tackles this problem, by fusing information from coordinates and features at the super-point level between point clouds. This formulation has remained unexplored primarily because it must guarantee rotation and translation invariance since point clouds reside in different and independent reference frames. We integrate the Gromov-Wasserstein distance into the cross-attention formulation to jointly compute distances between points across different point clouds and account for their geometric structure. By doing so, points from two distinct point clouds can attend to each other under arbitrary rigid transformations. At the point level, we also devise a self-attention mechanism that aggregates the local geometric structure information into point features for fine matching. Our formulation boosts the number of inlier correspondences, thereby yielding more precise registration results compared to state-of-the-art approaches. We have conducted an extensive evaluation on 3DMatch, 3DLoMatch, KITTI, and 3DCSR datasets. Project page: https://github.com/twowwj/FLAT.
Weijie Wang 0002, Guofeng Mei, Jian Zhang 0002, Nicu Sebe, Bruno Lepri, Fabio Poiesi
3DV3
2025 LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval
abstract
Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of expert annotation.While large language models (LLMs) offer a promising solution by enriching textual descriptions, their outputs frequently suffer from hallucinations or miss visually grounded details.To address these challenges, we propose C 3 , a data augmentation framework that enhances cross-modal retrieval performance by improving the completeness and consistency of LLM-generated descriptions.C 3 introduces a completeness evaluation module to assess semantic coverage using both visual cues and language-model outputs.Furthermore, to mitigate factual inconsistencies, we formulate a Markov Decision Process to supervise Chain-of-Thought reasoning, guiding consistency evaluation through adaptive query control.Experiments on the cultural heritage datasets CulTi and TimeTravel, as well as on general benchmarks MSCOCO and Flickr30K, demonstrate that C 3 achieves state-of-the-art performance in both fine-tuned and zero-shot settings.The code of this paper is available at https://github.com/JianZhang24/C-3.
Jian Zhang 0002, Junyi Guo, Junyi Yuan, Huanda Lu, Fangyu Wu 0001, Dongming Lu
EMNLP1
2025 PointGAC: Geometric-Aware Codebook for Masked Point Cloud Modeling
abstract
Most masked point cloud modeling (MPM) methods follow a regression paradigm to reconstruct the coordinate or feature of masked regions. However, they tend to over-constrain the model to learn the details of the masked region, resulting in failure to capture generalized features. To address this limitation, we propose \textbf{\textit{PointGAC}}, a novel clustering-based MPM method that aims to align the feature distribution of masked regions. Specially, it features an online codebook-guided teacher-student framework. Firstly, it presents a geometry-aware partitioning strategy to extract initial patches. Then, the teacher model updates a codebook via online k-means based on features extracted from the complete patches. This procedure facilitates codebook vectors to become cluster centers. Afterward, we assigns the unmasked features to their corresponding cluster centers, and the student model aligns the assignment for the reconstructed masked features. This strategy focuses on identifying the cluster centers to which the masked features belong, enabling the model to learn more generalized feature representations. Benefiting from a proposed codebook maintenance mechanism, codebook vectors are actively updated, which further increases the efficiency of semantic feature learning. Experiments validate the effectiveness of the proposed method on various downstream tasks. Code is available at https://github.com/LAB123-tech/PointGAC
Abiao Li, Chenlei Lv, Yuming Fang 0001, Yifan Zuo 0001, Jian Zhang 0002, Guofeng Mei
ICCV5
2025 Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
abstract
Part-level features are crucial for image understanding, but few studies focus on them because of the lack of fine-grained labels. Although unsupervised part discovery can eliminate the reliance on labels, most of them cannot maintain robustness across various categories and scenarios, which restricts their application range. To overcome this limitation, we present a more effective paradigm for unsupervised part discovery, named Masked Part Autoencoder (MPAE). It first learns part descriptors as well as a feature map from the inputs and produces patch features from a masked version of the original images. Then, the masked regions are filled with the learned part descriptors based on the similarity between the local features and descriptors. By restoring these masked patches using the part descriptors, they become better aligned with their part shapes, guided by appearance features from unmasked patches. Finally, MPAE robustly discovers meaningful parts that closely match the actual object shapes, even in complex scenarios. Moreover, several looser yet more effective constraints are proposed to enable MPAE to identify the presence of parts across various scenarios and categories in an unsupervised manner. This provides the foundation for addressing challenges posed by occlusion and for exploring part similarity across multiple categories. Extensive experiments demonstrate that our method robustly discovers meaningful parts across various categories and scenarios. The code is available at the project https://github.com/Jiahao-UTS/MPAE.
Jiahao Xia 0001, Yike Wu 0001, Wenjian Huang 0001, Jianguo Zhang 0001, Jian Zhang 0002
ICCV5
2025 Towards Cross-Modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
Junyi Yuan, Jian Zhang 0002, Fangyu Wu 0001, Huanda Lu, Dongming Lu, Qiufeng Wang 0001
ICDAR (4)2
2025 Enhancing origin-destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution
Piao Yu, Xu Zhang 0039, Yongshun Gong, Jian Zhang 0002, Haoliang Sun, Junjie Zhang 0002, Xinxin Zhang 0004, Yilong Yin
Expert Syst. Appl.4
2025 HFA-UNet: hybrid and full attention UNet for thyroid nodule segmentation
abstract
Ultrasound imaging is the most commonly used method for screening thyroid nodules due to its low cost and non-invasive nature. Thyroid nodule lesions have variable shapes, rich aspect ratios, unclear boundaries, calcified nodule-induced acoustic shadows, and noise interference, causing challenges in accurate segmentation. Recent methods ignore various scale features and details in different resolutions of images, leading to redundant or missing feature information and then affecting the segmentation performance. In this paper, we introduce a hybrid and full attention UNet model for ultrasound thyroid nodule segmentation. Self, spatial and channel attention are combined in a U-Net-like structure to extract global and local features simultaneously. A novel full attention multi-scale fusion stage is designed to enhance boundary features while suppressing noise features. At the same time, the model dynamically adjusts the number of skip connections corresponding to images of different resolutions to better utilize multi-scale features and detailed information. We evaluate our model on DDTI, TN3K and Stanford Cine-Clip datasets, including internal validation and cross-dataset testing. The results show that our proposed model for internal validation in the DDTI dataset increases the Dice score and mean intersection over union by 2.36 % and 1.04 % compared to the state-of-the-art model. In the TN3K dataset, they increase by 1.66 % and 3.05 %.
Yuanhao Zou, Xiangjian He, Qing Xu 0014, Ming Liu 0021, Shengji Jin, Qian Zhang 0018, Maggie M. He, Jian Zhang 0002
Knowl. Based Syst.9
2025 Fine-grained visual tracking via distribution-aware mask modeling and temporal propagation
Junjie Zhang 0002, Hongwen Yu, Fangyu Wu 0001, Xiaoshui Huang, Jian Zhang 0002
Knowl. Based Syst.6
2025 Consistent Image Inpainting With Pre-Perception and Cross-Perception Collaborative Processes
abstract
It has been proven that introducing multiple guidance sources boosts image inpainting performance. However, existing methods primarily focus on local relationships and neglect the holistic interplay between guidance and texture information. Moreover, they lack an effective feedback mechanism to adaptively update the guidance process as corrupted texture information is progressively restored, potentially resulting in inconsistent inpainting. To tackle this issue, we propose a novel scheme aligned with pre-perception and cross-perception collaborative processes in human drawing. To mimic the pre-perception process, we introduce a pre-perceptual transformer block that captures long-range contextual dependencies and activates meaningful information to individually optimize image structures, semantic layouts, and textures, thereby effectively controlling their respective generation. To mimic the cross-perception collaborative process, we propose a cyclic cross-perceptual interaction to maintain consistency across the entire image regarding structure, layout, and texture while progressively refining their details. This interaction accounts for the global attention relationship between texture and other guidance sources (including image structure and semantic layout) to enhance image texture, alongside integrating a dedicated feedback mechanism to update guidance information. The proposed components are alternately deployed in three-branch decoders of the new scheme from rough to fine-grained levels to achieve these two iterative processes of human drawing. Experimental results prove the superiority of the proposed scheme over state-of-the-art methods across three datasets.
Yongle Zhang 0001, Yimin Liu 0001, Hao Fan 0004, Ruotong Hu, Jian Zhang 0002, Qiang Wu 0001
IEEE Trans. Image Process.5
2025 DSENet++: A Coarse-to-Fine Framework for Enhanced Sub-Region Detection in Aerial Images
Xiangjie Wang, Liang Chen 0004, Junjie Zhang 0002, Jian Zhang 0002, Shiming Ge, Dan Zeng 0001
IEEE Trans. Multim.5
2025 Hierarchical Multi-Prototype Discrimination: Boosting Support-Query Matching for Few-Shot Segmentation
abstract
Few-shot segmentation (FSS) aims at training a model on base classes with sufficient annotations and then tasking the model with predicting a binary mask to identify novel class pixels with limited labeled images. Mainstream FSS methods adopt a support-query matching paradigm that activates target regions of the query image according to their similarity with a single support class prototype. However, this prototype vector is inclined to overfit the support images, leading to potential under-matching in latent query object regions and incorrect mismatches with base class features in the query image. To address these issues, this study reformulates conventional single foreground prototype matching to a multi-prototype matching paradigm. In this paradigm, query features exhibiting high confidence with non-target prototypes will be categorized as background. Specifically, the target query features are drawn closer to the novel class prototype through a Masked Cross-Image Encoding (MCE) module and a Semantic Multi-prototype Matching (SMM) module is incorporated to collaboratively filter unexpected base class regions on multi-scale features. Furthermore, we devise an adaptive class activation map, termed target-aware class activation map (TCAM) to preserve semantically coherent regions that might be inadvertently suppressed under pixel-wise matching guidance. Experimental results on PASCAL-5$^{i}$and COCO-20$^{i}$datasets demonstrate the advantage of the proposed novel modules, with the holistic approach outperforming compared state-of-the-art methods.
Wenbo Xu 0004, Huaxi Huang, Yongshun Gong, Litao Yu, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.6
2024 Label-Efficient Few-Shot Semantic Segmentation with Unsupervised Meta-Training
abstract
The goal of this paper is to alleviate the training cost for few-shot semantic segmentation (FSS) models. Despite that FSS in nature improves model generalization to new concepts using only a handful of test exemplars, it relies on strong supervision from a considerable amount of labeled training data for base classes. However, collecting pixel-level annotations is notoriously expensive and time-consuming, and small-scale training datasets convey low information density that limits test-time generalization. To resolve the issue, we take a pioneering step towards label-efficient training of FSS models from fully unlabeled training data, or additionally a few labeled samples to enhance the performance. This motivates an approach based on a novel unsupervised meta-training paradigm. In particular, the approach first distills pre-trained unsupervised pixel embedding into compact semantic clusters from which a massive number of pseudo meta-tasks is constructed. To mitigate the noise in the pseudo meta-tasks, we further advocate a robust Transformer-based FSS model with a novel prototype-based cross-attention design. Extensive experiments have been conducted on two standard benchmarks, i.e., PASCAL-5i and COCO-20i, and the results show that our method produces impressive performance without any annotations, and is comparable to fully supervised competitors even using only 20% of the annotations. Our code is available at: https://github.com/SSSKYue/UMTFSS.
Jianwu Li, Kaiyue Shi, Guosen Xie, Xiaofeng Liu 0006, Jian Zhang 0002, Tianfei Zhou
AAAI5
2024 DSENet: An Object-Wise Density-Informed Coarse-to-Fine Object Detector for Aerial Image
abstract
Object detection in aerial images remains formidable due to substantial object scale variations, and uneven object distributions. Previous methods widely adopt the coarse-to-fine methodology where detectors focus on large-scale objects coarsely. Sub-regions that contain densely distributed small ones are captured and detected finely. However, two pivotal assessment factors of sub-regions, positional precision, and detection difficulty, deserve further consideration. In this paper, we propose an object-wise density-informed DSENet including consecutive stages termed "Discernment, Selection, Elevation ". Specifically, the sophisticated object-wise density map that considers both object scales and angles, helps discern more positional-precise sub-regions. Then sub-regions with high detection difficulty are selected based on density intensities and coarse detections collaboratively. Finally, the fine detector head instead of the full detector, fine-tuned with selected sub-regions efficiently, elevates what and where coarse detections are mediocre. Extensive experiments show that DSENet achieves state-of-the-art performance on two popular aerial image datasets, VisDrone and DOTA-V1.5.
Xiangjie Wang, Junjie Zhang 0002, Jian Zhang 0002, Dan Zeng 0001
ICME4
2024 Densely Connected Transformer with Frequency Awareness and Sam Guidance for Semi-Supervised Hyperspectral Image Classification
abstract
Advancements in Hyperspectral Image (HSI) spatial resolution pose challenges in pixel-wise classification. Semi-supervised self-training shows potential by using pseudo-labels from unlabeled samples. However, the Hughes phenomenon and environmental factors often lead to spectral variability and undermine pseudo-label credibility. To address above issues, we propose a densely connected Transformer leveraging Discrete Wavelet Transform for extracting nuanced spatial-spectral features and redundancy removal, and we design a filtering strategy guided by the Segment Anything Model (SAM) to retain reliable pseudo labeled samples given the spatial and semantic consistency of HSI regions. Experiments show promising performance of proposed model on high-resolution HSIs compared to trending methods under limited supervision.
Yutao Rao, Liwei Sun, Junjie Zhang 0002, Jian Zhang 0002, Dan Zeng 0001
ICME5
2024 GSTran: Joint Geometric and Semantic Coherence for Point Cloud Segmentation
Abiao Li, Chenlei Lv, Guofeng Mei, Yifan Zuo 0001, Jian Zhang 0002, Yuming Fang 0001
ICPR (18)5
2024 Task Consistent Prototype Learning for Incremental Few-Shot Semantic Segmentation
Wenbo Xu 0004, Yang Wang 0002, Qiang Wu 0001, Jian Zhang 0002
ICPR (23)6
2024 Unsupervised Point Cloud Representation Learning by Clustering and Neural Rendering
abstract
Abstract Data augmentation has contributed to the rapid advancement of unsupervised learning on 3D point clouds. However, we argue that data augmentation is not ideal, as it requires a careful application-dependent selection of the types of augmentations to be performed, thus potentially biasing the information learned by the network during self-training. Moreover, several unsupervised methods only focus on uni-modal information, thus potentially introducing challenges in the case of sparse and textureless point clouds. To address these issues, we propose an augmentation-free unsupervised approach for point clouds, named CluRender, to learn transferable point-level features by leveraging uni-modal information for soft clustering and cross-modal information for neural rendering. Soft clustering enables self-training through a pseudo-label prediction task, where the affiliation of points to their clusters is used as a proxy under the constraint that these pseudo-labels divide the point cloud into approximate equal partitions. This allows us to formulate a clustering loss to minimize the standard cross-entropy between pseudo and predicted labels. Neural rendering generates photorealistic renderings from various viewpoints to transfer photometric cues from 2D images to the features. The consistency between rendered and real images is then measured to form a fitting loss, combined with the cross-entropy loss to self-train networks. Experiments on downstream applications, including 3D object detection, semantic segmentation, classification, part segmentation, and few-shot learning, demonstrate the effectiveness of our framework in outperforming state-of-the-art techniques.
Guofeng Mei, Cristiano Saltori, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001, Jian Zhang 0002, Fabio Poiesi
Int. J. Comput. Vis.6
2024 DCTracker: Rethinking MOT in soccer events under dual views via cascade association
Long Hu, Junjie Zhang 0002, Weiyi Lv, Yongshun Gong, Jingya Wang 0001, Jian Zhang 0002, Dan Zeng 0001
Knowl. Based Syst.6
2024 Weakly Supervised Semantic Segmentation With Consistency-Constrained Multiclass Attention for Remote Sensing Scenes
abstract
Obtaining image-level class labels for Remote Sensing (RS) images is a relatively straightforward process, sparking significant interest in Weakly Supervised Semantic Segmentation (WSSS). However, RS images present challenges beyond those encountered in generic WSSS, including complex backgrounds, densely distributed small objects, and considerable scale variations. To address above issues, we introduce a COnsistency-COnstrained Multi-Class Attention model, noted asCocoaNet. Specifically, CocoaNet endeavors to capture both semantic correlation and class distinctiveness using a Global-Local Adaptive Attention mechanism, which integrates the self-attention to model global correlation, complemented by a Local Perception branch that intensifies focus on local regions. The resulting class-specific attention weights and patch-level pairwise affinity weights are employed to optimize the initial CAMs. This mechanism proves highly effective in mitigating inter-class interference and managing the distribution of densely clustered small objects. Moreover, we invoke a Consistency Constraint to rectify activation inaccuracy. By utilizing a Siamese structure for the mutual supervision of features extracted from images at different scales, we address substantial scale variations in RS scenes. Simultaneously, a Class Contrast Loss is adopted to enhance the discriminativeness of class-specific features. Departing from the conventional CAM optimization, which is rather complex and time-consuming, we harness the prior knowledge from generic Segment Anything model to design a joint optimization strategy that refines target boundaries and further promotes discriminative visual features. We validate the effectiveness of our proposed approach on three benchmark datasets in multi-class RS scenarios, experimental results demonstrate that our model yield promising advancements compared to state-of-the-art methods.
Junjie Zhang 0002, Yongshun Gong, Jian Zhang 0002, Liang Chen 0004, Dan Zeng 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 PCL: Point Contrast and Labeling for Weakly Supervised Point Cloud Semantic Segmentation
abstract
Point cloud semantic segmentation is a fundamental task in 3D scene understanding and has recently achieved remarkable progress. The success of existing approaches is attributed to recent advanced deep networks for point clouds and the availability of a large amount of labeled training data. However, creating such fully annotated training datasets for supervised point cloud semantic segmentation methods is a time-consuming and labor-intensive process, which increases the difficulty of extending supervised approaches to new application scenarios. To alleviate the data-hungry nature of deep learning, we propose PCL, the point contrast and labeling framework for weakly supervised point cloud semantic segmentation with small percentages of point-level annotations. The core idea of this method is to exploit contrastive learning to help learn a larger number of discriminative feature representations with limited annotations. By introducing two types of contrastive relationships, cross-sample point contrast and low-level similarity-based point contrast, our proposed framework can directly regularize the learned feature space, considering not only the low-level similarity within each point cloud but also the discriminative semantics within and across point clouds on both labeled and unlabeled points via pseudo labels. In addition, we propose a pseudo label refinery module to generate robust and reliable pseudo labels online, reducing the negative impact of incorrect pseudo labels. Our method achieves state-of-the-art performance on a diverse set of label-efficient semantic segmentation tasks.
Anan Du, Tianfei Zhou, Shuchao Pang, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.5
2024 Focusing on Subtle Differences: A Feature Disentanglement Model for Series Photo Selection
abstract
Nowadays, capturing cherished moments results in an abundance of photos, which necessitates the selection of the finest one from a pool of akin images—a process both intricate and time-intensive. Thus, series photo selection (SPS) techniques have been developed to recommend the optimal moment from nearly identical photos through the use of aesthetic quality assessment. However, addressing SPS proves demanding due to the subtle nuances within such imagery. Existing approaches predominantly rely on diverse feature types (e.g., color, layout, generic features) extracted from original images to discern the qualified shot, yet they disregard disentangling generality and specificity at the feature level. This study aims to detect subtle aesthetic distinctions among akin photos. We propose a feature separation model that captures all label-relevant information through an encoder. We introduce Information Bottleneck (IB) learning to obtain non-redundant representations of image pairs and filter out noise information from the representations. Our model segregates image features into shared and specific attributes by employing feature constraints to boost mutual information across images and guide meaningful information within individual images. This process filters out extraneous data within individual images, thus significantly enhancing the representation of similar image pairs. Extensive experiments on the Phototriage dataset show that our model can accentuate subtle disparities and achieve superior results when compared to alternative methods.
Yongshun Gong, Xinxin Zhang 0004, Jian Zhang 0002, Yilong Yin
IEEE Trans. Multim.5
2024 Modeling Multiple Aesthetic Views for Series Photo Selection
abstract
Numerous photos are taken in daily life, and sorting them is laborious and time consuming. The large number of similar images exacerbates the difficulty of album management, under this scenario, serial photo selection (SPS) emerges. As an important branch of image aesthetic quality assessment, it focuses on identifying the best image among a series of almost identical photos. Currently, most existing SPS methods focus only on extracting features from the original image, while neglecting the fact that multiple views of the image can provide much more detailed aesthetic information. In this article, we propose a Siamese network structure called SPSNet to enhance the representation learning of multi-view features by acquiring the depth, generic, and handcrafted features of images. In specific, we implement a parallel structure to extract deep and shallow features, fusing local and global representations at different resolutions interactively. The aggregation of multiple views of image via a self-attentive module with adaptive weights enables the model to discriminate the importance of each view. Moreover, we employ a graph neural network to construct the relationships among the multi-view features. Our proposed method, which is trained by a Siamese network, can effectively distinguish the nuances of similar images, and thus, select the best one from a series of almost identical photos. Extensive experiments conducted on the aesthetic dataset demonstrate that our method outperforms other state-of-the-art SPS methods, which achieves the 75.36% accuracy on the Phototriage dataset. Besides, our model is up to 3.04% better than the baseline methods in terms of the average accuracy.
Yongshun Gong, Lu Zhang 0062, Jian Zhang 0002, Liqiang Nie, Yilong Yin
IEEE Trans. Multim.4
2024 Mutual Dual-Task Generator With Adaptive Attention Fusion for Image Inpainting
abstract
Image segmentation can reveal the semantic structure information in an image, which is helpful guidance information for image inpainting. Notably, it can help mitigate the artifacts on the boundaries of different semantic regions during the inpainting process. Existing semantic guidance-based image inpainting provides one-way guidance from the semantic segmentation task to the image inpainting task. There is no feedback from the inpainting results to adjust the guidance process, which causes inferior performance. To tackle this issue, this work proposes mutual dual-task generators to establish the interaction between image segmentation and image inpainting tasks. Thus, semantic segmentation guides image inpainting and also receives feedback from image inpainting. These two processes interact with each other and progressively improve the inpainting quality. The mutual dual-task generator consists of a shared encoder and mutual decoders with the bidirectional Cross-domain Feature DeNormalization (CFDN) module inside, which hierarchically models the Segmentation-guided image Texture (ST) generation and Texture-guided semantic Segmentation (TS) generation. At the end of mutual decoders, an Adaptive Attention Fusion (AAF) module is proposed to augment the texture and semantic class affinity between pixels, further refining the inpainted results. Experimental results demonstrate that the proposed mutual dual-task generator pipeline achieves superior inpainting performances over the state of the arts on three public datasets.
Yongle Zhang 0001, Yimin Liu 0001, Ruotong Hu, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.5
2023 Unsupervised Deep Probabilistic Approach for Partial Point Cloud Registration
abstract
Deep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior probability distributions of Gaussian mixture models (GMMs) from point clouds. To handle partial point cloud registration, we apply the Sinkhorn algorithm to predict the distribution-level correspondences under the constraint of the mixing weights of GMMs. To enable unsupervised learning, we design three distribution consistency-based losses: self-consistency, cross-consistency, and local contrastive. The self-consistency loss is formulated by encouraging GMMs in Euclidean and feature spaces to share identical posterior distributions. The cross-consistency loss derives from the fact that the points of two partially overlapping point clouds belonging to the same clusters share the cluster centroids. The cross-consistency loss allows the network to flexibly learn a transformation-invariant posterior distribution of two aligned point clouds. The local contrastive loss facilitates the network to extract discriminative local features. Our UDPReg achieves competitive performance on the 3DMatch/3DLoMatch and ModelNet/ModelLoNet benchmarks.
Guofeng Mei, Hao Tang 0005, Xiaoshui Huang, Weijie Wang 0002, Juan Liu 0006, Jian Zhang 0002, Luc Van Gool, Qiang Wu 0001
CVPR6
2023 Camera Proxy based Contrastive Learning with Hard Sampling for Unsupervised Person Re-identification
abstract
Because of the advantages of dealing with large-scale unlabelled data, unsupervised learning has recently attracted more attention for person re-identification. Particularly, the combination of the unsupervised learning paradigm with contrastive learning shows promising efficiency in network optimization. This work adopts the successful camera-aware contrastive learning approach and further explores its capability on the camera proxy level to improve the data pair consistency. Thus, it is more robust to the camera change, which still challenges the unsupervised person re-identification. This work proposed a Camera Proxy-based Contrastive Learning framework, which explicitly considers inter-camera scenario and intra-camera scenario. Moreover, this work is motivated by the strategy of selecting a hard negative sample in triplet loss learning and further extends it to contrastive learning for both negative and positive pair creation on the camera proxy level. Extensive experiments demonstrate the superiority of the proposed framework over state-of-the-art approaches on purely unsupervised re-identification.
Yimin Liu 0001, Meibin Qi, Qiang Wu 0001, Yanfang Yang, Xiaohong Li 0002, Jian Zhang 0002
ICME6
2023 Masked Cross-image Encoding for Few-shot Segmentation
abstract
Few-shot segmentation (FSS) is a dense prediction task that aims to infer the pixel-wise labels of unseen classes using only a limited number of annotated images. The key challenge in FSS is to classify the labels of query pixels using class prototypes learned from the few labeled support exemplars. Prior approaches to FSS have typically focused on learning class-wise descriptors independently from support images, thereby ignoring the rich contextual information and mutual dependencies among support-query features. To address this limitation, we propose a joint learning method termed Masked Cross-Image Encoding (MCE), which is designed to capture common visual properties that describe object details and to learn bidirectional inter-image dependencies that enhance feature interaction. MCE is more than a visual representation enrichment module; it also considers cross-image mutual dependencies and implicit guidance. Experiments on FSS benchmarks PASCAL-5iand COCO-20idemonstrate the advanced meta-learning ability of the proposed method.
Wenbo Xu 0004, Huaxi Huang, Litao Yu, Qiang Wu 0001, Jian Zhang 0002
ICME6
2023 An Intrusion Detection Method Based on Hash Function for Industrial Cloud Data
abstract
With industrial control systems (ICSs) commonly connected to the cloud, the security of ICS has received widespread attention. Intrusion detection systems (IDSs) are widely employed to protect ICSs, yet most existing intrusion detection models require expert systems to select the important features and large amount of storage space to store data, both result in increased costs. This paper proposes a preprocessing approach based on hash functions, which does not require prior knowledge and saves a lot of storage space. First, we compute a hash code for each piece of original data with the hash function. Next, the hash codes are converted to decimal and normalised. Finally, a number between 0 and 1 is obtained as a feature for intrusion detection. In experiments, we combine our method with various machine learning algorithms, and extensive experimental results illustrate that our method combined with support vector machine (SVM) and k-nearest neighbor (K-NN) achieve good detection results.
Yinchu Wang, Heng Zhang 0001, Hongran Li, Jian Zhang 0002
ICPADS5
2023 Automated Flock Density and Activity Recognition for Welfare Monitoring on Commercial Egg Farms
abstract
Monitoring poultry behaviour provides the opportunity to aid egg production and animal welfare. With the current development in machine learning and computer vision, automated content analysis has become a practical way for low-cost and continuous monitoring of animal behaviours. In this demo, we will show a simple yet effective flock monitoring system based on computer vision and machine learning techniques for egg farmers that allows them to reduce labour yet improve performance. This demo shows that it is possible to auto-analyse flock activities thereby providing early warning of welfare issues, by applying object detection, tracking and crowd-counting techniques. Summaries of individual bird activity and their distribution are closely related to the flock behaviour, which in turn reflects the welfare status. Specifically, the density and movement patterns of birds provide reliable information on the welfare status of the flock. For example, the real-time monitoring of density and movement can give early warnings of pile-ups. To observe these and other important flock activities, we developed a low-cost and easy-use system based on recent computer vision techniques to auto-estimate the density and movement of birds on commercial egg farms.
Litao Yu, Wenbo Xu 0004, Qiang Wu 0001, Jian Zhang 0002
MMSP4
2023 Overlap-guided Gaussian Mixture Models for Point Cloud Registration
abstract
Probabilistic 3D point cloud registration methods have shown competitive performance in overcoming noise, outliers, and density variations. However, registering point cloud pairs in the case of partial overlap is still a challenge. This paper proposes a novel overlap-guided probabilistic registration approach that computes the optimal transformation from matched Gaussian Mixture Model (GMM) parameters. We reformulate the registration problem as the problem of aligning two Gaussian mixtures such that a statistical discrepancy measure between the two corresponding mixtures is minimized. We introduce a Transformer-based detection module to detect overlapping regions, and represent the input point clouds using GMMs by guiding their alignment through overlap scores computed by this detection module. Experiments show that our method achieves superior registration accuracy and efficiency than state-of-the-art methods when handling point clouds with partial overlap and different densities on synthetic and real-world datasets. https://github.com/gfmei/ogmm
Guofeng Mei, Fabio Poiesi, Cristiano Saltori, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe
WACV4
2023 Cross-source point cloud registration: Challenges, progress and prospects
Xiaoshui Huang, Guofeng Mei, Jian Zhang 0002
Neurocomputing3
2023 Latent evolution model for change point detection in time-varying networks
Yongshun Gong, Jian Zhang 0002, Meng Chen 0003
Inf. Sci.3
2023 Missing Value Imputation for Multi-View Urban Statistical Data via Spatial Correlation Learning
abstract
As a developing trend of urbanization, massive amounts of urban statistical data with multiple views (e.g., views of Population and Economy) are increasingly collected and benefited to diverse domains, including transportation service, regional analysis, etc. Unfortunately, these statistical data that are divided into fine-grained regions usually suffer from missing value problem during the acquisition and storage processes. It is mianly caused by some inevitable circumstances, e.g., the document defacement, statistical difficulty in remote districts, and inaccurate information cleaning, etc. Those missing entries which make valuable information invisible may distort the further urban analysis. To improve the quality of missing data imputation, we propose an improved spatial multi-kernel learning method to guide the imputation process incorporating with the adaptive-weight non-negative matrix factorization strategy. Our model takes into account the regional latent similarities and the real geographical positions as well as the correlations among various views that are able to complete missing values precisely. We conduct intensive experiments to evaluate our method and compare with other state-of-the-art approaches on real-world datasets. All the empirical results show that the proposed model outperforms all the other state-of-the-art methods. Additionally, our model represents a strong generalization ability across multiple cities.
Yongshun Gong, Zhibin Li 0002, Jian Zhang 0002, Wei Liu 0007, Yilong Yin, Yu Zheng 0004
IEEE Trans. Knowl. Data Eng.3
2023 Boosting Robust Learning Via Leveraging Reusable Samples in Noisy Web Data
abstract
Webly-supervised fine-grained visual classification (FGVC) has attracted increasing attention in recent years because of the unaffordable cost of obtaining correctly-labeled large-scale fine-grained datasets. However, due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause accumulated errors. Sample selection methods identify clean (“easy”) samples based on the fact that small losses can alleviate the accumulated errors. However, “hard” and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the network. Furthermore, in order to endow our model with the capability to capture richer and more discriminative feature representations, we propose a cross-layer attention-based feature refinement (CLAR) block. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jian Zhang 0002, Xian-Sheng Hua 0001
IEEE Trans. Multim.5
2023 Improving Disentangled Representation Learning for Gait Recognition Using Group Supervision
abstract
In decades, gait has been gathering extensive interest for the advantage that it can be measured from a distance without physical contact. However, for image/video-based gait recognition, its performance can be remarkably influenced by exterior factors, such as viewing angles and clothing changes. Thus, in this paper, a group-supervised disentangled representation learning network is proposed for gait recognition to extract features invariant to these factors. First, sequences are explicitly disentangled into pose, gait, appearance, and view features through a generic encoder-decoder framework. To ensure the feature adaptability and independency, a disentanglement swap module is specifically adopted during our encode-decoder process through a series of swap operations based on the feature attributes. Following the feature disentanglement, a disentanglement aggregation module is also specially proposed for pose, gait, and appearance features to enhance their effectiveness. Finally, the enhanced three features are concatenated together for gait recognition. Relevant experiments certify that compared with other disentangled representation learning-based gait recognition methods, our proposed method enables to obtain a more excellent recognition result, despite fewer gait frames being utilized.
Lingxiang Yao, Worapan Kusakunniran, Peng Zhang 0057, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.5
2022 Data Augmentation-free Unsupervised Learning for 3D Point Cloud Understanding
Guofeng Mei, Cristiano Saltori, Fabio Poiesi, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001
BMVC4
2022 Unsupervised Point Cloud Pre-Training Via Contrasting and Clustering
abstract
The annotation for large-scale point clouds is still time-consuming and unavailable for many complex real-world tasks. Point cloud pre-training is a promising direction to auto-extract features without labeled data. Therefore, this paper proposes a general unsupervised approach, named ConClu for point cloud pre-training by jointly performing contrasting and clustering. Specifically, the contrasting is formulated by maximizing the similarity feature vectors produced by encoders fed with two augmentations of the same point cloud. The clustering simultaneously clusters the data while enforcing consistency between cluster assignments produced different augmentations. Experimental evaluations on downstream applications outperform state-of-the-art techniques, which demonstrates the effectiveness of our framework.
Guofeng Mei, Xiaoshui Huang, Juan Liu 0006, Jian Zhang 0002, Qiang Wu 0001
ICIP4
2022 Partial Point Cloud Registration Via Soft Segmentation
abstract
Most existing correspondence-free registration methods suffer from performance degradation in partial overlapped point clouds. To solve the partial overlapped point cloud registration, this paper proposes, SegReg, a soft Segmentation-based correspondence-free Registration approach. Specifically, we first softly segment both source and target point clouds into a discrete number of geometric partitions, respectively. Then registration is achieved through iteratively using the IC-LK algorithm to minimize the distance between the feature descriptors of the corresponded partitions. Extensive experiments on synthetic synthetic dataset ModelNet40 and real dataset 7Scene show that the proposed method achieves state-of-the-art performance.
Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001
ICIP3
2022 Series Photo Selection via Multi-View Graph Learning
abstract
Series photo selection (SPS) is an important branch of the image aesthetics quality assessment, which focuses on finding the best one from a series of nearly identical photos. While a great progress has been observed, most of the existing SPS approaches concentrate solely on extracting features from the original image, neglecting that multiple views, e.g, saturation level, color histogram and depth of field of the image, will be of benefit to successfully reflecting the subtle aesthetic changes. Taken multi-view into consideration, we leverage a graph neural network to construct the relationships between multi-view features. Besides, multiple views are aggregated with an adaptive-weight self-attention module to verify the significance of each view. Finally, a siamese network is proposed to select the best one from a series of nearly identical photos. Experimental results demonstrate that our model accomplish the highest success rates compared with competitive methods.
Lu Zhang 0062, Yongshun Gong, Jian Zhang 0002, Xiushan Nie, Yilong Yin
ICME4
2022 Overlap-Guided Coarse-to-Fine Correspondence Prediction for Point Cloud Registration
abstract
Establishing reliable correspondences between a pair of point clouds is essential for registration with partial overlaps. However, existing correspondence estimation works usually struggle to distinguish the points in overlap and non-overlap regions. This paper thus proposes an Overlap-guided Coarse-to-Fine Network, named OCFNet, which first establishes correspondences at a coarse level and then refines them at a point level. Specifically, at the coarse level, our model first aggregates two point clouds into smaller sets of super-points with associated features and overlap scores, followed by establishing coarse-level correspondences between the two sets of super-points under the guidance of overlap scores. On the fine stage, a decoder recovers the raw points while jointly learning the associated features and overlap scores. Coarse-level proposals are then expanded to patches, and point-level correspondences are sequentially refined from the corresponding patches. We conducted comprehensive experiments on 3DMatch, 3DLoMatch, and KITTI benchmarks to show the effectiveness of the proposed method. [code]
Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001
ICME3
2022 Robust real-world point cloud registration by inlier detection
Xiaoshui Huang, Yangfu Wang, Sheng Li 0020, Guofeng Mei, Zongyi Xu, Yucheng Wang 0003, Jian Zhang 0002, Mohammed Bennamoun
Comput. Vis. Image Underst.7
2022 Distribution-Aware Margin Calibration for Semantic Segmentation in Images
Litao Yu, Zhibin Li 0002, Min Xu 0001, Yongsheng Gao 0001, Jiebo Luo 0001, Jian Zhang 0002
Int. J. Comput. Vis.6
2022 Blockchain-Enabled Fish Provenance and Quality Tracking System
abstract
Accurate assessment of fish quality is difficult in practice due to the lack of trusted fish provenance and quality tracking information. Working with Sydney Fish Market (SFM), we develop a Blockchain-enabled fish provenance and quality tracking (BeFAQT) system. A multilayer Blockchain architecture based on attribute-based encryption (ABE) is proposed to tackle the privacy issue caused by applying Blockchain to secure supply chain data and achieve trusted and confidential data sharing among parties in fish supply chains. An Internet-of-Things (IoT) chain saves encrypted fish provenance and quality tracking data, and an ABE chain is specifically designed for the access control to the data in the IoT chain. Latest IoT and artificial intelligence (AI) technologies, including NarrowBand-IoT, image processing, and biosensing, are developed for fish origin proof, supply chain tracking, and objective fish quality assessment. As proven by field trials with SFM and a local fish supply chain, the BeFAQT is able to provide trusted and comprehensive fish provenance and quality tracking information in real time.
Xu Wang 0004, Guangsheng Yu, Ren Ping Liu 0001, Jian Zhang 0002, Qiang Wu 0001, Steven W. Su, Ying He 0011, Zongjian Zhang, Litao Yu, Taoping Liu, Wentian Zhang, Peter Loneragan, Eryk Dutkiewicz, Erik Poole, Nick Paton
IEEE Internet Things J.4
2022 TOAN: Target-Oriented Alignment Network for Fine-Grained Image Categorization With Few Labeled Samples
abstract
In this paper, we study the fine-grained categorization problem under the few-shot setting, i.e., each fine-grained class only contains a few labeled examples, termed Fine-Grained Few-Shot classification (FGFS). The core predicament in FGFS is the high intra-class variance yet low inter-class fluctuations in the dataset. In traditional fine-grained classification, the high intra-class variance can be somewhat relieved by conducting the supervised training on the abundant labeled samples. However, with few labeled examples, it is hard for the FGFS model to learn a robust class representation with the significantly higher intra-class variance. Moreover, the inter- and intra-class variance are closely related. The significant intra-class variance in FGFS often aggravates the low inter-class variance issue. To address the above challenges, we propose a Target-Oriented Alignment Network (TOAN) to tackle the FGFS problem from both intra- and inter-class perspective. To reduce the intra-class variance, we propose a target-oriented matching mechanism to reformulate the spatial features of each support image to match the query ones in the embedding space. To enhance the inter-class discrimination, we devise discriminative fine-grained features by integrating local compositional concept representations with the global second-order pooling. We conducted extensive experiments on four public datasets for fine-grained categorization, and the results show the proposed TOAN obtains the state-of-the-art.
Huaxi Huang, Junjie Zhang 0002, Litao Yu, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 Collaborative Feature Learning for Gait Recognition Under Cloth Changes
abstract
Since gait can be utilized to identify individuals from a far distance without their interaction and coordination, recently many gait recognition methods have been proposed. However, due to a real-world scenario of clothing changes, a degradation occurs for most of these methods. Thus in this paper, a more efficient gait recognition method is proposed to address the problem of clothing variances. First, part-based gait features are formulated from two different perspectives,i.e., the separated body parts that are more robust to clothing changes and the estimated human skeleton key-point regions. It is reasonable to formulate such features for cloth-changing gait recognition, because these two perspectives are both less vulnerable to clothing changes. Given that each feature has its own advantages and disadvantages, a more efficient gait feature is generated in this paper by assembling these two features together. Moreover, since local features are more discriminative than global features, in this paper more attention is focused on the local short-range features. Also, unlike most methods, in our method we treat the estimated key-point features as a set of word embeddings, and a transformer encoder is specifically used to learn the dependence of each correlative key-points. The robustness and effectiveness of our proposed method are certified by experiments on CASIA Gait Dataset B, and it has achieved the state-of-the-art performance on this dataset.
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2022 Self-Supervised Depth Completion From Direct Visual-LiDAR Odometry in Autonomous Driving
abstract
In this work, a simple yet effective deep neural network is proposed to generate the dense depth map of the scene by exploiting both LiDAR sparse point cloud and the monocular camera image. Specifically, a feature pyramid network is firstly employed to extract feature maps from images across time. Then the relative pose is calculated by minimizing the feature distance between aligned pixels from inter-frame feature maps. Finally, the feature maps and the relative pose are further applied to compute the feature-metric loss for training the depth completion network. The key novelty of this work lies in that a self-supervised mechanism is presented to train the depth completion network by directly using visual-LiDAR odometry between consecutive frames. Comprehensive experiments and ablation studies on benchmark dataset KITTI demonstrate the superior performance over other state-of-the-art methods in terms of pose estimation and depth completion. The detailed performance of the proposed approach (referred to asSelfCompDVLO) can be found on the KITTI depth completion benchmark. The source code, models, and data have been made available at GitHub.
Zhenbo Song, Jianfeng Lu 0003, Yazhou Yao, Jian Zhang 0002
IEEE Trans. Intell. Transp. Syst.4
2022 Online Spatio-Temporal Crowd Flow Distribution Prediction for Complex Metro System
abstract
As a key mission of the modern traffic management, crowd flow prediction (CFP) benefits in many tasks of intelligent transportation services. However, most existing techniques focus solely on forecasting entrance and exit flows of metro stations that do not provide enough useful knowledge for traffic management. In practical applications, managers desperately want to solve the problem of getting the potential passenger distributions to help authorities improve transport services, termed as crowd flow distribution (CFD) forecasts. Therefore, to improve the quality of transportation services, we proposed three spatiotemporal models to effectively address the network-wide CFD prediction problem based on the online latent space (OLS) strategy. Our models take into account the various trending patterns and climate influences, as well as the inherent similarities among different stations that are able to predict both CFD and entrance and exit flows precisely. In our online systems, a sequence of CFD snapshots is used as the training data. The latent attribute evolutions of different metro stations can be learned from the previous trend and do the next prediction based on the transition patterns. All the empirical results demonstrate that the three developed models outperform all the other state-of-the-art approaches on three large-scale real-world datasets.
Yongshun Gong, Zhibin Li 0002, Jian Zhang 0002, Wei Liu 0007, Yu Zheng 0004
IEEE Trans. Knowl. Data Eng.3
2022 Semantically Meaningful Class Prototype Learning for One-Shot Image Segmentation
abstract
One-shot semantic image segmentation aims to segment the object regions for the novel class with only one annotated image. Recent works adopt the episodic training strategy to mimic the expected situation at testing time. However, these existing approaches simulate the test conditions too strictly during the training process, and thus cannot make full use of the given label information. Besides, these approaches mainly focus on the foreground-background target class segmentation setting. They only utilize binary mask labels for training. In this paper, we propose to leverage the multi-class label information during the episodic training. It will encourage the network to generate more semantically meaningful features for each category. After integrating the target class cues into the query features, we then propose a pyramid feature fusion module to mine the fused features for the final classifier. Furthermore, to take more advantage of the support image-mask pair, we propose a self-prototype guidance branch to support image segmentation. It can constrain the network for generating more compact features and a robust prototype for each semantic class. For inference, we propose a fused prototype guidance branch for the segmentation of the query image. Specifically, we leverage the prediction of the query image to extract the pseudo-prototype and combine it with the initial prototype. Then we utilize the fused prototype to guide the final segmentation of the query image. Extensive experiments demonstrate the superiority of our proposed approach. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/SMCP.
Tao Chen 0012, Guosen Xie, Yazhou Yao, Qiong Wang 0003, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.7
2022 Exploiting Web Images for Fine-Grained Visual Recognition by Eliminating Open-Set Noise and Utilizing Hard Examples
abstract
Labeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention. However, the presence of label noise and hard examples in web images are two obstacles for training robust fine-grained recognition models. Therefore, in this paper, we propose a novel approach for removing irrelevant samples from real-world web images during training, while employing useful hard examples to update the network. Thus, our approach can alleviate the harmful effects of irrelevant noisy web images and hard examples to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Advanced-Softly-Update-Drop.
Huafeng Liu 0004, Chuanyi Zhang, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.7
2022 Dual Attention on Pyramid Feature Maps for Image Captioning
abstract
Generating natural sentences from images is a fundamental learning task for visual-semantic understanding in multimedia. In this paper, we propose to apply dual attention on pyramid image feature maps to fully explore the visual-semantic correlations and improve the quality of generated sentences. Specifically, with the full consideration of the contextual information provided by the hidden state of the RNN controller, the pyramid attention can better localize the visually indicative and semantically consistent regions in images. On the other hand, the contextual information can help re-calibrate the importance of feature components by learning the channel-wise dependencies, to improve the discriminative power of visual features for better content description. We conducted comprehensive experiments on three well-known datasets: Flickr8K, Flickr30 K and MS COCO, which achieved impressive results in generating descriptive and smooth natural sentences from images. Using either convolution visual features or more informative bottom-up attention features, the composite model can boost the performance of image-to-sentence translation, with a limited computational resource overhead. The proposed pyramid attention and dual attention methods are highly modular, which can be inserted into various image captioning modules to further improve the performance.
Litao Yu, Jian Zhang 0002, Qiang Wu 0001
IEEE Trans. Multim.2
2022 Guest Editorial: Learning From Noisy Multimedia Data
abstract
This special issue provides a premier forum for researchers in multimedia big data to share challenges and recent advancements in learning from noisy multimedia data. The multimedia age and its proliferation of devices and platforms is fueling exponential data growth. As computational power and deep learning algorithms rapidly evolve, the web has become a rich source of potential training data for robust machine learning, with search engines such as Google and Bing, Twitter, TikTok, Instagram, and short video sharing platforms offering large-scale data points in the hundreds of millions. The concurrent shift in the Internet to richer web data modalities such as text, audio, image, and video reveal further opportunities to leverage large-scale data for the automatic construction of a variety of datasets for model training and testing. However, the ubiquity of multimedia data means noise is a fundamental challenge, with ‘label noise’ and ‘domain mismatch’ the most critical issues in automatically collected datasets. Learning from noisy multimedia data tends towards poor performance, making it increasingly essential to address these challenges.
Jian Zhang 0002, Alan Hanjalic, Ramesh Jain 0001, Xian-Sheng Hua 0001, Shin'ichi Satoh 0001, Yazhou Yao, Dan Zeng 0001
IEEE Trans. Multim.1
2022 Multimodal Marketing Intent Analysis for Effective Targeted Advertising
abstract
People’s daily information sharing and acquisition through the Internet has become more and more popular. The comprehensive multimodal marketing advertorial generated by ‘We Media’ accounts besides the normal social news is gaining its importance on social media platforms. In order to achieve effective advertising, the marketing intent understanding is a key step towards generating targeted advertising strategies (push advertorials to specific people at a specific time). However, advertorials in real are usually designed to pretend as normal social news with a wide range of contents. This poses big challenges to the platforms on accurately recognizing and analyzing the marketing intents behind the advertorials. As a pioneering study, we address this new problem of multimodal-based marketing intent analysis and answer three core questions: (1) does a piece of social news contain marketing intent? (2) what is the topic of marketing intent? (3) what is the extent of marketing intent? Towards this end, we propose a novel Multimodal-based Marketing Intent Analysis scheme (MMIA) to estimate the marketing intent embedded in the multimodal contents. Specifically, a novel supervised neural autoregressive model (SmiDocNADE) is proposed to enhance the discriminative capacity of the learned hidden features so that a single system is capable of solving the three questions. In order to effectively model inter-correlations between images and text in advertorials, we fuse multimodal data and extract features by Graph Convolution Networks as an enhancement to SmiDocNADE. The extensive evaluations demonstrate the advantages of our proposed system in multimodal-based marketing intent analysis from multiple aspects.
Lu Zhang 0062, Jialie Shen 0001, Jian Zhang 0002, Jingsong Xu, Zhibin Li 0002, Yazhou Yao, Litao Yu
IEEE Trans. Multim.3
2022 Unsupervised Image and Text Fusion for Travel Information Enhancement
abstract
With the explosive growth of the shared information on social media platforms, people are increasingly interested in sharing and making their travel plans by referring to others’ travel experiences. However, different social media sources render the heterogeneity of these valuable data, bringing difficulties for data collection and fusion. Thus, facing massive information online, one of the biggest challenges to enhance travel information is how to integrate and match these multi-source data without clear labels. In this paper, we propose an unsupervised method to fuse and match images and travelogues. We first use the three textual components (title, tag, and description) of the descriptive texts of images as three criteria to embed travelogues and the descriptive texts of images, and further introduce images into our method by joint embedding texts and images. Finally, a multiple kernel clustering approach is adopted for matching travelogues and images. Extensive experiments conducted on the real dataset crawled from two websites (Flickr and TripAdvisor) demonstrate the effectiveness and robustness of our proposed method.
Lu Zhang 0062, Jingsong Xu, Yongshun Gong, Litao Yu, Jian Zhang 0002, Jialie Shen 0001
IEEE Trans. Multim.5
2022 Recognizing Gaits Across Walking and Running Speeds
abstract
For decades, very few methods were proposed for cross-mode (i.e., walking vs. running) gait recognition. Thus, it remains largely unexplored regarding how to recognize persons by the way they walk and run. Existing cross-mode methods handle the walking-versus-running problem in two ways, either by exploring the generic mapping relation between walking and running modes or by extracting gait features which are non-/less vulnerable to the changes across these two modes. However, for the first approach, a mapping relation fit for one person may not be applicable to another person. There is no generic mapping relation given that walking and running are two highly self-related motions. The second approach does not give more attention to the disparity between walking and running modes, since mode labels are not involved in their feature learning processes. Distinct from these existing cross-mode methods, in our method, mode labels are used in the feature learning process, and a mode-invariant gait descriptor is hybridized for cross-mode gait recognition to handle this walking-versus-running problem. Further research is organized in this article to investigate the disparity between walking and running. Running is different from walking not only in the speed variances but also, more significantly, in prominent gesture/motion changes. According to these rationales, in our proposed method, we give more attention to the differences between walking and running modes, and a robust gait descriptor is developed to hybridize the mode-invariant spatial and temporal features. Two multi-task learning-based networks are proposed in this method to explore these mode-invariant features. Spatial features describe the body parts non-/less affected by mode changes, and temporal features depict the instinct motion relation of each person. Mode labels are also adopted in the training phase to guide the network to give more attention to the disparity across walking and running modes. In addition, relevant experiments on OU-ISIR Treadmill Dataset A have affirmed the effectiveness and feasibility of the proposed method. A state-of-the-art result can be achieved by our proposed method on this dataset.
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.5
2021 PTN: A Poisson Transfer Network for Semi-supervised Few-shot Learning
abstract
The predicament in semi-supervised few-shot learning (SSFSL) is to maximize the value of the extra unlabeled data to boost the few-shot learner. In this paper, we propose a Poisson Transfer Network (PTN) to mine the unlabeled information for SSFSL from two aspects. First, the Poisson Merriman–Bence–Osher (MBO) model builds a bridge for the communications between labeled and unlabeled examples. This model serves as a more stable and informative classifier than traditional graph-based SSFSL methods in the message-passing process of the labels. Second, the extra unlabeled samples are employed to transfer the knowledge from base classes to novel classes through contrastive learning. Specifically, we force the augmented positive pairs close while push the negative ones distant. Our contrastive transfer scheme implicitly learns the novel-class embeddings to alleviate the over-fitting problem on the few labeled data. Thus, we can mitigate the degeneration of embedding generality in novel classes. Extensive experiments indicate that PTN outperforms the state-of-the-art few-shot and SSFSL models on miniImageNet and tieredImageNet benchmark datasets.
Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002
AAAI3
2021 Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of pseudo labels within the image’s salient region. In this work, we propose a non-salient region object mining approach for weakly supervised semantic segmentation. We introduce a graph-based global reasoning unit to strengthen the classification network’s ability to capture global relations among disjoint and distant regions. This helps the network activate the object features outside the salient area. To further mine the non-salient region objects, we propose to exert the segmentation network’s self-correction ability. Specifically, a potential object mining module is proposed to reduce the false-negative rate in pseudo labels. Moreover, we propose a non-salient region masking module for complex images to generate masked pseudo labels. Our non-salient region masking module helps further discover the objects in the non-salient region. Extensive experiments on the PASCAL VOC dataset demonstrate state-of-the-art results compared to current methods. The source codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/nsrom.
Yazhou Yao, Tao Chen 0012, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Zhenmin Tang, Jian Zhang 0002
CVPR8
2021 Jo-SRC: A Contrastive Approach for Combating Noisy Labels
abstract
Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature tends to perform sample selection within each mini-batch, neglecting the imbalance of noise ratios in different mini-batches. Moreover, valuable knowledge within high-loss samples is wasted. To this end, we propose a noise-robust approach named Jo-SRC (Joint Sample Selection and Model Regularization based on Consistency). Specifically, we train the network in a contrastive learning manner. Predictions from two different views of each sample are used to estimate its "likelihood" of being clean or out-of-distribution. Furthermore, we propose a joint loss to advance the model generalization performance by introducing consistency regularization. Extensive experiments have validated the superiority of our approach over existing state-of-the-art methods. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/Jo-SRC.
Yazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Jian Zhang 0002, Zhenmin Tang
CVPR6
2021 Webly Supervised Fine-Grained Recognition: Benchmark Datasets and An Approach
abstract
Learning from the web can ease the extreme dependence of deep learning on large-scale manually labeled datasets. Especially for fine-grained recognition, which targets at distinguishing subordinate categories, it will significantly reduce the labeling costs by leveraging free web data. Despite its significant practical and research value, the webly supervised fine-grained recognition problem is not extensively studied in the computer vision community, largely due to the lack of high-quality datasets. To fill this gap, in this paper we construct two new benchmark webly supervised fine-grained datasets, termed WebFG-496 and WebiNat-5089, respectively. In concretely, WebFG-496 consists of three sub-datasets containing a total of 53,339 web training images with 200 species of birds (Web-bird), 100 types of aircrafts (Web-aircraft), and 196 models of cars (Web-car). For WebiNat-5089, it contains 5089 sub-categories and more than 1.1 million web training images, which is the largest webly supervised fine-grained dataset ever. As a minor contribution, we also propose a novel webly supervised method (termed "Peer-learning") for benchmarking these datasets. Comprehensive experimental results and analyses on two new benchmark datasets demonstrate that the proposed method achieves superior performance over the competing baseline models and states-of-the-art. Our benchmark datasets and the source codes of Peer-learning have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/weblyFG-dataset.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jianxin Wu 0001, Jian Zhang 0002, Heng Tao Shen
ICCV7
2021 Incorporating Multimodal Cues for Advertorial Discovery
abstract
Commercial advertorials shared on websites are usually designed to pretend as normal social news for commercial benefits. The analysis of the commercial intents embedded in advertorials can greatly help media platforms personalize content. However, commercial intents are not only concealed in news texts but also conveyed by news images explicitly or implicitly. Consequently, how to effectively extract and incorporate the crucial cues of multiple modalities has been emerging as an important but challenging problem. Motivated by this observation, we propose a framework Multimodal Advertorial Discovery Model (MADM) to estimate the commercial intents embedded in the multimodal social news. Specifically, a novel Cross-graph Fusion (CGF) strategy is developed to achieve a soft assignment to incorporate images and text and generate comprehensive multimodal representations. The extensive evaluations demonstrate the superiority of our proposed system in multimodal-based advertorial detection and analysis.
Lu Zhang 0062, Jian Zhang 0002, Jialie Shen 0001, Jingsong Xu, Zhibin Li 0002, Litao Yu
ICME2
2021 A Multi-branch Hybrid Transformer Network for Corneal Endothelial Cell Segmentation
Yinglin Zhang, Risa Higashita, Huazhu Fu, Yanwu Xu 0001, Haofeng Liu, Jian Zhang 0002, Jiang Liu 0001
MICCAI (1)7
2021 Inferring the Importance of Product Appearance with Semi-supervised Multi-modal Enhancement: A Step Towards the Screenless Retailing
abstract
Nowadays, almost all the online orders were placed through screened devices such as mobile phones, tablets, and computers. With the rapid development of the Internet of Things (IoT) and smart appliances, more and more screenless smart devices, e.g., smart speaker and smart refrigerator, appear in our daily lives. They open up new means of interaction and may provide an excellent opportunity to reach new customers and increase sales. However, not all the items are suitable for screenless shopping, since some items' appearance play an important role in consumer decision making. Typical examples include clothes, dolls, bags, and shoes. In this paper, we aim to infer the significance of every item's appearance in consumer decision making and identify the group of items that are suitable for screenless shopping. Specifically, we formulate the problem as a classification task that predicts if an item's appearance has a significant impact on people's purchase behavior. To solve this problem, we extract multi-modal features from three different views, and collect a set of necessary labels via crowdsourcing. We then propose an iterative semi-supervised learning framework with a carefully designed multi-modal enhancement module. Experimental results verify the effectiveness of the proposed method.
Yongshun Gong, Jinfeng Yi, Jian Zhang 0002, Zhihua Zhou
ACM Multimedia4
2021 Structured discriminative tensor dictionary learning for unsupervised domain adaptation
Songsong Wu, Yan Yan 0002, Hao Tang 0005, Jianjun Qian, Jian Zhang 0002, Xiaoyuan Jing
Neurocomputing5
2021 GID-Net: Detecting human-object interaction with global and instance dependency
Dongming Yang, Yuexian Zou, Jian Zhang 0002, Ge Li 0002
Neurocomputing3
2021 Exploring the auxiliary learning for long-tailed visual recognition
Junjie Zhang 0002, Lingqiao Liu, Peng Wang 0023, Jian Zhang 0002
Neurocomputing4
2021 ASDN: A Deep Convolutional Network for Arbitrary Scale Image Super-Resolution
Jialiang Shen, Yucheng Wang 0003, Jian Zhang 0002
Mob. Networks Appl.3
2021 Constructing multilayer locality-constrained matrix regression framework for noise robust face super-resolution
Guangwei Gao, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Meng Yang 0001, Jian Zhang 0002
Pattern Recognit.6
2021 Exploiting textual queries for dynamically visual disambiguation
abstract
Due to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits the performance of current webly supervised models is the problem of visual polysemy. In this work, we present a novel framework that resolves visual polysemy by dynamically matching candidate text queries with retrieved images. Specifically, our proposed framework includes three major steps: we first discover and then dynamically select the text queries according to the keyword-based image search results, we employ the proposed saliency-guided deep multi-instance learning (MIL) network to remove outliers and learn classification models for visual disambiguation. Compared to existing methods, our proposed approach can figure out the right visual senses, adapt to dynamic changes in the search results, remove outliers, and jointly learn the classification models . Extensive experiments and ablation studies on CMU-Poly-30 and MIT-ISD datasets demonstrate the effectiveness of our proposed approach.
Zeren Sun, Yazhou Yao, Jimin Xiao, Lei Zhang 0054, Jian Zhang 0002, Zhenmin Tang
Pattern Recognit.5
2021 Robust gait recognition using hybrid descriptors based on Skeleton Gait Energy Image
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang, Wankou Yang
Pattern Recognit. Lett.4
2021 Low-Rank Pairwise Alignment Bilinear Network For Few-Shot Fine-Grained Image Classification
abstract
Deep neural networks have demonstrated advanced abilities on various visual classification tasks, which heavily rely on the large-scale training samples with annotated ground-truth. However, it is unrealistic always to require such annotation in real-world applications. Recently, Few-Shot learning (FS), as an attempt to address the shortage of training samples, has made significant progress in generic classification tasks. Nonetheless, it is still challenging for current FS models to distinguish the subtle differences between fine-grained categories given limited training data. To filling the classification gap, in this paper, we address the Few-Shot Fine-Grained (FSFG) classification problem, which focuses on tackling the fine-grained classification under the challenging few-shot learning setting. A novel low-rank pairwise bilinear pooling operation is proposed to capture the nuanced differences between the support and query images for learning an effective distance metric. Moreover, a feature alignment layer is designed to match the support image features with query ones before the comparison. We name the proposed model Low-Rank Pairwise Alignment Bilinear Network (LRPABN), which is trained in an end-to-end fashion. Comprehensive experimental results on four widely used fine-grained classification data sets demonstrate that our LRPABN model achieves the superior performances compared to state-of-the-art methods.
Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Jingsong Xu, Qiang Wu 0001
IEEE Trans. Multim.3
2021 Deep Unsupervised Self-Evolutionary Hashing for Image Retrieval
abstract
Hashing methods have proven to be effective in the field of large-scale image retrieval. In recent years, the performance of hashing algorithms based on deep learning has greatly exceeded that of non-deep methods. However, most of the outstanding hashing methods are supervised models that heavily rely on annotated labels. In order to circumvent the huge overhead of labeling large-scale datasets, some unsupervised hashing algorithms have been proposed, such as pseudo labels and pseudo pairs. Since the image labels are strictly unavailable, some hyper-parameters in these methods are difficult to be selected, e.g., the final result is very sensitive to the picked number of categories or the chosen threshold of similarity for pairs. In addition, the calculation of pseudo-labels in high-dimensional space is not only computationally complex, but also has low precision. Therefore, in order to alleviate these issues in this paper, we propose a simple but effective Deep Unsupervised Self-evolutionary Hashing (DUSH) algorithm, which utilizes a curriculum learning strategy to iteratively select pseudo pairs from easy to hard in low dimensional Hamming space. Extensive experiments are conducted on four popular datasets, including two single-label datasets and two multi-label datasets, and the results show that our method can significantly outperform the state-of-the-art methods.
Haofeng Zhang 0001, Yazhou Yao, Zheng Zhang 0006, Li Liu 0004, Jian Zhang 0002, Ling Shao 0001
IEEE Trans. Multim.6
2021 Parameter-Efficient Deep Neural Networks With Bilinear Projections
abstract
Recent research on deep neural networks (DNNs) has primarily focused on improving the model accuracy. Given a proper deep learning framework, it is generally possible to increase the depth or layer width to achieve a higher level of accuracy. However, the huge number of model parameters imposes more computational and memory usage overhead and leads to the parameter redundancy. In this article, we address the parameter redundancy problem in DNNs by replacing conventional full projections with bilinear projections (BPs). For a fully connected layer with D input nodes and D output nodes, applying BP can reduce the model space complexity fromO(D2) toO(2D), achieving a deep model with a sublinear layer size. However, the structured projection has a lower freedom of degree compared with the full projection, causing the underfitting problem. Therefore, we simply scale up the mapping size by increasing the number of output channels, which can keep and even boosts the model accuracy. This makes it very parameter-efficient and handy to deploy such deep models on mobile systems with memory limitations. Experiments on four benchmark data sets show that applying the proposed BP to DNNs can achieve even higher accuracies than conventional full DNNs while significantly reducing the model size.
Litao Yu, Yongsheng Gao 0001, Jun Zhou 0001, Jian Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.4
2020 Potential Passenger Flow Prediction: A Novel Study for Urban Transportation Development
abstract
Recently, practical applications for passenger flow prediction have brought many benefits to urban transportation development. With the development of urbanization, a real-world demand from transportation managers is to construct a new metro station in one city area that never planned before. Authorities are interested in the picture of the future volume of commuters before constructing a new station, and estimate how would it affect other areas. In this paper, this specific problem is termed as potential passenger flow (PPF) prediction, which is a novel and important study connected with urban computing and intelligent transportation systems. For example, an accurate PPF predictor can provide invaluable knowledge to designers, such as the advice of station scales and influences on other areas, etc. To address this problem, we propose a multi-view localized correlation learning method. The core idea of our strategy is to learn the passenger flow correlations between the target areas and their localized areas with adaptive-weight. To improve the prediction accuracy, other domain knowledge is involved via a multi-view learning process. We conduct intensive experiments to evaluate the effectiveness of our method with real-world official transportation datasets. The results demonstrate that our method can achieve excellent performance compared with other available baselines. Besides, our method can provide an effective solution to the cold-start problem in the recommender system as well, which proved by its outperformed experimental results.
Yongshun Gong, Zhibin Li 0002, Jian Zhang 0002, Wei Liu 0007, Jinfeng Yi
AAAI3
2020 Feature-Metric Registration: A Fast Semi-Supervised Approach for Robust Point Cloud Registration Without Correspondences
abstract
We present a fast feature-metric point cloud registration framework, which enforces the optimisation of registration by minimising a feature-metric projection error without correspondences. The advantage of the feature-metric projection error is robust to noise, outliers and density difference in contrast to the geometric projection error. Besides, minimising the feature-metric projection error does not need to search the correspondences so that the optimisation speed is fast. The principle behind the proposed method is that the feature difference is smallest if point clouds are aligned very well. We train the proposed method in a semi-supervised or unsupervised approach, which requires limited or no registration label data. Experiments demonstrate our method obtains higher accuracy and robustness than the state-of-the-art methods. Besides, experimental results show that the proposed method can handle significant noise and density difference, and solve both same-source and cross-source point cloud registration.
Xiaoshui Huang, Guofeng Mei, Jian Zhang 0002
CVPR3
2020 Exploring Long-Short-Term Context For Point Cloud Semantic Segmentation
abstract
Point cloud semantic segmentation attracts numerous attention following the success of the point-based convolution neural network. Due to the ambiguity of the point-based feature, many methods study on integrating contextual information to solve the ambiguous problem. However, the extracted context is severely limited to the small input blocks. Few prior works exploit contextual information beyond the blocks to capture long-range dependencies. To address this limitation, we propose a novel long-short-term context framework, which adopts a long-short-term feature bank to exploit both the local context within each block and the long-range context beyond the current task block. The proposed framework is flexible and easy to be combined with existing models, thereby enables existing models to capture the larger range context. Extensive experiments demonstrate that the proposed model achieves improved segmentation performance, and augmenting existing models with a long-short-term feature bank consistently increases the performance.
Anan Du, Shuchao Pang, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001
ICIP4
2020 Classification Constrained Discriminator For Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation aims to transfer knowledge from label-rich synthetic datasets to real-world images without any annotation. The traditional adversarial learning methods for domain adaptation learn to extract domain-invariant feature representations by aligning the feature distributions of both domains. However, these methods suffer from an imbalance in adversarial training and feature distortion. In this work, we propose a classification constrained discriminator to alleviate these problems. Specifically, we first propose to balance the adversarial training by eliminating any pooling layers or strided convolutions in the discriminator. Then, we propose to constrain the discriminator with an auxiliary classification loss to help the feature generator extract the domain-invariant features that are useful for segmentation rather than just ambiguous features to fool the domain discriminator. Extensive experiments demonstrate the superiority of our proposed approach. The source code and models have been made available at https://github.com/NUSTMachine-Intelligence-Laboratory/ccd.
Tao Chen 0012, Jian Zhang 0002, Guosen Xie, Yazhou Yao, Xiaoshui Huang, Zhenmin Tang
ICME2
2020 Web-Supervised Network for Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) is a tough task due to its high annotation cost of the fine-grained subcategories. To build a large-scale dataset at low manual cost, straightforwardly learning from web images for FGVC has attracted broad attention. However, there exist two characteristics in the need of concerning for the web dataset: 1) Noisy images; 2) A large proportion of hard examples. In this paper, we propose a simple yet effective approach to deal with noisy images and hard examples during training. Our method is a pure web-supervised method for FGVC. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to the state-of-the-art web-supervised methods. The data and source code of this work have been posted available at: https://github.com/NUST-Machine-Intelligence-Laboratory/WSNFG.
Chuanyi Zhang, Yazhou Yao, Jiachao Zhang, Jian Zhang 0002, Zhenmin Tang
ICME6
2020 Towards Better Graph Representation: Two-Branch Collaborative Graph Neural Networks For Multimodal Marketing Intention Detection
abstract
Inspired by the fact that spreading and collecting information through the Internet becomes the norm, more and more people choose to post for-profit contents (images and texts) in social networks. Due to the difficulty of network censors, malicious marketing may be capable of harming the society. Therefore, it is meaningful to detect marketing intentions online automatically. However, gaps between multimodal data make it difficult to fuse images and texts for content marketing detection. To this end, this paper proposes Two-Branch Collaborative Graph Neural Networks to collaboratively represent multimodal data by Graph Convolution Networks (GCNs) in an end-to-end fashion. We first separately embed groups of images and texts by GCNs layers from two views and further adopt the proposed multimodal fusion strategy to learn the graph representation collaboratively. Experimental results demonstrate that our proposed method achieves superior graph classification performance for marketing intention detection.
Lu Zhang 0062, Jian Zhang 0002, Zhibin Li 0002, Jingsong Xu
ICME2
2020 Part-based Collaborative Spatio-temporal Feature Learning for Cloth-changing Gait Recognition
abstract
In decades many gait recognition methods have been proposed using different techniques. However, due to a real-world scenario of clothing variations, a reduction of the recognition rate occurs for most of these methods. Thus in this paper, a part-based spatio-temporal feature learning method is proposed to tackle the problem of clothing variations for gait recognition. First, based on the anatomical properties, human bodies are segmented into two regions, which are affected and unaffected by clothing variations. A learning network is particularly proposed in this paper to grasp principal spatio-temporal features from those unaffected regions. Different from most part-based methods with spatial or temporal features solely being utilized, in our method these two features are associated in a more collaborative manner. Snapshots are created for each gait sequence from the H-W and T-W views. Stable spatial information is embedded in the H-W view and adequate temporal information is embedded in the T-W view. An inherent relationship exists between these two views. Thus, a collaborative spatio-temporal feature will be hybridized by concatenating these correlative spatial and temporal information. The robustness and efficiency of our proposed method are validated by experiments on CASIA Gait Dataset B and OU-ISIR Treadmill Gait Dataset B. Our proposed method can both achieve the state-of-the-art results on these two databases.
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Jingsong Xu
ICPR4
2020 A Spatial Missing Value Imputation Method for Multi-view Urban Statistical Data
abstract
Large volumes of urban statistical data with multiple views imply rich knowledge about the development degree of cities. These data present crucial statistics which play an irreplaceable role in the regional analysis and urban computing. In reality, however, the statistical data divided into fine-grained regions usually suffer from missing data problems. Those missing values hide the useful information that may result in a distorted data analysis. Thus, in this paper, we propose a spatial missing data imputation method for multi-view urban statistical data. To address this problem, we exploit an improved spatial multi-kernel clustering method to guide the imputation process cooperating with an adaptive-weight non-negative matrix factorization strategy. Intensive experiments are conducted with other state-of-the-art approaches on six real-world urban statistical datasets. The results not only show the superiority of our method against other comparative methods on different datasets, but also represent a strong generalizability of our model.
Yongshun Gong, Zhibin Li 0002, Jian Zhang 0002, Wei Liu 0007, Bei Chen 0008, Xiangjun Dong 0001
IJCAI3
2020 CRSSC: Salvage Reusable Samples from Noisy Data for Robust Learning
abstract
Due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause severe accumulated errors. Sample selection methods identify clean ("easy") samples based on the fact that small losses can alleviate the accumulated errors. However, "hard" and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the networks. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Xian-Sheng Hua 0001, Yazhou Yao, Xiu-Shen Wei, Guosheng Hu, Jian Zhang 0002
ACM Multimedia6
2020 Bridging the Web Data and Fine-Grained Visual Recognition via Alleviating Label Noise and Domain Mismatch
abstract
To distinguish the subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, we propose to directly leverage web images for fine-grained visual recognition. Our work mainly focuses on two critical issues including "label noise" and "domain mismatch" in the web images. Specifically, we propose an end-to-end deep denoising network (DDN) model to jointly solve these problems in the process of web images selection. To verify the effectiveness of our proposed approach, we first collect web images by using the labels in fine-grained datasets. Then we apply the proposed deep denoising network model for noise removal and domain mismatch alleviation. We leverage the selected web images as the training set for fine-grained categorization models learning. Extensive experiments and ablation studies demonstrate state-of-the-art performance gained by our proposed approach, which, at the same time, delivers a new pipeline for fine-grained visual categorization that is to be highly effective for real-world applications.
Yazhou Yao, Xian-Sheng Hua 0001, Guanyu Gao, Zeren Sun, Zhibin Li 0002, Jian Zhang 0002
ACM Multimedia6
2020 Field-wise Learning for Multi-field Categorical Data
abstract
We propose a new method for learning with multi-field categorical data. Multi-field categorical data are usually collected over many heterogeneous groups. These groups can reflect in the categories under a field. The existing methods try to learn a universal model that fits all data, which is challenging and inevitably results in learning a complex model. In contrast, we propose a field-wise learning method leveraging the natural structure of data to learn simple yet efficient one-to-one field-focused models with appropriate constraints. In doing this, the models can be fitted to each category and thus can better capture the underlying differences in data. We present a model that utilizes linear models with variance and low-rank constraints, to help it generalize better and reduce the number of parameters. The model is also interpretable in a field-wise manner. As the dimensionality of multi-field categorical data can be very high, the models applied to such data are mostly over-parameterized. Our theoretical analysis can potentially explain the effect of over-parametrization on the generalization of our model. It also supports the variance constraints in the learning objective. The experiment results on two large-scale datasets show the superior performance of our model, the trend of the generalization error bound, and the interpretability of learning outcomes. Our code is available at https://github.com/lzb5600/Field-wise-Learning.
Zhibin Li 0002, Jian Zhang 0002, Yongshun Gong, Yazhou Yao, Qiang Wu 0001
NeurIPS2
2020 Automatic Sheep Counting by Multi-object Tracking
abstract
Animal counting is a highly skilled yet tedious task in livestock transportation and trading. To effectively free up the human labour and provide accurate counts for sheep loading/unloading, we develop an auto sheep counting system based on multi-object detection, tracking and extrapolation techniques. Our system has demonstrated more than 99.9% accuracy with sheep moving freely in a race under optimal visual conditions.
Jingsong Xu, Litao Yu, Jian Zhang 0002, Qiang Wu 0001
VCIP3
2020 A Vision Based Fish Processing System
abstract
The digital fish provenance and quality tracking system is essential for the seafood supply chain. As a part of this system, we develop a vision-based fish processing system to automatically perform fish freshness estimation, size measurement and species classification. Under the constrained illumination environment, our system is able to auto-process the fish selection, thus greatly reduce the human labour and bring trust and efficiency to the seafood supply chain from catch to market.
Zongjian Zhang, Litao Yu, Jian Zhang 0002, Qiang Wu 0001
VCIP3
2020 Top-Push Constrained Modality-Adaptive Dictionary Learning for Cross-Modality Person Re-Identification
abstract
Person re-identification aims to match person captured by multiple non-overlapping cameras that mainly mean standard RGB cameras. In contemporary surveillance, cameras of different modalities such as infrared cameras and depth cameras are introduced because of their unique advantages in poor illumination scenarios. However, re-identifying the persons across such cameras of different modalities is extremely difficult and, unfortunately, seldom discussed. It is mainly caused by extremely different appearances of the person shown under such different camera modalities. In this paper, we tackle this challenging cross-modality people re-identification through a top-push constrained modality-adaptive dictionary learning. The proposed model asymmetrically projects the heterogeneous features from dissimilar modalities onto a common space. In this way, the modality-specific bias is mitigated. Thus, the heterogeneous data can be simultaneously enforced by a shared dictionary in a canonical space. Moreover, a top-push ranking graph regularization is embedded in the proposed model to improve the discriminability, which efficiently further boosts the matching accuracy. In order to implement the proposed model, an iterative process is developed in this paper to optimize these two processes jointly. Extensive experiments on the benchmark SYSU-MM01 and BIWI RGBD-ID person re-identification datasets show promising results which outperform state-of-the-art methods.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2020 Scale-Aware Crowd Counting via Depth-Embedded Convolutional Neural Networks
abstract
Scale variation of pedestrians in a crowd image presents a significant challenge for vision-based people counting systems. Such variations are mainly caused by perspective-related distortions due to the camera pose relative to the ground plane. Following the density-based counting paradigm, we postulate that generating density values adaptive to object scales plays a critical role in the accuracy of the final counting results. Motivated by this, we distill the underlying information from depth cues to obtain scale-aware representations that can respond to object scales considering the fact that the scale is inversely proportional to the object depth. Specifically, we propose a depth embedding module as add-ons into existing networks. This module exploits essential depth cues to spatially re-calibrate the magnitude of the original features. In this way, the objects, although in the same class, will attain distinct representations according to their scales, which directly benefits the estimation of scale-aware density values. We conduct a comprehensive analysis of the effects of the depth embedding module and validate that exploiting depth cues to perceive object scale variations in convolutional neural networks improves crowd counting performances. Our experiments demonstrate the effectiveness of the proposed approach on four popular benchmark datasets.
Muming Zhao, Jian Zhang 0002, Fatih Porikli, Bingbing Ni, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Towards Automatic Construction of Diverse, High-Quality Image Datasets
abstract
The availability of labeled image datasets has been shown critical for high-level image understanding, which continuously drives the progress of feature designing and models developing. However, constructing labeled image datasets is laborious and monotonous. To eliminate manual annotation, in this work, we propose a novel image dataset construction framework by employing multiple textual queries. We aim at collecting diverse and accurate images for given queries from the Web. Specifically, we formulate noisy textual queries removing and noisy images filtering as a multi-view and multi-instance learning problem separately. Our proposed approach not only improves the accuracy but also enhances the diversity of the selected images. To verify the effectiveness of our proposed approach, we construct an image dataset with 100 categories. The experiments show significant performance gains by using the generated data of our approach on several tasks, such as image classification, cross-dataset generalization, and object detection. The proposed method also consistently outperforms existing weakly supervised and web-supervised approaches.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Dongxiang Zhang, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.2
2020 Exploiting Web Images for Multi-Output Classification: From Category to Subcategories
abstract
Studies present that dividing categories into subcategories contributes to better image classification. Existing image subcategorization works relying on expert knowledge and labeled images are both time-consuming and labor-intensive. In this article, we propose to select and subsequently classify images into categories and subcategories. Specifically, we first obtain a list of candidate subcategory labels from untagged corpora. Then, we purify these subcategory labels through calculating the relevance to the target category. To suppress the search error and noisy subcategory label-induced outlier images, we formulate outlier images removing and the optimal classification models learning as a unified problem to jointly learn multiple classifiers, where the classifier for a category is obtained by combining multiple subcategory classifiers. Compared with the existing subcategorization works, our approach eliminates the dependence on expert knowledge and labeled images. Extensive experiments on image categorization and subcategorization demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Guosen Xie, Li Liu 0004, Fan Zhu 0001, Jian Zhang 0002, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.6
2019 Mind Your Neighbours: Image Annotation With Metadata Neighbourhood Graph Co-Attention Networks
abstract
As the visual reflections of our daily lives, images are frequently shared on the social network, which generates the abundant 'metadata' that records user interactions with images. Due to the diverse contents and complex styles, some images can be challenging to recognise when neglecting the context. Images with the similar metadata, such as 'relevant topics and textual descriptions', 'common friends of users' and 'nearby locations', form a neighbourhood for each image, which can be used to assist the annotation. In this paper, we propose a Metadata Neighbourhood Graph Co-Attention Network (MangoNet) to model the correlations between each target image and its neighbours. To accurately capture the visual clues from the neighbourhood, a co-attention mechanism is introduced to embed the target image and its neighbours as graph nodes, while the graph edges capture the node pair correlations. By reasoning on the neighbourhood graph, we obtain the graph representation to help annotate the target image. Experimental results on three benchmark datasets indicate that our proposed model achieves the best performance compared to the state-of-the-art methods.
Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003
CVPR3
2019 Leveraging Heterogeneous Auxiliary Tasks to Assist Crowd Counting
abstract
Crowd counting is a challenging task in the presence of drastic scale variations, the clutter background, and severe occlusions, etc. Existing CNN-based counting methods tackle these challenges mainly by fusing either multi-scale or multi-context features to generate robust representations. In this paper, we propose to address these issues by leveraging the heterogeneous attributes compounded in the density map. We identify three geometric/semantic/numeric attributes essentially important to the density estimation, and demonstrate how to effectively utilize these heterogeneous attributes to assist the crowd counting by formulating them into multiple auxiliary tasks. With the multi-fold regularization effects induced by the auxiliary tasks, the backbone CNN model is driven to embed desired properties explicitly and thus gains robust representations towards more accurate density estimation. Extensive experiments on three challenging crowd counting datasets have demonstrated the effectiveness of the proposed approach.
Muming Zhao, Jian Zhang 0002, Wenjun Zhang 0001
CVPR2
2019 Bi-Level Masked Multi-scale CNN-RNN Networks for Short Text Representation
abstract
Representing short text is becoming extremely important for a variety of valuable applications. However, representing short text is critical yet challenging because it involves lots of informal words and typos (i.e. the noise problem) but only few vocabularies in each text (i.e. the sparsity problem). Most of existing work on representing short text relies on noise recognition and sparsity expansion. However, the noises in short text are with various forms and changing fast, but, most of the current methods may fail to adaptively recognize the noise. Also, it is hard to explicitly expand a sparse text to a high-quality dense text. In this paper, we tackle the noise and sparsity problems in short text representation by learning multi-grain noise-tolerant patterns and then embedding the most significant patterns in a text as its representation. To achieve this goal, we propose a bi-level multi-scale masked CNN-RNN network to embed the most significant multi-grain noise-tolerant relations among words and characters in a text into a dense vector space. Comprehensive experiments on five large real-world data sets demonstrate our method significantly outperforms the state-of-the-art competitors.
Qian Li 0006, Qiang Wu 0001, Chengzhang Zhu, Jian Zhang 0002
ICDAR4
2019 Kpsnet: Keypoint Detection and Feature Extraction for Point Cloud Registration
abstract
This paper presents the KPSNet, a KeyPoint Siamese Network to simultaneously learn task-desirable keypoint detector and feature extractor. The keypoint detector is optimized to predict a score vector, which signifies the probability of each candidate being a keypoint. The feature extractor is optimized to learn robust features of keypoints by exploiting the correspondence between the keypoints generated from two inputs, respectively. For training, the KPSNet does not require to manually annotate keypoints and local patches pairwise. Instead, we design an alignment module to establish the correspondence between the two inputs and generate positive and negative samples on-the-fly. Therefore, our method can be easily extended to new scenes. We test the proposed method on the open-source benchmark and experiments show the validity of our method.
Anan Du, Xiaoshui Huang, Jian Zhang 0002, Lingxiang Yao, Qiang Wu 0001
ICIP3
2019 Fast Registration for Cross-Source Point Clouds by using Weak Regional Affinity and Pixel-Wise Refinement
abstract
Many types of 3D acquisition sensors have emerged in recent years and point cloud has been widely used in many areas. Accurate and fast registration of cross-source 3D point clouds from different sensors is an emerged research problem in computer vision. This problem is extremely challenging because cross-source point clouds contain a mixture of various variances, such as density, partial overlap, large noise and outliers, viewpoint changing. In this paper, an algorithm is proposed to align cross-source point clouds with both high accuracy and high efficiency. There are two main contributions: firstly, two components, the weak region affinity and pixel-wise refinement, are proposed to maintain the global and local information of 3D point clouds. Then, these two components are integrated into an iterative tensor-based registration algorithm to solve the cross-source point cloud registration problem. We conduct experiments on a synthetic cross-source benchmark dataset and real cross-source datasets. Comparison with six state-of-the-art methods, the proposed method obtains both higher efficiency and accuracy.
Xiaoshui Huang, Lixin Fan, Qiang Wu 0001, Jian Zhang 0002, Chun Yuan 0003
ICME4
2019 Compare More Nuanced: Pairwise Alignment Bilinear Network for Few-Shot Fine-Grained Learning
abstract
The recognition ability of human beings is developed in a progressive way. Usually, children learn to discriminate various objects from coarse to fine-grained with limited supervision. Inspired by this learning process, we propose a simple yet effective model for the Few-Shot Fine-Grained (FSFG) recognition, which tries to tackle the challenging fine-grained recognition task using meta-learning. The proposed method, named Pairwise Alignment Bilinear Network (PABN), is an end-to-end deep neural network. Unlike traditional deep bilinear networks for fine-grained classification, which adopt the self-bilinear pooling to capture the subtle features of images, the proposed model uses a novel pairwise bilinear pooling to compare the nuanced differences between base images and query images for learning a deep distance metric. In order to match base image features with query image features, we design feature alignment losses before the proposed pairwise bilinear pooling. Experiment results on four fine-grained classification datasets and one generic few-shot dataset demonstrate that the proposed model outperforms both the state-of-the-art few-shot fine-grained and general few-shot methods.
Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Jingsong Xu
ICME3
2019 The Influence of Different Urban and Rural Selection Methods on the Spatial Variation of Urban Heat Island Intensity
abstract
Few studies have compared the effects of different urban-rural selection methods on estimating urban heat island intensity (SUHII) spatial variation. This study aims to analyze the influence of different urban and rural selection methods (Whether water and elevation should be excluded, whether rural areas should be restricted, and whether urban-rural fringe should be considered) on the spatial variation of SUHII, to take 32 cities in China as an example. The main findings include: (1) Comparison of different methods, the SUHII calculated by these methods that include water bodies and elevations, unrestricted rural areas, and rural areas that exclude urban-rural fringe will be larger.(2) Estimates of SUHII in 32 cities are affected by different methods differently.(3)When using different methods to analyze SUHII in four types of humid and dry areas. The order of the large and small SUHII in the four regions obtained is different. Except for the arid regions, the other three regions showed the same regular characteristics as the conclusion (1). However, in the arid regions, SUHII obtained by these methods that unrestricted rural areas and excluded urban-rural fringe is smaller.(4)In the analysis of the seven geographical divisions. The relative size of SUHII in different geographical regions is different by the influence of the method, and the difference is obvious. This study compares the six urban and rural selection methods to find that different methods have an important impact on the spatial variation of SUHII. Therefore, in the study of SUHII, the emphasis on the selection of urban and rural partition methods should be improved.
Weiqi Zhou, Wenbo Xu 0004, Jian Zhang 0002
IGARSS5
2019 An Inferable Representation Learning for Fraud Review Detection with Cold-start Problem
abstract
Fraud review significantly damages the business reputation and also customers' trust to certain products. It has become a serious problem existing on the current social media. Various efforts have been put in to tackle such problems. However, in the case of cold-start where a review is posted by a new user who just pops up on the social media, common fraud detection methods may fail because most of them are heavily depended on the information about the user's historical behavior and its social relation to other users, yet such information is lacking in the cold-start case. This paper presents a novel Joint-bEhavior-and-Social-relaTion-infERable (JESTER) embedding method to leverage the user reviewing behavior and social relations for cold-start fraud review detection. JESTER embeds the deep characteristics of existing user behavior and social relations of users and items in an inferable user-item-review-rating representation space where the representation of a new user can be efficiently inferred by a closed-form solution and reflects the user's most probable behavior and social relations. Thus, a cold-start fraud review can be effectively detected accordingly. Our experiments show JESTER (i) performs significantly better in detecting fraud reviews on four real-life social media data sets, and (ii) effectively infers new user representation in the cold-start problem, compared to three state-of-the-art and two baseline competitors.
Qian Li 0006, Qiang Wu 0001, Chengzhang Zhu, Jian Zhang 0002
IJCNN4
2019 Sample Adaptive Multiple Kernel Learning for Failure Prediction of Railway Points
abstract
Railway points are among the key components of railway infrastructure. As a part of signal equipment, points control the routes of trains at railway junctions, having a significant impact on the reliability, capacity, and punctuality of rail transport. Meanwhile, they are also one of the most fragile parts in railway systems. Points failures cause a large portion of railway incidents. Traditionally, maintenance of points is based on a fixed time interval or raised after the equipment failures. Instead, it would be of great value if we could forecast points' failures and take action beforehand, minimising any negative effect. To date, most of the existing prediction methods are either lab-based or relying on specially installed sensors which makes them infeasible for large-scale implementation. Besides, they often use data from only one source. We, therefore, explore a new way that integrates multi-source data which are ready to hand to fulfil this task. We conducted our case study based on Sydney Trains rail network which is an extensive network of passenger and freight railways. Unfortunately, the real-world data are usually incomplete due to various reasons, e.g., faults in the database, operational errors or transmission faults. Besides, railway points differ in their locations, types and some other properties, which means it is hard to use a unified model to predict their failures. Aiming at this challenging task, we firstly constructed a dataset from multiple sources and selected key features with the help of domain experts. In this paper, we formulate our prediction task as a multiple kernel learning problem with missing kernels. We present a robust multiple kernel learning algorithm for predicting points failures. Our model takes into account the missing pattern of data as well as the inherent variance on different sets of railway points. Extensive experiments demonstrate the superiority of our algorithm compared with other state-of-the-art methods.
Zhibin Li 0002, Jian Zhang 0002, Qiang Wu 0001, Yongshun Gong, Jinfeng Yi, Christina Kirsch
KDD2
2019 Unsupervised User Behavior Representation for Fraud Review Detection with Cold-Start Problem
Qian Li 0006, Qiang Wu 0001, Chengzhang Zhu, Jian Zhang 0002
PAKDD (1)4
2019 Learn Image Object Co-segmentation with Multi-scale Feature Fusion
abstract
Image object co-segmentation aims to segment common objects in a group of images. This paper proposes a novel neural network, which extracts multi-scale convolutional features at multiple layers via a modified VGG network and fuses them both within and across images as the intra-image and the inter-image features. Then these two kinds of features are further fused at each scale as the multi-scale co-features of common objects, and finally the multi-scale co-features are summed up and upsampled to obtain the co-segmentation results. To simplify the network and reduce the rapidly rising resource cost along with the inputs, the reduced input size, less downsampling and dilation convolution are adopted in the proposed model. Experimental results on the public dataset demonstrate that the proposed model achieves a comparable performance to the state-of-the-art co-segmentation methods while the computation cost has been effectively reduced.
Zhi Liu 0003, Jian Zhang 0002, Xiaofei Zhou 0003
VCIP3
2019 Scale-Informed Density Estimation for Dense Crowd Counting
abstract
Dense crowd counting (DCC) remains challenging due to the scale variation and occlusion. Several deep learning based DCC methods have achieved the state-of-arts on public datasets. However, experimental results show that the scale variation is still the main factor to hinder the DCC performance. In this work, we propose a scale-informed dense crowd counting method focusing on handling the negative effect caused by scale variation. More specifically, we propose a method to obtain the scale information of the patch from its GT density maps via estimating the mean value of the Gaussian kernel width and then a scale-classifier is deigned and trained accordingly. Moreover, with the estimated scale information, two sub-nets are dedicatedly deigned to learn the density maps for large-scale head patch and small-scale patch separately. Experimental results validate the performance of our proposed method which achieves the best performance on three dense crowd datasets.
Yuexian Zou, Guoshuai Wang, Jian Zhang 0002
VCIP4
2019 Jointly learning perceptually heterogeneous features for blind 3D video quality assessment
Shuai Yuan 0008, Yun Zhu 0002, Jian Zhang 0002, Ping An 0001
Neurocomputing4
2019 C-RPNs: Promoting object detection in real world via a cascade structure of Region Proposal Networks
Dongming Yang, Yuexian Zou, Jian Zhang 0002, Ge Li 0002
Neurocomputing3
2019 Heritage image annotation via collective knowledge
Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003, Qiang Wu 0001
Pattern Recognit.3
2019 Robust Distracter-Resistive Tracker via Learning a Multi-Component Discriminative Dictionary
abstract
Discriminative dictionary learning (DDL) provides an appealing paradigm for appearance modeling in visual tracking. However, most existing DDL-based trackers cannot handle drastic appearance changes, especially for scenarios with background cluster and/or similar object interference. One reason is that they often suffer from the loss of subtle visual information, which is critical to distinguish an object from distracters. In this paper, we explore the use of activations from the convolutional layer of a convolutional neural network to improve the object representation and then propose a robust distracter-resistive tracker via learning a multi-component discriminative dictionary. The proposed method exploits both the intra-class and inter-class visual information to learn shared atoms and the class-specific atoms. By imposing several constraints into the objective function, the learned dictionary is reconstructive, compressive, and discriminative, and thus can better distinguish an object from the background. In addition, our convolutional features have structural information for object localization and balance the discriminative power and semantic information of the object. Tracking is carried out within a Bayesian inference framework where a joint decision measure is used to construct the observation model. To alleviate the drift problem, the reliable tracking results obtained online are accumulated to update the dictionary. Both the qualitative and quantitative results on the CVPR2013 benchmark, the VOT2015 data set, and the SPOT data set demonstrate that our tracker achieves substantially better overall performance against the state-of-the-art approaches.
Weichao Shen, Yuwei Wu 0001, Junsong Yuan 0001, Ling-Yu Duan, Jian Zhang 0002, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.5
2019 Multi-Pseudo Regularized Label for Generated Data in Person Re-Identification
abstract
Sufficient training data normally is required to train deeply learned models. However, due to the expensive manual process for labelling large number of images (i.e., annotation), the amount of available training data (i.e., real data) is always limited. To produce more data for training a deep network, Generative Adversarial Network (GAN) can be used to generate artificial sample data (i.e., generated data). However, the generated data usually does not have annotation labels. To solve this problem, in this paper, we propose a virtual label called Multi-pseudo Regularized Label (MpRL) and assign it to the generated data. With MpRL, the generated data will be used as the supplementary of real training data to train a deep neural network in a semi-supervised learning fashion. To build the corresponding relationship between the real data and generated data, MpRL assigns each generated data a proper virtual label which reflects the likelihood of the affiliation of the generated data to predefined training classes in the real data domain. Unlike the traditional label which usually is a single integral number, the virtual label proposed in this work is a set of weight-based values each individual of which is a number in (0,1] called multi-pseudo label and reflects the degree of relation between each generated data to every pre-defined class of real data. A comprehensive evaluation is carried out by adopting two state-of-the-art convolutional neural networks (CNNs) in our experiments to verify the effectiveness of MpRL. Experiments demonstrate that by assigning MpRL to generated data, we can further improve the person re-ID performance on five re-ID datasets, i.e., Market-1501, DukeMTMC-reID, CUHK03, VIPeR, and CUHK01. The proposed method obtains +6.29%, +6.30%, +5.58%, +5.84%, and +3.48% improvements in rank-1 accuracy over a strong CNN baseline on the five datasets respectively, and outperforms state-of-the-art methods.
Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Zhedong Zheng, Zhaoxiang Zhang 0001, Jian Zhang 0002
IEEE Trans. Image Process.6
2019 Extracting Privileged Information for Enhancing Classifier Learning
abstract
The accuracy of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), e.g., attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. Moreover, due to the limitations of personal knowledge, manually labeled PI may not be rich enough. To address these issues, we propose to enhance classifier learning by exploring PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data and obtain much richer PI. In detail, we treat each selected PI as a subcategory and learn one classifier for each subcategory independently. The classifiers for all subcategories are integrated together to form a more powerful category classifier. Particularly, we propose a novel instancelevel multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal SVM classifiers based on the selected images. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001
IEEE Trans. Image Process.3
2019 A Computational Model for Stereoscopic Visual Saliency Prediction
abstract
Depth information plays an important role in human vision as it provides additional cues that distinguish objects from their backgrounds. This paper explores depth information for analyzing stereoscopic saliency and presents a computational model that predicts stereoscopic visual saliency based on three aspects of human vision: 1) the pop-out effect; 2) comfort zones; and 3) background effects. Through an analysis of these three phenomena, we find that most of the stereoscopic saliency region can be explained. Our model comprises three modules, each describing one aspect of saliency distribution, and a control function that can be used to adjust the three models independently. The relationship between the three models is not mutually exclusive. One, two, or three phenomena may appear in one image. Therefore, to accurately determine which phenomena the image conforms to, we have devised a selection strategy that chooses the appropriate combination of models based on the content of the image. Our approach is implemented within a framework based on the multifeature analysis. The framework considers surrounding regions, color/depth contrast, and points of interest. The selection strategy can improve the performance of the framework. A series of experiments on two recent eye-tracking datasets shows that our proposed method outperforms several state-of-the-art saliency models.
Hao Cheng 0003, Jian Zhang 0002, Qiang Wu 0001, Ping An 0001
IEEE Trans. Multim.2
2019 Feature Affinity-Based Pseudo Labeling for Semi-Supervised Person Re-Identification
abstract
Vision-based person re-identification aims to match a person's identity across multiple images, which is a fundamental task in multimedia content analysis and retrieval. Deep neural networks have recently manifested great potential in this task. However, a major bottleneck of existing supervised deep networks is their reliance on a large amount of annotated training data. Manual labeling for person identities in large-scale surveillance camera systems is quite challenging and incurs significant costs. Some recent studies adopt generative model outputs as training data augmentation. To more effectively use these synthetic data for an improved feature learning and re-identification performance, this paper proposes a novel feature affinity-based pseudo labeling method with two possible label encodings. To the best of our knowledge, this is the first study that employs pseudo-labeling by measuring the affinity of unlabeled samples with the underlying clusters of labeled data samples using the intermediate feature representations from deep networks. We propose training the network with the joint supervision of cross-entropy loss together with a center regularization term, which not only ensures discriminative feature representation learning but also simultaneously predicts pseudo-labels for unlabeled data. We show that both label encodings can be learned in a unified manner and help improve the overall performance. Our extensive experiments on three person re-identification datasets: Market-1501, DukeMTMC-reID, and CUHK03, demonstrate significant performance boost over the state-of-the-art person re-identification approaches.
Guodong Ding, Shanshan Zhang 0001, Salman Khan 0001, Zhenmin Tang, Jian Zhang 0002, Fatih Porikli
IEEE Trans. Multim.5
2019 Extracting Multiple Visual Senses for Web Learning
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time consuming and labor intensive. To reduce the dependence on manually labeled data, there have been increasing research efforts on learning visual classifiers by directly exploiting web images. One issue that limits their performance is the problem of polysemy. Existing unsupervised approaches attempt to reduce the influence of visual polysemy by filtering out irrelevant images, but do not directly address polysemy. To this end, in this paper, we present a multimodal framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses from untagged corpora to retrieve sense-specific images. Then, we merge visual similar semantic senses and prune noise by using the retrieved images. Finally, we train one visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and reranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001
IEEE Trans. Multim.3
2018 Discovering and Distinguishing Multiple Visual Senses for Polysemous Words
abstract
To reduce the dependence on labeled data, there have been increasing research efforts on learning visual classifiers by exploiting web images. One issue that limits their performance is the problem of polysemy. To solve this problem, in this work, we present a novel framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses to retrieve sense-specific images. Then we merge visual similar semantic senses and prune noises by using the retrieved images. Finally, we train a visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and re-ranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Zhenmin Tang
AAAI2
2018 Kill Two Birds With One Stone: Weakly-Supervised Neural Network for Image Annotation and Tag Refinement
abstract
The number of social images has exploded by the wide adoption of social networks, and people like to share their comments about them. These comments can be a description of the image, or some objects, attributes, scenes in it, which are normally used as the user-provided tags. However, it is well-known that user-provided tags are incomplete and imprecise to some extent. Directly using them can damage the performance of related applications, such as the image annotation and retrieval. In this paper, we propose to learn an image annotation model and refine the user-provided tags simultaneously in a weakly-supervised manner. The deep neural network is utilized as the image feature learning and backbone annotation model, while visual consistency, semantic dependency, and user-error sparsity are introduced as the constraints at the batch level to alleviate the tag noise. Therefore, our model is highly flexible and stable to handle large-scale image sets. Experimental results on two benchmark datasets indicate that our proposed model achieves the best performance compared to the state-of-the-art methods.
Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003
AAAI3
2018 Towards Locally Consistent Object Counting with Constrained Multi-stage Convolutional Neural Networks
Muming Zhao, Jian Zhang 0002, Wenjun Zhang 0001
ACCV (6)2
2018 Network-wide Crowd Flow Prediction of Sydney Trains via Customized Online Non-negative Matrix Factorization
abstract
Crowd Flow Prediction (CFP) is one major challenge in the intelligent transportation systems of the Sydney Trains Network. However, most advanced CFP methods only focus on entrance and exit flows at the major stations or a few subway lines, neglecting Crowd Flow Distribution (CFD) forecasting problem across the entire city network. CFD prediction plays an irreplaceable role in metro management as a tool that can help authorities plan route schedules and avoid congestion. In this paper, we propose three online non-negative matrix factorization (ONMF) models. ONMF-AO incorporates an Average Optimization strategy that adapts to stable passenger flows. ONMF-MR captures the Most Recent trends to achieve better performance when sudden changes in crowd flow occur. The Hybrid model, ONMF-H, integrates both ONMF-AO and ONMF-MR to exploit the strengths of each model in different scenarios and enhance the models' applicability to real-world situations. Given a series of CFD snapshots, both models learn the latent attributes of the train stations and, therefore, are able to capture transition patterns from one timestamp to the next by combining historic guidance. Intensive experiments on a large-scale, real-world dataset containing transactional data demonstrate the superiority of our ONMF models.
Yongshun Gong, Zhibin Li 0002, Jian Zhang 0002, Wei Liu 0007, Yu Zheng 0004, Christina Kirsch
CIKM3
2018 Goal-Oriented Visual Question Generation via Intermediate Rewards
Junjie Zhang 0002, Qi Wu 0001, Chunhua Shen, Jian Zhang 0002, Jianfeng Lu 0003, Anton van den Hengel
ECCV (5)4
2018 Extracting Privileged Information from Untagged Corpora for Classifier Learning
abstract
The performance of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), \eg attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. To address this issue, we propose to enhance classifier learning by extracting PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data. In detail, we treat each selected PI as a subcategory and learn one classifier for per subcategory independently. The classifiers for all subcategories are then integrated together to form a more powerful category classifier. Particularly, we propose a new instance-level multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal classifiers based on the selected images. Extensive experiments demonstrate the superiority of our approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Xian-Sheng Hua 0001, Zhenmin Tang
IJCAI2
2018 Siamese Network Based Features Fusion for Adaptive Visual Tracking
Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen, Jian Zhang 0002
PRICAI (1)6
2018 Long-Term Person Re-identification Using True Motion from Videos
abstract
Most person re-identification approaches and benchmarks assume that pedestrians go across the surveillance network without significant appearance changes in a brief period, which explicitly restricts person re-identification to a short-term event and incurs inter-sample similarity measurement by appearance matching. However, pedestrians are likely to reappear in the surveillance network after a long-time interval (long-term) and change their wearing in many real-world scenarios. These scenarios inevitably cause appearances between subjects more ambiguous and indistinguishable. In this paper we consider these scenarios and propose a unified feature representation based on true motion cues from videos named FIne moTion encoDing (FITD). Our hypothesis is that people keep constant motion patterns under non-distraction walking condition. Therefore, the motion characteristics are more reliable than static appearance feature to describe a walking person. Particularly, we extract motion patterns hierarchically by encoding trajectory-aligned descriptors with Fisher vectors in a spatial-aligned pyramid. To verify benefits of the proposed FITD, we collect a new dataset typically for the long-term situations. Extensive experiments demonstrate the merits of our FITD especially for the long-term scenarios.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
WACV4
2018 Unsupervised image co-segmentation via guidance of simple images
Zhi Liu 0003, Jian Zhang 0002
Neurocomputing3
2018 A Coarse-to-Fine Algorithm for Matching and Registration in 3D Cross-Source Point Clouds
abstract
We propose an efficient method to deal with the matching and registration problem found in cross-source point clouds captured by different types of sensors. This task is especially challenging due to the presence of density variation, scale difference, a large proportion of noise and outliers, missing data, and viewpoint variation. The proposed method has two stages: in the coarse matching stage, we use the ensemble of shape functions descriptor to select potential K regions from the candidate point clouds for the target. In the fine stage, we propose a scale embedded generative Gaussian mixture models registration method to refine the results from the coarse matching stage. Following the fine stage, both the best region and accurate camera pose relationships between the candidates and target are found. We conduct experiments in which we apply the method to two applications: one is 3D object detection and localization in street-view outdoor (LiDAR/VSFM) cross-source point clouds and the other is 3D scene matching and registration in indoor (KinectFusion/VSFM) cross-source point clouds. The experiment results show that the proposed method performs well when compared with the existing methods. It also shows that the proposed method is robust under various sensing techniques, such as LiDAR, Kinect, and RGB camera.
Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001, Lixin Fan, Chun Yuan 0003
IEEE Trans. Circuits Syst. Video Technol.2
2018 Explicit Edge Inconsistency Evaluation Model for Color-Guided Depth Map Enhancement
abstract
Color-guided depth enhancement is used to refine depth maps according to the assumption that the depth edges and the color edges at the corresponding locations are consistent. In methods on such low-level vision tasks, the Markov random field (MRF), including its variants, is one of the major approaches that have dominated this area for several years. However, the assumption above is not always true. To tackle the problem, the state-of-the-art solutions are to adjust the weighting coefficient inside the smoothness term of the MRF model. These methods lack an explicit evaluation model to quantitatively measure the inconsistency between the depth edge map and the color edge map, so they cannot adaptively control the efforts of the guidance from the color image for depth enhancement, leading to various defects such as texture-copy artifacts and blurring depth edges. In this paper, we propose a quantitative measurement on such inconsistency and explicitly embed it into the smoothness term. The proposed method demonstrates promising experimental results compared with the benchmark and state-of-the-art methods on the Middlebury ToF-Mark, and NYU data sets.
Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Depth Super-Resolution on RGB-D Video Sequences With Large Displacement 3D Motion
abstract
To enhance the resolution and accuracy of depth data, some video-based depth super-resolution methods have been proposed which utilizes its neighboring depth images in the temporal domain. They often consist of two main stages: motion compensation of temporally neighboring depth images and fusion of compensated depth images. However, large displacement 3D motion often leads to compensation error, and the compensation error is further introduced into the fusion. A video-based depth super-resolution method with novel motion compensation and fusion approaches is proposed in this paper. We claim that, 3D Nearest Neighboring Field (NNF) is a better choice than using positions with true motion displacement for depth enhancements. To handle large displacement 3D motion, the compensation stage utilized 3D NNF instead of true motion used in previous methods. Next, the fusion approach is modeled as a regression problem to predict the super-resolution result efficiently for each depth image by using its compensated depth images. A new deep convolutional neural network architecture is designed for fusion, which is able to employ a large amount of video data for learning the complicated regression function. We comprehensively evaluate our method on various RGB-D video sequences to show its superior performance.
Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Zhengyou Zhang, Yunde Jia
IEEE Trans. Image Process.2
2018 Minimum Spanning Forest With Embedded Edge Inconsistency Measurement Model for Guided Depth Map Enhancement
abstract
Guided depth map enhancement based on Markov Random Field (MRF) normally assumes edge consistency between the color image and the corresponding depth map. Under this assumption, the low-quality depth edges can be refined according to the guidance from the high-quality color image. However, such consistency is not always true, which leads to texture-copying artifacts and blurring depth edges. In addition, the previous MRF-based models always calculate the guidance affinities in the regularization term via a non-structural scheme which ignores the local structure on the depth map. In this paper, a novel MRF-based method is proposed. It computes these affinities via the distance between pixels in a space consisting of the Minimum Spanning Trees (Forest) to better preserve depth edges. Furthermore, inside each Minimum Spanning Tree, the weights of edges are computed based on explicit edge inconsistency measurement model, which significantly mitigates texture-copying artifacts. To further tolerate the effects caused by noise and better preserve depth edges, a bandwidth adaption scheme is proposed. Our method is evaluated for depth map super-resolution and depth map completion problems on synthetic and real datasets including Middlebury, ToF-Mark and NYU. A comprehensive comparison against 16 state-of-the-art methods is carried out. Both qualitative and quantitative evaluation present the improved performances.
Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001
IEEE Trans. Image Process.3
2018 Multilabel Image Classification With Regional Latent Semantic Dependencies
abstract
Deep convolution neural networks (CNNs) have demonstrated advanced performance on single-label image classification, and various progress also has been made to apply CNN methods on multilabel image classification, which requires annotating objects, attributes, scene categories, etc., in a single shot. Recent state-of-the-art approaches to the multilabel image classification exploit the label dependencies in an image, at the global level, largely improving the labeling capacity. However, predicting small objects and visual concepts is still challenging due to the limited discrimination of the global visual features. In this paper, we propose a regional latent semantic dependencies model (RLSD) to address this problem. The utilized model includes a fully convolutional localization architecture to localize the regions that may contain multiple highly dependent labels. The localized regions are further sent to the recurrent neural networks to characterize the latent semantic dependencies at the regional level. Experimental results on several benchmark datasets show that our proposed model achieves the best performance compared to the state-of-the-art models, especially for predicting small objects occurring in the images. Also, we set up an upper bound model (RLSD+ft-RPN) using bounding-box coordinates during training, and the experimental results also show that our RLSD can approach the upper bound without using the bounding-box annotations, which is more realistic in the real world.
Junjie Zhang 0002, Qi Wu 0001, Chunhua Shen, Jian Zhang 0002, Jianfeng Lu 0003
IEEE Trans. Multim.4
2018 Learning deep facial expression features from image and optical flow sequences using 3D CNN
Jianfeng Zhao 0005, Xia Mao, Jian Zhang 0002
Vis. Comput.3
2017 Deep learning for robust outdoor vehicle visual tracking
abstract
Robust visual tracking for outdoor vehicle is still a challenging problem due to large appearance variations caused by illumination variation, occlusion and scale variation, etc. In this paper, a deep-learning-based approach for robust outdoor vehicle tracking is proposed. Firstly, a stacked denoising auto-encoder is pre-trained to learn the feature representation way of images. Then, a k-sparse constraint is added to the stacked denoising auto-encoder and the encoder of k-sparse stacked denoising auto-encoder (kSSDAE) is connected with a classification layer to construct a classification neural network. After fine-tuning, the classification neural network is applied to online tracking under particle filter framework. Extensive tracking experiments are conducted on a challenging single object online tracking evaluation platform benchmark to verify the effectiveness of our tracker. Experiments show that our tracker outperforms most state-of-the-art trackers.
Jing Xin, Jian Zhang 0002
ICME3
2017 Learning a perspective-embedded deconvolution network for crowd counting
abstract
We present a novel deep learning framework for crowd counting by learning a perspective-embedded deconvolution network. Perspective is an inherent property of most surveillance scenes. Unlike the traditional approaches that exploit the perspective as a separate normalization, we propose to fuse the perspective into a deconvolution network, aiming to obtain a robust, accurate and consistent crowd density map. Through layer-wise fusion, we merge perspective maps at different resolutions into the deconvolution network. With the injection of perspective, our network is driven to learn to combine the underlying scene geometric constraints adaptively, thus enabling an accurate interpretation from high-level feature maps to the pixel-wise crowd density map. In addition, our network allows generating density map for arbitrary-sized input in an end-to-end fashion. The proposed method achieves competitive result on the WorldExpo2010 crowd dataset.
Muming Zhao, Jian Zhang 0002, Fatih Porikli, Wenjun Zhang 0001
ICME2
2017 Minimum spanning forest with embedded edge inconsistency measurement for color-guided depth map upsampling
abstract
Color-guided depth map up-sampling, such as Markov-Random-Field-based (MRF-based) methods, is a popular depth map enhancement solution, which normally assumes edge consistency between color image and corresponding depth map. It calculates the coefficients of smoothness term in MRF according to such assumption. However, such consistency is not always true which leads to texture-copying artifacts and blurring depth edges. In this paper, we propose a novel coefficient computing scheme for smoothness term in MRF which is based on the distance between pixels in the Minimum Spanning Trees (Forest) to better preserve depth edges. The explicit edge inconsistency measurement is embedded into weights of edges in Minimum Spanning Trees, which significantly mitigates texture-copying artifacts. The proposed method is evaluated on Middlebury datasets and ToF-Mark datasets which demonstrates improved results compared with state-of-the-art methods.
Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001
ICME3
2017 RGB-D Tracking Based on Kernelized Correlation Filter with Deep Features
Yao Lu 0001, Lin Zhang 0033, Jian Zhang 0002
ICONIP (3)4
2017 User relationship strength modeling for friend recommendation on Instagram
Dongyan Guo, Jingsong Xu, Jian Zhang 0002, Min Xu 0001, Xiangjian He
Neurocomputing3
2017 A new web-supervised method for image dataset constructions
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
Neurocomputing2
2017 Region-based Mixture Models for human action recognition in low-resolution videos
Huijun Di, Jian Zhang 0002, Yao Lu 0001, Feng Lv
Neurocomputing3
2017 A Systematic Approach for Cross-Source Point Cloud Registration by Preserving Macro and Micro Structures
abstract
We propose a systematic approach for registering cross-source point clouds that come from different kinds of sensors. This task is especially challenging due to the presence of significant missing data, large variations in point density, scale difference, large proportion of noise, and outliers. The robustness of the method is attributed to the extraction of macro and micro structures. Macro structure is the overall structure that maintains similar geometric layout in cross-source point clouds. Micro structure is the element (e.g., local segment) being used to build the macro structure. We use graph to organize these structures and convert the registration into graph matching. With a novel proposed descriptor, we conduct the graph matching in a discriminative feature space. The graph matching problem is solved by an improved graph matching solution, which considers global geometrical constraints. Robust cross source registration results are obtained by incorporating graph matching outcome with RANSAC and ICP refinements. Compared with eight state-of-the-art registration algorithms, the proposed method invariably outperforms on Pisa Cathedral and other challenging cases. In order to compare quantitatively, we propose two challenging cross-source data sets and conduct comparative experiments on more than 27 cases, and the results show we obtain much better performance than other methods. The proposed method also shows high accuracy in same-source data sets.
Xiaoshui Huang, Jian Zhang 0002, Lixin Fan, Qiang Wu 0001, Chun Yuan 0003
IEEE Trans. Image Process.2
2017 Two-Stage Friend Recommendation Based on Network Alignment and Series Expansion of Probabilistic Topic Model
abstract
Precise friend recommendation is an important problem in social media. Although most social websites provide some kinds of auto friend searching functions, their accuracies are not satisfactory. In this paper, we propose a more precise auto friend recommendation method with two stages. In the first stage, by utilizing the information of the relationship between texts and users, as well as the friendship information between users, we align different social networks and choose some “possible friends.” In the second stage, with the relationship between image features and users, we build a topic model to further refine the recommendation results. Because some traditional methods, such as variational inference and Gibbs sampling, have their limitations in dealing with our problem, we develop a novel method to find out the solution of the topic model based on series expansion. We conduct experiments on the Flickr dataset to show that the proposed algorithm recommends friends more precisely and faster than traditional methods.
Shangrong Huang, Jian Zhang 0002, Dan Schonfeld, Lei Wang 0001, Xian-Sheng Hua 0001
IEEE Trans. Multim.2
2017 Exploiting Web Images for Dataset Construction: A Domain Robust Approach
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time-consuming and labor intensive. To reduce the cost of manual labeling, there has been increased research interest in automatically constructing image datasets by exploiting web images. Datasets constructed by existing methods tend to have a weak domain adaptation ability, which is known as the “dataset bias problem.” To address this issue, we present a novel image dataset construction framework that can be generalized well to unseen target domains. Specifically, the given queries are first expanded by searching the Google Books Ngrams Corpus to obtain a rich semantic description, from which the visually nonsalient and less relevant expansions are filtered out. By treating each selected expansion as a “bag” and the retrieved images as “instances,” image selection can be formulated as a multi-instance learning problem with constrained positive bags. We propose to solve the employed problems by the cutting-plane and concave-convex procedure algorithm. By using this approach, images from different distributions can be kept while noisy images are filtered out. To verify the effectiveness of our proposed approach, we build an image dataset with 20 categories. Extensive experiments on image classification, cross-dataset generalization, diversity comparison, and object detection demonstrate the domain robustness of our dataset.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
IEEE Trans. Multim.2
2016 Explicit measurement on depth-color inconsistency for depth completion
abstract
Color-guided depth completion is to refine depth map through structure light sensing by filling missing depth structure and de-nosing. It is based on the assumption that depth discontinuity and color edge at the corresponding location are consistent. Among all proposed methods, MRF-based method including its variants is one of major approaches. However, the assumption above is not always true, which causes texture-copy and depth discontinuity blurring artifacts. The state-of-the-art solutions usually are to modify the weighting inside smoothness term of MRF model. Because there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge, they cannot adaptively control the effect of guidance from color image when completing depth map. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. The proposed method is evaluated on NYU Kinect datasets and demonstrates promising results.
Yifan Zuo 0001, Qiang Wu 0001, Ping An 0001, Jian Zhang 0002
ICIP4
2016 Unsupervised visual domain adaptation via dictionary evolution
abstract
In real-word visual applications, distribution mismatch between samples from different domains may significantly degrade classification performance. To improve the generalization capability of classifier across domains, domain adaptation has attracted a lot of interest in computer vision. This work focuses on unsupervised domain adaptation which is still challenging because no labels are available in the target domain. Most of the attention has been dedicated to seeking domain-invariant feature by exploring the shared structure between domains, ignoring the valuable discriminative information contained in the labeled source data. In this paper, we propose a Dictionary Evolution (DE) approach to construct discriminative features robust to domain shift. Specifically, DE aims to adapt a discriminative dictionary learnt based on labeled source samples to unlabeled target samples through a gradual transition process. We show that the learnt dictionary is endowed with cross-domain data representation ability and powerful discriminant capability. Empirical results on real world data sets demonstrate the advantages of the proposed approach over competing methods.
Songsong Wu, Xiaoyuan Jing, Dong Yue 0001, Jian Zhang 0002, K. Jian Yang, Jing-Yu Yang 0001
ICME4
2016 Automatic image dataset construction with multiple textual metadata
abstract
The goal of this work is to automatically collect a large number of highly relevant images from the Internet for given queries. A novel image dataset construction framework is proposed by employing multiple textual metadata. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora to obtain a richer semantic description, from which the visually non-salient and less relevant expansions are then filtered. After retrieving images from the Internet with filtered expansions, we further filter noisy images by clustering and progressively Convolutional Neural Networks (CNN). To verify the effectiveness of our proposed method, we construct a dataset with 10 categories, which is not only much larger than but also have comparable cross-dataset generalization ability with manually labeled dataset STL-10 and CIFAR-10.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
ICME2
2016 Recognizing human actions from low-resolution videos by region-based mixture models
abstract
Recognizing human action from low-resolution (LR) videos is essential for many applications including large-scale video surveillance, sports video analysis and intelligent aerial vehicles. Currently, state-of-the-art performance in action recognition is achieved by the use of dense trajectories which are extracted by optical flow algorithms. However, the optical flow algorithms are far from perfect in LR videos. In addition, the spatial and temporal layout of features is a powerful cue for action discrimination. While, most existing methods encode the layout by previously segmenting body parts which is not feasible in LR videos. Addressing the problems, we adopt the Layered Elastic Motion Tracking (LEMT) method to extract a set of long-term motion trajectories and a long-term common shape from each video sequence, where the extracted trajectories are much denser than those of sparse interest points(SIPs); then we present a hybrid feature representation to integrate both of the shape and motion features; and finally we propose a Region-based Mixture Model (RMM) to be utilized for action classification. The RMM models the spatial layout of features without any needs of body parts segmentation. Experiments are conducted on two publicly available LR human action datasets. Among which, the UT-Tower dataset is very challenging because the average height of human figures is only about 20 pixels. The proposed approach attains near-perfect accuracy on both of the datasets.
Huijun Di, Jian Zhang 0002, Yao Lu 0001, Feng Lv
ICME3
2016 Video object segmentation aggregation
abstract
We present an approach for unsupervised object segmentation in unconstrained videos. Driven by the latest progress in this field, we argue that segmentation performance can be largely improved by aggregating the results generated by state-of-the-art algorithms. Initially, objects in individual frames are estimated through a per-frame aggregation procedure using majority voting. While this can predict relatively accurate object location, the initial estimation fails to cover the parts that are wrongly labeled by more than half of the algorithms. To address this, we build a holistic appearance model using non-local appearance cues by linear regression. Then, we integrate the appearance priors and spatio-temporal information into an energy minimization framework to refine the initial estimation. We evaluate our method on challenging benchmark videos and demonstrate that it outperforms state-of-the-art algorithms.
Tianfei Zhou, Yao Lu 0001, Huijun Di, Jian Zhang 0002
ICME4
2016 Explicit modeling on depth-color inconsistency for color-guided depth up-sampling
abstract
Color-guided depth up-sampling is to enhance the resolution of depth map according to the assumption that the depth discontinuity and color image edge at the corresponding location are consistent. Through all methods reported, MRF including its variants is one of major approaches, which has dominated in this area for several years. However, the assumption above is not always true. Solution usually is to adjust the weighting inside smoothness term in MRF model. But there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. Such solution has not been reported in the literature. The improved depth up-sampling based on the proposed method is evaluated on Middlebury datasets and ToFMark datasets and demonstrate promising results.
Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001
ICME3
2016 A Domain Robust Approach For Image Dataset Construction
abstract
There have been increasing research interests in automatically constructing image dataset by collecting images from the Internet. However, existing methods tend to have a weak domain adaptation ability, known as the "dataset bias problem". To address this issue, in this work, we propose a novel image dataset construction framework which can generalize well to unseen target domains. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora (GBNC) to obtain a richer semantic description, from which the noisy query expansions are then filtered out. By treating each expansion as a "bag" and the retrieved images therein as "instances", we formulate image filtering as a multi-instance learning (MIL) problem with constrained positive bags. By this approach, images from different data distributions will be kept while with noisy images filtered out. Comprehensive experiments on two challenging tasks demonstrate the effectiveness of our proposed approach.
Yazhou Yao, Xian-Sheng Hua 0001, Fumin Shen, Jian Zhang 0002, Zhenmin Tang
ACM Multimedia4
2016 Extracting Visual Knowledge from the Internet: Making Sense of Image Data
Yazhou Yao, Jian Zhang 0002, Xian-Sheng Hua 0001, Fumin Shen, Zhenmin Tang
MMM (1)2
2016 Special Issue on Individual and Group Activities in Video Event Analysis
Liang Wang 0001, Ioannis Patras, Jian Zhang 0002, Greg Mori, Larry Davis 0001
Comput. Vis. Image Underst.3
2016 Saliency Detection Via Similar Image Retrieval
abstract
This letter proposes a novel saliency detection framework by propagating saliency of similar images retrieved from large and diverse Internet image collections to boost saliency detection performance effectively. For the input image, a group of similar images is retrieved based on the saliency weighted color histograms and the Gist descriptor from Internet image collections. Then, a pixel-level correspondence process between images is performed to guide the saliency propagation from the retrieved images. Both initial saliency map and correspondence saliency map are exploited to select the training samples by using the graph cut-based segmentation. Finally, the training samples are input into a set of weak classifiers to learn the boosted classifier for generating the boosted saliency map, which is integrated with the initial saliency map to generate the final saliency map. Experimental results on two public image datasets demonstrate that the proposed model can achieve the better saliency detection performance than the state-of-the-art single-image saliency models and co-saliency models.
Linwei Ye, Zhi Liu 0003, Xiaofei Zhou 0003, Liquan Shen, Jian Zhang 0002
IEEE Signal Process. Lett.5
2016 Handling Occlusion and Large Displacement Through Improved RGB-D Scene Flow Estimation
abstract
The accuracy of scene flow is restricted by several challenges such as occlusion and large displacement motion. When occlusion happens, the positions inside the occluded regions lose their corresponding counterparts in preceding and succeeding frames. Large displacement motion will increase the complexity of motion modeling and computation. Moreover, occlusion and large displacement motion are highly related problems in scene flow estimation, e.g., large displacement motion often leads to considerably occluded regions in the scene. An improved dense scene flow method based on red-green-blue-depth (RGB-D) data is proposed in this paper. To handle occlusion, we model the occlusion status for each point in our problem formulation, and jointly estimate the scene flow and occluded regions. To deal with large displacement motion, we employ an over-parameterized scene flow representation to model both the rotation and translation components of the scene flow, since large displacement motion cannot be well approximated using translational motion only. Furthermore, we employ a two-stage optimization procedure for this overparameterized scene flow representation. In the first stage, we propose a new RGB-D PatchMatch method, which is mainly applied in the RGB-D image space to reduce the computational complexity introduced by the large displacement motion. According to the quantitative evaluation based on the Middlebury data set, our method outperforms other published methods. The improved performance is also comprehensively confirmed on the real data acquired by Kinect sensor.
Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Philip A. Chou, Zhengyou Zhang, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.2
2016 Social Friend Recommendation Based on Multiple Network Correlation
abstract
Friend recommendation is an important recommender application in social media. Major social websites such as Twitter and Facebook are all capable of recommending friends to individuals. However, most of these websites use simple friend recommendation algorithms such as similarity, popularity, or “friend's friends are friends,” which are intuitive but consider few of the characteristics of the social network. In this paper we investigate the structure of social networks and develop an algorithm for network correlation-based social friend recommendation (NC-based SFR). To accomplish this goal, we correlate different “social role” networks, find their relationships and make friend recommendations. NC-based SFR is characterized by two key components: 1) related networks are aligned by selecting important features from each network, and 2) the network structure should be maximally preserved before and after network alignment. After important feature selection has been made, we recommend friends based on these features. We conduct experiments on the Flickr network, which contains more than ten thousand nodes and over 30 thousand tags covering half a million photos, to show that the proposed algorithm recommends friends more precisely than reference methods.
Shangrong Huang, Jian Zhang 0002, Lei Wang 0001, Xian-Sheng Hua 0001
IEEE Trans. Multim.2
2015 Absent Multiple Kernel Learning
abstract
Multiple kernel learning (MKL) optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels are missing, which is common in practical applications. This paper proposes an absent MKL (AMKL) algorithm to address this issue. Different from existing approaches where missing channels are firstly imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithm directly classifies each sample with its observed channels. In specific, we define a margin for each sample in its own relevant space, which corresponds to the observed channels of that sample. The proposed AMKL algorithm then maximizes the minimum of all sample-based margins, and this leads to a difficult optimization problem. We show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. Extensive experiments are conducted on five MKL benchmark data sets to compare the proposed algorithm with existing imputation-based methods. As observed, our algorithm achieves superior performance and the improvement is more significant with the increasing missing ratio.
Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, Yong Dou, Jian Zhang 0002
AAAI5
2015 A Survey of Multiple Sequence Alignment Techniques
Xiao-Dan Wang, Jin-Xing Liu 0001, Yong Xu 0001, Jian Zhang 0002
ICIC (1)4
2015 Social Friend Recommendation Based on Network Correlation and Feature Co-Clustering
abstract
Friend recommendation is an important recommender application in social media. Major social websites such as Twitter and Facebook are all capable of recommending friends to individuals. However, friend recommendation is a difficult task and most social websites use simple friend recommendation algorithms such as similarity and popularity, whose level of accuracy does do not satisfy the majority of users. In this paper we propose a two-stage procedure for more accurate friend recommendation: In the first stage, based on the relationship of different social networks, the Flickr tag network and contact network are aligned to generate a "possible friend list"; In the second stage, making the assumption that "a friend's friends also tend to be friends", co-clustering is applied to the tag and image information of the list to refine the recommendation result in the first stage. Experimental results show that the proposed method achieves good performance and every stage contributes to the recommendation.
Shangrong Huang, Jian Zhang 0002, Shiyang Lu, Xian-Sheng Hua 0001
ICMR2
2015 Decorrelation-stretch based cloud detection for total sky images
abstract
Cloud detection plays an important role in total-sky images based solar forecasting and has received more attention in recent years. Accurate cloud detection for complicated total-sky images is especially changeling due to the low contrast and vague boundaries between cloud and sky regions. Unlike the existing cloud detection method without any preprocessing, one novel decorrelation-stretch (DS) based method is proposed in this work, where the total-sky images are preprocessed using the DS algorithm firstly. With this enhancement, color feature disparity of cloud and sky can be intensified notably, and then a more accurate threshold can be obtained by applying the Minimum Cross Entropy (MCE) to the preprocessed image. Experimental results demonstrated the proposed scheme achieves better performance than the existing cloud detection methods on total-sky images, especially for images with low contrast or vague boundaries between cloud and sky regions.
Muming Zhao, Wenjun Zhang 0001, Wei Li 0037, Jian Zhang 0002
VCIP5
2015 Multiple kernel extreme learning machine
Xinwang Liu 0002, Lei Wang 0001, Guang-Bin Huang, Jian Zhang 0002, Jianping Yin
Neurocomputing4
2015 A fast affine-invariant features for image stitching under large viewpoint changes
Xiaomin Ma, Jian Zhang 0002, Jing Xin
Neurocomputing3
2015 Abrupt motion tracking via nearest neighbor field driven stochastic sampling
Tianfei Zhou, Yao Lu 0001, Feng Lv, Huijun Di, Qingjie Zhao, Jian Zhang 0002
Neurocomputing6
2015 An efficient radius-incorporated MKL algorithm for Alzheimer's disease prediction
Xinwang Liu 0002, Luping Zhou, Lei Wang 0001, Jian Zhang 0002, Jianping Yin, Dinggang Shen
Pattern Recognit.4
2015 Robust facial landmark localization using classified random ferns and pose-based initialization
Jian Zhang 0002, Dongyan Guo, Zhong Jin
Signal Process.2
2015 Exploratory Product Image Search With Circle-to-Search Interaction
abstract
Exploratory search is emerging as a new form of information-seeking activity in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this paper, we investigate the challenges of understanding users' search interests from the product images being browsed and inferring their actual search intentions. We propose a novel interactive image exploring system for allowing users to lightly switch between browse and search processes, and naturally complete visual-based exploratory search tasks in an effective and efficient way. This system enables users to specify their visual search interests in product images by circling any visual objects in web pages, and then the system automatically infers users' underlying intent by analyzing the browsing context and by analyzing the same or similar product images obtained by large-scale image search technology. Users can then utilize the recommended queries to complete intent-specific exploratory tasks. The proposed solution is one of the first attempts to understand users' interests for a visual-based exploratory product search task by integrating the browse and search activities. We have evaluated our system performance based on five million product images. The evaluation study demonstrates that the proposed system provides accurate intent-driven search results and fast response to exploratory search demands compared with the conventional image search methods, and also, provides users with robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2015 Manifold Kernel Sparse Representation of Symmetric Positive-Definite Matrices and Its Applications
abstract
The symmetric positive-definite (SPD) matrix, as a connected Riemannian manifold, has become increasingly popular for encoding image information. Most existing sparse models are still primarily developed in the Euclidean space. They do not consider the non-linear geometrical structure of the data space, and thus are not directly applicable to the Riemannian manifold. In this paper, we propose a novel sparse representation method of SPD matrices in the data-dependent manifold kernel space. The graph Laplacian is incorporated into the kernel space to better reflect the underlying geometry of SPD matrices. Under the proposed framework, we design two different positive definite kernel functions that can be readily transformed to the corresponding manifold kernels. The sparse representation obtained has more discriminating power. Extensive experimental results demonstrate good performance of manifold kernel sparse codes in image classification, face recognition, and visual tracking.
Yuwei Wu 0001, Yunde Jia, Peihua Li, Jian Zhang 0002, Junsong Yuan 0001
IEEE Trans. Image Process.4
2015 Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR Descriptor
abstract
The appearance gap between sketches and photo- realistic images is a fundamental challenge in sketch-based image retrieval (SBIR) systems. The existence of noisy edges on photo- realistic images is a key factor in the enlargement of the appearance gap and significantly degrades retrieval performance . To bridge the gap, we propose a framework consisting of a new line segment -based descriptor named histogram of line relationship (HLR) and a new noise impact reduction algorithm known as object boundary selection . HLR treats sketches and extracted edges of photo- realistic images as a series of piece-wise line segments and captures the relationship between them. Based on the HLR, the object boundary selection algorithm aims to reduce the impact of noisy edges by selecting the shaping edges that best correspond to the object boundaries. Multiple hypotheses are generated for descriptors by hypothetical edge selection. The selection algorithm is formulated to find the best combination of hypotheses to maximize the retrieval score; a fast method is also proposed. To reduce the distraction of false matches in the scoring process, two constraints on spatial and coherent aspects are introduced . We tested the HLR descriptor and the proposed framework on public datasets and a new image dataset of three million images, which we recently collected for SBIR evaluation purposes. We compared the proposed HLR with state-of-the-art descriptors (SHoG, GF-HOG). The experimental results show that our HLR descriptor outperforms them. Combined with the object boundary selection algorithm, our framework significantly improves SBIR performance.
Jian Zhang 0002, Tony X. Han, Zhenjiang Miao
IEEE Trans. Multim.2
2014 Sample-Adaptive Multiple Kernel Learning
abstract
Existing multiple kernel learning (MKL) algorithms \textit{indiscriminately} apply a same set of kernel combination weights to all samples. However, the utility of base kernels could vary across samples and a base kernel useful for one sample could become noisy for another. In this case, rigidly applying a same set of kernel combination weights could adversely affect the learning performance. To improve this situation, we propose a sample-adaptive MKL algorithm, in which base kernels are allowed to be adaptively switched on/off with respect to each sample. We achieve this goal by assigning a latent binary variable to each base kernel when it is applied to a sample. The kernel combination weights and the latent variables are jointly optimized via margin maximization principle. As demonstrated on five benchmark data sets, the proposed algorithm consistently outperforms the comparable ones in the literature.
Xinwang Liu 0002, Lei Wang 0001, Jian Zhang 0002, Jianping Yin
AAAI3
2014 Fast Mode and Depth Decision Algorithm for Intra Prediction of Quality SHVC
Chun Yuan 0003, Yu Sun 0003, Jian Zhang 0002, Hanning Zhou
ICIC (1)4
2014 Street view cross-sourced point cloud matching and registration
abstract
Object registration has been widely discussed with the development of various range sensing technologies. In most work, however, the point clouds of reference and target are generated by the same technology, such as a Kinect range camera, LiDAR sensor, or Structure from Motion technique. Cases in which reference and target point clouds are generated by different technologies are rarely discussed. Due to the significant differences across various point cloud data in terms of point cloud density, sensing noise, scale, occlusion etc., object registration between such different point clouds becomes extremely difficult. In this study, we address for the first time an even more challenging case in which the differently-sourced point clouds are acquired from a real street view. One is generated on the basis of an image sequence through the SfM process, and the other is produced directly by the LiDAR system. We propose a two-stage matching and registration algorithm to achieve object registration between these two different point clouds. The experiments are based on real building object point cloud data and demonstrate the effectiveness and efficiency of the proposed solution. The newly proposed solution can be further developed to contribute to several related applications, such as Location Based Service.
Furong Peng, Qiang Wu 0001, Lixin Fan, Jian Zhang 0002, Yu You, Jianfeng Lu 0003, Jing-Yu Yang 0001
ICIP4
2014 Multiple Kernel Learning Based Multi-view Spectral Clustering
abstract
For a given data set, exploring their multi-view instances under a clustering framework is a practical way to boost the clustering performance. This is because that each view might reflect partial information for the existing data. Furthermore, due to the noise and other impact factors, exploring these instances from different views will enhance the mining of the real structure and feature information within the data set. In this paper, we propose a multiple kernel spectral clustering algorithm through the multi-view instances on the given data set. By combining the kernel matrix learning and the spectral clustering optimization into one process framework, the algorithm can determine the kernel weights and cluster the multi-view data simultaneously. We compare the proposed algorithm with some recent published methods on real-world datasets to show the efficiency of the proposed algorithm.
Dongyan Guo, Jian Zhang 0002, Xinwang Liu 0002, Chunxia Zhao
ICPR2
2014 A Method of Discriminative Information Preservation and In-Dimension Distance Minimization Method for Feature Selection
abstract
Preserving sample's pair wise similarity is essential for feature selection. In supervised learning, labels can be used as a direct measure to check whether two samples are similar with each other. In unsupervised learning, however, such similarity information is usually unavailable. In this paper, we propose a new feature selection method through spectral clustering based on discriminative information as an underlying data structure. Laplacian matrix is used to obtain more partitioning information than other previously proposed structures such as the Eigen space of original data. The high dimension of sample data is projected into a low dimensional space. The in-dimension distance is also considered to get a better compact clustering result. The proposed method can be solved efficiently by updating the projection matrix and its inverse normalized diagonal matrix. A comprehensive experimental study has demonstrated that the proposed method outperforms many state-of-the-art feature selection algorithms with different criterion including the accuracy of clustering/classification and Jaccard score.
Shangrong Huang, Jian Zhang 0002, Xinwang Liu 0002, Lei Wang 0001
ICPR2
2014 Depth Super-resolution by Fusing Depth Imaging and Stereo Vision with Structural Determinant Information Inference
abstract
In this paper, we present a depth super-resolution framework by fusing depth imaging and stereo vision for high-resolution and high-accuracy depth maps. Depth cameras and stereo vision have their own limitations in some aspects, but their characteristics of range sensing are complementary. Thus, combining both approaches can produce more satisfactory results than either one. Unlike previous fusion methods, we initially taking the noisy depth observation from depth camera as prior information of scene structure. The prior information of scene structure is also utilized to infer structural determinant information, like depth discontinuity and occlusion, which is essential to improve the quality of depth map in the fusion process. In succession, the prior knowledge helps to overcome difficulties of intensity inconsistency in image observation from stereo vision component. Experimental results demonstrate effectiveness and accuracy of the proposed method.
Yucheng Wang 0003, Huijun Di, Wei Liang 0008, Jian Zhang 0002, Yunde Jia
ICPR5
2014 A people counting method based on head detection and tracking
abstract
This paper proposes a novel people counting method based on head detection and tracking to evaluate the number of people who move under an over-head camera. There are four main parts in the proposed method: foreground extraction, head detection, head tracking, and crossing-line judgment. The proposed method first utilizes an effective foreground extraction method to obtain foreground regions of moving people, and some morphological operations are employed to optimize the foreground regions. Then it exploits a LBP feature based Adaboost classifier for head detection in the optimized foreground regions. After head detection is performed, the candidate head object is tracked by a local head tracking method based on Meanshift algorithm. Based on head tracking, the method finally uses crossing-line judgment to determine whether the candidate head object will be counted or not. Experiments show that our method can obtain promising people counting accuracy about 96% and acceptable computation speed under different circumstances.
Jian Zhang 0002, Zheng Zhang 0006, Yong Xu 0001
SMARTCOMP2
2014 Recognizing flu-like symptoms from videos
abstract
BACKGROUND: Vision-based surveillance and monitoring is a potential alternative for early detection of respiratory disease outbreaks in urban areas complementing molecular diagnostics and hospital and doctor visit-based alert systems. Visible actions representing typical flu-like symptoms include sneeze and cough that are associated with changing patterns of hand to head distances, among others. The technical difficulties lie in the high complexity and large variation of those actions as well as numerous similar background actions such as scratching head, cell phone use, eating, drinking and so on. RESULTS: In this paper, we make a first attempt at the challenging problem of recognizing flu-like symptoms from videos. Since there was no related dataset available, we created a new public health dataset for action recognition that includes two major flu-like symptom related actions (sneeze and cough) and a number of background actions. We also developed a suitable novel algorithm by introducing two types of Action Matching Kernels, where both types aim to integrate two aspects of local features, namely the space-time layout and the Bag-of-Words representations. In particular, we show that the Pyramid Match Kernel and Spatial Pyramid Matching are both special cases of our proposed kernels. Besides experimenting on standard testbed, the proposed algorithm is evaluated also on the new sneeze and cough set. Empirically, we observe that our approach achieves competitive performance compared to the state-of-the-arts, while recognition on the new public health dataset is shown to be a non-trivial task even with simple single person unobstructed view. CONCLUSIONS: Our sneeze and cough video dataset and newly developed action recognition algorithm is the first of its kind and aims to kick-start the field of action recognition of flu-like symptoms from videos. It will be challenging but necessary in future developments to consider more complex real-life scenario of detecting these actions simultaneously from multiple persons in possibly crowded environments.
Tuan Hue Thi, Li Wang 0033, Jian Zhang 0002, Sebastian Maurer-Stroh, Li Cheng 0001
BMC Bioinform.4
2014 Exploiting Universum data in AdaBoost using gradient descent
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
Image Vis. Comput.3
2014 A fast mode decision algorithm applied to Coarse-Grain quality Scalable Video Coding
Chun Yuan 0003, Yu Sun 0003, Jian Zhang 0002, Xin Jin 0002
J. Vis. Commun. Image Represent.4
2014 Metric Learning Based Structural Appearance Model for Robust Visual Tracking
abstract
Appearance modeling is a key issue for the success of a visual tracker. Sparse representation based appearance modeling has received an increasing amount of interest in recent years. However, most of existing work utilizes reconstruction errors to compute the observation likelihood under the generative framework, which may give poor performance, especially for significant appearance variations. In this paper, we advocate an approach to visual tracking that seeks an appropriate metric in the feature space of sparse codes and propose a metric learning based structural appearance model for more accurate matching of different appearances. This structural representation is acquired by performing multiscale max pooling on the weighted local sparse codes of image patches. An online multiple instance metric learning algorithm is proposed that learns a discriminative and adaptive metric, thereby better distinguishing the visual object of interest from the background. The multiple instance setting is able to alleviate the drift problem potentially caused by misaligned training examples. Tracking is then carried out within a Bayesian inference framework, in which the learned metric and the structure object representation are used to construct the observation model. Comprehensive experiments on challenging image sequences demonstrate qualitatively and quantitatively that the proposed algorithm outperforms the state-of-the-art methods.
Yuwei Wu 0001, Bo Ma 0001, Min Yang 0003, Jian Zhang 0002, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.4
2014 Boosting Separability in Semisupervised Learning for Object Classification
abstract
Boosting algorithms, especially AdaBoost, have attracted great attention in computer vision. In the early version of boosting algorithms, the weak classifier selection and the strong classifier learning are linked together. It has been demonstrated that decoupling of these two processes can provide more flexibility for training a better classifier. In these studies, linear discriminant analysis (LDA) has been adopted to select weak classifiers independently based on class separability rather than a training error that occurs normally in AdaBoost. It is observed that LDA is successful only if a large number of labeled training samples is available. However, a large-scale labeled training set is not always available in many computer vision applications such as object classification. To tackle this problem, this paper proposes semisupervised subspace learning combined with a boosting framework for object classification, through which unlabeled data can participate in the boosting training to compensate for the lack of enough labeled data. With the proposed framework, this paper develops three various approaches that utilize unlabeled data in different ways. According to the experiments on several public image data sets, the proposed methods achieve superior performance over AdaBoost and existing semisupervised algorithms.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang
IEEE Trans. Circuits Syst. Video Technol.3
2014 Recognizing Gaits Across Views Through Correlated Motion Co-Clustering
abstract
Human gait is an important biometric feature, which can be used to identify a person remotely. However, view change can cause significant difficulties for gait recognition because it will alter available visual features for matching substantially. Moreover, it is observed that different parts of gait will be affected differently by view change. By exploring relations between two gaits from two different views, it is also observed that a part of gait in one view is more related to a typical part than any other parts of gait in another view. A new method proposed in this paper considers such variance of correlations between gaits across views that is not explicitly analyzed in the other existing methods. In our method, a novel motion co-clustering is carried out to partition the most related parts of gaits from different views into the same group. In this way, relationships between gaits from different views will be more precisely described based on multiple groups of the motion co-clustering instead of a single correlation descriptor. Inside each group, a linear correlation between gait information across views is further maximized through canonical correlation analysis (CCA). Consequently, gait information in one view can be projected onto another view through a linear approximation under the trained CCA subspaces. In the end, a similarity between gaits originally recorded from different views can be measured under the approximately same view. Comprehensive experiments based on widely adopted gait databases have shown that our method outperforms the state-of-the-art.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li, Liang Wang 0001
IEEE Trans. Image Process.3
2014 Global and Local Structure Preservation for Feature Selection
abstract
The recent literature indicates that preserving global pairwise sample similarity is of great importance for feature selection and that many existing selection criteria essentially work in this way. In this paper, we argue that besides global pairwise sample similarity, the local geometric structure of data is also critical and that these two factors play different roles in different learning scenarios. In order to show this, we propose a global and local structure preservation framework for feature selection (GLSPFS) which integrates both global pairwise sample similarity and local geometric data structure to conduct feature selection. To demonstrate the generality of our framework, we employ methods that are well known in the literature to model the local geometric data structure and develop three specific GLSPFS-based feature selection algorithms. Also, we develop an efficient optimization algorithm with proven global convergence to solve the resulting feature selection problem. A comprehensive experimental study is then conducted in order to compare our feature selection algorithms with many state-of-the-art ones in supervised, unsupervised, and semisupervised learning scenarios. The result indicates that: 1) our framework consistently achieves statistically significant improvement in selection performance when compared with the currently used algorithms; 2) in supervised and semisupervised learning scenarios, preserving global pairwise similarity is more important than preserving local geometric data structure; 3) in the unsupervised scenario, preserving local geometric data structure becomes clearly more important; and 4) the best feature selection performance is always obtained when the two factors are appropriately integrated. In summary, this paper not only validates the advantages of the proposed GLSPFS framework but also gains more insight into the information to be preserved in different feature selection tasks.
Xinwang Liu 0002, Lei Wang 0001, Jian Zhang 0002, Jianping Yin, Huan Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2014 Browse-to-Search: Interactive Exploratory Search with Visual Entities
abstract
With the development of image search technology, users are no longer satisfied with searching for images using just metadata and textual descriptions. Instead, more search demands are focused on retrieving images based on similarities in their contents (textures, colors, shapes etc.). Nevertheless, one image may deliver rich or complex content and multiple interests. Sometimes users do not sufficiently define or describe their seeking demands for images even when general search interests appear, owing to a lack of specific knowledge to express their intents. A new form of information seeking activity, referred to as exploratory search, is emerging in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this work, we investigate the challenges of understanding users' search interests from the images being browsed and infer their actual search intentions. We develop a novel system to explore an effective and efficient way for allowing users to seamlessly switch between browse and search processes, and naturally complete visual-based exploratory search tasks. The system, called Browse-to-Search enables users to specify their visual search interests by circling any visual objects in the webpages being browsed, and then the system automatically forms the visual entities to represent users' underlying intent. One visual entity is not limited by the original image content, but also encapsulated by the textual-based browsing context and the associated heterogeneous attributes. We use large-scale image search technology to find the associated textual attributes from the repository. Users can then utilize the encapsulated visual entities to complete search tasks. The Browse-to-Search system is one of the first attempts to integrate browse and search activities for a visual-based exploratory search, which is characterized by four unique properties: (1) in session—searching is performed during browsing session and search results naturally accompany with browsing content; (2) in context—the pages being browsed provide text-based contextual cues for searching; (3) in focus—users can focus on the visual content of interest without worrying about the difficulties of query formulation, and visual entities will be automatically formed; and (4) intuitiveness—a touch and visual search-based user interface provides a natural user experience. We deploy the Browse-to-Search system on tablet devices and evaluate the system performance using millions of images. We demonstrate that it is effective and efficient in facilitating the user's exploratory search compared to the conventional image search methods and, more importantly, provides users with more robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
ACM Trans. Inf. Syst.4
2013 A new edge feature for head-shoulder detection
abstract
In this work, we introduce a new edge feature to improve the head-shoulder detection performance. Since Head-shoulder detection is much vulnerable to vague contour, our new edge feature is designed to extract and enhance the head-shoulder contour and suppress the other contours. The basic idea is that head-shoulder contour can be predicted by filtering edge image with edge patterns, which are generated from edge fragments through a learning process. This edge feature can significantly enhance the object contour such as human head and shoulder known as En-Contour. To evaluate the performance of the new En-Contour, we combine it with HOG+LBP [1] as HOG+LBP+En-Contour. The HOG+LBP is the state-of-the-art feature in pedestrian detection. Because the human head-shoulder detection is a special case of pedestrian detection, we also use it as our baseline. Our experiments have indicated that this new feature significantly improve the HOG+LBP.
Jian Zhang 0002, Zhenjiang Miao
ICIP2
2013 Training boosting-like algorithms with semi-supervised subspace learning
abstract
Boosting algorithms have attracted great attention since the first real-time face detector by Viola & Jones through feature selection and strong classifier learning simultaneously. On the other hand, researchers have proposed to decouple such two procedures to improve the performance of Boosting algorithms. Motivated by this, we propose a boosting-like algorithm framework by embedding semi-supervised subspace learning methods. It selects weak classifiers based on class-separability. Combination weights of selected weak classifiers can be obtained by subspace learning. Three typical algorithms are proposed under this framework and evaluated on public data sets. As shown by our experimental results, the proposed methods obtain superior performances over their supervised counterparts and AdaBoost.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang
ICIP3
2013 Attribute-based learning for large scale object classification
abstract
Scalability to large numbers of classes is an important challenge for multi-class classification. It can often be computationally infeasible at test phase when class prediction is performed by using every possible classifier trained for each individual class. This paper proposes an attribute-based learning method to overcome this limitation. First is to define attributes and their associations with object classes automatically and simultaneously. Such associations are learned based on greedy strategy under certain conditions. Second is to learn a classifier for each attribute instead of each class. Then, these trained classifiers are used to predict classes based on their attribute representations. The proposed method also allows trade-off between test-time complexity (which grows linearly with the number of attributes) and accuracy. Experiments based on Animals-with-Attributes and ILSVRC2010 datasets have shown that the performance of our method is promising when compared with the state-of-the-art.
Worapan Kusakunniran, Shin'ichi Satoh 0001, Jian Zhang 0002, Qiang Wu 0001
ICME3
2013 On Discovering the Correlated Relationship between Static and Dynamic Data in Clinical Gait Analysis
Yin Song, Jian Zhang 0002, Longbing Cao, Morgan Sangeux
ECML/PKDD (3)2
2013 Fast human action classification and VOI localization with enhanced sparse coding
Shiyang Lu, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng
J. Vis. Commun. Image Represent.2
2013 An Efficient Approach to Integrating Radius Information into Multiple Kernel Learning
abstract
Integrating radius information has been demonstrated by recent work on multiple kernel learning (MKL) as a promising way to improve kernel learning performance. Directly integrating the radius of the minimum enclosing ball (MEB) into MKL as it is, however, not only incurs significant computational overhead but also possibly adversely affects the kernel learning performance due to the notorious sensitivity of this radius to outliers. Inspired by the relationship between the radius of the MEB and the trace of total data scattering matrix, this paper proposes to incorporate the latter into MKL to improve the situation. In particular, in order to well justify the incorporation of radius information, we strictly comply with the radius-margin bound of support vector machines (SVMs) and thus focus on the l2-norm soft-margin SVM classifier. Detailed theoretical analysis is conducted to show how the proposed approach effectively preserves the merits of incorporating the radius of the MEB and how the resulting optimization is efficiently solved. Moreover, the proposed approach achieves the following advantages over its counterparts: 1) more robust in the presence of outliers or noisy training samples; 2) more computationally efficient by avoiding the quadratic optimization for computing the radius at each iteration; and 3) readily solvable by the existing off-the-shelf MKL packages. Comprehensive experiments are conducted on University of California, Irvine, protein subcellular localization, and Caltech-101 data sets, and the results well demonstrate the effectiveness and efficiency of our approach.
Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, En Zhu, Jian Zhang 0002
IEEE Trans. Cybern.5
2013 An Adaptive Approach to Learning Optimal Neighborhood Kernels
abstract
Learning an optimal kernel plays a pivotal role in kernel-based methods. Recently, an approach called optimal neighborhood kernel learning (ONKL) has been proposed, showing promising classification performance. It assumes that the optimal kernel will reside in the neighborhood of a "pre-specified" kernel. Nevertheless, how to specify such a kernel in a principled way remains unclear. To solve this issue, this paper treats the pre-specified kernel as an extra variable and jointly learns it with the optimal neighborhood kernel and the structure parameters of support vector machines. To avoid trivial solutions, we constrain the pre-specified kernel with a parameterized model. We first discuss the characteristics of our approach and in particular highlight its adaptivity. After that, two instantiations are demonstrated by modeling the pre-specified kernel as a common Gaussian radial basis function kernel and a linear combination of a set of base kernels in the way of multiple kernel learning (MKL), respectively. We show that the optimization in our approach is a min-max problem and can be efficiently solved by employing the extended level method and Nesterov's method. Also, we give the probabilistic interpretation for our approach and apply it to explain the existing kernel learning methods, providing another perspective for their commonness and differences. Comprehensive experimental results on 13 UCI data sets and another two real-world data sets show that via the joint learning process, our approach not only adaptively identifies the pre-specified kernel, but also achieves superior classification performance to the original ONKL and the related MKL algorithms.
Xinwang Liu 0002, Jianping Yin, Lei Wang 0001, Lingqiao Liu, Jun Liu 0003, Chenping Hou, Jian Zhang 0002
IEEE Trans. Cybern.7
2013 A New View-Invariant Feature for Cross-View Gait Recognition
abstract
Human gait is an important biometric feature which is able to identify a person remotely. However, change of view causes significant difficulties for recognizing gaits. This paper proposes a new framework to construct a new view-invariant feature for cross-view gait recognition. Our view-normalization process is performed in the input layer (i.e., on gait silhouettes) to normalize gaits from arbitrary views. That is, each sequence of gait silhouettes recorded from a certain view is transformed onto the common canonical view by using corresponding domain transformation obtained through invariant low-rank textures (TILTs). Then, an improved scheme of procrustes shape analysis (PSA) is proposed and applied on a sequence of the normalized gait silhouettes to extract a novel view-invariant gait feature based on procrustes mean shape (PMS) and consecutively measure a gait similarity based on procrustes distance (PD). Comprehensive experiments were carried out on widely adopted gait databases. It has been shown that the performance of the proposed method is promising when compared with other existing methods in the literature.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
IEEE Trans. Inf. Forensics Secur.3
2012 Object Detection Based on Co-occurrence GMuLBP Features
abstract
Image co-occurrence has shown great powers on object classification because it captures the characteristic of individual features and spatial relationship between them simultaneously. For example, Co-occurrence Histogram of Oriented Gradients (CoHOG) has achieved great success on human detection task. However, the gradient orientation in CoHOG is sensitive to noise. In addition, CoHOG does not take gradient magnitude into account which is a key component to reinforce the feature detection. In this paper, we propose a new LBP feature detector based image co-occurrence. Building on uniform Local Binary Patterns, the new feature detector detects Co-occurrence Orientation through Gradient Magnitude calculation. It is known as CoGMuLBP. An extension version of the GoGMuLBP is also presented. The experimental results on the UIUC car data set show that the proposed features outperform state-of-the-art methods.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
ICME3
2012 Browse-to-search
abstract
This demonstration presents a novel interactive online shopping application based on visual search technologies. When users want to buy something on a shopping site, they usually have the requirement of looking for related information from other web sites. Therefore users need to switch between the web page being browsed and other websites that provide search results. The proposed application enables users to naturally search products of interest when they browse a web page, and make their even causal purchase intent easily satisfied. The interactive shopping experience is characterized by: 1) in session---it allows users to specify the purchase intent in the browsing session, instead of leaving the current page and navigating to other websites; 2) in context---the browsed web page provides implicit context information which helps infer user purchase preferences; 3) in focus---users easily specify their search interest using gesture on touch devices and do not need to formulate queries in search box; 4) natural-gesture inputs and visual-based search provides users a natural shopping experience. The system is evaluated against a data set consisting of several millions commercial product images.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng, Jian-Tao Sun, Shipeng Li 0001
ACM Multimedia4
2012 Local visual words coding for low bit rate mobile visual search
abstract
Mobile visual search has attracted extensive attention for its huge potential for numerous applications. Research on this topic has been focused on two schemes: sending query images, and sending compact descriptors extracted on mobile phones. The first scheme requires about 30-40KB data to transmit, while the second can reduce the bit rate by 10 times. In this paper, we propose a third scheme for extremely low bit rate mobile visual search, which sends compressed visual words consisting of vocabulary tree histogram and descriptor orientations rather than descriptors. This scheme can further reduce the bit rate with few extra computational costs on the client. Specifically, we store a vocabulary tree and extract visual descriptors on the mobile client. A light-weight pre-retrieval is performed to obtain the visited leaf nodes in the vocabulary tree. The orientation of each local descriptor and the tree histogram are then encoded to be transmitted to server. Our new scheme transmits less than 1KB data, which reduces the bit rate in the second scheme by 3 times, and obtains about 30% improvement in terms of search accuracy over the traditional Bag-of-Words baseline. The time cost is only 1.5 secs on the client and 240 msecs on the server.
Shiyang Lu, Tao Mei 0001, Jian Zhang 0002, Shipeng Li 0001
ACM Multimedia4
2012 Integrating local action elements for action analysis
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
Comput. Vis. Image Underst.3
2012 Structured learning of local features for human action classification and localization
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
Image Vis. Comput.3
2012 Cross-view and multi-view gait recognitions based on view transformation model using multi-layer perceptron
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
Pattern Recognit. Lett.3
2012 Fast and Accurate Human Detection Using a Cascade of Boosted MS-LBP Features
abstract
In this letter, a new scheme for generating local binary patterns (LBP) is presented. This Modified Symmetric LBP (MS-LBP) feature takes advantage of LBP and gradient features. It is then applied into a boosted cascade framework for human detection. By combining MS-LBP with Haar-like feature into the boosted framework, the performances of heterogeneous features based detectors are evaluated for the best trade-off between accuracy and speed. Two feature training schemes, namely Single AdaBoost Training Scheme (SATS) and Dual AdaBoost Training Scheme (DATS) are proposed and compared. On the top of AdaBoost, two multidimensional feature projection methods are described. A comprehensive experiment is presented. Apart from obtaining higher detection accuracy, the detection speed based on DATS is 17 times faster than HOG method.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
IEEE Signal Process. Lett.3
2012 Gait Recognition Under Various Viewing Angles Based on Correlated Motion Regression
abstract
It is well recognized that gait is an important biometric feature to identify a person at a distance, e.g., in video surveillance application. However, in reality, change of viewing angle causes significant challenge for gait recognition. A novel approach using regression-based view transformation model (VTM) is proposed to address this challenge. Gait features from across views can be normalized into a common view using learned VTM(s). In principle, a VTM is used to transform gait feature from one viewing angle (source) into another viewing angle (target). It consists of multiple regression processes to explore correlated walking motions, which are encoded in gait features, between source and target views. In the learning processes, sparse regression based on the elastic net is adopted as the regression function, which is free from the problem of overfitting and results in more stable regression models for VTM construction. Based on widely adopted gait database, experimental results show that the proposed method significantly improves upon existing VTM-based methods and outperforms most other baseline methods reported in the literature. Several practical scenarios of applying the proposed method for gait recognition under various views are also discussed in this paper.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
IEEE Trans. Circuits Syst. Video Technol.3
2012 Gait Recognition Across Various Walking Speeds Using Higher Order Shape Configuration Based on a Differential Composition Model
abstract
Gait has been known as an effective biometric feature to identify a person at a distance. However, variation of walking speeds may lead to significant changes to human walking patterns. It causes many difficulties for gait recognition. A comprehensive analysis has been carried out in this paper to identify such effects. Based on the analysis, Procrustes shape analysis is adopted for gait signature description and relevant similarity measurement. To tackle the challenges raised by speed change, this paper proposes a higher order shape configuration for gait shape description, which deliberately conserves discriminative information in the gait signatures and is still able to tolerate the varying walking speed. Instead of simply measuring the similarity between two gaits by treating them as two unified objects, a differential composition model (DCM) is constructed. The DCM differentiates the different effects caused by walking speed changes on various human body parts. In the meantime, it also balances well the different discriminabilities of each body part on the overall gait similarity measurements. In this model, the Fisher discriminant ratio is adopted to calculate weights for each body part. Comprehensive experiments based on widely adopted gait databases demonstrate that our proposed method is efficient for cross-speed gait recognition and outperforms other state-of-the-art methods.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
IEEE Trans. Syst. Man Cybern. Part B3
2011 Pairwise Shape configuration-based PSA for gait recognition under small viewing angle change
abstract
Two main components of Procrustes Shape Analysis (PSA) are adopted and adapted specifically to address gait recognition under small viewing angle change: 1) Procrustes Mean Shape (PMS) for gait signature description; 2) Procrustes Distance (PD) for similarity measurement. Pairwise Shape Configuration (PSC) is proposed as a shape descriptor in place of existing Centroid Shape Configuration (CSC) in conventional PSA. PSC can better tolerate shape change caused by viewing angle change than CSC. Small variation of viewing angle makes large impact only on global gait appearance. Without major impact on local spatio-temporal motion, PSC which effectively embeds local shape information can generate robust view-invariant gait feature. To enhance gait recognition performance, a novel boundary re-sampling process is proposed. It provides only necessary re-sampled points to PSC description. In the meantime, it efficiently solves problems of boundary point correspondence, boundary normalization and boundary smoothness. This re-sampling process adopts prior knowledge of body pose structure. Comprehensive experiment is carried out on the CASIA gait database. The proposed method is shown to significantly improve performance of gait recognition under small viewing angle change without additional requirements of supervised learning, known viewing angle and multi-camera system, when compared with other methods in literatures.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
AVSS3
2011 Speed-invariant gait recognition based on Procrustes Shape Analysis using higher-order shape configuration
abstract
Walking speed change is considered a typical challenge hindering reliable human gait recognition. This paper proposes a novel method to extract speed-invariant gait feature based on Procrustes Shape Analysis (PSA). Two major components of PSA, i.e., Procrustes Mean Shape (PMS) and Procrustes Distance (PD), are adopted and adapted specifically for the purpose of speed-invariant gait recognition. One of our major contributions in this work is that, instead of using conventional Centroid Shape Configuration (CSC) which is not suitable to describe individual gait when body shape changes particularly due to change of walking speed, we propose a new descriptor named Higher-order derivative Shape Configuration (HSC) which can generate robust speed-invariant gait feature. From the first order to the higher order, derivative shape configuration contains gait shape information of different levels. Intuitively, the higher order of derivative is able to describe gait with shape change caused by the larger change of walking speed. Encouraging experimental results show that our proposed method is efficient for speed-invariant gait recognition and evidently outperforms other existing methods in the literatures.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
ICIP3
2011 Active learning for human action recognition with Gaussian Processes
abstract
This paper presents an active learning approach for recognizing human actions in videos based on multiple kernel combined method. We design the classifier based on Multiple Kernel Learning (MKL) through Gaussian Processes (GP) regression. This classifier is then trained in an active learning approach. In each iteration, one optimal sample is selected to be interactively annotated and incorporated into training set. The selection of the sample is based on the heuristic feedback of the GP classifier. To our knowledge, GP regression MKL based active learning methods have not been applied to address the human action recognition yet. We test this approach on standard benchmarks. This approach outperforms the state-of-the-art techniques in accuracy while requires significantly less training samples.
Xianghang Liu, Jian Zhang 0002
ICIP2
2011 SKRWM based descriptor for pedestrian detection in thermal images
abstract
Pedestrian detection in a thermal image is a difficult task due to intrinsic challenges:1) low image resolution, 2) thermal noising, 3) polarity changes, 4) lack of color, texture or depth information. To address these challenges, we propose a novel mid-level feature descriptor for pedestrian detection in thermal domain, which combines pixel-level Steering Kernel Regression Weights Matrix (SKRWM) with their corresponding covariances. SKRWM can properly capture the local structure of pixels, while the covariance computation can further provide the correlation of low level feature. This mid-level feature descriptor not only captures the pixel-level data difference and spatial differences of local structure, but also explores the correlations among low-level features. In the case of human detection, the proposed mid-level feature descriptor can discriminatively distinguish pedestrian from complexity. For testing the performance of proposed feature descriptor, a popular classifier framework based on Principal Component Analysis (PCA) and Support Vector Machine (SVM) is also built. Overall, our experimental results show that proposed approach has overcome the problems caused by background subtraction in [1] while attains comparable detection accuracy compared to the state-of-the-arts.
Qiang Wu 0001, Jian Zhang 0002, Glenn Geers
MMSP3
2011 Incremental Training of a Detector Using Online Sparse Eigendecomposition
abstract
The ability to efficiently and accurately detect objects plays a very crucial role for many computer vision tasks. Recently, offline object detectors have shown a tremendous success. However, one major drawback of offline techniques is that a complete set of training data has to be collected beforehand. In addition, once learned, an offline detector cannot make use of newly arriving data. To alleviate these drawbacks, online learning has been adopted with the following objectives: 1) the technique should be computationally and storage efficient; 2) the updated classifier must maintain its high classification accuracy. In this paper, we propose an effective and efficient framework for learning an adaptive online greedy sparse linear discriminant analysis model. Unlike many existing online boosting detectors, which usually apply exponential or logistic loss, our online algorithm makes use of linear discriminant analysis' learning criterion that not only aims to maximize the class-separation criterion but also incorporates the asymmetrical property of training data distributions. We provide a better alternative for online boosting algorithms in the context of training a visual object detector. We demonstrate the robustness and efficiency of our methods on handwritten digit and face data sets. Our results confirm that object detection tasks benefit significantly when trained in an online manner.
Sakrapee Paisitkriangkrai, Chunhua Shen, Jian Zhang 0002
IEEE Trans. Image Process.3
2011 Efficiently Learning a Detection Cascade With Sparse Eigenvectors
abstract
Real-time object detection has many computer vision applications. Since Viola and Jones proposed the first real-time AdaBoost based face detection system, much effort has been spent on improving the boosting method. In this work, we first show that feature selection methods other than boosting can also be used for training an efficient object detector. In particular, we introduce greedy sparse linear discriminant analysis (GSLDA) for its conceptual simplicity and computational efficiency; and slightly better detection performance is achieved compared with . Moreover, we propose a new technique, termed boosted greedy sparse linear discriminant analysis (BGSLDA), to efficiently train a detection cascade. BGSLDA exploits the sample reweighting property of boosting and the class-separability criterion of GSLDA. Experiments in the domain of highly skewed data distributions (e.g., face detection) demonstrate that classifiers trained with the proposed BGSLDA outperforms AdaBoost and its variants. This finding provides a significant opportunity to argue that AdaBoost and similar approaches are not the only methods that can achieve high detection results for real-time object detection.
Chunhua Shen, Sakrapee Paisitkriangkrai, Jian Zhang 0002
IEEE Trans. Image Process.3
2010 Face Detection with Effective Feature Extraction
Sakrapee Paisitkriangkrai, Chunhua Shen, Jian Zhang 0002
ACCV (3)3
2010 Human Action Recognition and Localization in Video Using Structured Learning of Local Space-Time Features
abstract
This paper presents a unified framework for human action classification and localization in video using structured learning of local space-time features. Each human action class is represented by a set of its own compact set of local patches. In our approach, we first use a discriminative hierarchical Bayesian classifier to select those space-time interest points that are constructive for each particular action. Those concise local features are then passed to a Support Vector Machine with Principal Component Analysis projection for the classification task. Meanwhile, the action localization is done using Dynamic Conditional Random Fields developed to incorporate the spatial and temporal structure constraints of superpixels extracted around those features. Each superpixel in the video is defined by the shape and motion information of its corresponding feature region. Compelling results obtained from experiments on KTH [22], Weizmann [1], HOHA [13] and TRECVid [23] datasets have proven the efficiency and robustness of our framework for the task of human action recognition and localization in video.
Tuan Hue Thi, Jian Zhang 0002, Li Cheng 0001, Li Wang 0033, Shin'ichi Satoh 0001
AVSS2
2010 Support vector regression for multi-view gait recognition based on local motion feature selection
abstract
Gait is a well recognized biometric feature that is used to identify a human at a distance. However, in real environment, appearance changes of individuals due to viewing angle changes cause many difficulties for gait recognition. This paper re-formulates this problem as a regression problem. A novel solution is proposed to create a View Transformation Model (VTM) from the different point of view using Support Vector Regression (SVR). To facilitate the process of regression, a new method is proposed to seek local Region of Interest (ROI) under one viewing angle for predicting the corresponding motion information under another viewing angle. Thus, the well constructed VTM is able to transfer gait information under one viewing angle into another viewing angle. This proposal can achieve view-independent gait recognition. It normalizes gait features under various viewing angles into a common viewing angle before similarity measurement is carried out. The extensive experimental results based on widely adopted benchmark dataset demonstrate that the proposed algorithm can achieve significantly better performance than the existing methods in literature.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
CVPR3
2010 CROSS-layer QoS-optimized EDCA adaptation for wireless video streaming
abstract
In this paper, we propose an adaptive cross layer technique that optimally enhance the QoS of wireless video transmission in an IEEE 802.11e WLAN. The optimization takes into account the unequal error protection characteristics of video streaming, the IEEE 802.11e EDCA parameters and the lossy nature of wireless channel. Our proposed technique makes use of two analytical models: a video distortion model and a channel throughput estimation model. The first model predicts the video quality in term of average PSNR of all decoded video frames. The second model estimates the channel throughput and packet loss rates of each MAC layers queue, which are then fed into the first model as inputs. The optimal EDCA parameters are selected by an optimization module based on the information from the analytical models. The accuracy of our optimized EDCA parameters is verified through the extensive simulation.
Werayut Saesue, Chun Tung Chou, Jian Zhang 0002
ICIP3
2010 Implicit Motion-Shape Model: A generic approach for action matching
abstract
We develop a robust technique to find similar matches of human actions in video. Given a query video, Motion History Images (MHI) are constructed for consecutive keyframes. This is followed by dividing the MHI into local Motion-Shape regions, which allows us to analyze the action as a set of sparse space-time patches in 3D. Inspired by the idea of Generalized Hough Transform, we develop the Implicit Motion-Shape Model that allows the integration of these local patches to describe the dynamic characteristics of the query action. In the same way we retrieve motion segments from video candidates, then project them onto the Hough Space built by the query model. This produces the matching score by running Parzen window density estimation under different scales. Empirical experiments on popular datasets demonstrate the efficiency of this approach, where highly accurate matches are returned within acceptable processing time.
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033
ICIP3
2010 Improved human detection and classification in thermal images
abstract
We present a new method for detecting pedestrians in thermal images. The method is based on the Shape Context Descriptor (SCD) with the Adaboost cascade classifier framework. Compared with standard optical images, thermal imaging cameras offer a clear advantage for night-time video surveillance. It is robust on the light changes in day-time. Experiments show that shape context features with boosting classification provide a significant improvement on human detection in thermal images. In this work, we have also compared our proposed method with rectangle features on the public dataset of thermal imagery. Results show that shape context features are much better than the conventional rectangular features on this task.
Jian Zhang 0002, Chunhua Shen
ICIP2
2010 Multi-view Gait Recognition Based on Motion Regression Using Multilayer Perceptron
abstract
It has been shown that gait is an efficient biometric feature for identifying a person at a distance. However, it is a challenging problem to obtain reliable gait feature when viewing angle changes because the body appearance can be different under the various viewing angles. In this paper, the problem above is formulated as a regression problem where a novel View Transformation Model (VTM) is constructed by adopting Multilayer Perceptron (MLP) as regression tool. It smoothly estimates gait feature under an unknown viewing angle based on motion information in a well selected Region of Interest (ROI) under other existing viewing angles. Thus, this proposal can normalize gait features under various viewing angles into a common viewing angle before gait similarity measurement is carried out. Encouraging experimental results have been obtained based on widely adopted benchmark database.
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li
ICPR3
2010 Weakly Supervised Action Recognition Using Implicit Shape Models
abstract
In this paper, we present a robust framework for action recognition in video, that is able to perform competitively against the state-of-the-art methods, yet does not rely on sophisticated background subtraction preprocess to remove background features. In particular, we extend the Implicit Shape Modeling (ISM) of [10] for object recognition to 3D to integrate local spatiotemporal features, which are produced by a weakly supervised Bayesian kernel filter. Experiments on benchmark datasets (including KTH and Weizmann) verifies the effectiveness of our approach.
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
ICPR3
2010 Video quality prediction in the presence of MAC contention and wireless channel error
abstract
This paper proposes an integrated model to predict the quality of video, expressed in terms of mean square error (MSE) of the received video frames, in an IEEE 802.1 le wireless network. The proposed system takes into account contention at the MAC layer, wireless channel error, queueing at the MAC layer, parameters of different 802.1 le access categories (ACs), and video characteristics of different H.264 data partitions (DPs). To the best of the authors' knowledge, this is the first system that takes these network and video characteristics into consideration to predict video quality in an IEEE 802.1 le network. The proposed system consists of two components. The first component predicts the packet loss rate of each H.264 data partition by using a multi-dimensional discrete-time Markov chain (DTMC) coupled to a M/G/l queue. The second component uses these packet loss rates and the video characteristics to predict the MSE of each received video frames. We verify the accuracy of our combination system by using discrete event simulation and real H.264 coded video sequences.
Werayut Saesue, Chun Tung Chou, Jian Zhang 0002
WOWMOM3
2009 Automatic Gait Recognition Using Weighted Binary Pattern on Video
abstract
Human identification by recognizing the spontaneous gait recorded in real-world setting is a tough and not yet fully resolved problem in biometrics research. Several issues have contributed to the difficulties of this task. They include various poses, different clothes, moderate to large changes of normal walking manner due to carrying diverse goods when walking, and the uncertainty of the environments where the people are walking. In order to achieve a better gait recognition, this paper proposes a new method based on Weighted Binary Pattern (WBP). WBP first constructs binary pattern from a sequence of aligned silhouettes. Then, adaptive weighting technique is applied to discriminate significances of the bits in gait signatures. Being compared with most of existing methods in the literatures, this method can better deal with gait frequency, local spatial-temporal human pose features, and global body shape statistics. The proposed method is validated on several well known benchmark databases. The extensive and encouraging experimental results show that the proposed algorithm achieves high accuracy, but with low complexity and computational time.
Worapan Kusakunniran, Qiang Wu 0001, Hongdong Li, Jian Zhang 0002
AVSS4
2009 Human Body Articulation for Action Recognition in Video Sequences
abstract
This paper presents a new technique for action recognition in video using human body part-based approach, combining both local feature description of each body part, and global graphical model structure of the human action. The human body is divided into elementary points from which a Decomposable Triangulated Graph will be built. The temporal variation of human activity is encoded in the velocity distribution of each node in the graph, while the graph structure shows the spatial configuration of all the nodes in the action. Tracking trajectories of unlabeled good feature points are correctly labeled using Maximum a Posterior probability. Dynamic Programming is then implemented to boost up the exhaustive search for the optimal labeling of unknown body parts and the best possible action. A simple and efficient technique for building the optimal structure of the human action graph is also implemented. Experimental results on the KTH dataset proves the success and potential applications of this proposed technique.
Tuan Hue Thi, Sijun Lu, Jian Zhang 0002, Li Cheng 0001, Li Wang 0033
AVSS3
2009 Efficiently training a better visual detector with sparse eigenvectors
abstract
Face detection plays an important role in many vision applications. Since Viola and Jones [1] proposed the first real-time AdaBoost based object detection system, much effort has been spent on improving the boosting method. In this work, we first show that feature selection methods other than boosting can also be used for training an efficient object detector. In particular, we have adopted Greedy Sparse Linear Discriminant Analysis (GSLDA) [2] for its computational efficiency; and slightly better detection performance is achieved compared with [1]. Moreover, we propose a new technique, termed Boosted Greedy Sparse Linear Discriminant Analysis (BGSLDA), to efficiently train object detectors. BGSLDA exploits the sample re-weighting property of boosting and the class-separability criterion of GSLDA. Experiments in the domain of highly skewed data distributions, e.g., face detection, demonstrates that classifiers trained with the proposed BGSLDA outperforms AdaBoost and its variants. This finding provides a significant opportunity to argue that Adaboost and similar approaches are not the only methods that can achieve high classification results for high dimensional data such as object detection.
Sakrapee Paisitkriangkrai, Chunhua Shen, Jian Zhang 0002
CVPR3
2009 An overview of fast pedestrian detection: Feature selection and cascade framework of boosted features
abstract
Efficiently and accurately detecting pedestrians plays a crucial role in many vision applications such as video surveillance, multimedia retrieval and smart car etc. In order to find the right feature for this task, we first present a comprehensive experimental study on pedestrian detection using state-of-the-art locally-extracted features. Building upon our findings, we propose a new, simpler pedestrian detecting framework based on the covariance features. We conduct feature selection and weak classifier training in the Euclidean space for faster computation. To this end, two machine learning algorithms have been designed: AdaBoost with weighted Fisher linear discriminant analysis (WLDA) based weak classifiers and Greedy Sparse Linear Discriminant Analysis (GSLDA). To further accelerate the detection, we employ a faster strategy, multiple cascade layers with heterogeneous features, to exploit the efficiency of the Haar-like features and the discriminative power of the covariance features. Experimental results shown on different datasets prove that the new pedestrian detection is not only comparable to the performance of the state-of-the-art pedestrian detectors but it also performs at a faster speed.
Jian Zhang 0002, Sakrapee Paisitkriangkrai, Chunhua Shen
ICME1
2009 Detecting Ghost and Left Objects in Surveillance Video
abstract
This paper proposes an efficient method for detecting ghost and left objects in surveillance video, which, if not identified, may lead to errors or wasted computational power in background modeling and object tracking in video surveillance systems. This method contains two main steps: the first one is to detect stationary objects, which narrows down the evaluation targets to a very small number of regions in the input image; the second step is to discriminate the candidates between ghost and left objects. For the first step, we introduce a novel stationary object detection method based on continuous object tracking and shape matching. For the second step, we propose a fast and robust inpainting method to differentiate between ghost and left objects by reconstructing the real background using the candidate's corresponding regions in the current input and background image. The effectiveness of our method has been validated by experiments over a variety of video sequences and comparisons with existing state-of-art methods.
Sijun Lu, Jian Zhang 0002, David Dagan Feng
Int. J. Pattern Recognit. Artif. Intell.2
2008 Face detection from few training examples
abstract
Face detection in images is very important for many multimedia applications. Haar-like wavelet features have become dominant in face detection because of their tremendous success since Viola and Jones [1] proposed their AdaBoost based detection system. While Haar features' simplicity makes rapid computation possible, its discriminative power is limited. As a consequence, a large training dataset is required to train a classifier. This may hamper its application in scenarios that a large labeled dataset is difficult to obtain. In this work, we address the problem of learning to detect faces from a small set of training examples. In particular, we propose to use co- variance features. Also for better classification performance, linear hyperplane classifier based on Fisher discriminant analysis (FDA) is proffered. Compared with the decision stump, FDA is more discriminative and therefore fewer weak learners are needed. We show that the detection rate can be significantly improved with covariance features on a small dataset (a few hundred positive examples), compared to Haar features used in current most face detection systems.
Chunhua Shen, Sakrapee Paisitkriangkrai, Jian Zhang 0002
ICIP3
2008 An experimental study on pedestrian classification using local features
abstract
This paper presents an experimental study on pedestrian detection using state-of-the-art local feature extraction and support vector machine (SVM) classifiers. The performance of pedestrian detection using region covariance, histogram of oriented gradients (HOG) and local receptive fields (LRF) feature descriptors is experimentally evaluated. The experiments are performed on both the benchmarking dataset used in [1] and the MIT CBCL dataset. Both can be publicly accessed. The experimental results show that region covariance features with radial basis function (RBF) kernel SVM and HOG features with quadratic kernel SVM outperform the combination of LRF features with quadratic kernel SVM reported in [1].
Sakrapee Paisitkriangkrai, Chunhua Shen, Jian Zhang 0002
ISCAS3
2008 Robust object tracking using the particle filtering and level set methods: A comparative experiment
abstract
Robust visual tracking has become an important topic of research in computer vision. A novel method for robust object tracking, GATE [11], improves object tracking in complex environments using the particle filtering and the level set-based active contour method. GATE creates a spatial prior in the state space using shape information of the tracked object to filter particles in the state space in order to reshape and refine the posterior distribution of the particle filtering. This paper describes a comparative experiment that applies GATE and the standard particle filtering to track the object of interest in complex environments using simple features. Image sequences captured by the hand held, stationary and the PTZ camera are utilised. The experimental results demonstrate that GATE is able to solve the ambiguous outlier problem of particle filters in order to deal with heavy clutters in the background, occlusion, low resolution and noisy images, and thus significantly improves the particle filtering in object tracking.
Cheng Luo 0003, Xiongcai Cai, Jian Zhang 0002
MMSP3
2008 Hybrid frame-recursive block-based distortion estimation model for wireless video transmission
abstract
In wireless environments, video quality can be severely degraded due to channel errors. Improving error robustness towards the impact of packet loss in error-prone network is considered as a critical concern in wireless video networking research. Data partitioning (DP) is an efficient error-resilient tool in video codec that is capable of reducing the effect of transmission errors by reorganizing the coded video bitstream into different partitions with different levels of importance. Significant video performance improvement can be achieved if DP is jointly optimized with unequal error protection (UEP). This paper proposes a fast and accurate frame-recursive block-based distortion estimation model for the DP tool in H.264.AVC. The accuracy of our model comes from appropriately approximating the error-concealment cross-correlation term (which is neglected in earlier work in order to reduce computation burden) as a function of the first moment of decoded pixels.Without increasing computation complexity, our proposed distortion model can be applied to both fixed and variable block size intra-prediction and motion compensation. Extensive simulation results are presented to show the accuracy of our estimation algorithm.
Werayut Saesue, Jian Zhang 0002, Chun Tung Chou
MMSP2
2008 Fast Pedestrian Detection Using a Cascade of Boosted Covariance Features
abstract
Efficiently and accurately detecting pedestrians plays a very important role in many computer vision applications such as video surveillance and smart cars. In order to find the right feature for this task, we first present a comprehensive experimental study on pedestrian detection using state-of-the-art locally extracted features (e.g., local receptive fields, histogram of oriented gradients, and region covariance). Building upon the findings of our experiments, we propose a new, simpler pedestrian detector using the covariance features. Unlike the work in [1], where the feature selection and weak classifier training are performed on the Riemannian manifold, we select features and train weak classifiers in the Euclidean space for faster computation. To this end, AdaBoost with weighted Fisher linear discriminant analysis-based weak classifiers are designed. A cascaded classifier structure is constructed for efficiency in the detection phase. Experiments on different datasets prove that the new pedestrian detector is not only comparable to the state-of-the-art pedestrian detectors but it also performs at a faster speed. To further accelerate the detection, we adopt a faster strategy-multiple layer boosting with heterogeneous features-to exploit the efficiency of the Haar feature and the discriminative power of the covariance feature. Experiments show that, by combining the Haar and covariance features, we speed up the original covariance feature detector [1] by up to an order of magnitude in detection time with a slight drop in detection performance.
Sakrapee Paisitkriangkrai, Chunhua Shen, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2007 An efficient method for detecting ghost and left objects in surveillance video
abstract
This paper proposes an efficient method for detecting ghost and left objects in surveillance video, which, if not identified, may lead to errors or wasted computation in background modeling and object tracking in surveillance systems. This method contains two main steps: the first one is to detect stationary objects, which narrows down the evaluation targets to a very small number of foreground blobs; the second step is to discriminate the candidates between ghost and left objects. For the first step, we introduce a novel stationary object detection method based on continuous object tracking and shape matching. For the second step, we propose a fast and robust inpainting method to differentiate between ghost and left objects by constructing the real background using the candidate 's corresponding regions in the input and the background images. The effectiveness of our method has been validated by experiments over a variety of video sequences.
Sijun Lu, Jian Zhang 0002, David Dagan Feng
AVSS2
2007 Long-term Trajectory Extraction for Moving Vehicles
abstract
In recent years, trajectory analysis of moving vehicles in video-based traffic monitoring systems has drawn the attention of many researchers. Trajectory extraction is a fundamental step that is required prior to trajectory analysis. Lots of previous work have focused on trajectory extraction via tracking. However, they often fail to achieve long-term consistent trajectories. In this paper, we propose a robust approach for extracting long-term trajectories of moving vehicles in traffic monitoring using SIFT-descriptor. Experimental results show that the proposed method outperforms tracking-based techniques.
Jie Xu 0008, Getian Ye, Jian Zhang 0002
MMSP3
2007 Detecting unattended packages through human activity recognition and object association
Sijun Lu, Jian Zhang 0002, David Dagan Feng
Pattern Recognit.2
2006 A Knowledge-Based Approach for Detecting Unattended Packages in Surveillance Video
abstract
This paper describes a novel approach for detecting unattended packages in surveillance video. Unlike the traditional approach to just detecting stationary objects in monitored scenes, our approach detects unattended packages based on accumulated knowledge about human and non-human objects from continuous object tracking and classification. We design different reasoning rules for detecting different scenarios of the unattended package events. In the case where a package is left unattended by a single person explicitly, a rule using human activity recognition is introduced to decide the package ownership. In the case where a suspicious package is dropped down by a group of humans or under heavy occlusions, a rule based on historic tracking and classification information is proposed. Furthermore, an additional rule is given to reduce false alarms that may happen with traditional stationary object detection methods.
Sijun Lu, Jian Zhang 0002, David Dagan Feng
AVSS2
2005 Classification of Moving Humans Using Eigen-Features and Support Vector Machines
Sijun Lu, Jian Zhang 0002, David Dagan Feng
CAIP2
2005 Detecting New Stable Objects In Surveillance Video
abstract
We describe a novel method to detect new stable objects in video. This includes detecting new objects that appear in a scene and remain stationary for a period of time. Examples include detecting a dropped bag or a parked car. Our method utilizes the state transition history (or a record of the "life cycle") of individual Gaussian distributions in a Gaussian Mixture Model (GMM) used to model the background. In typical implementations of the GMM, this state transition information is ignored however we show that by observing and retaining the history of state transitions of individual distributions, it is possible to detect long term changes in a scene. In particular we identify changes to the most probable background distribution and impose certain conditions on the characteristics and temporal behavior of this distribution. Results presented in this paper illustrate the success of the proposed method and its relevance to surveillance applications.
Reji Mathew, Zhenghua Yu, Jian Zhang 0002
MMSP3
2000 A cell-loss concealment technique for MPEG-2 coded video
abstract
Audio-visual and other multimedia services are seen as important sources of traffic for future telecommunication networks, including wireless networks. A major drawback with some wireless networks is that they introduce a significant number of transmission errors into the digital bitstream. For video, such errors can have the effect of degrading the quality of service to the point where it is unusable. We introduce a technique that allows for the concealment of the impact of these errors. Our work is based on MPEG-2 encoded video transmitted over a wireless network whose data structures are similar to those of asynchronous transfer mode (ATM) networks. Our simulations include the impact of the MPEG-2 systems layer and cover cell-loss rates up to 5%. This is substantially higher than those that have been discussed in the literature up to this time. We demonstrate that our new approach can significantly increase received video quality, but at the cost of a considerable computational overhead. We then extend our technique to allow for higher computational efficiency and demonstrate that a significant quality improvement is still possible.
Jian Zhang 0002, Michael R. Frater, John F. Arnold
IEEE Trans. Circuits Syst. Video Technol.1
1999 Error resilience in the MPEG-2 video coding standard for cell based networks - A review
John F. Arnold, Michael R. Frater, Jian Zhang 0002
Signal Process. Image Commun.3
1999 MPEG 2 video error resilience experiments: : The importance considering the impact of the systems layer
Michael R. Frater, John F. Arnold, Jian Zhang 0002
Signal Process. Image Commun.3
1997 MPEG 2 Video Services for Wireless ATM Networks
abstract
Audio-visual and other multimedia services are seen as an important source of traffic for future telecommunications networks, including wireless networks. We examine the impact of the properties of a 50 Mb/s asynchronous transfer mode (ATM)-based wireless local-area network (WLAN) on Moving Picture Experts Group phase 2 (MPEG 2) compressed video traffic, with emphasis on the network's error characteristics. The paper includes a description of the WLAN system used and its loss characteristics, a brief discussion of relevant aspects of the MPEG 2 standards and the associated error resilience techniques for minimizing the effect of transmission errors, and a description of the method by which the video data is organized for transmission on the network. We show results on the effect of cell loss due to transmission errors on the quality of the decoded video at the receiver, and demonstrate how error resilience techniques in both the systems and video layers of MPEG 2 can be used to improve the quality of service. Situations where up to 1% of the data is lost due to network transmission errors are examined. Most important among the findings are that error resilience experiments that do not take into account the effect of the MPEG 2 systems layer will tend to significantly overestimate the quality of received video, and that the error resilience techniques provided within the MPEG 2 standard are not sufficient to provide acceptable quality with acceptable overheads, but that this quality can be significantly increased by the addition of a small number of simple techniques.
Jian Zhang 0002, Michael R. Frater, John F. Arnold, Terence M. Percival
IEEE J. Sel. Areas Commun.1