Zhang Zhang 0001

dblp:94/2468-1 · DBLP profile ↗
← Back
78ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0001-9425-3065ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Law for Multimodal Large Language Model Supervised Fine-Tuning
abstract
YiFan Zhang, Tao Yu, Feng Li, Chaoyou Fu, Yibo Hu, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yifan Zhang 0004, Chaoyou Fu, Yibo Hu 0001, Qingsong Wen, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001
ACL (1)8
2026 Spectral decomposition and adaptation for non-stationary time series anomaly detection
Huanyu Zhang 0002, Yifan Zhang 0004, Jian Liang 0001, Zhang Zhang 0001, Liang Wang 0001
Neurocomputing4
2026 Layer-wise contrastive network for unsupervised graph representation learning
abstract
Unsupervised graph representation learning has emerged as a cornerstone for extracting meaningful insights from complex relational data. While contrastive learning paradigms have achieved notable success, they traditionally rely on stochastic data augmentations to generate multiple input views—a process that often incurs significant computational overhead and sensitivity to augmentation quality. To transcend these limitations, we propose the Layer-wise Contrastive Network (LCN), a novel and efficient paradigm that redefines the construction of contrastive views. Unlike conventional methods that rely on extrinsic data perturbations, LCN exploits the intrinsic architectural hierarchy of Graph Convolutional Networks. By treating distinct neural layers as different views of the same graph instance, we introduce a contrastive objective that enforces consistency between shallow and deep representations. This mechanism not only eliminates the need for expensive augmentation operations but also distills and preserves fundamental node characteristics from the original graph throughout deeper layers. Furthermore, LCN serves as a flexible plug-and-play framework, exemplified by its extension into Wide LCN, which integrates with traditional augmentation-based methods. Extensive evaluations across both transductive and inductive benchmarks demonstrate that our method achieves superior representational robustness and computational efficiency, offering a scalable and principled perspective for future graph-based contrastive learning. Our code is available at https://github.com/XiangluZhu/LCN.git . • We introduce a novel contrastive loss to learn graph representations by contrasting shallow and deep features. • Our method is flexible and can be readily combined with existing graph contrastive learning techniques that utilize data augmentation. • We demonstrate outstanding performance with our method across four benchmark datasets for node classification.
Xianglu Zhu, Zhang Zhang 0001, Zilei Wang, Liang Wang 0001, Tieniu Tan
Neurocomputing2
2026 Beyond LLaVA-HD: Diving Into High-Resolution Multimodal Large Language Models
abstract
Seeing clearly with high resolution is a foundation of Multimodal Large Language Models (MLLMs), which has been proven to be vital for visual perception and reasoning. Existing works usually employ a straightforward resolution upscaling method, where the image consists of global and local branches, with the latter being the sliced image patches but resized to the same resolution as the former. This means that higher resolution requires more local patches, resulting in exorbitant computational expenses, and meanwhile, the dominance of local image tokens may diminish the global context. In this paper, we dive into the problems and propose a new framework as well as an elaborate optimization strategy. Specifically, we extract contextual information from the global view using a mixture of adapters, based on the observation that different adapters excel at different tasks. With regard to local patches, learnable query embeddings are introduced to reduce image tokens, the important tokens most relevant to the user question will be further selected by a similarity-based selector. Our empirical results demonstrate a 'less is more' pattern, where utilizing fewer but more informative local image tokens leads to improved performance. Besides, a significant challenge lies in the training strategy, as simultaneous end-to-end training of the global mining block and local compression block does not yield optimal results. We thus advocate for an alternating training way, ensuring balanced learning between global and local aspects. Finally, we also introduce a challenging dataset with high requirements for image detail, enhancing the training of the local compression layer. The proposed method, termed MLLM with Sophisticated Tasks, Local image compression, and Mixture of global Experts (SliME), achieves leading performance across various benchmarks with only 2 million training data.
YiFan Zhang, Qingsong Wen, Chaoyou Fu, Kun Wang 0056, Xue Wang 0010, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Model-Free Test Time Adaptation for Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is essential for the reliability of ML models. Most existing methods for OOD detection learn a fixed decision criterion from a given in-distribution dataset and apply it universally to decide if a data point is OOD. Recent work Fang et al. (2022) shows that given only in-distribution data, it is impossible to reliably detect OOD data without extra assumptions. Motivated by the theoretical result and recent exploration of test-time adaptation methods, we propose a Non-Parametric Test Time Adaptation framework for Out-Of-Distribution Detection (AdaODD). Unlike conventional methods, AdaODD utilizes online test samples for model adaptation during testing, enhancing adaptability to changing data distributions. The framework incorporates detected OOD instances into decision-making, reducing false positive rates, particularly when ID and OOD distributions overlap significantly. We demonstrate the effectiveness of AdaODD through comprehensive experiments on multiple OOD detection benchmarks, extensive empirical studies show that AdaODD significantly improves the performance of OOD detection over state-of-the-art methods. Specifically, AdaODD reduces the false positive rate (FPR95) by 23.23% on the CIFAR-10 benchmarks and 38% on the ImageNet-1 k benchmarks compared to the advanced methods. Lastly, we theoretically verify the effectiveness of AdaODD.
Yifan Zhang 0004, Xue Wang 0010, Tian Zhou 0004, Kun Yuan 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Mitigating Catastrophic Forgetting in Online Continual Learning With Dual-Margin Contrastive Replay
abstract
Online Continual Learning (OCL) enables machine learning models to learn from a stream of non-stationary tasks, making it more aligned with real-world scenarios. However, OCL faces a significant challenge: catastrophic forgetting, wherein the model learned in previous tasks is substantially overwritten upon encountering new tasks, leading to a biased forgetting of prior knowledge. Among various OCL strategies, replay-based methods have proven particularly effective in mitigating catastrophic forgetting by maintaining a small buffer of past samples and retraining them alongside new data. However, due to strict memory constraints, these replay buffers often fail to adequately represent the true data distribution of previous tasks. This leads to distributional shifts in the feature space, amplifying forgetting and degrading model performance. To address the problem, in this paper, we propose a novel replay strategy, termed Dual-Margin Contrastive Replay (DMCR), to anchor the distribution of old tasks and reduce the negative transfer effects. First, we propose to select memory for more representative samples guided by constructed centroids in a data stream. Then, to keep the model from distribution chaos in biased replay, a two-level angular cross-task Contrastive Margin Loss (CML) is proposed, to encourage the intra-class and intra-task compactness, and increase the inter-class and inter-task discrepancy. Finally, to further suppress the distributional drift, we present an optional Centroid Distillation Loss (CDL) on the replay memory to anchor the knowledge in feature space for each previous old task. Extensive experimental results on five benchmark datasets validate that the proposed DMCR can effectively mitigate the catastrophic forgetting and achieve state-of-the-art (SOTA) performance in OCL.
Fan Lyu, Gongbo Cheng, Daofeng Liu, Linglan Zhao, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 GAIN: Global-Atomic INteraction Graph for Few-Shot Class-Incremental Learning
Fan Lyu, Linglan Zhao, Chengyan Liu, Yinying Mei, Zhang Zhang 0001, Baoqing Yu, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 MambaPTP: Exploring the Potential of Mamba for Pedestrian Trajectory Prediction
abstract
Pedestrian Trajectory Prediction (PTP) aims to predict the future trajectory of pedestrians based on a historical trajectory. Transformer-based approaches have demonstrated unparalleled performance for PTP tasks, encoding long-term temporal dependencies and heterogeneous spatial interactions of pedestrians. However, Transformer often involves redundant information and noisy interactions from irrelevant regions by considering all available trajectory features. Recently, the structured state space model, Mamba has been proposed, which captures long-range dependency in sequences with a selective mechanism to filter out redundant information. To further tap into the potential of the novel Mamba architecture for the PTP task, in this paper, we presentMambaPTP, which predicts future trajectories based purely on Mamba mechanisms, to mitigate the noisy interactions of irrelevant trajectory features and avoid repetitive trajectory modeling, while maintaining high-performance trajectory prediction. Specifically, we propose a new Bidirectional Gating Mamba (BGM) module with bidirectional state space models, which leverages the sparse gate mechanism to select informative temporal patterns and spatial interactions. Moreover, we design a Bidirectional Trajectory Alignment (BTA) module towards aligning the predicted trajectory to the ground truth, ensuring that the model to learn the effective sparse feature representation of trajectories. We conduct extensive experiments on several mainstream pedestrian trajectory prediction datasets. The results demonstrate that the proposed MambaPTP achieves competitive performance compared to advanced Transformer-based models. We hope this paper can further inspire research in Mamba for the PTP task, leading to a tighter integration of the Mamba and PTP communities.
Shuangqing Zhang, Gangming Zhao, Fan Lyu, Songping Wang, Zhang Zhang 0001, Fang Zhao 0006, Caifeng Shan, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Efficient RGBT Tracking via Early Fusion and Hierarchical Knowledge Distillation
Jinhu Wang, Mai Wen, Zhang Zhang 0001, Liang Wang 0001, Chenglong Li 0002
ICIG (2)3
2025 MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
abstract
Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult to measure the significant challenges that models face in the real world, including: 1) small data scale leads to a large performance variance; 2) reliance on model-based annotations results in restricted data quality; 3) insufficient task difficulty, especially caused by the limited image resolution. To tackle these issues, we introduce MME-RealWorld. Specifically, we collect more than $300$ K images from public datasets and the Internet, filtering $13,366$ high-quality images for annotation. This involves the efforts of professional $25$ annotators and $7$ experts in MLLMs, contributing to $29,429$ question-answer pairs that cover $43$ subtasks across $5$ real-world scenarios, extremely challenging even for humans. As far as we know, **MME-RealWorld is the largest manually annotated benchmark to date, featuring the highest resolution and a targeted focus on real-world applications**. We further conduct a thorough evaluation involving $29$ prominent MLLMs, such as GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet. Our results show that even the most advanced models struggle with our benchmarks, where none of them reach 60\% accuracy. The challenges of perceiving high-resolution images and understanding complex real-world scenarios remain urgent issues to be addressed. The data and evaluation code are released in our Project Page.
Yifan Zhang 0004, Huanyu Zhang 0002, Haochen Tian 0001, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Kun Wang 0056, Qingsong Wen, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001
ICLR10
2025 Controllable Continual Test-Time Adaptation
abstract
Continual Test-Time Adaptation (CTTA) is an emerging and challenging task where a model trained in a source domain must adapt to continuously changing conditions during testing, without access to the original source data. CTTA is prone to error accumulation due to uncontrollable domain shifts, leading to blurred decision boundaries between categories. Existing CTTA methods primarily focus on suppressing domain shifts, which proves inadequate during the unsupervised test phase. In contrast, we introduce a novel approach that guides rather than suppresses these shifts. Specifically, we propose Controllable Continual Test-Time Adaptation (C-CoTTA), which explicitly prevents any single category from encroaching on others, thereby mitigating the mutual influence between categories caused by uncontrollable shifts. Moreover, our method reduces the sensitivity of model to domain transformations, thereby minimizing the magnitude of category shifts. Extensive quantitative experiments demonstrate the effectiveness of our method, while qualitative analyses, such as t-SNE plots, confirm the theoretical validity of our approach. Our code is available at https://github.com/RenshengJi/C-CoTTA.
Ziqi Shi, Fan Lyu, Fanhua Shang, Fuyuan Hu, Wei Feng 0005, Zhang Zhang 0001, Liang Wang 0001
ICME7
2025 MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
abstract
Existing efforts to align multimodal large language models (MLLMs) with human preferences have only achieved progress in narrow areas, such as hallucination reduction, but remain limited in practical applicability and generalizability. To this end, we introduce **MM-RLHF**, a dataset containing **120k** fine-grained, human-annotated preference comparison pairs. This dataset represents a substantial advancement over existing resources, offering superior size, diversity, annotation granularity, and quality. Leveraging this dataset, we propose several key innovations to improve both the quality of reward models and the efficiency of alignment algorithms. Notably, we introduce the **Critique-Based Reward Model**, which generates critiques of model outputs before assigning scores, offering enhanced interpretability and more informative feedback compared to traditional scalar reward mechanisms. Additionally, we propose **Dynamic Reward Scaling**, a method that adjusts the loss weight of each sample according to the reward signal, thereby optimizing the use of high-quality comparison pairs. Our approach is rigorously evaluated across **10** distinct dimensions, encompassing **27** benchmarks, with results demonstrating significant and consistent improvements in model performance (Figure.1).
Yifan Zhang 0004, Haochen Tian 0001, Chaoyou Fu, Peiyan Li 0001, Jianshu Zeng, Wulin Xie, Yang Shi 0009, Huanyu Zhang 0002, Junkang Wu, Xue Wang 0010, Yibo Hu 0001, Tingting Gao, Zhang Zhang 0001, Fan Yang 0094, Di Zhang 0026, Liang Wang 0001, Rong Jin 0001
ICML15
2025 Debiasing Multimodal Large Language Models via Penalization of Language Priors
abstract
In the realms of computer vision and natural language processing, Multimodal Large Language Models (MLLMs) have become indispensable tools, proficient in generating textual responses based on visual inputs. Despite their advancements, our investigation reveals a noteworthy bias: the generated content is often driven more by the inherent priors of the underlying Large Language Models (LLMs) than by the input image. Empirical experiments underscore the persistence of this bias, as MLLMs often provide confident answers even in the absence of relevant images or given incongruent visual inputs. To rectify these biases and redirect the model's focus toward visual information, we propose two simple, training-free strategies. First, for tasks such as classification or multi-choice question answering, we introduce a ''Post-Hoc Debias'' method using an affine calibration step to adjust the output distribution. This approach ensures uniform answer scores when the image is absent, acting as an effective regularization technique to alleviate the influence of LLM priors. For more intricate open-ended generation tasks, we extend this method to ''Visual Debias Decoding'', which mitigates bias by contrasting token log-probabilities conditioned on a correct image versus a meaningless one. Additionally, our investigation sheds light on the instability of MLLMs across various decoding configurations. Through systematic exploration of different settings, we achieve significant performance improvements-surpassing previously reported results-and raise concerns about the fairness of current evaluation practices. Comprehensive experiments substantiate the effectiveness of our proposed strategies in mitigating biases. These strategies not only prove beneficial in minimizing hallucinations but also contribute to the generation of more helpful and precise illustrations.
Yifan Zhang 0004, Yang Shi 0009, Weichen Yu, Qingsong Wen, Xue Wang 0010, Wenjing Yang 0002, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001
ACM Multimedia7
2025 DAA: Amplifying Unknown Discrepancy for Test-Time Discovery
abstract
Test-Time Discovery (TTD) addresses the critical challenge of identifying and adapting to novel classes during inference while maintaining performance on known classes, which is a capability essential for dynamic real-world environments such as healthcare and autonomous driving. Recent TTD methods adopt training-free, memory-based strategies but rely on frozen models and static representations, resulting in poor generalization. In this paper, we propose a Discrepancy-Amplifying Adapter (DAA), a trainable module that enables real-time adaptation by amplifying feature-level discrepancies between known and unknown classes. During training, DAA is optimized using simulated unknowns and a novel warm-up strategy to enhance its discriminative capacity. To ensure continual adaptation at test time, we introduce a Short-Term Memory Renewal (STMR) mechanism, which maintains a queue-based memory for unknown classes and selectively refreshes prototypes using recent, reliable samples. DAA is further updated through self-supervised learning, promoting knowledge retention for known classes while improving discrimination of emerging categories. Extensive experiments show that our method maintains high adaptability and stability, and significantly improves novel class discovery performance. Our code will be available.
Fan Lyu, Chenggong Ni, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
NeurIPS4
2025 MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
abstract
Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios.
Yang Shi 0009, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yifan Zhang 0004, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Zhuoran Zhang 0003, Xinlong Chen, Bohan Zeng, Yushuo Guan, Zhang Zhang 0001, Liang Wang 0001, Haoxuan Li 0001, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan 0001, Haotian Wang 0001, Wenjing Yang 0002
NeurIPS16
2025 CSFRNet: Integrating Clothing Status Awareness for Long-Term Person Re-identification
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
Int. J. Comput. Vis.3
2025 TimeRAF: Retrieval-Augmented Foundation Model for Zero-Shot Time Series Forecasting
abstract
Time series forecasting plays a crucial role in data mining, driving rapid advancements across numerous industries. With the emergence of large models, time series foundation models (TSFMs) have exhibited remarkable generalization capabilities, such as zero-shot learning, through large-scale pre-training. Meanwhile, Retrieval-Augmented Generation (RAG) methods have been widely employed to enhance the performance of foundation models on unseen data, allowing models to access to external knowledge. In this paper, we introduceTimeRAF, aRetrieval-AugmentedForecasting model that enhance zero-shot time series forecasting through retrieval-augmented techniques. We develop customized time series knowledge bases that are tailored to the specific forecasting tasks. TimeRAF employs an end-to-end learnable retriever to extract valuable information from the knowledge base. Additionally, we propose Channel Prompting for knowledge integration, which effectively extracts relevant information from the retrieved knowledge along the channel dimension. Extensive experiments demonstrate the effectiveness of our model, showing significant improvement across various domains and datasets.
Huanyu Zhang 0002, Chang Xu 0008, Yifan Zhang 0004, Zhang Zhang 0001, Liang Wang 0001, Jiang Bian 0002
IEEE Trans. Knowl. Data Eng.4
2025 Multi-View Knowledge Guided Semantic Prototype Learning for Generalized Zero-Shot Action Recognition
abstract
Generalized zero-shot skeleton-based action recognition (GZSSAR) is an emerging and challenging problem in the computer vision community. It requires models to recognize human actions, including some classes that are unseen during training. Previous studies typically rely solely on action labels to bridge the gap between seen and unseen action classes. However, the limited action semantic information hinders the learning of comprehensive semantic prototypes, thereby restricting the model's ability to generalize to unseen classes. To address this issue, in addition to the original action labels, we explore four types of textual action descriptions (i.e., interpretive and motional descriptions derived from manual expert annotation and large language model) for each action class. In order to comprehensively utilize multi-view semantic information for zeroshot classification, an Attentional Multi-view Semantic Fusion (AMSF) model is proposed. It effectively integrates the multi-view semantic features and aligns visual and semantic features in a common space, subsequently realizing the recognition of unseen action classes. Furthermore, previous works typically evaluate models in settings that include specific unseen classes, which is insufficient for GZSSAR research. To thoroughly evaluate different models, we introduce two novel distinct experimental settings, termed the “easy setting” and the “hard setting”, based on the semantic similarities between action classes. Extensive experimental results on three large-scale skeleton-based action recognition benchmarks (PKU-MMD, NTU-60, and NTU-120) not only validate the advantages of the proposed multi-view action descriptions and the AMSF model but also demonstrate the rationality of the novel experimental settings. All the data and code of this paper are publicly available on GitHub.
Ming-Zhe Li, Zhang Zhang 0001, Yaoning Li, Zhanyu Ma, Liang Wang 0001
IEEE Trans. Multim.3
2024 Attribute-Guided Pedestrian Retrieval: Bridging Person Re-ID with Internal Attribute Variability
abstract
In various domains such as surveillance and smart retail, pedestrian retrieval, centering on person re-identification (Re-ID), plays a pivotal role. Existing Re-ID methodologies often overlook subtle internal attribute variations, which are crucial for accurately identifying individuals with changing appearances. In response, our paper introduces the Attribute-Guided Pedestrian Retrieval (AGPR) task, focusing on integrating specified attributes with query images to refine retrieval results. Although there has been progress in attribute-driven image retrieval, there remains a notable gap in effectively blending robust Re-ID models with intra-class attribute variations. To bridge this gap, we present the Attribute-Guided Transformer-based Pedestrian Retrieval (ATPR) framework. ATPR adeptly merges global ID recognition with local attribute learning, ensuring a co-hesive linkage between the two. Furthermore, to effectively handle the complexity of attribute interconnectivity, ATPR organizes attributes into distinct groups and applies both inter-group correlation and intra-group decorrelation regularizations. Our extensive experiments on a newly estab-lished benchmark using the RAP dataset [32] demonstrate the effectiveness of ATPR within the AGPR paradigm.
Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
CVPR2
2024 Learning Energy-Based Models for 3D Human Pose Estimation
abstract
Recently, 3D human pose estimation has attracted more attention due to its promising applications. In general, existing methods usually directly predict a target 3D pose for a given input using a Deep Neural Network (DNN), and train the DNN by minimizing the mean squared error (MSE) loss. Despite the impressive performance of these methods, they create a fixed-variance Gaussian model of the conditional target density (the distribution for the target 3D pose given the input) from a probabilistic perspective, which significantly restricts the expressive capabilities of the learned conditional target density. Thus, this hinders the complete utilization of the predictive potential embedded within the DNN. We tackle this problem by delving into the latest developments in conditional energy-based models (EBMs) for probabilistic regression. In this work, we design a simple yet effective network to learn an energy function from 2D and 3D joints pairs. Then a gradient-based refinement procedure is adopted to minimize the energy function to find the corresponding target 3D pose. In this way, we can apply the energy-based model to refine the initial 3D joints estimated by the state-of-the-art 3D human pose estimator. Extensive experiments are conducted on two popular benchmarks on human pose estimation and the results demonstrate the superiority of our method over existing state-of-the-art approaches.
Xianglu Zhu, Zhang Zhang 0001, Wei Wang 0115, Zilei Wang, Liang Wang 0001
IJCNN2
2024 Understanding Driving Risks via Prompt Learning
abstract
Understanding driving risks is crucial for enhancing driving safety. It is a challenging task to evaluate driving risks in various complex driving scenarios. Inspired by prompt-based learning, we propose an end-to-end approach for identifying the highest-risk object in the current driving scenario based on a learnable risk pool. Specifically, a method based on key-value pair matching is designed to build a memory system for learning a collection of risk prototypes. Extensive experiments on the DRAMA dataset show that the proposed method achieves an improvement of 18.6% in Mean-IOU and 3.0% in B4 score compared to the state-of-the-art (SOTA) methods, which indicates that our method can effectively localize risky objects and accurately describe the driving scenes.
Yubo Chang, Fan Lyu, Zhang Zhang 0001, Liang Wang 0001
SMC3
2024 HDD4DBP: A Large-Scale Multi-Modal Benchmark on Driving Behavior Prediction
abstract
Driving behavior prediction (DBP) plays an important role in autonomous driving. Accurately anticipating driving behaviors (e.g., left turn, right turn) of ego-vehicles 3–5 seconds before actual occurrences of maneuvers can help AI pilots planning safer trajectories or informing possible dangers to drivers. However, current DBP benchmark datasets, e.g., the Brain4Cars, are restricted by the limited amount of data samples and small number of behavior categories. Thus, there is an urgent need for creating a new larger-scale benchmark to help boost the studies of DBP. In this work, based on the annotations in the HRI Driving Dataset (HDD), we extract 5000+ clips from untrimmed multi-modal data sequences to form a new dataset, termed the HDD4DBP, as a convincing large-scale testbed for evaluating various DPB methods. Furthermore, inspired by the success of transformers on modeling long range dependence in sequences, we build a strong baseline with the vision transformer (ViT) backbone for predicting driving behaviors. Compared to previous representative baselines, a large margin performance gain can be achieved by our strong baseline on the HDD4DBP. Moreover, simply fine-tuning the pre-trained strong baseline can obtain the state-of-the-art performance on the Brain4Cars dataset, thereby futher confirming the benefits of the HDD4DBP. The dataset and source code will be released. Find data and code at https://github.com/HDD4DBP/MTL.
Qinghui Dong, Zhang Zhang 0001, Yubo Chang, Liang Wang 0001
SMC2
2024 Class Hierarchy-Guided Generalized Few-Shot Ship Detection in Remote Sensing Images
abstract
Fine-grained ship detection in remote sensing images (RSIs) depends heavily on numerous training data with expensive manual annotations. Learning novel ship categories from very few labeled samples and without forgetting the learned knowledge of seen categories is important to real-world applications. In this letter, we formulate fine-grained ship detection in RSIs as a problem of generalized few-shot object detection (G-FSOD). Existing methods often neglect the structured information in ship taxonomy, and thus result in mutually exclusive representations between base and novel classes and hinder the transfer of the learned knowledge to the novel concepts under the few-shot settings. To handle this problem, we propose to incorporate the inherent hierarchical taxonomy in ship classes into the generalized few-shot ship detection to leverage the shared knowledge among base and novel classes. In particular, a ship detector is trained based on the coarsest class labels and a multitask classification network is built to distinguish various ships at both coarse and fine-grained levels on base classes, which leads to a generalized ship representation between base classes to novel classes. To build the classifier of novel classes, a prototype bank is constructed with the few-shot samples of novel classes, without the wreck of the feature extractor so as to maintain the performance on base classes. Extensive experiments on two large-scale ship detection datasets demonstrate the effectiveness of our method against state-of-the-art methods.
Shuangqing Zhang, Zhang Zhang 0001, Da Li 0003, Chenglong Li 0002, Liang Wang 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 Customized meta-dataset for automatic classifier accuracy evaluation
Yan Huang 0023, Zhang Zhang 0001, Yan Huang 0008, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
Pattern Recognit.2
2024 Enhancing Person Re-Identification Performance Through In Vivo Learning
abstract
This research investigates the potential of in vivo learning to enhance visual representation learning for image-based person re-identification (re-ID). Compared to traditional self-supervised learning (which require external data), the introduced in vivo learning utilizes supervisory labels generated from pedestrian images to improve re-ID accuracy without relying on external data sources. Three carefully designed in vivo learning tasks, leveraging statistical regularities within images, are proposed without the need for laborious manual annotations. These tasks enable feature extractors to learn more comprehensive and discriminative person representations by jointly modeling various aspects of human biological structure information, contributing to enhanced re-ID performance. Notably, the method seamlessly integrates with existing re-ID frameworks, requiring minimal modifications and no additional data beyond the existing training set. Extensive experiments on diverse datasets, including Market1501, CUHK03-NP, Celeb-reID, Celeb-reid-light, PRCC, and LTCC, demonstrate substantial enhancements in rank-1 precision compared to state-of-the-art methods.
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
IEEE Trans. Image Process.3
2024 Meta Clothing Status Calibration for Long-Term Person Re-Identification
abstract
Recent studies have seen significant advancements in the field of long-term person re-identification (LT-reID) through the use of clothing-irrelevant or insensitive features. This work takes the field a step further by addressing a previously unexplored issue, the Clothing Status Distribution Shift (CSDS). CSDS refers to the differing ratios of samples with clothing changes to those without clothing changes between the training and test sets, leading to a decline in LT-reID performance. We establish a connection between the performance of LT-reID and CSDS, and argue that addressing CSDS can improve LT-reID performance. To that end, we propose a novel framework called Meta Clothing Status Calibration (MCSC), which uses meta-learning to optimize the LT-reID model. Specifically, MCSC simulates CSDS between meta-train and meta-test with meta-optimization objectives, optimizing the LT-reID model and making it robust to CSDS. This framework is designed to prevent overfitting and improve the generalization ability of the LT-reID model in the presence of CSDS. Comprehensive evaluations on seven datasets demonstrate that the proposed MCSC framework effectively handles CSDS and improves current state-of-the-art LT-reID methods on several LT-reID benchmarks.
Yan Huang 0023, Qiang Wu 0001, Zhang Zhang 0001, Caifeng Shan, Yan Huang 0008, Yi Zhong 0002, Liang Wang 0001
IEEE Trans. Image Process.3
2024 LogoRA: Local-Global Representation Alignment for Robust Time Series Classification
abstract
Unsupervised domain adaptation (UDA) of time series aims to teach models to identify consistent patterns across various temporal scenarios, disregarding domain-specific differences, which can maintain their predictive accuracy and effectively adapt to new domains. However, existing UDA methods struggle to adequately extract and align both global and local features in time series data. To address this issue, we propose theLocal-GlobalRepresentationAlignment framework (LogoRA), which employs a two-branch encoder–comprising a multi-scale convolutional branch and a patching transformer branch. The encoder enables the extraction of both local and global representations from time series. A fusion module is then introduced to integrate these representations, enhancing domain-invariant feature alignment from multi-scale perspectives. To achieve effective alignment, LogoRA employs strategies like invariant feature learning on the source domain, utilizing triplet loss for fine alignment and dynamic time warping-based feature alignment. Additionally, it reduces source-target domain gaps through adversarial training and per-class prototype alignment. Our evaluations on four time-series datasets demonstrate that LogoRA outperforms strong baselines by up to 12.52%, showcasing its superiority in time series UDA tasks.
Huanyu Zhang 0002, Yifan Zhang 0004, Zhang Zhang 0001, Qingsong Wen, Liang Wang 0001
IEEE Trans. Knowl. Data Eng.3
2024 Illumination Distillation Framework for Nighttime Person Re-Identification and a New Benchmark
abstract
Nighttime person Re-ID (person re-identification in the nighttime) is a very important and challenging task for visual surveillance but it has not been thoroughly investigated. Under the low illumination condition, the performance of person Re-ID methods usually sharply deteriorates. To address the low illumination challenge in nighttime person Re-ID, this paper proposes an Illumination Distillation Framework (IDF), which utilizes illumination enhancement and illumination distillation schemes to promote the learning of Re-ID models. Specifically, IDF consists of a master branch, an illumination enhancement branch, and an illumination distillation module. The master branch is used to extract the features from a nighttime image. The illumination enhancement branch first estimates an enhanced image from the nighttime image using a nonlinear curve mapping method and then extracts the enhanced features. However, nighttime and enhanced features usually contain data noise due to unstable lighting conditions and enhancement failures. To fully exploit the complementary benefits of nighttime and enhanced features while suppressing data noise, we propose an illumination distillation module. In particular, the illumination distillation module fuses the features from two branches through a bottleneck fusion model and then uses the fused features to guide the learning of both branches in a distillation manner. In addition, we build a real-world nighttime person Re-ID dataset, namedNight600, which contains 600 identities captured from different viewpoints and nighttime illumination conditions under complex outdoor environments. Experimental results demonstrate that our IDF can achieve state-of-the-art performance on two nighttime person Re-ID datasets (i.e.,Night600andKnight). We will release our code and dataset athttps://github.com/Alexadlu/IDF.
Andong Lu, Zhang Zhang 0001, Yan Huang 0023, Yifan Zhang 0004, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001
IEEE Trans. Multim.2
2024 Pedestrian Attribute Recognition via Spatio-temporal Relationship Learning for Visual Surveillance
abstract
Pedestrian attribute recognition (PAR) aims at predicting the visual attributes of a pedestrian image. PAR has been used as soft biometrics for visual surveillance and IoT security. Most of the current PAR methods are developed based on discrete images. However, it is challenging for the image-based method to handle the occlusion and action-related attributes in real-world applications. Recently, video-based PAR has attracted much attention in order to exploit the temporal cues in the video sequences for better PAR. Unfortunately, existing methods usually ignore the correlations among different attributes and the relations between attributes and spatio regions. To address this problem, we propose a novel method for video-based PAR by exploring the relationships among different attributes in both the spatio and temporal domains. More specifically, a spatio-temporal saliency module (STSM) is introduced to capture the key visual patterns from the video sequences, and a module for spatio-temporal attribute relationship learning (STARL) is proposed to mine the correlations among these patterns. Meanwhile, a large-scale benchmark for video-based PAR, RAP-Video, is built by extending the image-based dataset RAP-2, which contains 83,216 tracklets with 25 scenes. To the best of our knowledge, this is the largest dataset for video-based PAR. Extensive experiments are performed on the proposed benchmark as well as on MARS Attribute and DukeMTMC-Video Attribute. The superior performance demonstrates the effectiveness of the proposed method.
Da Li 0003, Zhang Zhang 0001, Peng Zhang 0057, Caifeng Shan, Jungong Han
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Multi-semantic Fusion Model For Generalized Zero-Shot Skeleton-Based Action Recognition
Ming-Zhe Li, Zhang Zhang 0001, Zhanyu Ma, Liang Wang 0001
ICIG (1)3
2023 Free Lunch for Domain Adversarial Training: Environment Label Smoothing
Yifan Zhang 0004, Xue Wang 0010, Jian Liang 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan
ICLR4
2023 AdaNPC: Exploring Non-Parametric Classifier for Test-Time Adaptation
abstract
Many recent machine learning tasks focus to develop models that can generalize to unseen distributions. Domain generalization (DG) has become one of the key topics in various fields. Several literatures show that DG can be arbitrarily hard without exploiting target domain information. To address this issue, test-time adaptive (TTA) methods are proposed. Existing TTA methods require offline target data or extra sophisticated optimization procedures during the inference stage. In this work, we adopt Non-Parametric Classifier to perform the test-time Adaptation (AdaNPC). In particular, we construct a memory that contains the feature and label pairs from training domains. During inference, given a test instance, AdaNPC first recalls $k$ closed samples from the memory to vote for the prediction, and then the test feature and predicted label are added to the memory. In this way, the sample distribution in the memory can be gradually changed from the training distribution towards the test distribution with very little extra computation cost. We theoretically justify the rationality behind the proposed method. Besides, we test our model on extensive numerical experiments. AdaNPC significantly outperforms competitive baselines on various DG benchmarks. In particular, when the adaptation target is a series of domains, the adaptation accuracy of AdaNPC is $50$% higher than advanced TTA methods.
Yifan Zhang 0004, Xue Wang 0010, Kexin Jin, Kun Yuan 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan
ICML5
2023 Domain-Specific Risk Minimization for Domain Generalization
abstract
Domain generalization (DG) approaches typically use the hypothesis learned on source domains for inference on the unseen target domain. However, such a hypothesis can be arbitrarily far from the optimal one for the target domain, induced by a gap termed ''adaptivity gap.'' Without exploiting the domain information from the unseen test samples, adaptivity gap estimation and minimization are intractable, which hinders us to robustify a model to any unknown distribution. In this paper, we first establish a generalization bound that explicitly considers the adaptivity gap. Our bound motivates two strategies to reduce the gap: the first one is ensembling multiple classifiers to enrich the hypothesis space, then we propose effective gap estimation methods for guiding the selection of a better hypothesis for the target. The other method is minimizing the gap directly by adapting model parameters using online target samples. We thus propose Domain-specific Risk Minimization (DRM). During training, DRM models the distributions of different source domains separately; for inference, DRM performs online model steering using the source hypothesis for each arriving target sample. Extensive experiments demonstrate the effectiveness of the proposed DRM for domain generalization. Code is available at: https://github.com/yfzhang114/AdaNPC.
Yifan Zhang 0004, Jindong Wang 0001, Jian Liang 0001, Zhang Zhang 0001, Baosheng Yu, Liang Wang 0001, Dacheng Tao, Xing Xie 0001
KDD4
2023 OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online Ensembling
abstract
Online updating of time series forecasting models aims to address the concept drifting problem by efficiently updating forecasting models based on streaming data. Many algorithms are designed for online time series forecasting, with some exploiting cross-variable dependency while others assume independence among variables. Given every data assumption has its own pros and cons in online time series modeling, we propose **On**line **e**nsembling **Net**work (**OneNet**). It dynamically updates and combines two models, with one focusing on modeling the dependency across the time dimension and the other on cross-variate dependency. Our method incorporates a reinforcement learning-based approach into the traditional online convex programming framework, allowing for the linear combination of the two models with dynamically adjusted weights. OneNet addresses the main shortcoming of classical online learning methods that tend to be slow in adapting to the concept drift. Empirical results show that OneNet reduces online forecasting error by more than $\mathbf{50}\\%$ compared to the State-Of-The-Art (SOTA) method.
Yifan Zhang 0004, Qingsong Wen, Xue Wang 0010, Liang Sun 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan
NeurIPS6
2023 Dual-focus transfer network for zero-shot learning
Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001, Tieniu Tan
Neurocomputing2
2023 Constructing Stronger and Faster Baselines for Skeleton-Based Action Recognition
abstract
One essential problem in skeleton-based action recognition is how to extract discriminative features over all skeleton joints. However, the complexity of the recent State-Of-The-Art (SOTA) models for this task tends to be exceedingly sophisticated and over-parameterized. The low efficiency in model training and inference has increased the validation costs of model architectures in large-scale datasets. To address the above issue, recent advanced separable convolutional layers are embedded into an early fused Multiple Input Branches (MIB) network, constructing an efficient Graph Convolutional Network (GCN) baseline for skeleton-based action recognition. In addition, based on such the baseline, we design a compound scaling strategy to expand the model's width and depth synchronously, and eventually obtain a family of efficient GCN baselines with high accuracies and small amounts of trainable parameters, termed EfficientGCN-Bx, where "x" denotes the scaling coefficient. On two large-scale datasets, i.e., NTU RGB+D 60 and 120, the proposed EfficientGCN-B4 baseline outperforms other SOTA methods, e.g., achieving 92.1% accuracy on the cross-subject benchmark of NTU 60 dataset, while being 5.82× smaller and 5.85× faster than MS-G3D, which is one of the SOTA methods. The source code in PyTorch version and the pretrained models are available at https://github.com/yfsong0709/EfficientGCNv1.
Yi-Fan Song, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Few-shot learning with unsupervised part discovery and part-aligned similarity
Zhang Zhang 0001, Wei Wang 0025, Liang Wang 0001, Zilei Wang, Tieniu Tan
Pattern Recognit.2
2023 Diag-IoU Loss for Object Detection
abstract
Existing IoU-based loss functions have achieved promising performance for bounding box regression in object detection. However, they cannot fully reflect the relation between the predicted and target boxes in the case of box inclusions, and might thus deteriorate detection accuracy and efficiency. In this paper, we design a novel similarity measurement based on the box diagonal called Diag-IoU to well represent the divergence between the predicted and target boxes even in the case of box inclusions, and thus achieve superior localization accuracy and fast convergence. In particular, we equivalently represent a rectangular box with its box diagonal, which contains exclusive and informative geometrical factors, and define the Diag-IoU based on the similarities of a set of sampled point pairs from the predicted and target box diagonals. Based on the Diag-IoU, we design a general Diag-IoU loss, which can provide holistic information in measuring two boxes and thus differentiate the two boxes in the case of box inclusions. To validate the effectiveness of the proposed method, we apply the Diag-IoU loss to several representative object detectors, including YOLO v5s, Faster R-CNN, and FCOS. Extensive experiments on the synthetic data and two challenging object detection benchmark datasets, i.e., MS COCO and PASCAL VOC, demonstrate the superior performance of the proposed Diag-IoU loss compared to previous IoU-based losses as well as other metrics.
Shuangqing Zhang, Chenglong Li 0002, Lei Liu 0049, Zhang Zhang 0001, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Incremental Pedestrian Attribute Recognition via Dual Uncertainty-Aware Pseudo-Labeling
abstract
Incremental pedestrian attribute recognition (IncPAR) aims to learn novel person attributes continuously and avoid the catastrophic forgetting, which is an essential problem for image forensic and security applications, e.g., suspect search. Different from the conventional continual learning for visual classification, we formulate the IncPAR as a problem of multi-label continual learning with incomplete labels (MCL-IL), where the training samples in a novel task are annotated with only a few categories of interest but may implicitly contain other attributes of previous tasks. The incomplete label assignments is a challenging and frequently-encountered issue in real-world multi-label classification applications due to a number of reasons, e.g., incomplete data collection, moderate budget for annotations, etc. To tackle the MCL-IL problem, we propose a self-training based approach via dual uncertainty-aware pseudo-labeling (DUAPL) to transfer the knowledge learned in previous tasks to novel tasks. Specially, both kinds of uncertainties, i.e., aleatoric uncertainty and epistemic uncertainty, are modeled to mitigate the negative influences of noisy pseudo labels induced by low quality samples and immature models learned by inadequate training in early tasks. Based on the DUAPL, more reliable supervision signals can be estimated to prevent the model evolution from forgetting attributes seen in previous tasks. For standard evaluations of MCL-IL methods, two benchmarks on IncPAR, termed RAP-CL and PETA-CL, are constructed by re-organizing public human attribute datasets. Extensive experiments have been performed on these benchmarks to compare the proposed method with multiple baselines. The superior performance in terms of both recognition accuracies and forgetting ratios demonstrate the effectiveness of the proposed DUAPL for IncPAR.
Da Li 0003, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2023 Learning Domain Invariant Representations for Generalizable Person Re-Identification
abstract
Generalizable person Re-Identification (ReID) aims to learn ready-to-use cross-domain representations for direct cross-data evaluation, which has attracted growing attention in the recent computer vision (CV) community. In this work, we construct a structural causal model (SCM) among identity labels, identity-specific factors (clothing/shoes color etc.), and domain-specific factors (background, viewpoints etc.). According to the causal analysis, we propose a novel Domain Invariant Representation Learning for generalizable person Re-Identification (DIR-ReID) framework. Specifically, we propose to disentangle the identity-specific and domain-specific factors into two independent feature spaces, based on which an effective backdoor adjustment approximate implementation is proposed for serving as a causal intervention towards the SCM. Extensive experiments have been conducted, showing that DIR-ReID outperforms state-of-the-art (SOTA) methods on large-scale domain generalization (DG) ReID benchmarks.
Yifan Zhang 0004, Zhang Zhang 0001, Da Li 0003, Liang Wang 0001, Tieniu Tan
IEEE Trans. Image Process.2
2022 Fusion Tree Network for RGBT Tracking
abstract
RGBT tracking is often affected by complex scenes (i.e., occlusions, scale changes, noisy background, etc). Existing works usually adopt a single-strategy RGBT tracking fusion scheme to handle modality fusion in all scenarios. However, due to the limitation of fusion model capacity, it is difficult to fully integrate the discriminative features between different modalities. To tackle this problem, we propose a Fusion Tree Network (FTNet), which provides a multi-strategy fusion model with high capacity to efficiently fuse different modalities. Specifically, we combine three kinds of attention modules (i.e., channel attention, spatial attention, and location attention) in a tree structure to achieve multi-path hybrid attention in the deeper convolutional stages of the object tracking network. Extensive experiments are performed on three RGBT tracking datasets, and the results show that our method achieves superior performance among state-of-the-art RGBT tracking models.
Zhiyuan Cheng 0012, Andong Lu, Zhang Zhang 0001, Chenglong Li 0002, Liang Wang 0001
AVSS3
2022 Graph Structure Learning Boosted Neural Network for Image Segmentation
abstract
Although Convolutional Neural Networks have made significant progress in image segmentation, it remains inadequate for exploring the structural relationships between image components and how graphs can be employed to guide image segmentation. To explore the structural relationships inherent in image components, the Graph Structure Learning Boosted Neural Network was proposed, which takes the contextual information generated by the CNN as features of the nodes and then uses a self-supervised graph generator to generate an adjacency matrix representing the image components connectivity. Then a Graph Neural Network (GNN) uses the adjacency matrix to fuse information between components according to their connectivity, thus transforming the CNN’s pixel classification problem into the GNN’s pixel classification problem. The whole model is lightweight and scalable, and extensive experiments have demonstrated the scalability of the model alongside the effectiveness of the method.
Jinde Liu, Zhang Zhang 0001
AVSS2
2022 Cross-Domain Cross-Set Few-Shot Learning via Learning Compact and Aligned Representations
Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan
ECCV (20)2
2022 Focal and efficient IOU loss for accurate bounding box regression
Yifan Zhang 0004, Weiqiang Ren, Zhang Zhang 0001, Liang Wang 0001, Tieniu Tan
Neurocomputing3
2022 Dual-branch self-attention network for pedestrian attribute recognition
Zhang Zhang 0001, Da Li 0003, Peng Zhang 0057, Caifeng Shan
Pattern Recognit. Lett.2
2021 Richly Activated Graph Convolutional Network for Robust Skeleton-Based Action Recognition
abstract
Current methods for skeleton-based human action recognition usually work with complete skeletons. However, in real scenarios, it is inevitable to capture incomplete or noisy skeletons, which could significantly deteriorate the performance of current methods when some informative joints are occluded or disturbed. To improve the robustness of action recognition models, a multi-stream graph convolutional network (GCN) is proposed to explore sufficient discriminative features spreading over all skeleton joints, so that the distributed redundant representation reduces the sensitivity of the action models to non-standard skeletons. Concretely, the backbone GCN is extended by a series of ordered streams which is responsible for learning discriminative features from the joints less activated by preceding streams. Here, the activation degrees of skeleton joints of each GCN stream are measured by the class activation maps (CAM), and only the information from the unactivated joints will be passed to the next stream, by which rich features over all active joints are obtained. Thus, the proposed method is termed richly activated GCN (RA-GCN). Compared to the state-of-the-art (SOTA) methods, the RA-GCN achieves comparable performance on the standard NTU RGB+D 60 and 120 datasets. More crucially, on the synthetic occlusion and jittering datasets, the performance deterioration due to the occluded and disturbed joints can be significantly alleviated by utilizing the proposed RA-GCN.
Yi-Fan Song, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Meta-USR: A Unified Super-Resolution Network for Multiple Degradation Parameters
abstract
Recent research on single image super-resolution (SISR) has achieved great success due to the development of deep convolutional neural networks. However, most existing SISR methods merely focus on super-resolution of a single fixed integer scale factor. This simplified assumption does not meet the complex conditions for real-world images which often suffer from various blur kernels or various levels of noise. More importantly, previous methods lack the ability to cope with arbitrary degradation parameters (scale factors, blur kernels, and noise levels) with a single model. A few methods can handle multiple degradation factors, e.g., noninteger scale factors, blurring, and noise, simultaneously within a single SISR model. In this work, we propose a simple yet powerful method termed meta-USR which is the first unified super-resolution network for arbitrary degradation parameters with meta-learning. In Meta-USR, a meta-restoration module (MRM) is proposed to enhance the traditional upscale module with the capability to adaptively predict the weights of the convolution filters for various combinations of degradation parameters. Thus, the MRM can not only upscale the feature maps with arbitrary scale factors but also restore the SR image with different blur kernels and noise levels. Moreover, the lightweight MRM can be placed at the end of the network, which makes it very efficient for iteratively/repeatedly searching the various degradation factors. We evaluate the proposed method through extensive experiments on several widely used benchmark data sets on SISR. The qualitative and quantitative experimental results show the superiority of our Meta-USR.
Xuecai Hu, Zhang Zhang 0001, Caifeng Shan, Zilei Wang, Liang Wang 0001, Tieniu Tan
IEEE Trans. Neural Networks Learn. Syst.2
2020 Stronger, Faster and More Explainable: A Graph Convolutional Baseline for Skeleton-based Action Recognition
abstract
One essential problem in skeleton-based action recognition is how to extract discriminative features over all skeleton joints. However, the complexity of the State-Of-The-Art (SOTA) models of this task tends to be exceedingly sophisticated and over-parameterized, where the low efficiency in model training and inference has obstructed the development in the field, especially for large-scale action datasets. In this work, we propose an efficient but strong baseline based on Graph Convolutional Network (GCN), where three main improvements are aggregated, i.e., early fused Multiple Input Branches (MIB), Residual GCN (ResGCN) with bottleneck structure and Part-wise Attention (PartAtt) block. Firstly, an MIB is designed to enrich informative skeleton features and remain compact representations at an early fusion stage. Then, inspired by the success of the ResNet architecture in Convolutional Neural Network (CNN), a ResGCN module is introduced in GCN to alleviate computational costs and reduce learning difficulties in model training while maintain the model accuracy. Finally, a PartAtt block is proposed to discover the most essential body parts over a whole action sequence and obtain more explainable representations for different skeleton action sequences. Extensive experiments on two large-scale datasets, i.e., NTU RGB+D 60 and 120, validate that the proposed baseline slightly outperforms other SOTA models and meanwhile requires much fewer parameters during training and inference procedures, e.g., at most 34 times less than DGNN, which is one of the best SOTA methods.
Yi-Fan Song, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001
ACM Multimedia2
2020 Kinematic skeleton graph augmented network for human parsing
Jinde Liu, Zhang Zhang 0001, Caifeng Shan, Tieniu Tan
Neurocomputing2
2020 Multi angle optimal pattern-based deep learning for automatic facial expression recognition
Deepak Kumar Jain 0001, Zhang Zhang 0001, Kaiqi Huang
Pattern Recognit. Lett.2
2020 Deep Unbiased Embedding Transfer for Zero-Shot Learning
abstract
Zero-shot learning aims to recognize objects which do not appear in the training dataset. Previous prevalent mapping-based zero-shot learning methods suffer from the projection domain shift problem due to the lack of image classes in the training stage. In order to alleviate the projection domain shift problem, a deep unbiased embedding transfer (DUET) model is proposed in this paper. The DUET model is composed of a deep embedding transfer (DET) module and an unseen visual feature generation (UVG) module. In the DET module, a novel combined embedding transfer net which integrates the complementary merits of the linear and nonlinear embedding mapping functions is proposed to connect the visual space and semantic space. What's more, the end-to-end joint training process is implemented to train the visual feature extractor and the combined embedding transfer net simultaneously. In the UVG module, a visual feature generator trained with a conditional generative adversarial framework is used to synthesize the visual features of the unseen classes to ease the disturbance of the projection domain shift problem. Furthermore, a quantitative index, namely the score of resistance on domain shift (ScoreRDS), is proposed to evaluate different models regarding their resistance capability on the projection domain shift problem. The experiments on five zero-shot learning benchmarks verify the effectiveness of the proposed DUET model. As demonstrated by the qualitative and quantitative analysis, the unseen class visual feature generation, the combined embedding transfer net and the end-to-end joint training process all contribute to alleviating projection domain shift in zero-shot learning.
Zhang Zhang 0001, Liang Wang 0001, Caifeng Shan, Tieniu Tan
IEEE Trans. Image Process.2
2020 Deep Fusion Feature Representation Learning With Hard Mining Center-Triplet Loss for Person Re-Identification
abstract
Person re-identification (Re-ID) is a challenging task in the field of computer vision and focuses on matching people across images from different cameras. The extraction of robust feature representations from pedestrian images through CNNs with a single deterministic pooling operation is problematic as the features in real pedestrian images are complex and diverse. To address this problem, we propose a novel center-triplet (CT) model that combines the learning of robust feature representation and the optimization of metric loss function. Firstly, we design a fusion feature learning network (FFLN) with a novel fusion strategy consisting of max pooling and average pooling. Instead of adopting a single deterministic pooling operation, the FFLN combines two pooling operations that can learn high response values, bright features, and low response values, discriminative features simultaneously. Our model obtains more discriminative fusion features by adaptively learning the weights of the features learned by the corresponding pooling operations. In addition, we design a hard mining center-triplet loss (HCTL), a novel improved triplet loss, which effectively optimizes the intra/inter-class distance and reduces the cost of computing and mining hard training samples simultaneously, thereby enhancing the learning of robust feature representation. Finally, we proved our method can learn robust and discriminative feature representations for complex pedestrian images in real scenes. The experimental results also illustrate that our method achieves an 81.8% mAP and a 93.8% rank-1 accuracy on Market1501, a 68.2% mAP and an 83.3% rank-1 accuracy on DukeMTMC-ReID, and a 43.6% mAP and a 74.3% rank-1 accuracy on MSMT17, outperforming most state-of-the-art methods and achieving better performance for person re-identification.
Cairong Zhao, Xinbi Lv, Zhang Zhang 0001, Wangmeng Zuo, Jun Wu 0006, Duoqian Miao 0001
IEEE Trans. Multim.3
2019 A Comprehensive Study on Large-Scale Person Retrieval in Real Surveillance Scenarios
abstract
Person retrieval is a hot research topic due to its important application potential for public security. Though existing algorithms have achieved impressive progresses on current public datasets, it is still a challenging task in the real surveillance scenarios due to the various viewpoints, pose variations and occlusions. Moreover, few of the existing works study the problem of person retrieval on large-scale gallery set, where lots of distractions may deteriorate the retrieval results heavily. To have a deep understanding on the above challenges, we perform a comprehensive study on current state-of-the-art person retrieval algorithms with a large-scale benchmark in real surveillance scenarios. In the study, two kinds of techniques, i.e., attribute recognition and person re-identification, including eight algorithms, are evaluated at both algorithm level and system level. Here, the system-level evaluations investigate the effects of the combinations of the above algorithms with the module of person detection, where lots of distractions in person detection results pose a big challenge for person retrieval in real scenes. Extensive evaluations with large gallery sizes (up to 243k) and comprehensive analyses are presented in the study, which will guide researchers to develop more advanced algorithms in future.
Da Li 0003, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001, Tieniu Tan
AVSS2
2019 Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-Identification
abstract
The fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically, a Class Activation Maps (CAM) augmentation model is proposed to expand the activation scope of baseline Re-ID model to explore rich visual cues, where the backbone network is extended by a series of ordered branches which share the same input but output complementary CAM. A novel Overlapped Activation Penalty is proposed to force the new branch to pay more attention to the image regions less activated by the old ones, such that spatial diverse visual features can be discovered. The proposed model achieves state-of-the-art results on three person Re-ID benchmarks. Moreover, a visualization approach termed ranking activation map (RAM) is proposed to explicitly interpret the ranking results in the test stage, which gives qualitative validations of the proposed method.
Wenjie Yang 0005, Houjing Huang, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang, Shu Zhang 0001
CVPR3
2019 Unsupervised Cross-Domain Person Re-Identification: A New Framework
abstract
Although existing person Re-IDentification (ReID) methods have achieved great progress with large-scale labeled data, it is still hard to generalize to unseen scenarios without any la-beled person identities. To alleviate this problem, this paper proposes a new framework to take full advantage of the label information of source domain and the data distribution geometry of unlabeled target domain to improve the ReID performance in the unlabeled target domain. Instead of direct model transfer, the data transfer is first adopted where the identity preserving samples are generated from the labeled source domain to unlabeled target domain. Accordingly, a better initialized target domain adapted ReID model could be obtained with the generated samples. The fine-grained part-level features are then learned instead of global features to better mine new persons in the unlabeled target domain. Finally, the proposed framework iteratively updates the ReID model with the generated persons and the mined persons in last iteration, and explores new persons from the unlabeled target domain. The state-of-the-art experimental results are achieved on Market1501 and DukeMTMC-reID in terms of unsupervised cross-domain person ReID.
Da Li 0003, Dangwei Li, Zhang Zhang 0001, Liang Wang 0001, Tieniu Tan
ICIP3
2019 Richly Activated Graph Convolutional Network for Action Recognition with Incomplete Skeletons
abstract
Current methods for skeleton-based human action recognition usually work with completely observed skeletons. However, in real scenarios, it is prone to capture incomplete and noisy skeletons, which will deteriorate the performance of traditional models. To enhance the robustness of action recognition models to incomplete skeletons, we propose a multi-stream graph convolutional network (GCN) for exploring sufficient discriminative features distributed over all skeleton joints. Here, each stream of the network is only responsible for learning features from currently unactivated joints, which are distinguished by the class activation maps (CAM) obtained by preceding streams, so that the activated joints of the proposed method are obviously more than traditional methods. Thus, the proposed method is termed richly activated GCN (RA-GCN), where the richly discovered features will improve the robustness of the model. Compared to the state-of-the-art methods, the RA-GCN achieves comparable performance on the NTU RGB+D dataset. Moreover, on a synthetic occlusion dataset, the performance deterioration can be alleviated by the RA-GCN significantly.
Yi-Fan Song, Zhang Zhang 0001, Liang Wang 0001
ICIP2
2019 MAPNet: Multi-modal attentive pooling network for RGB-D indoor scene classification
Yabei Li, Zhang Zhang 0001, Yanhua Cheng, Liang Wang 0001, Tieniu Tan
Pattern Recognit.2
2019 Corrigendum to "MAPNet: Multi-modal attentive pooling network for RGB-D indoor scene classification" [Pattern Recognition 90 (2019) 436-449]
Yabei Li, Zhang Zhang 0001, Yanhua Cheng, Liang Wang 0001, Tieniu Tan
Pattern Recognit.2
2019 A Richly Annotated Pedestrian Dataset for Person Retrieval in Real Surveillance Scenarios
abstract
Retrieving specific persons with various types of queries, e.g., a set of attributes or a portrait photo has great application potential in large-scale intelligent surveillance systems. In this paper, we propose a richly annotated pedestrian (RAP) dataset which serves as a unified benchmark for both attribute-based and image-based person retrieval in real surveillance scenarios. Typically, previous datasets have three improvable aspects, including limited data scale and annotation types, heterogeneous data source, and controlled scenarios. Differently, RAP is a large-scale dataset which contains 84928 images with 72 types of attributes and additional tags of viewpoint, occlusion, body parts, and 2589 person identities. It is collected in the real uncontrolled scene and has complex visual variations in pedestrian samples due to the change of viewpoints, pedestrian postures, and cloth appearance. Towards a high-quality person retrieval benchmark, an amount of state-of-the-art algorithms on pedestrian attribute recognition and person re-identification (ReID), are performed for quantitative analysis with three evaluation tasks, i.e., attribute recognition, attribute-based and image-based person retrieval, where a new instance-based metric is proposed to measure the dependency of the prediction of multiple attributes. Finally, some interesting problems, e.g., the joint feature learning of attribute recognition and ReID, and the problem of cross-day person ReID, are explored to show the challenges and future directions in person retrieval.
Dangwei Li, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang
IEEE Trans. Image Process.2
2019 ISEE: An Intelligent Scene Exploration and Evaluation Platform for Large-Scale Visual Surveillance
abstract
Intelligent video surveillance (IVS) is always an interesting research topic to utilize visual analysis algorithms for exploring richly structured information from big surveillance data. However, existing IVS systems either struggle to utilize computing resources adequately to improve the efficiency of large-scale video analysis, or present a customized system for specific video analytic functions. It still lacks of a comprehensive computing architecture to enhance efficiency, extensibility and flexibility of IVS system. Moreover, it is also an open problem to study the effect of the combinations of multiple vision modules on the final performance of end applications of IVS system. Motivated by these challenges, we develop an Intelligent Scene Exploration and Evaluation (ISEE) platform based on a heterogeneous CPU-GPU cluster and some distributed computing tools, where Spark Streaming serves as the computing engine for efficient large-scale video processing and Kafka is adopted as a middle-ware message center to decouple different analysis modules flexibly. To validate the efficiency of the ISEE and study the evaluation problem on composable systems, we instantiate the ISEE for an end application on person retrieval with three visual analysis modules, including pedestrian detection with tracking, attribute recognition and re-identification. Extensive experiments are performed on a large-scale surveillance video dataset involving 25 camera scenes, totally 587 hours 720p synchronous videos, where a two-stage question-answering procedure is proposed to measure the performance of execution pipelines composed of multiple visual analysis algorithms based on millions of attribute-based and relationship-based queries. The case study of system-level evaluations may inspire researchers to improve visual analysis algorithms and combining strategies from the view of a scalable and composable system in the future.
Da Li 0003, Zhang Zhang 0001, Kai Yu 0003, Kaiqi Huang, Tieniu Tan
IEEE Trans. Parallel Distributed Syst.2
2018 Adversarially Occluded Samples for Person Re-Identification
abstract
Person re-identification (ReID) is the task of retrieving particular persons across different cameras. Despite its great progress in recent years, it is still confronted with challenges like pose variation, occlusion, and similar appearance among different persons. The large gap between training and testing performance with existing models implies the insufficiency of generalization. Considering this fact, we propose to augment the variation of training data by introducing Adversarially Occluded Samples. These special samples are both a) meaningful in that they resemble real-scene occlusions, and b) effective in that they are tough for the original model and thus provide the momentum to jump out of local optimum. We mine these samples based on a trained ReID model and with the help of network visualization techniques. Extensive experiments show that the proposed samples help the model discover new discriminative clues on the body and generalize much better at test time. Our strategy makes significant improvement over strong baselines on three large-scale ReID datasets, Market1501, CUHK03 and DukeMTMC-reID.
Houjing Huang, Dangwei Li, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang
CVPR3
2018 Pose Guided Deep Model for Pedestrian Attribute Recognition in Surveillance Scenarios
abstract
Recognizing pedestrian attributes, such as gender, backpack, and cloth types, has obtained increasing attention recently due to its great potential in intelligent video surveillance. Existing methods usually solve it with end-to-end multi-label deep neural networks, while the structure knowledge of pedestrian body has been little utilized. Considering that attributes have strong spatial correlations with human structures, e.g. glasses are around the head, in this paper, we introduce pedestrian body structure into this task and propose a Pose Guided Deep Model (PGDM) to improve attribute recognition. The PGDM consists of three main components: 1) coarse pose estimation which distillates the pose knowledge from a pre-trained pose estimation model, 2) body parts localization which adaptively locates informative image regions with only image-level supervision, 3) multiple features fusion which combines the part-based features for attribute recognition. In the inference stage, we fuse the part-based PGDM results with global body based results for final attribute prediction and the performance can be consistently improved. Compared with state-of-the-art models, the performances on three large-scale pedestrian attribute datasets, i.e., PETA, RAP, and PA-100K, demonstrate the effectiveness of the proposed method.
Dangwei Li, Xiaotang Chen, Zhang Zhang 0001, Kaiqi Huang
ICME3
2018 Gestalt laws based tracklets analysis for human crowd understanding
Zhang Zhang 0001, Kaiqi Huang
Pattern Recognit.2
2018 Random walk-based feature learning for micro-expression recognition
Deepak Kumar Jain 0001, Zhang Zhang 0001, Kaiqi Huang
Pattern Recognit. Lett.2
2017 Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
Kai Yu 0003, Biao Leng, Zhang Zhang 0001, Dangwei Li, Kaiqi Huang
BMVC4
2017 Learning Deep Context-Aware Features over Body and Latent Parts for Person Re-identification
abstract
Person Re-identification (ReID) is to identify the same person across different cameras. It is a challenging task due to the large variations in person pose, occlusion, background clutter, etc. How to extract powerful features is a fundamental problem in ReID and is still an open problem today. In this paper, we design a Multi-Scale Context-Aware Network (MSCAN) to learn powerful features over full body and body parts, which can well capture the local context knowledge by stacking multi-scale convolutions in each layer. Moreover, instead of using predefined rigid parts, we propose to learn and localize deformable pedestrian parts using Spatial Transformer Networks (STN) with novel spatial constraints. The learned body parts can release some difficulties, e.g. pose variations and background clutters, in part-based representation. Finally, we integrate the representation learning processes of full body and body parts into a unified framework for person ReID through multi-class person identification tasks. Extensive evaluations on current challenging large-scale person ReID datasets, including the image-based Market1501, CUHK03 and sequence-based MARS datasets, show that the proposed method achieves the state-of-the-art results.
Dangwei Li, Xiaotang Chen, Zhang Zhang 0001, Kaiqi Huang
CVPR3
2016 ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory Clustering
abstract
For spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simultaneously. Given an initial graph with only a few nearest but most reliable pairwise relations, new reliable relations are discovered by an assumption of reliability preservation, i.e., the reliable relations will preserve their reliabilities in the learnt projection subspace. We formulate the idea as a cross entropy (CE) minimization problem to reduce the discrepancy between two Bernoulli distributions parameterized by the updated distances and the existing relation graph respectively. Furthermore, to overcome the imbalanced distribution of samples, a Boosting-like strategy is proposed to balance the discovered relations over all clusters. To evaluate the proposed method, extensive experiments are performed with various trajectory clustering tasks, including motion segmentation, time series clustering and crowd detection. The results demonstrate that ReDSFA can discover reliable intra-cluster relations with high precision, and competitive clustering performance can be achieved in comparison with state-of-the-art.
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Peipei Yang, Jun Li 0010
CVPR1
2016 Joint crowd detection and semantic scene modeling using a Gestalt laws-based similarity
abstract
This paper presents a novel approach to detecting crowd groups and learning semantic regions with a Gestalt laws-based similarity. Different from the existing approaches based on optical flows or complete trajectories, our model adopts tracklets as the original input, because they carry more detailed information. Though those tracklets do not appear in the same duration, they are more robust to noise in crowd scene. According to the Gestalt laws of grouping, we propose three priors to define a unified similarity measure to calculate the affinities of pairs of original tracklets and pairs of representative tracklets in crowd groups. Therefore, the short-term crowd groups and the long-term semantic paths in crowded scene can be detected by a bottom-up hierarchical clustering algorithm simultaneously. Extensive experiments on hundreds of video clips demonstrate that our approach is effective and reliable for crowd detection and semantic scene understanding.
Zhang Zhang 0001, Kaiqi Huang
ICIP2
2015 Adaptive Slice Representation for Human Action Classification
abstract
Common action recognition methods describe an action sequence along with its time axis, i.e., first extracting features from the x y plane, and then modeling the dynamic changes along with the time axis. Other than the ordinary x y plane-based representation, other views, e.g., xt slice-based representation, may be more efficient to distinguish different actions. In this paper, we investigate different slicing views of the spatiotemporal volume to organize action sequences and propose an efficient slice representation for human action recognition. First, a minimum average entropy principle is proposed to select the optimal slicing angle for each action sequence adaptively. This allows the foreground pixels to be distributed in the fewest slices so as to reduce more uncertainty caused by the information dispersed in different slices. Then, the obtained slice sequence is transformed into a pair of 1-D signals to describe the distribution of foreground pixels along the time axis. Finally, the mel frequency cepstrum coefficient features are calculated to describe the spectrum characteristics of the 1-D signals over time. Thus, a 3-D spatiotemporal action volume is efficiently transformed into a low-dimensional spectrum features. Extensive experiments on the 2-D human action data sets (the UIUC and the WEIZMANN) as well as the Microsoft Research (MSR) Action3-D depth data set demonstrate the effectiveness of the slice-based representation, where the recognition performance can reach to the state-of-the-art level with high efficiency.
Yanhu Shan, Zhang Zhang 0001, Peipei Yang, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.2
2012 Interest Point Selection with Spatio-temporal Context for Realistic Action Recognition
abstract
Spatio-Temporal Interest Point (STIP) has been widely used for human action recognition. However, the performance of the STIP based methods are still limited in realistic datasets which often include large variations in illuminations, viewpoints and camera motions. One reason of the low performance is that the STIPs only reflect the local change in videos, which is not enough to obtain stable informative features for action representation in realistic scene. To tackle the problem, we proposed an approach to selecting the "stable STIPs" with the spatio-temporal distribution of STIPs in neighbor region. Then, BoW feature is constructed to represent actions with these selected points. The experimental results on KTH dataset and HMDB (the largest realistic human action dataset) demonstrate that the proposed approach has obvious effect on improving the recognition rates of realistic data.
Yanhu Shan, Zhang Zhang 0001, Junge Zhang, Kaiqi Huang, Oh Se Hyun
AVSS2
2012 Baseline Results for Violence Detection in Still Images
abstract
Recognizing objectionable content draws more and more attention nowadays given the rapid proliferation of images and videos on the Internet. Although there are some investigations about violence video detection and pornographic information filtering, very few existing methods touch on the problem of violence detection in still images. However, given its potential use in violence webpage filtering, online public opinion monitoring and some other aspects, recognizing violence in still images is worth being deeply investigated. To this end, we first establish a new database containing 500 violence images and 1500 non-violence images. And we use the Bag-of-Words (BoW) model which is frequently adopted in image classification domain to discriminate violence images and non-violence images. The effectiveness of four different feature representations are tested within the BoW framework. Finally the baseline results for violence image detection on our newly built database are reported.
Dong Wang 0004, Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
AVSS2
2012 Continuum regression for cross-modal multimedia retrieval
abstract
Understanding the relationship among different modalities is a challenging task. The frequently used canonical correlation analysis (CCA) and its variants have proved effective for building a common space in which the correlation between different modalities is maximized. In this paper, we show that CCA and its variants may cause information dissipation when switching the modals, and thus propose to use the continuum regression (CR) model to handle this problem. In particular, the CR model with a fixed variance coefficient of 1/2 is adopted here. We also apply the multinomial logistic regression model for further classification task. To evaluate the CR model, we perform a series of cross-modal retrieval experiments in terms of two kinds of modals, namely image and text. Compared with previous methods, experimental results show that the CR model has achieved the best retrieval precision, which demonstrates the potential of our method for real internet search applications.
Yongming Chen, Liang Wang 0001, Wei Wang 0025, Zhang Zhang 0001
ICIP4
2012 Segment-Based Features for Time Series Classification
abstract
In this paper, we propose an approach termed segment-based features (SBFs) to classify time series. The approach is inspired by the success of the component- or part-based methods of object recognition in computer vision, in which a visual object is described as a number of characteristic parts and the relations among the parts. Utilizing this idea in the problem of time series classification, a time series is represented as a set of segments and the corresponding temporal relations. First, a number of interest segments are extracted by interest point detection with automatic scale selection. Then, a number of feature prototypes are collected by random sampling from the segment set, where each feature prototype may include single segment or multiple ordered segments. Subsequently, each time series is transformed to a standard feature vector, i.e. SBF, where each entry in the SBF is calculated as the maximum response (maximum similarity) of the corresponding feature prototype to the segment set of the time series. Based on the original SBF, an incremental feature selection algorithm is conducted to form a compact and discriminative feature representation. Finally, a multi-class support vector machine is trained to classify the test time series. Extensive experiments on different time series datasets, including one synthetic control dataset, two sign language datasets and one gait dynamics dataset, have been performed to evaluate the proposed SBF method. Compared with other state-of-the-art methods, our approach achieves superior classification performance, which clearly validates the advantages of the proposed method.
Zhang Zhang 0001, Jun Cheng 0002, Jun Li 0010, Wei Bian 0003, Dacheng Tao
Comput. J.1
2012 Slow Feature Analysis for Human Action Recognition
abstract
Slow Feature Analysis (SFA) extracts slowly varying features from a quickly varying input signal. It has been successfully applied to modeling the visual receptive fields of the cortical neurons. Sufficient experimental results in neuroscience suggest that the temporal slowness principle is a general learning principle in visual perception. In this paper, we introduce the SFA framework to the problem of human action recognition by incorporating the discriminative information with SFA learning and considering the spatial relationship of body parts. In particular, we consider four kinds of SFA learning strategies, including the original unsupervised SFA (U-SFA), the supervised SFA (S-SFA), the discriminative SFA (D-SFA), and the spatial discriminative SFA (SD-SFA), to extract slow feature functions from a large amount of training cuboids which are obtained by random sampling in motion boundaries. Afterward, to represent action sequences, the squared first order temporal derivatives are accumulated over all transformed cuboids into one feature vector, which is termed the Accumulated Squared Derivative (ASD) feature. The ASD feature encodes the statistical distribution of slow features in an action sequence. Finally, a linear support vector machine (SVM) is trained to classify actions represented by ASD features. We conduct extensive experiments, including two sets of control experiments, two sets of large scale experiments on the KTH and Weizmann databases, and two sets of experiments on the CASIA and UT-interaction databases, to demonstrate the effectiveness of SFA for human action recognition. Experimental results suggest that the SFA-based approach (1) is able to extract useful motion patterns and improves the recognition performance, (2) requires less intermediate processing steps but achieves comparable or even better performance, and (3) has good potential to recognize complex multiperson activities.
Zhang Zhang 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 An Extended Grammar System for Learning and Recognizing Complex Visual Events
abstract
For a grammar-based approach to the recognition of visual events, there are two major limitations that prevent it from real application. One is that the event rules are predefined by domain experts, which means huge manual cost. The other is that the commonly used grammar can only handle sequential relations between subevents, which is inadequate to recognize more complex events involving parallel subevents. To solve these problems, we propose an extended grammar approach to modeling and recognizing complex visual events. First, motion trajectories as original features are transformed into a set of basic motion patterns of a single moving object, namely, primitives (terminals) in the grammar system. Then, a Minimum Description Length (MDL) based rule induction algorithm is performed to discover the hidden temporal structures in primitive stream, where Stochastic Context-Free Grammar (SCFG) is extended by Allen's temporal logic to model the complex temporal relations between subevents. Finally, a Multithread Parsing (MTP) algorithm is adopted to recognize interesting complex events in a given primitive stream, where a Viterbi-like error recovery strategy is also proposed to handle large-scale errors, e.g., insertion and deletion errors. Extensive experiments, including gymnastic exercises, traffic light events, and multi-agent interactions, have been executed to validate the effectiveness of the proposed approach.
Zhang Zhang 0001, Tieniu Tan, Kaiqi Huang
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Multi-thread Parsing for Recognizing Complex Events in Videos
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan
ECCV (3)1
2007 Trajectory Series Analysis based Event Rule Induction for Visual Surveillance
abstract
In this paper, a generic rule induction framework based on trajectory series analysis is proposed to learn the event rules. First the trajectories acquired by a tracking system are mapped into a set of primitive events that represent some basic motion patterns of moving object. Then a minimum description length (MDL) principle based grammar induction algorithm is adopted to infer the meaningful rules from the primitive event series. Compared with previous grammar rule based work on event recognition where the rules are all defined manually, our work aims to learn the event rules automatically. Experiments in a traffic crossroad have demonstrated the effectiveness of our methods. Shown in the experimental results, most of the grammar rules obtained by our algorithm are consistent with the actual traffic events in the crossroad. Furthermore the traffic lights rule in the crossroad can also be leaned correctly with the help of eliminating the irrelevant trajectories.
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Liangsheng Wang
CVPR1
2006 Complex Activity Representation and Recognition by Extended Stochastic Grammar
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan
ACCV (1)1