Hao Zhou 0014

dblp:63/778-14 · DBLP profile ↗
← Back
21ranked-venue papers
14as first author
16since 2021 · last 2026
0000-0002-8167-7018ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical kernel decoupling for graph convolution: Enhancing skeleton-based action recognition through structured representation
Ying Li 0016, Hao Zhou 0014, Chuanping Hu, Mingzhou Lu, Yan Luo 0003
Pattern Recognit.3
2026 Collaborative Observation Imputation and Trajectory Prediction via Consistency Evaluation
abstract
Pedestrian trajectory prediction is a critical task in various mobile computing applications, such as video surveillance, robot navigation, autonomous driving, and human mobility analysis. Although significant progress has been made by current methods, the challenge of observation deficiency in pedestrian trajectory prediction remains largely unaddressed. Since most existing methods focus on optimizing prediction accuracy under the assumption of complete observations, while ignoring the potential for observation deficiency caused by failures in detection or tracking algorithms. To overcome this challenge, we propose a collaborative observation imputation and trajectory prediction framework, which employs consistency evaluation to jointly perform the imputation and prediction tasks. Specifically, we first build a consistency evaluation module to align features between observed and future trajectory pairs using contrastive learning. Then, we design a trajectory imputation and prediction baseline, which adopts a parallel paradigm, to mitigate the impact of coarse imputations on trajectory prediction when performing initial imputation and prediction. Next, we introduce a consistency-guided Skip-Diffusion module, which leverages consistency evaluation between initial imputations and ground truth future trajectories to refine the initial imputations. Finally, we propose a consistency-driven Cross-Mamba module, which uses consistency evaluation between ground-truth observations and initial predictions to refine the initial predictions. Extensive experiments demonstrate the effectiveness of the proposed framework in both imputation and prediction tasks.
Hao Zhou 0014, Mingyu Fan, Xu Yang 0004, Hai Huang 0004, Kaishun Wu, Lu Wang 0002, Fei Luo 0003
IEEE Trans. Mob. Comput.1
2025 Three-Dimensional Trajectory Prediction with 3DMoTraj Dataset
abstract
With the growing interest in embodied and spatial intelligence, accurately predicting trajectories in 3D environments has become increasingly critical. However, no datasets have been explicitly designed to study 3D trajectory prediction. To this end, we contribute a 3D motion trajectory (3DMoTraj) dataset collected from unmanned underwater vehicles (UUVs) operating in oceanic environments. Mathematically, trajectory prediction becomes significantly more complex when transitioning from 2D to 3D. To tackle this challenge, we analyze the prediction complexity of 3D trajectories and propose a new method consisting of two key components: decoupled trajectory prediction and correlated trajectory refinement. The former decouples inter-axis correlations, thereby reducing prediction complexity and generating coarse predictions. The latter refines the coarse predictions by modeling their inter-axis correlations. Extensive experiments show that our method significantly improves 3D trajectory prediction accuracy and outperforms state-of-the-art methods. Both the 3DMoTraj dataset and the method are available at https://github.com/zhouhao94/3DMoTraj.
Hao Zhou 0014, Xu Yang 0004, Mingyu Fan, Lu Qi 0001, Xiangtai Li, Ming-Hsuan Yang 0001, Fei Luo 0003
ICML1
2025 Rethinking Evaluation Metrics of Open-Vocabulary Segmentation
abstract
This paper highlights a problem of evaluation metrics adopted in the open-vocabulary segmentation. The evaluation process relies heavily on closed-set metrics on zero-shot or cross-dataset pipelines without considering the similarity between predicted and ground truth categories. We first survey eleven similarity measurements between two categorical words using WordNet linguistics statistics, text embedding, or language models by comprehensive quantitative analysis and user study to tackle this issue. Based on those explored measurements, we design novel evaluation metrics, Open mIoU, Open AP, and Open PQ, tailored for three open-vocabulary segmentation tasks. We benchmark the proposed evaluation metrics on twelve open-vocabulary methods in three segmentation tasks. Despite the relative subjectivity of similarity distance, we demonstrate that our metrics can still well evaluate the open ability of the existing open-vocabulary segmentation methods. We hope our work can bring the community new thinking about evaluating model ability for open-vocabulary segmentation.
Hao Zhou 0014, Lu Qi 0001, Tiancheng Shen, Hai Huang 0004, Xu Yang 0004, Xiangtai Li, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
Hao Zhou 0014, Pengfei Xing, Long Zhao 0003, Junwei Liang 0001, Alex Hauptmann 0001, Ting Liu 0005, Andrew C. Gallagher
ECCV (29)2
2024 VideoPrism: A Foundational Visual Encoder for Video Understanding
abstract
We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding by global-local distillation of semantic video embeddings and a token shuffling scheme, enabling VideoPrism to focus primarily on the video modality while leveraging the invaluable text associated with videos. We extensively test VideoPrism on four broad groups of video understanding tasks, from web video question answering to CV for science, achieving state-of-the-art performance on 31 out of 33 video understanding benchmarks.
Long Zhao 0003, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou 0014, Shen Yan 0008, Jennifer J. Sun, Luke Friedman, Rui Qian 0003, Tobias Weyand, Yue Zhao 0006, Rachel Hornung, Florian Schroff, Ming-Hsuan Yang 0001, David A. Ross, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko, Ting Liu 0005, Boqing Gong
ICML4
2023 Static-dynamic global graph representation for pedestrian trajectory prediction
Hao Zhou 0014, Xu Yang 0004, Mingyu Fan, Hai Huang 0004, Dongchun Ren, Huaxia Xia
Knowl. Based Syst.1
2023 CSR: Cascade Conditional Variational Auto Encoder with Socially-aware Regression for Pedestrian Trajectory Prediction
Hao Zhou 0014, Dongchun Ren, Xu Yang 0004, Mingyu Fan, Hai Huang 0004
Pattern Recognit.1
2023 CSIR: Cascaded Sliding CVAEs With Iterative Socially-Aware Rethinking for Trajectory Prediction
abstract
Pedestrian trajectory prediction is a hot research topic in many applications, such as video surveillance and autonomous driving. Although many efforts have been done on this topic, there are still many challenges, including accumulated prediction errors, insufficient training data usage, and future-past incompatibility. To overcome these challenges, we propose a novel trajectory prediction method, called CSIR, which consists of a cascaded sliding conditional variational autoencoder (CS-CVAE) module and an iterative future-past social compatible rethinking (I-SCR) module. The CS-CVAE module reduces the accumulated prediction errors by using cascaded prediction models for the early future time steps. In this way, the training losses of the early time steps are separately considered and minimized from the later losses. For the following time steps in CS-CVAE, a sliding prediction model with a longer observation time span is used and additional data from the future time span can be collected for training. On the other hand, the I-SCR module generates offsets to improve the predictions iteratively by checking the interaction compatibility between the predicted trajectories and the past trajectories, which resembles with the human rethinking mechanism in motion planning. Experiments results on two widely explored pedestrian trajectory prediction datasets, Stanford Drone Dataset (SDD) and ETH/UCY, show that the proposed method surpasses previous state-of-the-art methods by notable margins.
Hao Zhou 0014, Xu Yang 0004, Dongchun Ren, Hai Huang 0004, Mingyu Fan
IEEE Trans. Intell. Transp. Syst.1
2022 Out-of-Distribution Identification: Let Detector Tell Which I Am Not Sure
Ruoqi Li, Hao Zhou 0014, Yan Luo 0003
ECCV (10)3
2022 Informed Patch Enhanced HyperGCN for skeleton-based action recognition
Ying Li 0016, Hao Zhou 0014, Yan Luo 0003, Chuanping Hu
Inf. Process. Manag.4
2022 CANet: Co-attention network for RGB-D semantic segmentation
Hao Zhou 0014, Lu Qi 0001, Hai Huang 0004, Xu Yang 0004, Zhaoliang Wan, Xianglong Wen
Pattern Recognit.1
2022 Thinking Inside Uncertainty: Interest Moment Perception for Diverse Temporal Grounding
abstract
Given a language query, temporal grounding task is to localize temporal boundaries of the described event in an untrimmed video. There is a long-standing challenge that multiple moments may be associated with one same video-query pair, termed label uncertainty. However, existing methods struggle to localize diverse moments due to the lack of multi-label annotations. In this paper, we propose a novel Diverse Temporal Grounding framework (DTG) to achieve diverse moment localization with only single-label annotations. By delving into the label uncertainty, we find the diverse moments retrieved tend to involve similar actions/objects, driving us to perceive these interest moments. Specifically, we construct soft multi-label through semantic similarity of multiple video-query pairs. These soft labels reveal whether multiple moments in the intra-videos contain similar verbs/nouns, thereby guiding interest moment generation. Meanwhile, we put forward a diverse moment regression network (DMRNet) to achieve multiple predictions in a single pass, where plausible moments are dynamically picked out from the interest moments for joint optimization. Moreover, we introduce new metrics that better reveal multi-output performance. Extensive experiments conducted on Charades-STA and ActivityNet Captions show that our method achieves state-of-the-art performance in terms of both standard and new metrics.
Hao Zhou 0014, Yan Luo 0003, Chuanping Hu, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal Grounding
abstract
Temporal grounding aims to localize temporal boundaries within untrimmed videos by language queries, but it faces the challenge of two types of inevitable human uncertainties: query uncertainty and label uncertainty. The two uncertainties stem from human subjectivity, leading to limited generalization ability of temporal grounding. In this work, we propose a novel DeNet (Decoupling and Debias) to embrace human uncertainty: Decoupling — We explicitly disentangle each query into a relation feature and a modified feature. The relation feature, which is mainly based on skeleton-like words (including nouns and verbs), aims to extract basic and consistent information in the presence of query uncertainty. Meanwhile, modified feature assigned with style-like words (including adjectives, adverbs, etc) represents the subjective information, and thus brings personalized predictions; De-bias — We propose a de-bias mechanism to generate diverse predictions, aim to alleviate the bias caused by single-style annotations in the presence of label uncertainty. Moreover, we put forward new multi-label metrics to diversify the performance evaluation. Extensive experiments show that our approach is more effective and robust than state-of-the-arts on Charades-STA and ActivityNet Captions datasets.
Hao Zhou 0014, Yan Luo 0003, Chuanping Hu
CVPR1
2021 AST-GNN: An attention-based spatio-temporal graph neural network for Interaction-aware pedestrian trajectory prediction
Hao Zhou 0014, Dongchun Ren, Huaxia Xia, Mingyu Fan, Xu Yang 0004, Hai Huang 0004
Neurocomputing1
2021 Improving Visual Relationship Detection With Two-Stage Correlation Exploitation
abstract
Visual relationship detection, as a challenging task used to find and distinguish interactions between object-pairs in one image, has received much attention recently. In this work, we devise a unified visual relationship detection framework with two types of correlation exploitation to address the combination explosion problem in the object-pairs proposing stage and the non-exclusive label problem in the predicate recognition stage. In the object-pairs proposing stage, with the exploitation of relative location correlation between two objects in one pair, one location-embedded rating module (LRM) is developed to effectively select plausible proposals. In the predicate recognition stage, one label-correlation graph module (LGM) is introduced to measure the implicit semantic correlation among predicates; and then assign discrete distributed labels to predicates to improve the precision of top-n recall. Experiments on the two widely used VRD and VG datasets show that our proposed method outperforms current state-of-the-art methods.
Hao Zhou 0014, Muming Zhao, Yan Luo 0003, Chuanping Hu
IEEE Trans. Circuits Syst. Video Technol.1
2020 RGB-D Co-attention Network for Semantic Segmentation
Hao Zhou 0014, Lu Qi 0001, Zhaoliang Wan, Hai Huang 0004, Xu Yang 0004
ACCV (1)1
2020 An Attention-Based Interaction-Aware Spatio-Temporal Graph Neural Network for Trajectory Prediction
Hao Zhou 0014, Dongchun Ren, Huaxia Xia, Mingyu Fan, Xu Yang 0004, Hai Huang 0004
ICONIP (5)1
2019 Visual Relationship Detection with Relative Location Mining
abstract
Visual relationship detection, as a challenging task used to find and distinguish the interactions between object pairs in one image, has received much attention recently. In this work, we propose a novel visual relationship detection framework by deeply mining and utilizing relative location of object-pair in every stage of the procedure. In both the stages, relative location information of each object-pair is abstracted and encoded as auxiliary feature to improve the distinguishing capability of object-pairs proposing and predicate recognition, respectively; Moreover, one Gated Graph Neural Network(GGNN) is introduced to mine and measure the relevance of predicates using relative location. With the location-based GGNN, those non-exclusive predicates with similar spatial position can be clustered firstly and then be smoothed with close classification scores, thus the accuracy of top n recall can be increased further. Experiments on two widely used datasets VRD and VG show that, with the deeply mining and exploiting of relative location information, our proposed model significantly outperforms the current state-of-the-art.
Hao Zhou 0014, Chuanping Hu
ACM Multimedia1
2019 Faster R-CNN for marine organisms detection and recognition using data augmentation
Hai Huang 0004, Hao Zhou 0014, Xu Yang 0004, Lu Zhang 0054, Lu Qi 0001, Ai-Yun Zang
Neurocomputing2
2018 Single Shot Feature Aggregation Network for Underwater Object Detection
abstract
The rapidly developing ocean exploration and observation make the demand for underwater object detection become increasingly urgent. Recently, deep convolutional neural networks (CNN) have shown strong ability in feature representation and CNN-based detectors also achieve remarkable performance, but still facing the big challenge when detecting multi-scale objects in a complex underwater environment. To address this challenge, we propose a novel underwater object detector, introducing multiscale features and complementary context information for better classification and location ability. In the auto-grabbing contest of 2017 Underwater Robot Picking Contest sponsored by National Natural Science Foundation of China (NSFC), we won the 1-st place by using proposed method for real coastal underwater object detection.
Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001, Lu Qi 0001, Hao Zhou 0014, Charles Chiu
ICPR5