Jiachang Hao

dblp:210/8416 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-6842-4721ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 From Static to Dynamic: GNNs-Driven Clinical Decision-Making Assistance
Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao, Jiachang Hao, Haifeng Sun 0001
DASFAA (2)6
2023 Two Heads Are Better than One: Image-Point Cloud Network for Depth-Based 3D Hand Pose Estimation
abstract
Depth images and point clouds are the two most commonly used data representations for depth-based 3D hand pose estimation. Benefiting from the structuring of image data and the inherent inductive biases of the 2D Convolutional Neural Network (CNN), image-based methods are highly efficient and effective. However, treating the depth data as a 2D image inevitably ignores the 3D nature of depth data. Point cloud-based methods can better mine the 3D geometric structure of depth data. However, these methods suffer from the disorder and non-structure of point cloud data, which is computationally inefficient. In this paper, we propose an Image-Point cloud Network (IPNet) for accurate and robust 3D hand pose estimation. IPNet utilizes 2D CNN to extract visual representations in 2D image space and performs iterative correction in 3D point cloud space to exploit the 3D geometry information of depth data. In particular, we propose a sparse anchor-based "aggregation-interaction-propagation'' paradigm to enhance point cloud features and refine the hand pose, which reduces irregular data access. Furthermore, we introduce a 3D hand model to the iterative correction process, which significantly improves the robustness of IPNet to occlusion and depth holes. Experiments show that IPNet outperforms state-of-the-art methods on three challenging hand datasets.
Pengfei Ren 0001, Jiachang Hao, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
AAAI3
2023 Reasoning Guided by a Manual: Context-Aware Image Captioning with Novel Objects
abstract
Novel object captioning task aims at describing objects that are absent from training data. Due to the scarcity of novel objects, it’s challenging to find a way to utilize external data to improve model’s reasoning ability. While previously designed methods all follow a deep learning approach, we boost novel object captioning by incorporating reasoning with traditional deep learning framework. We design a manual from dictionaries that provides our model with sufficient and accurate external information on novel objects. We propose Manual-guided Context-aware Novel Object Captioning model (MC-NOC) that utilizes image and caption context to generate novel object captions. It contains a Manual-Guided Novel Object Reasoning module to reason about novel objects based on other objects of the given image and a Caption Reconstruction module to incorporate novel objects into generated captions according to caption context. We validate MC-NOC with state-of-the-art performance on the challenging Held-out COCO and Nocaps dataset, leading their leaderboard. In particular, we improved the CIDER metric by 6.4 points on the held-out coco dataset. Comprehensive experiments demonstrate our model’s reasoning capability and the quality of generated captions.
Peiyao Hua, Haifeng Sun 0001, Jiachang Hao, Cong Liu 0046, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao
ECAI3
2023 Turn on the Right Track: Weakly Supervised Video Moment Retrieval with Self-Improving Query Reconstruction
abstract
Existing weakly-supervised temporal sentence grounding methods typically regard query reconstruction as the pretext task in place of the absent temporal supervision. However, their approaches suffer from two flaws, i.e. insignificant reconstruction and discrepancy in alignment. Insignificant reconstruction indicates the randomly masked words may not be discriminative enough to distinguish the target event from unrelated events in the video. Discrepancy in alignment indicates the incorrect partial alignment built by query reconstruction task. The flaws undermine the reliability of current reconstruction-based methods. To this end, we propose a novel Self-improving Query ReconstrucTion (SQRT) framework for weakly-supervised temporal sentence grounding. To deal with insignificant reconstruction, we devise a key words mining strategy to determine the important words for language grounding. To attain better moment-query alignment, we introduce inter-sample contrast to tackle the partial alignment built by query reconstruction. The self-improving framework utilizes query reconstruction for language grounding and alleviates the discrepancy in alignment, thus turning on the right track. Experiments on two popular datasets show that SQRT achieves state-of-the-art performance on Charades-STA and comparable performance to the state-of-the-art on ActivityNet Captions.
Haifeng Sun 0001, Jiachang Hao, Jing Wang 0039, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
ECAI3
2023 Pose-Guided Hierarchical Graph Reasoning for 3-D Hand Pose Estimation From a Single Depth Image
abstract
Estimating 3-D hand pose estimation from a single depth image is important for human-computer interaction. Although depth-based 3-D hand pose estimation has made great progress in recent years, it is still difficult to deal with some complex scenes, especially the issues of serious self-occlusion and high self-similarity of fingers. Inspired by the fact that multipart context is critical to alleviate ambiguity, and constraint relations contained in the hand structure are important for the robust estimation, we attempt to explicitly model the correlations between different hand parts. In this article, we propose a pose-guided hierarchical graph convolution (PHG) module, which is embedded into the pixelwise regression framework to enhance the convolutional feature maps by exploring the complex dependencies between different hand parts. Specifically, the PHG module first extracts hierarchical fine-grained node features under the guidance of hand pose and then uses graph convolution to perform hierarchical message passing between nodes according to the hand structure. Finally, the enhanced node features are used to generate dynamic convolution kernels to generate hierarchical structure-aware feature maps. Our method achieves state-of-the-art performance or comparable performance with the state-of-the-art methods on five 3-D hand pose datasets: 1) HANDS 2019; 2) HANDS 2017; 3) NYU; 4) ICVL; and 5) MSRA.
Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
IEEE Trans. Cybern.3
2023 Fine-Grained Text-to-Video Temporal Grounding from Coarse Boundary
abstract
Text-to-video temporal grounding aims to locate a target video moment that semantically corresponds to the given sentence query in an untrimmed video. In this task, fully supervised works require text descriptions for each event along with its temporal segment coordinate for training, which is labor-consuming. Existing weakly supervised works require only video-sentence pairs but cannot achieve satisfactory performance. However, many available annotations in the form of coarse temporal boundaries for sentences are ignored and unexploited. These coarse boundaries are common in streaming media platform and can be collected in a mechanical manner. We propose a novel approach to perform fine-grained text-to-video temporal grounding from these coarse boundaries. We take dense video captioning as base task and leverage the trained captioning model to identify the relevance of each video frame to the sentence query according to the frame participation in event captioning. To quantify the frame participation in event captioning, we proposeevent activation sequence, a simple method that highlights the temporal regions which have high correlations to the text modality in videos. Experiments on modified ActivityNet Captions and a use case demonstrate the promising fine-grained performance of our approach.
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Mining Multi-View Information: A Strong Self-Supervised Framework for Depth-based 3D Hand Pose and Mesh Estimation
abstract
In this work, we study the cross-view information fusion problem in the task of self-supervised 3D hand pose estimation from the depth image. Previous methods usually adopt a hand-crafted rule to generate pseudo labels from multi-view estimations in order to supervise the network training in each view. However, these methods ignore the rich semantic information in each view and ignore the complex dependencies between different regions of different views. To solve these problems, we propose a cross-view fusion network to fully exploit and adaptively aggregate multi-view information. We encode diverse semantic information in each view into multiple compact nodes. Then, we introduce the graph convolution to model the complex dependencies between nodes and perform cross-view information interaction. Based on the cross-view fusion network, we propose a strong self-supervised framework for 3D hand pose and hand mesh estimation. Furthermore, we propose a pseudo multi-view training strategy to extend our framework to a more general scenario in which only single-view training data is used. Results on NYU dataset demonstrate that our method outperforms the previous self-supervised methods by 17.5% and 30.3% in multi-view and single-view scenarios. Meanwhile, our framework achieves comparable re-sults to several strongly supervised methods.
Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao
CVPR3
2022 Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao
ECCV (36)1
2022 Query-aware video encoder for video moment retrieval
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao
Neurocomputing1
2022 A Dual-Branch Self-Boosting Framework for Self-Supervised 3D Hand Pose Estimation
abstract
Although 3D hand pose estimation has made significant progress in recent years with the development of the deep neural network, most learning-based methods require a large amount of labeled data that is time-consuming to collect. In this paper, we propose a dual-branch self-boosting framework for self-supervised 3D hand pose estimation from depth images. First, we adopt a simple yet effective image-to-image translation technology to generate realistic depth images from synthetic data for network pre-training. Second, we propose a dual-branch network to perform 3D hand model estimation and pixel-wise pose estimation in a decoupled way. Through a part-aware model-fitting loss, the network can be updated according to the fine-grained differences between the hand model and the unlabeled real image. Through an inter-branch loss, the two complementary branches can boost each other continuously during self-supervised learning. Furthermore, we adopt a refinement stage to better utilize the prior structure information in the estimated hand model for a more accurate and robust estimation. Our method outperforms previous self-supervised methods by a large margin without using paired multi-view images and achieves comparable results to strongly supervised methods. Besides, by adopting our regenerated pose annotations, the performance of the skeleton-based gesture recognition is significantly improved.
Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
IEEE Trans. Image Process.3
2021 Spatial-aware stacked regression network for real-time 3D hand pose estimation
Pengfei Ren 0001, Haifeng Sun 0001, Weiting Huang, Jiachang Hao, Daixuan Cheng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
Neurocomputing4