Hongjie Zhang 0002

dblp:89/3863-2 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-5085-0765ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
abstract
Video temporal grounding (VTG) aims to locate precise segments in videos based on language queries, which is a fundamental challenge in video understanding. While recent Multimodal Large Language Models (MLLMs) have shown promise in tackling VTG through reinforcement learning (RL), they overlook the challenges arising from both the quality and difficulty of training samples. (1) Partially annotated samples. Many manually annotated samples contain relevant segments beyond the annotated interval, introducing ambiguous supervision. (2) Hard-to-ground samples. Samples with poor zero-shot performance produce consistently low and indistinguishable rewards during RL training, exhibiting no clear preference among multiple outputs and thus hindering learning efficiency. To address these challenges, we propose VideoTG-R1, a novel curriculum RL framework with reflected boundary annotations, enabling data-efficient training. Specifically, we propose a Boundary Reflection Agent that utilizes MLLMs to predict query-relevant timestamps outside the annotated intervals, allowing us to identify and filter out partially annotated samples, thereby reducing ambiguity. Furthermore, we introduce a Difficulty Estimation Agent to assess the training difficulty of each sample and design a curriculum RL strategy that dynamically masks the videos of hard-to-ground samples according to the training steps, easing the training difficulty and providing clearer preference. Experiments on the VTG and grounded VideoQA tasks demonstrate the effectiveness of our method. Remarkably, with only 10% of the training samples and 21% of the computational budget, VideoTG-R1 outperforms full-data counterparts under both group relative policy optimization (GRPO) and supervised fine-tuning (SFT). The code is available at https://github.com/ldong1111/VideoTG-R1.
Lu Dong 0005, Ziang Yan, Xiangyu Zeng 0004, Hongjie Zhang 0002, Yifei Huang 0006, Yi Wang 0033, Zhen-Hua Ling, Limin Wang 0002, Yali Wang 0001
ICMR6
2025 ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol Spotting
abstract
Recognizing symbols in architectural CAD drawings is critical for various advanced engineering applications. In this paper, we propose a novel CAD data annotation engine that leverages intrinsic attributes from systematically archived CAD drawings to automatically generate high-quality annotations, thus significantly reducing manual labeling efforts. Utilizing this engine, we construct ArchCAD-400K, a large-scale CAD dataset consisting of 413,062 chunks from 5538 highly standardized drawings, making it over 26 times larger than the largest existing CAD dataset. ArchCAD-400K boasts an extended drawing diversity and broader categories, offering line-grained annotations. Furthermore, we present a new baseline model for panoptic symbol spotting, termed Dual-Pathway Symbol Spotter (DPSS). It incorporates an adaptive fusion module to enhance primitive features with complementary image features, achieving state-of-the-art performance and enhanced robustness. Extensive experiments validate the effectiveness of DPSS, demonstrating the value of ArchCAD-400K and its potential to drive innovation in architectural design and construction.
Ruifeng Luo, Zhengjie Liu, Tianxiao Cheng, Tongjie Wang, Fu Chai, Xingguang Wei, Haomin Wang 0002, Shenglong Ye, Wenhai Wang, Yu Qiao 0001, Hongjie Zhang 0002, Xianzhong Zhao
NeurIPS15
2025 Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
abstract
We study the task of panoptic symbol spotting, which involves identifying both individual instances of countable \textit{things} and the semantic regions of uncountable \textit{stuff} in computer-aided design (CAD) drawings composed of vector graphical primitives. Existing methods typically rely on image rasterization, graph construction, or point-based representation, but these approaches often suffer from high computational costs, limited generality, and loss of geometric structural information. In this paper, we propose \textit{VecFormer}, a novel method that addresses these challenges through \textit{line-based representation} of primitives. This design preserves the geometric continuity of the original primitive, enabling more accurate shape representation while maintaining a computation-friendly structure, making it well-suited for vector graphic understanding tasks. To further enhance prediction reliability, we introduce a \textit{Branch Fusion Refinement} module that effectively integrates instance and semantic predictions, resolving their inconsistencies for more coherent panoptic outputs. Extensive experiments demonstrate that our method establishes a new state-of-the-art, achieving 91.1 PQ, with Stuff-PQ improved by 9.6 and 21.2 points over the second-best results under settings with and without prior information, respectively—highlighting the strong potential of line-based representation as a foundation for vector graphic understanding.
Xingguang Wei, Haomin Wang 0002, Shenglong Ye, Ruifeng Luo, Lixin Gu, Jifeng Dai, Yu Qiao 0001, Wenhai Wang, Hongjie Zhang 0002
NeurIPS10
2025 LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Hongjie Zhang 0002, Lu Dong 0005, Yi Liu 0081, Yifei Huang 0002, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001
Int. J. Comput. Vis.1
2025 Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
abstract
The task of weakly supervised temporal sentence grounding (WSTSG) aims to detect temporal intervals corresponding to a language description from untrimmed videos with only video-level video-language correspondence. For an anchor sample, most existing approaches generate negative samples either from other videos or within the same video for contrastive learning. However, some training samples are highly similar to the anchor sample, directly regarding them as negative samples leads to difficulties for optimization and ignores the correlations between these similar samples and the anchor sample. To address this, we propose Positive Sample Mining (PSM), a novel framework that mines positive samples from the training set to provide more discriminative supervision. Specifically, for a given anchor sample, we partition the remaining training set into semantically similar and dissimilar subsets based on the similarity of their text queries. To effectively leverage these correlations, we introduce a PSM-guided contrastive loss to ensure that the anchor proposal is closer to similar samples and further from dissimilar ones. Additionally, we design a PSM-guided rank loss to ensure that similar samples are closer to the anchor proposal than to the negative intra-video proposal, aiming to distinguish the anchor proposal and the negative intra-video proposal. Experiments on the WSTSG and grounded VideoQA tasks demonstrate the effectiveness and superiority of our method.
Lu Dong 0005, Hongjie Zhang 0002, Yifei Huang 0002, Zhen-Hua Ling, Yu Qiao 0001, Limin Wang 0002, Yali Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
abstract
Being able to map the activities of others into one's own point of view is a fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that emulates the human demonstration following process, in which individuals record egocentric videos as they execute tasks guided by exocentric-view demonstration videos. Focusing on the potential applications in daily assistance and professional support, EgoExoLearn contains egocentric and demonstration video data spanning 120 hours captured in daily life scenarios and specialized laboratories. Along with the videos we record high-quality gaze data and provide detailed multimodal annotations, formulating a playground for modeling the human ability to bridge asynchronous procedural actions from different viewpoints. To this end, we present benchmarks such as crossview association, cross-view action planning, and crossview referenced skill assessment, along with detailed analysis. We expect EgoExoLearn can serve as an important resource for bridging the actions across views, thus paving the way for creating AI agents capable of seamlessly learning by observing humans in the real world. The dataset and benchmark codes are available at https://github.com/OpenGVLab/EgoExoLearn.
Yifei Huang 0002, Guo Chen 0006, Jilan Xu, Mingfang Zhang 0002, Lijin Yang, Baoqi Pei, Hongjie Zhang 0002, Lu Dong 0005, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001
CVPR7
2024 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Yi Wang 0074, Kunchang Li 0002, Xinhao Li 0004, Jiashuo Yu, Yinan He, Guo Chen 0006, Baoqi Pei, Rongkun Zheng, Zun Wang 0001, Yansong Shi, Tianxiang Jiang, Jilan Xu, Hongjie Zhang 0002, Yifei Huang 0002, Yu Qiao 0001, Yali Wang 0001, Limin Wang 0002
ECCV (85)14
2024 Matching Compound Prototypes for Few-Shot Action Recognition
abstract
Abstract The task of few-shot action recognition aims to recognize novel action classes using only a small number of labeled training samples. How to better describe the action in each video and how to compare the similarity between videos are two of the most critical factors in this task. Directly describing the video globally or by its individual frames cannot well represent the spatiotemporal dependencies within an action. On the other hand, naively matching the global representations of two videos is also not optimal since action can happen at different locations in a video with different speeds. In this work, we propose a novel approach that describes each video using multiple types of prototypes and then computes the video similarity with a particular matching strategy for each type of prototypes. To better model the spatiotemporal dependency, we describe the video by generating prototypes that model the multi-level spatiotemporal relations via transformers. There are a total of three types of prototypes. The first type of prototypes are trained to describe specific aspects of the action in the video e.g., the start of the action, regardless of its timestamp. These prototypes are directly matched one-to-one between two videos to compare their similarity. The second type of prototypes are the timestamp-centered prototypes that are trained to focus on specific timestamps of the video. To deal with the temporal variation of actions in a video, we apply bipartite matching to allow the matching of prototypes of different timestamps. The third type of prototypes are generated from the timestamp-centered prototypes, which regularize their temporal consistency while serving as an auxiliary summarization of the whole video. Experiments demonstrate that our proposed method achieves state-of-the-art results on multiple benchmarks.
Yifei Huang 0002, Lijin Yang, Guo Chen 0006, Hongjie Zhang 0002, Feng Lu 0005, Yoichi Sato 0001
Int. J. Comput. Vis.4
2023 Learning Discriminative Feature Representation for Open Set Action Recognition
abstract
Open set action recognition (OSAR) is a challenging task that requires a classifier to identify actions that do not belong to any of the classes in its training set. Existing methods employ the Evidential Neural Network (ENN) as an open-set classifier, which is trained in a supervised manner on feature representations from known classes to quantify the predictive uncertainty of human actions. In this paper, we propose a novel framework for OSAR that enriches the discriminative representation from a backbone with a reconstructive one to further improve performance. Our approach involves augmenting the input features with their reconstruction obtained from a reconstruction-based model in unsupervised training on known classes. We then use the correspondence between the two features to learn the open-set classifier, forcing it to associate low correspondence both when the feature is from unknown classes as well as when the input feature and its reconstruction variant are inconsistent with each other. Our experimental results on standard OSAR benchmarks demonstrate that our end-to-end trained model significantly outperforms state-of-the-art methods. Our proposed approach shows the effectiveness of combining discriminative and reconstructive representations for OSAR.
Hongjie Zhang 0002, Yi Liu 0081, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001
ACM Multimedia1
2023 Elastic temporal alignment for few-shot action recognition
abstract
Abstract Few‐shot action recognition aims to learn a classification model with good generalisation ability when trained with only a few labelled videos. However, it is difficult to learn discriminative feature representations for videos in such a setting. The Elastic Temporal Alignment (ETA) for few‐shot action recognition is proposed. First, a convolutional neural network is employed to extract feature representations of video frames sparsely sampled from videos. In order to obtain the similarity of two videos, a temporal alignment estimation function is utilised to estimate the matching score between each pair of frames from the two videos through an elastic alignment mechanism. The analysis shows that when we judge whether two frames from respective videos are matched, multiple adjacent frames in the videos should be considered, so as to embody the temporal information. Thus, before feeding per‐frame feature vectors of videos into the temporal alignment estimation function, a temporal message passing function is leveraged to propagate the information of per‐frame features in the temporal domain. The method has been evaluated on four action recognition datasets, including Kinetics, Something‐Something V2, HMDB51, and UCF101. The experimental results verify the effectiveness of ETA and show its superiority over state‐of‐the‐art methods.
Chunlei Xu, Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001
IET Comput. Vis.3
2023 Improving Open Set Domain Adaptation Using Image-to-Image Translation and Instance-Weighted Adversarial Learning
Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001
J. Comput. Sci. Technol.1
2020 Hybrid Models for Open Set Recognition
Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001
ECCV (3)1
2019 Improving Open Set Domain Adaptation Using Image-to-Image Translation
abstract
The open set domain adaptation problem was rarely studied and its existing solutions are mostly based on learning a joint latent space which may encounter issues when the domains differ significantly from each other. This work is driven by the question whether or not it is beneficial to operate the source images to another image domain as close to the target as possible. We propose to address the open set domain adaptation problem by aligning sample at both feature space and pixel space. Our approach, called Open Set Translation and Adaptation Network (Ostan), consists of two main components: translation and adaptation. The translation model is a cycle-consistent generative adversarial network, which translates any source sample to the "style" of a target domain. The adaptation network is built upon OpenBP, an open set domain adaptation framework, and trained using both (labeled) translated source images and (unlabeled) target images. The proposed Ostan model significantly outperforms the state-of-the-art open set domain adaptation methods on multiple public datasets. Our experiment also demonstrates that an image-to-image translation component can further improve the decision boundaries for both known and unknown classes.
Hongjie Zhang 0002, Yang Zhang 0053, Yanwen Guo 0001
ICME1
2019 Viewpoint Assessment and Recommendation for Photographing Architectures
abstract
This paper studies the problem of how to assess the quality of photographing viewpoints and how to choose good viewpoints for taking photographs of architectures. We achieve this by learning from photographs of world famous landmarks that are available on the Internet and their viewpoint quality ranked by online user annotation. Unlike previous efforts devoted to photo quality assessment which mainly rely on 2D image features, we show in this paper combining 2D image features extracted from images with 3D geometric features computed on the 3D models can result in more reliable evaluation of viewpoint quality. Specifically, we collect a set of photographs for each of 15 world famous architectures as well as their 3D models from the Internet. Viewpoint recovery for images is carried out through an image-model registration process, after which a newly proposed viewpoint clustering strategy is exploited to validate users' viewpoint preferences when photographing landmarks. Finally, we extract a number of 2D and 3D features for each image based on multiple visual and geometric cues and perform viewpoint recommendation by learning from both 2D and 3D features using a specifically designed SVM-2K multi-view learner, achieving superior performance over using solely 2D or 3D features. We show the effectiveness of the proposed approach through extensive experiments. The experiments also demonstrate that our system can be used to recommend viewpoints for rendering textured 3D models of buildings for the use of architectural design, in addition to viewpoint evaluation of photographs and recommendation of viewpoints for photographing architectures in practice.
Jingwu He, Linbo Wang 0001, Wenzhe Zhou, Hongjie Zhang 0002, Xiufen Cui, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.4
2018 Correlation-Preserving Photo Collage
abstract
A new method is presented for producing photo collages that preserve content correlation of photos. We use deep learning techniques to find correlation among given photos to facilitate their embedding on the canvas, and develop an efficient combinatorial optimization technique to make correlated photos stay close to each other. To make efficient use of canvas space, our method first extracts salient regions of photos and packs only these salient regions. We allow the salient regions to have arbitrary shapes, therefore yielding informative, yet more compact collages than by other similar collage methods based on salient regions. We present extensive experimental results, user study results, and comparisons against the state-of-the-art methods to show the superiority of our method.
Lingjie Liu, Hongjie Zhang 0002, Guangmei Jing, Yanwen Guo 0001, Zhonggui Chen, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.2