Zhilong Ou

dblp:310/6507 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-8587-1854ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Modality-Aware Shot Relating and Comparing for Video Scene Detection
abstract
Video scene detection involves assessing whether each shot and its surroundings belong to the same scene. Achieving this requires meticulously correlating multi-modal cues, e.g., visual entity and place modalities, among shots and comparing semantic changes around each shot. However, most methods treat multi-modal semantics equally and do not examine contextual differences between the two sides of a shot, leading to sub-optimal detection performance. In this paper, we propose the Modality-Aware Shot Relating and Comparing approach (MASRC), which enables relating shots per their own characteristics of visual entity and place modalities, as well as comparing multi-shots similarities to have scene changes explicitly encoded. Specifically, to fully harness the potential of visual entity and place modalities in modeling shot relations, we mine long-term shot correlations from entity semantics while simultaneously revealing short-term shot correlations from place semantics. In this way, we can learn distinctive shot features that consolidate coherence within scenes and amplify distinguishability across scenes. Once equipped with distinctive shot features, we further encode the relations between preceding and succeeding shots of each target shot by similarity convolution, aiding in the identification of scene ending shots. We validate the broad applicability of the proposed components in MASRC. Extensive experimental results on public benchmark datasets demonstrate that the proposed MASRC significantly advances video scene detection.
Jiawei Tan, Hongxing Wang 0001, Kang Dang, Zhilong Ou
AAAI5
2025 Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2D
abstract
Video moment retrieval aims to locate specific moments from a video according to the query text. This task presents two main challenges: i) aligning the query and video frames at the feature level, and ii) projecting the query-aligned frame features to the start and end boundaries of the matching interval. Previous work commonly involves all frames in feature alignment, easy to cause aligning irrelevant frames with the query. Furthermore, they forcibly map visual features to interval boundaries but ignoring the information gap between them, yielding suboptimal performance. In this study, to reduce distraction from irrelevant frames, we designate an anchor frame as that with the maximum query-frame relevance measured by the established Vision-Language Model. Via similarity comparison between the anchor frame and the others, we produce a semantically compact segment around the anchor frame, which serves as a guide to align features of query and related frames. We observe that such a feature alignment will make similarity cohesive between target frames, which enables us to predict the interval boundaries by a single point detection in the 2D semantic similarity space of frames, thus well bridging the information gap between frame semantics and temporal boundaries. Experimental results across various datasets demonstrate that our approach significantly improves the alignment between queries and video frames while effectively predicting temporal moment boundaries. Especially, on QVHighlights Test and ActivityNet Captions datasets, our proposed approach achieves 3.8% and 7.4% respectively higher than current state-of-the-art [email protected] performance. The code is available at https://github.com/ExMorgan-Alter/AFAFSGD.
Jiawei Tan, Hongxing Wang 0001, Junwu Weng, Zhilong Ou, Kang Dang
CVPR5
2025 NCD: Normal-Guided Chamfer Distance Loss for Watertight Mesh Reconstruction from Unoriented Point Clouds
abstract
Abstract As a widely used loss function in learnable watertight mesh reconstruction from unoriented point clouds, Chamfer Distance (CD) efficiently quantifies the alignment between the sampled point cloud from the reconstructed mesh and its corresponding input point cloud. Occasionally, to enhance reconstruction fidelity, CD incorporates a normal consistency term, albeit at the cost of efficiency. In this context, normal estimation for unoriented point clouds requires computationally intensive matrix decomposition or specialized pre‐trained models, whereas deriving normals for mesh‐sampled points can be readily achieved using the cross product of mesh vertices. However, the reconstruction models employing CD and its variants typically rely solely on the spatial coordinates of the points, which omits normal information in favor of efficiency and deployability. To tackle this challenge, we propose a novel loss function for watertight mesh reconstruction from unoriented point clouds, termed Normal‐guided Chamfer Distance (NCD). Building upon CD, NCD introduces a normal‐steered weighting mechanism based on the angle between the normal at each mesh‐sampled point and the vector to its corresponding input point, offering several advantages: (i) it leverages readily available mesh‐sampled point normals to weight coordinate‐based Euclidean distances, thus extending the capability of CD; (ii) it eliminates the need for normal estimation from input unoriented point clouds; (iii) it incurs a negligible increase in computational complexity compared to CD. We employ NCD as the training loss for point‐to‐mesh reconstruction with multiple models and initial watertight meshes on benchmark datasets, demonstrating its superiority over state‐of‐the‐art CD variants.
Jiawei Tan, Zhilong Ou, Hongxing Wang 0001
Comput. Graph. Forum3
2025 Label refinement for change detection in remote sensing
Zhilong Ou, Hongxing Wang 0001, Jiawei Tan, Zhangbin Qian
Image Vis. Comput.1
2025 Aligning Instance-Semantic Sparse Representation Towards Unsupervised Object Segmentation and Shape Abstraction With Repeatable Primitives
abstract
Understanding 3D object shapes necessitates shape representation by object parts abstracted from results of instance and semantic segmentation. Promising shape representations enable computers to interpret a shape with meaningful parts and identify their repeatability. However, supervised shape representations depend on costly annotation efforts, while current unsupervised methods work under strong semantic priors and involve multi-stage training, thereby limiting their generalization and deployment in shape reasoning and understanding. Driven by the tendency of high-dimensional semantically similar features to lie in or near low-dimensional subspaces, we introduce a one-stage, fully unsupervised framework towards semantic-aware shape representation. This framework produces joint instance segmentation, semantic segmentation, and shape abstraction through sparse representation and feature alignment of object parts in a high-dimensional space. For sparse representation, we devise a sparse latent membership pursuit method that models each object part feature as a sparse convex combination of point features at either the semantic or instance level, promoting part features in the same subspace to exhibit similar semantics. For feature alignment, we customize an attention-based strategy in the feature space to align instance- and semantic-level object part features and reconstruct the input shape using both of them, ensuring geometric reusability and semantic consistency of object parts. To firm up semantic disambiguation, we construct cascade unfrozen learning on geometric parameters of object parts. Experiments conducted on benchmark datasets confirm that our approach results in instance- and semantic-level joint segmentation and shape abstraction with repeatable primitives, providing coherent semantic interpretations of 3D object shapes across categories in a one-stage, fully unsupervised manner, without relying on annotations or heuristic semantic priors.
Hongxing Wang 0001, Jiawei Tan, Zhilong Ou, Junsong Yuan 0001
IEEE Trans. Vis. Comput. Graph.4
2024 Neighbor Relations Matter in Video Scene Detection
abstract
Video scene detection aims to temporally link shots for obtaining semantically compact scenes. It is essential for this task to capture scene-distinguishable affinity among shots by similarity assessment. However, most methods relies on ordinary shot-to-shot similarities, which may inveigle similar shots into being linked even though they are from different scenes, and meanwhile hinder dissimilar shots from being blended into a complete scene. In this paper, we propose NeighborNet to inject shot contexts into shot-to-shot similarities through carefully exploring the relations between semantic/temporal neighbors of shots over a local time period. In this way, shot-to-shot similarities are remeasured as semantic/temporal neighbor-aware similarities so that NeighborNet can learn context embedding into shot features using graph convolutional network. As a result, not only do the learned shot features suppress the affinity among similar shots from different scenes, but they also promote the affinity among dissimilar shots in the same scene. Experimental results on public benchmark datasets show that our proposed NeighborNet yields substantial improvements in video scene detection, especially outperforms released state-of-the-arts by at least 6% in Average Precision (AP). The code is available at https://github.com/ExMorgan-Alter/NeighborNet.
Jiawei Tan, Hongxing Wang 0001, Zhilong Ou, Zhangbin Qian
CVPR4
2024 CLIP-Driven Multi-Scale Instance Learning for Weakly Supervised Video Anomaly Detection
abstract
Existing weakly supervised video anomaly detection methods mainly employ Multiple Instance Learning (MIL) to identify abnormal snippets in untrimmed videos. However, the semantics and presentations of anomalies frequently exhibit ambiguity that MIL is difficult to tackle. Moreover, MIL suffers from false alarms due to its independent optimization of each instance, neglecting temporal correlation between adjacent snippets. Consequently, we badly need to better connect abnormal presentations and their semantics, as well as to enable multi-temporal-scale anomaly discovery. This paper proposes a CLIP-Driven Multi-Scale Instance Learning (CMSIL) framework with two branches including Vision-Language (VL) and Multi-Scale Instance Learning (MSIL). The VL branch leverages the powerful visual concept priors from Contrastive Language-Image Pre-training (CLIP) to generate pseudo anomalies, thereby providing suspected anomaly cues for model training guidance. The MSIL branch utilizes a feature pyramid to fully mine fine-grained temporal dependencies by employing MIL within each pyramid level to learn anomalous patterns across different temporal scales. By collaborating with the two branches, the proposed CMSIL shows better proficiency in handling anomalies with varying durations. Extensive experiments on the XD-Violence and UCF-Crime datasets demonstrate the superior performance of our method. The code is available at https://github.com/casperZB/CMSIL.
Zhangbin Qian, Jiawei Tan, Zhilong Ou, Hongxing Wang 0001
ICME3
2024 Object Recognition Consistency in Regression for Active Detection
Ming Jing, Zhilong Ou, Hongxing Wang 0001
Mach. Vis. Appl.2
2024 Ultra-FastNet: an end-to-end learnable network for multi-person posture prediction
Tiandi Peng, Yanmin Luo 0001, Zhilong Ou, Jixiang Du, Gonggeng Lin
J. Supercomput.3
2023 A perception-enhancement network for accurate multi-person 2D pose estimation
Yanmin Luo 0001, Zhilong Ou, Zhiqian Zhang, Jin Gou, Jing-Ming Gou
Appl. Intell.2
2022 FastNet: Fast high-resolution network for human pose estimation
Yanmin Luo 0001, Zhilong Ou, TianJun Wan, Jing-Ming Guo
Image Vis. Comput.2
2022 SRFNet: selective receptive field network for human pose estimation
Zhilong Ou, Yanmin Luo 0001
J. Supercomput.1