Zheyun Qin

dblp:256/8991 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-2564-071XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval
Xuri Ge, Chunhao Wang, Xindi Wang 0001, Zheyun Qin, Zhumin Chen, Xin Xin 0003
WWW4
2026 Fairness-aware graph representation learning through bias disentanglement
Zheyun Qin, Zhaohui Peng
Inf. Softw. Technol.2
2025 Sliced Wasserstein Bridge for Open-Vocabulary Video Instance Segmentation
Zheyun Qin, Deng Yu, Chuanchen Luo, Zhumin Chen
ICCV1
2025 Video Instance Segmentation by Weighted Structure Inference
abstract
Video instance segmentation presents significant challenges in complex and dynamic environments, where instances experience progressive occlusion, either from objects obstructing each other or due to changes in the camera's viewpoint. Current state-of-the-art methods rely on memory bank mechanisms, but we still look forward to new paradigms that have the ability to capture and utilize structural information, the ability to model complex relationships, and the flexibility to adapt to dynamic scenarios. To this end, we propose the Weighted Structure Inference method for Video Instance Segmentation. We build on high-order structural relationships by constructing hypergraphs for each video frame, enabling the capture of complex interactions that go beyond traditional pairwise methods. To model intricate dynamics, we introduce Weighted Sheaf Hypergraph Convolution, which enhances the hierarchical and structural information embedded in the hypergraph. Furthermore, we ensure spatio-temporal consistency by employing a dynamic inference mechanism based on Weighted Sliced Wasserstein distance to compare structural features across adjacent frames. Our method preserves the topological characteristics of occlusion instances and improves the reliability of instance tracking across frames. Experimental results demonstrate that our method outperforms existing video instance segmentation frameworks in both Video Instance and Panoptic Segmentation tasks.
Zheyun Qin, Deng Yu, Qiangchang Wang, Zhumin Chen
ACM Multimedia1
2024 Distribution-Aware Contrastive Learning for Robust Medical Image Segmentation
abstract
Medical image segmentation is pivotal in quantifying tissue volumes, facilitating diagnoses, and enabling other critical medical applications. However, accurately segmenting medical images can be challenging because the complex intensity distribution inherent in the data arises from the highly complex interaction of many latent factors (data heterogeneity). In this context, we propose a novel method called Distribution-aware Contrastive Learning for Robust Segmentation (DCL-Seg) to address the inconsistency in medical image segmentation. Based on the assumption of content separability, we use learnable parameters to construct positive samples with a potential structure invariance via contrastive learning. In this way, our method can mitigate the negative effects of data heterogeneity to separate overlapped class distribution and structural solid boundary. We are in one public dataset and two clinical datasets for Breast tumor and Retinal vessel segmentation, which have achieved excellent results and widely proved the superiority of our method.
Zheyun Qin, Xiaoming Xi, Yilong Yin
ICASSP1
2023 Exposing the Self-Supervised Space-Time Correspondence Learning via Graph Kernels
abstract
Self-supervised space-time correspondence learning is emerging as a promising way of leveraging unlabeled video. Currently, most methods adapt contrastive learning with mining negative samples or reconstruction adapted from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence predictive models require mining more inherent properties in videos, such as structural information. In this work, we propose the VideoHiGraph, a space-time correspondence framework based on a learnable graph kernel. Concerning the video as the spatial-temporal graph, the learning objectives of VideoHiGraph are emanated in a self-supervised manner for predicting unobserved hidden graphs via graph kernel manner. We learn a representation of the temporal coherence across frames in which pairwise similarity defines the structured hidden graph, such that a biased random walk graph kernel along the sub-graph can predict long-range correspondence. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The self-supervision of the model training is formed by the structural and temporal consistency of the graph. VideoHiGraph achieves superior performance and demonstrates its robustness across the benchmark of label propagation tasks involving objects, semantic parts, keypoints, and instances. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/VideoHiGraph.
Zheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin, Jianbing Shen
AAAI1
2023 Unified 3D Segmenter As Prototypical Classifiers
abstract
The task of point cloud segmentation, comprising semantic, instance, and panoptic segmentation, has been mainly tackled by designing task-specific network architectures, which often lack the flexibility to generalize across tasks, thus resulting in a fragmented research landscape. In this paper, we introduce ProtoSEG, a prototype-based model that unifies semantic, instance, and panoptic segmentation tasks. Our approach treats these three homogeneous tasks as a classification problem with different levels of granularity. By leveraging a Transformer architecture, we extract point embeddings to optimize prototype-class distances and dynamically learn class prototypes to accommodate the end tasks. Our prototypical design enjoys simplicity and transparency, powerful representational learning, and ad-hoc explainability. Empirical results demonstrate that ProtoSEG outperforms concurrent well-known specialized architectures on 3D point cloud benchmarks, achieving 72.3%, 76.4% and 74.2% mIoU for semantic segmentation on S3DIS, ScanNet V2 and SemanticKITTI, 66.8% mCov and 51.2% mAP for instance segmentation on S3DIS and ScanNet V2, 62.4% PQ for panoptic segmentation on SemanticKITTI, validating the strength of our concept and the effectiveness of our algorithm. The code and models are available at https://github.com/zyqin19/PROTOSEG.
Zheyun Qin, Cheng Han 0001, Qifan Wang 0001, Xiushan Nie, Yilong Yin, Xiankai Lu
NeurIPS1
2023 Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence Learning
abstract
Self-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph.
Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui
IEEE Trans. Image Process.1
2022 Dpnet: end-to-end Aerial Image Segmentation Via Deformable Point Network
abstract
Aerial image Segmentation segmentation faces intrinsic foreground-background imbalance and background clutter distraction. To guide the segmentation model to learn more discriminative foreground ability and more invariant back-ground representation features, we design a Deformable Point Network (DPNet). It is an end-to-end segmentation network and consists of a multi-head deformable attention module that simultaneously considers foreground object information and background suppression. Specifically, we first employ a feature pyramid network to aggregate multiple-layer features to handle scale variants. And then, we further investigate deformable convolution to select some representative points for each layer and propose a differential module to implement it automatically instead of traditional dense fusion. Moreover, we incorporate the multi-head mechanism in the feature fusion to focus on the key contents from different representation regions. Experimental results on the representative iSAID, Vaihingen, and Postdam datasets demonstrate that our DPNet achieves competitive performance. Also, the multiple-head deformable attention facilitates the network convergence significantly.
Yiyou Guo, Zheyun Qin, Yongtai Yang, Xiankai Lu, Huan Xie 0001, Xiaohua Tong
IGARSS2
2022 Difficulty-aware bi-network with spatial attention constrained graph for axillary lymph node segmentation
Xiaoming Xi, Xianjing Meng, Zheyun Qin, Xiushan Nie, Yongjian Wu 0001, Chenglong Li 0004, Yilong Yin
Sci. China Inf. Sci.4
2022 Learning disentangled representation for self-supervised video object segmentation
Wenjie Hou, Zheyun Qin, Xiaoming Xi, Xiankai Lu, Yilong Yin
Neurocomputing2
2021 Learning Hierarchical Embedding for Video Instance Segmentation
abstract
In this paper, we address video instance segmentation using a new generative model that learns effective representations of the target and background appearance. We propose to exploit hierarchical structural embedding over spatio-temporal space, which is compact, powerful, and flexible in contrast to current tracking-by-detection methods. Specifically, our model segments and tracks instances across space and time in a single forward pass, which is formulated as hierarchical embedding learning. The model is trained to locate the pixels belonging to specific instances over a video clip. We firstly take advantage of a novel mixing function to better fuse spatio-temporal embeddings. Moreover, we introduce normalizing flows to further improve the robustness of the learned appearance embedding, which theoretically extends conventional generative flows to a factorized conditional scheme. Comprehensive experiments on the video instance segmentation benchmark, i.e., YouTube-VIS, demonstrate the effectiveness of the proposed approach. Furthermore, we evaluate our method on an unsupervised video object segmentation dataset to demonstrate its generalizability.
Zheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen, Yilong Yin
ACM Multimedia1