VLDB 2026 Research / reviewers in the wild / expert
Zun Li 0001
dblp:215/9536-1
· DBLP profile ↗
25ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0001-6100-2788ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCAHCL: Multi-Level Hypergraph Contrastive Learning for Connected Component AwarenessabstractHypergraph contrastive learning has emerged as a powerful unsupervised paradigm for hypergraph representation learning. Traditional hypergraph contrastive learning methods typically leverage neighbor aggregation strategy to obtain entity (node and hyperedge) representations within each connected component, and then utilize contrastive losses (e.g., node- or hyperedge-level) to update the encoders. However, since entities are usually focused equally on their respective losses, large connected components with numerous entities tend to provide a dominant contribution to the whole learning process, which inevitably hinders the effective learning of entity representations within small connected components. To address this issue, we propose a novel Connected-Component-Aware Hypergraph Contrastive Learning method (CCAHCL). Different from previous methods that only construct node or hyperedge representations, our method additionally constructs the connected component representations, and accordingly designs a hierarchical contrastive loss to balance the model's focus on different scales of connected components. Specifically, we first use the traditional neighbor aggregation strategy to aggregate and update entity (node and hyperedge) representations. Then, these entity representations are further aggregated to generate the connected component representations, where entity features are incorporated into connected components and their structural information is propagated back to enrich their corresponding entities. Afterwards, we employ node-level and hyperedge-level losses to learn the enriched entity representations, and further propose a novel connected-component-level contrastive loss to balance the model's focus on all different connected components, naturally avoiding the learning bias on large connected components. Extensive experiments on various datasets demonstrate that our proposed model achieves superior performance against other state-of-the-art methods. Gengyu Lyu, Yuena Lin, Zhen Yang 0004, Zun Li 0001 |
AAAI | 7 |
| 2025 | VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity RecognitionabstractMost of existing weakly supervised GAR methods are typically bottom-up, automatically mining key areas by the attention mechanism. Due to the lack of a semantic connection to individual actions, some regions associated with these actions may be omitted, potentially impacting performance. In fact, a group activity is a combination of multiple individual actions, and the prototype of a specific action can be obtained from visual representations of individuals performing it, denoted as visual conceptual knowledge. In this paper, we propose a Visual Conceptual Knowledge Guided Action Map framework. It uses prototypes to produce individual action maps that indicate the likelihood of actions occurring at different locations. In some scenarios, the spatial distribution of actions shows strong regularity, which we compile as A-A Maps to enhance individual action maps. The action maps are integrated with action semantic representations for group activity recognition. Extensive experiments on two public benchmarks, the Volleyball and the NBA datasets, demonstrate the effectiveness of our proposed method, even in cases of limited training data. Zhuming Wang, Yihao Zheng 0002, Jiarui Li 0002, Yaofei Wu, Yan Huang 0008, Zun Li 0001, Lifang Wu, Liang Wang 0001 |
ACM Multimedia | 6 |
| 2025 | Statistical Information Assisted Interaction Reasoning for skeleton-only group activity recognition
Zhuming Wang, Zun Li 0001, Yihao Zheng 0002, Lifang Wu |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Multi-scale motion-based relational reasoning for group activity recognition
Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Zun Li 0001, Ye Xiang |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Dual Consistency Regularization for Generalized Face Anti-SpoofingabstractRecent Face Anti-Spoofing (FAS) methods have improved generalization to unseen domains by leveraging domain generalization techniques. However, they overlooked the semantic relationships between local features, resulting in suboptimal feature alignment and limited performance. To this end, pixel-wise supervision has been introduced to offer contextual guidance for better feature alignment. Unfortunately, the semantic ambiguity in coarsely designed pixel-wise supervision often leads to misalignment. This paper proposes a novel Dual Consistency Regularization Network (DCRN). It promotes the fine-grained alignment of local features with dense semantic correspondence for FAS. Specifically, a Dual Consistency Learning module (DCL) is devised to capture the inter- and intra-similarity between each region of sample pairs. In this module, a dual consistency regularization learning objective enhances the semantic consistency of local features by minimizing both the variance of inter-similarity and the distance between inter- and intra-similarity. Further, a weight matrix is estimated based on the inter-similarity, representing the possibility that each region belongs to the living class. Based on this weight matrix, WMSE loss is designed to guide the model in avoiding mapping the live regions to the spoofing class, thus alleviating semantic ambiguity in pixel-wise supervision. Extensive experiments on four widely used datasets clearly demonstrate the superiority and high generalization of the proposed DCRN. Yongluo Liu, Zun Li 0001, Lifang Wu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Face Anti-Spoofing via Interaction Learning with Face Image Quality AlignmentabstractFace Anti-Spoofing is critical to secure face recognition systems from presentation attacks. Existing methods often suffer from performance degradation due to image quality issues, such as blurring, overexposure, or varied background, which cause distribution deviations of face images in the quality space, and hinder the learning of effective liveness features. In this paper, we propose a novel method that interactively co-reinforces the liveness and Face Quality representations for Face Anti-Spoofing (FQ-FAS). Specifically, to enhance the discrimination of face quality representation, FQ-FAS first designs a face quality learning module that naturally mitigates the interference from background. Subsequently, a quality-spoofing feature interaction module is devised to co-reinforce both liveness and face quality representations. Meanwhile, we propose a quality aware triplet loss to align the distribution of face images from two aspects: one is to pull the homogeneous face images with different quality together, while the other is to push the inhomogeneous samples with similar quality away in the feature space. In this way, FQ-FAS can learn reliable and discriminative representations for face anti-spoofing. Extensive intra-dataset and cross-dataset experiments clearly demonstrate that our method obtains better performance than previous state-of-the-art methods. Yongluo Liu, Zun Li 0001, Zhuming Wang, Lifang Wu |
FG | 2 |
| 2024 | SkatingVerse: A large-scale benchmark for comprehensive evaluation on human action understandingabstractAbstract Human action understanding (HAU) is a broad topic that involves specific tasks, such as action localisation, recognition, and assessment. However, most popular HAU datasets are bound to one task based on particular actions. Combining different but relevant HAU tasks to establish a unified action understanding system is challenging due to the disparate actions across datasets. A large‐scale and comprehensive benchmark, namely SkatingVerse is constructed for action recognition, segmentation, proposal, and assessment. SkatingVerse focus on fine‐grained sport action, hence figure skating is chosen as the task object, which eliminates the biases of the object, scene, and space that exist in most previous datasets. In addition, skating actions have inherent complexity and similarity, which is an enormous challenge for current algorithms. A total of 1687 official figure skating competition videos was collected with a total of 184.4 h, exceeding four times over other datasets with a similar topic. SkatingVerse enables to formulate a unified task to output fine‐grained human action classification and assessment results from a raw figure skating competition video. In addition, SkatingVerse can facilitate the study of HAU foundation model due to its large scale and abundant categories. Moreover, image modality is incorporated for human pose estimation task into SkatingVerse . Extensive experimental results show that (1) SkatingVerse significantly helps the training and evaluation of HAU methods, (2) the performance of existing HAU methods has much room to improve, and SkatingVerse helps to reduce such gaps, and (3) unifying relevant tasks in HAU through a uniform dataset can facilitate more practical applications. SkatingVerse will be publicly available to facilitate further studies on relevant problems. Ziliang Gan, Lei Jin 0003, Yu Cheng 0009, Yinglei Teng, Zun Li 0001, Yawen Li 0001, Wenhan Yang, Junliang Xing, Jian Zhao 0006 |
IET Comput. Vis. | 6 |
| 2024 | Quality-Invariant Domain Generalization for Face Anti-Spoofing
Yongluo Liu, Zun Li 0001, Yaowen Xu, Zhizhi Guo, Zhaofan Zou, Lifang Wu |
Int. J. Comput. Vis. | 2 |
| 2024 | Light dual hypergraph convolution for collaborative filtering
Meng Jian, Langchen Lang, Zun Li 0001, Tuo Wang 0001, Lifang Wu |
Pattern Recognit. | 4 |
| 2024 | GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng 0002, Kunming Luo, Shuaicheng Liu, Zun Li 0001, Ye Xiang, Lifang Wu, Bing Zeng 0001, Chang Wen Chen |
Pattern Recognit. Lett. | 4 |
| 2024 | Knowledge Augmented Relation Inference for Group Activity RecognitionabstractGroup activity recognition is a challenging task because it involves diverse individual actions and complex relations. Most existing methods enhance individual representation by introducing relation inference using appearance features. Some methods utilize extra knowledge, such as action labels, to enhance relation inference and refine the individual representation, but the knowledge they explored is simple and insufficient. In this paper, we propose a novel idea of knowledge concretization and further develop a Knowledge Augmented Relation Inference framework (KARI) for group activity recognition. Specifically, we first concretize knowledge from training data, and then represent them as Class-Class co-occurrence Map (C-C Map) and Class-Position distribution Map (C-P Map). On top of them, KARI explores concretized knowledge to integrate visual and semantic representation in a unified architecture for group activity recognition. Experimental results on two public datasets show that the proposed framework performs favorably compared with state-of-the-art approaches. Zhuming Wang, Zun Li 0001, Xianglong Lang, Yihao Zheng 0002, Lifang Wu, Liang Wang 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Modality Meets Long-Term Tracker: A Siamese Dual Fusion Framework for Tracking UAVabstractTracking an Unmanned Aerial Vehicle (UAV) to obtain its locations and trajectory is a crucial task to avoid the unlawful use of UAVs. However, most existing UAV tracking methods fail when facing cluster environments, out-of-view, and occlusions because of their insufficient representation of global context information capacity. To mitigate these issues, we propose a new tracker, namely SiamFusion, to innovate a dual fusion procedure that leverages the advantages in both the feature and decision levels. In particular, we propose a novel feature fusion module named Modality-Fusion to utilize multi-modal information, enhancing the perception of the target. From the decision level, we further develop a local-global converter based on a multi-modal fusion decision-making mechanism to reduce the accumulation during tracking, which significantly increases the robustness of the tracking process. Extensive experiments demonstrate the superiority of the proposed SiamFusion, which achieves the best performance on Anti-UAV in terms of accuracy and speed. In particular, we exceed the state-of-the-art tracking algorithm in the tracking accuracy by 4.2% at a similar frame rate. Our source codes, pre-trained models, and online demos will be released upon acceptance. Lei Jin 0003, Shengjie Li 0003, Jianqiang Xia, Jun Wang 0041, Zun Li 0001, Wenhan Yang, Pengfei Zhang 0016, Jian Zhao 0006, Bo Zhang 0007 |
ICIP | 6 |
| 2023 | Depth guided feature selection for RGBD salient object detection
Zun Li 0001, Congyan Lang, Guanqin Li, Tao Wang 0011, Yidong Li |
Neurocomputing | 1 |
| 2023 | Dual-stream correlation exploration for face anti-Spoofing
Yongluo Liu, Lifang Wu, Zun Li 0001, Zhuming Wang |
Pattern Recognit. Lett. | 3 |
| 2023 | Active Spatial Positions Based Hierarchical Relation Inference for Group Activity RecognitionabstractGroup activity recognition aims to recognize behaviors characterized by multiple individuals within a scene. Existing schemes rely on individual relation inference and usually take the individuals as tokens. Essentially they select the most relevant region of the group activity from the entire image while filtering out irrelevant background noises. However, these schemes require individual bounding box labeling in both training and testing stages. Since individuals have usually been presented at one scale, multi-scale individuals cannot be combined in an effective way. In this paper, we present a novel end-to-end hierarchical relation inference framework based on active spatial positions for group activity recognition. This framework is designed to locate active spatial positions and use them as visual tokens to infer the relations for token embeddings. It requires individual bounding box labeling only in the training stage while automatically eliminating the background after locating active spatial positions from the entire scene. The hierarchical relations can be naturally inferred based on the visual tokens at different scales, contributing to further performance improvement. Experimental results demonstrate that the proposed framework is competitive against existing schemes that require more laboring and computation to generate labels in both the training and testing stage. Lifang Wu, Xianglong Lang, Ye Xiang, Chang Wen Chen, Zun Li 0001, Zhuming Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Improving Face Anti-spoofing via Advanced Multi-perspective Feature LearningabstractFace anti-spoofing (FAS) plays a vital role in securing face recognition systems. Previous approaches usually learn spoofing features from a single perspective, in which only universal cues shared by all attack types are explored. However, such single-perspective-based approaches ignore the differences among various attacks and commonness between certain attacks and bona fides, thus tending to neglect some non-universal cues that contain strong discernibility against certain types. As a result, when dealing with multiple types of attacks, the above approaches may suffer from the uncomprehensive representation of bona fides and spoof faces. In this work, we propose a novel Advanced Multi-Perspective Feature Learning network (AMPFL), in which multiple perspectives are adopted to learn discriminative features, to improve the performance of FAS. Specifically, the proposed network first learns universal cues and several perspective-specific cues from multiple perspectives, then aggregates the above features and further enhances them to perform face anti-spoofing. In this way, AMPFL obtains features that are difficult to be captured by single-perspective-based methods and provides more comprehensive information on bona fides and spoof faces, thus achieving better performance for FAS. Experimental results show that our AMPFL achieves promising results in public databases, and it effectively solves the issues of single-perspective-based approaches. Zhuming Wang, Yaowen Xu, Lifang Wu, Hu Han 0001, Zun Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Multi-Sequence Dilated Network for Object DetectionabstractScale variation is one of the key challenges in the object detection. Most previous object detectors remedy this by using dilated convolution to enlarge the receptive fields of the vanilla convolutional layers. However, these methods focus on either the spatial information of small objects or the semantics of middle and large objects, which still fail to effectively adapt the scale variance of different objects, resulting in a sub-optimal performance for the object detection. In this paper, we propose a novel Multi-Sequence Dilated Network (MSDN) that stacks different dilated convolutions with different orders in parallel for improving the performance of the object detection. Concretely, MSDN contains a sequential dilated module and a dilated attention module. The former aims to generate scale-specific feature maps with fine-spatial and semantic information of objects at different scales, while the latter further selects more powerful information to adaptively enlarge the receptive fields of object features at different scales. Facilitated with these modules, MSDN well obtains the fine-spatial and semantic information of objects at different scales, thus solving the problem of the scale variation. Comprehensive experimental results over two public object detection benchmarks clearly demonstrate the effectiveness of our proposed MSDN. Particularly, on the COCO dataset, the mAP value of MSDN is 48.7%, outperforming existing state-of-the-art methods in a single model manner. Zun Li 0001, Chang Xin, Lifang Wu, Yongluo Liu |
MMSP | 2 |
| 2022 | Pedestrian attribute recognition based on attribute correlation
Ruijie Zhao 0007, Congyan Lang, Zun Li 0001, Liqian Liang, Songhe Feng, Tao Wang 0011 |
Multim. Syst. | 3 |
| 2022 | Dense Attentive Feature Enhancement for Salient Object DetectionabstractAttention mechanisms have been proven highly effective for salient object detection. Most previous works utilize attention as a self-gated module to reweigh the feature maps at different levels independently. However, they are limited to certain-level guidance and could not satisfy the need of both accurately detecting intact objects and maintaining their detailed boundaries. In this paper, we build dense attention upon features from multiple levels simultaneously and propose a novel Dense Attentive Feature Enhancement (DAFE) module for efficient feature enhancement in saliency detection. DAFE stacks several attentional units and densely connects attentive feature output from current unit to its all subsequent units. This allows feature maps at deep units to absorb attentive information from shallow units, thus more discriminative information can be efficiently selected at the final output. Note that DAFE is plug and play, which can be effortlessly inserted into any saliency or video saliency models for their performance improvements. We further instantiate a highly effective Dense Attentive Feature Enhancement Network (DAFE-Net) for accurate salient object detection. DAFE-Net constructs DAFE over the aggregation feature that contains both semantics and saliency details, the entire salient objects and their boundaries can be well retained through dense attentions. Extensive experiments demonstrate that the proposed DAFE module is highly effective, and the DAFE-Net performs favorably compared with state-of-the-art approaches. Zun Li 0001, Congyan Lang, Liqian Liang, Jian Zhao 0006, Songhe Feng, Qibin Hou, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Seeing Crucial Parts: Vehicle Model Verification via a Discriminative Representation ModelabstractWidely used surveillance cameras have promoted large amounts of street scene data, which contains one important but long-neglected object: the vehicle. Here we focus on the challenging problem of vehicle model verification. Most previous works usually employ global features (e.g., fully connected features) to further perform vehicle-level deep metric learning (e.g., triplet-based network). However, we argue that it is noteworthy to investigate the distinctiveness of local features and consider vehicle-part-level metric learning by reducing the intra-class variance as much as possible. In this article, we introduce a simple yet powerful deep model—the enforced intra-class alignment network (EIA-Net)—which can learn a more discriminative image representation by localizing key vehicle parts and jointly incorporating two distance metrics: vehicle-level embedding and vehicle-part-sensitive embedding. For learning features, we propose an effective feature extraction module that is composed of two components: the regional proposal network (RPN)-based network and part-based CNN. The RPN is used to define key vehicle regions and aggregate local features on these regions, whereas part-based CNN offers supplementary global features for the RPN-based network. The fusion features learned by feature extraction module are cast into the deep metric learning module. Especially, we derived an enforced intra-class alignment loss by re-utilizing key vehicle part information to enhance reducing intra-class variance. Furthermore, we modify the coupled cluster loss to model the vehicle-level embedding by enlarging the inter-class variance while shortening intra-class variance. Extensive experiments over benchmark datasets VehicleID and CompCars have shown that the proposed EIA-Net significantly outperforms the state-of-the-art approaches for vehicle model verification. Furthermore, we also conduct comprehensive experiments on vehicle re-identification datasets (i.e., VehicleID and VeRi776) to validate the generalization ability effectiveness of our proposed method. Liqian Liang, Congyan Lang, Zun Li 0001, Jian Zhao 0006, Tao Wang 0011, Songhe Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Deep spatio-frequency saliency detection
Zun Li 0001, Congyan Lang, Tao Wang 0011, Yidong Li, Jiashi Feng |
Neurocomputing | 1 |
| 2021 | Cross-Layer Feature Pyramid Network for Salient Object DetectionabstractFeature pyramid network (FPN) based models, which fuse the semantics and salient details in a progressive manner, have been proven highly effective in salient object detection. However, it is observed that these models often generate saliency maps with incomplete object structures or unclear object boundaries, due to the indirect information propagation among distant layers that makes such fusion structure less effective. In this work, we propose a novel Cross-layer Feature Pyramid Network (CFPN), in which direct cross-layer communication is enabled to improve the progressive fusion in salient object detection. Specifically, the proposed network first aggregates multi-scale features from different layers into feature maps that have access to both the high- and low- level information. Then, it distributes the aggregated features to all the involved layers to gain access to richer context. In this way, the distributed features per layer own both semantics and salient details from all other layers simultaneously, and suffer reduced loss of important information during the progressive feature fusion. At last, CFPN fuses the distributed features of each layer stage-by-stage. This way, the high-level features that contain context useful for locating complete objects are preserved until the final output layer, and the low-level features that contain spatial structure details are embedded into each layer to preserve spatial structural details. Extensive experimental results over six widely used salient object detection benchmarks and with three popular backbones clearly demonstrate that CFPN can accurately locate fairly complete salient regions and effectively segment the object boundaries. Zun Li 0001, Congyan Lang, Jun Hao Liew, Yidong Li, Qibin Hou, Jiashi Feng |
IEEE Trans. Image Process. | 1 |
| 2019 | Recurrent convolutional network for video-based smoke detection
Mengxia Yin, Congyan Lang, Zun Li 0001, Songhe Feng, Tao Wang 0011 |
Multim. Tools Appl. | 3 |
| 2019 | Co-saliency Detection with Graph MatchingabstractRecently, co-saliency detection, which aims to automatically discover common and salient objects appeared in several relevant images, has attracted increased interest in the computer vision community. In this article, we present a novel graph-matching based model for co-saliency detection in image pairs. A solution of graph matching is proposed to integrate the visual appearance, saliency coherence, and spatial structural continuity for detecting co-saliency collaboratively. Since the saliency and the visual similarity have been seamlessly integrated, such a joint inference schema is able to produce more accurate and reliable results. More concretely, the proposed model first computes the intra-saliency for each image by aggregating multiple saliency cues. The common and salient regions across multiple images are thus discovered via a graph matching procedure. Then, a graph reconstruction scheme is proposed to refine the intra-saliency iteratively. Compared to existing co-saliency detection methods that only utilize visual appearance cues, our proposed model can effectively exploit both visual appearance and structure information to better guide co-saliency detection. Extensive experiments on several challenging image pair databases demonstrate that our model outperforms state-of-the-art baselines significantly. Zun Li 0001, Congyan Lang, Jiashi Feng, Yidong Li, Tao Wang 0011, Songhe Feng |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2018 | Saliency ranker: A new salient object detection method
Zun Li 0001, Congyan Lang, Songhe Feng, Tao Wang 0011 |
J. Vis. Commun. Image Represent. | 1 |