VLDB 2026 Research / reviewers in the wild / expert
Yukun Qiu
dblp:269/4613
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0001-9490-5502ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Patch-Based Privacy Attention for Weakly-Supervised Privacy-Preserving Action RecognitionabstractPrivacy-preserving action recognition aims to prevent privacy leakage by learning to anonymize video frames for action recognition. However, apart from video-level action labels, supervised methods require costly frame-level privacy labels. To achieve privacy-preserving action recognition in the absence of privacy labels, weakly-supervised privacy-preserving action recognition proposes to merely utilize action labels for learning without using any privacy labels during training, and therefore the challenge lies in removing the privacy information without privacy annotations. Inspired by the fact that private information such as the identity of the participant is not significant to action recognition, our main idea is to utilize the attention mechanism to automatically discover the action-sensitive information in frames and remove the other information to prevent privacy leakage without using the privacy labels. Our method contains a novel patch-based privacy attention module and an action recognition module. The patch-based privacy attention module splits raw frames into patches and exploits self-attention among patches to adaptively discover action-sensitive but privacy-less information. Our patch-based privacy attention module mines the action-sensitive information from both individual frames and adjacent frames to generate anonymized frames. A distance correlation loss is introduced to enforce that the generated anonymized frames are distinguished from the original frames and contain less private information. In addition, the action recognition module learns to recognize actions based on anonymized frames. Extensive experiments demonstrate that our model can effectively alleviate privacy leakage and maintain the performance of action recognition without using privacy labels. Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Wei-Shi Zheng 0001 |
FG | 2 |
| 2024 | Privacy-Preserving Action Recognition: A Survey
Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001 |
PRCV (7) | 2 |
| 2024 | TwinFormer: Fine-to-Coarse Temporal Modeling for Long-Term Action RecognitionabstractThe long-term action in untrimmed video generally contains multiple sub-actions, among which various semantic patterns exist (e.g., the co-occurrence or sequentiality between sub-actions). These semantic patterns are temporally coarse, and correlated with multiple local contexts which encode the local temporal evolution of visual elements (e.g., hands, objects) in videos. The local contexts and semantic patterns form the inherent fine-to-coarse temporal structure of long-term actions, which is neglected by existing works. Accordingly, in this work we propose TwinFormer, which exploits a novel fine-to-coarse temporal modeling manner to uncover the temporal structure of long-term actions. The proposed TwinFormer consists of a pair of twin encoders with the same structural design, namely Localcontext Encoder and Semantic-pattern Encoder, and a Temporalbridged Attention to bridge the two twin encoders. The Localcontext Encoder aims to model the local contexts in the longterm action. And the Temporal-bridged Attention is designed to correlate the local contexts with semantic patterns. Furthermore, the Semantic-pattern Encoder reveals the temporal evolution of semantic patterns. Experimental results on three benchmarks demonstrate the effectiveness of the proposed model. Kun-Yu Lin, Yukun Qiu, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Inner-Outer Aware Reconstruction Model for Monocular 3D Scene ReconstructionabstractMonocular 3D scene reconstruction aims to reconstruct the 3D structure of scenes based on posed images. Recent volumetric-based methods directly predict the truncated signed distance function (TSDF) volume and have achieved promising results. The memory cost of volumetric-based methods will grow cubically as the volume size increases, so a coarse-to-fine strategy is necessary for saving memory. Specifically, the coarse-to-fine strategy distinguishes surface voxels from non-surface voxels, and only potential surface voxels are considered in the succeeding procedure. However, the non-surface voxels have various features, and in particular, the voxels on the inner side of the surface are quite different from those on the outer side since there exists an intrinsic gap between them. Therefore, grouping inner-surface and outer-surface voxels into the same class will force the classifier to spend its capacity to bridge the gap. By contrast, it is relatively easy for the classifier to distinguish inner-surface and outer-surface voxels due to the intrinsic gap. Inspired by this, we propose the inner-outer aware reconstruction (IOAR) model. IOAR explores a new coarse-to-fine strategy to classify outer-surface, inner-surface and surface voxels. In addition, IOAR separates occupancy branches from TSDF branches to avoid mutual interference between them. Since our model can better classify the surface, outer-surface and inner-surface voxels, it can predict more precise meshes than existing methods. Experiment results on ScanNet, ICL-NUIM and TUM-RGBD datasets demonstrate the effectiveness and generalization of our model. The code is available at https://github.com/YorkQiu/InnerOuterAwareReconstruction. Yukun Qiu, Guo-Hao Xu, Wei-Shi Zheng 0001 |
NeurIPS | 1 |
| 2023 | Learning Relation Models to Detect Important People in Still ImagesabstractImportant people detection aims to identify the most important people (i.e., the people who play the main roles in scenes) in images, which is challenging since people's importance in images depends not only on their appearance but also on their interactions with others (i.e., relations among people) and their roles in the scene (i.e., relations between people and underlying events). In this work, we propose the People Relation Network (PRN) to solve this problem. PRN consists of three modules (i.e., the feature representation, relation and classification modules) to extract visual features, model relations and estimate people's importance, respectively. The relation module contains two submodules to model two types of relations, namely, the person-person relation submodule and the person-event relation submodule. The person-person relation submodule infers the relations among people from the interaction graph and the person-event relation submodule models the relations between people and events by considering the spatial correspondence between features. With the help of them, PRN can effectively distinguish important people from other individuals. Extensive experiments on the Multi-Scene Important People (MS) and NCAA Basketball Image (NCAA) datasets show that PRN achieves state-of-the-art performance and generalizes well when available data is limited. Yukun Qiu, Fa-Ting Hong, Wei-Hong Li 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Adversarial Partial Domain Adaptation by Cycle Inconsistency
Kun-Yu Lin, Yukun Qiu, Wei-Shi Zheng 0001 |
ECCV (33) | 3 |
| 2020 | Neighbor Combinatorial Attention for Critical Structure MiningabstractGraph convolutional networks (GCNs) have been widely used to process graph-structured data. However, existing GNN methods do not explicitly extract critical structures, which reflect the intrinsic property of a graph. In this work, we propose a novel GCN module named Neighbor Combinatorial ATtention (NCAT) to find critical structure in graph-structured data. NCAT attempts to match combinatorial neighbors with learnable patterns and assigns different weights to each combination based on the matching degree between the patterns and combinations. By stacking several NCAT modules, we can extract hierarchical structures that is helpful for down-stream tasks. Our experimental results show that NCAT achieves state-of-the-art performance on several benchmark graph classification datasets. In addition, we interpret what kind of features our model learned by visualizing the extracted critical structures. Tanli Zuo, Yukun Qiu, Wei-Shi Zheng 0001 |
IJCAI | 2 |