EDBT 2026 Demo / reviewers in the wild / expert
Xumin Yu
dblp:237/0070
· DBLP profile ↗
21ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point2Seq: Quantized Serialization Encoding for Object Point Cloud Pretraining
Xumin Yu, Zuyan Liu, Jie Zhou 0001, Jiwen Lu |
Int. J. Comput. Vis. | 1 |
| 2026 | ProtoComp++: Diverse Point Cloud Completion With Controllable PrototypeabstractPoint cloud completion aims to reconstruct the geometry of partial point clouds captured by various sensors. Traditionally, point cloud models are trained on synthetic datasets that feature limited categories and differ significantly from real-world scenarios. This gap often causes existing methods to struggle when faced with unfamiliar categories and severe incompleteness in real-world applications. In this paper, we propose PrototypeCompletion, a novel prototype-based approach for point cloud completion. The method begins by generating rough prototypes, which are then refined with additional geometric details to make the final prediction. We introduce two distinct approaches for integrating prototypes into the network: explicit prototypes and implicit prototypes. Our approach demonstrates strong generalization capabilities, allowing it to handle point cloud completion for a variety of unseen categories beyond the training data. We demonstrate that incorporating language prompts into the training of point cloud completion models significantly expands their applicability and enhances their performance in diverse point cloud completion tasks. Furthermore, we propose a new evaluation metric and a test benchmark based on ScanNet200 and KITTI, designed to assess the model's performance in real-world scenarios and foster future research in the field. Experimental results show that our method outperforms state-of-the-art models on the existing PCN and ShapeNet34 benchmarks and also excels in various real-world settings, handling different object categories and sensor types effectively. The code will be made publicly available. Xumin Yu, Zuyan Liu, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | FlowTurbo: Accelerating Flow-Based Image Generation Models via Multi-Stage RefinementabstractBuilding on the. success of diffusion models in visual generation, flow-based models reemerge as another prominent family of generative models that have achieved competitive or better performance in terms of both visual quality and inference speed. By learning the velocity field through flow-matching, flow-based models tend to produce a straighter sampling trajectory, which is advantageous during the sampling process. However, unlike diffusion models for which fast samplers are well-developed, efficient sampling of flow-based generative models has been rarely explored. In this paper, we propose a framework called FlowTurbo to accelerate the sampling of flow-based models while still enhancing the sampling quality. Our primary observation is that the velocity predictor's outputs in the flow-based models will become stable during the sampling, enabling the estimation of velocity via a lightweight velocity refiner. Additionally, we introduce several techniques including a pseudo corrector and sample-aware compilation to further reduce inference time. Since FlowTurbo does not change the multi-step sampling paradigm, it can be effectively applied for various tasks such as image editing, inpainting, etc. Besides, we propose a new multi-stage refinement technique that is designed to reduce the inference costs with large flow-based image generation models. Specifically, the multi-stage refinement split the whole generation procedure on different resolutions, forming a coarse-to-fine text-to-image pipeline. We further adopt a stage-aware deployment strategy that can maximize the inference speed in terms of both latency and throughput. By integrating FlowTurbo into different flow-based models, we obtain an acceleration ratio of 53.1%$\sim$∼58.3% on class-conditional generation and 29.8%$\sim$∼38.5% on text-to-image generation. Notably, FlowTurbo reaches an FID of 2.12 on ImageNet with 100 (ms/img) and FID of 3.93 with 38 (ms/img), achieving the real-time image generation and establishing the new state-of-the-art. Equipped with the recent SD 3.5 Large, we achieved FID of 28.05 with a speed improvement of around 50% on NVIDIA 3090 GPU. Wenliang Zhao, Minglei Shi, Xumin Yu, Zengyi Qin, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | LLM-assisted Deep Reinforcement Learning for Satellite Range SchedulingabstractSatellite networks have emerged as a critical infrastructure in communication systems, providing wide coverage and global connectivity. However, the growing scale and cross-domain nature of the satellite range scheduling (SRS) problem make efficient satellite and ground station resources scheduling ever more challenging, especially for emergency tasks. Conventional scheduling based on deep reinforcement learning (DRL) suffers from sparse rewards in dynamic cross-domain satellite network, limiting both training speed and solution quality. In this study, we propose RAPPO, a Retrieval-Augmented Proximal Policy Optimization framework that leverages retrieval-augmented generation (RAG) to feed external SRS knowledge into a large language model (LLM) for iterative reward function generation. RAPPO comprises two modules: (1) LLM prompt initialization, which retrieves knowledge of SRS via RAG; and (2) reward function optimization, where the LLM generates and continuously optimizes the DRL reward function based on scene descriptions and reward-shaping knowledge retrieved during optimization. In simulation experiments, RAPPO achieves an approximate 48.23% improvement in convergence speed, with load-balance and transmission timeliness enhanced by about 30.83% and 31.04%, respectively, compared to conventional PPO. These results demonstrate that integrating RAG-enhanced LLM reasoning into reward design substantially improves DRL training efficiency and scheduling effectiveness for cross-domain SRS. Mingying Yang, Xumin Yu |
GLOBECOM | 5 |
| 2025 | Vision Generalist Model: A Survey
Ziyi Wang 0007, Yongming Rao, Shuofeng Sun, Xinrun Liu, Yi Wei 0003, Xumin Yu, Zuyan Liu, Hongmin Liu 0001, Jie Zhou 0001, Jiwen Lu |
Int. J. Comput. Vis. | 6 |
| 2024 | ProtoComp: Diverse Point Cloud Completion with Controllable Prototype
Xumin Yu, Jie Zhou 0001, Jiwen Lu |
ECCV (50) | 1 |
| 2024 | XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic SegmentationabstractExisting methodologies in open vocabulary 3D semantic segmentation primarily concentrate on establishing a unified feature space encompassing 3D, 2D, and textual modalities. Nevertheless, traditional techniques such as global feature alignment or vision-language model distillation tend to impose only approximate correspondence, struggling notably with delineating fine-grained segmentation boundaries. To address this gap, we propose a more meticulous mask-level alignment between 3D features and the 2D-text embedding space through a cross-modal mask reasoning framework, XMask3D. In our approach, we developed a mask generator based on the denoising UNet from a pre-trained diffusion model, leveraging its capability for precise textual control over dense pixel representations and enhancing the open-world adaptability of the generated masks. We further integrate 3D global features as implicit conditions into the pre-trained 2D denoising UNet, enabling the generation of segmentation masks with additional 3D geometry awareness. Subsequently, the generated 2D masks are employed to align mask-level 3D representations with the vision-language feature space, thereby augmenting the open vocabulary capability of 3D geometry embeddings. Finally, we fuse complementary 2D and 3D mask features, resulting in competitive performance across multiple benchmarks for 3D open vocabulary semantic segmentation. Code is available at https://github.com/wangzy22/XMask3D. Ziyi Wang 0007, Xumin Yu, Jie Zhou 0001, Jiwen Lu |
NeurIPS | 3 |
| 2024 | FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity RefinerabstractBuilding on the success of diffusion models in visual generation, flow-based models reemerge as another prominent family of generative models that have achieved competitive or better performance in terms of both visual quality and inference speed. By learning the velocity field through flow-matching, flow-based models tend to produce a straighter sampling trajectory, which is advantageous during the sampling process. However, unlike diffusion models for which fast samplers are well-developed, efficient sampling of flow-based generative models has been rarely explored. In this paper, we propose a framework called FlowTurbo to accelerate the sampling of flow-based models while still enhancing the sampling quality. Our primary observation is that the velocity predictor's outputs in the flow-based models will become stable during the sampling, enabling the estimation of velocity via a lightweight velocity refiner. Additionally, we introduce several techniques including a pseudo corrector and sample-aware compilation to further reduce inference time. Since FlowTurbo does not change the multi-step sampling paradigm, it can be effectively applied for various tasks such as image editing, inpainting, etc. By integrating FlowTurbo into different flow-based models, we obtain an acceleration ratio of 53.1\%$\sim$58.3\% on class-conditional generation and 29.8\%$\sim$38.5\% on text-to-image generation. Notably, FlowTurbo reaches an FID of 2.12 on ImageNet with 100 (ms / img) and FID of 3.93 with 38 (ms / img), achieving the real-time image generation and establishing the new state-of-the-art. Code is available at https://github.com/shiml20/FlowTurbo. Wenliang Zhao, Minglei Shi, Xumin Yu, Jie Zhou 0001, Jiwen Lu |
NeurIPS | 3 |
| 2024 | Point-to-Pixel Prompting for Point Cloud Analysis With Pre-Trained Image ModelsabstractNowadays, pre-training big models on large-scale datasets has achieved great success and dominated many downstream tasks in natural language processing and 2D vision, while pre-training in 3D vision is still under development. In this paper, we provide a new perspective of transferring the pre-trained knowledge from 2D domain to 3D domain with Point-to-Pixel Prompting in data space and Pixel-to-Point distillation in feature space, exploiting shared knowledge in images and point clouds that display the same visual world. Following the principle of prompting engineering, Point-to-Pixel Prompting transforms point clouds into colorful images with geometry-preserved projection and geometry-aware coloring. Then the pre-trained image models can be directly implemented for point cloud tasks without structural changes or weight modifications. With projection correspondence in feature space, Pixel-to-Point distillation further regards pre-trained image models as the teacher model and distills pre-trained 2D knowledge to student point cloud models, remarkably enhancing inference efficiency and model capacity for point cloud analysis. We conduct extensive experiments in both object classification and scene segmentation under various settings to demonstrate the superiority of our method. In object classification, we reveal the important scale-up trend of Point-to-Pixel Prompting and attain 90.3% accuracy on ScanObjectNN dataset, surpassing previous literature by a large margin. In scene-level semantic segmentation, our method outperforms traditional 3D analysis approaches and shows competitive capacity in dense prediction tasks. Ziyi Wang 0007, Yongming Rao, Xumin Yu, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud ModelsabstractWith the overwhelming trend of mask image modeling led by MAE, generative pre-training has shown a remarkable potential to boost the performance of fundamental models in 2D vision. However, in 3D vision, the over-reliance on Transformer-based backbones and the unordered nature of point clouds have restricted the further development of generative pre-training. In this paper, we propose a novel 3D-to-2D generative pre-training method that is adaptable to any point cloud model. We propose to generate view images from different instructed poses via the cross-attention mechanism as the pre-training scheme. Generating view images has more precise supervision than its point cloud counterpart, thus assisting 3D backbones to have a finer comprehension of the geometrical structure and stereoscopic relations of the point cloud. Experimental results have proved the superiority of our proposed 3D-to-2D generative pre-training over previous pre-training methods. Our method is also effective in boosting the performance of architecture-oriented approaches, achieving state-of-the-art performance when fine-tuning on ScanObjectNN classification and ShapeNet-Part segmentation tasks. Code is available at https://github.com/wangzy22/TakeAPhoto. Ziyi Wang 0007, Xumin Yu, Yongming Rao, Jie Zhou 0001, Jiwen Lu |
ICCV | 2 |
| 2023 | Robust Signature-Based Hyperspectral Target Detection Using Dual NetworksabstractThe training of deep networks for hyperspectral target detection (HTD) is usually confronted with the problem of limited samples and in extreme cases, there might be only one target sample available. To address this challenge, we propose a novel approach with dual networks in this letter. First, a training set that is not fully accurate but representative enough regarding both targets and backgrounds is built through predetection and clustering. Then, two types of neural networks, that is, one generative adversarial network (GAN) and one convolutional neural network (CNN), which focus on spectral and spatial features of hyperspectral images (HSIs), are utilized for target detection. After that, the results of the two networks are fused, with the final detection result obtained. Experiments on real HSIs indicate that the proposed approach manages to perform HTD with only one target sample and is able to yield a more robust detection performance compared to other approaches. Yanlong Gao, Yan Feng 0005, Xumin Yu, Shaohui Mei |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | AdaPoinTr: Diverse Point Cloud Completion With Adaptive Geometry-Aware TransformersabstractIn this paper, we propose a Transformer encoder-decoder architecture, called PoinTr, which reformulates point cloud completion as a set-to-set translation problem and employs a geometry-aware block to model local geometric relationships explicitly. The migration of Transformers enables our model to better learn structural knowledge and preserve detailed information for point cloud completion. Taking a step towards more complicated and diverse situations, we further propose AdaPoinTr by developing an adaptive query generation mechanism and designing a novel denoising task during completing a point cloud. Coupling these two techniques enables us to train the model efficiently and effectively: we reduce training time (by 15x or more) and improve completion performance (over 20%). Additionally, we propose two more challenging benchmarks with more diverse incomplete point clouds that can better reflect real-world scenarios to promote future research. We also show our method can be extended to the scene-level point cloud completion scenario by designing a new geometry-enhanced semantic scene completion framework. Extensive experiments on the existing and newly-proposed datasets demonstrate the effectiveness of our method, which attains 6.53 CD on PCN, 0.81 CD on ShapeNet-55 and 0.392 MMD on real-world KITTI, surpassing other work by a large margin and establishing new state-of-the-arts on various benchmarks. Most notably, AdaPoinTr can achieve such promising performance with higher throughputs and fewer FLOPs compared with the previous best methods in practice. Xumin Yu, Yongming Rao, Ziyi Wang 0007, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | SemAffiNet: Semantic-Affine Transformation for Point Cloud SegmentationabstractConventional point cloud semantic segmentation methods usually employ an encoder-decoder architecture, where mid-level features are locally aggregated to extract geometric information. However, the over-reliance on these class-agnostic local geometric representations may raise confusion between local parts from different categories that are similar in appearance or spatially adjacent. To address this issue, we argue that mid-level features can be further enhanced with semantic information, and propose semantic-affine transformation that transforms features of mid-level points belonging to different categories with class-specific affine parameters. Based on this technique, we propose SemAffiNet for point cloud semantic segmentation, which utilizes the attention mechanism in the Transformer module to implicitly and explicitly capture global structural knowledge within local parts for overall comprehension of each category. We conduct extensive experiments on the ScanNetV2 and NYUv2 datasets, and evaluate semantic-affine transformation on various 3D point cloud and 2D image segmentation baselines, where both qualitative and quantitative results demonstrate the superiority and generalization ability of our proposed approach. Code is available at https://github.com/wangzy22/SemAffiNet. Ziyi Wang 0007, Yongming Rao, Xumin Yu, Jie Zhou 0001, Jiwen Lu |
CVPR | 3 |
| 2022 | FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentabstractMost existing action quality assessment methods rely on the deep features of an entire video to predict the score, which is less reliable due to the non-transparent inference process and poor interpretability. We argue that understanding both high-level semantics and internal temporal structures of actions in competitive sports videos is the key to making predictions accurate and interpretable. Towards this goal, we construct a new fine-grained dataset, called FineDiving, developed on diverse diving events with detailed annotations on action procedures. We also propose a procedure-aware approach for action quality assessment, learned by a new Temporal Segmentation Attention module. Specifically, we propose to parse pairwise query and exemplar action instances into consecutive steps with diverse semantic and temporal correspondences. The procedure-aware cross-attention is proposed to learn embeddings between query and exemplar steps to discover their semantic, spatial, and temporal correspondences, and further serve for fine-grained contrastive regression to derive a reliable scoring mechanism. Extensive experiments demonstrate that our approach achieves substantial improvements over the state-of-the-art methods with better interpretability. The dataset and code are available at https://github.com/xujinglin/FineDiving. Jinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen 0002, Jie Zhou 0001, Jiwen Lu |
CVPR | 3 |
| 2022 | Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingabstractWe present Point-BERT, a new paradigm for learning Transformers to generalize the concept of BERT [8] to 3D point cloud. Inspired by BERT, we devise a Masked Point Modeling (MPM) task to pre-train point cloud Transformers. Specifically, we first divide a point cloud into several local point patches, and a point cloud Tokenizer with a discrete Variational AutoEncoder (dVAE) is designed to generate discrete point tokens containing meaningful local information. Then, we randomly mask out some patches of input point clouds and feed them into the backbone Transformers. The pre-training objective is to recover the original point tokens at the masked locations under the supervision of point tokens obtained by the Tokenizer. Extensive experiments demonstrate that the proposed BERT-style pre-training strategy significantly improves the performance of standard point cloud Transformers. Equipped with our pre-training strategy, we show that a pure Transformer architecture attains 93.8% accuracy on ModelNet40 and 83.1% accuracy on the hardest setting of ScanObjectNN, surpassing carefully designed point cloud models with much fewer hand-made designs. We also demonstrate that the representations learned by Point-BERT transfer well to new tasks and domains, where our models largely advance the state-of-the-art of few-shot point cloud classification task. The code and pre-trained models are available at https://github.com/lulutang0608/Point-BERT. Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 0001, Jie Zhou 0001, Jiwen Lu |
CVPR | 1 |
| 2022 | P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel PromptingabstractNowadays, pre-training big models on large-scale datasets has become a crucial topic in deep learning. The pre-trained models with high representation ability and transferability achieve a great success and dominate many downstream tasks in natural language processing and 2D vision. However, it is non-trivial to promote such a pretraining-tuning paradigm to the 3D vision, given the limited training data that are relatively inconvenient to collect. In this paper, we provide a new perspective of leveraging pre-trained 2D knowledge in 3D domain to tackle this problem, tuning pre-trained image models with the novel Point-to-Pixel prompting for point cloud analysis at a minor parameter cost. Following the principle of prompting engineering, we transform point clouds into colorful images with geometry-preserved projection and geometry-aware coloring to adapt to pre-trained image models, whose weights are kept frozen during the end-to-end optimization of point cloud analysis tasks. We conduct extensive experiments to demonstrate that cooperating with our proposed Point-to-Pixel Prompting, better pre-trained image model will lead to consistently better performance in 3D vision. Enjoying prosperous development from image pre-training field, our method attains 89.3% accuracy on the hardest setting of ScanObjectNN, surpassing conventional point cloud models with much fewer trainable parameters. Our framework also exhibits very competitive performance on ModelNet classification and ShapeNet Part Segmentation. Code is available at https://github.com/wangzy22/P2P. Ziyi Wang 0007, Xumin Yu, Yongming Rao, Jie Zhou 0001, Jiwen Lu |
NeurIPS | 2 |
| 2022 | Learning from Temporal Spatial Cubism for Cross-Dataset Skeleton-based Action RecognitionabstractRapid progress and superior performance have been achieved for skeleton-based action recognition recently. In this article, we investigate this problem under a cross-dataset setting, which is a new, pragmatic, and challenging task in real-world scenarios. Following the unsupervised domain adaptation (UDA) paradigm, the action labels are only available on a source dataset, but unavailable on a target dataset in the training stage. Different from the conventional adversarial learning-based approaches for UDA, we utilize a self-supervision scheme to reduce the domain shift between two skeleton-based action datasets. Our inspiration is drawn from Cubism, an art genre from the early 20th century, which breaks and reassembles the objects to convey a greater context. By segmenting and permuting temporal segments or human body parts, we design two self-supervised learning classification tasks to explore the temporal and spatial dependency of a skeleton-based action and improve the generalization ability of the model. We conduct experiments on six datasets for skeleton-based action recognition, including three large-scale datasets (NTU RGB+D, PKU-MMD, and Kinetics) where new cross-dataset settings and benchmarks are established. Extensive results demonstrate that our method outperforms state-of-the-art approaches. The source codes of our model and all the compared methods are available at https://github.com/shanice-l/st-cubism. Yansong Tang, Xumin Yu, Jiwen Lu, Jie Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | PoinTr: Diverse Point Cloud Completion with Geometry-Aware TransformersabstractPoint clouds captured in real-world applications are of-ten incomplete due to the limited sensor resolution, single viewpoint, and occlusion. Therefore, recovering the complete point clouds from partial ones becomes an indispensable task in many practical applications. In this paper, we present a new method that reformulates point cloud completion as a set-to-set translation problem and design a new model, called PoinTr that adopts a transformer encoder-decoder architecture for point cloud completion. By rep-resenting the point cloud as a set of unordered groups of points with position embeddings, we convert the point cloud to a sequence of point proxies and employ the transformers for point cloud generation. To facilitate transformers to better leverage the inductive bias about 3D geometric structures of point clouds, we further devise a geometry-aware block that models the local geometric relationships explicitly. The migration of transformers enables our model to better learn structural knowledge and preserve detailed information for point cloud completion. Furthermore, we propose two more challenging benchmarks with more diverse incomplete point clouds that can better reflect the real-world scenarios to promote future research. Experimental results show that our method outperforms state-of-the-art methods by a large margin on both the new bench-marks and the existing ones. Code is available at https://github.com/yuxumin/PoinTr. Xumin Yu, Yongming Rao, Ziyi Wang 0007, Zuyan Liu, Jiwen Lu, Jie Zhou 0001 |
ICCV | 1 |
| 2021 | Group-aware Contrastive Regression for Action Quality AssessmentabstractAssessing action quality is challenging due to the subtle differences between videos and large variations in scores. Most existing approaches tackle this problem by regressing a quality score from a single video, suffering a lot from the large inter-video score variations. In this paper, we show that the relations among videos can provide important clues for more accurate action quality assessment during both training and inference. Specifically, we reformulate the problem of action quality assessment as regressing the relative scores with reference to another video that has shared attributes (e.g., category and difficulty), instead of learning unreferenced scores. Following this formulation, we propose a new Contrastive Regression (CoRe) framework to learn the relative scores by pair-wise comparison, which highlights the differences between videos and guides the models to learn the key hints for assessment. In order to further exploit the relative information between two videos, we devise a group-aware regression tree to convert the conventional score regression into two easier sub-problems: coarse-to-fine classification and regression in small intervals. To demonstrate the effectiveness of CoRe, we conduct extensive experiments on three mainstream AQA datasets including AQA-7, MTL-AQA and JIGSAWS. Our approach outperforms previous methods by a large margin and establishes new state-of-the-art on all three benchmarks. Xumin Yu, Yongming Rao, Wenliang Zhao, Jiwen Lu, Jie Zhou 0001 |
ICCV | 1 |
| 2020 | Feature Extraction and Classification of Hyperspectral Images Using Hierarchical NetworkabstractIn recent years, researchers have frequently utilized convolutional neural networks (CNNs) to classify hyperspectral images and have, indeed, embraced exciting achievements. However, most of the existing approaches tend to handle images block by block, which is less efficient as image blocks need to be fed into the network for many times. With this in mind, this letter presents a novel hierarchical CNN that adopts raw images as the input and extracts useful features for classification. Specifically, we adopt several hierarchical convolutional neural layers as a feature extractor and adopt the support vector machine instead of the classifying layer in the original network as the final classifier. Experiments show the proposed approach can work efficiently and exhibit competitive performance when compared to some other approaches based on deep networks. Yanlong Gao, Yan Feng 0005, Xumin Yu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Graph Interaction Networks for Relation Transfer in Human Activity VideosabstractRecent years have witnessed rapid progress in employing graph convolutional networks (GCNs) for various video analysis tasks where graph-based data abound. However, exploring the transferable knowledge between different graphs, which is a direction with wide and potential applications, has been rarely studied. To address this issue, we propose a graph interaction networks (GINs) model for transferring relation knowledge across two graphs. Different from conventional domain adaptation or knowledge distillation approaches, our GINs focus on a “self-learned” weight matrix, which is a higher-level representation of the input data. And each element of the weight matrix represents the pair-wise relation among different nodes within the graph. Moreover, we guide the networks to transfer the knowledge across the weight matrices by designing a task-specific loss function, so that the relation information is well preserved during transfer. We conduct experiments on two different scenarios for video analysis, including a new proposed setting for unsupervised skeleton-based action recognition across different datasets, and supervised group activity recognition with multi-modal inputs. Extensive experiments on six widely used datasets illustrate that our GINs achieve very competitive performance in comparison with the state-of-the-arts. Yansong Tang, Yi Wei 0003, Xumin Yu, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |