VLDB 2026 Research / reviewers in the wild / expert
Kangning Yin
dblp:273/8676
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-8652-0151ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedSpike: A personalized federated learning method for energy-efficient spiking neural networks in edge intelligence
Kangning Yin, Zhen Ding, Shaoqi Hou, Ye Li 0024, Yujian Du |
Inf. Sci. | 1 |
| 2026 | Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding SpaceabstractMotion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding space for motion retrieval from text or visual modality. However, these methods lack a more intuitive and user-friendly interaction mode and often overlook the sequential representation of most modalities for improved retrieval performance. To address these limitations, we propose a framework that aligns four modalities—text, audio, video, and motion—within a fine-grained joint embedding space, incorporating audio for the first time in motion retrieval to enhance user immersion and convenience. This fine-grained space is achieved through a sequence-level contrastive learning approach, which captures critical details across modalities for better alignment. To evaluate our framework, we augment existing text-motion datasets with synthetic but diverse audio recordings, creating two multi-modal motion retrieval datasets. Experimental results demonstrate superior performance over state-of-the-art methods across multiple sub-tasks, including an 10.16% improvement in R@10 for text-to-motion retrieval and a 25.43% improvement in R@1 for video-to-motion retrieval on the HumanML3D dataset. Furthermore, our results show that our 4-modal framework significantly outperforms its 3-modal counterpart, underscoring the potential of multi-modal motion retrieval for advancing motion acquisition. Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian 0002, Weixin Si, Shihao Zou |
IEEE Trans. Multim. | 3 |
| 2025 | Towards heterogeneous tasks conflict avoidance for cross-modal federated learning via knowledge distillation
Kangning Yin, Xinhui Ji, Zhen Ding, Shaoqi Hou, Zhiguo Wang 0004 |
Inf. Sci. | 1 |
| 2025 | Continual adaptation Person re-identification via vision-language fusion with enhanced annotation robustness
Xiuchuan Cheng, Kangning Yin, Zhen Ding, Guisong Liu, Zhiguo Wang 0004 |
Multim. Syst. | 2 |
| 2025 | Self-attention fusion and adaptive continual updating for multimodal federated learning with heterogeneous data
Kangning Yin, Zhen Ding, Xinhui Ji, Zhiguo Wang 0004 |
Neural Networks | 1 |
| 2024 | DHFM-FLM: A Dynamic Hierarchical Federated Learning Mechanism for Financial Models under Client Resource HeterogeneityabstractFederated Learning (FL) is an emerging distributed machine learning technology. However, in practical applications, it frequently encounters the challenge of client resource heterogeneity. This can result in long wait times or even model training failures during the communication process of FL. To address this problem, we propose a dynamic hierarchical federated learning mechanism for financial models (DHFM-FLM). The local client adopts a dynamic model training design that leverages the property of resource heterogeneity to enhance the performance of the local model. To avoid prolonged wait times for failing clients, a dynamic communication detection design is proposed at three critical junctures. In each round, the model hierarchical reservation communication design is employed to collect models in segments, thus reducing communication congestion and preventing malicious attacks on the communication process. Experiments with heterogeneous computation and communication resources demonstrate that utilizing the DHFM-FLM boosts model performance by approximately 5-8% and reduces communication time by about 15%. Additionally, DHFM-FLM increases the success rate of the FL task by approximately 6%. Kangning Yin, Zhen Ding, Shaoqi Hou, Xinhui Ji, Guangqiang Yin, Zhiguo Wang 0004 |
IEEE Big Data | 1 |
| 2024 | Tri-Modal Motion Retrieval by Learning a Joint Embedding SpaceabstractInformation retrieval is an ever-evolving and crucial re-search domain. The substantial demand for high-quality human motion data especially in online acquirement has led to a surge in human motion research works. Prior works have mainly concentrated on dual-modality learning, such as text and motion tasks, but three-modality learning has been rarely explored. Intuitively, an extra introduced modality can enrich a model's application scenario, and more importantly, an adequate choice of the extra modality can also act as an intermediary and enhance the alignment between the other two disparate modalities. In this work, we introduce LAVIMO (LAnguage-VIdeo-MOtion alignment), a novel framework for three-modality learning integrating human-centric videos as an additional modality, thereby ef-fectively bridging the gap between text and motion. More-over, our approach leverages a specially designed attention mechanism to foster enhanced alignment and synergistic effects among text, video, and motion modalities. Empirically, our results on the HumanML3D and KIT-ML datasets show that LAVIMO achieves state-of-the-art performance in various motion-related cross-modal retrieval tasks, in-cluding text-to-motion, motion-to-text, video-to-motion and motion-to-video. Our project webpage can be found in https://lavimo2023.github.io/LAVIMO/. Kangning Yin, Shihao Zou, Yuxuan Ge, Zheng Tian 0002 |
CVPR | 1 |
| 2024 | RACon: Retrieval-Augmented Simulated Character Locomotion ControlabstractIn computer animation, driving a simulated character with lifelike motion is challenging. Current generative models, though able to generalize to diverse motions, often pose challenges to the responsiveness of end-user control. To address these issues, we introduce RACon: Retrieval-Augmented Simulated Character Locomotion Control. Our end-to-end hierarchical reinforcement learning method utilizes a retriever and a motion controller. The retriever searches motion experts from a user-specified database in a task-oriented fashion, which boosts the responsiveness to the user’s control. The selected motion experts and the manipulation signal are then transferred to the controller to drive the simulated character. In addition, a retrieval-augmented discriminator is designed to stabilize the training process. Our method surpasses existing techniques in both quality and quantity in locomotion control, as demonstrated in our empirical study. Moreover, by switching extensive databases for retrieval, it can adapt to distinctive motion types at run time. We will release our code upon acceptance. Yuxuan Mu, Shihao Zou, Kangning Yin, Zheng Tian 0002, Li Cheng 0001, Weinan Zhang 0001, Jun Wang 0012 |
ICME | 3 |
| 2023 | A multitask joint framework for real-time person search
Ye Li 0024, Kangning Yin, Zhuofu Tan, Xinzhong Wang, Guangqiang Yin, Zhiguo Wang 0004 |
Multim. Syst. | 2 |
| 2021 | Person Re-identification Algorithm Based on Spatial Attention Network
Shaoqi Hou, Kangning Yin, Guangqiang Yin |
WASA (3) | 3 |
| 2021 | Pedestrian re-identification based on attribute mining and reasoningabstractAbstract The high‐level semantic information extracted from the pedestrian attribute feature is an important element for pedestrian recognition. Pedestrian attribute recognition plays an important role in both intelligent video surveillance and pedestrian re‐identification promoting the convenience of searching and performance of model. This paper tries finding a practical method to improve the performance of the pedestrian re‐identification by combining pedestrian attributes and identities. The multi‐task learning method combines pedestrian recognition and attribute information in a direct way that considers the correlation between pedestrian attributes and identities but ignores the principle and degree of such correlation. To solve this problem, a new pedestrian recognition framework based on attribute mining and reasoning is proposed in this paper. To enhance the expression ability of attribute features, it designs spatial channel attention module (SCAM) based on attention mechanism to extract features from every attribute. SCAM can not only locate the attributes on the feature map, but also effectively mine channel features with a higher degree of association with attributes. In addition, both spatial attention model and channel attention model are integrated by multiple groups of parallel branches, which further improve the network performance. Finally, using the semantic reasoning and information transmission function of graph convolutional network, the relationship between attribute features and pedestrian features can be mined. Besides, pedestrian features with stronger expression ability can also be obtained. Experiment work is conducted in two databases, DukeMTMC‐reID and Market‐1501, which are commonly used in pedestrian recognition tasks. On the Market‐1501 dataset, the final effect of the algorithm model CMC‐1 can reach 94.74%, and mAP can reach 87.02%; on the DukeMTMC‐reID dataset, CMC‐1 can reach 87.03%, and mAP can reach 77.11%. The results show that our method is at the top of the existing pedestrian recognition methods. Chao Li 0053, Xiaoyu Yang 0008, Kangning Yin, Yifan Chang, Zhiguo Wang 0004, Guangqiang Yin |
IET Image Process. | 3 |
| 2021 | SAN-GAL: Spatial Attention Network Guided by Attribute Label for Person Re-identificationabstractPerson Re‐identification (Re‐ID) is aimed at solving the matching problem of the same pedestrian at a different time and in different places. Due to the cross‐device condition, the appearance of different pedestrians may have a high degree of similarity; at this time, using the global features of pedestrians to match often cannot achieve good results. In order to solve these problems, we designed a Spatial Attention Network Guided by Attribute Label (SAN‐GAL), which is a dual‐trace network containing both attribute classification and Re‐ID. Different from the previous approach of simply adding a branch of attribute binary classification network, our SAN‐GAL is mainly divided into two connecting steps. First, with attribute labels as guidance, we generate Attribute Attention Heat map (AAH) through Grad‐CAM algorithm to accurately locate fine‐grained attribute areas of pedestrians. Then, the Attribute Spatial Attention Module (ASAM) is constructed according to the AHH which is taken as the prior knowledge and introduced into the Re‐ID network to assist in the discrimination of the Re‐ID task. In particular, our SAN‐GAL network can integrate the local attribute information and global ID information of pedestrians without introducing additional attribute region annotation, which has good flexibility and adaptability. The test results on Market1501 and DukeMTMC‐reID show that our SAN‐GAL can achieve good results and can achieve 85.8% Rank‐1 accuracy on DukeMTMC‐reID dataset, which is obviously competitive compared with most Re‐ID algorithms. Shaoqi Hou, Kangning Yin, Yiyin Ding, Zhiguo Wang 0004, Guangqiang Yin |
Wirel. Commun. Mob. Comput. | 3 |
| 2020 | A Real-Time Vehicle Logo Detection Method Based on Improved YOLOv2
Kangning Yin, Shaoqi Hou, Ye Li 0024, Chao Li 0053, Guangqiang Yin |
WASA (1) | 1 |