VLDB 2026 Research / reviewers in the wild / expert
Kun Yang 0010
dblp:63/1587-10
· DBLP profile ↗
30ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0002-9956-2200ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 14 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR AdvancementabstractRecent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this,we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model’s performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods. Weitao Jia, Jinghui Lu, Haiyang Yu 0004, Guozhi Tang, An-Lan Wang, Weijie Yin, Dingkang Yang, Yuxiang Nie, Bin Shan, Hao Feng 0009, Irene Li, Kun Yang 0010, Jingqun Tang, Teng Fu 0001, Changhong Jin, Xiaohui Lv, Can Huang 0002 |
AAAI | 13 |
| 2026 | MeritFL: Self-Regulating Federated Learning via Merit-Gated Communication
Zhengliang Guo, Kun Yang 0010, Linxiao Gong, Yang Liu 0246, Jing Liu 0050 |
ISCAS | 2 |
| 2026 | A knowledge-driven self-supervised learning method for enhancing EEG-based emotion recognition
Hanqi Wang, Peng Ye 0006, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003 |
Neural Networks | 4 |
| 2026 | Multimodal human video generation with uncertainty-aware pose guidance
Kun Yang 0010, Yuanyuan Meng, Juncen Guo, Yanda Meng, Songwen Pei, Jing Liu 0050, Yang Liu 0246 |
Pattern Recognit. | 3 |
| 2026 | FourierMask: Explain EEG-Based End-to-End Deep Learning Models in the Frequency DomainabstractThe rise of EEG-based end-to-end deep learning models has underscored the need to elucidate how these models process time-series raw EEG signals to generate predictions. The frequency domain provides a more suitable perspective for this task due to two key advantages: the strong correlation with cognitive states and the inherent capacity to model long-range temporal dependencies. However, this perspective remains underexplored in existing research. To bridge this gap, we propose FourierMask, the first mask perturbation framework specifically designed for frequency-domain explanation of EEG-based end-to-end models. Our method introduces three key innovations. First, the Fourier-based domain transformation enables direct manipulation of spectral components. Second, A learnable mask mechanism jointly models the spectral-spatial couplings relationship for EEG explanation. Third, a perturbation generator constrained by a target alignment loss ensures natural perturbations by minimizing distribution shift via cluster-aware regularization. We validate our method through experiments on an EEG benchmark dataset across EEGNet, TSCeption, and DeepConvNet models. Our method reaches a 36.0% average accuracy drop gap (vs. 8.6% for LIME and 6.6% for easyPEASI) at the group-level. And, it reaches a 17.8% average accuracy drop gap (vs. 8.9% for LIME and 9.9% for easyPEASI) at the instance-level. Our model-agnostic framework provides a plug-and-play solution for enhancing transparency of EEG-based end-to-end deep learning models. It links model decisions to frequency biomarkers, with potential applications in neuromedicine and brain-computer interfaces. Hanqi Wang, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | UniqueNFT: Uniqueness Protection of Digital Assets in Decentralized WebabstractWith the rapid evolution of the Decentralized Web (DWeb), decentralized technologies have paved new avenues for Web3 applications and the authentication of digital assets. Among them, Non-Fungible Tokens (NFTs) have gained significant popularity due to their immutability and uniqueness, reshaping the landscape of artistic creation, marketing, and intellectual property protection. However, current blockchain-based NFT implementations still face core challenges within decentralized architecture: how to maintain decentralization while ensuring the visual uniqueness of digital assets and reducing storage costs. The rampant issue of duplication undermines the scarcity of digital art and erodes market confidence in copyright authenticity. Moreover, high gas fees and energy consumption further hinder the widespread adoption of NFTs, while reliance on external storage solutions like InterPlanetary File System (IPFS) introduces risks of data instability and loss. To address these challenges, this article presents the UniqueNFT framework, a novel architecture that deeply integrates blockchain oracles with decentralized storage verification mechanisms. The framework achieves three key technological breakthroughs: Using image inversion and generation techniques based on Encoder for Editing (E4E) and StyleGAN3, it extracts compact and expressive semantic features from NFT images, enabling efficient data compression and significantly reducing on-chain storage volume; The Crypto-Mask algorithm, by utilizing the hash value of blockchain user information (user-controlled SHA-256 digest of Ethereum address, user nickname, and registration time), ensures the visual uniqueness of NFTs; A smart contract extension compatible with the ERC721 standard, demonstrating UniqueNFT’s seamless integration within the blockchain ecosystem. By leveraging the technologies of the Decentralized Web, our framework represents an important step forward in enhancing the security and uniqueness of digital assets. It not only innovatively resolves the issues of NFT duplication and homogenization but also injects new vitality and long-term momentum into the creation of a trusted, sustainable blockchain-based digital asset ecosystem. Kun Yang 0010, Haihan Duan, Runhao Zeng, Xiping Hu |
ACM Trans. Web | 1 |
| 2025 | Robust Multi-Agent Collaborative Perception via Spatio-Temporal AwarenessabstractAs an emerging application in autonomous driving, multi-agent collaborative perception has recently received significant attention. Despite promising advances from previous efforts, several unavoidable challenges that cause performance bottlenecks remain, including the single-frame detection dilemma, communication redundancy, and defective collaboration process. To this end, we proposeSCOPE++, a versatile collaborative perception framework aggregating spatio-temporal information across on-road agents to tackle these issues. We introduce four components inSCOPE++for robust collaboration by seeking a reasonable trade-off between perception performance and communication bandwidth. First, we devise a context-aware information aggregation to capture valuable semantic cues in the temporal context and enhance the current local representation of the ego agent. Second, an exclusivity-aware sparse communication is introduced to filter perceptually unnecessary information from collaborators and transmit complementary features relative to the ego agent. Third, we present an importance-aware cross-agent collaboration to incorporate semantic representations of spatially critical locations across agents flexibly. Finally, a contribution-aware adaptive fusion is designed to integrate multi-source representations based on dynamic contributions. Our framework is evaluated on multiple LiDAR-based collaborative detection datasets in real-world and simulated scenarios, and comprehensive experiments show thatSCOPE++outperforms state-of-the-art methods on all datasets. Kun Yang 0010, Zhi Xu 0010, Dingkang Yang, Lihua Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Efficient and Robust Collaborative Perception via Cross-Vehicle Spatio-Temporal Feature SelectingabstractCollaborative perception systems enhance the perception capabilities of individual vehicles by facilitating information exchange between neighbouring vehicles. This approach effectively addresses challenges like occlusions and long-range perceptions that single vehicle cannot manage alone. However, practical applications often face difficulties due to constraints in wireless communication resources and reliability, which limit the effectiveness of latency-sensitive collaborative perception. To overcome these barriers, we introduce CERCP, a Communication Efficient and Robust Collaborative Perception framework. CERCP comprises two core modules: a cross-vehicle spatio-temporal feature selection module, which minimizes communication by transmitting only essential sensor regions with spatio-temporal complementarity, and a global-aware feature synchronization module, which mitigates data delays due to communication latency. To our knowledge, CERCP is the first general collaborative perception framework designed for efficient communication and is applicable across various tasks and modalities. We comprehensively evaluate CERCP on three datasets from real-world and simulated scenarios, using two sensor modalities (LiDAR and camera) and two perception tasks (3D object detection and BEV semantic segmentation). Extensive experiments demonstrate the superior performance of our method. Kun Yang 0010, Hanqi Wang, Peng Sun 0007 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesabstractMultimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion approaches. To this end, we propose a Unified multimodal Missing modality self-Distillation Framework (UMDF) to handle the problem of uncertain missing modalities in MSA. Specifically, a unified self-distillation mechanism in UMDF drives a single network to automatically learn robust inherent representations from the consistent distribution of multimodal data. Moreover, we present a multi-grained crossmodal interaction module to deeply mine the complementary semantics among modalities through coarse- and fine-grained crossmodal attention. Eventually, a dynamic feature integration module is introduced to enhance the beneficial semantics in incomplete modalities while filtering the redundant information therein to obtain a refined and robust multimodal representation. Comprehensive experiments on three datasets demonstrate that our framework significantly improves MSA performance under both uncertain missing-modality and complete-modality testing conditions. Mingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang 0001, Shuaibing Wang, Liuzhen Su, Kun Yang 0010, Lihua Zhang 0002 |
AAAI | 7 |
| 2024 | Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete ModalitiesabstractMultimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause uncertain modality missingness, which drastically degrades the model's performance. To this end, we propose a Correlation-decoupled Knowledge Distillation (CorrKD) framework for the MSA task under uncertain missing modalities. Specifically, we present a sample-level contrastive distillation mechanism that transfers comprehensive knowledge containing cross-sample correlations to reconstruct missing semantics. Moreover, a category-guided prototype distillation mechanism is introduced to capture cross-category correlations using category prototypes to align feature distributions and generate favorable joint representations. Eventually, we design a response-disentangled consistency distillation strategy to optimize the sentiment decision boundaries of the student network through response disentanglement and mutual information maximization. Comprehensive experiments on three datasets indicate that our framework can achieve favorable improvements compared with several baselines. Mingcheng Li, Dingkang Yang, Shuaibing Wang, Yan Wang 0068, Kun Yang 0010, Dongliang Kou, Ziyun Qian, Lihua Zhang 0002 |
CVPR | 6 |
| 2024 | Robust Emotion Recognition in Context DebiasingabstractContext-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the target person's emotional state. Despite advancements, the biggest challenge remains due to context bias interference. The harmful bias forces the models to rely on spurious correlations between background contexts and emotion labels in likelihood estimation, causing severe performance bottlenecks and confounding valuable context priors. In this paper, we propose a counterfactual emotion inference (CLEF) framework to address the above issue. Specifically, we first formulate a generalized causal graph to decouple the causal relationships among the variables in CAER. Following the causal graph, CLEF introduces a non-invasive context branch to capture the adverse direct effect caused by the context bias. During the inference, we eliminate the direct context effect from the total causal effect by comparing factual and counterfactual outcomes, resulting in bias mitigation and robust prediction. As a model-agnostic framework, CLEF can be readily integrated into existing methods, bringing consistent performance gains. Dingkang Yang, Kun Yang 0010, Mingcheng Li, Shunli Wang 0001, Shuaibing Wang, Lihua Zhang 0002 |
CVPR | 2 |
| 2024 | ERMVP: Communication-Efficient and Collaboration-Robust Multi-Vehicle Perception in Challenging EnvironmentsabstractCollaborative perception enhances perception performance by enabling autonomous vehicles to exchange complementary information. Despite its potential to revolutionize the mobile industry, challenges in various environments, such as communication bandwidth limitations, localization errors and information aggregation inefficiencies, hinder its implementation in practical applications. In this work, we propose ERMVP, a communication-Efficient and collaboration-Robust Multi-Vehicle Perception method in challenging environments. Specifically, ERMVP has three distinct strengths: i) It utilizes the hierarchical feature sampling strategy to abstract a representative set of feature vectors, using less communication overhead for efficient communication; ii) It employs the sparse consensus features to execute precise spatial location calibrations, effectively mitigating the implications of vehicle localization errors; iii) A pioneering feature fusion and interaction paradigm is introduced to integrate holistic spatial semantics among different vehicles and data sources. To thoroughly validate our method, we conduct extensive experiments on real-world and simulated datasets. The results demonstrate that the proposed ERMVP is significantly superior to the state-of-the-art collaborative perception methods. Kun Yang 0010, Hanqi Wang, Peng Sun 0007 |
CVPR | 2 |
| 2024 | Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002 |
ECCV (58) | 5 |
| 2024 | Align Before Collaborate: Mitigating Feature Misalignment for Robust Multi-agent Perception
Kun Yang 0010, Dingkang Yang, Ke Li 0015, Dongling Xiao, Zedian Shao, Peng Sun 0007 |
ECCV (4) | 1 |
| 2024 | Guidelines for Parameter Selection in Traffic Light Control Methods Using Reinforcement Learning: Insights from Empirical StudiesabstractThe ever-changing traffic dynamics make the traditional traffic signal control methods unable to adapt to the environment. Meanwhile, deep reinforcement learning (DRL) has the property of interacting with the environment and adapting to changes in the environment. Therefore, in recent years, researchers have usually solved traffic signal control (TSC) problems through DRL methods. They have not only improved the design of neural networks, but also improved the ability of models to understand traffic conditions and learn corresponding task requests by designing different states and rewards. However, although the existing TSC algorithms based on DRL have proposed many well-designed states and reward strategies, which combinations of states and rewards should be adopted in practice to achieve the performance margin of models remains a question that researchers are seeking the answer to. Therefore, we introduce a general simulation platform to test and compare experimental performance under different combinations of states and rewards. Specifically, we test and analyze the experimental effects under different combinations of multiple traffic states and rewards through various TSC methods with a set of unified model settings. We further design and test some new state representations and reward strategies based on more detailed traffic information. The test results show that when researchers design the state and reward, refining the traffic state like vehicle running condition and making the state and reward match can make the experimental performance better than other combinations in most cases. We hope these results have some implications for the state and reward choice when researchers conduct experiments on TSC problem or other traffic decision management problems. Lang Qian, Peng Sun 0007, Kun Yang 0010, Azzedine Boukerche |
IWQoS | 3 |
| 2024 | A novel hierarchical distributed vehicular edge computing framework for supporting intelligent driving
Kun Yang 0010, Peng Sun 0007, Dingkang Yang, Jieyu Lin, Azzedine Boukerche |
Ad Hoc Networks | 1 |
| 2024 | The evolution of detection systems and their application for intelligent transportation systems: From solo to symphony
Zedian Shao, Kun Yang 0010, Peng Sun 0007, Yulin Hu, Azzedine Boukerche |
Comput. Commun. | 2 |
| 2024 | Towards Context-Aware Emotion Recognition Debiasing From a Causal Demystification Perspective via De-Confounded TrainingabstractUnderstanding emotions from diverse contexts has received widespread attention in computer vision communities. The core philosophy of Context-Aware Emotion Recognition (CAER) is to provide valuable semantic cues for recognizing the emotions of target persons by leveraging rich contextual information. Current approaches invariably focus on designing sophisticated structures to extract perceptually critical representations from contexts. Nevertheless, a long-neglected dilemma is that a severe context bias in existing datasets results in an unbalanced distribution of emotional states among different contexts, causing biased visual representation learning. From a causal demystification perspective, the harmful bias is identified as a confounder that misleads existing models to learn spurious correlations based on likelihood estimation, limiting the models' performance. To address the issue, we embrace causal inference to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a customized causal graph. Subsequently, we present a Contextual Causal Intervention Module (CCIM) to de-confound the confounder, which is built upon backdoor adjustment theory to facilitate seeking approximate causal effects during model training. As a plug-and-play component, CCIM can easily integrate with existing approaches and bring significant improvements. Systematic experiments on three datasets demonstrate the effectiveness of our CCIM. Dingkang Yang, Kun Yang 0010, Haopeng Kuang, Zhaoyu Chen 0001, Lihua Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Towards Asynchronous Multimodal Signal Interaction and Fusion via Tailored TransformersabstractThe signals from human expressions are usually multimodal, including natural language, facial gestures, and acoustic behaviors. A key challenge is how to fuse multimodal time-series signals with temporal asynchrony. To this end, we present a Transformer-driven Signal Interaction and Fusion (TSIF) approach to effectively model asynchronous multimodal signal sequences. TSIF consists of linear and cross-modal transformer modules with different duties. The linear transformer module efficiently performs the global interaction for multimodal signals, and the vital philosophy is to replace the dot product similarity with the Exponential Kernel while achieving linear complexity by a low-rank matrix decomposition. By targeting the language modality, the cross-modal transformer module aims to capture reliable element correlations among distinct signals and mitigate noise interference in audio and visual modalities. Numerous experiments on two multimodal benchmarks show that our TSIF comparably outperforms previous state-of-the-art models with lower space-time complexities. The systematic analysis also proves the effectiveness of the proposed modules. Dingkang Yang, Haopeng Kuang, Kun Yang 0010, Mingcheng Li, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic RepresentationsabstractUnderstanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach. Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | AMP-Net: Appearance-Motion Prototype Network Assisted Automatic Video Anomaly Detection SystemabstractAs essential tools for industry safety protection, automatic video anomaly detection systems (AVADS) are designed to detect anomalous events of concern in surveillance videos. Existing VAD methods lack effective exploration of the prototypical appearance and motion features leading to poor performance in realistic scenarios. Specifically, they either misreport regular events as anomalies due to insufficient representation power, or lead to missed detections with over-power generalization. In this regard, we propose an appearance-motion prototype network (AMP-net) that uses external memories to record prototype features and augments the appearance-motion prototype with a spatial-temporal fusion. In addition, AMP-net sequentially fuses appearance features from deep to shallow to utilize multiscale spatial context. Additionally, we introduce temporal attention to capture important dynamics and enhance AMP-net for representing regular motion. The proposed method achieves a delicate balance of effective representation of normal events and limited generalization to anomalies. Experiments on three benchmark datasets demonstrate that our method can accurately detect anomalous events, achieving performance comparable to state-of-the-art methods with frame-level AUCs of 98.7%, 92.4%, and 78.8% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. Moreover, we conducted a case study on the self-collected industrial dataset, and the results indicate that our AMP-net can cope with complex industrial scenarios and outperform existing methods. Yang Liu 0246, Jing Liu 0050, Kun Yang 0010, Bobo Ju, Siao Liu, Dingkang Yang, Peng Sun 0007 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | A Novel Efficient Multi-View Traffic-Related Object Detection FrameworkabstractWith the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Accordingly, we propose a novel traffic-related framework named CEVAS to achieve efficient object detection using multi-view video data. Briefly, a fine-grained input filtering policy is introduced to produce a reasonable region of interest from the captured images. Also, we design a sharing object manager to manage the information of objects with spatial redundancy and share their results with other vehicles. We further derive a content-aware model selection policy to select detection methods adaptively. Experimental results show that our framework significantly reduces response latency while achieving the same detection accuracy as the state-of-the-art methods. Kun Yang 0010, Jing Liu 0050, Dingkang Yang, Hanqi Wang, Peng Sun 0007 |
ICASSP | 1 |
| 2023 | A Novel Multi-Factor Aware Online Scheduling Method for Improving Vehicular Edge Computing EfficiencyabstractVehicular Edge Computing (VEC), as one of the major components of Intelligent Transportation Systems, improves road safety by providing computing services to safety-related applications on vehicles. Currently, the existing fine-grained computing scheduling algorithms are normally designed based on some simple scheduling policies. Due to the heterogeneous nature of tasks offloaded from various applications, they may not effectively satisfy various performance requirements of the real system, thereby leading to the problem that the short-term residual computing power cannot be effectively utilized when computing-costly tasks occupy the server. Therefore, improving the overall system performance and the efficiency of utilizing computing power is a critical issue. Accordingly, in this paper, we study the problem of computing scheduling inside edge servers in VEC, where multiple tasks can be offloaded to Road Side Units (RSUs). We analyze the role played by multiple evaluation metrics in the existing methods for ensuring the quality of service (QoS) and further design a novel online multi-factor aware task offloading algorithm with a hierarchical fine-grained computing scheduling scheme inside the edge server. We evaluate it by conducting intensive simulation tests and comparing the results with some state-of-the-art approaches. Numerical results show that the proposed algorithm outperforms the methods in the control group in different aspects and achieves the best overall performance. Lang Qian, Peng Sun 0007, Kun Yang 0010, Azzedine Boukerche |
ICC | 3 |
| 2023 | AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionabstractDriver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE. Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002 |
ICCV | 9 |
| 2023 | Spatio-Temporal Domain Awareness for Multi-Agent Collaborative PerceptionabstractMulti-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this emerging research. In this paper, we propose SCOPE, a novel collaborative perception frame-work that aggregates the spatio-temporal awareness characteristics across on-road agents in an end-to-end manner. Specifically, SCOPE has three distinct strengths: i) it considers effective semantic cues of the temporal context to enhance current representations of the target agent; ii) it aggregates perceptually critical spatial information from heterogeneous agents and overcomes localization errors via multi-scale feature interactions; iii) it integrates multi-source representations of the target agent based on their complementary contributions by an adaptive fusion paradigm. To thoroughly evaluate SCOPE, we consider both real-world and simulated scenarios of collaborative 3D object detection tasks on three datasets. Extensive experiments show the superiority of our approach and the necessity of the proposed components. The project link is https://ydk122024.github.io/SCOPE/. Kun Yang 0010, Dingkang Yang, Mingcheng Li, Yang Liu 0246, Jing Liu 0050, Hanqi Wang, Peng Sun 0007 |
ICCV | 1 |
| 2023 | A Novel Robust Reinforcement Learning-based Dependent Task Offloading Algorithm for Mobile Edge IntelligenceabstractWith the rise of advanced applications based on Artificial Intelligence (AI) and Internet-of-Things (IoT), mobile devices have become more intelligent, introducing a novel concept, Mobile Edge Intelligence. But the limited on-board resources often hinder the capabilities of mobile devices. Mobile Edge Computing (MEC), regarded as an effective method to expand device capability, effectively overcomes this barrier. However, the dynamic networks driven by mobility and the dependency on applications pose significant challenges for offloading, which can degrade MEC’s overall performance. Therefore, how to effectively combine the above points to achieve a stable and effective sharing of computing resources between devices and servers is a critical issue. In this paper, we consider a multi-slot MEC system with device mobility and multiple applications of unknown arrival. To improve application completion rate while reducing task delay, we introduce a novel, robust distributed offloading algorithm, which calls the Multi-Attention Pointer network-based Reinforcement Learning algorithm (MAPRL), for the dynamic and unstable resource offloading scenario. Numerous experiments have been carried out to demonstrate that, compared with the existing methods, MAPRL exhibits robustness when facing the changing scenario, it can adapt to the unknown workload and dynamic network connections to enhance the offloading performance. Peng Sun 0007, Kun Yang 0010, Gaoyun Fang, Azzedine Boukerche |
ICPADS | 3 |
| 2023 | What2comm: Towards Communication-efficient Collaborative Perception via Feature DecouplingabstractMulti-agent collaborative perception has received increasing attention recently as an emerging application in driving scenarios. Despite advancements in previous approaches, challenges remain due to redundant communication patterns and vulnerable collaboration processes. To address these issues, we propose What2comm, an end-to-end collaborative perception framework to achieve a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we design an efficient communication mechanism based on feature decoupling to transmit exclusive and common feature maps among heterogeneous agents to provide perceptually holistic messages. Secondly, a spatio-temporal collaboration module is introduced to integrate complementary information from collaborators and temporal ego cues, leading to a robust collaboration procedure against transmission delay and localization errors. Ultimately, we propose a common-aware fusion strategy to refine final representations with informative common features. Comprehensive experiments in real-world and simulated scenarios demonstrate the effectiveness of What2comm. Kun Yang 0010, Dingkang Yang, Hanqi Wang, Peng Sun 0007 |
ACM Multimedia | 1 |
| 2023 | How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionabstractMulti-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay, and collaboration heterogeneity. To tackle these issues, we propose \textit{How2comm}, a collaborative perception framework that seeks a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we devise a mutual information-aware communication mechanism to maximally sustain the informative features shared by collaborators. The spatial-channel filtering is adopted to perform effective feature sparsification for efficient communication. Second, we present a flow-guided delay compensation strategy to predict future characteristics from collaborators and eliminate feature misalignment due to temporal asynchrony. Ultimately, a pragmatic collaboration transformer is introduced to integrate holistic spatial semantics and temporal context clues among agents. Our framework is thoroughly evaluated on several LiDAR-based collaborative detection datasets in real-world and simulated scenarios. Comprehensive experiments demonstrate the superiority of How2comm and the effectiveness of all its vital components. The code will be released at https://github.com/ydk122024/How2comm. Dingkang Yang, Kun Yang 0010, Jing Liu 0050, Zhi Xu 0010, Rongbin Yin, Peng Zhai, Lihua Zhang 0002 |
NeurIPS | 2 |
| 2023 | Target and source modality co-reinforcement for emotion understanding from asynchronous multimodal sequences
Dingkang Yang, Yang Liu 0246, Can Huang 0002, Mingcheng Li, Kun Yang 0010, Yan Wang 0068, Peng Zhai, Lihua Zhang 0002 |
Knowl. Based Syst. | 7 |
| 2022 | A Novel Distributed Task Scheduling Framework for Supporting Vehicular Edge IntelligenceabstractIn recent years, data-driven intelligent transportation systems (ITS) have developed rapidly and brought various AI-assisted applications to improve traffic efficiency. However, these applications are constrained by their inherent high computing demand and the limitation of vehicular computing power. Vehicular edge computing (VEC) has shown great potential to support these applications by providing computing and storage capacity in close proximity. For facing the heterogeneous nature of in-vehicle applications and the highly dynamic network topology in the Internet-of-Vehicle (IoV) environment, how to achieve efficient scheduling of computational tasks is a critical problem. Accordingly, we design a two-layer distributed online task scheduling framework to maximize the task acceptance ratio (TAR) under various QoS requirements when facing unbalanced task distribution. Briefly, we implement the computation offloading and transmission scheduling policies for the vehicles to optimize the onboard computational task scheduling. Meanwhile, in the edge computing layer, a new distributed task dispatching policy is developed to maximize the utilization of system computing power and minimize the data transmission delay caused by vehicle motion. Through single-vehicle and multi-vehicle simulations, we evaluate the performance of our framework, and the experimental results show that our method outperforms the state-of-the-art algorithms. Moreover, we conduct ablation experiments to validate the effectiveness of our core algorithms. Kun Yang 0010, Peng Sun 0007, Jieyu Lin, Azzedine Boukerche |
ICDCS | 1 |