VLDB 2026 Research / reviewers in the wild / expert
Yujun Ma
dblp:152/5036
· DBLP profile ↗
21ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-2733-8813ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Computer networks · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Action Recognition in RGB-D Videos via an End-to-End Cross-Modal Attention TransformerabstractRecognizing actions in RGB-D videos necessitates an in-depth understanding of spatial and temporal information from both RGB and depth modalities, as well as an efficient fusion of these data streams. Existing late fusion approaches often suffer from modality collapse due to high semantic similarity between modalities, where the fusion diminishes the distinct contributions of each modality and compromises overall effectiveness. In this context, we propose a novel End-to-end Cross-Modal Attention Transformer (E-CMAT) model. Our model processes RGB and depth inputs through two distinct expert encoders that utilize factorized spatio-temporal representations to capture dimension-independent features, enhancing motion understanding. The extracted RGB and depth tokens are subsequently fused via our innovative cross-modal attention, which is applied iteratively to ensure a robust integration that accentuates critical patterns and discrepancies between the modalities, thereby facilitating more precise action recognition. Our cross-modal attention utilizes one class token per expert for inter-expert information exchange, focusing attention on salient features and fostering synergistic predictions while reducing computational complexity from quadratic to linear. Furthermore, extensive experiments conducted on widely used benchmark datasets, such as NTU RGB-D 60, NTU RGB-D 120, and THU-READ, demonstrate the favorable performance of E-CMAT compared to state-of-the-art models. Yujun Ma, Benjia Zhou, Hong Zhang 0050, Xizheng Zhang, Xiangjian He, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Scale-Aware Attention and Multi-Modal Prompt Learning With Fusion Adapter for RGBT TrackingabstractFusing visible (RGB) and thermal (T) images for RGBT tracking has received growing interest in the field of computer vision. However, how to improve the robustness of the tracker to target scale variety, effectively apply visual prompts to multimodal tracking tasks, and enhance the multimodal fusion effectiveness are still urgent challenges in the field of RGBT tracking. To this purpose, this work proposes an RGBT tracking framework integrating scale-aware dilation attention, multimodal prompt interaction learning, and cross- fusion adapter, named MPANet. Firstly, a scale-aware dilation attention (SADA) module is put forward to enhance the flexibility of the tracker in the presence of target scale variations by embedding convolutions with different dilation rates into the self-attention. Subsequently, a multimodal prompt interaction learning (MPIL) module is constructed, which combines global token adaptive attention and spatial attention to efficiently learn visual prompts from different modalities and achieve intermodal prompt interactions. Finally, a cross-fusion adapter (CFA) is developed to facilitate the adaptability of the network to different modalities in the process of multimodal information fusion through the adapter mechanism. Extensive experiments on public RGBT benchmark tracking datasets such as GTOT, RGBT234, LasHeR and VTUAV demonstrate that the proposed method outperforms existing advanced trackers and achieves state-of-the-art performance. Victor S. Sheng, Yujun Ma, Xiaoguo Liang, Guanbo Wang |
IEEE Trans. Multim. | 4 |
| 2025 | Laplacian eigenmaps based manifold regularized CNN for visual recognition
Ming Zong, Zhizhong Ma, Fangyi Zhu, Yujun Ma, Ruili Wang 0001 |
Inf. Sci. | 4 |
| 2025 | Survey on deep learning in multimodal medical imaging for cancer detection
Zhaocheng Xu, Yujun Ma, Weiping Ding 0001, Ruili Wang 0001, Zhihong Gao, Guohua Cheng, Linyang He, Xuran Zhao |
Neural Comput. Appl. | 3 |
| 2025 | DPMNet: A Remote Sensing Forest Fire Real-Time Detection Network Driven by Dual Pathways and Multidimensional Interactions of FeaturesabstractA fundamental challenge in remote sensing-based forest fire detection lies in accurately discerning fire characteristics on various scales against the backdrop of intricate and heterogeneous forest landscapes. In response to this challenge, we propose a dual-path network (DPMNet) with multidimensional feature interaction for real time remote sensing forest fire detection. Initially, a dual-path backbone network is designed, integrating coarse-grained and fine-grained parallel pathways, working in tandem to capture both global visual features and nuanced local texture details. Subsequently, we develop the Multidimensional Interactive Feature Pyramid Network (MiFPN), a novel structure that amalgamates information streams from varied levels through a three-branch structure and engenders profound fusion and dynamic interaction of features across multiple scales. Thereafter, the Context-Enriched Adaptive Fusion Module (CEAFM) is proposed, which emerges to meticulously blend macroscopic visual elements harvested via coarse-grained conduits, employing a multi-faceted pathway strategy to bolster the model’s overarching comprehension and precision in forest fire detection. Finally, the Enhanced Contextual Pooling Bottleneck (ECPB) is put forward, an integration that augments the model’s spatial perception and contextual acumen through the incorporation of dilated convolution and global pooling techniques. Extensive experiments are conducted on the remote sensing forest fire dataset in order to confirm the efficacy of DPMNet. The experimental results demonstrate that our DPMNet achieves satisfactory performance in terms of real-time performance as well as accuracy and provides an effective solution for real-time detection of remote sensing forest fires based on UAVs. Guanbo Wang, Victor S. Sheng, Yujun Ma, Hongwei Ding 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Convolutional transformer network for fine-grained action recognition
Yujun Ma, Ruili Wang 0001, Ming Zong, Wanting Ji, Yi Wang 0037 |
Neurocomputing | 1 |
| 2024 | k-NN attention-based video vision transformer for action recognitionabstractAction Recognition aims to understand human behavior and predict a label for each action. Recently, Vision Transformer (ViT) has achieved remarkable performance on action recognition, which models the long sequences token over spatial and temporal index in a video. The fully-connected self-attention layer is the fundamental key in the vanilla Transformer. However, the redundant architecture of the vision Transformer model ignores the locality of video frame patches, which involves non-informative tokens and potentially leads to increased computational complexity. To solve this problem, we propose a k-NN attention-based Video Vision Transformer (k-ViViT) network for action recognition. We adopt k-NN attention to Video Vision Transformer (ViViT) instead of original self-attention, which can optimize the training process and neglect the irrelevant or noisy tokens in the input sequence. We conduct experiments on the UCF101 and HMDB51 datasets to verify the effectiveness of our model. The experimental results illustrate that the proposed k-ViViT achieves superior accuracy compared to several state-of-the-art models on these action recognition datasets. Weirong Sun, Yujun Ma, Ruili Wang 0001 |
Neurocomputing | 2 |
| 2024 | Relative-position embedding based spatially and temporally decoupled Transformer for action recognition
Yujun Ma, Ruili Wang 0001 |
Pattern Recognit. | 1 |
| 2023 | Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture RecognitionabstractRGB-D action and gesture recognition remain an interesting topic in human-centered scene understanding, primarily due to the multiple granularities and large variation in human motion. Although many RGB-D based action and gesture recognition approaches have demonstrated remarkable results by utilizing highly integrated spatio-temporal representations across multiple modalities (i.e., RGB and depth data), they still encounter several challenges. Firstly, vanilla 3D convolution makes it hard to capture fine-grained motion differences between local clips under different modalities. Secondly, the intricate nature of highly integrated spatio-temporal modeling can lead to optimization difficulties. Thirdly, duplicate and unnecessary information can add complexity and complicate entangled spatio-temporal modeling. To address the above issues, we propose an innovative heuristic architecture called Multi-stage Factorized Spatio-Temporal (MFST) for RGB-D action and gesture recognition. The proposed MFST model comprises a 3D Central Difference Convolution Stem (CDC-Stem) module and multiple factorized spatio-temporal stages. The CDC-Stem enriches fine-grained temporal perception, and the multiple hierarchical spatio-temporal stages construct dimension-independent higher-order semantic primitives. Specifically, the CDC-Stem module captures bottom-level spatio-temporal features and passes them successively to the following spatio-temporal factored stages to capture the hierarchical spatial and temporal features through the Multi-Scale Convolution and Transformer (MSC-Trans) hybrid block and Weight-shared Multi-Scale Transformer (WMS-Trans) block. The seamless integration of these innovative designs results in a robust spatio-temporal representation that outperforms state-of-the-art approaches on RGB-D action and gesture recognition datasets. Yujun Ma, Benjia Zhou, Ruili Wang 0001, Pichao Wang |
ACM Multimedia | 1 |
| 2023 | Multi-head attention-based two-stream EfficientNet for action recognitionabstractAbstract Recent years have witnessed the popularity of using two-stream convolutional neural networks for action recognition. However, existing two-stream convolutional neural network-based action recognition approaches are incapable of distinguishing some roughly similar actions in videos such as sneezing and yawning. To solve this problem, we propose a Multi-head Attention-based Two-stream EfficientNet (MAT-EffNet) for action recognition, which can take advantage of the efficient feature extraction of EfficientNet. The proposed network consists of two streams (i.e., a spatial stream and a temporal stream), which first extract the spatial and temporal features from consecutive frames by using EfficientNet. Then, a multi-head attention mechanism is utilized on the two streams to capture the key action information from the extracted features. The final prediction is obtained via a late average fusion, which averages the softmax score of spatial and temporal streams. The proposed MAT-EffNet can focus on the key action information at different frames and compute the attention multiple times, in parallel, to distinguish similar actions. We test the proposed network on the UCF101, HMDB51 and Kinetics-400 datasets. Experimental results show that the MAT-EffNet outperforms other state-of-the-art approaches for action recognition. Aihua Zhou, Yujun Ma, Wanting Ji, Ming Zong, Mingzhe Liu 0001 |
Multim. Syst. | 2 |
| 2022 | Spatial-temporal interaction learning based two-stream network for action recognition
Yujun Ma, Wenhan Yang, Wanting Ji, Ruili Wang 0001 |
Inf. Sci. | 2 |
| 2021 | Medical-Level Suicide Risk Analysis: A Novel Standard and Evaluation ModelabstractThe frequent occurrence of suicides in modern society constitutes a serious public health issue. While the motives, methods, and consequences of suicide are quite complicated, if people at risk of suicide can be identified and intervened in time, the loss of life can be reduced. Through analyses based on combining a large number of suicide texts and professional medical literature, a dictionary of potential suicide risk impact factors has been established in this article. Based on this dictionary, a novel medical-level suicide risk standard is proposed to monitor suicide risk from point-to-surface under the timeline baseline. In order to solve the problem of insufficient Chinese suicide data sets, the manually assisted method based on knowledge perception is adopted to annotate the data set with corresponding to risk level. At the same time, a Bert evaluation model based on knowledge perception was established for the classification of risk level. The experimental results showed that proposed method has a 56% recognition accuracy in the prediction of 10-Label suicide risk level proposed in this article, and the classification performance is better than traditional machine learning algorithms. Therefore, the results showed that the classification standard and evaluation model can be effectively used for the identification and early warning of suicide risk, which can discover high suicide risk groups to reduce the occurrence of suicide. It is of great significance to people’s emotion care monitoring. Rui Wang 0077, Bing Xiang Yang, Yujun Ma, Qiao Yu 0002, Xiaofen Zong, Simeng Ma, Long Hu, Kai Hwang 0001, Zhongchun Liu |
IEEE Internet Things J. | 3 |
| 2021 | Multi-cue based four-stream 3D ResNets for video-based action recognition
Ming Zong, Yujun Ma, Wanting Ji, Mingzhe Liu 0001, Ruili Wang 0001 |
Inf. Sci. | 4 |
| 2021 | Optimal Location Privacy Preserving and Service Quality Guaranteed Task Allocation in Vehicle-Based Crowdsensing NetworksabstractWith increasing popularity of related applications of mobile crowdsensing, especially in the field of Internet of Vehicles (IoV), task allocation has attracted wide attention. How to select appropriate participants is a key problem in vehicle-based crowdsensing networks. Some traditional methods choose participants based on minimizing distance, which requires participants to submit their current locations. In this case, participants' location privacy is violated, which influences disclosure of participants' sensitive information. Many privacy preserving task allocation mechanisms have been proposed to encourage users to participate in mobile crowdsensing. However, most of them assume that different participants' task completion quality is the same, which is not reasonable in reality. In this paper, we propose an optimal location privacy preserving and service quality guaranteed task allocation in vehicle-based crowdsensing networks. Specifically, we utilize differential privacy to preserve participants' location privacy, where every participant can submit the obfuscated location to the platform instead of the real one. Based on the obfuscated locations, we design an optimal problem to minimize the moving distance and maximize the task completion quality simultaneously. In order to solve this problem, we decompose it into two linear optimization problems. We conduct extensive experiments to demonstrate the effectiveness of our proposed mechanism. Yongfeng Qian, Yujun Ma, Jing Chen 0003, Di Wu 0001, Daxin Tian, Kai Hwang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Cognitive Wearable Robotics for Autism Perception EnhancementabstractAutism spectrum disorder (ASD) is a serious hazard to the physical and mental health of children, which limits the social activities of patients throughout their lives and places a heavy burden on families and society. The developments of communication techniques and artificial intelligence (AI) have provided new potential methods for the treatment of autism. The existing treatment systems based on AI for children with ASD focus on detecting health status and developing social skills. However, the contradiction between the terminal interaction capability and availability cannot meet the needs for real application scenarios. At the same time, the lack of diverse data cannot provide individualized care for autistic children. To explore this robot-based approach, a novel AI-based first-view-robot architecture is proposed in this article. By providing care from the first-person perspective, the proposed wearable robot overcomes the difficulty of the absence of cognitive ability in the third-view of traditional robotics and improves the social interaction ability of children with ASD. The first-view-robot architecture meets the requirements of dynamic, individualized, and highly immersed interaction services for autistic children. First, the multi-modal and multi-scene data collection processes of standard, static, and dynamic datasets are introduced in detail. Then, to comprehensively evaluate the learning ability of children with ASD through mental states and external performances, a learning assessment model with emotion correction is proposed. Besides, a wearable robot-assisted environment perception and expression enhancement mechanism for children with ASD is realized by reinforcement learning, which can be adapted to interactive environments with optimal action policies. An interactive testbed for children with ASD treatments is demonstrated and experimental cases for test subjects are presented. Last, three open issues are discussed from data processing, robot designing, and service responding perspectives. Min Chen 0003, Wenjing Xiao, Long Hu, Yujun Ma, Yin Zhang 0002, Guangming Tao |
ACM Trans. Internet Techn. | 4 |
| 2016 | Green data center with IoT sensing and cloud-assisted smart temperature control system
Yujun Ma, Musaed Alhussein, Yin Zhang 0002, Limei Peng |
Comput. Networks | 2 |
| 2016 | Smart Clothing: Connecting Human with Clouds and Big Data for Sustainable Health Monitoring
Min Chen 0003, Yujun Ma, Jeungeun Song 0001, Chin-Feng Lai, Bin Hu 0001 |
Mob. Networks Appl. | 2 |
| 2016 | CGMP: cloud-assisted green multimedia processing
Yujun Ma, Yin Zhang 0002, Zhengguo Sheng, Ruan Hang |
Multim. Tools Appl. | 1 |
| 2015 | Android-based intelligent mobile robot for indoor healthcareabstractThe intelligent mobile robot platform based on Android features multiple wireless communication functions, which, via Internet and WLAN, may control the movement of the robot. The platform integrates cloud speech recognition and offline speech recognition technologies to control the movement of the robot and have simple man-machine conversation. In addition, with a camera, this platform may realize the remote real-time transmission of video. Since the platform is characterized with sound hardware compatibility and expansibility, we may conduct the research of the robot and rapidly develop an intelligent mobile robot which applies to a specific application scenario on the platform. Yujun Ma, Dengming Xiao, Ruan Hang, Junlong Zhao, Yin Zhang 0002 |
HealthCom | 1 |
| 2015 | PWDGR: Pair-Wise Directional Geographical Routing Based on Wireless Sensor NetworkabstractMultipath routing in wireless multimedia sensor network makes it possible to transfer data simultaneously so as to reduce delay and congestion and it is worth researching. However, the current multipath routing strategy may cause problem that the node energy near sink becomes obviously higher than other nodes which makes the network invalid and dead. It also has serious impact on the performance of wireless multimedia sensor network (WMSN). In this paper, we propose a pair-wise directional geographical routing (PWDGR) strategy to solve the energy bottleneck problem. First, the source node can send the data to the pair-wise node around the sink node in accordance with certain algorithm and then it will send the data to the sink node. These pair-wise nodes are equally selected in 360° scope around sink according to a certain algorithm. Therefore, it can effectively relieve the serious energy burden around Sink and also make a balance between energy consumption and end-to-end delay. Theoretical analysis and a lot of simulation experiments on PWDGR have been done and the results indicate that PWDGR is superior to the proposed strategies of the similar strategies both in the view of the theory and the results of those simulation experiments. With respect to the strategies of the same kind, PWDGR is able to prolong 70% network life. The delay time is also measured and it is only increased by 8.1% compared with the similar strategies. Yin Zhang 0002, Jialun Wang, Yujun Ma, Min Chen 0003 |
IEEE Internet Things J. | 4 |
| 2013 | Enabling comfortable sports therapy for patient: A novel lightweight durable and portable ECG monitoring systemabstractIn developing countries, people's living pressure is increasing with the society's development by inefficient economic growth mode. Recently, the number of people who suffer from sudden cardiac death is progressively increasing, and cardiovascular disease (CVD) has become one great killer which threats the life and health of people. However, at present, we are confronted with one problem: when a patient has chest distress or chest pain, etc., he/she hurries to the hospital to go through electrocardiograph (ECG) examination but the abnormal ECG signal disappears. Therefore, the opportunity to timely capture the ECG status of a patient and make an accurate judgment is lost. Thus, extensive efforts have been made to design various systems for patient monitoring at anytime and anywhere, in order to have real-time records and analysis on vital signal of patents, so as to provide early detection before the occurrence of adverse effect. However, the mobility of the patient is limited in most existing healthcare systems. While sporting is beneficial to improve patient's health, designing a comfortable and durable healthcare system for facilitating patient's movement is a critical issue. This paper presents a novel comfortable and durable portable ECG monitoring system to have real-time monitoring and analysis on a moving user. In the meantime, its special low power and on-demand data collection design alleviates the problem that the current wearable ECG monitoring equipment could not be used for a long time due to the constraint of its battery life. Min Chen 0003, Yujun Ma, Jialun Wang, Ong Mau Dung, Enmin Song |
Healthcom | 2 |