Junxiao Xue

dblp:38/1915 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0003-1569-5362ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Space Computing Constellation: System Architecture, Implementations, and Challenges
abstract
Low Earth Orbit (LEO) satellite constellations have experienced rapid growth in recent years, driven by their potential to deliver global, high-bandwidth Internet services with low latency. Beyond connectivity, LEO constellations also offer promising opportunities to enable in-orbit processing of space-native data to support a wide range of emerging space applications. In this context, the concept of space computing has been proposed, a paradigm that seamlessly integrates networking and computing to provide computing-as-a-service anytime and anywhere in space. However, the inherent characteristics of satellite constellations, such as dynamic network topologies, constrained system resources, and the harsh space environment, pose significant challenges in achieving this vision. This paper outlines the system architecture and the key enabling technologies for space computing, including spaceborne computers, laser communications, spaceborne router, distributed operating systems, and onboard AI. We also present the implementation of an open space computing platform, the 3-Body computing constellation, along with the in-orbit experimental results that demonstrate the advantages of multi-satellite distributed computing. Furthermore, we outline future research directions essential for advancing toward a truly interconnected, autonomous, and intelligent space computing system.
Hua Wang 0011, Kelu Yao, Luqi Gong, Yichao Jin 0001, Yuan Liu 0030, Junxiao Xue, Zhiguo Wan, Chao Li 0028, Zhifeng Zhao
IEEE Internet Things J.8
2026 A Benchmark of Microvideos for Public Opinion Analysis
Junxiao Xue, Peifu Yang, Bin Wu 0019
IEEE Trans. Comput. Soc. Syst.1
2025 Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
Rui Zhang 0055, Shuailong Li, Junxiao Xue, Diying Yan, Xiaoran Yan
CGI (2)3
2025 AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised Learning
abstract
Affective Video Facial Analysis (AVFA) is important for advancing emotion-aware AI, yet the persistent data scarcity in AVFA presents challenges. Recently, the self-supervised learning (SSL) technique of Masked Autoencoders (MAE) has gained significant attention, particularly in its audio-visual adaptation. Insights from general domains suggest that scaling is vital for unlocking impressive improvements, though its effects on AVFA remain largely unexplored. Additionally, capturing both intra- and inter-modal correlations through scalable representations is a crucial challenge in this field. To tackle these gaps, we introduce AVF-MAE++, a series audio-visual MAE designed to explore the impact of scaling on AVFA with a focus on advanced correlation modeling. Our method incorporates a novel audio-visual dual masking strategy and an improved modality encoder with a holistic view to better support scalable pre-training. Furthermore, we propose the Iteratively Audio-Visual Correlations Learning Module to improve correlations capture within the SSL framework, bridging the limitations of prior methods. To support smooth adaptation and mitigate overfitting, we also introduce a progressive semantics injection strategy, which structures training in three stages. Extensive experiments across 17 datasets, spanning three key AVFA tasks, demonstrate the superior performance of AVFMAE++, establishing new state-of-the-art outcomes. Ablation studies provide further insights into the critical design choices driving these gains. Code is released at this URL.
Heli Sun, Jiayu Nie, Junxiao Xue, Liang He 0006
CVPR7
2025 InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
abstract
Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to variability across sequences and the uneven distribution of information within each sequence. To tackle this, we introduce InfoSyncNet, a non-uniform sequence modeling network enhanced by tailored data augmentation techniques. Central to InfoSyncNet is a non-uniform quantization module positioned between the encoder and decoder, enabling dynamic adjustment to the network’s focus and effectively handling the natural inconsistencies in visual speech data. Additionally, multiple training strategies are incorporated to enhance the model’s capability to handle variations in lighting and the speaker’s orientation. Comprehensive experiments on the LRW and LRW1000 datasets confirm the superiority of InfoSyncNet, achieving new state-of-the-art accuracies of 92.0% and 60.7% Top-1 ACC1.
Junxiao Xue, Fei Yu 0012
IJCNN1
2025 Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
abstract
Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs. Considering the lack of SVs emotion data, we introduce a large-scale dataset named eMotions, comprising 27,996 videos. Meanwhile, we alleviate the impact of subjectivities on labeling quality by emphasizing better personnel allocations and multi-stage annotations. In addition, we provide the category-balanced and test-oriented variants through targeted data sampling. Some commonly used videos, such as facial expressions, have been well studied. However, it is still challenging to analysis the emotions in SVs. Since the broader content diversity brings more distinct semantic gaps and difficulties in learning emotion-related features, and there exists local biases and collective information gaps caused by the emotion inconsistence under the prevalently audio-visual co-expressions. To tackle these challenges, we present an end-to-end audio-visual baseline AV-CANet which employs the video transformer to better learn semantically relevant representations. We further design the Local-Global Fusion Module to progressively capture the correlations of audio-visual features. The EP-CE Loss is then introduced to guide model optimization. Extensive experimental results across seven datasets demonstrate the effectiveness of AV-CANet, while providing broad insights for future works. Besides, we explore the key components of AV-CANet by ablation studies. Datasets and code are released at https://github.com/XuecWu/eMotions.
Heli Sun, Junxiao Xue, Jiayu Nie, Xiangyan Kong, Ruofan Zhai, Danlei Huang, Liang He 0006
ICMR3
2025 HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
abstract
Advances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set.
Heli Sun, Danlei Huang, Xinyi Yin, Hao Wang 0182, Jia Zhang 0016, Fei Wang 0128, Peihao Guo, Suyu Xing, Junxiao Xue, Liang He 0006
ACM Multimedia11
2025 3A-YOLO : New Real-Time Object Detectors with Triple Discriminative Awareness and Coordinated Representations
abstract
Recent research on general real-time object detectors (e.g., YOLO series) has demonstrated the effectiveness of attention mechanisms for elevating model performance. Nevertheless, existing methods often neglect to unifiedly deploy hierarchical attention mechanisms to construct a more discriminative YOLO head which is enriched with more useful intermediate features. To tackle these gaps, this work aims to leverage multiple attention mechanisms to hierarchically enhance the triple discriminative awareness of the YOLO detection head and complementarily learn the coordinated intermediate representations, resulting in a new series detector 3A-YOLO for the general scenarios. Specifically, taking YOLOv4 and YOLOv8 as baselines, we first propose a new detection head denoted TDA-YOLO Module, which unifiedly enhance the representations learning capability of scale-awareness, spatial-awareness, and task-awareness. Secondly, we steer the intermediate features to coordinatedly learn the inter-channel relationships and precise positional information. Finally, we perform neck network improvements followed by introducing various tricks to boost the adaptability of our 3A-YOLO. Extensive experiments across COCO and VOC benchmarks indicate the effectiveness of our detectors. Ablation studies provide further insights into the critical design choices driving the performance gains.
Junxiao Xue, Liangyu Fu, Jiayu Nie, Danlei Huang, Xinyi Yin, Tingqi Hu
SMC2
2025 Orientation control via G3 parametric quaternion interpolation spline curves
Aihua Liang, Junxiao Xue
Comput. Graph.5
2025 Improved DDPG based on enhancing decision evaluation for path planning in high-density environments
Junxiao Xue, Mengyang He, Jinpu Chen, Bowei Dong, Yuanxun Zheng
Expert Syst. Appl.1
2025 Efficient Resource Allocation in Computing Power Networks Considering Similar Task Merging: A Lyapunov Optimization-Based DRL Approach
abstract
The cloud-edge–terminal architecture relies on hierarchy for resource allocation but lacks global optimization. The computing power network (CPN) introduces a new distributed computing paradigm, integrating cross-domain, heterogeneous resources for global scheduling. However, most CPN research focuses on task optimization during resource allocation, while neglecting the similarity of random tasks before the allocation stage. Additionally, fragmented CPN resources and complex task demands pose challenges to global load balancing. This article proposes a deep reinforcement learning framework with task merging and congestion avoidance for on-demand resource allocation. Specifically, a low-complexity similar task merging algorithm reduces redundant resource consumption during task preprocessing. In task offloading, the principal neighborhood aggregated graph neural network captures CPN’s intricate features. Lyapunov optimization, integrated into a multithreaded training framework, minimizes resource backlog congestion. A carefully designed reward function balances multiple objectives, enhancing computing resource utilization efficiency and ensuring system stability. Theoretical analysis shows that with control parameter V, the tradeoff between resource utilization efficiency and system stability follows the relationship [O(1/V), O(V)]. Extensive experiments demonstrate a 33.5% improvement in resource utilization efficiency and a 62.7% increase in task offloading success rates with respect to those in state-of-the-art algorithms. The proposed algorithm exhibits robustness and effectiveness, particularly in high-load and real network topologies.
Zhonghai Jia, Junxiao Xue, Lei Shi 0001, Jie Li 0002, Mengyang He
IEEE Internet Things J.2
2024 Enhancing Human Action Recognition with Fine-grained Body Movement Attention
abstract
In the field of vision-language models (VLMs), human action recognition models, while effective, always rely on large pre-trained models or high-resolution inputs, leading to computational challenges. To address this, we propose a novel VLM approach with fine-grained attention to body movements. Unlike methods relying on coarse video-text matching, we guide the model to infer actions from fine-grained body part movements using two techniques: fine-tuning pre-trained encoders at the fine-grained level and matching labels from language and vision perspectives at the coarse-grained level. Experiments show our model excels in fully-supervised, few-shot, and zero-shot scenarios with just 8 random frames and a ViT-B/32 backbone. It outperforms most ViT-L/14 based models, demonstrating effectiveness while saving computational resources. The largest Top-1 accuracy improvement over second-best approaches is 6.8%.
Rui Zhang 0055, Junxiao Xue, Pavel Smirnov 0005, Xiaoran Yan
ICME2
2024 Resolving the Resource Decision-Making Dilemma of Leaderless Group-Based Multiagent Systems and Repeated Games
abstract
Leaderless rational individuals often lead the group into a resource decision dilemma in resource competition. Reducing the cost of resource competition while avoiding group decision dilemmas is a challenging task. Inspired by multiagent systems (MASs) and repeated games, we propose a decision-making reward discrimination (DRD) framework to address the resource competition dilemma of leaderless group formation. We aim to model the leaderless group’s resource gaming process using MAS and achieve optimal rewards for the group while minimizing conflict in resource competition. The proposed framework consists of three modules: 1) the decision-making module; 2) the reward module; and 3) the discriminative module. The decision-making module defines the agents and models the decision-making process, while the reward module calculates the group reward in each round using the reward matrix. The discriminative module compares the group reward with the target reward while providing the agent with environmental information. We verify the feasibility of the model through numerous experiments. The results show that agents adopt a revenge strategy to avoid resource competition dilemmas and achieve group reward optimality.
Junxiao Xue, Mingchuang Zhang, Bowei Dong, Lei Shi 0001, Andrés Adolfo Navarro Newball
IEEE Trans. Syst. Man Cybern. Syst.1
2023 A Multimodal Spatio-temporal Model for Micro-Video Emotion Classification
abstract
Spatio-temporal features are information unique to video compared to images, and are an important source of features for video content analysis. The introduction of spatio-temporal features has brought a significant change to affective computing, which is no longer limited to single-image and ignores temporal affective changes. However, existing affective computing models still do not introduce spatio-temporal features. For example, facial emotion classification still uses single-image as input to the model and ignores spatio-temporal features. Therefore, we propose a novel neural network that takes video as the input of the model and extracts feature information containing spatio-temporal features in it. The extracted feature information consists of two modalities, which are audio spatio-temporal feature information and image spatio-temporal feature information. Decision-level fusion of the two modal information is performed to achieve emotion classification. Where a feature extractor based on VGG implementation is used as the audio emotion feature extraction module, which extracts the spectrogram feature information while focusing on its temporal feature information. And using the feature extractor implemented based on I3D as the image emotion feature extraction module, which consists of 3D convolutional blocks. Compared with 2D convolutional blocks, it not only focuses on image features, but also performs convolutional extraction for both spatial features and temporal features of the images within the blocks to obtain image features containing spatio-temporal features. By fusing multiple spatio-temporal feature classifiers in the best way, our model achieves an accuracy of 62.41%. The experimental results show that feature information containing spatio-temporal features is more effective in sentiment classifiers than feature information containing only image features.
Junxiao Xue, Yuanxun Zheng
MDM1
2023 Fine-grained sequence-to-sequence lip reading based on self-attention and self-distillation
Junxiao Xue, Shibo Huang, Huawei Song, Lei Shi 0001
Frontiers Comput. Sci.1
2023 Cross-modal information fusion for voice spoofing detection
Junxiao Xue, Huawei Song, Bin Wu 0019, Lei Shi 0001
Speech Commun.1
2023 Efficient Top-k Matching for Publish/Subscribe Ride Hitching
abstract
With the continued proliferation of mobile Internet and geo-locating technologies, carpooling as a green transport mode is widely accepted and becoming tremendously popular worldwide. In this paper, we focus on a popular carpooling service calledride hitching, which is typically implemented using a publish/subscribe approach. In a ride hitching service, drivers subscribe ride orders published by riders and continuously receive matching ride orders until one is picked. The current systems (e.g., Didi Hitch) adopt a threshold-based approach to filter ride orders. That is, a new ride order will be sent to all subscribing drivers whose planned trips can match the ride order within a pre-defined detour threshold. A limitation of this approach is that it is difficult for drivers to specify a reasonable detour threshold in practice. In addressing this problem, we propose a novel type of top-$k$subscription queries calledTop-$k$kRideSubscription (TkRS)query, which continuously returns the best$k$ride orders that match drivers’ trip plans to them. We propose two efficient algorithms to enable the top-$k$result maintenance. We also design a novel hybrid grid index and a two-level buffer structure to efficiently track the top-$k$results for allTkRSqueries. Finally, extensive experiments on real-life datasets suggest that our proposed algorithms are capable of achieving desirable performance in practical settings.
Hongyan Gu, Rui Chen 0012, Jianliang Xu, Shangwei Guo, Junxiao Xue, Mingliang Xu 0001
IEEE Trans. Knowl. Data Eng.6
2022 Naturalistic Driving Scenario Recognition with Multimodal Data
abstract
Driving Scenario recognition is one fundamental technology of automated driving systems or advanced driver assistance systems. A common practice of driving scenario recog-nition is to conduct classification tasks with the data collected by in-vehicle data acquisition system or driving simulator. In most existing works, visual data were used since the relevant information of driving scenarios is usually inferable from their visual appearance. However, the non-visual information, e.g. physiological state of the driver, also provide complementary information for the scenario recognition task especially when the visual appearance is insufficient to differentiate similar naturalistic scenarios in some cases. In this paper, we propose a hybrid driving scenario recognition model with multimodal input. The model consists of a convolutional neural network based visual data sub-model, a stacked autoencoder based physiological data sub-model, and a fusion sub-model that combines the extracted features from both visual and physiological sub-models. Besides, a post-processing is adopted to correct the recognition results of some ambiguous scenarios. Experimental results on the trip data collected in the naturalistic driving context demonstrated the effectiveness of the proposed method.
Junxiao Xue, Hao Liu 0125
MDM5
2022 Physiological-physical feature fusion for automatic voice spoofing detection
Junxiao Xue
Frontiers Comput. Sci.1
2022 Agent-Based Campus Novel Coronavirus Infection and Control Simulation
abstract
Corona Virus Disease 2019 (COVID-19), due to its extremely high infectivity, has been spreading rapidly around the world and bringing huge influence to socioeconomic development and people’s daily life. Taking for example the virus transmission that may occur after college students return to school, we analyze the quantitative influence of the key factors on the virus spread, including crowd density and self-protection. One Campus Virus Infection and Control Simulation (CVICS) model of the novel coronavirus is proposed in this article, fully considering the characteristics of repeated contact and strong mobility of crowd in the closed environment. Specifically, we build an agent-based infection model, introduce the mean field theory to calculate the probability of virus transmission, and microsimulate the daily prevalence of infection among individuals. The experimental results show that the proposed model in this article efficiently simulates how the virus spreads in the dense crowd in frequent contact under a closed environment. Furthermore, preventive and control measures, such as self-protection, crowd decentralization, and isolation during the epidemic, can effectively delay the arrival of infection peak, reduce the prevalence, and, finally, lower the risk of COVID-19 transmission after the students return to school.
Pei Lv, Boya Xu, Ran Feng, Chaochao Li, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Comput. Soc. Syst.6
2022 Modeling the Impact of Social Distancing on the COVID-19 Pandemic in a Low Transmission Setting
abstract
According to the World Health Organization and the CDC, social distancing is currently one of the most effective ways to slow the transmission of COVID-19. However, most existing epidemic models do not consider the impact of social distancing on the COVID-19 pandemic. In this article, we propose a new method to deterministic modeling of the effects of social distancing on the COVID-19 pandemic in a low transmission setting. Our model dynamic is expressed by a single predictive variable that satisfies an integro-differential equation. Once the dynamic variable is calculated, the process of agents from the normal state, infection state to rehabilitation state, or death state can be explored. Besides, an important parameter is added to the model to measure the impact of social distancing on epidemic transmission. We performed qualitative and quantitative experiments on various scenarios, and the results showed that 2 m is a safe social distancing on the COVID-19 pandemic in a low transmission setting.
Junxiao Xue, Mingchuang Zhang, Mingliang Xu 0001
IEEE Trans. Comput. Soc. Syst.1
2021 Detecting fake news by exploring the consistency of multimodal data
Junxiao Xue, Yichen Tian, Lei Shi 0001
Inf. Process. Manag.1
2019 Crowd queuing simulation with an improved emotional contagion model
Junxiao Xue, Pei Lv, Mingliang Xu 0001
Sci. China Inf. Sci.1
2019 Disassembling a 3D mechanism for efficient packing
abstract
Abstract This paper introduces a disassemble‐and‐pack algorithm to disassemble a mechanical 3D model in groups that can be efficiently packed within a box, with the objective of reassembling them easily after delivery. Its key feature is that, mostly, the mechanism can be disassembled at the joint and each part can be an adjusted motion structure based on its joint type. Our system consists of two steps: disassembling the mechanical object into a group set and packing them within a box efficiently. The first step consists in the creation of a hierarchy of possible group set of parts that can be tightly packed within their minimum bounding boxes. Use the breadth‐first search algorithm to traverse the hierarchy of possible group set in order to disconnect the joints and get the group set. In the second step, according to the reverse order of volume, each group in the set is inserted into the specified box. The fact that mechanism disassembly and shape packing are both an NP‐complete problem justifies finding approximated solutions according to efficacy and efficiency. Experimental results show that our approach can really efficiently pack a range of mechanisms from a simple model to complex objects.
Xiaoheng Jiang, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001
Comput. Animat. Virtual Worlds5
2019 Wheat ear growth modeling based on a polygon
abstract
Visual inspection of wheat growth has been a useful tool for understanding and implementing agricultural techniques and a way to accurately predict the growth status of wheat yields for economists and policy decision makers. In this paper, we present a polygonal approach for modeling the growth process of wheat ears. The grain, lemma, and palea of wheat ears are represented as editable polygonal models, which can be re-polygonized to detect collision during the growth process. We then rotate and move the colliding grain to resolve the collision problem. A linear interpolation and a spherical interpolation are developed to simulate the growth of wheat grain, performed in the process of heading and growth of wheat grain. Experimental results show that the method has a good modeling effect and can realize the modeling of wheat ears at different growth stages.
Junxiao Xue, Chen-yang Sun, Junjin Cheng, Mingliang Xu 0001, Shui Yu 0001
Frontiers Inf. Technol. Electron. Eng.1
2017 Mechanical Assembly Packing Problem Using Joint Constraints
Mingliang Xu 0001, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003
J. Comput. Sci. Technol.5
2008 Layered deformation of solid model using conformal mapping
Zhongxuan Luo, Junxiao Xue
Comput. Graph.2
2007 3D Model Deformation via Conformal Mapping
abstract
In this paper, we introduce conformal mapping to 3D model deformation. The surface of the solid model is represented by base-patch and cross-section function. Accordingly, the shape of the 3D model can be modified by interactive means, such as changing the corresponding base-patch with conformal mapping and adjusting a cross-section function. There are two features of the proposed method. First, due to the properties of conformal mapping, the scheme satisfies the local demands on rigidity and still allows for considerable global deformation. Second, the deformation effect is predictable and the transformation function of the deformation can be expressed analytically. The proposed transformation of 3D deformation is continuous, not self- intersection, and topologically consistent. Some examples are presented in the end of the paper to illustrate the efficiency of our approach.
Zhongxuan Luo, Junxiao Xue
CAD/Graphics2