VLDB 2026 Research / reviewers in the wild / expert
Zixuan Tang
dblp:304/1309
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A power-aware graph domain adaptation framework for variable-power fault detection of self-powered neutron detectors in nuclear power plants
Zixuan Tang, Aidong Xu, Bingjun Yan, Mingxu Gang |
Expert Syst. Appl. | 1 |
| 2026 | STAR: Skeletal Token Alignment and Rearrangement for Interaction RecognitionabstractUnderstanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing methods rely on skeleton sequences-effective in low-light and privacy-sensitive environments-they face two major challenges: 1) learning and effectively exploiting interaction cues from skeletal data, and 2) compensating for the lack of visual information absent in skeletons alone. To address these challenges, we propose skeletal token alignment and rearrangement (STAR) for human-robot and human-human interaction recognition. It learns interaction-specific skeleton features and enriches them using visual cues by aligning skeleton and RGB video representations in a shared latent space. Specifically, STAR consists of three key components. First, we design a skeleton encoder that captures fine-grained interdependencies using Entity Rearrangement (ER) and Interactive Spatiotemporal Tokens (ISTs). Second, we present Visual Interaction Encoding that introduces a Focus on Interactions (FoI) strategy to attend to spatiotemporal regions relevant to interactions in RGB videos. Finally, these representations are aligned via a contrastive learning objective, with a refinement head further refines predictions. During training, STAR leverages both skeleton and RGB video data to learn robust, discriminative interaction representations. At inference time, it operates on skeletons alone, retaining visual-informed benefits while preserving skeleton-only efficiency. Extensive experiments on Chico, HARPER, NTU Mutual 11 and 26 datasets consistently validate our approach by demonstrating superior performance over state-of-the-art methods. Our code is publicly available athttps://github.com/Necolizer/STAR. Yuhang Wen 0001, Mengyuan Liu 0001, Zixuan Tang, Junsong Yuan 0001, Beichen Ding |
IEEE Trans. Multim. | 3 |
| 2025 | A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive LearningabstractGraph-based models have emerged as a powerful paradigm for modeling multimodal urban data and learning region representations for various downstream tasks. However, existing approaches face two major limitations. (1) They typically employ identical graph neural network architectures across all modalities, failing to capture modality-specific structures and characteristics. (2) During the fusion stage, they often neglect spatial heterogeneity by assuming that the aggregation weights of different modalities remain invariant across regions, resulting in suboptimal representations. To address these issues, we propose MTGRR, a modality-tailored graph modeling framework for urban region representation, built upon a multimodal dataset comprising point of interest (POI), taxi mobility, land use, road element, remote sensing, and street view images. (1) MTGRR categorizes modalities into two groups based on spatial density and data characteristics: aggregated-level and point-level modalities. For aggregated-level modalities, MTGRR employs a mixture-of-experts (MoE) graph architecture, where each modality is processed by a dedicated expert GNN to capture distinct modality-specific characteristics. For the point-level modality, a dual-level GNN is constructed to extract fine-grained visual semantic features. (2) To obtain effective region representations under spatial heterogeneity, a spatially-aware multimodal fusion mechanism is designed to dynamically infer region-specific modality fusion weights. Building on this graph modeling framework, MTGRR further employs a joint contrastive learning strategy that integrates region aggregated-level, point-level, and fusion-level objectives to optimize region representations. Experiments on two real-world datasets across six modalities and three tasks demonstrate that MTGRR consistently outperforms state-of-the-art baselines, validating its effectiveness. Yaya Zhao, Kaiqi Zhao 0001, Zixuan Tang, Xiaoling Lu, Yalei Du |
ECAI | 3 |
| 2025 | Data Synchronization and Redundancy Mechanism for Virtual PLCs in Industrial Control SystemsabstractVirtual Programmable Logic Controllers (vPLCs), as a newborn technology, are becoming increasingly important in modern industrial automation due to their flexibility and scalability. There is lack of researches on data synchronization and redundancy mechanisms for vPLCs, limiting applications of vPLCs in critical industrial scenarios. This paper designs and implements a data synchronization and redundancy mechanism between vPLCs based on heartbeat detection to enhance the reliability of vPLC systems. The mechanism continuously monitors for failures and synchronizes data between vPLCs to ensure seamless control task takeover in the event of a failure. Experimental results demonstrate the mechanism’s high effectiveness in fault detection and recovery, achieving a redundancy switchover time that meets industrial application requirements. Zixuan Tang, Dong Li 0009, Yu Liu 0011, Dapeng Lan, Peng Bo 0004, Zhibo Pang |
INDIN | 1 |
| 2025 | MedSoft-Diffusion: Medical Semantic-Guided Diffusion Model with Soft Mask Conditioning for Vertebral Disease Diagnosis
Shidan He, Enyuan Hu, Zixuan Tang, Dongdong Yu, Yuan Hong 0004, Zhenzhong Liu, Mengtang Li |
MICCAI (15) | 3 |
| 2025 | MIBF-Net: Multi-Modal Information Balanced Fusion Network for Clinical Diagnosis via Patient Narratives and Lesion Image
Zixuan Tang, Bai Sun, Shidan He, Yuan Hong 0004, Dongdong Yu, Zhenzhong Liu, Mengtang Li |
MICCAI (1) | 1 |
| 2025 | GraphJCL: A Dual-Perspective Graph-Based Framework for Urban Region Representation via Joint Contrastive Learning
Yaya Zhao, Kaiqi Zhao 0001, Zixuan Tang, Xiaoling Lu, Yuanyuan Zhang 0010, Yalei Du |
ECML/PKDD (3) | 3 |
| 2024 | Progressive deep snake for instance boundary extraction in medical images
Zixuan Tang, Bin Chen 0029, An Zeng |
Expert Syst. Appl. | 1 |
| 2024 | Facial Prior Guided Micro-Expression GenerationabstractThis paper focuses on the facial micro-expression (FME) generation task, which has potential application in enlarging digital FME datasets, thereby alleviating the lack of training data with labels in existing micro-expression datasets. Despite obvious progress in the image animation task, FME generation remains challenging because existing image animation methods can hardly encode subtle and short-term facial motion information. To this end, we present a facial-prior-guided FME generation framework that takes advantage of facial priors for facial motion generation. Specifically, we first estimate the geometric locations of action units (AUs) with detected facial landmarks. We further calculate an adaptive weighted prior (AWP) map, which alleviates the estimation error of AUs while efficiently capturing subtle facial motion patterns. To achieve smooth and realistic synthesis results, we use our proposed facial prior module to guide motion representation and generation modules in mainstream image animation frameworks. Extensive experiments on three benchmark datasets consistently show that our proposed facial prior module can be adopted in image animation frameworks and significantly improve their performance on micro-expression generation. Moreover, we use the generation technique to enlarge existing datasets, thereby improving the performance of general action recognition backbones on the FME recognition task. Our code is available at https://github.com/sysu19351158/FPB-FOMM. Xinhua Xu, Youjun Zhao, Yuhang Wen 0001, Zixuan Tang, Mengyuan Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Interactive Spatiotemporal Token Attention Network for Skeleton-Based General Interactive Action RecognitionabstractRecognizing interactive action plays an important role in human-robot interaction and collaboration. Previous methods use late fusion and co-attention mechanism to capture interactive relations, which have limited learning capability or inefficiency to adapt to more interacting entities. With assumption that priors of each entity are already known, they also lack evaluations on a more general setting addressing the diversity of subjects. To address these problems, we propose an Interactive Spatiotemporal Token Attention Network (ISTA-Net), which simultaneously model spatial, temporal, and interactive relations. Specifically, our network contains a tokenizer to partition Interactive Spatiotemporal Tokens (ISTs), which is a unified way to represent motions of multiple diverse entities. By extending the entity dimension, ISTs provide better interactive representations. To jointly learn along three dimensions in ISTs, multi-head self-attention blocks integrated with 3D convolutions are designed to capture inter-token correlations. When modeling correlations, a strict entity ordering is usually irrelevant for recognizing interactive actions. To this end, Entity Rearrangement is proposed to eliminate the orderliness in ISTs for interchangeable entities. Extensive experiments on four datasets verify the effectiveness of ISTA-Net by outperforming state-of-the-art methods. Our code is publicly available at https://github.com/Necolizer/ISTA-Net. Yuhang Wen 0001, Zixuan Tang, Yunsheng Pang, Beichen Ding, Mengyuan Liu 0001 |
IROS | 2 |
| 2023 | Digital Twin-Driven Collaborative Scheduling for Heterogeneous Task and Edge-End Resource via Multi-Agent Deep Reinforcement LearningabstractWith the interdisciplinary advances of mobile communication and edge computing, massive heterogeneous tasks are accessing wireless networks and competing for the edge-end computing and communication resources. Digital twin (DT), which establishes the digital models of physical objects for simulation, analysis and optimization, provides a promising method for network scheduling and management. This paper proposes a DT-driven edge-end collaborative scheduling algorithm for heterogeneous tasks and heterogeneous computing/communication resources. Specifically, multiple end devices (EDs) cooperate with each other to accomplish a complex job, where each ED can offload individual task to multiple edge servers (ESs) for parallel computing. By fully considering deadline requirements of heterogeneous tasks, maximum computing capabilities of ESs and EDs, computing resource estimation deviations of DT, maximum transmit powers of EDs and tolerable peak interference powers to coexisting EDs, we formulate a job completion time minimization problem to jointly optimize the edge-end task division, transmit power control, computing resource type matching and allocation. To solve this non-convex problem, we first reformulate it by multi-agent Markov decision process, where a compound reward leveraging latency reward and deadline reward according to the task criticality is designed. Then, we propose a multi-agent deep reinforcement learning-based scheduling algorithm, where Actor-Critic framework with estimation and target networks is designed for policy and value iterations. Meanwhile, a step-by-step ϵ-greedy algorithm is proposed to balance exploration and exploitation, avoiding local optimal trap. Through offline centralized training by DT and online distributed execution by EDs, we realize edge-end collaborative computing for heterogeneous tasks. Experimental results demonstrate that, comparing with typical benchmark algorithms, the proposed algorithm converges with the highest reward and achieves the smallest job completion time, where the deadlines of heterogeneous tasks can be well satisfied respectively. Chi Xu 0001, Zixuan Tang, Peng Zeng 0001, Linghe Kong |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | Facial Prior Based First Order Motion Model for Micro-expression GenerationabstractSpotting facial micro-expression from videos finds various potential applications in fields including clinical diagnosis and interrogation, meanwhile this task is still difficult due to the limited scale of training data. To solve this problem, this paper tries to formulate a new task called micro-expression generation and then presents a strong baseline which combines the first order motion model with facial prior knowledge. Given a target face, we intend to drive the face to generate micro-expression videos according to the motion patterns of source videos. Specifically, our new model involves three modules. First, we extract facial prior features from a region focusing module. Second, we estimate facial motion using key points and local affine transformations with a motion prediction module. Third, expression generation module is used to drive the target face to generate videos. We train our model on public CASME II, SAMM and SMIC datasets and then use the model to generate new micro-expression videos for evaluation. Our model achieves the first place in the Facial Micro-Expression Challenge 2021 (MEGC2021), where our superior performance is verified by three experts with Facial Action Coding System certification. Source code is provided in https://github.com/Necolizer/Facial-Prior-Based-FOMM. Youjun Zhao, Yuhang Wen 0001, Zixuan Tang, Xinhua Xu, Mengyuan Liu 0001 |
ACM Multimedia | 4 |