VLDB 2026 Research / reviewers in the wild / expert
Jiajia Tang
dblp:234/3242
· DBLP profile ↗
19ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structure-aware coarse-to-fine upsampling network for arbitrary-scale super-resolution of remote sensing images
Shi-Yan Wang, Jiajia Tang |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Multiscale spatial-spectral attention network for arbitrary-scale hyperspectral image super-resolution
Shiyan Wang, Jiajia Tang |
Neurocomputing | 3 |
| 2026 | Adaptive Multimodal Semantic Balancing Framework for Sentiment Analysis
Jiajia Tang, Feiwei Zhou, Xiping Wang, Qibin Zhao, Yu Ding 0001, Wanzeng Kong |
IEEE Trans. Multim. | 1 |
| 2026 | Cognition-driven Adaptive Semantic Decoding Framework for Multimodal Sentiment AnalysisabstractIn real-world scenarios, multimodal sentiment analysis faces significant challenges, particularly in cross-scenario generalization. Existing works fail to effectively deal with the variability in evaluation frameworks and modality combinations, which results in poor transfer performance across different application contexts. In this article, the cognition-driven adaptive semantic decoding framework (CASDF) is proposed to realize an evaluation system and modality-independent multimodal sentiment analysis. Specifically, the adaptive modality association module is proposed to construct the adaptive modality mapping space, which allows us to dynamically adapt to arbitrary modality combinations. This indeed breaks through the limitation of the modality number and effectively deals with the modality gap. Furthermore, similar to the human hierarchical cognition (“perception-concept-decision”), the evaluation system progressive alignment module is presented to establish the unified evaluation system. This consists of the perception, concept, and decision analysis, which contributes to the adaptive cross-task analysis from the discrete sentiment space to the continuous sentiment space. The above joint analysis of the evaluation system and modality number indeed leads to the more flexible and generable multimodal sentiment semantic decoding paradigm. The experiments demonstrate that our sentiment semantic analysis network can achieve state-of-the-art performance. Jiajia Tang, Honggang Liu, Xuanyu Jin, Wanzeng Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Single-cell multi-omics and machine learning for dissecting stemness in cancerabstractCancer stem cells (CSCs) are a subpopulation of tumor cells with self-renewal capacity and the ability to drive tumor growth, metastasis, and relapse. They are widely recognized as major contributors to therapeutic resistance. Despite extensive efforts to characterize and target CSCs, their elusive nature continues to drive therapeutic resistance and relapse in epithelial malignancies. Single-cell RNA sequencing (scRNA-seq) has transformed our understanding of tumor biology. It enables high-resolution profiling of rare subpopulations (<5%) and reveals the functional heterogeneity that contributes to treatment failure. In this review, we discuss evolving evidence for a paradigm shift, enabled by rapidly advancing single-cell technologies, from a static, marker-based definition of CSCs to a dynamic and functional perspective. We explore how trajectory inference and spatial transcriptomics redefine stemness by context-dependent dynamic-state modelling. We also highlight emerging platforms, including artificial intelligence-driven predictive modelling, multi-omics integration, and functional CRISPR screens. These approaches have the potential to uncover new vulnerabilities in CSC populations. Together, these advances should lead to new precision medicine strategies for disrupting CSC plasticity, niche adaptation, and immune evasion. Shenghui Huang, Chiara Reina, Berina Sabanovic, Miriam Roberto, Alexandra Aicher, Jiajia Tang, Christopher Heeschen |
Briefings Bioinform. | 7 |
| 2025 | Backdoor Attack and Defense on Deep Learning: A SurveyabstractDeep learning, as an important branch of machine learning, has been widely applied in computer vision, natural language processing, speech recognition, and more. However, recent studies have revealed that deep learning systems are vulnerable to backdoor attacks. Backdoor attackers inject a hidden backdoor into the deep learning model, such that the predictions of the infected model will be maliciously changed if the hidden backdoor is activated by input with a backdoor trigger while behaving normally on any benign sample. This kind of attack can potentially result in severe consequences in the real world. Therefore, research on defending against backdoor attacks has emerged rapidly. In this article, we have provided a comprehensive survey of backdoor attacks, detections, and defenses previously demonstrated on deep learning. We have investigated widely used model architectures, benchmark datasets, and metrics in backdoor research and have classified attacks, detections and defenses based on different criteria. Furthermore, we have analyzed some limitations in existing methods and, based on this, pointed out several promising future research directions. Through this survey, beginners can gain a preliminary understanding of backdoor attacks and defenses. Furthermore, we anticipate that this work will provide new perspectives and inspire extra research into the backdoor attack and defense methods in deep learning. Yang Bai 0011, Gaojie Xing, Zhihong Rao, Chuan Ma 0001, Shiping Wang, Xiaolei Liu 0001, Yimin Zhou 0002, Jiajia Tang, Kaijun Huang, Jiale Kang |
IEEE Trans. Comput. Soc. Syst. | 9 |
| 2025 | MECG: modality-enhanced convolutional graph for unbalanced multimodal representations
Jiajia Tang, Binbin Ni, Yutao Yang, Yu Ding 0001, Wanzeng Kong |
J. Supercomput. | 1 |
| 2025 | Fine-grained Semantic Disentanglement Network for Multimodal Sarcasm AnalysisabstractMultimodal sarcasm analysis is one of the most challenging research branch of the sentiment analysis area, due to the presence of cross-modality incongruity. However, existing works mainly attend to the coarse-grained incongruity analysis, and totally ignore the sentiment semantic coupling issue. This indeed limits the discriminate capability and robustness of the sarcasm analysis model. In order to address the above issue, we propose a novel Fine-grained Semantic Disentanglement Network (FSDN). Specifically, the intra-modality semantic disentanglement is performed to investigate the more intrinsic semantic cues of the same modality. Additionally, the inter-modality semantic disentanglement is leveraged to simultaneously facilitate the common and intrinsic semantic cues across modalities. Furthermore, the dual-spatial semantic interaction block is presented to explore the long-range cross-spatial semantic context between the obtained verbal and non-verbal semantic space with the global view. The above semantic disentanglement processes with both local and global views significantly unleash much more robustness even for the sarcasm case consisting of multiple semantic message. Various experiments indicate that the FSDN can receive state-of-the-art or competitive performance. Jiajia Tang, Binbin Ni, Feiwei Zhou, Dongjun Liu, Yu Ding 0001, Yong Peng 0001, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Fine-grained Dual-space Context Analysis Network for Multimodal Emotion RecognitionabstractMultimodal emotion recognition has received widespread attention in variety of domains, which can utilize multiple modalities of emotional information to improve the performance of emotion recognition. However, existing coarse-grained works mainly attend to the temporal domain context and totally ignore the temporal-spatial domain context, which results in the significant deterioration of the emotion analysis performance. In this work, the fine-grained dual-space context analysis network (FDCAN) is proposed to fully investigate the fine-grained emotion context among the joint temporal-spatial emotion representative space. Specifically, the depthwise separable convolution based operation is leveraged to exploit the temporal and spatial space from EEG and EOG modality. Furthermore, the correlation analysis based technique is introduced to investigate cross-modality correlation messages from the above obtained spaces, leading to the coupled temporal-spatial representative space. Additionally, the attention mechanism based procedure is performed to deal with the comprehensive and sophisticated multi-modality emotion context from the coupled representative space. Note that, the above carefully designed hierarchical dual-space emotion context procedure indeed provide a new detection and bears the strong potential to facilitate the emotion analysis. We validated the effectiveness of the proposed framework on the popular and public multimodal emotion analysis benchmark DEAP. The experimental results demonstrated that our model achieved better performance of 97.4% and 96.9% for binary and four-class emotion classification task. Jiajia Tang, Wanzeng Kong, Yanyan Ying |
IJCNN | 2 |
| 2024 | Unbiased Semantic Representation Learning Based on Causal Disentanglement for Domain GeneralizationabstractDomain generalization primarily mitigates domain shift among multiple source domains, generalizing the trained model to an unseen target domain. However, the spurious correlation usually caused by context prior (e.g., background) makes it challenging to get rid of the domain shift. Therefore, it is critical to model the intrinsic causal mechanism. The existing domain generalization methods only attend to disentangle the semantic and context-related features by modeling the causation between input and labels, which totally ignores the unidentifiable but important confounders. In this article, a Causal Disentangled Intervention Model (CDIM) is proposed for the first time, to the best of our knowledge, to construct confounders via causal intervention. Specifically, a generative model is employed to disentangle the semantic and context-related features. The contextual information of each domain from generative model can be considered as a confounder layer, and the center of all context-related features is utilized for fine-grained hierarchical modeling of confounders. Then the semantic and confounding features from each layer are combined to train an unbiased classifier, which exhibits both transferability and robustness across an unknown distribution domain. CDIM is evaluated on three widely recognized benchmark datasets, namely, Digit-DG, PACS, and NICO, through extensive ablation studies. The experimental results clearly demonstrate that the proposed model achieves state-of-the-art performance. Xuanyu Jin, Wanzeng Kong, Jiajia Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | BAFN: Bi-Direction Attention Based Fusion Network for Multimodal Sentiment AnalysisabstractAttention-based networks currently identify their effectiveness in multimodal sentiment analysis. However, existing methods ignore the redundancy of auxiliary modalities. More importantly, existing methods only attend to top-down attention (static process) or down-top attention (implicit process), leading to the coarse-grained multimodal sentiment context. In this paper, during the preprocessing period, we first propose the multimodal dynamic enhanced block to capture the intra-modality sentiment context. This can effectively decrease the intra-modality redundancy of auxiliary modalities. Furthermore, the bi-direction attention block is proposed to capture fine-grained multimodal sentiment context via the novel bi-direction multimodal dynamic routing mechanism. Specifically, the bi-direction attention block first highlights the explicit and low-level multimodal sentiment context. Then, the low-level multimodal context is transmitted to a carefully designed bi-direction multimodal dynamic routing procedure. This allows us to dynamically update and investigate high-level and much more fine-grained multimodal sentiment contexts. The experiments demonstrate that our fusion network can achieve state-of-the-art performance. Notably, our model outperforms the best baseline on the metric ‘Acc-7’ with an improvement of 6.9%. Jiajia Tang, Dongjun Liu, Xuanyu Jin, Yong Peng 0001, Qibin Zhao, Yu Ding 0001, Wanzeng Kong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | MMT: Multi-way Multi-modal Transformer for Multimodal LearningabstractThe heart of multimodal learning research lies the challenge of effectively exploiting fusion representations among multiple modalities.However, existing two-way cross-modality unidirectional attention could only exploit the intermodal interactions from one source to one target modality. This indeed fails to unleash the complete expressive power of multimodal fusion with restricted number of modalities and fixed interactive direction.In this work, the multiway multimodal transformer (MMT) is proposed to simultaneously explore multiway multimodal intercorrelations for each modality via single block rather than multiple stacked cross-modality blocks. The core idea of MMT is the multiway multimodal attention, where the multiple modalities are leveraged to compute the multiway attention tensor. This naturally benefits us to exploit comprehensive many-to-many multimodal interactive paths. Specifically, the multiway tensor is comprised of multiple interconnected modality-aware core tensors that consist of the intramodal interactions. Additionally, the tensor contraction operation is utilized to investigate intermodal dependencies between distinct core tensors.Essentially, our tensor-based multiway structure allows for easily extending MMT to the case associated with an arbitrary number of modalities. Taking MMT as the basis, the hierarchical network is further established to recursively transmit the low-level multiway multimodal interactions to high-level ones. The experiments demonstrate that MMT can achieve state-of-the-art or comparable performance. Jiajia Tang, Xuanyu Jin, Wanzeng Kong, Yu Ding 0001, Qibin Zhao |
IJCAI | 1 |
| 2022 | Dynamically Adjust Word Representations Using Unaligned Multimodal InformationabstractMultimodal Sentiment Analysis is a promising research area for modeling multiple heterogeneous modalities. Two major challenges that exist in this area are a) multimodal data is unaligned in nature due to the different sampling rates of each modality, and b) long-range dependencies between elements across modalities. These challenges increase the difficulty of conducting efficient multimodal fusion. In this work, we propose a novel end-to-end network named Cross Hyper-modality Fusion Network (CHFN). The CHFN is an interpretable Transformer-based neural model that provides an efficient framework for fusing unaligned multimodal sequences. The heart of our model is to dynamically adjust word representations in different non-verbal contexts using unaligned multimodal sequences. It is concerned with the influence of non-verbal behavioral information at the scale of the entire utterances and then integrates this influence into verbal expression. We conducted experiments on both publicly available multimodal sentiment analysis datasets CMU-MOSI and CMU-MOSEI. The experiment results demonstrate that our model surpasses state-of-the-art models. In addition, we visualize the learned interactions between language modality and non-verbal behavior information and explore the underlying dynamics of multimodal language data. Jiwei Guo, Jiajia Tang, Weichen Dai 0001, Yu Ding 0001, Wanzeng Kong |
ACM Multimedia | 2 |
| 2021 | CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion NetworkabstractJiajia Tang, Kang Li, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jiajia Tang, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong |
ACL/IJCNLP (1) | 1 |
| 2021 | Knowledge-distillation-aided Lightweight Neural Network for Massive MIMO CSI FeedbackabstractIn massive multiple-input multiple-output (MIMO) systems, channel state information (CSI) is required by the base station (BS) to achieve high-performance gains. In frequency division duplexing (FDD) systems, the downlink CSI matrix should be sent back to the BS; unfortunately, the computational and overhead cost of this task is inherently high. Recently, deep learning has been increasingly applied in the space of CSI feedback. However, neural networks entail extra memory and computational requirements, which undermines the deployment of CSI feedback neural networks at the user equipment (UE) side. The conventional lightweight methods such as pruning and quantization requires heavy workload of experiments and difficulty of individually designing training methods for each neural network (NN). In this paper, a novel network lightweight method utilizing knowledge distillation as a training method is introduced to lighten the computation burden of the encoder at the UEs. Knowledge distillation (KD) aims at transferring knowledge from a complex network to a simple network and improving the performance of the simple network close to the complex network. Our numerical experiments demonstrate that the performance of the proposed network can be improved with KD. Huaze Tang, Jiajia Tang, Michail Matthaiou, Chao-Kai Wen, Shi Jin 0002 |
VTC Fall | 2 |
| 2021 | Delay and Energy Consumption Optimization Oriented Multi-service Cloud Edge Collaborative Computing Mechanism in IoTabstractThe rapid development of the Internet of Things has put forward higher requirements for the processing capacity of the network. The adoption of cloud edge collaboration technology can make full use of computing resources and improve the processing capacity of the network. However, in the cloud edge collaboration technology, how to design a collaborative assignment strategy among different devices to minimize the system cost is still a challenging work. In this paper, a task collaborative assignment algorithm based on genetic algorithm and simulated annealing algorithm is proposed. Firstly, the task collaborative assignment framework of cloud edge collaboration is constructed. Secondly, the problem of task assignment strategy was transformed into a function optimization problem with the objective of minimizing the time delay and energy consumption cost. To solve this problem, a task assignment algorithm combining the improved genetic algorithm and simulated annealing algorithm was proposed, and the optimal task assignment strategy was obtained. Finally, the simulation results show that compared with the traditional cloud computing, the proposed method can improve the system efficiency by more than 25%. Sujie Shao, Jiajia Tang, Jianong Li, Shao-Yong Guo 0001, Feng Qi 0004 |
J. Web Eng. | 2 |
| 2019 | Deep Multimodal Multilinear Fusion with High-order Polynomial PoolingabstractTensor-based multimodal fusion techniques have exhibited great predictive performance. However, one limitation is that existing approaches only consider bilinear or trilinear pooling, which fails to unleash the complete expressive power of multilinear fusion with restricted orders of interactions. More importantly, simply fusing features all at once ignores the complex local intercorrelations, leading to the deterioration of prediction. In this work, we first propose a polynomial tensor pooling (PTP) block for integrating multimodal features by considering high-order moments, followed by a tensorized fully connected layer. Treating PTP as a building block, we further establish a hierarchical polynomial fusion network (HPFN) to recursively transmit local correlations into global ones. By stacking multiple PTPs, the expressivity capacity of HPFN enjoys an exponential growth w.r.t. the number of layers, which is shown by the equivalence to a very deep convolutional arithmetic circuits. Various experiments demonstrate that it can achieve the state-of-the-art performance. Jiajia Tang, Wanzeng Kong, Qibin Zhao |
NeurIPS | 2 |
| 2018 | A New Method for Brain Death Diagnosis Based on Phase Synchronization Analysis With EEG
Jianting Cao, Wanzeng Kong, Jiajia Tang, Yong Peng 0001 |
BIBM | 5 |
| 2018 | Emotional-state brain network analysis revealed by minimum spanning tree using EEG signals
Shaokai Zhao, Jiajia Tang, Tao Zhang 0062, Yong Peng 0001, Wanzeng Kong |
BIBM | 4 |