Jiandian Zeng

dblp:193/9904 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-1349-8077ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Matrix as Plan: Structured Logical Reasoning with Feedback-Driven Replanning
abstract
As knowledge and semantics on the web grow increasingly complex, enhancing Large Language Models (LLMs)' comprehension and reasoning capabilities has become particularly important. Chain-of-Thought (CoT) prompting has been shown to enhance the reasoning capabilities of LLMs. However, it still falls short on logical reasoning tasks that rely on symbolic expressions and strict deductive rules. Neuro-symbolic methods address this gap by enforcing formal correctness through external solvers. Yet these solvers are highly format-sensitive, and small instabilities in model outputs can lead to frequent processing failures. The LLM-driven approaches avoid parsing brittleness, but they lack structured representations and process-level error-correction mechanisms. To further enhance the logical reasoning capabilities of LLMs, we propose MatrixCoT, a structured CoT framework with a matrix-based plan. Specifically, we normalize and type natural language expressions and attach explicit citation fields, and introduce a matrix-based planning method to preserve global relations among steps. The plan thus becomes a verifiable artifact and execution becomes more stable. For verification, we also add a feedback-driven replanning mechanism. Under semantic-equivalence constraints, it identifies omissions and defects, rewrites and compresses the dependency matrix, and produces a more trustworthy final answer. Experiments on five logical-reasoning benchmarks and five LLMs show that, without relying on external solvers, MatrixCoT enhances both the robustness and interpretability of LLMs when tackling complex symbolic reasoning tasks, while maintaining competitive performance.
Jiandian Zeng, Zihao Peng, Guangxue Zhang, Tian Wang 0001
WWW2
2026 Cloud-edge Collaboration for Robust Network Embeddings
abstract
Learning network representations, also known as network embeddings, has attracted significant attention in recent years. Real-world scenarios often involve networks with multiple views, where each view captures a distinct aspect of the network’s structure. Existing network embedding methods mainly focus on the global information from each view, neglecting the implied relations among multiple views. Additionally, maintaining the scalability of node embeddings while adapting to changes in network topology remains a major challenge. To this end, this article proposes a Cloud-edge Collaboration Network (CC-Net) to learn robust node embeddings in multi-view networks. Specifically, we design a decomposition and regrouping module to capture implied relations within multi-view networks, enabling the generation of comprehensive node representations that integrate information from all sub-networks. Besides, by leveraging the hybrid approach of cloud and edge computing, our proposed CC-Net can efficiently handle the complexities and dynamics of multi-view networks without retraining the entire network. Extensive experiments and analyses on real-world Twitter and YouTube datasets demonstrate the superiority of our approach compared to several benchmark methods, and validate its effectiveness in capturing implied relations and generating robust node embeddings.
Jiandian Zeng, Gunagxue Zhang, Yang Li 0049, Jiantao Zhou 0001, Tian Wang 0001, Weijia Jia 0001
ACM Trans. Internet Techn.1
2026 Dynamic Grouping and Aggregation Weight Optimization for Hierarchical Federated Learning With Quantization
abstract
Hierarchical Federated Learning (HFL) alleviates communication bottlenecks by organizing the system into multiple layers: client, intermediate aggregator, and server. Clients and aggregation layers form groups based on connection patterns, and the methods used for grouping and aggregation directly affect convergence performance. Currently, some studies have proposed grouping algorithms to address the non-independent and identically distributed (non-IID) characteristics of client data to improve performance. However, these methods do not account for network heterogeneity, such as clients using different quantization levels or adaptive quantization strategies to minimize communication overhead. Moreover, most methods rely on heuristics that blindly explore the combinatorial grouping space, incurring substantial computational overhead. In this paper, we conduct a rigorous convergence analysis and frame the dual challenges of heterogeneous data and quantization heterogeneity in HFL as the joint optimization of aggregation weights and grouping strategy. Specifically, we derive the optimal closed-form solution for the aggregation weights and propose an Alternating Optimization Hierarchical Optimal Weights (AO-HOW) algorithm to compute these weights. Building on this result, we propose DyGHFL—Dynamic Exclusion–Reallocation Grouping for Hierarchical Federated Learning, a structu-reaware and efficient greedy algorithm that reorganizes groups by maximizing structured gain and updating weight coefficients according to current system conditions. Experiments on multiple datasets show that DyGHFL consistently outperforms existing baselines, demonstrating its effectiveness in HFL.
Zihao Peng, Nan Zou, Jiandian Zeng, Shengbo Chen, Tian Wang 0001, Weijia Jia 0001
IEEE Trans. Netw.3
2026 Fine-Grained Lifetime Control for Heterogeneous Service Provisioning in Energy-Constrained Edge-Edge Systems
abstract
To support delay-critical applications, migrating services from the cloud to edge servers (ESs) can effectively reduce service delay. However, in such a resource-constrained scenario, providing heterogeneous services to meet diverse user needs poses significant challenges. Specifically, ESs have limited computational resources and are often energy-constrained, complicating service placement and provisioning. Existing studies propose collaboration schemes to improve resource utilization and reduce service delay. Nevertheless, these works heavily depend on full-service coverage by the cloud or rely on coarse-grained (e.g., time-cycle level) strategies that inevitably waste energy in idle time slots, which are unsuitable for energy-constrained settings and dynamic edge environments. To address these challenges, we propose a novel edge-edge collaboration method tailored for heterogeneous edge service provisioning in energy-constrained networks. First, we formulate the heterogeneous service provisioning problem with the objective of delay minimization under energy constraints and prove its NP-hardness. Our framework leverages latest task statistics to decide the service lifetime in a fine-grained time slot level, so that enables edge collaboration to maximize resource utilization and adaptability in resource-limited conditions. Specifically, we decompose the problem into three subproblems: service placement, service lifetime decision, and task scheduling, and we design targeted lightweight solutions for each, ensuring low delay and efficient energy usage. Finally, we validate our method through comprehensive simulations on real-world datasets and implement it on a testbed with three ESs. Results demonstrate that our approach reduces service delays by an average of 69.6% across various energy-constrained scenarios.
Haodong Zou, Jianxiong Guo, Jiandian Zeng, Yupeng Li 0001, Changfu Xu, Haipeng Dai 0001, Jiannong Cao 0001, Tian Wang 0001
IEEE Trans. Netw.3
2026 Zero-Shot Real-Time Pedestrian Re-Identification Services via Edge-to-Edge Cooperation
abstract
Real-time person Re-Identification (ReID) services in video surveillance remain challenging due to real-time requirements and limited generalization of supervised models. Traditional supervised methods typically require extensive labeled training data and are confined to closed identities, struggling to generalize to unseen individuals or new environments. To overcome these limitations, recent work leverages large Vision-Language Models (VLMs) for zero-shot ReID, utilizing their rich visual-textual representations to identify pedestrians without prior training. However, directly deploying such VLMs in the cloud inherently introduces high latency, further hindering real-time performance. To tackle the above challenges, we propose VL-ZSReID, a novel edge-collaborative inference framework that deploys VLMs on resource-constrained edge devices to enable zero-shot ReID in complex multi-camera campus surveillance environments. The framework partitions the ReID task into a Tracking Service and a VLM Service, which are coordinated by a Cache-assisted Heuristic Edge Scheduling Strategy (CHESS) to reduce latency and improve throughput across heterogeneous edge devices. To enable practical deployment of VLMs on resource-constrained edge devices, VL-ZSReID adopts several system-level design choices. It incorporates a Tracking ID Discriminator (TID) as an intelligent caching module, an Attribute Caption Generation (ACG) mechanism that produces standardized feature descriptions, and a Temporal-Camera Feature Generator (TCFG) that leverages spatiotemporal context to improve accuracy. Implemented and evaluated in a real campus deployment, VL-ZSReID achieves 19.98%–195.47% higher throughput and 38.00%–82.27% lower latency than baseline methods, while improving Rank-1 and mAP by 2.44%–44.68% and 4.11%–60.89%, respectively.
Jiandian Zeng, Zihao Peng, Tian Wang 0001
IEEE Trans. Serv. Comput.2
2026 Enhancing AIGC Service Efficiency With Adaptive Multi-Edge Collaboration in a Distributed System
abstract
The Artificial Intelligence Generated Content (AIGC) technique has gained significant traction for producing diverse content. However, existing AIGC services typically operate within a centralized framework, resulting in high response times. To address this issue, we integrate collaborative Mobile Edge Computing (MEC) technology to reduce processing delays for AIGC services. Current collaborative MEC methods primarily support single-server offloading or facilitate interactions among fixed Edge Servers (ESs), limiting flexibility and resource utilization across all ESs to meet the varying computing and networking requirements of AIGC services. We propose AMCoEdge, an adaptive multi-server collaborative MEC approach to enhancing AIGC service efficiency. The AMCoEdge fully utilizes the computing and networking resources across all ESs through adaptive multi-ES selection and dynamic workload allocation, thereby minimizing the offloading make-span of AIGC services. Our design features an online distributed algorithm based on deep reinforcement learning, accompanied by theoretical analyses that confirm an approximate linear time complexity. Simulation results show that our method outperforms state-of-the-art baselines, achieving at least an$11.04\%$reduction in task offloading make-span and a$44.86\%$decrease in failure rate. Additionally, we develop a distributed prototype system to implement and evaluate our AMCoEdge method for real AIGC service execution, demonstrating service delays that are$9.23\% - 31.98\%$lower than the three representative methods.
Changfu Xu, Jianxiong Guo, Jiandian Zeng, Houming Qiu, Tian Wang 0001, Xiaowen Chu 0001, Jiannong Cao 0001
IEEE Trans. Serv. Comput.3
2025 As-Stg: Spatio-Temporal Graph Learning with Active Sampling for Dynamic IoT Sensing
abstract
Efficient sensing is critical for Internet of Things (IoT) applications, such as environmental monitoring and traffic management, where high quality sensing data is essential for decision-making. Traditional sensing methods, however, are often plagued by high deployment costs and incomplete data coverage, significantly limiting their practicality. Despite recent progress, these methods continue to face challenges in maintaining data accuracy, ultimately degrading the Quality of Service (QoS) for IoT applications. To address these limitations, we propose ASSTG, a novel framework that combines an Active Sampling strategy with Spatio-Temporal Graph learning to enable efficient and accurate IoT sensing. At its core, AS-STG is designed to minimize the sampling cost while ensuring the accuracy of the data. The framework begins by analyzing historical data to determine the minimum sampling requirements for accurate inference in subsequent time slots. It then constructs a spatio-temporal graph to model the complex relationships between sensing grids, capturing both spatial and temporal dynamics. To supplement the spatio-temporal information and further optimize representations, we introduce two contrastive learning tasks. Leveraging the refined representation, AS-STG strategically selects informationrich regions for sampling, ensuring that even a sparse subset of samples can provide comprehensive coverage of the entire sensing area. Finally, AS-STG employs matrix completion techniques to reconstruct the complete sensing data from these sparse samples. Extensive experiments on real-world datasets demonstrate that AS-STG significantly outperforms baselines in terms of inference accuracy, cost-efficiency, and scalability. By effectively reducing sampling costs without compromising QoS, AS-STG offers a robust and scalable solution for dynamic IoT sensing systems.
Yaxin Mei, Jiandian Zeng, Huiling Qin, Guangxue Zhang, James Xi Zheng, Qin Liu 0001, Tian Wang 0001
IWQoS2
2025 E2EC: Edge-to-Edge Collaboration for Efficient Real-Time Video Surveillance Inference
abstract
In smart cities, Multi-Camera Multi-Target pedestrian tracking and Re-identification (MCMT-ReID) is essential for effective surveillance, particularly in real-time scenarios, as it demands significant computational resources. Current edgecloud collaboration methods encounter issues such as high latency and potential data leakage due to the physical distance between cloud servers and cameras. To address these issues, we propose a novel Edge-to-Edge Collaboration (E2EC) system that fully utilizes collaboration between heterogeneous edge devices. E2EC partitions the MCMT-ReID task into two modular applications: Tracking and Re-identification (ReID), and employs a customized Kafka communication protocol to optimize data exchange efficiency. Moreover, E2EC dynamically orchestrates intermediate inference flows and transmits features instead of pedestrian detection frames to avoid data leakage. To enhance ReID accuracy, we introduce a real-time ReID Loop Confirmation (ReLC) algorithm, which continuously validates identities to boost reliability and accuracy. E2EC has been deployed and tested in a real-world campus environment to validate its effectiveness. Experimental results demonstrate that E2EC enhances the Rank-1 accuracy and mAP of pedestrian ReID by 36.88% and 46.00%, respectively. Furthermore, it achieves an increase of about 6.35%-12.66% in throughput and reduces latency by 35.01%-57.83% compared to baselines, ensuring realtime performance under dynamic workloads.
Jiandian Zeng, Zihao Peng, Yuzhu Liang, James Xi Zheng, Tian Wang 0001
IEEE Trans. Mob. Comput.2
2025 Enhancing QoE in Collaborative Edge Systems With Feedback Diffusion Generative Scheduling
Changfu Xu, Jianxiong Guo, Yuzhu Liang, Haodong Zou, Jiandian Zeng, Haipeng Dai 0001, Weijia Jia 0001, Jiannong Cao 0001, Tian Wang 0001
IEEE Trans. Mob. Comput.5
2024 DifAttack: Query-Efficient Black-Box Adversarial Attack via Disentangled Feature Space
abstract
This work investigates efficient score-based black-box adversarial attacks with high Attack Success Rate (ASR) and good generalizability. We design a novel attack method based on a Disentangled Feature space, called DifAttack, which differs significantly from the existing ones operating over the entire feature space. Specifically, DifAttack firstly disentangles an image's latent feature into an adversarial feature and a visual feature, where the former dominates the adversarial capability of an image, while the latter largely determines its visual appearance. We train an autoencoder for the disentanglement by using pairs of clean images and their Adversarial Examples (AEs) generated from available surrogate models via white-box attack methods. Eventually, DifAttack iteratively optimizes the adversarial feature according to the query feedback from the victim model until a successful AE is generated, while keeping the visual feature unaltered. In addition, due to the avoidance of using surrogate models' gradient information when optimizing AEs for black-box models, our proposed DifAttack inherently possesses better attack capability in the open-set scenario, where the training dataset of the victim model is unknown. Extensive experimental results demonstrate that our method achieves significant improvements in ASR and query efficiency simultaneously, especially in the targeted attack and open-set scenarios. The code is available The code is available at https://github.com/csjunjun/DifAttack.git.
Jun Liu 0071, Jiantao Zhou 0001, Jiandian Zeng, Jinyu Tian 0001
AAAI3
2024 Distributed and Efficient Request Scheduling in Collaborative Edge Computing
abstract
Cloud computing typically involves transferring users' requests to centralized cloud servers, a process that is inherently fraught with substantial delays due to the unpredictable nature of network transmissions. This inherent latency issue presents considerable challenges to applications that are highly sensitive to delay. We propose leveraging edge collaboration to minimize latency by enabling efficient user request scheduling within geographical proximities. However, cross-regional edge collaboration faces challenges due to the lack of real-time resource knowledge across regions, a problem we have identified as NP-hard. To address this, we introduce a model to connect edge nodes globally, thereby accurately reflecting their resource status. By employing an enhanced Dijkstra algorithm, we optimize the request routing process, achieving a notable reduction in delays compared to baseline methods, thus enhancing performance across various test scenarios.
Yuzhu Liang, Yaxin Mei, Guangxue Zhang, Jiandian Zeng, Tian Wang 0001
ICDCS5
2024 Enhancing AI-Generated Content Efficiency Through Adaptive Multi-Edge Collaboration
abstract
The Artificial Intelligence-Generated Content (AIGC) technique has gained significant popularity in creating diverse content. However, the current deployment of AIGC services in a centralized framework leads to high response times. To address this issue, we propose the integration of collaborative Mobile Edge Computing (MEC) technology to decrease the processing delay of AIGC services. Nevertheless, existing collaborative MEC methods only facilitate collaborative processing among fixed Edge Servers (ESs), limiting flexibility and resource utilization across heterogeneous ESs for different computing and networking requirements associated with AIGC tasks. This poses challenges for efficient resource allocation. We present an adaptive multi-server collaborative MEC approach tailored for heterogeneous edge environments to achieve efficient AIGC by dynamically allocating task workload across multiple ESs. We formulate our problem as an online linear programming problem aiming to minimize task offloading make-span. This problem is proved to be NP-hard and we propose an online adaptive multi-server selection and allocation algorithm based on deep reinforcement learning that effectively addresses this problem. Additionally, we provide theoretical performance analysis, demonstrating that our algorithm achieves near-optimal solutions within approximate linear time complexity bounds. Finally, experimental results validate the effectiveness of our method by showcasing at least 11.04% reduction in task offloading make-span and a 44.86 % decrease in failure rate compared to state-of-the-art methods.
Changfu Xu, Jianxiong Guo, Jiandian Zeng, Shengguang Meng, Xiaowen Chu 0001, Jiannong Cao 0001, Tian Wang 0001
ICDCS3
2024 Fine-Grained Service Lifetime Optimization for Energy-Constrained Edge-Edge Collaboration
abstract
Collaborative edge computing has been widely advo-cated by network operators and service providers to promote the quality of service (QoS), provisioning diverse delay-sensitive and computation-intensive applications. Existing studies mainly focus on cloud-edge collaboration, since cloud servers have massive resources to provide diverse services and edge servers can provide low-delay services with close proximity to end users. However, in scenarios that capture privacy, e.g., personal bioinformation and business areas, there is a great need for zero cloud involvement. Moreover, current edge servers are typically energy-constrained, which poses great challenges in enabling high-QoS services in ever-densely deployed edge networks. To tackle these issues, in this paper, we study the energy-constrained edge-edge collaboration problem. First, we formulate the edge-edge collaboration with delay minimization and energy reduction aims and prove its NP-hardness. Second, we propose a novel Fine-Grained Service Lifetime Optimization (FGSLO) scheme as a possible solution. The problem is then transformed and decoupled into three sub-problems, namely service placement, service lifetime decision, and task scheduling, which are solved by our proposed method, respectively. Finally, real-world data-driven experimental results show that FGSLO is capable of reducing 21.4%~90.1 % system delay in different energy-constrained scenarios, compared to baselines without service lifetime control.
Haodong Zou, Jianxiong Guo, Jiandian Zeng, Yupeng Li 0001, Jiannong Cao 0001, Tian Wang 0001
ICDCS3
2024 Incorporating Startup Delay into Collaborative Edge Computing for Superior Task Efficiency
abstract
Collaborative edge computing enables low service delay for many delay-sensitive Internet of Things applications through edge-edge and edge-cloud collaborations. Due to the limited edge resources and varying task demands, optimizing Joint Service Placement and Task Offloading (JSPTO) becomes crucial in minimizing overall processing delays. However, existing JSPTO methods overlook the impact of service startup delay, which may undermine total latency reduction, especially in scenarios with large startup delays. This paper introduces an online JSPTO method that integrates the consideration of service startup delay to enhance task offloading efficiency. However, a significant challenge is ensuring timely service response with large startup delays. We formulate this problem as an integer linear programming problem, aiming to minimize the total service startup and task processing delay. We propose a novel algorithm called SD-JSPTO, which performs online JSPTO in the presence of large startup delays. Theoretical performance analyses reveal that SD-JSPTO attains a near-optimal solution within polynomial time, demonstrating a competitive ratio of $1 + \frac{{{A_2}}}{{V{T^{{\text{opt}}}}}}$. Experimental evaluations demonstrate that our method significantly reduces the total delay by no less than 18.72% compared to state-of-the-art baseline methods while preserving system stability.
Changfu Xu, Jianxiong Guo, Jiandian Zeng, Yupeng Li 0001, Jiannong Cao 0001, Tian Wang 0001
IWQoS3
2024 PVConvNet: Pixel-Voxel Sparse Convolution for multimodal 3D object detection
Huaijin Liu, Jixiang Du, Yong Zhang 0066, Hongbo Zhang 0002, Jiandian Zeng
Pattern Recognit.5
2024 MSSA: Multi-Representation Semantics-Augmented Set Abstraction for 3D Object Detection
abstract
Accurate recognition and localization of 3D objects is a fundamental research problem in 3D computer vision. Benefiting from transformation-free point cloud processing and flexible receptive fields, point-based methods have become accurate in 3D point cloud modeling, but still fall behind voxel-based competitors in 3D detection. We observe that the set abstraction module, commonly utilized by point-based methods for downsampling points, tends to retain excessive irrelevant background information, thus hindering the effective learning of features for object detection tasks. To address this issue, we propose MSSA, a Multi-representation Semantics-augmented Set Abstraction for 3D object detection. Specifically, we first design a backbone network to encode different representation features of point clouds, which extracts point-wise features through PointNet to preserve fine-grained geometric structure features, and adopts VoxelNet to extract voxel features and BEV features to enhance the semantic features of key points. Second, to efficiently fuse different representation features of keypoints, we propose a Point feature-guided Voxel feature and BEV feature fusion (PVB-Fusion) module to adaptively fuse multi-representation features and remove noise. At last, a novel Multi-representation Semantic-guided Farthest Point Sampling (MS-FPS) algorithm is designed to help set abstraction modules progressively downsample point clouds, thereby improving instance recall and detection performance with more important foreground points. We evaluate MSSA on the widely used KITTI dataset and the more challenging nuScenes dataset. Experimental results show that compared to PointRCNN, our method improves the AP of “moderate” level for three classes of objects by 7.02%, 6.76%, and 5.44%, respectively. Compared to the advanced point-voxel-based method PV-RCNN, our method improves the AP of “moderate” level by 1.23%, 2.84%, and 0.55% for the three classes, respectively.
Huaijin Liu, Jixiang Du, Yong Zhang 0066, Hongbo Zhang 0002, Jiandian Zeng
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Exploring Semantic Relations for Social Media Sentiment Analysis
abstract
With the massive social media data available online, the conventional single modality emotion classification has developed into more complex models of multimodal sentiment analysis. Most existing works simply extracted image features at a coarse level, resulting in the absence of partially detailed visual features. Besides, social media data usually contain multiple images, while existing works considered a single image case and used only one image for representing visual features. In fact, it is nontrivial to extend the single image case to the multiple images case, due to the complex relations among multiple images. To solve the above issues, in this paper, we propose aGatedFusionSemanticRelation (GFSR) network to explore semantic relations for social media sentiment analysis. In addition to inter-relations between visual and textual modalities, we also exploit intra-relations among multiple images, potentially improving the sentiment analysis performance. Specifically, we design a gated fusion network to fuse global image embeddings and the corresponding local Adjective Noun Pair (ANP) embeddings. Then, apart from textual relations and cross-modal relations, we employ the multi-head cross attention mechanism between images and ANPs to capture similar semantic contents. Eventually, the updated textual and visual representations are concatenated for the final sentiment prediction. Extensive experiments are conducted on real-worldYelpandFlickr30kdatasets, showing that our GFSR can improve about 0.10% to 3.66% in terms of accuracy on theYelpdataset with multiple images, and achieve the best accuracy for two classes and the best macro F1 for three classes on theFlickr30kdataset with a single image.
Jiandian Zeng, Jiantao Zhou 0001, Caishi Huang
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Robust Multimodal Sentiment Analysis via Tag Encoding of Uncertain Missing Modalities
abstract
Multimodal sentiment analysis aims to extract emotions with multiple data sources, usually under the assumption that all modalities are available. In practice, such a strong assumption does not always hold, and most of multimodal sentiment analysis methods may fail when partial modalities are missing. Some existing works have started to address the missing modality problem; but only considered the single modality missing case, while ignoring the practically more general cases of multiple modalities missing. To this end, in this paper, we propose a Tag-Assisted Transformer Encoder (TATE) network to handle the problem of missing uncertain modalities. Specifically, we design a tag encoding module to cover both the single modality and multiple modalities missing cases, so as to guide the network's attention to those missing modalities. Besides, a new space projection pattern is adopted to align common vectors, taking into account the different importance of each modality. Afterwards, a Transformer encoder-decoder network is utilized to learn the missing modality features, and the outputs of the Transformer encoder are extracted for the final sentiment classification. Extensive experiments and analyses are conducted on CMU-MOSI, IEMOCAP, and MELD datasets, which show that the proposed method can achieve significant improvements compared with several baselines.
Jiandian Zeng, Jiantao Zhou 0001
IEEE Trans. Multim.1
2022 Mitigating Inconsistencies in Multimodal Sentiment Analysis under Uncertain Missing Modalities
abstract
For the missing modality problem in Multimodal Sentiment Analysis (MSA), the inconsistency phenomenon occurs when the sentiment changes due to the absence of a modality.The absent modality that determines the overall semantic can be considered as a key missing modality.However, previous works all ignored the inconsistency phenomenon, simply discarding missing modalities or solely generating associated features from available modalities.The neglect of the key missing modality case may lead to incorrect semantic results.To tackle the issue, we propose an Ensemble-based Missing Modality Reconstruction (EMMR) network to detect and recover semantic features of the key missing modality.Specifically, we first learn joint representations with remaining modalities via a backbone encoder-decoder network.Then, based on the recovered features, we check the semantic consistency to determine whether the absent modality is crucial to the overall sentiment polarity.Once the inconsistency problem due to the key missing modality exists, we integrate several encoder-decoder approaches for better decision making.Extensive experiments and analyses are conducted on CMU-MOSI and IEMOCAP datasets, validating the superiority of the proposed method.
Jiandian Zeng, Jiantao Zhou 0001
EMNLP1
2022 Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing Modalities
abstract
Multimodal sentiment analysis has been studied under the assumption that all modalities are available. However, such a strong assumption does not always hold in practice, and most of multimodal fusion models may fail when partial modalities are missing. Several works have addressed the missing modality problem; but most of them only considered the single modality missing case, and ignored the practically more general cases of multiple modalities missing. To this end, in this paper, we propose a Tag-Assisted Transformer Encoder (TATE) network to handle the problem of missing uncertain modalities. Specifically, we design a tag encoding module to cover both the single modality and multiple modalities missing cases, so as to guide the network's attention to those missing modalities. Besides, we adopt a new space projection pattern to align common vectors. Then, a Transformer encoder-decoder network is utilized to learn the missing modality features. At last, the outputs of the Transformer encoder are used for the final sentiment classification. Extensive experiments are conducted on CMU-MOSI and IEMOCAP datasets, showing that our method can achieve significant improvements compared with several baselines.
Jiandian Zeng, Jiantao Zhou 0001
SIGIR1
2022 Relation construction for aspect-level sentiment classification
Jiandian Zeng, Weijia Jia 0001, Jiantao Zhou 0001
Inf. Sci.1
2021 Fine-grained Question-Answer sentiment classification with hierarchical graph attention network
Jiandian Zeng, Weijia Jia 0001, Jiantao Zhou 0001
Neurocomputing1
2020 Data collection from WSNs to the cloud based on mobile Fog elements
Tian Wang 0001, Jiandian Zeng, Yongxuan Lai, Yiqiao Cai, Hui Tian 0002, Baowei Wang
Future Gener. Comput. Syst.2
2018 Energy-efficient relay tracking with multiple mobile camera sensors
Tian Wang 0001, Jiandian Zeng, Md. Zakirul Alam Bhuiyan, Yiqiao Cai, Hui Tian 0002, Mande Xie
Comput. Networks2