Diandian Guo

dblp:347/8565 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
15since 2021 · last 2026
0009-0002-8468-3285ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Two Streams, One Sarcasm: Orthogonal Expert Tuning for Holistic Multimodal Sarcasm Understanding
abstract
Diandian Guo, Cong Cao, Fangfang Yuan, Pin Xu, Cheng Hu, Zhicheng Zhang, Yu Liu, Yanbing Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Diandian Guo, Cong Cao 0001, Fangfang Yuan, Pin Xu, Yanbing Liu 0007
ACL (1)1
2026 MSD-NR: A Noise Rectification Method for Multimodal Sarcasm Detection
Xiangyu Tian, Diandian Guo, Cong Cao 0001, Yueshan Wang, Fangfang Yuan, Yanbing Liu 0007
ICIC2
2026 MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialogues
Diandian Guo, Fangfang Yuan, Cong Cao 0001, Xixun Lin, Chuan Zhou 0001, Hao Peng 0001, Yanan Cao 0006, Yanbing Liu 0007
WWW1
2026 LungRes80: Towards tangled surgical workflow recognition in video-assisted thoracoscopic surgery
abstract
Video-Assisted Thoracoscopic Surgery (VATS) is a minimally invasive procedure developed to remove specific lung segments for the treatment of early-stage lung diseases. The surgical procedure involves intricate vascular and bronchial anatomy to preserve as much lung tissue as possible, minimizing impact on the pulmonary function. To assist in monitoring and early warning of this high-risk surgical workflow, we build a new dataset, LungRes80, including 269,806 video frames with phase annotations sampled from 80 VATS cases. LungRes80 presents unique challenges for hierarchical temporal modeling due to diverse short-term transitions between segmentectomy phases and latent long-term causal relations. To this end, we introduce an online baseline model termed LungReco. This framework employs Masked Causal Reasoning (MCR) to perform causal reasoning with semantic modeling from continuously updated memories along with pre-trained Large Language Models (LLMs), and combines it with Concurrent Spatial-Temporal encoding (CoST) for holistic bi-modal co-spatial-temporal aggregation across short- and long-term memories. Furthermore, a new metric, called the Attentional Distraction Coefficient (ADC), is proposed to quantify the costs of intraoperative distraction and postoperative corrections by wrong predictions. We establish a comprehensive benchmark for surgical workflow recognition by evaluating representative models on LungRes80, AutoLaparo, and Cholec80, where our method consistently achieves state-of-the-art performance. Code and data are available at LungRes80.
Diandian Guo, Jialun Pei, Jiaao Li, Yanhui Wan, Hao Chen 0011, Pheng-Ann Heng
Medical Image Anal.1
2025 Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resection with Pringle Maneuver
abstract
Pringle maneuver (PM) in laparoscopic liver resection aims to reduce blood loss and provide a clear surgical view by intermittently blocking blood inflow of the liver, whereas prolonged PM may cause ischemic injury. To comprehensively monitor this surgical procedure and provide timely warnings of ineffective and prolonged blocking, we suggest two complementary AI-assisted surgical monitoring tasks: workflow recognition and blocking effectiveness detection in liver resections. The former presents challenges in real-time capturing of short-term PM, while the latter involves the intraoperative discrimination of long-term liver ischemia states. To address these challenges, we meticulously collect a novel dataset, called PmLR50, consisting of 25,037 video frames covering various surgical phases from 50 laparoscopic liver resection procedures. Additionally, we develop an online baseline for PmLR50, termed PmNet. This model embraces Masked Temporal Encoding (MTE) and Compressed Sequence Modeling (CSM) for efficient short-term and long-term temporal information modeling, and embeds Contrastive Prototype Separation (CPS) to enhance action discrimination between similar intraoperative operations. Experimental results demonstrate that PmNet outperforms existing state-of-the-art surgical workflow recognition methods on the PmLR50 benchmark. Our research offers potential clinical applications for the laparoscopic liver surgery community.
Diandian Guo, Weixin Si, Zhixi Li, Jialun Pei, Pheng-Ann Heng
AAAI1
2025 Multi-View Incongruity Learning for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environments. Regarding this phenomenon, this paper undertakes several initiatives. Firstly, we identify two primary causes that lead to the reliance of spurious correlations. Secondly, we address these challenges by proposing a novel method that integrate Multimodal Incongruities via Contrastive Learning (MICL) for multimodal sarcasm detection. Specifically, we first leverage incongruity to drive multi-view learning from three views: token-patch, entity-object, and sentiment. Then, we introduce extensive data augmentation to mitigate the biased learning of the textual modality. Additionally, we construct a test set, SPMSD, which consists potential spurious correlations to evaluate the the model’s generalizability. Experimental results demonstrate the superiority of MICL on benchmark datasets, along with the analyses showcasing MICL’s advancement in mitigating the effect of spurious correlation.
Diandian Guo, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007, Guangjie Zeng, Hao Peng 0001, Philip S. Yu
COLING1
2025 CASD: Counterfactual Augmentation for Social Bot Detection on Twitter
abstract
Social bot detection has become increasingly important with the rise of social media platforms, as social bots could be used for malicious activities such as spreading misinformation. Recent advances mainly utilize Graph Neural Networks (GNNs) for bot detection. However, these detection methods may overlook the issue of social bots’ disguise. An advanced bot can highly mimic human users and homogenize their features with human accounts, ultimately weakening the effectiveness of social bot detectors. In this paper, we propose CASD, a counterfactual-based social bot detection method. Specifically, we design a subgraph generation module to enhance node feature learning by reducing interference from irrelevant global information in sparse networks. Then, we apply counterfactual reasoning to simulate the disguising of bots and identify their distinguishable features. Finally, we combine the counterfactual feature with its raw counterpart for robust bot detection. Experimental results on multiple datasets demonstrate that our method outperforms existing state-of-the-art (SOTA) methods. Our code is publicly available at1.
Pin Xu, Fangfang Yuan, Yueshan Wang, Diandian Guo, Cong Cao 0001, Yanbing Liu 0007
ICME4
2025 Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
abstract
Recently, Deep Learning (DL) models have been increasingly deployed on end-user devices as On-Device AI, offering improved efficiency and privacy. However, this deployment trend poses more serious Intellectual Property (IP) risks, as models are distributed on numerous local devices, making them vulnerable to theft and redistribution. Most existing ownership protection solutions (e.g., backdoor-based watermarking) are designed for cloud-based AI-as-a-Service (AIaaS) and are not directly applicable to large-scale distribution scenarios, where each user-specific model instance must carry a unique watermark. These methods typically embed a fixed watermark, and modifying the embedded watermark requires retraining the model. To address these challenges, we propose Hot-Swap MarkBoard, an efficient watermarking method. It encodes user-specific n-bit binary signatures by independently embedding multiple watermarks into a multi-branch Low-Rank Adaptation (LoRA) module, enabling efficient watermark customization without retraining through branch swapping. A parameter obfuscation mechanism further entangles the watermark weights with those of the base model, preventing removal without degrading model performance. The method supports black-box verification and is compatible with various model architectures and DL tasks, including classification, image generation, and text generation. Extensive experiments across three types of tasks and six backbone models demonstrate our method's superior efficiency and adaptability compared to existing approaches, achieving 100% verification accuracy.
Zhicheng Zhang 0002, Peizhuo Lv, Mengke Wan, Jiang Fang, Diandian Guo, Yezeng Chen, Yinlong Liu, Jiyan Sun, Liru Geng
ACM Multimedia5
2025 See Better, Say Better: Vision-Augmented Decoding for Mitigating Hallucinations in Large Vision-Language Models
Xinyi Sun, Diandian Guo, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007
NLPCC (1)2
2025 S²Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
abstract
Scene graph generation (SGG) of surgical procedures is crucial in enhancing holistically cognitive intelligence in the operating room (OR). However, previous works have primarily relied on multi-stage learning, where the generated semantic scene graphs depend on intermediate processes with pose estimation and object detection. This pipeline may potentially compromise the flexibility of learning multimodal representations, consequently constraining the overall effectiveness. In this study, we introduce a novel single-stage bi-modal transformer framework for SGG in the OR, termed S2Former-OR, aimed to complementally leverage multi-view 2D scenes and 3D point clouds for SGG in an end-to-end manner. Concretely, our model embraces a View-Sync Transfusion scheme to encourage multi-view visual information interaction. Concurrently, a Geometry-Visual Cohesion operation is designed to integrate the synergic 2D semantic features into 3D point cloud features. Moreover, based on the augmented feature, we propose a novel relation-sensitive transformer decoder that embeds dynamic entity-pair queries and relational trait priors, which enables the direct prediction of entity-pair relations for graph generation without intermediate steps. Extensive experiments have validated the superior SGG performance and lower computational cost of S2Former-OR on 4D-OR benchmark, compared with current OR-SGG methods, e.g., 3 percentage points Precision increase and 24.2M reduction in model parameters. We further compared our method with generic single-stage SGG methods with broader metrics for a comprehensive evaluation, with consistently better performance achieved. Our source code can be made available at: https://github.com/PJLallen/S2Former-OR.
Jialun Pei, Diandian Guo, Jingyang Zhang, Manxi Lin, Yueming Jin, Pheng-Ann Heng
IEEE Trans. Medical Imaging2
2024 Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
abstract
The estimation of implicit cross-frame correspondences and the high computational cost have long been major chal-lenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes, feature propagation, or cross-frame attention to address these issues. By contrast, we are the first to harness vanishing point (VP) priors for more effective segmentation. Intuitively, objects near VPs (i.e., away from the vehicle) are less discernible. Moreover, they tend to move radially away from the VP over time in the usual case of a forward-facing camera, a straight road, and linear forward motion of the vehicle. Our novel, efficient network for VSS, named VPSeg, incor-porates two modules that utilize exactly this pair of static and dynamic VP priors: sparse-to-dense feature mining (DenseVP) and VP-guided motion fusion (MotionVP). MotionVP employs VP-guided motion estimation to establish explicit correspondences across frames and help attend to the most relevant features from neighboring frames, while Dense Vp enhances weak dynamic features in distant re-gions around VPs. These modules operate within a context-detail framework, which separates contextual features from high-resolution local features at different input resolutions to reduce computational costs. Contextual and local fea-tures are integrated through contextualized motion attention (CMA) for the final prediction. Extensive experiments on two popular driving segmentation benchmarks, Cityscapes and ACDC, demonstrate that VPSeg outperforms previous SOTA methods, with only modest computational overhead. The resources are available at https://github.com/RascalGdd/VPSeg.
Diandian Guo, Deng-Ping Fan, Tongyu Lu, Christos Sakaridis, Luc Van Gool
CVPR1
2024 Tri-Modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
Diandian Guo, Manxi Lin, Jialun Pei, He Tang 0002, Yueming Jin, Pheng-Ann Heng
MICCAI (6)1
2023 Curvature-Driven Knowledge Graph Embedding for Link Prediction
abstract
Knowledge Graph Embedding (KGE) aims to learn how to represent the low-dimensional vectors for entities and relations based on the observed triplets in knowledge graph. Most of the existing models use simple structural features, such as node degrees and directed edges, and pay little attention to advanced inherent information of structured knowledge. In this paper, we propose CD-GCN, a curvature-driven KGE method for link prediction. Specifically, we first apply Ricci curvature to knowledge graph. Then, we use curvature information to drive the state update, which aims to further exploit the graph-structured information. Finally, we use a ConvE scoring function to output the link prediction results. Through extensive experiments on public datasets FB15k-237 and WN18RR, CD-GCN has achieved state-of-the-art results compared with all baseline models.
Diandian Guo, Majing Su, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007, Jianhui Fu
CSCWD1
2023 A Semi-Paired Approach for Label-to-Image Translation
abstract
Data efficiency, or the ability to generalize from a few labeled data, remains a major challenge in deep learning. Semi-supervised learning has thrived in traditional recognition tasks alleviating the need for large amounts of labeled data, yet it remains understudied in image-to-image translation (I2I) tasks. In this work, we introduce the first semi-supervised (semi-paired) framework for label-to-image translation, a challenging subtask of I2I which generates photorealistic images from semantic label maps. In the semi-paired setting, the model has access to a small set of paired data and a larger set of unpaired images and labels. Instead of using geometrical transformations as a pretext task like previous works, we leverage an input reconstruction task by exploiting the conditional discriminator on the paired data as a reverse generator. We propose a training algorithm for this shared network, and we present a rare classes sampling algorithm to focus on under-represented classes. Experiments on 3 standard benchmarks show that the proposed model outperforms state-of-the-art unsupervised and semi-supervised approaches, as well as some fully supervised approaches while using a much smaller number of paired samples.
George Eskandar, Mohamed Abdelsamad, Mark Youssef, Diandian Guo, Bin Yang 0009
ICIP5
2023 Towards Pragmatic Semantic Image Synthesis for Urban Scenes
abstract
The need for large amounts of training and validation data is a huge concern in scaling AI algorithms for autonomous driving. Semantic Image Synthesis (SIS), or label-to-image translation, promises to address this issue by translating semantic layouts to images, providing a controllable generation of photorealistic data. However, they require a large amount of paired data, incurring extra costs. In this work, we present a new task: given a dataset with synthetic images and labels and a dataset with unlabeled real images, our goal is to learn a model that can generate images with the content of the input mask and the appearance of real images. This new task reframes the well-known unsupervised SIS task in a more practical setting, where we leverage cheaply available synthetic data from a driving simulator to learn how to generate photorealistic images of urban scenes. This stands in contrast to previous works, which assume that labels and images come from the same domain but are unpaired during training. We find that previous unsupervised works underperform on this task, as they do not handle distribution shifts between two different domains. To bypass these problems, we propose a novel framework with two main contributions. First, we leverage the synthetic image as a guide to the content of the generated image by penalizing the difference between their high-level features on a patch level. Second, in contrast to previous works which employ one discriminator that overfits the target domain semantic distribution, we employ a discriminator for the whole image and multiscale discriminators on the image patches. Extensive comparisons on the benchmarks benchmarks GTA-V → Cityscapes and GTA-V → Mapillary show the superior performance of the proposed model against state-of-the-art on this task.
George Eskandar, Diandian Guo, Karim Guirguis, Bin Yang 0009
IV2