VLDB 2026 Research / reviewers in the wild / expert
Guoshun Nan
dblp:154/9547
· DBLP profile ↗
65ranked-venue papers
6as first author
57since 2021 · last 2026
0000-0002-1987-2736ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Security and privacy · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fusing Situations of Massive Mobile Nodes Improves the LLM-Based Attack Prediction for AI-Native Edges
Rushan Li, Zhuoran Duan, Guoshun Nan, Qimei Cui, Xiaofeng Tao 0001 |
WCNC | 4 |
| 2026 | Advancing LLM-Based Security Automation With Customized Group Relative Policy Optimization for Zero-Touch NetworksabstractZero-Touch Networks (ZTNs) represent a transformative paradigm toward fully automated and intelligent network management, providing the scalability and adaptability required for the complexity of sixth-generation (6G) networks. However, the distributed architecture, high openness, and deep heterogeneity of 6G networks expand the attack surface and pose unprecedented security challenges. To address this, security automation aims to enable intelligent security management across dynamic and complex environments, serving as a key capability for securing 6G ZTNs. Despite its promise, implementing security automation in 6G ZTNs presents two primary challenges: 1) automating the lifecycle from security strategy generation to validation and update under real-world, parallel, and adversarial conditions, and 2) adapting security strategies to evolving threats and dynamic environments. This motivates us to propose SecLoop and SA-GRPO. SecLoop constitutes the first fully automated framework that integrates large language models (LLMs) across the entire lifecycle of security strategy generation, orchestration, response, and feedback, enabling intelligent and adaptive defenses in dynamic network environments, thus tackling the first challenge. Furthermore, we propose SA-GRPO, a novel security-aware group relative policy optimization algorithm that iteratively refines security strategies by contrasting group feedback collected from parallel SecLoop executions, thereby addressing the second challenge. Extensive real-world experiments on five benchmarks, including 11 MITRE ATT&CK processes and over 20 types of attacks, demonstrate the superiority of the proposed SecLoop and SA-GRPO. We will release our platform to the community, facilitating the advancement of security automation towards next generation communications. Xinye Cao, Yihan Lin 0001, Guoshun Nan, Qinchuan Zhou, Yuhang Luo, Yurui Gao, Haolang Lu, Qimei Cui, Yan-Zhao Hou, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 3 |
| 2026 | Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning FrameworkabstractWhile virtualization and resource pooling empower cloud networks with structural flexibility and elastic scalability, they inevitably expand the attack surface and challenge cyber resilience. Reinforcement Learning (RL)-based defense strategies have been developed to optimize resource deployment and isolation policies under adversarial conditions, aiming to enhance system resilience by maintaining and restoring network availability. However, existing approaches lack robustness as they require retraining to adapt to dynamic changes in network structure, node scale, attack strategies, and attack intensity. Furthermore, the lack of Human-in-the-Loop (HITL) support limits interpretability and flexibility. To address these limitations, we propose CyberOps-Bots, a hierarchical multi agent reinforcement learning framework empowered by Large Language Models (LLMs). Inspired by MITRE ATT&CK's “Tactics-Techniques” model, CyberOps-Bots features a two-layer architecture: (1) An upper-level LLM agent with four mod ules—ReAct planning, IPDRR-based perception, long-short term memory, and action/tool integration—performs global awareness, human intent recognition, and tactical planning; (2) Lower-level RL agents, developed via heterogeneous separated pre-training, execute atomic defense actions within localized network regions. This synergy preserves LLM adaptability and interpretability while ensuring reliable RL execution. Experiments on real cloud datasets show that, compared to state-of-the-art algorithms, CyberOps-Bots maintains network availability 68.5% higher and achieves a 34.7% jumpstart performance gain when shifting the scenarios without retraining. To our knowledge, this is the first study to establish a robust LLM-RL framework with HITL support for cloud defense. Yixiao Peng, Hao Hu 0005, Feiyang Li, Xinye Cao, Yingchang Jiang, Jipeng Tang, Guoshun Nan |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | Secure and Efficient Model Training Framework for Multiuser Semantic Communications via Over-the-Air MixupabstractOnline model training is pivotal for enabling multiuser semantic communication systems to adapt to dynamic channel conditions. However, conventional frameworks suffer from prohibitive communication overhead and vulnerabilities to privacy attacks, hindering practical deployment. This paper proposes semantic information mixup (SIMix), a secure and efficient training framework that integrates Over-the-Air Mixup (OAM) with label-aware user grouping to jointly optimize spectral efficiency and semantic security. The OAM mixes semantic features of multiple users via wireless channels, inherently obfuscating sensitive data while reducing communication overhead. A closed-form Tx-Rx scaling optimization minimizes the mean square error (MSE) of over-the-air computation under channel noise, ensuring stable convergence in low-SNR regimes. Furthermore, an extended max-clique algorithm dynamically partitions users into groups with minimal intra-label similarity, reducing model inversion attack success rates. Experiments on CIFAR-10 and Tiny ImageNet demonstrate that the proposed approach is superior in terms of communication efficiency and security, reducing communication overhead by up to 25% and attaining 17.58 dB PSNR (20.98 dB reduction) under inversion attack and reducing 13.44% attack success rate under label inference attack, while achieving comparable transmission accuracy. Xun Ma, Xinchen Lyu, Chenshan Ren, Guoshun Nan, Qimei Cui |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs With Minimal Human InterventionsabstractRecent AI agents, such as ChatGPT and LLaMA, primarily rely on instruction tuning and reinforcement learning to calibrate the output of large language models (LLMs) with human intentions, ensuring the outputs are harmless and helpful. Existing methods heavily depend on the manual annotation of high-quality positive samples, while contending with issues such as noisy labels and minimal distinctions between preferred and dispreferred response data. However, readily available toxic samples with clear safety distinctions are often filtered out, removing valuable negative references that could aid LLMs in safety alignment. In response, we propose Positive–Toxic Self-Alignment (PT-ALIGN), a novel safety self-alignment approach that minimizes human supervision by automatically refining positive and toxic samples and performing fine-grained dual instruction tuning. Positive samples are harmless responses, while toxic samples deliberately contain extremely harmful content, serving as a new supervisory signal. Specifically, we utilize LLM itself to iteratively generate and refine training instances by only exploring fewer than 50 human annotations. We then employ two losses, i.e., maximum likelihood estimation (MLE) and fine-grained unlikelihood training (UT), to jointly learn to enhance the LLM’s safety. The MLE loss encourages an LLM to maximize the generation of harmless content based on positive samples. Conversely, the fine-grained UT loss guides the LLM to minimize the output of harmful words based on toxic samples at the token-level, thereby guiding the model to decouple safety from effectiveness, directing it toward safer fine-tuning objectives, and increasing the likelihood of generating helpful and reliable content. Experiments on 9 popular open-source LLMs demonstrate the effectiveness of our PT-ALIGN for safety alignment, while maintaining comparable levels of helpfulness and usefulness. Jingxin Xu, Guoshun Nan, Sheng Guan, Sicong Leng, Yilian Liu, Yuyang Ma, Yan-Zhao Hou, Xiaofeng Tao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | SinColor: Uncertainty-Guided Single-Step Diffusion for Image ColorizationabstractImage colorization is a fundamental yet challenging task in computer vision, aiming to recover plausible and spatially coherent colors from grayscale images. Recent advancements in diffusion models have enabled significant progress in this field, yet existing methods predominantly rely on multi-step diffusion processes. While effective for generating high-frequency details, these approaches are suboptimal for colorization, as color information is inherently low-frequency, spatially smooth, and globally consistent. This mismatch leads to two critical limitations: 1) color artifacts and inconsistency due to excessive noise in the color space, and 2) high computational cost that hinders practical application. In this work, we propose a novel single-step diffusion framework for efficient and high-quality image colorization. We introduce a color uncertainty estimation (CUE) module to identify reliable and uncertain regions in the image, allowing the model to prioritize local certainty while reasoning about confused regions. To focus the model on low-frequency color generation, we directly encode the grayscale image into a latent representation, remove structural components in the output, and reconstruct the final image via efficient decoding. Extensive experiments on ImageNet, COCO-Stuff, and Extended COCO-Stuff demonstrate that our approach achieves state-of-the-art performance while reducing inference time by 98% and trainable parameters by 97% compared to leading multi-step diffusion methods. Our contributions include a systematic analysis of diffusion-based colorization, a lightweight yet effective uncertainty-aware framework, and comprehensive validation of its efficiency and effectiveness. Yutong Gao 0001, Congyan Lang, Yidian Liu, Fayao Liu, Guoshun Nan, Yunchao Wei |
IEEE Trans. Image Process. | 7 |
| 2025 | E-MHSAC: Physical Layer Key Generation in MIMO-RIS Systems Using Deep Reinforcement Learning
Tingyu Xie, Guoshun Nan, Qimei Cui, Huici Wu, Xiaofeng Tao 0001 |
GLOBECOM | 2 |
| 2025 | Robust Secure MIMO Integrated Sensing and Communications with Sensing-Assisted Wiretap Channel AwarenessabstractA major limitation of physical layer security is the requirement for prior knowledge of the potential eavesdroppers' (Eves) channels, which makes its practical implementation challenging. Integrated sensing and communication systems can be exploited to address this issue by enabling both communication and sensing functionalities, sensing the physical environment to estimate Eves' directions and amplitudes, and further reconstructing their channel state information (CSI) to achieve sensing-assisted wiretap channel awareness, thus realizing secure communication. We analyze a more practical scenario where not only is reconstructed Eves' CSI inaccurate, but also the CSI of the legitimate user equipments (UEs) estimated by the base station is imperfect, and Eves are equipped with multiple antennas. To resolve this, we propose a robust optimization problem that jointly maximizes the sensing accuracy and the secrecy rate by co-designing the beamforming and the artificial noise matrix. For the highly challenging non-convex infinite objective function, the S-Procedure is introduced, which is then solved by an efficient algorithm based on alternating optimization and successive convex approximation. Numerical results validate the effectiveness of the proposed algorithm and demonstrate that the robust beamforming scheme can mitigate the impact of channel uncertainty (both UEs' and Eves') on the system performance. Xinyuan Ma, Zengbao Zhu, Qimei Cui, Guoshun Nan, Na Li 0001, Xiaofeng Tao 0001 |
ICC | 4 |
| 2025 | VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-Based Group Relative Policy OptimizationabstractUnderstanding hour-long videos with multi-modal large language models (MM-LLMs) enriches the landscape of human-centered AI applications. However, for end-to-end video understanding with LLMs, uniformly sampling video frames results in LLMs being overwhelmed by a vast amount of irrelevant information as video length increases. Existing hierarchical key frame extraction methods improve the accuracy of video understanding but still face two critical challenges. 1) How can the interference of extensive redundant information in long videos be mitigated? 2) How can a model dynamically adapt to complex hierarchical structures while accurately identifying key frames? To address these issues, we propose VideoMiner, which iteratively segments, captions, and clusters long videos, forming a hierarchical tree structure. The proposed VideoMiner progresses from long videos to events to frames while preserving temporal coherence, effectively addressing the first challenge. To precisely locate key frames, we introduce T-GRPO, a tree-based group relative policy optimization in reinforcement learning method that guides the exploration of the VideoMiner. The proposed T-GRPO is specifically designed for tree structures, integrating spatiotemporal information at the event level while being guided by the question, thus solving the second challenge. We achieve superior performance in all long-video understanding tasks and uncover several interesting insights. Our proposed T-GRPO surprisingly incentivizes the model to spontaneously generate a reasoning chain. Additionally, the designed tree growth auxin dynamically adjusts the expansion depth, obtaining accuracy and efficiency gains. The code is publicly available at https://github.com/caoxinye/VideoMiner. Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Yuqi Pan, Tianhao Hou, Yutong Gao 0001 |
ICCV | 4 |
| 2025 | From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image ReasoningabstractMulti-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challenges beyond single-image or non-interleaved multi-image tasks. While current multi-image benchmarks overlook interleaved textual contexts and neglect distinct relationships between individual images and their associated texts, enabling models to reason over multi-image interleaved data may significantly enhance their comprehension of complex scenes and better capture cross-modal correlations. To bridge this gap, we introduce a novel benchmark MIR, requiring joint reasoning over multiple images accompanied by interleaved textual contexts to accurately associate image regions with corresponding texts and logically connect information across images. To enhance MLLMs ability to comprehend multi-image interleaved data, we introduce reasoning steps for each instance within the benchmark and propose a stage-wise curriculum learning strategy. This strategy follows an "easy to hard" approach, progressively guiding models from simple to complex scenarios, thereby enhancing their ability to handle challenging tasks. Extensive experiments benchmarking multiple MLLMs demonstrate that our method significantly enhances models reasoning performance on MIR and other established benchmarks. We believe that MIR will encourage further research into multi-image interleaved reasoning, facilitating advancements in MLLMs capability to handle complex inter-modal tasks. Guoshun Nan, Wendi Deng, Zhenyan Chen, Xiao Wang 0002, Yuqi Pan, Tao Qi 0001, Sicong Leng |
ICCV | 3 |
| 2025 | KGMark: A Diffusion Watermark for Knowledge GraphsabstractKnowledge graphs (KGs) are ubiquitous in numerous real-world applications, and watermarking facilitates protecting intellectual property and preventing potential harm from AI-generated content. Existing watermarking methods mainly focus on static plain text or image data, while they can hardly be applied to dynamic graphs due to spatial and temporal variations of structured data. This motivates us to propose KGMark, the first graph watermarking framework that aims to generate robust, detectable, and transparent diffusion fingerprints for dynamic KG data. Specifically, we propose a novel clustering-based alignment method to adapt the watermark to spatial variations. Meanwhile, we present a redundant embedding strategy to harden the diffusion watermark against various attacks, facilitating the robustness of the watermark to the temporal variations. Additionally, we introduce a novel learnable mask matrix to improve the transparency of diffusion fingerprints. By doing so, our KGMark properly tackles the variation challenges of structured data. Experiments on various public benchmarks show the effectiveness of our proposed KGMark. Hongrui Peng, Haolang Lu, Yuanlong Yu 0002, Weiye Fu, Kun Wang 0056, Guoshun Nan |
ICML | 6 |
| 2025 | EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation
Zhijiang Li, Haoran He, Guoshun Nan |
ICONIP (5) | 3 |
| 2025 | SkinMamba: Segmentation and Classification of Skin Cancer with Multi-level Context UnderstandingabstractSkin cancer accounts for nearly 40% of all cancer cases. Segmentation and classification of the lesions can help medical professionals delineate the boundaries of skin lesions and then categorize the type, ensuring timely and efficient intervention. However, the nuance of lesions, hairs over the skin, and blurred boundaries make such a task quite challenging. This motivates us to propose SkinMamba, a novel method that explores Mamba to learn the multi-level context of skin lesions, thereby enabling more accurate segmentation of affected areas and determination of lesion type. Specifically, we introduce a novel encoder termed SkinBlock, by integrating convolutional layers with the Mamba approach, and such an encoder can effectively capture the global and local context of the lesions. We feed the output segmentation clues and three features with different focal areas to the classifier. The classifier consists of the proposed SkinBlock and ResNet50. By doing so, our SkinMamba can properly tackle the challenges mentioned above. Experiments on two public benchmarks show the effectiveness of the proposed SkinMamba. Guoshun Nan, Chengyao Jia, Yutong Gao 0001, Zuye Xiao |
IJCNN | 2 |
| 2025 | Advancing Expert Specialization for Better MoEabstractMixture-of-Experts (MoE) models enable efficient scaling of large language models (LLMs) by activating only a subset of experts per input.
However, we observe that the commonly used auxiliary load balancing loss often leads to expert overlap and overly uniform routing, which hinders expert specialization and degrades overall performance during post-training.
To address this, we propose a simple yet effective solution that introduces two complementary objectives: (1) an orthogonality loss to encourage experts to process distinct types of tokens, and (2) a variance loss to encourage more discriminative routing decisions.
Gradient-level analysis demonstrates that these objectives are compatible with the existing auxiliary loss and contribute to optimizing the training process.
Experimental results over various model architectures and across multiple benchmarks show that our method significantly enhances expert specialization.
Notably, our method improves classic MoE baselines with auxiliary loss by up to 23.79\%, while also maintaining load balancing in downstream tasks, without any architectural modifications or additional components. We will release our code to contribute to the community. Hongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu, Jialin Zhuang, Wenhao Che, Xinye Cao, Sicong Leng, Qimei Cui |
NeurIPS | 3 |
| 2025 | Auditing Meta-Cognitive Hallucinations in Reasoning Large Language ModelsabstractThe development of Reasoning Large Language Models (RLLMs) has significantly improved multi-step reasoning capabilities, but it has also made hallucination problems more frequent and harder to eliminate. While existing approaches address hallucination through external knowledge integration, model parameter analysis, or self-verification mechanisms, they fail to provide a comprehensive insight into how hallucinations **emerge** and **evolve** throughout the reasoning chain. In this work, we investigate hallucination causality under constrained knowledge domains by auditing the Chain-of-Thought (CoT) trajectory and assessing the model's cognitive confidence in potentially erroneous or biased claims.
Analysis reveals that in long-CoT settings, RLLMs may iteratively reinforce biases and errors through flawed reflective processes, ultimately inducing hallucinated reasoning paths.
Counterintuitively, even with interventions at hallucination origins, reasoning chains display pronounced ''chain disloyalty'', resisting correction and sustaining flawed trajectories.
We further point out that existing hallucination detection methods are *less reliable and interpretable than previously assumed*, especially in complex multi-step reasoning contexts.
Unlike circuit tracing that requires access to model parameters, our auditing **enables more interpretable long-chain hallucination attribution in black-box settings**, demonstrating stronger generalizability and practical utility.
Our code is available at [this link](https://github.com/Winnie-Lian/AHa_Meta_Cognitive). Haolang Lu, Yilian Liu, Jingxin Xu, Guoshun Nan, Yuanlong Yu 0002, Zhican Chen, Kun Wang 0056 |
NeurIPS | 4 |
| 2025 | Overview of AI and communication for 6G network: fundamentals, challenges, and future research opportunitiesabstractAbstract With the growing demand for seamless connectivity and intelligent communication, the integration of artificial intelligence (AI) and sixth-generation (6G) communication networks has emerged as a transformative paradigm. By embedding AI capabilities across various network layers, this integration enables optimized resource allocation, improved efficiency, and enhanced system robust performance. This paper presents a comprehensive overview of AI and communication for 6G networks, with a focus on their foundational principles, inherent challenges, and future research opportunities. We first review the integration of AI and communications in the context of 6G, exploring the driving factors behind incorporating AI into wireless communications, as well as the vision for the convergence of AI and 6G. The discourse then transitions to a detailed exposition of the envisioned integration of AI within 6G networks, divided into three progressive stages. The first stage, AI for network, focuses on employing AI to augment network performance, optimize efficiency, and enhance user service experiences. The second stage, network for AI, highlights the role of the network in facilitating and buttressing AI operations and presents key enabling technologies. We compare wireless network large models with conventional large language models (LLMs), and identify key design principles and components for building wireless network architectures. In the final stage, AI as a service, it is anticipated that future 6G networks will innately provide AI functions as services, supporting application scenarios like immersive communication and intelligent industrial robots. Specifically, we define the quality of AI service, which refers to a framework for measuring AI services within the network. We further summarize the standardization process of AI for wireless networks, highlighting key milestones and ongoing efforts. In addition, we analyze the critical challenges faced by the integration of AI and communications in 6G. Finally, we outline promising future research opportunities that are expected to drive the development and refinement of AI and 6G communications. Qimei Cui, Xiaohu You 0001, Wei Ni 0001, Guoshun Nan, Xuefei Zhang 0003, Jianhua Zhang 0001, Xinchen Lyu, Ming Ai, Xiaofeng Tao 0001, Zhiyong Feng 0001, Ping Zhang 0003, Qingqing Wu 0001, Meixia Tao, Yongming Huang 0001, Chongwen Huang, Guangyi Liu 0001, Chenghui Peng, Zhiwen Pan, Dusit Niyato, Tao Chen 0011, Muhammad Khurram Khan, Abbas Jamalipour, Mohsen Guizani, Chau Yuen |
Sci. China Inf. Sci. | 4 |
| 2025 | Accountable Distributed Access Control With Privacy Preservation for Blockchain-Enabled Internet of Things Systems: A Zero-Trust Security SchemeabstractWhile being able to avoid single point failures, emerging decentralized security techniques are facing new challenges of reliability, robustness, and privacy preservation in blockchain-enabled Internet of Things (IoT) systems. To circumvent these issues, a zero-trust security scheme is proposed through distributed access control, enhanced authentication, dynamic authorization, and privacy preservation enabled by the consortium blockchain. The proposed scheme integrates three key components, i.e., a distributed recommendation mechanism, where multiple authorized nodes are utilized as referrers to efficiently confer their trust on a new public entity for enhanced authentication; an anonymous credential generation strategy, which is developed for the new entity to further protect its privacy from linking attacks; and an adaptive reputation update strategy, which is proposed for evaluating the nodes’ behaviors in the system for accountability and dynamic multiple-level authorization. The proposed scheme is implemented in a Hyperledge Fabric and the results show that it significantly enhances security and protects private information. He Fang, Li Xu 0002, Guoshun Nan, Danyang Zheng 0001, Haitao Zhao 0004, Xianbin Wang 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal CommunicationsabstractInteractive multimodal applications (IMAs), such as route planning in the Internet of Vehicles, enrich users’ personalized experiences by integrating various forms of data over wireless networks. Recent advances in large language models (LLMs) utilize mixture-of-experts (MoE) mechanisms to empower multiple IMAs, with each LLM trained individually for a specific task that presents different business workflows. In contrast to existing approaches that rely on multiple LLMs for IMAs, this paper presents a novel paradigm that accomplishes various IMAs using a single compositional LLM over wireless networks. The two primary challenges include 1) guiding a single LLM to adapt to diverse IMA objectives and 2) ensuring the flexibility and efficiency of the LLM in resource-constrained mobile environments. To tackle the first challenge, we propose ContextLoRA, a novel method that guides an LLM to learn the rich structured context among IMAs by constructing a task dependency graph. We partition the learnable parameter matrix of neural layers for each IMA to facilitate LLM composition. Then, we develop a step-by-step fine-tuning procedure guided by task relations, including training, freezing, and masking phases. This allows the LLM to learn to reason among tasks for better adaptation, capturing the latent dependencies between tasks. For the second challenge, we introduce ContextGear, a scheduling strategy to optimize the training procedure of ContextLoRA, aiming to minimize computational and communication costs through a strategic grouping mechanism. Experiments on three benchmarks show the superiority of the proposed ContextLoRA and ContextGear. Furthermore, we prototype our proposed paradigm on a real-world wireless testbed, demonstrating its practical applicability for various IMAs. We will release our code to the community. Xinye Cao, Hongcan Guo, Guoshun Nan, Jiaoyang Cui, Haoting Qian, Yihan Lin 0001, Yilin Peng, Diyang Zhang, Yan-Zhao Hou, Huici Wu, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Exploring LLM-Based Multi-Agent Situation Awareness for Zero-Trust Space-Air-Ground Integrated NetworkabstractSpace-air-ground integrated network (SAGIN), which integrates satellite systems, aerial networks, and terrestrial communications, offers ubiquitous coverage for a multitude of applications. Nevertheless, the highly dynamic and open nature of SAGIN increases the network’s vulnerability. Hence, zero-trust security, operating on the principle of “never trust, always verify”, holds the significant potential of securing SAGIN. However, implementing zero-trust SAGIN in practice presents three primary challenges: 1) understanding massive unstructured threat information across diverse domains, 2) performing adaptive security assessments, and 3) making in-depth security decisions. This motivates us to propose SAG-Attack and LLM-SA to enhance zero-trust SAGIN. SAG-Attack serves as a simulator that aims to mimic various attacks in SAGIN. Our LLM-SA is a novel situation awareness method that explores the multiple agents of large language model (LLM). Specifically, the output logs of SAG-Attack will be fed into LLM-SA, and LLM-SA fuses vast amounts of heterogeneous threat information from various domains, thus tackling the first challenge. Then, our LLM-SA relies on multiple LLM-based agents to perform adaptive security assessments, utilizing the chain-of-thought capabilities of LLMs to automatically generate in-depth defense strategies, thereby addressing the second and third challenges. Experiments on five benchmarks demonstrate the superiority of the proposed SAG-Attack and LLM-SA. Notably, our method based on open-sourced Llama3-8B even outperforms ChatGPT-4 under the same setting, despite involving significantly fewer parameters. To foster further research in this area, we will release our platform to the community, facilitating the advancement of zero-trust SAGIN. Xinye Cao, Guoshun Nan, Hongcan Guo, Hanqing Mu, Yihan Lin 0001, Qinchuan Zhou, Baohua Qin, Qimei Cui, Xiaofeng Tao 0001, He Fang, Haitao Du, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | A Novel Indicator for Quantifying and Minimizing Information Utility Loss of Robot TeamsabstractThe timely exchange of information among robots within a team is vital, but it can be constrained by limited wireless capacity. The inability to deliver information promptly can result in estimation errors that impact collaborative efforts among robots. In this paper, we propose a new metric termed Loss of Information Utility (LoIU) to quantify the freshness and utility of information critical for cooperation. The metric enables robots to prioritize information transmissions within bandwidth constraints. We also propose the estimation of LoIU using belief distributions and accordingly optimize both transmission schedule and resource allocation strategy for device-to-device transmissions to minimize the time-average LoIU within a robot team. A semi-decentralized Multi-Agent Deep Deterministic Policy Gradient framework is developed, where each robot functions as an actor responsible for scheduling transmissions among its collaborators while a central critic periodically evaluates and refines the actors in response to mobility and interference. Simulations validate the effectiveness of our approach, demonstrating an enhancement of information freshness and utility by 98%, compared to alternative methods. Xiyu Zhao, Qimei Cui, Wei Ni 0001, Quan Z. Sheng, Abbas Jamalipour, Guoshun Nan, Xiaofeng Tao 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Disentangled Dynamic Intrusion DetectionabstractNetwork-based intrusion detection system (NIDS) monitors network traffic for malicious activities, formingthe frontline defense against increasing attacks over information infrastructures. Although promising, our quantitative analysis shows that existing methods perform inconsistently in attacks (e.g., 18% F1 for the MITM and 93% F1 for DDoS by a GCN-based state-of-the-art method), and perform poorly in few-shot intrusion detections (e.g., dramatically drops from 91% to 36% in 3D-IDS, and drops from 89% to 20% in E-GraphSAGE). We reveal that the underlying cause is entangled distributions of flow features. This motivates us to propose DIDS-MFL, a disentangled intrusion detection approach for various scenarios. DIDS-MFL involves two key components: a double Disentanglement-based Intrusion Detection System (DIDS) and a plug-and-play Multi-scale Few-shot Learning-based (MFL) intrusion detection module. Specifically, the proposed DIDS first disentangles traffic features by a non-parameterized optimization, automatically differentiating tens and hundreds of complex features. Such differentiated features will be further disentangled to highlight the attack-specific features. Our DIDS additionally uses a novel graph diffusion method that dynamically fuses the network topology for spatial-temporal aggregation in evolving data streams. Furthermore, the proposed MFL involves an alternating optimization framework to address the entangled representations in few-shot traffic threats with rigorous derivation. MFL first captures multi-scale information in latent space to distinguish attack-specific information and then optimizes the disentanglement term to highlight the attack-specific information. Finally, MFL fuses and alternately solves them in an end-to-end way. To the best of our knowledge, DIDS-MFL takes the first step toward disentangled dynamic intrusion detection under various attack scenarios. Equipped with DIDS-MFL, administrators can effectively identify various attacks in encrypted traffic, including known, unknown, and few-shot threats that are not easily detected. Comprehensive experiments show the superiority of our proposed DIDS-MFL. For few-shot NIDS, our DIDS-MFL achieves a 71.91% -125.19% improvement in average F1-score over 14 baselines and shows versatility in multiple baselines and multiple tasks. Chenyang Qiu 0001, Guoshun Nan, Hongrui Xia, Zheng Weng, Meng Shen 0001, Xiaofeng Tao 0001, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Google Map-Based Password Authentication Systems Using Tolerant Distance and Homomorphic EncryptionabstractPasswords are widely used for authentication in Internet applications. Recently, users tend to adopt graphical passwords instead of traditional alphanumeric passwords, since it is much easier for humans to remember images than verbal representations. However, the existing graphical password authentication systems generally suffer from three main issues. 1) It is required to remember and perform complicated operations during the registration/login phases, which significantly limits the systems’ usability; 2) The users’ passwords are simply stored as plaintexts in servers, and thus the security is compromised; 3) The users need to register/login to each server separately when they are applied in multi-server environment. To address the above issues, we propose a user-friendly and secure Google map-based graphical password (FS-GMGP) system using tolerant distance and homomorphic encryption. By using a homomorphic encryption scheme, each user encrypts his password point and response point selected on Google map, while the servers compute and decrypt the distance between the two encrypted points and then compare the resulting value with a tolerant distance for authentication. Moreover, the FS-GMGP system is extended for multi-server environment. The evaluation results and security analysis show that the FS-GMGP and its extended version achieve desirable usability and security in single-server environment and multi-server environment, respectively. Zhili Zhou 0001, Ching-Nung Yang, Shaowei Wang 0003, Guoshun Nan, Stelvio Cimato, Yifeng Zheng 0001, Qian Wang 0002 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Malsight: Exploring Malicious Source Code and Benign Pseudocode for Iterative Binary Malware SummarizationabstractBinary malware summarization aims to automatically generate human-readable descriptions of malware behaviors from executable files, facilitating tasks like malware cracking and detection. Previous methods based on Large Language Models (LLMs) have shown great promise. However, they still face significant issues, including poor usability, inaccurate explanations, and incomplete summaries, primarily due to the obscure pseudocode structure and the lack of malware training summaries. Further, calling relationships between functions, which involve the rich interactions within a binary malware, remain largely underexplored. To this end, we propose MALSIGHT, a novel code summarization framework that can iteratively generate descriptions of binary malware by exploring malicious source code and benign pseudocode. Specifically, we construct the first malware summary dataset, MalS and MalP, using an LLM and manually refine this dataset with human effort. At the training stage, we tune our proposed MalT5, a novel LLM-based code model, on the MalS and benign pseudocode datasets. Then, at the test stage, we iteratively feed the pseudocode functions into MalT5 to obtain the summary. Such a procedure facilitates the understanding of pseudocode structure and captures the intricate interactions between functions, thereby benefiting summaries’ usability, accuracy, and completeness. Additionally, we propose a novel evaluation benchmark, BLEURT-sum, to measure the quality of summaries. Experiments on three datasets show the effectiveness of the proposed MALSIGHT. Notably, our proposed MalT5, with only 0.77B parameters, delivers comparable performance to much larger Code-Llama. Haolang Lu, Hongrui Peng, Guoshun Nan, Jiaoyang Cui, Weifei Jin, Shengli Pan 0001, Xiaofeng Tao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Mining Multi-Scale Spatial-Frequency Clues for Unsupervised Intrusion DetectionabstractUnsupervised network-based intrusion detection system (UNIDS) identifies suspicious traffic and alerts administrators without using any traffic labels. Existing Graph Convolutional Network (GCN)-based UNIDS approaches show great potential with collaboratively utilizing traffic features and network topologies. However, these methods suffer from an excessively high false-positive rate (FPR), e.g., 3.27% FPR in a supervised NIDS approach, while increases dramatically to 19.37% under the unsupervised setting. We reveal that the high FPR stems from a single-scale spatial-frequency learning paradigm, which blurs the distinction between benign and malicious traffic and misleads UNIDS systems. Therefore, we propose a Multi-scale Spatial-Frequency Intrusion Detection System (MSF-IDS) to mitigate the high FPR. Specifically, we propose multi-scale frequency encoders, thereby differentiating abnormal feature patterns. Then we propose NAPH as a spatial encoder, mining intrinsic abnormal topology patterns by tracking tens of thousands of evolving traffic nodes. To the best of our knowledge, NAPH takes the first step toward differentiable persistent homology analysis over dynamic network data. We also develop an executable application for NAPH to provide easy-access visualization insights. Finally, a self-supervised representation augmentor and an intrusion detector are proposed to refine and highlight the attack-specific information. Equipped with MSF-IDS, administrators effectively identify the unknown attack traffic, freeing security staff from labor-intensive engineering. Extensive experiments demonstrate the superiority of MSF-IDS, including binary classification, multi-classification, online intrusion detection, and visualized discussions. Our codes, datasets, and an executable application are available at https://github.com/qcydm/MSF-IDS. Chenyang Qiu 0001, Guoshun Nan, Caiyi Zhang, Chenrui Liang, Ruiqi Dai, Hongchen Yang, Changhua Pei, Xiaofeng Tao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Semantic Entropy Can Simultaneously Benefit Transmission Efficiency and Channel Security of Wireless Semantic CommunicationsabstractRecently proliferated deep learning-based semantic communications (DLSC) focus on how transmitted symbols efficiently convey a desired meaning to the destination. However, the sensitivity of neural models and the openness of wireless channels cause the DLSC system to be extremely fragile to various malicious attacks. This inspires us to ask a question: “Can we further exploit the advantages of transmission efficiency in wireless semantic communications while also alleviating its security disadvantages?”. Keeping this in mind, we propose SemEntropy, a novel method that answers the above question by exploring the semantics of data for both adaptive transmission and physical layer encryption. Specifically, we first introduce semantic entropy, which indicates the expectation of various semantic scores regarding the transmission goal of the DLSC. Equipped with such semantic entropy, we can dynamically assign informative semantics to Orthogonal Frequency Division Multiplexing (OFDM) subcarriers with better channel conditions in a fine-grained manner. We also use the entropy to guide semantic key generation to safeguard communications over open wireless channels. By doing so, both transmission efficiency and channel security can be simultaneously improved. Extensive experiments over various benchmarks show the effectiveness of the proposed SemEntropy. We discuss the reason why our proposed method benefits secure transmission of DLSC, and also give some interesting findings, e.g., SemEntropy can keep the semantic accuracy remain 95% with 60% less transmission. Yankai Rong, Guoshun Nan, Minwei Zhang, Xuefei Zhang 0003, Nan Ma 0014, Shixun Gong, Zhaohui Yang 0001, Qimei Cui, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Secret Key Generation With Untrusted Internal Eavesdropper: Token-Based Anti-EavesdroppingabstractPhysical layer (PHY) secret key generation (SKG) has been widely studied as a promising approach to achieving One-Time-Pad security. The improvement of SKG rate is quite a huge challenge, especially in scenarios with untrusted internal helpers or eavesdroppers that aim to wiretap the negotiated secret keys between legitimate parties. In this paper, we propose a token-based SKG scheme to deal with the problem of information leakage with internal eavesdropping attacks. The basic idea is to cover random pilots with protective tokens to confuse eavesdroppers. Three scenarios including passive external eavesdropping, active internal eavesdropping with a reconfigurable intelligent surface (RIS)-assisted untrusted helper, and active internal eavesdropping with an untrusted relay are considered and analyzed to evaluate the performance of the proposed anti-eavesdropping scheme. Theoretical analysis shows that the proposed token-based SKG scheme can perfectly secure the key negotiation, achieving zero information leakage even in the untrusted relaying scenario without a direct link between Alice and Bob. Moreover, closed-form expressions for secret key capacity (SKC) are obtained. Finally, numerical results indicate that the proposed scheme outperforms the state-of-the-art methods. Using a token-generation mapping function with greater diversity in amplitude and phase, our approach achieves enhanced SKC performance across various scenarios, including those with a passive eavesdropper, a RIS-assisted untrusted helper, and an untrusted relay. Huici Wu, Na Li 0001, Xin Yuan 0004, Zhiqing Wei, Guoshun Nan, Xiaofeng Tao 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | ReTrial: Robust Encrypted Malicious Traffic Detection via Discriminative Relation Incorporation and Misleading Relation CorrectionabstractEncryption techniques greatly ensure the confidentiality and integrity of network communications. However, they also allow attackers to conceal malicious activities within encrypted traffic, posing severe cybersecurity challenges. Current detection methods primarily rely on statistics and correlation analysis. However, both statistical features and inter-entity relations can be easily obfuscated. Moreover, issues with low-quality data and fixed feature sets limit the generalizability and adaptability to defend against various evasion techniques. Robustifying encrypted malicious traffic detection in adverse conditions is still an open problem. In this paper, we propose ReTrial, a robust encrypted malicious traffic detection system via discriminative relation incorporation and misleading relation correction. The key motivations behind ReTrialare to accurately leverage the rich relations among flows for contextual analysis, and correct misleading ones for robust threat detection. Specifically, we construct a relational multigraph and develop a tailored Graph Attention Network (GAT) to selectively incorporate contextual information. Then we retrieve multi-order neighborhood similarity graphs as observations for adaptive relation correction. Following an iterative scheme, both detector performance and graph topology mutually optimize. To validate the robustness of ReTrial, we simulate various adverse conditions by randomly dropping packets and greedily injecting perturbation edges. The experimental results show that ReTrialis competitive in ideal condition. Under adverse conditions, though the performances of other state-of-the-art methods degrade significantly, ReTrialconsistently exhibits superior performance with a maximum reduction of only 5.88% in F1, highlighting its robustness in threat detection. Jianjin Zhao, Qi Li 0057, Zewei Han, Junsong Fu 0001, Guoshun Nan, Meng Shen 0001, Bharat K. Bhargava |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Data-Driven Cyber-Physical Anomaly Detection With GAN in Federated Smart FactoriesabstractResilient operation of a wireless networked multirobot system (MRS) in a smart factory relies on the effective detection of physical anomalies from robots and cyber anomalies from wireless transmission errors or imprecise artificial intelligence decisions, which leads to a new technological frontier in data-driven industrial informatics: cyber-physical anomaly detection (AD). Furthermore, data patterns in a single smart factory are unlikely enough to train high-quality learning models for this new cyber-physical AD, which suggests the necessity to utilize operating data from multiple smart factories while keeping the privacy of each factory's data. To overcome the aforementioned technical challenges for cyber-physical AD in smart factories, this article proposes an integral mechanism of generative adversarial networks, federated learning, and fuzzy clustering acceleration. Generative adversarial networks facilitate data imputation to regenerate complete datasets alleviating anomalies caused by wireless communications. Federated learning enables rich privacy-preserving datasets to be jointly used among multiple collaborative factories. Furthermore, fuzzy clustering acceleration is embedded to speed up the factory selection algorithm such that efficient training and real-time physical AD in the large-scale operation of multiple smart factories can be achieved. Extensive computational experiments based on the KDD-99 dataset demonstrate the effective and efficient cyber-physical AD of wireless networked MRS in collaborative multiple smart factories. Yaxin Liao, Yingze Wang, Qimei Cui, Kwang-Cheng Chen, Guoshun Nan, Xiaofeng Tao 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Plugging and Breathing on the Air: A Practical Defense System for Deep Learning-Based Wireless Semantic CommunicationsabstractDeep learning-based semantic communications (DLSC) leverage deep neural networks in transmitters and receivers, pushing the boundaries beyond Shannon limit. However, DLSC is extremely vulnerable to malicious physical-layer adversarial attacks due to the openness of wireless channels. Meanwhile, existing defense approaches still suffer from two challenges for robust DLSC. First, most methods require offline DLSC retraining to defend against various attacks, causing interruptions of online service. Second, they struggle to achieve effective defense in real-world time-varying channels, thus limiting DLSC reliability. We propose PBNet, integrating a pluggable protector and an adaptive protector to respectively address the above two challenges. First, the pluggable protector utilizes a novel denoising module to safeguard the transmitted signals, enabling hot-pluggable deployment without interrupting communication. Second, the adaptive protector leverages a novel alternating adaption strategy to achieve effective defense in time-varying channels, ensuring robust performances under real-world dynamic conditions. Evaluations involving symbols, images, texts, and speeches show the efficacy of our PBNet, which has respectively achieved an impressive 72.22% and 73.71% accuracy improvement in defending against unknown$l_{0}$-norm and$l_{2}$-norm attacks on image-based DLSC. Furthermore, we developed two real-world radio systems of PBNet to perform over-the-air signal generation, integrating hardware and software such as FPGA chips and GNU radio. We also implemented an interactive UI of PBNet based on QT5, aiming to demonstrate the effect of attacks and defense visually. This work achieves robust DLSC performances under various attacks and time-varying channels, taking a significant step towards the practical defense scheme for robust DLSC. Chenyang Qiu 0001, Guoshun Nan, Ruiwen Liang, Wendi Deng, Yuchong Gao, Di Wang 0011, Meng Qu, Zhuoran Duan, Qianlong Sun, Qimei Cui, Xiaodong Xu 0001, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm UnderstandingabstractMultimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks and approaches usually focus on sentence-level MSU. In document-level news, sarcasm clues are sparse or small and are often concealed in long text. Moreover, compared to sentence-level comments like tweets, which mainly focus on only a few trends or hot topics (e.g., sports events), content in the news is considerably diverse. Models created for sentence-level MSU may fail to capture sarcasm clues in document-level news. To fill this gap, we present a comprehensive benchmark for Document-level Multimodal Sarcasm Understanding (DocMSU). Our dataset contains 102,588 pieces of news with text-image pairs, covering 9 diverse topics such as health, business, etc. The proposed large-scale and diverse DocMSU significantly facilitates the research of document-level MSU in real-world scenarios. To take on the new challenges posed by DocMSU, we introduce a fine-grained sarcasm comprehension method to properly align the pixel-level image features with word-level textual features in documents. Experiments demonstrate the effectiveness of our method, showing that it can serve as a baseline approach to the challenging DocMSU. Guoshun Nan, Binzhu Xie, Junrui Xu, Hehe Fan, Qimei Cui, Xiaofeng Tao 0001 |
AAAI | 2 |
| 2024 | Refining Latent Homophilic Structures over Heterophilic Graphs for Robust Graph Convolution NetworksabstractGraph convolution networks (GCNs) are extensively utilized in various graph tasks to mine knowledge from spatial data. Our study marks the pioneering attempt to quantitatively investigate the GCN robustness over omnipresent heterophilic graphs for node classification. We uncover that the predominant vulnerability is caused by the structural out-of-distribution (OOD) issue. This finding motivates us to present a novel method that aims to harden GCNs by automatically learning Latent Homophilic Structures over heterophilic graphs. We term such a methodology as LHS. To elaborate, our initial step involves learning a latent structure by employing a novel self-expressive technique based on multi-node interactions. Subsequently, the structure is refined using a pairwisely constrained dual-view contrastive learning approach. We iteratively perform the above procedure, enabling a GCN model to aggregate information in a homophilic way on heterophilic graphs. Armed with such an adaptable structure, we can properly mitigate the structural OOD threats over heterophilic graphs. Experiments on various benchmarks show the effectiveness of the proposed LHS approach for robust GCNs. Chenyang Qiu 0001, Guoshun Nan, Tianyu Xiong, Wendi Deng, Di Wang 0011, Zhiyang Teng, Qimei Cui, Xiaofeng Tao 0001 |
AAAI | 2 |
| 2024 | Empowering Seamless Handover Authentication for High-speed UEs via Dual-blockchain over STINsabstractSatellite-terrestrial integrated networks (STINs) offer wide coverage and ubiquitous connectivity for massive mobile equipments (UEs). However, the security issues in STINs should not be ignored. Authentication and key agreement is the fundamental way to protect STINs from unauthorized access. Nevertheless, high-speed UEs, such as those in vehicles and trains, may encounter frequent handovers among heterogeneous access points in STINs, resulting in reduced authentication efficiency and an increased rate of access failures. To tackle these challenging issues, this paper proposes DBC-Auth, a seamless and universal handover authentication scheme that relies on a dual-blockchain architecture supporting all handover scenarios over STINs. Specifically, the proposed dual-blockchain architecture involves two chains that reside on the ground base stations and satellites, respectively, to efficiently manage the handover authentication in different segments of STINs. Leveraging the blockchain’s global availability and tamper-resistance, the authentication receipts can be recorded on the blockchain in advance, which can significantly improve the handover authentication efficiency and alleviate the handover failures. Security analysis demonstrates that our scheme achieves various security requirements, and performance evaluation shows superior efficiency compared to existing works. Shiyun Xie, Guoshun Nan, Xiaofeng Tao 0001 |
APCC | 2 |
| 2024 | Towards Robust Temporal Activity Localization Learning with Noisy LabelsabstractThis paper addresses the task of temporal activity localization (TAL). Although recent works have made significant progress in TAL research, almost all of them implicitly assume that the dense frame-level correspondences in each video-query pair are correctly annotated. However, in reality, such an assumption is extremely expensive and even impossible to satisfy due to subjective labeling. To alleviate this issue, in this paper, we explore a new TAL setting termed Noisy Temporal activity localization (NTAL), where a TAL model should be robust to the mixed training data with noisy moment boundaries. Inspired by the memorization effect of neural networks, we propose a novel method called Co-Teaching Regularizer (CTR) for NTAL. Specifically, we first learn a Gaussian Mixture Model to divide the mixed training data into preliminary clean and noisy subsets. Subsequently, we refine the labels of the two subsets by an adaptive prediction function so that their true positive and false positive samples could be identified. To avoid single model being prone to its mistakes learned by the mixed data, we adopt a co-teaching paradigm, which utilizes two models sharing the same framework to teach each other for robust learning. A curriculum strategy is further introduced to gradually learn the moment confidence from easy to hard. Experiments on three datasets demonstrate that our CTR is significantly more robust to the noisy training data compared to the existing methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Guoshun Nan, Keke Tang, Wanlong Fang, Yu Cheng 0001 |
LREC/COLING | 6 |
| 2024 | Uncovering what, why and How: A Comprehensive Benchmark for Causation Understanding of Video AnomalyabstractVideo anomaly understanding (VAU) aims to automat-ically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial manufacturing. While existing VAU benchmarks primarily concentrate on anomaly detection and localization, our focus is on more practicality, prompting us to raise the following crucial questions: “what anomaly occurred?”,”why did it happen?”, and “how severe is this abnormal event?”. In pursuit of these answers, we present a comprehensive benchmark for Causation Understanding of Video Anomaly (CUVA). Specifically, each instance of the proposed benchmark involves three sets of human annotations to indicate the”what”, “why” and “how” of an anomaly, including 1) anomaly type, start and end times, and event descriptions, 2) natural language explanations for the cause of an anomaly, and 3) free text reflecting the effect of the abnormality. In addition, we also introduce MMEval, a novel evaluation metric designed to better align with human preferences for CUVA, facilitating the measurement of existing LLMs in comprehending the underlying cause and corresponding effect of video anoma-lies. Finally, we propose a novel prompt-based method that can serve as a baseline approach for the challenging CUVA. We conduct extensive experiments to show the superiority of our evaluation metric and the prompt-based approach. Our code and dataset are available at https://github.com/fesvhtr/CUVA. Binzhu Xie, Guoshun Nan, Junrui Xu, Hangyu Liu 0001, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Linli Chen, Xuhuan Li, Jianhang Chen, Qimei Cui, Xiaofeng Tao 0001 |
CVPR | 4 |
| 2024 | GROSS: One-time Secret Sharing Can Make Group-based Authentication More EfficientabstractGroup-based authentication allows users within a single domain and group to access networks without repeating an individual authentication instance, greatly reducing the energy consumption of low-resource mobile devices in IoT and M2M communications. Nevertheless, in the case of the upcoming 6G massive communications with an exponentially larger number of connections, the computation and communication overhead of existing approaches on mobile devices are still significant. To this end, we propose GROSS, a novel GRoup-based authentication and key agreement (AKA) protocol that uses a One-time Secure Secret-sharing mechanism for more efficient authentication over massive wireless communications. Specifically, we employ a lightweight cryptographic operation for the above one-time secret sharing. The proposed GROSS significantly reduces both computation and communication overhead by consistently maintaining the validity of credentials for group-based authentication, thus enabling efficient verification of device legitimacy within a group. We also implement a simulation platform on JAVA for energy consumption evaluations for massive wireless communications. Our platform facilitates the flexible configuration of various energy components for authentication and supports up to million-level wireless connections. We conduct extensive experiments to show the effectiveness of our proposed GROSS. Yuandong Wu, Guoshun Nan, Jianlong Ban, Hanqing Mu, He Fang, Qimei Cui, Xiaofeng Tao 0001, Pengxuan Mao, Tianyuan Yang |
GLOBECOM | 2 |
| 2024 | Modeling and Analysis of Over-the-Air Attack with QoS-Aware Scheduling: Queuing-based ApproachabstractOver-the-air (OTA) attacks, such as flooding and jamming, significantly compromise the availability and reliability of radio access networks, especially with Quality of Service (QoS)-aware scheduling. Despite extensive research aimed at enhancing network performance and efficiency, the impact of OTA attacks on QoS performance has been insufficiently addressed. This paper conducts a thorough analysis of how OTA attacks influence the network performance. By leveraging a multi-class M/M/1 queuing model with selection and feedback mechanisms, the complex relationship between attack strategies and prioritized network scheduling is analyzed. Expressions for the delay and throughput are derived based on the result of the stationary distribution of Continuous Time Markov Chain (CTMC) model. Finally, numerical and simulation results are demonstrated to validate the theoretical analysis and to analyze the impact of OTA attack on the scheduling performance. Results show that the attacker can compromise the QoS of low-priority traffic by blocking high-priority traffic. Moreover, the QoS of high-priority traffic can be compromised by jamming low-priority traffic. Huici Wu, Guoshun Nan, Xiaofeng Tao 0001 |
GLOBECOM | 4 |
| 2024 | Can We Improve Channel Reciprocity via Loop-back Compensation for RIS-assisted Physical Layer Key GenerationabstractReconfigurable intelligent surface (RIS) facilitates the extraction of unpredictable channel features for physical layer key generation (PKG), securing communications among legitimate users with symmetric keys. Previous works have demonstrated that channel reciprocity plays a crucial role in generating symmetric keys in PKG systems, whereas, in reality, reciprocity is greatly affected by hardware interference and RIS-based jamming attacks. This motivates us to propose LoCKey, a novel approach that aims to improve channel reciprocity by mitigating interferences and attacks with a loop-back compensation scheme, thus maximizing the secrecy performance of the PKG system. Specifically, our proposed LoCKey is capable of effectively compensating for the CSI non-reciprocity by the combination of transmit-back signal value and error minimization module. Firstly, we introduce the entire flowchart of LoCKey and provide an in-depth discussion of each step. Following that, we delve into a theoretical analysis of the performance optimizations when our LoCKey is applied for CSI reciprocity enhancement. Finally, we conduct experiments to verify the effectiveness of the proposed LoCKey in improving channel reciprocity under various interferences for RIS-assisted wireless communications. The results demonstrate a significant improvement in both the rate of key generation assisted by the RIS and the consistency of the generated keys, showing great potential for the practical deployment of our LoCKey in future wireless systems. Ningya Xu, Guoshun Nan, Xiaofeng Tao 0001, Na Li 0001, Pengxuan Mao, Tianyuan Yang |
ICC | 2 |
| 2024 | Secure Transmission for MISO Integrated Sensing and Communication Secrecy SystemsabstractIntegrated Sensing and Communication (ISAC) can reuse the same spectrum and hardware resources for communication and radar sensing, which is a promising technology to alleviate spectrum congestion. Recently, there has been an increasing research interest in the physical layer secrecy aspects of ISAC systems. This paper studies a basic multiple-input single-output ISAC secrecy system, where a multi-antenna base station transmits a unified signal to sense a point target and communicate to a single-antenna receiver in the presence of$K$single antenna eavesdroppers. Under this setup, a sensing signal-to-noise ratio (SNR) constrained secrecy rate maximization problem has been formulated. We showed that the problem can be reformulated as a convex semidefinite programming problem and proved that beamforming is the optimal transmission scheme. We derived two low-complexity semiclosed-form optimal beamforming solutions and one suboptimal closed-form solution. Numerical results demonstrate that the proposed solutions attain the optimum of the sensing SNR constrained secrecy rate maximization problem and have lower computational complexity than the existing method. Zengbao Zhu, Qimei Cui, Guoshun Nan, Xuefei Zhang 0003 |
ICC | 4 |
| 2024 | Anti-Quantum Certificateless Group Authentication for Massive Accessing IoT DevicesabstractInternet of Things (IoT) is one of the most representative application scenarios in the 5G and 6G era. The concurrent access of massive IoT devices definitely poses enormous communication, computation, and certificate management challenges to the wireless authentication. Moreover, the emergence of quantum computing makes classical cryptography-based authentication protocols, such as 5G-AKA, more easier to be broken. Facing the challenges posed by the massive concurrent authentication and quantum attacks, this paper proposes a lattice cryptography based group authentication scheme, where lattice-based aggregate signature algorithm and identity-based encryption (IBE) are leveraged to achieve simultaneous authentication of concurrent accessed devices. The proposed authentication scheme eliminates the process of public key certificate management, greatly reducing the storage overhead of core network. Moreover, the utilization of lattice cryptography enables the resistance of quantum attacks. The proposed solution does not rely on additional security assumptions such as security channel or trusted group center, making it more flexible to be deployed in actual network scenario. Finally, formal security analysis of the proposed protocol is provided with the tool ProVerif. It is demonstrated that the proposed protocol can satisfy the goals of identity privacy, authentication, data confidentiality and forward secrecy. In addition, compared with existing advanced solutions, the outperformance of the proposed scheme in terms of computation overhead, signaling overhead, communication overhead, and security properties is validated with simulations. Pengbo Xu, Huici Wu, Xiaofeng Tao 0001, Chenyu Wang 0002, Dajiang Chen, Guoshun Nan |
IEEE Internet Things J. | 6 |
| 2024 | Meta Security Metric Learning for Secure Deep Image HidingabstractDeep Image Hiding (DIH) aims to imperceptibly hide images within image. To improve its security performance, some DIH methods design Security Metrics (SMs) to guide the learning of their hiding networks. However, these methods focus on optimizing their anti-steganalysis ability on specific SMs, resulting in inferior generalization ability. To overcome these limitations, in this paper, we introduce meta-learning into DIH and propose Meta Security Metric-based DIH (MSM-DIH). In the MSM-DIH, the Invertible Neural Network (INN)-based hiding network is learned under the guidance of a learnable meta SM generalized from multiple fixed source SMs, and each SM is composed of a metric network and a contrastive loss function. Specifically, MSM-DIH is trained with bi-level optimization. In the outer optimization, a meta SM is learned to assign higher security scores for more advanced stego images. Besides, the domain knowledge of steganalysis is transferred from the multiple pre-trained source metric networks to the meta metric network, so as to enhance the generalization ability of the meta SM. In the inner optimization, the hiding network is learned to generate more secure stego images according to the learned meta SM. Experimental results show that our MSM-DIH has achieved the best security performance in most cases. Weixuan Tang 0004, Zhili Zhou 0001, Ruohan Meng, Guoshun Nan, Yun Q. Shi 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed VideosabstractGiven an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Although previous respectable works have made decent success, they only focus on high-level visual features extracted from the consecutive decoded frames and fail to handle the compressed videos for query modelling, suffering from insufficient representation capability and significant computational complexity during training and testing. In this paper, we pose a new setting, compressed-domain TSG, which directly utilizes compressed videos rather than fully-decompressed frames as the visual input. To handle the raw video bit-stream input, we propose a novel Three-branch Compressed-domain Spatial-temporal Fusion (TCSF) framework, which extracts and aggregates three kinds of low-level visual features (I-frame, motion vector and residual features) for effective and efficient grounding. Particularly, instead of encoding the whole decoded frames like previous works, we capture the appearance representation by only learning the I-frame feature to reduce delay or latency. Besides, we explore the motion information not only by learning the motion vector feature, but also by exploring the relations of neighboring frames via the residual feature. In this way, a three-branch spatial-temporal attention layer with an adaptive motion-appearance fusion module is further designed to extract and aggregate both appearance and motion information for the final grounding. Experiments on three challenging datasets shows that our TCSF achieves better performance than other state-of-the-art methods with lower complexity. Daizong Liu, Pan Zhou 0001, Guoshun Nan |
CVPR | 4 |
| 2023 | Passive Eavesdropping Can Significantly Slow Down RIS-Assisted Secret Key GenerationabstractReconfigurable Intelligent Surface (RIS) assisted physical layer key generation has shown great potential to secure wireless communications by smartly controlling signals such as phase and amplitude. However, previous studies mainly focus on RIS adjustment under ideal conditions, while the correlation between the eavesdropping channel and the legitimate channel, a more practical setting in the real world, is still largely under-explored for the key generation. To fill this gap, this paper aims to maximize the RIS-assisted physical-layer secret key generation by optimizing the RIS units switching under the eavesdropping channel. Firstly, we theoretically show that passive eavesdropping significantly reduces RIS-assisted secret key generation. Keeping this in mind, we then introduce a mathematical formulation to maximize the key generation rate and provide a step-by-step analysis. Extensive experiments show the effectiveness of our method in benefiting the secret key capacity under the eavesdropping channel. We also observe that the key randomness, and unmatched key rate, two metrics that measure the secret key quality, are also significantly improved, potentially paving the way to RIS-assisted key generation in real-world scenarios. Ningya Xu, Guoshun Nan, Xiaofeng Tao 0001 |
GLOBECOM | 2 |
| 2023 | Boosting Physical Layer Black-Box Attacks with Semantic Adversaries in Semantic CommunicationsabstractEnd-to-end semantic communication (ESC) system is able to improve communication efficiency by only transmitting the semantics of the input rather than raw bits. Although promising, ESC has also been shown susceptible to the crafted physical layer adversarial perturbations due to the openness of wireless channels and the sensitivity of neural models. Previous works focus more on the physical layer white-box attacks, while the challenging black-box ones, as more practical adversaries in real-world cases, are still largely under-explored. To this end, we present SemBLK, a novel method that can learn to generate destructive physical layer semantic attacks for an ESC system under the black-box setting, where the adversaries are imperceptible to humans. Specifically, 1) we first introduce a surrogate semantic encoder and train its parameters by exploring a limited number of queries to an existing ESC system. 2) Equipped with such a surrogate encoder, we then propose a novel semantic perturbation generation method to learn to boost the physical layer attacks with semantic adversaries. Experiments on two public datasets show the effectiveness of our proposed SemBLK in attacking the ESC system under the black-box setting. Finally, we provide case studies to visually justify the superiority of our physical layer semantic perturbations. Zeju Li, Xinghan Liu, Guoshun Nan, Jinfei Zhou, Xinchen Lyu, Qimei Cui, Xiaofeng Tao 0001 |
ICC | 3 |
| 2023 | Securing Semantic Communications with Physical-Layer Semantic Encryption and ObfuscationabstractDeep learning based semantic communication (DLSC) systems have shown great potential of making wireless networks significantly more efficient by only transmitting the semantics of the data. However, the open nature of wireless channel and fragileness of neural models cause DLSC systems extremely vulnerable to various attacks. Traditional wireless physical layer key (PLK), which relies on reciprocal channel and randomness characteristics between two legitimate users, holds the promise of securing DLSC. The main challenge lies in generating secret keys in the static environment with ultra-low/zero rate. Different from prior efforts that use relays or reconfigurable intelligent surfaces (RIS) to manipulate wireless channels, this paper proposes a novel physical layer semantic encryption scheme by exploring the randomness of bilingual evaluation understudy (BLEU) scores in the field of machine translation, and additionally presents a novel semantic obfuscation mechanism to provide further physical layer protections. Specifically, 1) we calculate the BLEU scores and corresponding weights of the DLSC system. Then, we generate semantic keys (SKey) by feeding the weighted sum of the scores into a hash function. 2) Equipped with the SKey, our proposed subcarrier obfuscation is able to further secure semantic communications with a dynamic dummy data insertion mechanism. Experiments show the effectiveness of our method, especially in the static wireless environment. Yankai Rong, Guoshun Nan, Shaokang Wu, Xuefei Zhang 0003, Qimei Cui, Xiaofeng Tao 0001 |
ICC | 3 |
| 2023 | 3D-IDS: Doubly Disentangled Dynamic Intrusion DetectionabstractNetwork-based intrusion detection system (NIDS) monitors network traffic for malicious activities, forming the frontline defense against increasing attacks over information infrastructures. Although promising, our quantitative analysis shows that existing methods perform inconsistently in declaring various unknown attacks (e.g., 9% and 35% F1 respectively for two distinct unknown threats for an SVM-based method) or detecting diverse known attacks (e.g., 31% F1 for the Backdoor and 93% F1 for DDoS for a GCN-based state-of-the-art method), and reveals that the underlying cause is entangled distributions of flow features. This motivates us to propose 3D-IDS, a novel method that aims to tackle the above issues through two-step feature disentanglements and a dynamic graph diffusion scheme. Specifically, we first disentangle traffic features by a non-parameterized optimization based on mutual information, automatically differentiating tens and hundreds of complex features of various attacks. Such differentiated features will be fed into a memory model to generate representations, which are further disentangled to highlight the attack-specific features. Finally, we use a novel graph diffusion method that dynamically fuses the network topology for spatial-temporal aggregation in evolving data streams. By doing so, we can effectively identify various attacks in encrypted traffics, including unknown threats and known ones that are not easily detected. Experiments show the superiority of our 3D-IDS. We also demonstrate that our two-step feature disentanglements benefit the explainability of NIDS. Chenyang Qiu 0001, Yingsheng Geng, Junrui Lu, Kaida Chen, Shitong Zhu, Ya Su, Guoshun Nan, Junsong Fu 0001, Qimei Cui, Xiaofeng Tao 0001 |
KDD | 7 |
| 2023 | Filling the Information Gap between Video and Query for Language-Driven Moment RetrievalabstractThis paper addresses the challenging task of language-driven moment retrieval. Previous methods are typically trained to localize the target moment corresponding to a single sentence query in a complicated video. However, this specific moment generally delivers richer contents than the query, i.e., the semantics of one query may miss certain object details or actions in the complex foreground-background visual contents. Such information imbalance between two modalities makes it difficult to finely align their representations. To this end, instead of training with a single query, we propose to utilize the diversity and complementarity among different queries corresponding to the same video moment for enriching the textual semantics. Specifically, we develop a Teacher-Student Moment Retrieval (TSMR) framework to fill this cross-modal information gap. A teacher model is trained to not only encode a certain query but also capture extra complementary queries to aggregate contextual semantics for obtaining more comprehensive moment-related query representations. Since the additional queries are inaccessible during inference, we further introduce an adaptive knowledge distillation mechanism to train a student model with a single query input by selectively absorbing the knowledge from the teacher model. In this manner, the student model is more robust to the cross-modal information gap during the moment retrieval guided by a single query. Experimental results on two benchmarks demonstrate the effectiveness of our proposed method. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan, Pan Zhou 0001, Zichuan Xu, Lixing Chen, Yu Cheng 0001 |
ACM Multimedia | 4 |
| 2023 | Mining KPI correlations for non-parametric anomaly diagnosis in wireless networks
Tengfei Sui, Xiaofeng Tao 0001, Huici Wu, Xuefei Zhang 0003, Jin Xu 0001, Guoshun Nan |
Sci. China Inf. Sci. | 6 |
| 2023 | Physical-Layer Adversarial Robustness for Deep Learning-Based Semantic CommunicationsabstractEnd-to-end semantic communications (ESC) rely on deep neural networks (DNN) to boost communication efficiency by only transmitting the semantics of data, showing great potential for high-demand mobile applications. We argue that central to the success of ESC is the robust interpretation of conveyed semantics at the receiver side, especially for security-critical applications such as automatic driving and smart healthcare. However, robustifying semantic interpretation is challenging as ESC is extremely vulnerable to physical-layer adversarial attacks due to the openness of wireless channels and the fragileness of neural models. Toward ESC robustness in practice, we ask the following two questions: Q1: For attacks, is it possible to generate semantic-oriented physical-layer adversarial attacks that are imperceptible, input-agnostic and controllable? Q2: Can we develop a defense strategy against such semantic distortions and previously proposed adversaries? To this end, we first presentMobileSC, a novel semantic communication framework that considers the computation and memory efficiency in wireless environments. Equipped with this framework, we proposeSemAdv, a physical-layer adversarial perturbation generator that aims to craft semantic adversaries over the air with the abovementioned criteria, thus answering the Q1. To better characterize the real-world effects for robust training and evaluation, we further introduce a novel adversarial training method$\texttt {SemMixed}$to harden the ESC againstSemAdvattacks and existing strong threats, thus answering the Q2. Extensive experiments on three public benchmarks verify the effectiveness of our proposed methods against various physical adversarial attacks. We also show some interesting findings, e.g., ourMobileSCcan even be more robust than classical block-wise communication systems in the low SNR regime. Guoshun Nan, Zhichun Li, Jinli Zhai, Qimei Cui, Gong Chen 0012, Xuefei Zhang 0003, Xiaofeng Tao 0001, Zhu Han 0001, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Doubled coupling for image emotion distribution learning
Huiyan Wu, Guoshun Nan |
Knowl. Based Syst. | 3 |
| 2022 | Semantic Reasoning with NLI for Assertion Detection in Medical TextabstractAssertion information is of crucial importance for constructing an intelligent diagnosis system as it contains clinical findings and decision basis of clinicians in the electronic medical records (EMRs), e.g., whether a symptom is present or not. Current work mainly treats assertion detection as a sequence labeling task, or constructs rule-based methods. However, there is a challenge that to detect assertions embedded in the context with long-range dependencies, needs considering the whole text to capture the complex linguistic and underlying semantic information. To tackle the above issues, we consider assertion detection as a semantic reasoning task based on natural language inference (NLI). First, we generate candidate spans with boundary detection on the basis of which we can enrich the training corpus with external knowledge such as assertion definitions. Then we detect the assertions through the NLI-based classification. To the best of our knowledge, we build the first Chinese assertion dataset, which contains 4237 sentences on privacy de-identified ophthalmology remote reading reports. Extensive experiments demonstrate that our proposed method achieves the best results on both of the English dataset i2b2 and the Chinese privacy dataset. Zongxin Du, Xiaohong Liu 0007, Guoshun Nan |
BIBM | 4 |
| 2022 | SemBAT: Physical Layer Black-box Adversarial Attacks for Deep Learning-based Semantic Communication SystemsabstractDeep learning-based semantic communications (DLSC) replace the physical blocks in traditional communication systems as end-to-end neural networks. DLSC significantly boost communication efficiency by only transmitting the meaning of data, showing great potentials for applications like automatic driving, digital twin and smart health. However, DLSC are fragile to black-box adversarial attacks due to the openness of wireless channel and sensitivities of neural models. To this end, this paper proposes SemBAT, a novel approach for crafting physical layer black-box adversarial attacks for semantic communication systems. The key ingredients of our method include the training of surrogate encoder and generation of adversarial perturbations. Specifically, we train our surrogate encoder by directly estimating the gradients based on Jacobian-matrixs, and then generate the adversarial perturbations by the particle swarm optimizations. Extensive experiments on a public benchmark show the effectiveness of our proposed SemBAT. We observe that our SemBAT with black-box adversaries can sharply decrease the classification accuracy of the semantic communication system from 78.4% to 11.6%. Meanwhile, such attacks are also imperceptible in terms of image quality metrics measured by the Structural similarity index measure (SSIM) and Peak Signal to Noise Ratio(PSNR). Zeju Li, Jinfei Zhou, Guoshun Nan, Zhichun Li, Qimei Cui, Xiaofeng Tao 0001 |
VTC Fall | 3 |
| 2022 | SemKey: Boosting Secret Key Generation for RIS-assisted Semantic Communication SystemsabstractDeep learning-based semantic communications (DLSC) significantly improve communication efficiency by only transmitting the meaning of the data rather than a raw message. Such a novel paradigm can brace the high-demand applications with massive data transmission and connectivities, such as automatic driving and internet-of-things. However, DLSC are also highly vulnerable to various attacks, such as eavesdropping, surveillance, and spoofing, due to the openness of wireless channels and the fragility of neural models. To tackle this problem, we present SemKey, a novel physical layer key generation (PKG) scheme that aims to secure the DLSC by exploring the underlying randomness of deep learning-based semantic communication systems. To boost the generation rate of the secret key, we introduce a reconfigurable intelligent surface (RIS) and tune its elements with the randomness of semantic drifts between a transmitter and a receiver. Precisely, we first extract the random features of the semantic communication system to form the randomly varying switch sequence of the RIS-assisted channel and then employ the parallel factor-based channel detection method to perform the channel detection under RIS assistance. Experimental results show that our proposed SemKey significantly improves the secret key generation rate, potentially paving the way for physical layer security for DLSC. Ningya Xu, Guoshun Nan, Qimei Cui, Xiaofeng Tao 0001 |
VTC Fall | 4 |
| 2021 | Learning Discriminative and Unbiased Representations for Few-Shot Relation ExtractionabstractFew-shot relation extraction (FSRE) aims to predict the relation for a pair of entities in a sentence by exploring a few labeled instances for each relation type. Current methods mainly rely on meta-learning to learn generalized representations by optimizing the network parameters based on various collections of tasks sampled from training data. However, these methods may suffer from two main issues. 1) Insufficient supervision of meta-learning to learn discriminative representations on very few training instances, which are sampled from a large amount of base class data. 2) Spurious correlations between entities and relation types due to the biased training procedure that focuses more on entity pair rather than context. To learn more discriminative and unbiased representations for FSRE, this paper proposes a two-stage approach via supervised contrastive learning and sentence- and entity-level prototypical networks. In the first (pre-training) stage, we introduce a supervised contrastive pre-training method, which is able to yield more discriminative representations by learning from the entire training instances, such that the semantically related representations are close to each other, and far away otherwise. In the second (meta-learning) stage, we propose a novel sentence- and entity-level prototypical network equipped with fine-grained feature-wise fusion strategy to learn unbiased representations, where the networks are initialized with the parameters trained in the first stage. Specifically, the proposed network consists of a sentence branch and an entity branch, taking entire sentences and entity mentions as inputs, respectively. The entity branch explicitly captures the correlation between entity pairs and relations, and then dynamically adjusts the sentence branch's prediction distributions. By doing so, the spurious correlations issue caused by biased training samples can be properly mitigated. Extensive experiments on two FSRE benchmarks demonstrate the effectiveness of our approach. Jiale Han 0001, Bo Cheng 0001, Guoshun Nan |
CIKM | 3 |
| 2021 | Interventional Video Grounding With Dual Contrastive LearningabstractVideo grounding aims to localize a moment from an untrimmed video for a given textual query. Existing approaches focus more on the alignment of visual and language stimuli with various likelihood-based matching or regression strategies, i.e., P(Y |X). Consequently, these models may suffer from spurious correlations between the language and video features due to the selection bias of the dataset. 1) To uncover the causality behind the model and data, we first propose a novel paradigm from the perspective of the causal inference, i.e., interventional video grounding (IVG) that leverages backdoor adjustment to deconfound the selection bias based on structured causal model (SCM) and do-calculus P(Y |do(X)). Then, we present a simple yet effective method to approximate the unobserved confounder as it cannot be directly sampled from the dataset. 2) Meanwhile, we introduce a dual contrastive learning approach (DCL) to better align the text and video by maximizing the mutual information (MI) between query and video clips, and the MI between start/end frames of a target moment and the others within a video to learn more informative visual representations. Experiments on three standard benchmarks show the effectiveness of our approaches. Guoshun Nan, Rui Qiao 0006, Jun Liu 0036, Sicong Leng, Hao Zhang 0048, Wei Lu 0011 |
CVPR | 1 |
| 2021 | Uncovering Main Causalities for Long-tailed Information ExtractionabstractInformation Extraction (IE) aims to extract structural information from unstructured texts.In practice, long-tailed distributions caused by the selection bias of a dataset, may lead to incorrect correlations, also known as spurious correlations, between entities and labels in the conventional likelihood models.This motivates us to propose counterfactual IE (CFIE), a novel framework that aims to uncover the main causalities behind data in the view of causal inference.Specifically, 1) we first introduce a unified structural causal model (SCM) for various IE tasks, describing the relationships among variables; 2) with our SCM, we then generate counterfactuals based on an explicit language structure to better calculate the direct causal effect during the inference stage; 3) we further propose a novel debiasing approach to yield more robust predictions.Experiments on three IE tasks across five public datasets show the effectiveness of our CFIE model in mitigating the spurious correlation issues. Guoshun Nan, Jiaqi Zeng, Rui Qiao 0006, Zhijiang Guo, Wei Lu 0011 |
EMNLP (1) | 1 |
| 2021 | Integrating Subgraph-Aware Relation and Direction Reasoning for Question AnsweringabstractQuestion Answering (QA) models over Knowledge Bases (KBs) are capable of providing more precise answers by utilizing relation information among entities. Although effective, most of these models solely rely on fixed relation representations to obtain answers for different question-related KB subgraphs. Hence, the rich structured information of these subgraphs may be overlooked by the relation representation vectors. Meanwhile, the direction information of reasoning, which has been proven effective for the answer prediction on graphs, has not been fully explored in existing work. To address these challenges, we propose a novel neural model, Relation-updated Direction-guided Answer Selector (RDAS), which converts relations in each subgraph to additional nodes to learn structure information. Additionally, we utilize direction information to enhance the reasoning ability. Experimental results show that our model yields substantial improvements on two widely used datasets. Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Hao Yang 0006, Ivan Sekulic, Guoshun Nan |
ICASSP | 8 |
| 2021 | Video Corpus Moment Retrieval with Contrastive LearningabstractGiven a collection of untrimmed and unsegmented videos, video corpus moment retrieval (VCMR) is to retrieve a temporal moment (i.e., a fraction of a video) that semantically corresponds to a given text query. As video and text are from two distinct feature spaces, there are two general approaches to address VCMR: (i) to separately encode each modality representations, then align the two modality representations for query processing, and (ii) to adopt fine-grained cross-modal interaction to learn multi-modal representations for query processing. While the second approach often leads to better retrieval accuracy, the first approach is far more efficient. In this paper, we propose a Retrieval and Localization Network with Contrastive Learning (ReLoCLNet) for VCMR. We adopt the first approach and introduce two contrastive learning objectives to refine video encoder and text encoder to learn video and text representations separately but with better alignment for VCMR. The video contrastive learning (VideoCL) is to maximize mutual information between query and candidate video at video-level. The frame contrastive learning (FrameCL) aims to highlight the moment region corresponds to the query at frame-level, within a video. Experimental results show that, although ReLoCLNet encodes text and video separately for efficiency, its retrieval accuracy is comparable with baselines adopting cross-modal interaction learning. Hao Zhang 0048, Aixin Sun, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh |
SIGIR | 4 |
| 2020 | HGMAN: Multi-Hop and Multi-Answer Question Answering Based on Heterogeneous Knowledge Graph (Student Abstract)abstractMulti-hop question answering models based on knowledge graph have been extensively studied. Most existing models predict a single answer with the highest probability by ranking candidate answers. However, they are stuck in predicting all the right answers caused by the ranking method. In this paper, we propose a novel model that converts the ranking of candidate answers into individual predictions for each candidate, named heterogeneous knowledge graph based multi-hop and multi-answer model (HGMAN). HGMAN is capable of capturing more informative representations for relations assisted by our heterogeneous graph, which consists of multiple entity nodes and relation nodes. We rely on graph convolutional network for multi-hop reasoning and then binary classification for each node to get multiple answers. Experimental results on MetaQA dataset show the performance of our proposed model over all baselines. Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Hao Yang 0006, Guoshun Nan |
AAAI | 7 |
| 2020 | Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionabstractDocument-level relation extraction requires integrating information within and across multiple sentences of a document and capturing complex interactions between inter-sentence entities.However, effective aggregation of relevant information in the document remains a challenging research question.Existing approaches construct static document-level graphs based on syntactic trees, co-references or heuristics from the unstructured text to model the dependencies.Unlike previous methods that may not be able to capture rich non-local interactions for inference, we propose a novel model that empowers the relational reasoning across sentences by automatically inducing the latent document-level graph.We further develop a refinement strategy, which enables the model to incrementally aggregate relevant information for multi-hop reasoning.Specifically, our model achieves an F 1 score of 59.05 on a large-scale documentlevel dataset (DocRED), significantly improving over the previous results, and also yields new state-of-the-art results on the CDR and GDA dataset.Furthermore, extensive analyses show that the model is able to discover more accurate inter-sentence relations. Guoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei Lu 0011 |
ACL | 1 |
| 2020 | Learning Latent Forests for Medical Relation ExtractionabstractThe goal of medical relation extraction is to detect relations among entities, such as genes, mutations and drugs in medical texts. Dependency tree structures have been proven useful for this task. Existing approaches to such relation extraction leverage off-the-shelf dependency parsers to obtain a syntactic tree or forest for the text. However, for the medical domain, low parsing accuracy may lead to error propagation downstream the relation extraction pipeline. In this work, we propose a novel model which treats the dependency structure as a latent variable and induces it from the unstructured text in an end-to-end fashion. Our model can be understood as composing task-specific dependency forests that capture non-local interactions for better relation extraction. Extensive results on four datasets show that our model is able to significantly outperform state-of-the-art systems without relying on any direct tree supervision or pre-training. Zhijiang Guo, Guoshun Nan, Wei Lu 0011, Shay B. Cohen |
IJCAI | 2 |
| 2020 | Interest packets scheduling and size-based flow control mechanism for content-centric networking web servers
Xiuquan Qiao, Pei Ren, Yukai Tu, Guoshun Nan, Junliang Chen 0001, M. Brian Blake |
Future Gener. Comput. Syst. | 5 |
| 2018 | The Frame Latency of Personalized Livestreaming Can Be Significantly Slowed Down by WiFiabstractThe popular personalized livestreaming (PL) in China, arguably the largest PL market in the world, is more monetized than PL in US and hence demands much lower interactive latencies to ensure a good quality of user experience. However, our pilot experiment shows that the video frame latency, dominant component of PL's interactive latency, can be significantly slowed down by WiFi, the primary Internet access method for PL. Understanding and further improving the frame latency over WiFi, however, have difficulties in 1) measuring end-to-end latency; 2) parsing encrypted PL's traffic and 3) modeling complex relationships between WiFi radio factors and the latency. To tackle these challenges, we design and prototype Latency Doctor (LTDr), a practical system which aims to model and optimize PL's video frame latency over WiFi. We deploy LTDr in our campus and obtain several key observations based on 13.9M video frames extracted from 12K individual views on three leading PLs in China. We observe that 40% frame latencies over WiFi hop are more than 30ms, and channel utilization should be less than 64% for low latency. Then we build a predictive model based on the dataset using the machine learning methodologies. Two real cases show that the median frame latencies are decreased by LTDr from 130ms to 22ms, and 50ms to 12ms respectively over WiFi networks. Guoshun Nan, Xiuquan Qiao, Jiting Wang, Zeyan Li 0001, Jiahao Bu, Changhua Pei, Mengyu Zhou, Dan Pei |
IPCCC | 1 |
| 2015 | Design and Implementation: the Native Web Browser and Server for Content-Centric NetworkingabstractContent-Centric Networking (CCN) has recently emerged as a clean-slate Future Internet architecture which has a completely different communication pattern compared with exiting IP network. Since the World Wide Web has become one of the most popular and important applications on the Internet, how to effectively support the dominant browser and server based web applications is a key to the success of CCN. However, the existing web browsers and servers are mainly designed for the HTTP protocol over TCP/IP networks and cannot directly support CCN-based web applications. Existing research mainly focuses on plug-in or proxy/gateway approaches at client and server sides, and these schemes seriously impact the service performance due to multiple protocol conversions. To address above problems, we designed and implemented a CCN web browser and a CCN web server to natively support CCN protocol. To facilitate the smooth evolution from IP networks to CCN, CCNBrowser and CCNxTomcat also support the HTTP protocol besides the CCN. Experimental results show that CCNBrowser and CCNxTomcat outperform existing implementations. Finally, a real CCN-based web application is deployed on a CCN experimental testbed, which validates the applicability of CCNBrowser and CCNxTomcat. Guoshun Nan, Xiuquan Qiao, Yukai Tu, Wei Tan 0001, Junliang Chen 0001 |
SIGCOMM | 1 |
| 2015 | NDNBrowser: An extended web browser for named data networking
Xiuquan Qiao, Guoshun Nan, Yunlei Sun, Junliang Chen 0001 |
J. Netw. Comput. Appl. | 2 |
| 2014 | CCNxTomcat: An extended web server for Content-Centric Networking
Xiuquan Qiao, Guoshun Nan, Wei Tan 0001, Junliang Chen 0001, Yukai Tu |
Comput. Networks | 2 |