Runqing Zhang

dblp:70/9540 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
abstract
Text-to-Image Person Retrieval (TIPR) aims to retrieve person images based on natural language descriptions. Although many TIPR methods have achieved promising results, sometimes textual queries cannot accurately and comprehensively reflect the content of the image, leading to poor cross-modal alignment and overfitting to limited datasets. Moreover, the inherent modality gap between text and image further amplifies these issues, making accurate cross-modal retrieval even more challenging. To address these limitations, we propose the Generation-Enhanced Alignment (GEA) from a generative perspective. GEA contains two parallel modules: 1) Text-Guided Token Enhancement (TGTE), which introduces diffusion-generated images as intermediate semantic representations to bridge the gap between text and visual patterns. These generated images enrich the semantic representation of text and facilitate cross-modal alignment. 2) Generative Intermediate Fusion (GIF) module, which combines cross-attention between generated images, original images, and text features to generate a unified representation optimized by triplet alignment loss. We conduct extensive experiments on three public TIPR datasets, CUHK-PEDES, RSTPReid, and ICFG-PEDES, to evaluate the performance of GEA. The remarkable results justify the efficacy of our method. More implementation details and extended results are available at https://github.com/sugelamyd123/Sup-for-GEA.
Runqing Zhang, Jianxiao Zou
ECAI2
2025 Malicious DoH Tunnel Traffic Identification Framework Based on DiffFlow-CNN
abstract
This study proposes DiffFlow-CNN, a novel framework for identifying malicious DNS over HTTPS (DoH) tunnel traffic, addressing the critical challenge of data imbalance in network security. By transforming network traffic into grayscale images, the framework can leverage context-rich spatial feature extraction to improve the detection accuracy. A diffusion model is employed for data augmentation, generating diverse, highquality malicious traffic samples to mitigate class imbalance. The augmented data is processed by a 2D Convolutional Neural Network (CNN), which effectively classifies traffic into NonDoH, benign DoH, and malicious DoH categories, with further differentiation of malicious types such as Iodine, dns2tcp, and DNSCat2. Experimental results on the CIRA-CIC-DoHBrw-2020 dataset demonstrate that DiffFlow-CNN achieves near-perfect performance, with an accuracy of 99.97%, precision of 99.99%, recall of 99.09%, and F1-score of 99.52% with DiffFlow-CNN. Comparative analysis highlights the superiority of bidirectional flow representations, particularly the Session+MFR method, which leverages early packet information for optimal feature capture. The framework significantly enhances the detection of covert malicious DoH traffic, offering a robust solution for network security management.
Weilin Gai, Runqing Zhang, Yunjun Ma, Peng Zhang 0044, Ruoxing Wang
HPCC3
2025 Optimized Dynamic Watermarking for Audio DNNs with Adaptive Embedding and Boundary Sampling
abstract
The intensified concerns arising from the widespread adoption of deep learning have led to increased scrutiny of intellectual property protection in DNN models. Existing audio watermarking techniques, predominantly based on traditional signal processing methods, struggle to balance robustness, imperceptibility, and defense resistance in the face of evolving adversarial attacks. These limitations underscore the urgent need for more effective watermarking solutions in the audio domain. In this paper, we propose a dynamic audio watermarking framework that introduces an optimization-based approach to attach robust and adaptable triggers at arbitrary positions within audio signals, and innovatively integrates boundary sample selection driven by forgetting events and an adaptive watermark trigger embedding technique based on the SNR. Comprehensive experimental results reveal that our scheme preserves high model performance while maintaining remarkable stealthiness and robustness, offering a secure and reliable solution for safeguarding intellectual property in the audio domain and advancing the field of DNN watermarking.
Hao Fei 0007, Hewang Nie, Songfeng Lu, Ling Qian, Dunbo Cai, Zhiguo Huang, Runqing Zhang
ICASSP9
2025 FedDiT: Federated Learning by Distillation Token Enhanced Vision Transformer
abstract
Federated learning (FL) is a promising approach for privacy-preserving machine learning, enabling collaborative model training across distributed devices without sharing raw data. However, FL faces significant challenges due to the nonindependent and identically distributed (non-IID) nature of data across devices, leading to difficulties in model convergence and generalization. In this paper, we propose FedDiT, a novel federated learning framework that combines knowledge distillation with vision transformers. FedDiT introduces the Distilled Vision Transformer (DTViT) model on the client side, incorporating a distillation token to enhance local learning and knowledge transfer. This approach significantly improves the robustness and performance of FL in non-IID environments. We validated FedDiT through extensive experiments on public datasets, and the results show that it outperforms existing FL methods in both accuracy and smoother convergence. Additionally, FedDiT achieves higher throughput compared to standard transformers and knowledge distillation methods, making it more efficient for practical deployment in federated learning scenarios.
Jue Xiao, Zepu Yi, Hewang Nie, Xueming Tang, Songfeng Lu, Zhiguo Huang, Runqing Zhang
ICASSP8
2025 AMNS: Attention-Weighted Selective Mask and Noise Label Suppression for Text-to-Image Person Retrieval
abstract
Most existing text-to-image person retrieval methods usually assume that the training image-text pairs are perfectly aligned; however, the noisy correspondence(NC) issue (i.e., incorrect or unreliable alignment) exists due to poor image quality and labeling errors. Additionally, random masking augmentation may inadvertently discard critical semantic content, introducing noisy matches between images and text descriptions. To address the above two challenges, we propose a noise label suppression method to mitigate NC and an Attention-Weighted Selective Mask (AWM) strategy to resolve the issues caused by random masking. Specifically, the Bidirectional Similarity Distribution Matching (BSDM) loss enables the model to effectively learn from positive pairs while preventing it from over-relying on them, thereby mitigating the risk of overfitting to noisy labels. In conjunction with this, Weight Adjustment Focal (WAF) loss improves the model’s ability to handle hard samples. Furthermore, AWM processes raw images through an EMA version of the image encoder, selectively retaining tokens with strong semantic connections to the text, enabling better feature extraction. Extensive experiments demonstrate the effectiveness of our approach in addressing noise-related issues and improving retrieval performance.
Runqing Zhang
ICASSP1
2025 ProfilE: Self-Supervised Communication Relationship Profiling for Imbalanced Smart Grid Anomaly Detection
abstract
Smart grids, as critical national infrastructure, require robust cybersecurity. However, traditional anomaly detection methods face inherent challenges: the scarcity of labeled attack samples due to power grid traffic characteristics, which results in highly imbalanced datasets. To overcome these limitations, this paper introduces ProfilE, a novel self-supervised anomaly detection model. The framework consists of three core modules: an encoder module is responsible for fusing historical, topological, and statistical features to generate multi-dimensional feature embeddings; a self-supervised module generates corrupted samples to mitigate the reliance on labeled data; and a Profile Module utilizes these embeddings to learn and construct a communication relationship profile that characterizes normal network behavior. By using the reconstruction error from the profile module for anomaly detection, ProfilE fundamentally eliminates reliance on labeled data. Experimental evaluations on the CICModbus2023 and CICIDS2017 datasets demonstrate that ProfilE performs better than existing baseline methods in imbalanced smart grid environments. Notably, in scenarios with scarce attack samples, its Macro F1 outperforms baselines, fully validating its effectiveness and robustness in real-world power grid settings.
Haimiao Li, Changbo Tian, Runqing Zhang, Shiyi Yuan
TrustCom3
2025 Encryption Traffic Classification Based on Mining Traffic Context and Transport Relationship
abstract
This paper proposes a novel ETC-MTCTR, which is designed to enable more accurate, versatile and efficient traffic classification in the context of multi-scenario, low-resource encrypted traffic. Through three modules of Datagram Token conversion, pretraining and fine-tuning, the method uses large-scale unlabeled encrypted traffic for pretraining, mining and learning the traffic context and transmission relationship of encrypted traffic classification tasks, so that a small number of labeled data samples can be effectively used in the fine-tuning stage. Significantly improve the performance of the model on specific downstream classification tasks, enhance the accuracy, adaptability and robustness of the model in diverse environments, limited resources and new encryption security protocols, and realize efficient encryption traffic classification in multi-scenario and low-resource background. The results show that ETC-MTCTR achieves the best performance on three tasks: encryption malware classification, VPN encrypted traffic classification, and TLS 1.3 encryption application classification. Its F1 score is improved by 0.22% in the classification task of encrypted malware, 1.4% in the classification task of VPN encrypted traffic App, 4.56% in the classification task of VPN encrypted traffic Service, and 9.89% in the classification task of TLS 1.3 encrypted application, which is significantly better than other comparison methods.
Weilin Gai, Runqing Zhang, Peng Zhang 0044
WCNC2
2025 MT-Agent: Constructing a GUI Agent via Modality Enhancement and Text-Guided Fusion
abstract
Graphical User Interfaces (GUIs) play a crucial role in facilitating user-computer interactions, making them an essential focus of research. However, current automated GUI agents face significant challenges in effectively associating task implementations with specific visual elements, and the resolution constraints of Vision-Language Models (VLMs) also limit the richness of visual information. To this end, we propose a novel multi-modal agent named MT-Agent, which enhances both textual and visual input modalities to enable the model to perceive visual elements in GUIs more effectively. Specifically, Textual Modality Enhancement improves the semantic richness of input text by capturing task-specific details via an external VLM, while Visual Modality Enhancement incorporates fine-grained visual details to better represent critical GUI elements. In addition, we introduce an innovative text-guided directional feature fusion mechanism, which leverages enriched text features to guide the integration with visual information. In experiments, MT-Agent demonstrated exceptional performance on AITZ dataset, achieving an action type prediction accuracy of 84.80% and a step prediction accuracy of 58.07%, surpassing previous state-of-the-art models. Furthermore, on the GUI Odyssey benchmark, MT-Agent achieves performance comparable to previous state-of-the-art models while using only about 1/20 of their trainable parameters. Our codes, demos, and relevant data will be released to facilitate further research and validation within the scientific community.
Jinhan Dong, Lei Jin 0003, Zhihong Zhang 0006, Runqing Zhang, Liqiang Xu, Junliang Xing
IEEE Internet Things J.5
2024 WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and Operations
abstract
Graph Neural Network (GNN) has emerged as an important workload for learning on graphs. With the size of graph data and the complexity of GNN model architectures increasing, developing an efficient GNN system grows more important. As GNN has heavy neural computation workloads on a large graph, it is crucial to partition the entire workload into smaller parts for parallel execution and optimization. However, existing approaches separately partition graph data and GNN operations, resulting in inefficiency and large data movement overhead.
Kezhao Huang, Jidong Zhai, Liyan Zheng 0001, Haojie Wang 0004, Yuyang Jin 0001, Qihao Zhang, Runqing Zhang, Zhen Zheng, Youngmin Yi, Xipeng Shen
EuroSys7
2024 Edit3D: Elevating 3D Scene Editing with Attention-Driven Multi-Turn Interactivity
abstract
With the rise of new 3D representations like NeRF and 3D Gaussian splatting, creating realistic 3D scenes is easier than ever before. However, the incompatibility of these 3D representations with existing editing software has also introduced unprecedented challenges to 3D editing tasks. Although recent advances in text-to-image generative models have made some progress in 3D editing, these methods either lack precision or require users to manually specify the editing areas in 3D space, complicating the editing process. To overcome these issues, we propose Edit3D, an innovative 3D editing method designed to enhance editing quality. Specifically, we propose a multi-turn editing framework and introduce an attention-driven open-set segmentation (ADSS) technique within this framework. ADSS allows for more precise segmentation of parts, which enhances the editing precision and minimizes interference with pixels in areas that are not being edited. Additionally, we propose a fine-tuning phase, intended to further improve the overall editing quality without compromising the training efficiency. Experiments demonstrate that Edit3D effectively adjusts 3D scenes based on textual instructions. Through continuous and multiple turns of editing, it achieves more intricate combinations, enhancing the diversity of 3D editing effects. Code is available at https://github.com/PeterouZh/Edit3D.
Peng Zhou 0010, Dunbo Cai, Yujian Du, Runqing Zhang, Bingbing Ni, Jie Qin 0004, Ling Qian
ACM Multimedia4
2023 DiffusionTracker: Targets Denoising Based on Diffusion Model for Visual Tracking
Runqing Zhang, Dunbo Cai, Ling Qian, Yujian Du, Huijun Lu
PRCV (12)1
2023 SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization
Mingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong, Runqing Zhang, Jidong Zhai
USENIX ATC5
2022 M-CoTransT: Adaptive spatial continuity in visual tracking
abstract
Abstract Visual tracking is an important area in computer vision. Based on the Siamese network, current tracking methods employ the self‐attention block in convolutional networks to extract semantic features containing the image structure information of an object. However, spatial continuity is a point of contradiction between two seemingly unrelated challenges, that is, occlusion and similar distractor, in tracking methods. At the same time, it is a spatially discontinuous task to locate a target reappearing after occlusion accurately. The prediction of bounding boxes should be constrained by spatial continuity to prevent them from jumping into similar distractors. This study proposes a novel tracking method for introducing spatial continuity in visual tracking called M‐CoTransT; the novel tracking method is developed through the confidence‐based adaptive Markov motion model (M‐model) and a novel correlation‐based feature fusion network (CoTransT). In particular, the M‐model provides confidence for the nodes of the Markov motion model to estimate the motion state continuity. It also predicts a more accurate search region for CoTransT, which then adds a cross‐correlation branch into the self‐attention tracking network to enhance the continuity of target appearance in the feature space. Extensive experiments on five challenging datasets (LaSOT, GOT‐10k, TrackingNet, OTB‐2015 and UAV123) demonstrated the effectiveness of the proposed M‐CoTransT in visual tracking.
Chunxiao Fan 0001, Runqing Zhang, Yue Ming 0001
IET Comput. Vis.2
2022 MP-LN: motion state prediction and localization network for visual object tracking
Chunxiao Fan 0001, Runqing Zhang, Yue Ming 0001
Vis. Comput.2
2021 Re-Identify Deformable Targets for Visual Tracking
Runqing Zhang, Chunxiao Fan 0001, Yue Ming 0001
PRCV (1)1
2020 An Effective Hierarchical Resolution Learning Method for Low-Resolution Targets Tracking
abstract
Suffering from the low-resolution target's visual quality, the precisions of visual object trackers are reduced. This paper proposes an effective hierarchical resolution learning method for low-resolution targets tracking, abbreviated as HRT. We adopt a hierarchical structure to exploit information from different resolution levels. (1) At the high level: the super-resolution (SR) images, determining the target's shape, contains richer image textures and clearer target contours, and transmits the search region to the low level. (2) At the low level: low-resolution (LR) images maintain the spatial structure information of the original target, providing the precise center coordinates of the target. Experimental results demonstrate the effectiveness of the proposed tracker, which HRT achieves 90.3% precision on OTB100 LR sequences and 78.5% precision on LR sequences from UAV123 datasets, gaining 2.0%, 2.4% improvement over state-of-the-art trackers respectively.
Runqing Zhang, Chunxiao Fan 0001, Yue Ming 0001, Hao Fu 0013, Xuyang Meng
ICIP1
2020 MD-ST: Monocular Depth Estimation Based on Spatio-Temporal Correlation Features
Xuyang Meng, Chunxiao Fan 0001, Yue Ming 0001, Runqing Zhang, Panzi Zhao
PRCV (1)4