VLDB 2026 Research / reviewers in the wild / expert
Zisen Qi
dblp:262/4850
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seedcap: Semantic Expansion and Entity-Driven Zero-Shot Image CaptioningabstractZero-shot image captioning (ZIC) generates natural language descriptions for images using only textual data during training. Recent progress leverages text-to-image models to construct pseudo image-text pairs, alleviating the modality mismatch between text-only training and image-based inference. However, the distribution gap between synthetic and real images can compromise performance when models trained on generated images are deployed in real-world scenarios. To address this challenge, we propose SEEDCap, a zero-shot captioning framework featuring SEED (Semantic Expansion and Entity-Driven), a parameter-free module that performs dual-level semantic expansion. SEED enhances the correspondence between image features and retrieved textual cues at both entity and sentence levels, producing enriched semantic embeddings that effectively reduce the distributional gap between synthetic and real inputs. These enhanced features are then integrated with the global visual representation through a modality fusion module, which is subsequently decoded by a large language model to generate accurate and contextually rich captions. Extensive experiments demonstrate that SEEDCap achieves state-of-the-art zero-shot results on both in-domain and cross-domain benchmarks, underscoring its robustness and practical utility. Binbin Li 0003, Shupei Xiao, Dayan Wu, Gengqi Yang, Siyu Jia, Zisen Qi |
ICMR | 6 |
| 2026 | MD-TabPFN: Multi-domain feature fusion for modulation recognition based on tabular prior data fitting networkabstractCommunication signal modulation recognition faces challenges including label scarcity, high dimensionality, nonlinearity, and real-time constraints. In recent years, deep learning-based modulation recognition has emerged as a prominent research focus; however, insufficient sample sizes frequently lead to overfitting, poor generalization, and the curse of dimensionality. This paper proposes a modulation recognition framework based on the Multi-Domain Feature Fusion and Tabular Prior-Data Fitted Network (MD-TabPFN). Adopting an optimization paradigm of general model + domain adaptation, this study represents the first attempt to transfer the tabular foundation model TabPFN to the field of communication signal recognition, enabling high-precision signal classification with only a single forward pass. The framework leverages deep multi-domain feature fusion to construct a tailored dataset that supports model inference, and introduces a Domain Attention (DA) module to dynamically allocate feature weights, thereby focusing on critical domain-specific information. Extensive experiments demonstrate that in scenarios with limited labeled samples (1–200 samples per class(SPC)), the proposed MD-TabPFN achieves state-of-the-art modulation recognition accuracy while maintaining highly efficient real-time inference speed, significantly outperforming comparative methods. Liaoyang Li, Zisen Qi, Guangwei Yang |
Neurocomputing | 3 |
| 2026 | Surpassing the Master: A Joint Channel and Power Jamming Decision-Making Approach Under Imperfect DemonstrationsabstractThe rapid advancement in anti-jamming technologies presents significant challenges to the communication of jamming UAVs. Although deep reinforcement learning (DRL) offers a promising solution, existing methods still rely to some extent on prior knowledge, such as the communication channel states, and they lack robustness against dynamic communication policies. To respond, this paper develops an intelligent jamming decision-making scheme against unknown communication policies. Firstly, we formulate the communication countermeasure as a partially observable random game (POSG), where the jammer operates without the communication channel state information. Then, we design a proximal policy optimization with policy feedback (PPO-PF) algorithm to derive the best response jamming strategy (BRJS). In contrast to conventional algorithms based on actor-critic (AC), PPO-PF directly uses policy value to update the value function, enabling joint optimization of policy and value function. An ensemble evaluation network is introduced to enhance value estimation accuracy in the case of limited training data. To enhance the robustness of the algorithm for dynamic communication policies, we propose a teacher-student framework (TSF) that leverages imperfect demonstrations interventions to enhance the learning efficiency. The results demonstrate our approach’s superior adaptability to diverse communication policies, achieving significant performance improvements over baselines. Notably, the TSF exhibits two key advantages: (1) effective utilization of imperfect teacher demonstrations to accelerate learning, and (2) guaranteed student policy performance independent of teacher policy quality. In the testing faced with communication strategy μDQN, our approach achieves performance improvements of 48.83%, 14.79%, and 17.85% in relation to various levels of PPO teacher policies. Zisen Qi, Dan Wang 0018, Ning Rao |
IEEE Internet Things J. | 3 |
| 2025 | DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual PromptsabstractRecent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object details or complex scenes. To address this limitation, we propose DualCap, a novel approach that enriches the visual representation by generating a visual prompt from retrieved similar images. Our model employs a dual retrieval mechanism, using standard image-to-text retrieval for text prompts and a novel image-to-image retrieval to source visually analogous scenes. Specifically, salient keywords and phrases are derived from the captions of visually similar scenes to capture key objects and similar details. These textual features are then encoded and integrated with the original image features through a lightweight, trainable feature fusion network. Extensive experiments demonstrate that our method achieves competitive performance while requiring fewer trainable parameters compared to previous visual-prompting captioning approaches. The source code is available at https://github.com/mungeryang/DualCap. Binbin Li 0003, Guimiao Yang, Zisen Qi, Haiping Wang 0003 |
MMAsia | 3 |
| 2025 | Energy-efficient strategy generation for smart jammer in non-zero-sum games: A deep reinforcement learning approach
Zisen Qi, Dan Wang 0018, Yue Zhang 0022, Yiqiong Pang |
Comput. Networks | 3 |
| 2025 | The pupil outdoes the master: Imperfect demonstration-assisted trust region jamming policy optimization against frequency-hopping spread spectrum
Ning Rao, Zisen Qi, Dan Wang 0018, Yue Zhang 0022 |
Comput. Commun. | 3 |
| 2025 | Optimization for Paralyzing G2A Communication Network: A DRL-Based Joint Path Planning and Jamming Power Allocation ApproachabstractThis letter investigates the jammer path planning and jamming power allocation problem during airborne deterrence operation (ADO) in highly dynamic environments. In response to airborne threats posed by enemy aircraft formations, jammers must rely on perceptual information to plan trajectories and emit jamming signals to paralyze the ground-to-air (G2A) communication networks. Unlike traditional static scenarios, the high mobility of both sides presents significant challenges. Most works only study jamming solutions for static ground or single airborne targets, failing to address multiple airborne targets. We propose a joint path planning and jamming power allocation approach based on deep reinforcement learning (JPPJPA-DRL). This approach considers the impact of flight paths on receiving antenna gain, models the ADO as a Markov Decision Process (MDP), and uses the proximal policy optimization (PPO) algorithm to generate optimized path points and jamming power allocation schemes. In addition, a scientific reward function is designed to guide the learning process, and a visual communication countermeasure simulation platform is developed. The results show that the proposed approach can efficiently paralyze G2A communication networks, outperforming the baseline. Zisen Qi, Dan Wang 0018, Yiqiong Pang |
IEEE Signal Process. Lett. | 3 |
| 2024 | Triple GNNs: Introducing Syntactic and Semantic Information for Conversational Aspect-Based Quadruple Sentiment AnalysisabstractConversational Aspect-Based Sentiment Analysis (DiaASQ) aims to detect quadruples {target, aspect, opinion, sentiment polarity} from given dialogues. In DiaASQ, elements constituting these quadruples are not necessarily confined to individual sentences but may span across multiple utterances within a dialogue. This necessitates a dual focus on both the syntactic information of individual utterances and the semantic interaction among them. However, previous studies have primarily focused on coarse-grained relationships between utterances, thus overlooking the potential benefits of detailed intra-utterance syntactic information and the granularity of inter-utterance relationships. This paper introduces the Triple GNNs network to enhance DiaAsQ. It employs a Graph Convolutional Network (GCN) for modeling syntactic dependencies within utterances and a Dual Graph Attention Network (DualGATs) to construct interactions between utterances. Experiments on two standard datasets reveal that our model significantly outperforms stateof-the-art baselines. The code is available at https://github.com/ nlperi2b/Triple-GNNs-. Binbin Li 0003, Siyu Jia, Bingnan Ma, Zisen Qi, Xingbang Tan, Menghan Guo, Shenghui Liu |
CSCWD | 6 |
| 2024 | Dynamic Multi-Scale Context Aggregation for Conversational Aspect-Based Sentiment Quadruple AnalysisabstractConversational aspect-based sentiment quadruple analysis, namely DiaASQ, aims to extract the quadruple of target-aspect-opinion-sentiment within a dialogue . In DiaASQ, a quadruple’s elements often cross multiple utterances. This situation complicates the extraction process, emphasizing the need for an adequate understanding of conversational context and interactions. However, existing work independently encodes each utterance, thereby struggling to capture long-range conversational context and overlooking the deep inter-utterance dependencies. In this work, we propose a novel Dynamic Multi-scale Context Aggregation network (DMCA) to address the challenges. Specifically, we first utilize dialogue structure to generate multi-scale utterance windows for capturing rich contextual information. After that, we design a Dynamic Hierarchical Aggregation module(DHA) to integrate progressive cues between them. In addition, we form a multi-stage loss strategy to improve model performance and generalization ability. Extensive experimental results show that the DMCA model outperforms baselines significantly and achieves state-of-the-art performance1. Wenyuan Zhang 0002, Binbin Li 0003, Siyu Jia, Zisen Qi, Xingbang Tan |
ICASSP | 5 |
| 2023 | UCWSC: A unified cross-modal weakly supervised classification ensemble frameworkabstractIn recent years, Internet data has grown exponentially, but due to the lack of labels, the data that can be used is still relatively small. To solve this problem, research on weak supervision has emerged. However, common weakly supervised research often focuses on either single-modal data or multi-modal data research, which cannot be compatible with both types of data at the same time. Motivated by this observation, we propose a unified cross-modal weakly supervised classification ensemble framework (UCWSC) to tackle this issue. Especially, our proposed framework is based on high-order feature information of different modes. First, We introduce a feature fusion method based on high-order features to increase the amount of acquired information. Then we propose a modified Feature MixMatch algorithm with learning from feature representations. We propose feature fusion and decision fusion methods for weakly supervised classification of multi-modal data with voting and weighting mechanisms as discriminators to obtain the final classification results, respectively. We demonstrate the compatibility of these techniques, our classification accuracy can reach around 99% on the Wikipedia dataset and 78% on the MVSA-Multiple dataset. Huiyang Chang, Binbin Li 0003, Guangjun Wu, Haiping Wang 0003, Siyu Jia, Zisen Qi, Xiaohua Jiang |
CSCWD | 7 |
| 2022 | An Incremental Malware Classification Approach Based on Few-Shot LearningabstractMalware classification plays a fundamental role among all the related tasks. Researchers and anti-virus vendors have proposed deep learning (DL) methods to deal with the fast emerging-malware families and samples. However, for ordinary deep learning methods, once the model is trained, the set of families that can be recognized is fixed. This is troublesome in practice to deal with emerging malware families or unknown families with scarce samples. To resolve this issue, we propose an incremental classification approach called IMC (Incremental Malware Classification) based on few-shot learning, and the classifier is implemented as a cosine similarity function between extracted features and feature vectors of target classes. IMC can efficiently extend the pre-trained model to unknown families dynamically with a handful of samples without losing the ability to recognize the families it has “seen”. In the process of adaptation, no re-training is needed and fast inference is realized by a single forward pass. We extensively evaluate our approach on a dataset named APIMDS where the framework achieves incremental ability to classify the unknown families with high accuracy while maintaining the ability to recognize the known families. To our best knowledge, this is the first approach to meet the requirements to unify the classification of both unknown and known malware families in a few-shot manner. Qian Qiang, Mian Cheng, Yuan Zhou 0008, Zisen Qi, Fei Jiao |
ICC | 7 |
| 2022 | Towards Better Personalization: A Meta-Learning Approach for Federated Recommender Systems
Zhengyang Ai, Guangjun Wu, Zisen Qi, Yong Wang 0032 |
KSEM (2) | 4 |
| 2022 | Cost-Effective Malware Classification Based on Deep Active Learning
Qian Qiang, Tianning Zang, Mian Cheng, Quanbo Pan, Zisen Qi |
SecureComm | 8 |
| 2021 | MALUP: A Malware Classification Framework using Convolutional Neural Network with Deep Unsupervised Pre-trainingabstractMalware is becoming the main threat during the development of the Internet. Driven by the increasing cost due to the complexity and diversity of malware, machine learning is more and more popular for related tasks. Among these ML methods, convolutional neural network has gradually become the mainstream for its excellent performance. At the same time, the scale of labeled data needed for training a ConvNet is so large that labeling malware samples has become a huge burden which hinders the practical application. Meanwhile, better performance means deeper network which requires more resources and more labeled samples for training. In this paper, we propose a novel and effective framework called MALUP for malware classification of guaranteed performance improvement which uses ConvNet with deep unsupervised pre-training. The framework is implemented with deepCluster as the pre-training method to deal with unlabeled malware images and the pre-trained ConvNet is then fine-tuned with labeled samples. We evaluate the effectiveness of MALUP by measuring the performance on a benchmark provided by Microsoft in different conditions. The experiments demonstrate the superior performance of MALUP compared to the ConvNet of the same architecture, especially for a shallower network. Additionally, we evaluate the impact of the number of clusters, the volume of unlabeled data, a simple balance restriction optimization during pre-training, and the percentage of labeled samples as well. Those variables form the optimization space of MALUP and help the proposed framework outperform the transfer learning model pre-trained with ImageNet in deepCluster. Qian Qiang, Mian Cheng, Zisen Qi |
TrustCom | 5 |