VLDB 2026 Research / reviewers in the wild / expert
Bingzhen Wu
dblp:272/8156
· DBLP profile ↗
17ranked-venue papers
0as first author
16since 2021 · last 2026
0009-0006-7267-839XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Orion: Steering Personalized Web Agents via Global-Micro Profiling and Adaptive Intent TrackingabstractRecently, Large Language Models (LLMs) based Web Agents have shown significant potential in web understanding and interaction tasks. However, their personalization ability and user experience remain limited by the ambiguity and dynamic nature of user intent, struggling to model diverse user interests and track intent changes over time. To address these challenges, this paper proposes Orion, a novel personalized Web Agent. Orion adopts a global-micro profiling mechanism to balance users' long-term stable preferences and scenario-based needs, and introduces context-aware interest retrieval to enhance personalization. Additionally, we design adaptive profile tracking and proactive disambiguation mechanisms to effectively address the continuous evolution of user intent in multi-turn interactions. Orion is optimized through end-to-end online reinforcement learning, improving personalized reasoning and decision-making ability in real interactive scenarios. Experiments demonstrate that Orion significantly outperforms state-of-the-art baselines in personalized understanding and task efficiency. Die Hu 0004, Jingguo Ge, Weitao Tang, He Kong 0003, Liangxiong Li, Bingzhen Wu |
AAAI | 6 |
| 2026 | Learning general-purpose and robust representations of microservice system states from multi-modal data
Jingguo Ge, Yulei Wu, Hui Li 0098, Bingzhen Wu, Tong Li 0012 |
Inf. Process. Manag. | 6 |
| 2026 | ReID: Re-ranking through image description for object re-identification
Xiukang Yang, Jingguo Ge, Hui Li 0098, Liangxiong Li, Bingzhen Wu |
Pattern Recognit. | 5 |
| 2025 | QShield: Universal Defense Framework Against QUIC Client-Side Attacks with eBPFabstractQUIC adapts well to complex network situations due to mechanisms such as 0/1-RTT handshake and fast retransmission. It has now become a new star in the era of IoT. However, a phenomenon reveals the security risks of QUIC. Studies indicate that the internet is exposed to an average of four QUIC flood attacks per hour. The efficiency-oriented features of QUIC introduce vulnerabilities, making it particularly susceptible to attacks. In IoT scenarios, protocol security often depends on the design of the protocol itself and the middlewares. However, QUIC's design does not prioritize security as its highest concern and due to the ossification of middleboxes, the server-side defense is the only option. Therefore, we propose the QShield framework to seek breakthroughs from an engineering perspective. Based on the SDN principles, QShield consists of three parts. At the application layer, QUIC applications can utilize this library to implement strategy development and enable data sharing. A user-space library operates as the control layer, facilitating real-time bidirectional data transfers between the kernel and user space via a suite of APIs. In the data layer, QShield core blocks attack packets at the lowest layer of the Linux network protocol stack using eBPF technology. QShield effectively resists client-side attacks, and experimental results show that QShield can reduce the server's CPU usage for processing attack packets by nearly 50%, thereby essentially restoring normal Queries Per Second and bandwidth, reducing the bandwidth amplification by about 68% on the specific QUIC implementation. Yulin Ni, Yuepeng E, Jingguo Ge, Bingzhen Wu |
CSCWD | 4 |
| 2025 | Deep Incremental Cross-Modal Hashing Network for Fast RetrievalabstractCross-modal hashing techniques provide an effective method for large-scale cross-modal search due to their ability to handle multiple data types and their efficient storage and computation performance. To achieve excellent performance, deep supervised cross-modal hashing methods require extensive training data from various classes. However, when new classes emerge in the database, existing cross-modal hashing methods typically need to retrain the image and text encoders and regenerate hash codes for all data, which is impractical for large-scale retrieval systems. In this paper, we introduce an innovative cross-modal incremental hashing framework called Deep Incremental Cross-Modal Hashing Network (DICMHN), which can learn hash codes incrementally. The DICMHN framework is capable of directly learning the hash codes of newly emerging images and texts while maintaining the integrity of existing hash codes. Moreover, this framework ensures the accuracy of query results by modeling the correlation between query images and matching texts, as well as between query texts and matching images, while preserving the distinctions between images and texts. Extensive experiments conducted on multiple cross-modal benchmark datasets demonstrate that our proposed DICMHN framework significantly reduces training time and outperforms existing state-of-the-art methods. Xiukang Yang, Jingguo Ge, Liangxiong Li, Bingzhen Wu |
CSCWD | 4 |
| 2025 | Multi-Dimensional Series Forecasting for Multi-Node Microservices: Leveraging Specialized Embedding in LLMsabstractCurrently, time series prediction in microservice systems suffers from inaccurate forecasts due to the complex interdependencies and highly dynamic workload characteristics. Traditional small-scale models struggle to understand the temporal patterns embedded within multi-dimensional metrics, leading to suboptimal performance. The advent of large language models (LLMs) offers a promising solution, as their powerful representation learning capabilities can effectively capture these complex temporal patterns. In this study, we propose a novel approach tailored for multi-dimensional time series forecasting in microservice environments. Our method leverages specialized embedding techniques that combine dynamic receptive field convolution and adaptive attention masks to capture temporal dependencies and feature relationships across multiple nodes. Additionally, we fine-tune a pre-trained LLaMA model to enhance its applicability for time series forecasting within microservice contexts. Experimental results demonstrate that our approach achieves higher prediction accuracy compared to baseline methods in different datasets. This research's achievements in time series forecasting provide new insights for downstream tasks such as resource allocation and fault prediction. Lefan Cheng, Jingguo Ge, Quanfeng Lv, Tong Li 0012, Bingzhen Wu |
HPCC | 5 |
| 2025 | WebSurfer: Enhancing LLM Agents with Web-Wise Feedback for Web NavigationabstractAs the Internet’s complexity and information volume surge, the need for efficient web automation becomes critical. Traditional web agents struggle with redundant web content, which disrupts their understanding of the environment. They also face inefficiencies in multi-task scenarios due to handcrafted exemplars and encounter error accumulation in long-horizon tasks, exacerbated by web-specific complexities like nested structures and interactive elements. To address these issues, we introduce WebSurfer, a novel web agent designed to filter, learn, and adapt in complex environments. WebSurfer refines task-oriented states for clearer observations and employs an exemplar retrieval and ordering strategy to enhance LLMs’ understanding and adaptability to current tasks. Notably,WebSurfer features a novel web-wise insight feedback mechanism that enables continuous adaptation and strategy refinement. Evaluations demonstrate that WebSurfer outperforms state-of-the-art (SOTA) methods on realistic tasks, achieving higher accuracy and enhancing longterm adaptability. Die Hu 0004, Jingguo Ge, Weitao Tang, Guoyi Li, Liangxiong Li, Bingzhen Wu |
ICASSP | 6 |
| 2025 | Reinforcement Learning-Based Multi-Teacher Knowledge Distillation for Enhancing Retrieval Ranking ConsistencyabstractKnowledge Distillation, an effective model compression technique, transfers knowledge from a large teacher model to a smaller student model, reducing computational costs while maintaining model performance. In large-scale retrieval tasks, maintaining the consistency of retrieval result rankings is crucial. However, traditional distillation methods focus on aligning the feature vectors extracted by the student model with those of the teacher model, which often fails to preserve ranking consistency in complex retrieval tasks. To address this issue, we propose a reinforcement learning-based multi-teacher knowledge distillation framework to optimize ranking consistency. By incorporating reinforcement learning strategies, the framework dynamically selects and adjusts the weights of multiple teacher models, enabling the student model to better learn from different teachers and accurately maintain retrieval rankings. Experimental results demonstrate that the proposed method significantly improves ranking consistency and retrieval performance on several benchmark datasets. Xiukang Yang, Jingguo Ge, Liangxiong Li, Bingzhen Wu |
ICASSP | 4 |
| 2025 | Integrating S1 &S2 Framework for Enhanced Semantic Match in Person Re-identification
Xiukang Yang, Jingguo Ge, Hui Li 0098, Liangxiong Li, Bingzhen Wu |
MMM (2) | 5 |
| 2024 | A time-sensitive cloud-native network based on eBPFabstractThe evolution of cloud computing and microservices is gradually supplanting traditional network deployment schemes within data centers. As applications deploy substantial computing resources in data centers, a fierce competition for network services ensues, marked by stringent quality requirements. Simultaneously, safeguarding the time sensitivity of the main flow becomes imperative. However, prevailing container network solutions primarily ensure service quality through scheduling and orchestration, neglecting the influence of computing resources on network service quality under intense resource competition. Consequently, our focus revolves around exploring the preservation of time-sensitive attributes of primary service network links, aiming to enhance the service quality of container networks in highly competitive computing resource environments. This paper introduces a novel container network solution designed to meet the quality of service requirements for time-sensitive data in container networks. Implemented on the Kubernetes platform, this solution establishes an underlay network structure based on Cilium for transmitting network packets requiring performance guarantees and exhibiting time sensitivity. Utilizing eBPF programs with adjusted CPU affinity for packet forwarding, the solution records packets necessitating quality of service guarantees. Network service quality is ensured through algorithms such as Multiqueue Priority, Earliest TxTime First, Enhancements for Scheduled Traffic, etc. The network packets requiring performance guarantees and time sensitivity refer to the TSN (Time-Sensitive Networking) standard. To assess the solution’s effectiveness, we deployed Kubernetes on two directly connected physical servers. Measurements were conducted in scenarios of both idle and highly competitive computing resources, evaluating bandwidth, latency, and jitter of container access packets across different hosts. The results confirm a noteworthy enhancement in container network service quality under highly competitive computing resource environments. Jifei Wen, Jingguo Ge, Hui Li 0098, Yuepeng E, Bingzhen Wu |
CSCWD | 6 |
| 2024 | A Latency-Predictable Cloud-Native Network Architecture based on XDPabstractCloud computing and microservices are increasingly supplanting traditional network deployment strategies within data centers due to their inherent flexibility in infrastructure management and rapid scalability. Nevertheless, current approaches often fall short in addressing quality of service (QoS) for virtual networks, especially under conditions of intense resource competition. This shortfall prevents the fulfillment of critical services’ requirements for low latency and predictability. To mitigate this challenge, we propose a cloud-native network service architecture designed to deliver consistently low latency and predictable packet arrival times. This architecture dynamically coordinates and reserves computational resources even during high contention periods, thereby maintaining container network QoS and ensuring the stable availability of microservice applications under extreme conditions. We validated the effectiveness of our proposed solution by deploying multiple container nodes across two directly connected servers and establishing a Kubernetes cluster. By launching a large volume of tasks within a constrained time frame, we assessed the container network’s load, latency, and jitter during peak usage periods. Our results indicate that, although our solution exhibits marginally reduced performance compared to the open-source Kubernetes network plugin under low-load conditions, it significantly outperforms the plugin under high-load scenarios. Specifically, when host CPU usage surpasses 90% and memory usage exceeds 80%, the open-source plugin experiences notable packet loss and long-tail latency distributions. In contrast, our solution demonstrates an 84% reduction in jitter compared to the open-source CNI. Jifei Wen, Jingguo Ge, Yuepeng E, Bingzhen Wu |
ISPA | 5 |
| 2024 | SIKGC: Structural Information Prompt Based Knowledge Graph Completion with Large Language ModelsabstractKnowledge Graph Completion (KGC) aims to enrich and complete the knowledge graph by discovering missing information from existing fact triples. However, existing KGC methods often overlook the utilization of structured knowledge within the knowledge base. In this paper, we propose a novel Large Language Models-based Knowledge Graph Completion framework, called SIKGC, which builds the structural information prompt to assist the knowledge graph completion tasks. Specifically, we arrange the triples in the knowledge graph as the sequences of text. By fusing the descriptions of entities, relations and their structural information as task-aware prompts, we input such prompts into large language models and regard the responses as prediction tasks. The experimental results on various public datasets show that the proposed method outperforms all baseline methods for the three knowledge completion tasks and attains state-of-the-art in triple classification. We also demonstrate that fine-tuning the smaller large language models (e.g., Baichuan2-13B, LLaMA2-13B, ChatGLM3-6B) with relevant data markedly enhances their KGC capabilities and significantly outperforms GPT-4. Jingguo Ge, Weihua Feng, Liangxiong Li, Bingzhen Wu |
SMC | 6 |
| 2023 | CDANER: Contrastive Learning with Cross-domain Attention for Few-shot Named Entity RecognitionabstractFew-shot Named Entity Recognition (NER) aims to recognize unseen name entities based on a tiny support set that consists of seen name entities and labels, which is obviously different from traditional supervised NER methods. Contrastive learning has become a popular solution for few-shot NER, which improves the robustness of NER to handle unlabeled entities by learning a similarity metric to measure the semantic similarity between test samples and entity labels. However, existing contrastive learning based NER methods individually learn the word embedding in source and target domains, ignoring connections between entities with the same label and limiting the effectiveness of contrast learning. In this paper, we propose a novel few-shot NER framework that jointly models different domain texts and optimizes a generalized objective of differentiating between words in all stages. The proposed model builds the cross-domain attention layer to enhance the feature representations of words and transfer the entity similarity information from the source domain to the target domain. This significantly reduces the divergence between entities with same label. Experimental results on the largest Few-shot NER dataset show that CDANER significantly outperforms all baseline methods, which verifies the effectiveness and robustness of the proposed model. Hui Li 0098, Jingguo Ge, Lei Zhang 0116, Liangxiong Li, Bingzhen Wu |
IJCNN | 6 |
| 2023 | A New Federated Learning Model for Host Intrusion Detection System Under Non-IID DataabstractHost Intrusion Detection System (HIDS) is an important research topic in the field of cyberspace security. With the explosion in the number of malicious attacks in recent years, machine learning-based detection method is now the most common and efficient approach. While traditional centralized machine learning needs to transmit data to the central server for training, which not only requires the central server to have large computing resources, but also causes problems such as sensitive data leakage and communication overhead. As a distributed machine learning paradigm, Federated Learning (FL) can achieve multi-party collaborative training and aggregate a unified global model without data sharing, which can well alleviate these problems. It is worth noting that existing studies on the use of FL in HIDS are all conducted in the scenario where the data is independent and identically distributed (IID). However, due to the different context of hosts, the data generated by hosts is usually non-independent and identically distributed (Non-IID) in reality. Therefore, We investigate the impact of Non-IID data with different skew levels on FL in HIDS. On this basis, we propose a data augmentation FL algorithm based on Synthetic Minority Over-Sampling Technique (SMOTE) to reduce the impact of Non-IID data. We also develop a data collection module using extended Berkeley Packet Filter (eBPF) technology to collect a dataset for experiments. Experimental results show that our proposed FL algorithm can effectively improve the performance of HIDS under Non-IID data. Yongfei Liu, Lanxue Zhang, Liangxiong Li, Tong Li 0012, Bingzhen Wu |
SMC | 7 |
| 2022 | EDP: An eBPF-based Dynamic Perimeter for SDP in Data CenterabstractIn recent years, the concept of Zero Trust Networks (ZTN) has been proposed to overcome unrealistic security assumptions, e.g., what lies in private networks (such as data centers) is always trusted and safe. In ZTN, no device or user is assumed to be secure, instead all connections have to be authenticated and authorized before being established. Software Defined Perimeter (SDP) is one of the most promising solution for ZTN, where the gateway allows clients to access services only after receiving legitimate Single Packet Authorization (SPA) data. However, existing SDP solutions either (1) need to decouple the SPA from the connection request, resulting in redundant communication processes and impersonation attacks; or (2) need to copy the SPA data to the user space from sniffers, causing the packets to enter the protocol stack repeatedly. Due to the large number of short-lived streams in the data center, inefficiency and insecurity of the SPA process lead to severe connection delays and network attacks (e.g., DDoS). To this end, we propose an eBPF-based Dynamic Perimeter (EDP) to enhance the security and performance of SDP. By using EDP, authentication data can be efficiently embedded into every packet and checked before entering the receiver's protocol stack. Experimental results show that the connection delay of EDP is 80% less than that of the existing state-of-the-art solutions. Lei Zhang 0116, Hui Li 0098, Jingguo Ge, Yulei Wu, Liangxiong Li, Bingzhen Wu, Haojiang Deng |
APNOMS | 6 |
| 2022 | Social Relationship Recognition Based on Relational Self-Attention MechanismabstractSocial relations are closely related to each of us and are a crucial part of society. Recognizing the social relationships of people in pictures can improve AI’s understanding of human behavior, thereby facilitating collaborative interactions between computers and humans. Previous work only focused on a single picture, so too little information can be obtained. In this paper, we proposed Picture Reasoning Model(PRM) to achieve relationship classification, which innovatively uses the self-attention method to learn the association between relationships. The association between relationships is at the social level, thus using it to assist relationship recognition can get rid of the problem of insufficient information in a single picture. In addition, the model also adopts a two-stream approach, extracting both characters and global features for getting multiple perspectives information. We conduct extensive experiments on two benchmark datasets PIPA and PISC. Experimental results show that our model has improved the accuracy metric of the datasets compared with SOTA. On the PIPA dataset, the accuracy increases from 64.4% to 65.6%, and on the PISC dataset, the mAP raises from 72.7% to 73.2%, which validates the effectiveness of our proposals. Deming Lin, Laifu Wang, Guoshui Shi, Hui Li 0098, Bingzhen Wu, Jingguo Ge |
CSCWD | 6 |
| 2020 | Mining DApp Repositories: Towards In-Depth Comprehension and Accurate Classification
Yeming Lin, Tong Li 0012, Jingguo Ge, Bingzhen Wu |
SEKE | 5 |