VLDB 2026 Research / reviewers in the wild / expert
Bingjie Zhu
dblp:206/7901 · also Bing-Jie Zhu
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Edge Inference for Large Language Models with Speculative Decoding
Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
ICC | 1 |
| 2026 | Physical-Constraint-Embedded Deep Reinforcement Learning Approach for Flight Control of Flying-Wing VehicleabstractDue to the uncertain nonlinear dynamics and the coupling motions of the longitudinal and lateral axes, the attitude control of flying-wing vehicles presents significant challenges. Proportional-Integral-Derivative (PID) control with gain scheduling relies on small disturbance linearization theory, which limits control performance under substantial nonlinear behavior. Nonlinear control techniques, such as Model Predictive Control (MPC) and Nonlinear Dynamic Inversion (NDI), are heavily dependent on accurate models, and their robustness diminishes when aerodynamic models are mismatched. To address these limitations, this paper proposes an attitude control algorithm based on Deep Reinforcement Learning (DRL), which considers the coupling of lateral and longitudinal motions as well as airspeed. The algorithm establishes a clear mapping between ’State-Action-Reward’ by utilizing an error-related state space. The proposed algorithm significantly enhances training efficiency and test success rates compared to DRL algorithms that incorporate attitude angle. Furthermore, it introduces physical constraints on the control surface deflection and an additional loss function term, resulting in smooth actuation signals and attitude responses. In simulated environments, the algorithm was compared with other state-of-the-art algorithms, and the results demonstrate its strong robustness in the presence of severe turbulence disturbances and model uncertainties. The algorithm is also assessed through real vehicle experiments, which fully corroborate its effectiveness. Shijing Zhang, Bingjie Zhu, Yafei Lu, Peng Wang 0176 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Efficient LLM Inference Over Heterogeneous Edge Networks With Speculative Decoding
Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Commun. | 1 |
| 2026 | Enabling Efficient Large Language Model Inference Over Wireless Networks With CachingabstractWith the proliferation of large language models (LLMs), cloud-based LLM serving mechanisms may cause network congestion and high serving delay. Edge computing offers a solution to alleviate backhaul pressure and reduce serving delay by deploying LLMs on edge servers and providing LLM inference services in users’ proximity. However, user accuracy requirements vary over time, and mismatches between these requirements and the deployed LLMs at the edge may lead to inefficient resource usage and increased serving delay. To address this, we formulate a joint LLM caching, inference task scheduling, and network resource allocation problem to minimize LLM serving delay under unknown time-varying user accuracy requirements. To solve the problem, we first derive closed-form solutions for optimal computation and communication resource allocation under any LLM caching and task scheduling policies. Then, we employ an improved branch-and-bound algorithm to obtain optimal task scheduling policies under any LLM caching strategies. Finally, we propose an improved double deep Q-network (DDQN)-based algorithm to determine the LLM caching decisions. It incorporates a state coding and action aggregation (SCAA) mechanism within the deep neural networks (DNNs) of the traditional DDQN. The SCAA-DNNs involve an input-layer gating mechanism to encode users’ request states for LLMs and a two-layer output architecture that dynamically aggregates LLM caching actions to generate the corresponding state-action values, thereby improving learning efficiency and accelerating convergence in large discrete action spaces. Experimental results show that the proposed scheme could rapidly converge and reduce average user delay by up to 20.8% compared to benchmarks. Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Joint Caching and Inference for Large Language Models in Wireless NetworksabstractTo reduce the serving delay of large language model (LLM)-based applications, the edge-based LLM serving mechanism offers a promising solution by caching LLMs at the edge to provide LLM inference services closer to users. Motivated by this, we propose an edge-based LLM caching and inference framework to support low-delay LLM-based services. Based on the framework, we formulate a joint LLM caching, inference task scheduling, and computation resource allocation optimization problem to minimize LLM serving delay, where time-varying LLM popularity is considered. Given an LLM caching policy, we first obtain the optimal solution for the computation resource allocation and task scheduling by using traditional optimization methods. Then, we propose an improved double deep Q-network (IDDQN) algorithm that effectively learns the optimal LLM caching strategy under unknown LLM popularity. The IDDQN algorithm integrates a state coding and action aggregation (SCAA) mechanism in the deep neural network structure, enabling it to efficiently capture users' preferences for LLMs and mitigate the slow convergence issues due to the large action space. Simulation results indicate that the proposed scheme achieves both lower average user delay and faster convergence than other benchmarks. Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
ICC | 1 |
| 2025 | Advancing ADMET prediction through multiscale fragment-aware pretraining with MSformer-ADMETabstractAbsorption, distribution, metabolism, excretion, and toxicity (ADMET) properties are critical determinants of the pharmacokinetic and safety profiles of drug candidates. Accurate and early-stage prediction of ADMET characteristics is essential for reducing late-stage attrition rates, lowering development costs, and accelerating the drug discovery process. Recent advances in deep learning have shown great promise in molecular property prediction, especially with the emergence of Transformer-based architectures that can effectively model long-range dependencies in molecular representations. However, most existing methods rely heavily on atom-level encodings (e.g. smiles or molecular graphs), which often lack structural interpretability and generalization across heterogeneous tasks. Previously, we developed a de novo and flexible molecular representation framework named MSformer (available at https://github.com/ZJUFanLab/MSformer), which demonstrated success in bioactivity prediction. We have now adapted and specialized this architecture for ADMET property prediction. This adapted implementation, designated as MSformer-ADMET, extends the framework's capabilities to pharmacokinetic and toxicity endpoints while maintaining its flexible, fragmentation-based approach to molecular representation learning. MSformer-ADMET is fine-tuned on 22 tasks collected from the Therapeutics Data Commons (TDC), covering both classification and regression settings. Results demonstrate that MSformer-ADMET achieves superior performance across a wide range of ADMET endpoints, consistently outperforming conventional smiles-based and graph-based models. Notably, we further conducted interpretability analyses by leveraging the model's attention distributions and fragment-to-atom mappings, allowing the identification of key structural fragments that are highly associated with molecular properties. This post hoc interpretability provides more transparent insights into the structure-property relationship. Collectively, results demonstrate that MSformer-ADMET is a highly effective and broadly applicable model for ADMET prediction. Bingjie Zhu, Shuyang Nie, Yugang Lin, Minjie Shen, Yanrong Zheng |
Briefings Bioinform. | 2 |
| 2024 | Cost-Efficient Cooperative Video Caching Over Edge NetworksabstractCooperative caching has emerged as an efficient way to alleviate backhaul traffic and enhance user experience by proactively prefetching popular videos at the network edge. However, it is challenging to achieve the optimal design of video caching, sharing, and delivery within storage-limited edge networks due to the growing diversity of videos, unpredictable video requirements, and dynamic user preferences. To address this challenge, this work explores cost-efficient cooperative video caching via video compression techniques while considering unknown video popularity. Firstly, we formulate the joint video caching, sharing, and delivery problem to capture a balance between user delay and system operative cost under unknown time-varying video popularity. To solve this problem, we develop a two-layer decentralized reinforcement learning algorithm, which effectively reduces the action space and tackles the coupling among video caching, sharing, and delivery decisions compared to the conventional algorithms. Specifically, the outer layer produces the optimal decisions for video caching and communication resource allocation by employing a multi-agent deep deterministic policy gradient algorithm. Meanwhile, the optimal video sharing and computation resource allocation are determined in each agent’s inner layer using the alternating optimization algorithm. Numerical results show that the proposed algorithm outperforms benchmarks in terms of the cache hit rate, delay of users and system operative cost, and effectively strikes a trade-off between system operative cost and users’ delay. Bingjie Zhu, Wenqiang Yi, Zhixiong Chen 0003, Arumugam Nallanathan |
IEEE Internet Things J. | 1 |
| 2023 | Design and Implementation of Holistic Service-Based End-to-end Network Slicing for 6GabstractWith the diversified development of the vertical industry, it's urgent to enhance end-to-end (E2E) network slicing for 6G. In addition, by introducing artificial intelligence (AI), the performance of E2E network slicing can be improved with limited radio resources. Therefore, we propose a holistic service-based E2E network slicing. Firstly, the E2E network slicing is abstracted into four layers and three planes, i.e., infrastructure, virtualization, function, and application layer at the horizontal, as well as control, AI, and MANO plane at the vertical. Especially, with reference to service-based architecture (SBA) in 5G core network (5GC), all the three planes are decoupled into independent functions, which are connected through a uniform service-based interface (SBI). Secondly, we design the network slicing templates for some typical applications and instantiate the templates to provide customized services for users. Finally, the experimental results show that our proposed holistic service-based E2E network slicing can ensure the isolation among network slices effectively and reduce the service response time, as well as improve the reliability of the system. Chang Qin, Tao Sun 0010, Mengtian Liu, Bingjie Zhu, Haiyan Tu, Manhua Zhu |
VTC2023-Spring | 4 |