Yunzhe Li 0001

dblp:220/3441-1 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0003-4146-8542ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
Yunzhe Li 0001, Hongzi Zhu, James Lin 0001, Shan Chang, Minyi Guo
NDSS1
2026 On the Availability Risks of Production LLM Services Under Unbounded Inference
abstract
Large Language Models (LLMs) have become foundational components in a wide range of applications, including natural language understanding and generation, embodied intelligence, and scientific discovery. As their computational requirements continue to grow, these models are increasingly deployed as cloud-based services, allowing users to access powerful LLMs via the Internet. However, this deployment model introduces a new class of threat: denial-of-service (DoS) attacks via unbounded reasoning, where adversaries craft specially designed inputs that cause the model to enter excessively long or infinite generation loops. These attacks can exhaust backend compute resources, degrading or denying service to legitimate users. To mitigate such risks, many LLM providers adopt a closed-source, black box setting to obscure model internals. In this paper, we propose ThinkTrap, a novel input-space optimization framework for DoS attacks against LLM services even in black-box environments. The core idea of ThinkTrap is to first map discrete tokens into a continuous embedding space, then undertake efficient black-box optimization in a low-dimensional subspace exploiting input sparsity. The goal of this optimization is to identify adversarial prompts that induce extended or non-terminating generation across several state-of-the-art LLMs, achieving DoS with minimal token overhead. We evaluate ThinkTrap across multiple commercial, closed-source LLM services and observe that it can consistently induce abnormally long outputs and noticeable response-side degradation under black-box access. To further quantify the system-level impact of the attack, we conduct controlled experiments on private LLM deployments, where ThinkTrap reduces service throughput to as low as 1% of its original capacity and, in extreme cases, induces complete service failure due to resource exhaustion.
Yunzhe Li 0001, Hongzi Zhu, James Lin 0001, Shan Chang, Minyi Guo
IEEE Trans. Dependable Secur. Comput.1
2026 Online Edge-Assisted Video Analytics on Mobile Agents via Differential Video Encoding
abstract
Ensuring stable and high-quality real-time video analytics for computationally constrained mobile agents is essential. However, limited computing resources and network bandwidth present significant challenges in meeting the objective of low response time and high inference accuracy. In this paper, we present DiVE, an edge-assisted video analytics system that utilizes motion vectors calculated by video codec to extract foregrounds and differentially encode frames. DiVE removes rotational components from motion vectors by solving over-determined linear equations and filters noisy motion vectors based on the observation that motion vectors of static objects point to the same point when the ego agent purely translates. To distinguish foregrounds from backgrounds, DiVE estimates the ground based on observations that all foregrounds stand on the ground and motion vectors on static objects at the same height have the same normalized magnitude. DiVE then uses region-growing-based clustering to identify foreground objects. An adaptive bitrate allocation method is applied to optimize accuracy under estimated bandwidth. We implement a prototype and conduct extensive experiments to evaluate the performance of DiVE. The results demonstrate that DiVE can improve detection accuracy by up to 19.0% and reduce response time by up to 56.0% compared with other video analytics schemes in real-world traces.
Hongzi Zhu, Jiangang Shen, Liang Zhang 0027, Yunzhe Li 0001, Shan Chang, Jie Wu 0001, Minyi Guo
IEEE Trans. Mob. Comput.4
2025 Saga: Capturing Multi-granularity Semantics from Massive Unlabelled IMU Data
abstract
Inertial measurement units (IMUs), have been prevalently used in a wide range of mobile perception applications such as activity recognition and user authentication, where a large amount of labelled data are normally required to train a satisfactory model. However, it is difficult to label micro-activities in massive IMU data due to the hardness of understanding raw IMU data and the lack of ground truth. In this paper, we propose a novel fine-grained user perception approach, called Saga, which only needs a small amount of labelled IMU data to achieve stunning user perception accuracy. The core idea of Saga is to first pre-train a backbone feature extraction model, utilizing the rich semantic information of different levels embedded in the massive unlabelled IMU data. Meanwhile, for a specific downstream user perception application, Bayesian Optimization is employed to determine the optimal weights for pre-training tasks involving different semantic levels. We implement Saga on five typical mobile phones and evaluate Saga on three typical tasks on three IMU datasets. Results show that when only using about 100 training samples per class, Saga can achieve over 90% accuracy of the full-fledged model trained on over ten thousands training samples with no additional system overhead.
Yunzhe Li 0001, Facheng Hu, Hongzi Zhu, Shifan Zhang, Liang Zhang 0027, Shan Chang, Minyi Guo
ICDCS1
2025 DiVE: Differential Video Encoding for Online Edge-assisted Video Analytics on Mobile Agents
Jiangang Shen, Hongzi Zhu, Liang Zhang 0027, Yunzhe Li 0001, Shan Chang, Jie Wu 0001, Minyi Guo
ICDCS4
2025 CoPe: Taming Collaborative 3D Perception via Lite Network Attention across Mobile Agents
abstract
To extend the receptive field of a mobile agent in complex scenarios, it is essential for multiple agents to cooperate with each other. However, it is challenging to achieve comprehensive 3D perception at the minimal computational and communication costs. In this paper, we propose CoPe, a lightweight and efficient collaborative 3D perception scheme for mobile agents. The main idea of CoPe is for an ego agent to query the most helpful information from its neighboring agents through a lightweight network attention mechanism. To this end, at each agent, we first leverage Singular Value Decomposition (SVD) to decompose a full-size point cloud feature into components. Meanwhile, with the novel self-attention and cross-attention algorithms, we respectively select the key component of an ego agent that best represent the point cloud of the ego agent as a query, and valuable components of each helper agent that are most relevant to the query as the answer. After feature reconstruction and aggregation, an ego agent can have a comprehensive understanding about the scene and make accurate predictions on downstream tasks. CoPe is lightweight and can be easily implemented on mobile devices. Results of extensive experiments conducted on both real-world and simulation datasets demonstrate that CoPe can achieve superior 3D object detection accuracy while significantly reducing the incurred computational and communication costs.
Shifan Zhang, Hongzi Zhu, Yunzhe Li 0001, Liang Zhang 0027, Shan Chang, Minyi Guo
ICDCS3
2025 Prism: Mining Task-aware Domains in Non-i.i.d. IMU Data for Flexible User Perception
abstract
A wide range of user perception applications leverage inertial measurement unit (IMU) data for online prediction. However, restricted by the non-i.i.d. nature of IMU data collected from mobile devices, most systems work well only in a controlled setting (e.g., for a specific user in particular postures), limiting application scenarios. To achieve uncontrolled online prediction on mobile devices, referred to as the flexible user perception (FUP) problem, is attractive but hard. In this paper, we propose a novel scheme, called Prism, which can obtain high FUP accuracy on mobile devices. The core of Prism is to discover task-aware domains embedded in IMU dataset, and to train a domain-aware model on each identified domain. To this end, we design an expectation-maximization (EM) algorithm to estimate latent domains with respect to the specific downstream perception task. Finally, the best-fit model can be automatically selected for use by comparing the test sample and all identified domains in the feature space. We implement Prism on various mobile devices and conduct extensive experiments. Results demonstrate that Prism can achieve the best FUP performance with a low latency.
Yunzhe Li 0001, Facheng Hu, Hongzi Zhu, Quan Liu 0006, Xiaoke Zhao, Jiangang Shen, Shan Chang, Minyi Guo
INFOCOM1
2025 LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation
Yunzhe Li 0001, Junting Wang 0001, Hari Sundaram, Zhining Liu 0002
RecSys1
2025 Exploiting Ground Depth Estimation for Mobile Monocular 3D Object Detection
abstract
Detecting 3D objects from a monocular camera in mobile applications, such as on a vehicle, drone, or robot, is a crucial but challenging task. The monocular vision's near-far disparity and the camera's constantly changing position make it difficult to achieve high accuracy, especially for distant objects. In this paper, we propose a new Mono3D framework named MoGDE, which takes inspiration from the observation that an object's depth can be inferred from the ground's depth underneath it. MoGDE estimates the corresponding ground depth of an image and utilizes this information to guide Mono3D. We use a pose detection network to estimate the camera's orientation and construct a feature map that represents pixel-level ground depth based on the 3D-to-2D perspective geometry. To further improve Mono3D with the estimated ground depth, we design an RGB-D feature fusion network based on transformer architecture. The long-range self-attention mechanism is utilized to identify ground-contacting points and pin the corresponding ground depth to the image feature map. We evaluate MoGDE on the KITTI dataset, and the results show that it significantly improves the accuracy and robustness of Mono3D for both near and far objects. MoGDE outperforms state-of-the-art methods and ranks first among the pure image-based methods on the KITTI 3D benchmark.
Yunsong Zhou, Quan Liu 0006, Hongzi Zhu, Yunzhe Li 0001, Shan Chang, Minyi Guo
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 A Scene-Aware Model Adaptation Scheme for Cross-Scene Online Inference on Mobile Devices
abstract
Emerging Artificial Intelligence of Things (AIoT) applications desire online prediction using deep neural network (DNN) models on mobile devices. However, due to the movement of devices,unfamiliartest samples constantly appear, significantly affecting the prediction accuracy of a pre-trained DNN. In addition, unstable network connection calls for local model inference. In this paper, we propose a light-weight scheme, calledAnole, to cope with the local DNN model inference on mobile devices. The core idea of Anole is to first establish an army of compact DNN models, and then adaptively select the model fitting the current test sample best for online inference. The key is to automatically identifymodel-friendlyscenes for training scene-specific DNN models. To this end, we design a weakly-supervised scene representation learning algorithm by combining both human heuristics and feature similarity in separating scenes. Moreover, we further train a model classifier to predict the best-fit scene-specific DNN model for each test sample. We implement Anole on different types of mobile devices and conduct extensive trace-driven and real-world experiments based on unmanned aerial vehicles (UAVs). The results demonstrate that Anole outwits the method of using a versatile large DNN in terms of prediction accuracy (4.5% higher), response time (33.1% faster) and power consumption (45.1% lower).
Yunzhe Li 0001, Hongzi Zhu, Zhuohong Deng, Yunlong Cheng, Zimu Zheng, Liang Zhang 0027, Shan Chang, Minyi Guo
IEEE Trans. Mob. Comput.1
2024 CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
abstract
Weixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li, Qian Chen, Wen Wang, Tingyu Lin, Weishan Zhao, Li Zhu, Hari Sundaram, Shuiguang Deng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Weixiang Yan, Yunkun Wang, Yunzhe Li 0001, Qian Chen 0003, Wen Wang 0001, Tingyu Lin 0002, Weishan Zhao, Hari Sundaram, Shuiguang Deng
ACL (1)4
2024 Advancing Precise Outline-Conditioned Text Generation with Task Duality and Explicit Outline Control
abstract
Yunzhe Li, Qian Chen, Weixiang Yan, Wen Wang, Qinglin Zhang, Hari Sundaram. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yunzhe Li 0001, Qian Chen 0003, Weixiang Yan, Wen Wang 0001, Hari Sundaram
EACL (1)1
2024 Anole: Adapting Diverse Compressed Models for Cross-Scene Prediction on Mobile Devices
abstract
Emerging Artificial Intelligence of Things (AIoT) applications desire online prediction using deep neural network (DNN) models on mobile devices. However, due to the movement of devices, unfamiliar test samples constantly appear, significantly affecting the prediction accuracy of a pre-trained DNN. In addition, unstable network connection calls for local model inference. In this paper, we propose a light-weight scheme, called Anole, to cope with the local DNN model inference on mobile devices. The core idea of Anole is to first establish an army of compact DNN models, and then adaptively select the model fitting the current test sample best for online inference. The key is to automatically identify model-friendly scenes for training scene-specific DNN models. To this end, we design a weakly-supervised scene representation learning algorithm by combining both human heuristics and feature similarity in separating scenes. Moreover, we further train a model classifier to predict the best-fit scene-specific DNN model for each test sample. We implement Anole on different types of mobile devices and conduct extensive trace-driven and real-world experiments based on unmanned aerial vehicles (UAV s). The results demonstrate that Anole outwits the method of using a versatile large DNN in terms of prediction accuracy (4.5 % higher), response time (33.1 % faster) and power consumption (45.1 % lower).
Yunzhe Li 0001, Hongzi Zhu, Zhuohong Deng, Yunlong Cheng, Liang Zhang 0027, Shan Chang, Minyi Guo
ICDCS1
2024 The Blind and the Elephant: A Preference-aware Edge Video Analytics Scheduler for Maximizing System Benefit
abstract
Video analytics is the killer workload in edge computing, which involves the scheduler’s complex decisions to balance analysis performance (latency and accuracy) and resource consumption (network, computation, and energy). Traditional schedulers address this as a single-objective optimization problem with fixed weights, unable to precisely capture unknown system preferences due to intricate pricing rules across various service levels and resource costs, consequently leading to suboptimal system benefit like monetary gain. In this paper, we propose a Bayesian optimization-driven multi-objective scheduler, PaMO, that can proactively explore the system pricing preference by pairwise comparing outcome vectors of all objectives. Moreover, PaMO designs a heuristic scheduling algorithm with a zero-delay jitter guarantee to avoid performance degradation caused by resource contention and uses a revised Bayesian optimization algorithm to make video configuration and scheduling decisions. Experiments on real video analytics workloads show that PaMO can achieve up to 53.9% benefit gain compared to state-of-the-art scheduling methods.
Liang Zhang 0027, Hongzi Zhu, Yunzhe Li 0001, Jiangang Shen, Minyi Guo
ICPP3
2024 Novas: Tackling Online Dynamic Video Analytics With Service Adaptation at Mobile Edge Servers
abstract
Video analytics at mobile edge servers offers significant benefits like reduced response time and enhanced privacy. However, guaranteeing various quality-of-service (QoS) requirements of dynamic video analysis requests on heterogeneous edge devices remains challenging. In this paper, we propose a scalable online video analytics scheme, called Novas, which automatically makes precise service configuration adjustments upon constant video content changes. Specifically, Novas leverages the filtered confidence sum and a two-window t-test to online detect accuracy fluctuations without ground truth information. In such cases, Novas efficiently estimates the performance of all potential service configurations through a singular value decomposition (SVD)-based collaborative filtering method. Finally, given the NP-hardness of the optimal scheduling problem, a heuristic scheduling strategy that maximizes the minimum remaining resources is devised to schedule the most suitable configurations to servers for execution. We evaluate the effectiveness of Novas through extensive hybrid experiments conducted on a dedicated testbed. Results show that Novas can achieve a substantial over 27$\times$improvement in satisfying the accuracy requirements compared with existing methods adopting fixed configurations, while ensuring latency requirements. Moreover, Novas improves the goodput of the system by an average of 37.86% compared to existing state-of-the-art scheduling solutions.
Liang Zhang 0027, Hongzi Zhu, Wen Fei, Yunzhe Li 0001, Mingjin Zhang, Jiannong Cao 0001, Minyi Guo
IEEE Trans. Computers4
2022 Numerical Feature Representation with Hybrid N-ary Encoding
abstract
Numerical features (e.g., statistical features) are widely used in recommender systems and online advertising. Existing approaches for numerical feature representation in industry are primarily based on discretization. However, hard-discretization based methods (e.g., Equal Distance Discretization) are deficient in continuity while soft-discretization based methods (e.g., AutoDis) lack discriminability. To emphasize both continuity and discriminability for numerical features, we propose an end-to-end representation learning framework named NaryDis. Specifically, NaryDis first leverages hybrid n-ary encoding as an automatic discretization module to generate hybrid-grained discretization results (multiple encoded sequences). Each position of the encoded sequence is assigned with a positional embedding and an intra-ary attention network is leveraged to aggregate the positional embeddings for obtaining ary-wise representations. Then an inter-ary attention is adopted to assemble these representations, which are further constrained by a self-supervised regularization module. Comprehensive experiments on two public datasets are conducted to show the superiority and compatibility of NaryDis. Besides, we deeply investigate the properties of continuity and discriminability. Moreover, we further verify the effectiveness of NaryDis on a large-scale industrial advertisement dataset.
Bo Chen 0023, Huifeng Guo, Weiwen Liu, Yue Ding 0001, Yunzhe Li 0001, Wei Guo 0006, Yichao Wang 0002, Zhicheng He 0001, Ruiming Tang, Rui Zhang 0003
CIKM5
2022 MoGDE: Boosting Mobile Monocular 3D Object Detection with Ground Depth Estimation
abstract
Monocular 3D object detection (Mono3D) in mobile settings (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Due to the near-far disparity phenomenon of monocular vision and the ever-changing camera pose, it is hard to acquire high detection accuracy, especially for far objects. Inspired by the insight that the depth of an object can be well determined according to the depth of the ground where it stands, in this paper, we propose a novel Mono3D framework, called MoGDE, which constantly estimates the corresponding ground depth of an image and then utilizes the estimated ground depth information to guide Mono3D. To this end, we utilize a pose detection network to estimate the pose of the camera and then construct a feature map portraying pixel-level ground depth according to the 3D-to-2D perspective geometry. Moreover, to improve Mono3D with the estimated ground depth, we design an RGB-D feature fusion network based on the transformer structure, where the long-range self-attention mechanism is utilized to effectively identify ground-contacting points and pin the corresponding ground depth to the image feature map. We conduct extensive experiments on the real-world KITTI dataset. The results demonstrate that MoGDE can effectively improve the Mono3D accuracy and robustness for both near and far objects. MoGDE yields the best performance compared with the state-of-the-art methods by a large margin and is ranked number one on the KITTI 3D benchmark.
Yunsong Zhou, Quan Liu 0006, Hongzi Zhu, Yunzhe Li 0001, Shan Chang, Minyi Guo
NeurIPS4
2021 Extracting Attentive Social Temporal Excitation for Sequential Recommendation
abstract
In collaborative filtering, it is an important way to make full use of social information to improve the recommendation quality, which has been proved to be effective because user behavior will be affected by her friends. However, existing works leverage the social relationship to aggregate user features from friends' historical behavior sequences in a user-levelindirect paradigm. A significant defect of the indirect paradigm is that it ignores the temporal relationships between behavior events across users. In this paper, we propose a novel time-aware sequential recommendation framework called Social Temporal Excitation Networks (STEN), which introduces temporal point processes to model the fine-grained impact of friends' behaviors on the user's dynamic interests in an event-leveldirect paradigm. Moreover, we propose to decompose the temporal effect in sequential recommendation into social mutual temporal effect and ego temporal effect. Specifically, we employ a social heterogeneous graph embedding layer to refine user representation via structural information. To enhance temporal information propagation, STEN directly extracts the fine-grained temporal mutual influence of friends' behaviors through themutually exciting temporal network. Besides, user's dynamic interests are captured through theself-exciting temporal network. Extensive experiments on three real-world datasets show that STEN outperforms state-of-the-art baseline methods. Moreover, STEN provides event-level recommendation explainability, which is also illustrated experimentally.
Yunzhe Li 0001, Yue Ding 0001, Bo Chen 0023, Xin Xin 0003, Yule Wang, Yuxiang Shi, Ruiming Tang, Dong Wang 0024
CIKM1
2021 Machine learning-based prediction of survival prognosis in cervical cancer
abstract
BACKGROUND: Accurately forecasting the prognosis could improve cervical cancer management, however, the currently used clinical features are difficult to provide enough information. The aim of this study is to improve forecasting capability by developing a miRNAs-based machine learning survival prediction model. RESULTS: The expression characteristics of miRNAs were chosen as features for model development. The cervical cancer miRNA expression data was obtained from The Cancer Genome Atlas database. Preprocessing, including unquantified data removal, missing value imputation, samples normalization, log transformation, and feature scaling, was performed. In total, 42 survival-related miRNAs were identified by Cox Proportional-Hazards analysis. The patients were optimally clustered into four groups with three different 5-years survival outcome (≥ 90%, ≈ 65%, ≤ 40%) by K-means clustering algorithm base on top 10 survival-related miRNAs. According to the K-means clustering result, a prediction model with high performance was established. The pathways analysis indicated that the miRNAs used play roles involved in the regulation of cancer stem cells. CONCLUSION: A miRNAs-based machine learning cervical cancer survival prediction model was developed that robustly stratifies cervical cancer patients into high survival rate (5-years survival rate ≥ 90%), moderate survival rate (5-years survival rate ≈ 65%), and low survival rate (5-years survival rate ≤ 40%).
Dongyan Ding, Tingyuan Lang, Dongling Zou, Jiawei Tan, Dong Wang 0024, Yunzhe Li 0001, Jingshu Liu, Cui Ma
BMC Bioinform.9
2021 AIRec: Attentive intersection model for tag-aware recommendation
Bo Chen 0023, Yue Ding 0001, Xin Xin 0003, Yunzhe Li 0001, Yule Wang, Dong Wang 0024
Neurocomputing4