Hao Zhang 0016

dblp:55/2270-16 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-6769-2115ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Stop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the Spot
abstract
Vision-Language Models (VLMs) have achieved impressive performance across various tasks, but often struggle to apply newly introduced visual concepts during inference. A common failure pattern is what we call Mixing Things Up: VLMs frequently confuse concept names, resulting in vague descriptions and failure to ground the concept correctly. Existing approaches mainly address person-related concepts through text prompts or tokenizer modifications. However, VLMs still miss or misinterpret untrained visual concepts, underscoring the need to learn new concepts directly from visual input, without relying on prior textual injection. To overcome these limitations, we propose BISCUIT (Basis-aligned Inference through Structured Concept Unification and Identification-aware Tuning), a two-step training method. Step I proposes a dual-stream structure-aware vision encoder that fuses RGB and edge-based embeddings within a shared basis space to enhance concept recognition. Step II enhances generation quality through identification-aware tuning, which encourages alignment between the generated text and the newly introduced visual concepts. Existing methods mainly focus on person concepts and lack comprehensive evaluation across diverse visual categories. We further propose a benchmark BiscuitVQA to evaluate VLMs performance on recognizing and applying novel image-introduced concepts across diverse concept types and task types, including real people, cartoons, animals, and symbolic content. We apply BISCUIT to LLaVA-1.5 and Qwen2.5-VL, achieving competitive results among open-source models and narrowing the gap to Gemini-2.5 and GPT-4o. Interestingly, our BISCUIT maintains strong generalization, showing minimal degradation on other downstream tasks.
Jiahua Bao, Siyao Cheng, Jiaxing Du, Yuhang Jia, Boyang Niu, Zeming Lang, Changjiang He, Hao Zhang 0016, Jie Liu 0001
AAAI8
2026 An online multi-agent path finding algorithm for large-scale puzzle-based conveyor system
Mingrui Yin, Hao Zhang 0016, Chenxin Cai, Jie Liu 0001
Expert Syst. Appl.2
2026 A Lyapunov-based client selection approach to handle system-induced heterogeneity in federated learning
Tian Ren, Hao Zhang 0016, Weilin Liao, Siyao Cheng, Jie Liu 0001
Knowl. Based Syst.2
2026 Chain-of-Detection enables robust and efficient jailbreak defense
Tingting Wu 0007, Hao Zhang 0016
Neural Networks2
2026 Model-Heterogeneous Federated Learning With Bidirectional Knowledge Distillation
Hao Zhang 0016, Yaolin Zhu, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001
IEEE Trans. Mob. Comput.1
2025 BOLT: Fewer Tokens but More Performance Retention for Efficient Vision-Language Models Inference
abstract
Vision-Language Models (VLMs) have achieved significant advances across various downstream tasks. However, as their performance improves, the increasing number of parameters results in slower prefilling speeds and longer inference times. To overcome these limitations, we observe that most VLMs do not require a large number of image tokens for inference, we propose BOLT (Basis-Oriented Lightweight Token-Trimming), a training-free and cross-attention-free token compression method. Unlike existing approaches, BOLT addresses the challenge of insufficient visual cues in textual prompts by leveraging token internal data distributions. We categorize tokens into three types: key tokens, proxy tokens, and remaining tokens. Then, by applying basis space similarity, we merge and filter the remaining tokens with the proxy tokens to retain the most informative ones. To account for the differences in VLM architectures and model sizes, we evaluate BOLT on LLaVA-Next-Llama3 and LLaVA-1.5 (7B and 13B). Our results show that BOLT achieves state-of-the-art performance, with a 90% token compression ratio leading to a 3.3× increase in pre-filling speed and a 1.5× improvement in inference speed, outperforming other methods.
Jiahua Bao, Siyao Cheng, Jiaxing Du, Changjiang He, Zeming Lang, Hao Zhang 0016, Jie Liu 0001
ACM Multimedia6
2025 Dynamic clustered federated learning via adaptive distribution similarity computation
Tian Ren, Siyao Cheng, Hao Zhang 0016, Jie Liu 0001
Comput. Networks3
2025 Omniveyor: An Assembled Logistics Sorting System Powered by Reinforcement Learning
abstract
To improve logistics efficiency, smart logistics sorting is an inevitable trend in logistics development. Existing smart logistics sorting systems suffer from high construction costs or limited scalability. To solve these problems, we design a brand-new two-dimensional conveyor system called Omniveyor, which transports and sorts high-density packages within a limited space. It is assembled from multiple repetitive square conveyor modules, achieving the goal of cost-effectiveness and easy maintenance. To realize automatic sorting, we model the planning problem on Omniveyor and propose a scheduling strategy named MMPPO by reinforcement learning. Unlike traditional path planning, MMPPO assigns actions to modules rather than packages, which reduces scheduling overhead in high-throughput scenarios. Furthermore, we develop a simulation environment to inspect the effectiveness of our method, which solves an intractable package-following problem that has plagued simulation implementation in this field. Experimental results show that MMPPO outperforms baselines in terms of throughput and overall consumption at high densities. Besides, we implement a physical prototype of Omniveyor to validate its feasibility. Note to Practitioners—The motivation of the paper is to solve the sorting problems in the logistics system. Existing logistics sorting systems often have problems such as low sorting efficiency and high costs, making them unsuitable for small-sized warehouses. In this paper, we present a modular 2D desktop logistics system that is both scalable and efficient for sorting. We mathematically describe the platform package transportation process and propose an algorithm that addresses scheduling and planning problems, capable of continuous planning under pipeline input. Preliminary simulation tests indicate that our approach is feasible, and we have also built a small-scale prototype. In future research, we will further expand the scale of the platform and conduct research.
Mingrui Yin, Hao Zhang 0016, Chenxin Cai, Meiyan Liang, Jie Liu 0001
IEEE Trans Autom. Sci. Eng.2
2024 Federated Edge Learning with Blurred or Pseudo Data Sharing
abstract
Edge servers and mobile devices are often assigned a large number of computing tasks. However, the data involved in computing tasks is often sensitive in terms of privacy. Our initial proposal is a federated edge learning strategy based on real-world scenarios, which combines blurred data or pseudo shared data. Federated learning is used to train device models with the aim of protecting privacy while enabling mobile devices to more effectively utilize data for decision-making. In the case of limited energy on mobile devices, we propose a federated edge learning algorithm with blurred data sharing. This algorithm can generate more accurate models by uploading partially blurred data. In order to further improve model accuracy and protect privacy of mobile devices, we propose a federated edge learning algorithm with pseudo data sharing based on dataset distillation and generative adversarial networks (GANs) in scenarios with relatively sufficient energy. The experimental results on several traditional datasets show that our proposed algorithms outperform traditional algorithms in terms of accuracy and energy consumption.
Yinlong Li, Hao Zhang 0016, Siyao Cheng, Jie Liu 0001
ICPP2
2024 StressViT: Splitting and Compressing Vision Transformer Through Edge-Cloud Collaboration
Changyao Lin, Yi Liu 0096, Chengxiang Li, Hao Zhang 0016, Jing Jin 0003, Jie Liu 0001
ICPR (5)5
2024 A Stackelberg-Game-Based Framework for Edge Pricing and Resource Allocation in Mobile Edge Computing
abstract
Nowadays, Mobile Edge Computing (MEC) appears as a new computing paradigm with its ability to utilize the computing power of both local devices and edge servers. In MEC, edge pricing and resource allocation are two important problems. Edge servers make a profit by selling computing services to users. To maximize their revenue, they need to determine an appropriate price for each user, and decide the amount of resources allocated to each user. However, none of the existing works consider the effect of users’ task assignment strategy on the revenue of the edge. In fact, edge pricing and resource allocation will affect the users’ task offloading decision, as they expect to minimize their total cost. In turn, the users’ decision will also influence the revenue of the edge. Therefore, the interaction between mobile users and edge servers should be considered carefully and the interests of both sides need to be maximized simultaneously. In this paper, we model the interaction between the two sides as a Stackelberg game. First, given a specified edge pricing and resource allocation strategy, we derive a near-optimal task assignment strategy for each user to minimize the total cost based on a greedy algorithm UTA-G. Then, by applying the backward induction method, two pricing and resource allocation schemes with different granularity, i.e., EPRA-U and EPRA-T are proposed to bring higher revenue to the edge. Experimental results demonstrate that all the proposed algorithms can have good performance in task-intensive, resource-deficient and workload-heavy scenarios.
Siyao Cheng, Tian Ren, Hao Zhang 0016, Jiayan Huang, Jie Liu 0001
IEEE Internet Things J.3
2024 CC-FedAvg: Computationally Customized Federated Averaging
abstract
Federated learning (FL) is an emerging paradigm to train model with distributed data from numerous Internet of Things (IoT) devices. It inherently assumes a uniform capacity among participants. However, due to different conditions such as differing energy budgets or executing parallel unrelated tasks, participants have diverse computational resources in practice. Participants with insufficient computation budgets must plan for the use of restricted computational resources appropriately; otherwise, they would be unable to complete the entire training procedure, resulting in model performance decline. To address this issue, we propose a strategy for estimating local models without computationally intensive iterations. Based on it, we propose computationally customized federated averaging (CC-FedAvg), which allows participants to determine whether to perform traditional local training or model estimation in each round based on their current computational budgets. Both theoretical analysis and exhaustive experiments indicate that CC-FedAvg has the same convergence rate and comparable performance as FedAvg without resource constraints. Furthermore, CC-FedAvg can be viewed as a computation-efficient version of FedAvg that retains model performance while considerably lowering computation overhead.
Hao Zhang 0016, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001
IEEE Internet Things J.1
2024 Uncertainty-guided label correction with wavelet-transformed discriminative representation enhancement
Tingting Wu 0007, Hao Zhang 0016, Minji Tang, Bing Qin 0001, Ting Liu 0001
Neural Networks3
2024 DiscrimLoss: A Universal Loss for Hard Samples and Incorrect Samples Discrimination
abstract
Given data with label noise (i.e., incorrect data), deep neural networks would gradually memorize the label noise and impair model performance. To relieve this issue, curriculum learning is proposed to improve model performance and generalization by ordering training samples in a meaningful (e.g., easy to hard) sequence. Previous work takes incorrect samples as generic hard ones without discriminating between hard samples (i.e., hard samples in correct data) and incorrect samples. Indeed, a model should learn from hard samples to promote generalization rather than overfit to incorrect ones. In this article, we address this problem by appending a novel loss functionDiscrimLoss, on top of the existing task loss. Its main effect is to automatically and stably estimate the importance of easy samples and difficult samples (including hard and incorrect samples) at the early stages of training to improve the model performance. Then, during the following stages, DiscrimLoss is dedicated to discriminating between hard and incorrect samples to improve the model generalization. Such a training strategy can be formulated dynamically in a self-supervised manner, effectively mimicking the main principle of curriculum learning. Experiments on image classification, image regression, text sequence regression, and event relation reasoning demonstrate the versatility and effectiveness of our method, particularly in the presence of diversified noise levels.
Tingting Wu 0007, Hao Zhang 0016, Jinglong Gao, Minji Tang, Bing Qin 0001, Ting Liu 0001
IEEE Trans. Multim.3
2023 ACIGS: An automated large-scale crops image generation system based on large visual language multi-modal models
abstract
Smart agriculture requires an extensive convergence of information technology and agriculture. Attaining intelligence mandates an enormous amount of data to train models. However, it is challenging to acquire a large number of crop image data, limiting the application and growth of computer vision technology in agriculture. To address this problem, we designed a crop image generation system that combines a large language model with visual language multi-modal large models to augment the scale, variety, and resolution of crop image data. First, the system inputs existing real crop images into the visual language multimodal model to extract features and represent crop images in text form. Then, the system passes the crop text representation to the language model for cleaning and processing, which generates prompts to create crop images. The prompts are input into the visual language multi-modal model to generate crop images based on text representation of crops. The resulting crop images undergo image quality evaluation in the visual language multimodal model, and high-quality crop images are saved to the crop image dataset based on the quality evaluation. These steps lead to the formation of the final generated crop image dataset. The experimental results indicate that the crop images generated using the proposed system are similar to but different from the example images. This characteristic enables the expansion of crop data while circumventing redundancy and allowing for resolution control, which is crucial for dense segmentation tasks. Using this method, the existing data can be enlarged up to 7.5 times.
Bolong Liu, Hao Zhang 0016, Jie Liu 0001, Qiang Wang 0001
SECON2
2023 Dynamic adaptive workload offloading strategy in mobile edge computing networks
Yinlong Li, Siyao Cheng, Hao Zhang 0016, Jie Liu 0001
Comput. Networks3
2023 Dynamic layer-wise sparsification for distributed deep learning
Hao Zhang 0016, Tingting Wu 0007, Zhifeng Ma, Jie Liu 0001
Future Gener. Comput. Syst.1
2023 Data-Augmentation-Based Federated Learning
abstract
With the rapid growth of the number of devices generating and collecting data, dispersion becomes an important feature of data in Internet of Things. Federated learning (FL) provides a feasible way to mine information in such distributed data. It involves training machine learning models over multiple distributed participants without raw data transmission. However, due to the data heterogeneity among participants, the performance of the FL model degrades dramatically. Currently, improved methods mainly reduce data heterogeneity from the perspective of modifying the process of model training, which usually have problems, such as high-resource consumption or the need for auxiliary data. In this article, we enhance FL model from another perspective, focusing on data rather than model training. We reduce data heterogeneity by enhancing the trained local data to improve FL performance. Specifically, we propose an FL method based on data augmentation (abbreviated as FedM-UNE), implementing the classic data augmentation method MixUp in federated scenarios without transferring raw data. Furthermore, in order to adapt this method to regression tasks, we first modify MixUp by bilateral neighborhood expansion (MixUp-BNE), and then propose a federated data augmentation method named FedM-BNE based on it. Compared with the conventional FL method, both FedM-UNE and FedM-BNE increase negligible overhead. To demonstrate the effectiveness, we conduct exhaustive experiments on six data sets employing a variety of loss functions. The results indicate that FedM-UNE and FedM-BNE consistently improve the performance of the FL model. Moreover, our methods are compatible with existing FL enhancements, which yield further improvements in performance.
Hao Zhang 0016, Qingying Hou, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001
IEEE Internet Things J.1
2023 FedCos: A Scene-Adaptive Enhancement for Federated Learning
abstract
Federated learning (FL) training global machine learning models over distributed edge devices has attracted sustained attentions. However, the heterogeneity of client data severely degrades the performance of FL compared with that in centralized training. On the one hand, it slows down or even stalls global updates, leading to inefficient communication. On the other hand, it enlarges the distances between local models, resulting in an aggregated global model with poor performance. Fortunately, these shortcomings can be mitigated by reducing the angle between the directions in which a local model move. Based on this observation, we propose FedCos, which reduces the directional inconsistency of local models by introducing a cosine-similarity penalty. It promotes local model iterations toward an auxiliary global direction. Moreover, our approach is auto-adapted to various non-identically and independently distributed (IID) settings without an elaborate selection of hyperparameters. Experimental results on both vision and language tasks with a variety of models (including CNN, ResNet, LSTM, etc.) show that FedCos outperforms the well-known baselines and can enhance them under a variety of FL scenes, including varying degrees of data heterogeneity, different number of participants, and cross-silo and cross-device settings. Besides, FedCos improves the communication efficiency by 2–5 times. With the help of FedCos, multiple FL methods require significantly fewer communication rounds than before to obtain a comparable model.
Hao Zhang 0016, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001
IEEE Internet Things J.1
2023 MM-RNN: A Multimodal RNN for Precipitation Nowcasting
abstract
Precipitation nowcasting, the high-resolution forecasting of precipitation in a short term, is essential in various applications in the real world. Previous deep learning methods use huge samples to learn potential laws, and the learning process lacks regularity, making it difficult to model the complex nonlinear precipitation phenomenon. Inspired by traditional numerical weather prediction models, we propose the MultiModal RNN (MM-RNN), which introduces knowledge of elements to guide precipitation prediction. This constraint forces the movement of precipitation to follow the underlying atmospheric motion laws. MM-RNN not only can provide accurate precipitation nowcasting but other meteorological elements predictions. Besides, it has high flexibility and is compatible with multiple RNN models, such as ConvLSTM, PredRNN, MIM, MotionRNN, etc. We conduct experiments on two multimodal datasets (MeteoNet and RAIN-F) and the results indicate that MM-RNN is superior to common RNN (MultiScale RNN, MS-RNN) using a single radar modality. For the MeteoNet, compared to MS-MotionRNN, the CSI (R ⩾ 10) of MM-MotionRNN increases by 23.4%, and the MSE of MM-MotionRNN decreases by 6.7%. For the RAIN-F, compared to MS-MIM, the HSS (R ⩾ 5) of MM-MIM increases by 209.4%, and the B-MSE of MM-MIM decreases by 4.6%.
Zhifeng Ma, Hao Zhang 0016, Jie Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 STGN: an Implicit Regularization Method for Learning with Noisy Labels in Natural Language Processing
abstract
Noisy labels are ubiquitous in natural language processing (NLP) tasks.Existing work, namely learning with noisy labels in NLP, is often limited to dedicated tasks or specific training procedures, making it hard to be widely used.To address this issue, SGD noise has been explored to provide a more general way to alleviate the effect of noisy labels by involving benign noise in the process of stochastic gradient descent.However, previous studies exert identical perturbation for all samples, which may cause overfitting on incorrect ones or optimizing correct ones inadequately.To facilitate this, we propose a novel stochastic tailor-made gradient noise (STGN), mitigating the effect of inherent label noise by introducing tailor-made benign noise for each sample.Specifically, we investigate multiple principles to precisely and stably discriminate correct samples from incorrect ones and thus apply different intensities of perturbation to them.A detailed theoretical analysis shows that STGN has good properties, beneficial for model generalization.Experiments on three different NLP tasks demonstrate the effectiveness and versatility of STGN.Also, STGN can boost existing robust training methods. 1
Tingting Wu 0007, Minji Tang, Hao Zhang 0016, Bing Qin 0001, Ting Liu 0001
EMNLP4
2022 Aperiodic Local SGD: Beyond Local SGD
abstract
Variations of stochastic gradient decedent (SGD) methods are at the core of training deep neural network models. However, in distributed deep learning, where multiple computing devices and data segments are employed in the training process, the performance of SGD can be significantly limited by the overhead of gradient communication. Local SGD methods are designed to overcome this bottleneck by averaging individual gradients trained over parallel workers after multiple local iterations. Currently, both for theoretical analyses and for practical applications, most studies employ periodic synchronization scheme by default, while few of them focus on the aperiodic schemes to obtain better performance models with limited computation and communication overhead. In this paper, we investigate local SGD with an arbitrary synchronization scheme to answer two questions: (1) Is the periodic synchronization scheme best? (2) If not, what is the optimal one? First, for any synchronization scheme, we derive the performance boundary with fixed overhead, and formulate the performance optimization under given computation and communication constraints. Then we find a succinct property of the optimal scheme that the local iteration number decreases as training continues, which indicates the periodic one is suboptimal. Furthermore, with some reasonable approximations, we obtain an explicit form of the optimal scheme and propose Aperiodic Local SGD (ALSGD) as an improved substitute for local SGD without any overhead increment. Our experiments also confirm that with the same computation and communication overhead, ALSGD outperforms local SGD in performance, especially for heterogeneous data.
Hao Zhang 0016, Tingting Wu 0007, Siyao Cheng, Jie Liu 0001
ICPP1
2022 PrecipLSTM: A Meteorological Spatiotemporal LSTM for Precipitation Nowcasting
abstract
Accurate and timely nowcasting precipitation has huge social and economic benefits. However, the changes of clouds including expansion, dissipation, and distortion are extremely complex, which exacerbates the difficulty of forecasting. Fortunately, it still follows certain meteorological laws, which can be explored based on spatiotemporal information but are not fully considered by previous models. In this paper, we design two modules to focus on these messages based on atmospheric characteristics. Specifically, the SLAM (Spatial Local Attention Memory) module combines local attention and memory mechanism to capture the meteorological spatial relationship, while the TDM (Time Difference Memory) module combines differential technology and memory mechanism to capture the meteorological temporal variants. We combine these two modules with PredRNN and propose PrecipLSTM to sufficiently capture the spatiotemporal dependencies of radar data. We do exhaustive experiments with five baselines on four radar datasets. It is verified that PrecipLSTM achieves state-of-the-art results with fewer parameters than the previous state-of-the-art method.
Zhifeng Ma, Hao Zhang 0016, Jie Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 GMDA: An Automatic Data Analysis System for Industrial Production
Zhiyu Liang, Hongzhi Wang 0001, Hao Zhang 0016, Hengyu Guo
DASFAA (3)3
2015 Mobile cloud computing based privacy protection in location-based information survey applications
abstract
Abstract Nowadays, location‐based service (LBS) has become pervasive. Given its high utility value, LBS, however, presents serious privacy concerns for cautious users. In this paper, we investigate privacy preserving for location‐based information survey application, which calculates the geographic distribution of user's information. The design objective is twofold: (i) calculate an information distribution for a pool of mobile users and (ii) protect the location and value privacy of individual user, in the presence of malicious servers and possible corrupted users. Our proposed solution leverages a mobile cloud computing paradigm, in which each mobile device is replicated with a system‐level clone in cloud. The computing of distribution function is distributed among the set of cloud clones via a P2P protocol. We further enhance our basic scheme with the multiple aggregation mechanism, aiming to protect the correctness of the aggregate result from the active attacker. Compared to the approaches based on centralized server or aggregate proxy, our proposed scheme and its enhanced version are advantageous in avoiding single point of failure/attack, load balancing, and overhead reduction. Simulation results verify these advantages and the protection to the correctness of aggregate result and suggest that our proposed scheme is suitable for large‐scale applications. Copyright © 2014 John Wiley & Sons, Ltd.
Hao Zhang 0016, Nenghai Yu, Yonggang Wen 0001
Secur. Commun. Networks1
2014 Towards optimal noise distribution for privacy preserving in data aggregation
Hao Zhang 0016, Nenghai Yu, Yonggang Wen 0001, Weiming Zhang 0001
Comput. Secur.1