Zhonghao Hu

dblp:25/9868 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Snoopy: Effective and Efficient Semantic Join Discovery Via Proxy Columns (Extended Abstract)
Yuxiang Guo 0003, Yuren Mao, Zhonghao Hu, Lu Chen 0001, Yunjun Gao
ICDE3
2026 ESA: Privacy-Preserving Data Sharing Framework With Efficient Fuzzy Search and Access Control
Lan Zhang 0002, Chen Tang 0002, Zhonghao Hu
IEEE Trans. Knowl. Data Eng.6
2026 Gproxy: Communication-Efficient Federated Graph Learning With Efficient Adaptive Proxying
abstract
Federated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinders the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors and develop a novel federated graph learning framework namedGproxy.Gproxyutilizes representation similarity and class correlation to select local proxies for external neighbors. Additionally, we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also design a proxy cache to accelerate the search process by reusing proxy search outcomes for similar external neighbors. Furthermore, we provide a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show thatGproxysignificantly reduces communication cost while maintaining model performance compared to strong baselines.
Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002, Yunhao Yao, Zhonghao Hu
IEEE Trans. Mob. Comput.7
2025 A Controller Placement Approach for SDN with Cooperative Optimization Objectives Using Improved Kepler Optimization Algorithm
abstract
Considering the Controller Placement Problem (CPP) for Software-Defined Networking (SDN), this paper tries to study this problem from the perspective of cooperation. Firstly, total transmission latency of the controller and the switch, the load balancing of the controller and the placement cost of controller are taken as cooperative optimization objectives for SDN CPP. Secondly, in order to minimize these cooperative optimization objectives, kepler optimization algorithm is improved by tent mapping, Levy flight and golden sine strategy to promote both the convergence speed and the optimization ability, and an SDN controller placement approach based on the improved kepler optimization algorithm is then proposed. Finally, the position of the controller as well as the mapping between the controller and the switch obtained by the proposed approach are compared with approaches based on other optimization algorithms. The experimental results show that, the proposed approach can effectively reduce the transmission latency, balance the controller load and reduce the placement cost of the controller.
Zhonghao Hu
CSCWD2
2025 A survey on LoRA of large language models
abstract
Abstract Low-Rank Adaptation (LoRA), which updates the dense neural network layers with pluggable low-rank matrices, is one of the best performed parameter efficient fine-tuning paradigms. Furthermore, it has significant advantages in cross-task generalization and privacy-preserving. Hence, LoRA has gained much attention recently, and the number of related literature demonstrates exponential growth. It is necessary to conduct a comprehensive overview of the current progress on LoRA. This survey categorizes and reviews the progress from the perspectives of (1) downstream adaptation improving variants that improve LoRA’s performance on downstream tasks; (2) cross-task generalization methods that mix multiple LoRA plugins to achieve cross-task generalization; (3) efficiency-improving methods that boost the computation-efficiency of LoRA; (4) data privacy-preserving methods that use LoRA in federated learning; (5) application. Besides, this survey also discusses the future directions in this field.
Yuren Mao, Yuhang Ge, Yijiang Fan, Yu Mi, Zhonghao Hu, Yunjun Gao
Frontiers Comput. Sci.6
2025 BIRDIE: Natural Language-Driven Table Discovery Using Differentiable Search Index
abstract
Natural language (NL)-driven table discovery identifies relevant tables from large table repositories based on NL queries. While current deep-learning-based methods using the traditional dense vector search pipeline, i.e., representation-index-search , achieve remarkable accuracy, they face several limitations that impede further performance improvements: (i) the errors accumulated during the table representation and indexing phases affect the subsequent search accuracy; and (ii) insufficient query-table interaction hinders effective semantic alignment, impeding accuracy improvements. In this paper, we propose a novel framework Birdie, using a differentiate search index. It unifies the indexing and search into a single encoder-decoder language model, thus getting rid of error accumulations. Birdie first assigns each table a prefix-aware identifier and leverages a large language model-based query generator to create synthetic queries for each table. It then encodes the mapping between synthetic queries/tables and their corresponding table identifiers into the parameters of an encoder-decoder language model, enabling deep query-table interactions. During search, the trained model directly generates table identifiers for a given query. To accommodate the continual indexing of dynamic tables, we introduce an index update strategy via parameter isolation, which mitigates the issue of catastrophic forgetting. Extensive experiments demonstrate that Birdie outperforms state-of-the-art dense methods by 16.8% in accuracy, and reduces forgetting by over 90% compared to other continual learning approaches.
Yuxiang Guo 0003, Zhonghao Hu, Yuren Mao, Baihua Zheng, Yunjun Gao, Mingwei Zhou
Proc. VLDB Endow.2
2025 Snoopy: Effective and Efficient Semantic Join Discovery via Proxy Columns
abstract
Semantic join discovery, which aims to find columns in a table repository with high semantic joinabilities to a query column, is crucial for dataset discovery. Existing methods can be divided into two categories: cell-level methods and column-level methods. However, neither of them ensures both effectiveness and efficiency simultaneously. Cell-level methods, which compute the joinability by counting cell matches between columns, enjoy ideal effectiveness but suffer poor efficiency. In contrast, column-level methods, which determine joinability only by computing the similarity of column embeddings, enjoy proper efficiency but suffer poor effectiveness due to the issues occurring in their column embeddings: (i) semantics-joinability-gap, (ii) size limit, and (iii) permutation sensitivity. To address these issues, this paper proposes to compute column embeddings via proxy columns; furthermore, a novel column-level semantic join discovery framework,${\sf Snoopy}$, is presented, leveraging proxy-column-based embeddings to bridge effectiveness and efficiency. Specifically, the proposed column embeddings are derived from the implicit column-to-proxy-column relationships, which are captured by the lightweight approximate-graph-matching-based column projection. To acquire good proxy columns for guiding the column projection, we introduce a rank-aware contrastive learning paradigm. Extensive experiments on four real-world datasets demonstrate that${\sf Snoopy}$outperforms SOTA column-level methods by 16% in Recall@25 and 10% in NDCG@25, and achieves superior efficiency—being at least 5 orders of magnitude faster than cell-level solutions, and 3.5× faster than existing column-level methods.
Yuxiang Guo 0003, Yuren Mao, Zhonghao Hu, Lu Chen 0001, Yunjun Gao
IEEE Trans. Knowl. Data Eng.3