Jiaquan Zhang

dblp:188/5932 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 61% Language models and text generation · 30% Multi-agent systems · 9%
Databases, data mining, and information retrieval
2 papers
Data mining · 65% Database system architecture and tuning · 25% Machine learning and data management · 10%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 91% Cloud and datacenter computing · 9%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
agent memory
1.012026
Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models
1.012026
Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026
Data mining
clustering
1.012026
Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026
Data mining › clustering
federated clustering
1.012026
Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026
Privacy and data protection › differential privacy
local differential privacy
1.012026
Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026
Privacy and data protection › privacy-preserving data analysis
privacy-preserving clustering
1.012026
Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026
Storage systems › scalable storage › big data storage
data lake storage
0.812024
Separation Is for Better Reunion: Data Lake Storage at Huawei · ICDE 2024
Storage systems › distributed storage
disaggregated storage
0.812024
Separation Is for Better Reunion: Data Lake Storage at Huawei · ICDE 2024
Storage systems
storage reliability
0.812024
Separation Is for Better Reunion: Data Lake Storage at Huawei · ICDE 2024
Knowledge, reasoning and agents › Multi-agent systems
agentic AI
0.312026
Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026
Machine learning and data management › privacy-preserving machine learning
federated learning
0.312026
Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026

Methods — techniques the papers use, named apart from their topics

topological aggregation · 2.0persistent homology · 2.0gravitational potential field · 2.0tiered storage · 1.5metadata acceleration · 1.5erasure coding · 1.5small language model · 1.0memory management · 1.0
YearPublicationVenuePosition
2026 Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy
abstract
Clustering non-independent and identically distributed (non-IID) data under local differential privacy (LDP) in federated settings presents a critical challenge: preserving privacy while maintaining accuracy without iterative communication. Existing one-shot methods rely on unstable pairwise centroid distances or neighborhood rankings, degrading severely under strong LDP noise and data heterogeneity. We present Gravitational Federated Clustering (GFC), a novel approach to privacy-preserving federated clustering that overcomes the limitations of distance-based methods under varying LDP. Addressing the critical challenge of clustering non-IID data with diverse privacy guarantees, GFC transforms privatized client centroids into a global gravitational potential field where true cluster centers emerge as topologically persistent singularities. Our framework introduces two key innovations: (1) a client-side compactness-aware perturbation mechanism that encodes local cluster geometry as "mass" values, and (2) a server-side topological aggregation phase that extracts stable centroids through persistent homology analysis of the potential field's superlevel sets. Theoretically, we establish a closed-form bound between the privacy budget ε and centroid estimation error, proving the potential field's Lipschitz smoothing properties exponentially suppress noise in high-density regions. Empirically, GFC outperforms state-of-the-art methods on ten benchmarks, especially under strong LDP constraints (ε < 1), while maintaining comparable performance at lower privacy budgets. By reformulating federated clustering as a topological persistence problem in a synthetic physics-inspired space, GFC achieves unprecedented privacy-accuracy trade-offs without iterative communication, providing a new perspective for privacy-preserving distributed learning.
Yunbo Long, Jiaquan Zhang, Alexandra Brintrup
AAAI2
2026 Lightweight LLM Agent Memory with Small Language Models
abstract
Jiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Pengcheng Zheng, Zhicheng Wang, Ping Guo, Fan Mo, Sung-Ho Bae, Jie Zou, Jiwei Wei, Yang Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Sung-Ho Bae, Jie Zou 0001, Jiwei Wei, Yang Yang 0002
ACL (1)1
2024 Separation Is for Better Reunion: Data Lake Storage at Huawei
abstract
Huawei collaborates with some Chinese large busi-ness companies to store and process exabytes of nationwide operational data in data lake storage to provide business insights. Specifically, our customers will ask to store and process massive log message data to support their real-time and decision-making applications. Thus, we need computation and storage components in the analytic platform to process and store these data cost-efficiently. To meet these user requirements, we have designed a storage system in data lake, StreamLake, which introduces a novel design to serve log message streaming and batch data processing in distributed storage, with high scalability, efficiency, reliability and low cost. Specifically, we introduce a stream (storage) object as a storage abstraction for message streaming data to achieve the storage-disaggregated architecture with high scalability and reliability. Moreover, we utilize the erasure coding and tiered storage to save the storage cost, and furthermore, the stream object can be automatically converted to a table object such that cost-effective stream and batch data processing can be achieved. For tabular data, we implement the lakehouse functionality to support ACID via the table object, with a metadata acceleration to improve the efficiency of data access between the compute and storage engines. Also, we design a LakeBrain optimizer at the storage side to optimize the query performance and resource utilization under the storage-disaggregated architecture. Finally, we have also deployed StreamLake in China Mobile, the world's largest mobile network operator to serve over 20PB production data, and the results demonstrate improvements of 30% to 4x in terms of query performance and over 37% in terms of cost saving.
Chengliang Chai, Haohai Ma, Zhenyong Fan, Jiaquan Zhang, Rui Zhang 0003, Duanshun Li, Keji Huang, Guangbin Meng, Yuefeng Zhou, Lirong Jian, Jiwu Shu, Ye Yuan 0001, Guoren Wang, Guoliang Li 0001
ICDE8
2022 Self-adaptive Multi-scale Aggregation Network for Stereo Matching
abstract
Nowadays stereo matching architectures based on convolutional neural network has achieved remarkable performance. However the existing methods still lack the capability to find the correspondence in ill-posed regions. In this paper we present Self-adaptive Multi-scale Aggregation Network (SMA-Net) for stereo matching. First of all, we construct the cost volume through multi-channels group-wise correlation, which divided the features into groups with different number of channels to enhance the ability of measuring the similarities of features from stereo images. Secondly the self-adaptive cost aggregation is used to regularize the two scale cost volumes from different aggregation branches with intermediate supervise. We conduct comprehensive experiments on SceneFlow, KITTI2012, and KITTI2015 datasets. The competitive results prove that the approach in this paper outperforms many other stereo matching algorithms especially in ill-posed regions.
Shuiqiang Ye, Jiaquan Zhang, Xin'an Wang, Qifei Dai, Zhengzhong Yu, Fuchi Li, Yong Zhao 0010
ICPR3
2021 A Data-Driven Analysis of K-12 Students' Participation and Learning Performance on an Online Supplementary Learning Platform
abstract
Due to the limitation of public school systems, many students pursue private supplementary tutoring for improving their academic performance. Different from public schools, the private online education provides diverse courses and satisfy differentiated demands of the students. Students’ behavior and performance in online supplementary learning are relevant to not only personal attributes, but also some factors such as city levels, grades and family situation. Existing studies mostly rely on panel survey/questionnaire data and few studied online private tutoring. In this paper, with 11,392 anonymous K-12 students’ 3-year learning data from one of the world’s largest online extra-curricular education platforms, we uncover students’ online learning behaviors and infer the impact of students’ home location, family socioeconomic situation and attended school’s reputation/rank on the students’ private tutoring course participation and learning outcomes. Further analysis suggests that such impact may be largely attributed to the inequality of access to educational resources in different cities and the inequality in family socioeconomic status. Finally, we study the predictability of students’ performance and behaviors using machine learning algorithms with different groups of features, showing students’ online learning performance can be predicted with MAE< 10%.
Jiaquan Zhang, Xin Gao 0034, Hui Chen 0017, Jar-der Luo, Xiaoming Fu 0001
ICCCN1
2020 Identifying unfamiliar callers' professions from privacy-preserving mobile phone data
abstract
Identifying an unfamiliar caller's profession is important to protect citizens' personal safety and property. Due to limited data protection of many popular online services in some countries such as taxi hailing or takeouts ordering, many users encounter an increasing number of phone calls from strangers. This may aggravate the situation that criminals pretend to be delivery staff or taxi drivers, bringing threats to the society. Additionally, many people nowadays suffer from excessive digital marketing and fraud phone calls because of personal information leakage. However, previous works on malicious call detection only focused on binary classification, and do not work for identification of multiple professions. We observed that web service requests issued from users' mobile phones which may show their Apps preferences, spatial and temporal patterns, and other profession related information. This offers us a hint to identify unfamiliar callers. In fact, some previous works already leveraged raw data from mobile phones (which includes sensitive information) for personality studies. However, accessing users' mobile phone raw data may violate the more and more strict private data protection policies or regulations (e.g. GDPR 71). Using appropriate statistical methods to eliminate private information and preserve personal characteristics, provides a way to identify mobile phone callers without privacy concern. In this paper, we exploit privacy-preserving mobile data to develop a model which can automatically identify the callers who are divided into four categories of users: normal users (other professions), taxi drivers, delivery and takeouts staffs, telemarketers and fraudsters. The validation results over an anonymized dataset of 1,282 users with a period of 3 months in Shanghai City prove that the proposed model could achieve an accuracy of 75+%.
Jiaquan Zhang, Xiaoming Yao, Xiaoming Fu 0001
MSN1