EDBT 2026 Demo / reviewers in the wild / expert
Jiaquan Zhang
dblp:188/5932
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 61% Language models and text generation · 30% Multi-agent systems · 9% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 65% Database system architecture and tuning · 25% Machine learning and data management · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 91% Cloud and datacenter computing · 9% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › LLM agents
agent memory |
1.0 | 1 | 2026 | Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models |
1.0 | 1 | 2026 | Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026 |
Data mining
clustering |
1.0 | 1 | 2026 | Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026 |
Data mining › clustering
federated clustering |
1.0 | 1 | 2026 | Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026 |
Privacy and data protection › differential privacy
local differential privacy |
1.0 | 1 | 2026 | Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026 |
Privacy and data protection › privacy-preserving data analysis
privacy-preserving clustering |
1.0 | 1 | 2026 | Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026 |
Storage systems › scalable storage › big data storage
data lake storage |
0.8 | 1 | 2024 | Separation Is for Better Reunion: Data Lake Storage at Huawei · ICDE 2024 |
Storage systems › distributed storage
disaggregated storage |
0.8 | 1 | 2024 | Separation Is for Better Reunion: Data Lake Storage at Huawei · ICDE 2024 |
Storage systems
storage reliability |
0.8 | 1 | 2024 | Separation Is for Better Reunion: Data Lake Storage at Huawei · ICDE 2024 |
Knowledge, reasoning and agents › Multi-agent systems
agentic AI |
0.3 | 1 | 2026 | Lightweight LLM Agent Memory with Small Language Models · ACL (1) 2026 |
Machine learning and data management › privacy-preserving machine learning
federated learning |
0.3 | 1 | 2026 | Topological Federated Clustering via Gravitational Potential Fields Under Local Differential Privacy · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
topological aggregation · 2.0persistent homology · 2.0gravitational potential field · 2.0tiered storage · 1.5metadata acceleration · 1.5erasure coding · 1.5small language model · 1.0memory management · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topological Federated Clustering via Gravitational Potential Fields Under Local Differential PrivacyabstractClustering non-independent and identically distributed (non-IID) data under local differential privacy (LDP) in federated settings presents a critical challenge: preserving privacy while maintaining accuracy without iterative communication. Existing one-shot methods rely on unstable pairwise centroid distances or neighborhood rankings, degrading severely under strong LDP noise and data heterogeneity. We present Gravitational Federated Clustering (GFC), a novel approach to privacy-preserving federated clustering that overcomes the limitations of distance-based methods under varying LDP. Addressing the critical challenge of clustering non-IID data with diverse privacy guarantees, GFC transforms privatized client centroids into a global gravitational potential field where true cluster centers emerge as topologically persistent singularities. Our framework introduces two key innovations: (1) a client-side compactness-aware perturbation mechanism that encodes local cluster geometry as "mass" values, and (2) a server-side topological aggregation phase that extracts stable centroids through persistent homology analysis of the potential field's superlevel sets. Theoretically, we establish a closed-form bound between the privacy budget ε and centroid estimation error, proving the potential field's Lipschitz smoothing properties exponentially suppress noise in high-density regions. Empirically, GFC outperforms state-of-the-art methods on ten benchmarks, especially under strong LDP constraints (ε < 1), while maintaining comparable performance at lower privacy budgets. By reformulating federated clustering as a topological persistence problem in a synthetic physics-inspired space, GFC achieves unprecedented privacy-accuracy trade-offs without iterative communication, providing a new perspective for privacy-preserving distributed learning. Yunbo Long, Jiaquan Zhang, Alexandra Brintrup |
AAAI | 2 |
| 2026 | Lightweight LLM Agent Memory with Small Language ModelsabstractJiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Pengcheng Zheng, Zhicheng Wang, Ping Guo, Fan Mo, Sung-Ho Bae, Jie Zou, Jiwei Wei, Yang Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Sung-Ho Bae, Jie Zou 0001, Jiwei Wei, Yang Yang 0002 |
ACL (1) | 1 |
| 2024 | Separation Is for Better Reunion: Data Lake Storage at HuaweiabstractHuawei collaborates with some Chinese large busi-ness companies to store and process exabytes of nationwide operational data in data lake storage to provide business insights. Specifically, our customers will ask to store and process massive log message data to support their real-time and decision-making applications. Thus, we need computation and storage components in the analytic platform to process and store these data cost-efficiently. To meet these user requirements, we have designed a storage system in data lake, StreamLake, which introduces a novel design to serve log message streaming and batch data processing in distributed storage, with high scalability, efficiency, reliability and low cost. Specifically, we introduce a stream (storage) object as a storage abstraction for message streaming data to achieve the storage-disaggregated architecture with high scalability and reliability. Moreover, we utilize the erasure coding and tiered storage to save the storage cost, and furthermore, the stream object can be automatically converted to a table object such that cost-effective stream and batch data processing can be achieved. For tabular data, we implement the lakehouse functionality to support ACID via the table object, with a metadata acceleration to improve the efficiency of data access between the compute and storage engines. Also, we design a LakeBrain optimizer at the storage side to optimize the query performance and resource utilization under the storage-disaggregated architecture. Finally, we have also deployed StreamLake in China Mobile, the world's largest mobile network operator to serve over 20PB production data, and the results demonstrate improvements of 30% to 4x in terms of query performance and over 37% in terms of cost saving. Chengliang Chai, Haohai Ma, Zhenyong Fan, Jiaquan Zhang, Rui Zhang 0003, Duanshun Li, Keji Huang, Guangbin Meng, Yuefeng Zhou, Lirong Jian, Jiwu Shu, Ye Yuan 0001, Guoren Wang, Guoliang Li 0001 |
ICDE | 8 |
| 2022 | Self-adaptive Multi-scale Aggregation Network for Stereo MatchingabstractNowadays stereo matching architectures based on convolutional neural network has achieved remarkable performance. However the existing methods still lack the capability to find the correspondence in ill-posed regions. In this paper we present Self-adaptive Multi-scale Aggregation Network (SMA-Net) for stereo matching. First of all, we construct the cost volume through multi-channels group-wise correlation, which divided the features into groups with different number of channels to enhance the ability of measuring the similarities of features from stereo images. Secondly the self-adaptive cost aggregation is used to regularize the two scale cost volumes from different aggregation branches with intermediate supervise. We conduct comprehensive experiments on SceneFlow, KITTI2012, and KITTI2015 datasets. The competitive results prove that the approach in this paper outperforms many other stereo matching algorithms especially in ill-posed regions. Shuiqiang Ye, Jiaquan Zhang, Xin'an Wang, Qifei Dai, Zhengzhong Yu, Fuchi Li, Yong Zhao 0010 |
ICPR | 3 |
| 2021 | A Data-Driven Analysis of K-12 Students' Participation and Learning Performance on an Online Supplementary Learning PlatformabstractDue to the limitation of public school systems, many students pursue private supplementary tutoring for improving their academic performance. Different from public schools, the private online education provides diverse courses and satisfy differentiated demands of the students. Students’ behavior and performance in online supplementary learning are relevant to not only personal attributes, but also some factors such as city levels, grades and family situation. Existing studies mostly rely on panel survey/questionnaire data and few studied online private tutoring. In this paper, with 11,392 anonymous K-12 students’ 3-year learning data from one of the world’s largest online extra-curricular education platforms, we uncover students’ online learning behaviors and infer the impact of students’ home location, family socioeconomic situation and attended school’s reputation/rank on the students’ private tutoring course participation and learning outcomes. Further analysis suggests that such impact may be largely attributed to the inequality of access to educational resources in different cities and the inequality in family socioeconomic status. Finally, we study the predictability of students’ performance and behaviors using machine learning algorithms with different groups of features, showing students’ online learning performance can be predicted with MAE< 10%. Jiaquan Zhang, Xin Gao 0034, Hui Chen 0017, Jar-der Luo, Xiaoming Fu 0001 |
ICCCN | 1 |
| 2020 | Identifying unfamiliar callers' professions from privacy-preserving mobile phone dataabstractIdentifying an unfamiliar caller's profession is important to protect citizens' personal safety and property. Due to limited data protection of many popular online services in some countries such as taxi hailing or takeouts ordering, many users encounter an increasing number of phone calls from strangers. This may aggravate the situation that criminals pretend to be delivery staff or taxi drivers, bringing threats to the society. Additionally, many people nowadays suffer from excessive digital marketing and fraud phone calls because of personal information leakage. However, previous works on malicious call detection only focused on binary classification, and do not work for identification of multiple professions. We observed that web service requests issued from users' mobile phones which may show their Apps preferences, spatial and temporal patterns, and other profession related information. This offers us a hint to identify unfamiliar callers. In fact, some previous works already leveraged raw data from mobile phones (which includes sensitive information) for personality studies. However, accessing users' mobile phone raw data may violate the more and more strict private data protection policies or regulations (e.g. GDPR 71). Using appropriate statistical methods to eliminate private information and preserve personal characteristics, provides a way to identify mobile phone callers without privacy concern. In this paper, we exploit privacy-preserving mobile data to develop a model which can automatically identify the callers who are divided into four categories of users: normal users (other professions), taxi drivers, delivery and takeouts staffs, telemarketers and fraudsters. The validation results over an anonymized dataset of 1,282 users with a period of 3 months in Shanghai City prove that the proposed model could achieve an accuracy of 75+%. Jiaquan Zhang, Xiaoming Yao, Xiaoming Fu 0001 |
MSN | 1 |