EDBT 2026 Demo / reviewers in the wild / expert
Danyang Zhu
dblp:203/3368
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M100: An Orchestrated Dataflow Architecture Powering General AI Computing
Changkui Mao, Changsong Wu, Chao Suo, Danyang Zhu, Hengchang Xiong, Hongzhan Lu, Hongzhen Liu, Junfeng Tang, Lipeng Ge, Shaodong Yang, Shibin Tang, Xiaobo Du, Yihua Jin, Zhiyuan Man, Zhongxiao Yao |
ISCA | 8 |
| 2026 | ICL: In-loop continual learning framework for language model pre-training for E-commerceabstractPre-trained language models have become a critical natural language processing component in many E-commerce applications. As businesses continue to evolve, the pre-trained models should be able to adopt new domain knowledge and new tasks. This paper proposes a novel sequential multi-task pre-trained language framework, ICL-BERT (In-loop Continual Learning BERT), which enables evolving the current model with new knowledge and new tasks. The contributions of ICL-BERT are (1) vocabularies and entities are optimized on E-commerce corpus; (2) a new glyph embedding is introduced to learn glyph information for vocabularies and entities; (3) specific and general tasks are designed to encode E-commerce knowledge for pre-training ICL-BERT; and (4) a new task-gating mechanism, called ICL (In-loop continual Learning), is proposed for sequential multi-task learning, which evolves the current model effectively and efficiently. Our evaluation results demonstrate that ICL-BERT outperforms existing models in both CLUE and e-commerce tasks, with an average accuracy improvement of 1.73% and 3.5%, respectively. Furthermore, ICL-BERT serves as a fundamental pre-trained language model that runs online in JingDong’s daily business. Chiman Wong, Sanpeng Wang, Danyang Zhu, Chi-Man Vong |
Intell. Data Anal. | 5 |
| 2025 | An Efficient FPGA Implementation of Approximate Nearest Neighbor SearchabstractApproximate nearest neighbor search (ANNS) plays an important role in modern artificial intelligence (AI) systems, being extensively utilized in search engines, advertising, and recommendation systems. With the advent of large language models (LLMs), ANNS is increasingly finding applications in edge scenarios such as personal assistants. The demand for efficient and fast ANNS solutions is, therefore, more pressing than ever. In this article, we propose a scalable and efficient field-programmable gate array (FPGA) implementation of ANNS based on the inverted file with product quantization (IVF-PQ) algorithm, thus marking the first hardware implementation supporting up to 1024-D datasets. First, we devise a novel architecture for the Top-Kmodule, capable of processing multiple input data streams simultaneously and linearly increasing throughput. Second, we adjust the data precision in several parts of our design, thus achieving obvious performance improvement without losing much recall. Moreover, we introduce a flexible distance calculation (Distance Cal) module that can be reused for various computational tasks at different query stages. We code our design in Verilog and implement it on Xilinx Alveo U280. The experimental results show that our search latency can be as low as 0.0071 ms at a 94% recall, while the power is 19.80 W. Compared to the state-of-the-art application-specified integrated circuit (ASIC) implementations, our design delivers a$4.5\times $speedup in latency and a 20% reduction in energy consumption. Yifeng Song, Chenjie Liu, Danyang Zhu, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Low-latency Hardware Architecture for VDF Evaluation in Class GroupsabstractThe verifiable delay function (VDF), as a kind of cryptographic primitives, has recently been adopted quite often in decentralized systems. Highly correlated to the security of VDFs, the fastest implementation for VDF evaluation is generally desired to be publicly known. In this paper, for the first time, we propose a low-latency hardware implementation for the complete VDF evaluation in the class group by jointly exploiting optimizations. On one side, we reduce the required computational cycles by decreasing the hardware-unfriendly divisions and increase the parallelism of computations by reducing the data dependency. On the other side, we provide low-latency large-number divisors, multipliers, and adders, respectively, while those operators are generally very hard to be accelerated. Besides, we carefully schedule the sub-modules and devise the low-latency architecture for the complete VDF evaluation. Finally, the proposed design is coded and synthesized under the TSMC 28-nm CMOS technology. The experimental results show that our design can achieve a speedup of 3.5x compared to the optimal C++ implementation for the VDF evaluation over an advanced CPU. Moreover, compared to the state-of-the-art hardware implementation for the squaring, a key step of VDF, we achieve about 2x speedup. Danyang Zhu, Jing Tian 0004, Minghao Li 0001, Zhongfeng Wang 0001 |
IEEE Trans. Computers | 1 |
| 2022 | Location privacy preservation through kernel transformationabstractSummary The frequent data leak scandals of recent years indicate that service providers who hold personal data may not be reliable as they claim. We assert that sensitive user information must be sanitized locally before it is sent to service providers if it is to be protected. The LPPK privacy‐preserving framework presented in this article is a local sanitization scheme, for location‐based services (LBSs). It applies a fog‐computing structure in which a private map is generated by the LBS server with kernel transformation for each user. A fog device then provides location services for each user according to the private map. Without colluding, neither the LBS server nor the fog device can deduce a user's real location. Experiments conducted on real‐world data sets demonstrate that LPPK delivers sufficient query accuracy at a level significantly higher than existing approaches while preserving location privacy. Lefeng Zhang, Guanghua Song, Danyang Zhu, Wei Ren 0002, Ping Xiong 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Low-Latency Architecture for the Parallel Extended GCD Algorithm of Large NumbersabstractThe extended Greatest Common Divisor (GCD) is an extension of the GCD operation, which computes not only the GCD of integers a and b but also the Bezout's coefficients that are integers x and y such that ax + by =3D GCD(a,b). Recently, the large-number extended GCD algorithm is used in the core function of the next-generation blockchain systems and served as the most time-consuming operation. Considering the efficiency, speeding up this operation is urgently desired. However, the extended GCD, which is rarely explored in literature, is extremely hard to parallelize because of long serial operations with strong data dependency. In this paper, we propose a low- latency architecture for the extended GCD of large numbers by utilizing many algorithmic transformations and architectural optimizations. Firstly, a parallel extended GCD algorithm is well studied and modified to be practical in hardware. Secondly, a high-parallel architecture is designed for the selected extended GCD, where the trade-off is well evaluated between computation latency and power consumption. Finally, the architecture is coded using Verilog language and synthesized under the TSMC 28- nm CMOS technology. The experimental results for the 1024-bit extended GCD show that our design significantly outperforms the prior arts. Danyang Zhu, Jing Tian 0004, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2019 | Optimizing rewards allocation for privacy-preserving spatial crowdsourcing
Ping Xiong 0001, Danyang Zhu, Lefeng Zhang, Wei Ren 0002, Tianqing Zhu |
Comput. Commun. | 2 |
| 2017 | Mobility-aware multimedia data transfer using Multipath TCP in Vehicular NetworkabstractThis paper proposed a mobility-aware multimedia data transfer mechanism using Multipath TCP in Vehicular Network. Since high transmission rate and low latency are the two key factors for multimedia data transmission that could provide stable video streaming services, therefore, we first adopted Multipath Transport Control Protocol which can transfer data concurrently for improving transmission rate and designed Quality-aware Data Distribution to dynamically allocate the data to different subflows. Moreover, in Vehicular Network, the mobile terminal, which means the vehicle, can communicate with remote server through roadside unit (RSU). However, the communication link between terminals and remote server will disrupt while the vehicle exceeding the communication range of roadside unit; and there also exists the situation that the vehicle is in the communication range of several RSU. Accordingly, we exploited a mobility-aware distance measurement for checking whether the vehicle has moved out of the communication range of any RSU or it is in multi-RSU's communication range. Afterwards, we designed a handover mechanism which transfer data to connected path (4G) for stable transmission using MPTCP when the vehicle exceeded the communication range of RSU; while one mobile can communicate with more than one RSU, we exploited a mechanism which can trigger new path for multipath data transmission and non-corporation Nash Equilibria was employed for solving the problem of fairness in the same kind of network technology multipath transmission. Simulation results show how mobility-aware multimedia data transfer mechanism improve the performance of transmission comparison with state-of-art solution. Danyang Zhu, Changqiao Xu, Jiuren Qin, Zan Zhou 0001, Jianfeng Guan |
IWCMC | 1 |