Aoyang Zhang

dblp:186/1165 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prefetching for Short Video Streaming: Experiences from a Longitudinal Evolution at Planetary Scale
abstract
Short-form video streaming is characterized by fast-paced, scrolling-driven user interactions. This poses unique challenges for its streaming algorithm design. To our knowledge, there is little understanding of how short video streaming algorithms perform in large-scale production platforms. To bridge this gap, this paper reports our two-year experience in evolving the prefetching algorithm, a critical algorithmic component for short video streaming, deployed in a leading global short video service. We adopt an iterative, production-driven approach, progressively evolving the design from simple heuristics to optimization-based and data-driven algorithms, with each iteration validated through large-scale A/B tests on hundreds of millions of users. Our evolution advances two core components: the prefetching logic and the viewing time estimation, including a lightweight on-device personalization mechanism. Through carefully balancing startup delay, mid-playback stalls, bandwidth usage, runtime overhead, and estimation accuracy, we achieve a 0.38% increase in user stay time - our key engagement metric - while simultaneously reducing bandwidth consumption by 14.7% throughout the evolution. We distill actionable insights from real-world deployment, highlighting the importance of startup latency, bandwidth efficiency, and low-overhead design for short video streaming at scale.
Yinjie Zhang, Yuming Hu, Aoyang Zhang, Zhixiang Luo, Zhendong Zhong, Haiqing Tao, Lan Xie, Shenglan Huang, Feng Qian 0001
SIGCOMM4
2025 Chameleon-SAT: An Adaptive Boolean Satisfiability Accelerator Using Mixed-Signal In-Memory Computing for Versatile SAT Problems
abstract
Boolean satisfiability (SAT), the first proven nondeterministic polynominal-complete problem, is crucial in dataintensive applications. Different applications have a wide spectrum of SAT problem sets (scale, complexity) and also various solution requirements (algorithm completeness, speed). Current SAT solvers are insufficient for providing ideal solutions under different scenarios. This work presents the Chameleon-SAT, the first ASIC-based SAT accelerator that can support local search, Davis-Putnam- Logemann-Lovel, Conflict-Driven Clause Learning algorithms, while leveraging the efficient mixed-signal inmemory computing architecture to achieve orders-of-magnitude improvements in speed compared to the prior SAT solvers. By judiciously selecting the reconfiguration mode, Chameleon-SAT is able to solve a wide range of the SAT problems to achieve smallscale, high-complexity cases ($\geq 90 \times$ for 20 variables/ 86 clauses, satisfiable problems), medium-scale, structured cases ($\geq 19 \times$ for 50 variables/ 215 clauses, unsatisfiable problems), and largescale, high-complexity cases ($\geq 7 \times$ for 100 variables/ 430 clauses, satisfiable problems).
Iris Ying Chou, Hao Kong 0003, Yi Huang 0036, Jianfeng Zhu 0001, Wenping Zhu, Shaojun Wei, Aoyang Zhang, Leibo Liu
DAC7
2025 EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
abstract
Fully Homomorphic Encryption (FHE) is a set of powerful cryptographic schemes that allows computation to be performed directly on encrypted data with an unlimited depth. Despite FHE’s promising in privacy-preserving computing, yet in most FHE schemes, ciphertext generally blows up thousands of times compared to the original message, and the massive amount of data load from off-chip memory for bootstrapping and privacy-preserving machine learning applications (such as HELR, ResNet-20), both degrade the performance of FHE-based computation. Several hardware designs have been proposed to address this issue, however, most of them require enormous resources and power. An acceleration platform with easy programmability, high efficiency, and low overhead is a prerequisite for practical application. This paper proposes EFFACT, a highly efficient full-stack FHE acceleration platform with a compiler that provides comprehensive optimizations and vector-friendly hardware. We start by examining the computational overhead across different real-world benchmarks to highlight the potential benefits of reallocating computing resources for efficiency enhancement. Then we make a design space exploration to find an optimal SRAM size with high utilization and low cost. On the other hand, EFFACT features a novel optimization named streaming memory access which is proposed to enable high throughput with limited SRAMs. Regarding the software-side optimization, we also propose a circuit-level function unit reuse scheme, to substantially reduce the computing resources without performance degradation. Moreover, we design novel NTT and automorphism units that are suitable for a cost-sensitive and highly efficient architecture, leading to low area. For generality, EFFACT is also equipped with an ISA and a compiler backend that can support several FHE schemes like CKKS, BGV, and BFV. We provide both FPGA and ASIC versions of EFFACT. On account of our full stack design, FPGA-EFFACT outperforms the SOTA FPGA accelerators in gmean by $1.22 \times$. Meanwhile, ASIC-EFFACT shows increased improvements in terms of the performance per chip area and the performance per Watt compared with the SOTA ASIC works.
Yi Huang 0036, Xinsheng Gong, Dibei Chen, Jianfeng Zhu 0001, Wenping Zhu, Liangwei Li, Mingyu Gao 0001, Shaojun Wei, Aoyang Zhang, Leibo Liu
HPCA10
2025 Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory
abstract
With the widespread use of large language models (LLMs), and with the privacy and cost concerns on cloud-based services, vendors are now pushing LLM inference to consumer devices. However, current attempts only enable real-time inference of low-quality small-sized LLMs. Large-sized LLMs have to load most of their weights from Flash storage for every execution iteration, which dominates the execution time of both the prefill and the generation phase. This performance bottleneck is attributed to both the low internal Flash memory bandwidth and the low transmission bandwidth between Flash and the Neural Processing Unit (NPU). To tackle these two challenges, we present Lincoln, a device-architecture co-design solution with LPDDR-interfaced, Compute-Enabled Flash Memory. On the device level, we boost the Flash internal bandwidth by improving upon existing array shrinking methods, to enable lower read latency and more parallel Flash planes within each Flash die. We specifically leverage 3D hybrid bonding, which is already adopted in consumer Flash products, to maintain high area efficiency and low density loss. On the architecture level, to leverage such increased internal bandwidth for resolving the transmission bottleneck, we propose two solutions for the two distinct phases of LLMs. For the compute-intensive prefill phase, we let Flash devices use the existing high-speed LPDDR interface (originally for DRAM), which offers much higher transmission bandwidth to the NPU than the conventional Flash interface, while maintaining good cost and area efficiency. For the memory-intensive generation phase, we rely on hybrid-bonding-based near-Flash computing to fully utilize the internal Flash bandwidth, and further equip with speculative decoding to eventually reach the real-time latency goal. Our evaluation shows that Lincoln enables real-time inference, with up to $13.23 \times$ and $254.1 \times$ speedups for LLM prefill and generation phases over conventional SSD-based systems.
Weiyi Sun, Mingyu Gao 0001, Zhaoshi Li, Aoyang Zhang, Iris Ying Chou, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu
HPCA4
2023 Counterfactual Video Recommendation for Duration Debiasing
abstract
Duration bias widely exists in video recommendations, where models tend to recommend short videos for the higher ratio of finish playing and thus possibly fail to capture users' true interests. In this paper, we eliminate the duration bias from both data and model. First, based on the extensive data analysis, we observe that play completion rate of videos with the same duration presents a bimodal distribution. Hence, we propose to perform threshold division to construct binary labels as training labels for alleviating the drawback of finish playing labels overly biased towards short videos. Algorithmically, we resort to causal inference, which enables us to inspect causal relationships of video recommendations with a causal graph. We identify that duration has two kinds of effect on prediction: direct and indirect. Duration bias lies in the direct effect, while the indirect effect benefits prediction. To this end, we design a model-agnostic Counterfactual Video Recommendation for Duration Debiasing (CVRDD) framework, which incorporates multi-task learning to estimate different causal effect during training. In the inference phase, we perform counterfactual inference to remove the direct effect of duration for unbiased prediction. We conduct experiments on two industrial datasets, and in addition to achieving highly promising results on traditional top-k recommendation metrics, CVRDD also improves the user watch time.
Shisong Tang, Qing Li 0006, Dingmin Wang, Ci Gao, Wentao Xiao, Dan Zhao 0003, Yong Jiang 0001, Aoyang Zhang
KDD9
2023 A Super-Resolution Flexible Video Coding Solution for Improving Live Streaming Quality
abstract
In the context of the latest growing popularity of live video streaming, ensuring high video quality has become one of the most significant challenges faced by all live streaming platforms. Insufficient uplink bandwidth is an important factor that influences these live video transmissions, affecting their bitrate and latency and consequently the associated video streaming quality. This paper proposes a novel flexible super-resolution-based video coding and uploading framework (FlexSRVC) that improves the quality of live video streaming in limited uplink network bandwidth conditions. FlexSRVC includes a flexible video coding scheme, which compresses high-resolution key and non-key video frames to a lower bitrate in order to reduce the upload delay. A new flexible bitrate adaptation algorithm is also proposed to select dynamically the number of frames to be compressed and the compression ratio by jointly considering uplink network conditions and available cloud computing resources. Trace-driven emulations demonstrate that FlexSRVC provides the same quality while reducing up to 25% of the required bandwidth compared to the original encoding method (H.264). FlexSRVC improves users' QoE by at least 50% compared to a super resolution-based method which employs reconstruction of all video frames in uplink bandwidth constrained conditions.
Qing Li 0006, Aoyang Zhang, Yong Jiang 0001, Longhao Zou, Zhimin Xu 0001, Gabriel-Miro Muntean
IEEE Trans. Multim.3
2022 Knowledge-based Temporal Fusion Network for Interpretable Online Video Popularity Prediction
abstract
Predicting the popularity of online videos has many real-world applications, such as recommendation, precise advertising, and edge caching strategies. Despite many efforts have been dedicated to the online video popularity prediction, there still exist several challenges: (1) The meta-data from online videos is usually sparse and noisy, which makes it difficult to learn a stable and robust representation. (2) The influence of content features and temporal features in different life cycles of online videos is dynamically changing, so it is necessary to build a model that can capture the dynamics. (3) Besides, there is a great need to interpret the predictive behavior of the model to assist administrators of video platforms in the subsequent decision-making.
Shisong Tang, Qing Li 0006, Xiaoteng Ma, Ci Gao, Dingmin Wang, Yong Jiang 0001, Aoyang Zhang, Hechang Chen
WWW8
2021 Higher quality live streaming under lower uplink bandwidth: an approach of super-resolution based video coding
abstract
With the growing popularity of live streaming, high video quality and low latency with limited uplink bandwidth have become a significant challenge. In this study, we propose Live Super-Resolution Based Video Coding (LiveSRVC), a novel video uploading framework that improves the quality of live streaming with low latency under limited uplink bandwidth. We design a new super-resolution-based key frame coding module to improve the coding compression efficiency. LiveSRVC dynamically selects the bitrate and the compression ratio of key frames, mitigating the influence of uplink bandwidth capacity on live streaming quality. Trace-driven emulations verify that LiveSRVC can provide the same quality while reducing up to 50% of the required bandwidth compared to the original encoding method (H.264). LiveSRVC consumes at least 10X less GPU occupation time compared to the method of reconstructing all frames with super-resolution.
Qing Li 0006, Aoyang Zhang, Longhao Zou, Yong Jiang 0001, Zhimin Xu 0001, Zhenhui Yuan
NOSSDAV3