Zifeng Zhao

dblp:233/8096 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MoEA: A Mixture of Experts Accelerator with Direct Token Access and Dynamic Expert Scheduling
abstract
Transformer-based large language models (LLMs) have achieved widespread adoption, but their growing model sizes impose substantial computational costs. Mixture-of-Experts (MoE) mechanism alleviates this by sparsely activating a subset of experts for each token. However, it still faces two critical challenges: (1) token rearrangement incurs non-trivial overhead of off-chip data movement, and (2) static execution flow fails to exploit inter-expert token reuse. In this paper, we present MoEA, a specialized accelerator designed to address the inefficiencies in MoE inference. MoEA introduces two key innovations: a Direct Token Access mechanism that leverages a hardware-managed metadata queue to eliminate off-chip token rearrangement, and a Dynamic Expert Scheduler that captures inter-expert token reuse patterns and optimizes expert execution order to maximize token reuse. Evaluations on representative MoE benchmarks show that MoEA reduces off-chip memory access by 12.06% and achieves $259.24 \times, 9.63 \times$ and $1.16 \times$ speedups over CPU, GPU and EdgeMoE accelerator, respectively.
Zifeng Zhao, Jiewen Zheng, Tianxing Xie, Xinghao Zhu, Gengsheng Chen
ASP-DAC1
2026 FlexEdge: A Hardware-Software Co-Design for Flexible Edge Transformer Inference
Jiewen Zheng, Zifeng Zhao, Gengsheng Chen, Wenbo Yin
ISCAS2
2025 RUNoC: Re-inject into the Underground Network to Alleviate Congestion in Large-Scale NoC
abstract
In modem high-performance systems, the demand for large-scale Networks-on-Chip (NoC) has grown rapidly. As NoCs are scaled up, the issue of network congestion becomes increasingly critical and complex to address. Existing researches utilize partition-based NoC and Two-Level Network (TLN) techniques to alleviate congestion in large-scale NoCs. However, most of these solutions have limitations concerning universality, hardware overhead and workload balancing due to their inherent complexity and design constraints. In this article, we propose RUNoC, a new partition-based TLN architecture consisting of a Main Network for normal transmission and a sparse Underground Network enabling fast transmission. A special hardware unit is designed and integrated into each Main Network Router to decide when a packet needs to be routed to Underground Network based on network congestion information and the distance to the packet's destination, thereby alleviating congestion in Main Network and ensuring a subtle load balance between the two networks. Additionally, RUNoC is further enhanced by Shared Row Buffers (SRBs) and specialized network interface to guarantee deadlock and livelock freedom. Evaluation results indicate that RUNoC improves up to 60% in performance compared to XY routing scheme and up to 34% in performance-area ratio compared to existing researches.
Xinghao Zhu, Jiyuan Bai, Zifeng Zhao, Qirong Yu, Gengsheng Chen
ASP-DAC3
2025 Transfer Learning for Nonparametric Contextual Dynamic Pricing
abstract
Dynamic pricing strategies are crucial for firms to maximize revenue by adjusting prices based on market conditions and customer characteristics. However, designing optimal pricing strategies becomes challenging when historical data are limited, as is often the case when launching new products or entering new markets. One promising approach to overcome this limitation is to leverage information from related products or markets to inform the focal pricing decisions. In this paper, we explore transfer learning for nonparametric contextual dynamic pricing under a covariate shift model, where the marginal distributions of covariates differ between source and target domains while the reward functions remain the same. We propose a novel Transfer Learning for Dynamic Pricing (TLDP) algorithm that can effectively leverage pre-collected data from a source domain to enhance pricing decisions in the target domain. The regret upper bound of TLDP is established under a simple Lipschitz condition on the reward function. To establish the optimality of TLDP, we further derive a matching minimax lower bound, which includes the target-only scenario as a special case and is presented for the first time in the literature. Extensive numerical experiments validate our approach, demonstrating its superiority over existing methods and highlighting its practical utility in real-world applications.
Feiyu Jiang, Zifeng Zhao, Yi Yu 0016
ICML3
2025 Learning Personalized Ad Impact via Contextual Reinforcement Learning under Delayed Rewards
abstract
Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed and long-term effects, cumulative ad impacts such as reinforcement or fatigue, and customer heterogeneity. However, these effects are often not jointly addressed in previous studies. To capture these factors, we model ad bidding as a Contextual Markov Decision Process (CMDP) with delayed Poisson rewards. For efficient estimation, we propose a two-stage maximum likelihood estimator combined with data-splitting strategies, ensuring controlled estimation error based on the first-stage estimator's (in)accuracy. Building on this, we design a reinforcement learning algorithm to derive efficient personalized bidding strategies. This approach achieves a near-optimal regret bound of $\tilde{\mathcal{O}}(dH^2\sqrt{T})$, where $d$ is the contextual dimension, $H$ is the number of rounds, and $T$ is the number of customers. Our theoretical findings are validated by simulation experiments.
Yuwei Cheng, Zifeng Zhao
NeurIPS2
2024 PAIR: Periodically Alternate the Identity of Routers to Ensure Deadlock Freedom in NoC
abstract
Network-on-Chip (NoC) has become widely adopted in multi/many-core systems for on-chip communication. Avoiding deadlock is a critical issue in NoC design. Recently proposed techniques partition the network resources into one or more express paths. Each path allows a blocked packet to move forward to its destination, thereby breaking deadlocks. However, as the scaling up of NoC, the existing methods are experiencing a decline in efficiency due to the coarse-grained partitioning strategy. In this work, we present PAIR, a novel scheme that guarantees deadlock freedom. PAIR adopts a fine-grained resource partitioning approach, significantly increasing the number of express paths available. The express paths in PAIR not only allow blocked packets but also permit non-blocked packets to be forwarded to the next hop, minimizing the performance degradation for the normal flow control. We implemented and evaluated PAIR on the mesh network using classic synthetic traffic patterns. Our experiments show that PAIR significantly improves throughput performance, achieving an average increase of 80.36% compared with state-of-the-art deadlock-free solutions. Furthermore, PAIR outperforms the latest deadlock-free scheme by 37.78% in terms of throughput.
Zifeng Zhao, Xinghao Zhu, Jiyuan Bai, Gengsheng Chen
ASPDAC1
2022 Speaker-Aware Mixture of Mixtures Training for Weakly Supervised Speaker Extraction
abstract
Dominant researches adopt supervised training for speaker extraction, while the scarcity of ideally clean corpus and channel mismatch problem are rarely considered.To this end, we propose speaker-aware mixture of mixtures training (SAMoM), utilizing the consistency of speaker identity among target source, enrollment utterance and target estimate to weakly supervise the training of a deep speaker extractor.In SAMoM, the input is constructed by mixing up different speaker-aware mixtures (SAMs), each contains multiple speakers with their identities known and enrollment utterances available.Informed by enrollment utterances, target speech is extracted from the input one by one, such that the estimated targets can approximate the original SAMs after a remix in accordance with the identity consistency.Moreover, using SAMoM in a semi-supervised setting with a certain amount of clean sources enables application in noisy scenarios.Extensive experiments on Libri2Mix show that the proposed method achieves promising results without access to any clean sources (11.06dBSI-SDRi) 1 .With a domain adaptation, our approach even outperformed supervised framework in a cross-domain evaluation on AISHELL-1.
Zifeng Zhao, Rongzhi Gu, Dongchao Yang, Jinchuan Tian, Yuexian Zou
INTERSPEECH1
2022 Target Confusion in End-to-end Speaker Extraction: Analysis and Approaches
abstract
Recently, end-to-end speaker extraction has attracted increasing attention and shown promising results.However, its performance is often inferior to that of a blind source separation (BSS) counterpart with a similar network architecture, due to the auxiliary speaker encoder may sometimes generate ambiguous speaker embeddings.Such ambiguous guidance information may confuse the separation network and hence lead to wrong extraction results, which deteriorates the overall performance.We refer to this as the target confusion problem.In this paper, we conduct an analysis of such an issue and solve it in two stages.In the training phase, we propose to integrate metric learning methods to improve the distinguishability of embeddings produced by the speaker encoder.While for inference, a novel post-filtering strategy is designed to revise the wrong results.Specifically, we first identify these confusion samples by measuring the similarities between output estimates and enrollment utterances, after which the true target sources are recovered by a subtraction operation.Experiments show that performance improvement of more than 1dB SI-SDRi can be brought, which validates the effectiveness of our methods and emphasizes the impact of the target confusion problem 1 .
Zifeng Zhao, Dongchao Yang, Rongzhi Gu, Yuexian Zou
INTERSPEECH1
2022 Change-point Detection for Sparse and Dense Functional Data in General Dimensions
abstract
We study the problem of change-point detection and localisation for functional data sequentially observed on a general $d$-dimensional space, where we allow the functional curves to be either sparsely or densely sampled. Data of this form naturally arise in a wide range of applications such as biology, neuroscience, climatology and finance. To achieve such a task, we propose a kernel-based algorithm named functional seeded binary segmentation (FSBS). FSBS is computationally efficient, can handle discretely observed functional data, and is theoretically sound for heavy-tailed and temporally-dependent observations. Moreover, FSBS works for a general $d$-dimensional domain, which is the first in the literature of change-point estimation for functional data. We show the consistency of FSBS for multiple change-point estimation and further provide a sharp localisation error rate, which reveals an interesting phase transition phenomenon depending on the number of functional curves observed and the sampling frequency for each curve. Extensive numerical experiments illustrate the effectiveness of FSBS and its advantage over existing methods in the literature under various settings. A real data application is further conducted, where FSBS localises change-points of sea surface temperature patterns in the south Pacific attributed to El Ni\~{n}o.
Carlos Misael Madrid Padilla, Daren Wang, Zifeng Zhao, Yi Yu 0016
NeurIPS3
2022 Functional Linear Regression with Mixed Predictors
abstract
We study a functional linear regression model that deals with functional responses and allows for both functional covariates and high-dimensional vector covariates. The proposed model is flexible and nests several functional regression models in the literature as special cases. Based on the theory of reproducing kernel Hilbert spaces (RKHS), we propose a penalized least squares estimator that can accommodate functional variables observed on discrete sample points. Besides a conventional smoothness penalty, a group Lasso-type penalty is further imposed to induce sparsity in the high-dimensional vector predictors. We derive finite sample theoretical guarantees and show that the excess prediction risk of our estimator is minimax optimal. Furthermore, our analysis reveals an interesting phase transition phenomenon that the optimal excess risk is determined jointly by the smoothness and the sparsity of the functional regression coefficients. A novel efficient optimization algorithm based on iterative coordinate descent is devised to handle the smoothness and group penalties simultaneously. Simulation studies and real data applications illustrate the promising performance of the proposed approach compared to the state-of-the-art methods in the literature.
Daren Wang, Zifeng Zhao, Yi Yu 0016, Rebecca Willett
J. Mach. Learn. Res.2
2021 Knowledge Learning of Insurance Risks Using Dependence Models
abstract
Learning the customers’ experience and behavior creates competitive advantages for any company over its rivals. The insurance industry is an essential sector in any developed economy and a better understanding of customers’ risk profile is critical to decision making in all aspects of insurance operations. In this paper, we explore the idea of using copula-based dependence models to learn the hidden risk of policyholders in property insurance. Specifically, we build a novel copula model to accommodate the dependence over time and over space among spatially clustered property risks. To tackle the computational challenge caused by the discreteness feature of large-scale insurance data, we propose an efficient multilevel composite likelihood approach for parameter estimation. Provided that latent risk induces correlation, the proposed customer learning method offers improved predictive analytics by allowing insurers to borrow strength from related risks in predicting new risks and also helps reveal the relative importance of the multiple sources of unobserved heterogeneity in updating policyholders’ risk profile. In the empirical study, we examine the loss cost of a portfolio of entities insured by a government property insurance program in Wisconsin. We find both significant temporal and spatial association among property risks. However, their effects on the predictive distribution of loss cost are different for the new and renewal policyholders. The two sources of dependence are complements for the former and substitutes for the latter. These findings are shown to have substantial managerial implications in key insurance operations such as experience rating, capital allocation, and reinsurance arrangement.
Zifeng Zhao, Peng Shi 0003, Xiaoping Feng
INFORMS J. Comput.1
2021 Statistically and Computationally Efficient Change Point Localization in Regression Settings
abstract
Detecting when the underlying distribution changes for the observed time series is a fundamental problem arising in a broad spectrum of applications. In this paper, we study multiple change-point localization in the high-dimensional regression setting, which is particularly challenging as no direct observations of the parameter of interest is available. Specifically, we assume we observe $\{ x_t, y_t\}_{t=1}^n$ where $ \{ x_t\}_{t=1}^n $ are $p$-dimensional covariates, $\{y_t\}_{t=1}^n$ are the univariate responses satisfying $\mathbb{E}(y_t) = x_t^\top \beta_t^* \text{ for } 1\le t \le n $ and $\{\beta_t^*\}_{t=1}^n $ are the unobserved regression coefficients that change over time in a piecewise constant manner. We propose a novel projection-based algorithm, Variance Projected Wild Binary Segmentation~(VPWBS), which transforms the original (difficult) problem of change-point detection in $p$-dimensional regression to a simpler problem of change-point detection in mean of a one-dimensional time series. VPWBS is shown to achieve sharp localization rate $O_p(1/n)$ up to a log factor, a significant improvement from the best rate $O_p(1/\sqrt{n})$ known in the existing literature for multiple change-point localization in high-dimensional regression. Extensive numerical experiments are conducted to demonstrate the robust and favorable performance of VPWBS over two state-of-the-art algorithms, especially when the size of change in the regression coefficients $\{\beta_t^*\}_{t=1}^n $ is small.
Daren Wang, Zifeng Zhao, Kevin Z. Lin, Rebecca Willett
J. Mach. Learn. Res.2
2017 Reduction of common-mode interference under different working modes of induction machine based on improved predictive torque control method
abstract
The common-mode voltage (CMV) could cause severe hazards to the induction machine winding insulation and bearings, reducing the machine lifetime. This paper proposes an improved predictive torque control (PTC) method with CMV reduction strategy which cost little and much smaller than traditional filter method. A simulation based on SIMULINK is analyzed which verify the effectiveness of the proposed method. This paper discusses CMV in start mode, reversal mode and steady modes of motor. By analyzing CMV in different modes, we can find that which mode contributes more to the CMV and the specific CMV reduction effect in each modes, which can guide the operating process of the motor.
Zifeng Zhao, Liyu Dai
IECON3