Junmei Yang

dblp:157/9330 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Computer networks · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Diffinformer: Diffusion informer model for long sequence time-series forecasting
Wei Chen 0165, Yican Liu, Junmei Yang, Zhiheng Zhou 0001, Delu Zeng
Expert Syst. Appl.4
2025 Bayesian Gaussian Process ODEs via Double Normalizing Flows
abstract
Gaussian processes have been used to model the vector field of continuous dynamical systems, which are characterized by a probabilistic ordinary differential equation (GP-ODE). Bayesian inference for these models has been extensively studied and applied in tasks such as time series prediction. However, the use of standard GPs with basic kernels like squared exponential kernels has been common in GP-ODE research, limiting the model’s ability to represent complex scenarios. To address this limitation, we introduce normalizing flows to reparameterize the ODE vector field, resulting in a data-driven prior distribution, thereby increasing flexibility and expressive power. We develop a variational inference algorithm that utilizes analytically tractable probability density functions of normalizing flows. Additionally, we also apply normalizing flows to the posterior inference of GP-ODEs to resolve the issue of strong mean-field assumptions. By applying normalizing flows in these ways, our model improves accuracy and uncertainty estimates for Bayesian GP-ODEs. We validate the effectiveness of our approach on simulated dynamical systems and real-world human motion data, including time series prediction and missing data recovery tasks.
Jian Xu 0021, Shian Du, Junmei Yang, Xinghao Ding, Delu Zeng, John W. Paisley
AISTATS3
2025 Dequantified Diffusion-Schrödinger Bridge for Density Ratio Estimation
abstract
Density ratio estimation is fundamental to tasks involving f-divergences, yet existing methods often fail under significantly different distributions or inadequately overlapping supports --- the density-chasm and the support-chasm problems. Additionally, prior approaches yield divergent time scores near boundaries, leading to instability. We design $\textbf{D}^3\textbf{RE}$, a unified framework for robust, stable and efficient density ratio estimation. We propose the dequantified diffusion bridge interpolant (DDBI), which expands support coverage and stabilizes time scores via diffusion bridges and Gaussian dequantization. Building on DDBI, the proposed dequantified Schr{\"o}dinger bridge interpolant (DSBI) incorporates optimal transport to solve the Schr{\"o}dinger bridge problem, enhancing accuracy and efficiency. Our method offers uniform approximation and bounded time scores in theory, and outperforms baselines empirically in mutual information and density estimation tasks.
Wei Chen 0165, Shigui Li, Junmei Yang, John W. Paisley, Delu Zeng
ICML4
2025 Variational Learning of Gaussian Process Latent Variable Models through Stochastic Gradient Annealed Importance Sampling
abstract
Gaussian Process Latent Variable Models (GPLVMs) have become increasingly popular for unsupervised tasks such as dimensionality reduction and missing data recovery due to their flexibility and non-linear nature. An importance-weighted version of the Bayesian GPLVMs has been proposed to obtain a tighter variational bound. However, this version of the approach is primarily limited to analyzing simple data structures, as the generation of an effective proposal distribution can become quite challenging in high-dimensional spaces or with complex data sets. In this work, we propose VAIS-GPLVM, a variational Annealed Importance Sampling method that leverages time-inhomogeneous unadjusted Langevin dynamics to construct the variational posterior. By transforming the posterior into a sequence of intermediate distributions using annealing, we combine the strengths of Sequential Monte Carlo samplers and VI to explore a wider range of posterior distributions and gradually approach the target distribution. We further propose an efficient algorithm by reparameterizing all variables in the evidence lower bound (ELBO). Experimental results on both toy and image datasets demonstrate that our method outperforms state-of-the-art methods in terms of tighter variational bounds, higher log-likelihoods, and more robust convergence.
Jian Xu 0021, Shian Du, Junmei Yang, Qianli Ma 0001, Delu Zeng, John W. Paisley
UAI3
2025 Fully Bayesian differential Gaussian processes through stochastic differential equations
Jian Xu 0021, Junmei Yang, Delu Zeng, John W. Paisley
Knowl. Based Syst.4
2025 Neural Operator Variational Inference Based on Regularized Stein Discrepancy for Deep Gaussian Processes
abstract
Deep Gaussian process (DGP) models offer a powerful nonparametric approach for Bayesian inference, but exact inference is typically intractable, motivating the use of various approximations. However, existing approaches, such as mean-field Gaussian assumptions, limit the expressiveness and efficacy of DGP models, while stochastic approximation can be computationally expensive. To tackle these challenges, we introduce neural operator variational inference (NOVI) for DGPs. NOVI uses a neural generator to obtain a sampler and minimizes the regularized Stein discrepancy (RSD) between the generated distribution and true posterior in $\mathcal {L}_{2}$ space. We solve the minimax problem using Monte Carlo estimation and subsampling stochastic optimization techniques and demonstrate that the bias introduced by our method can be controlled by multiplying the Fisher divergence with a constant, which leads to robust error control and ensures the stability and precision of the algorithm. Our experiments on datasets ranging from hundreds to millions demonstrate the effectiveness and the faster convergence rate of the proposed method. We achieve a classification accuracy of 93.56 on the CIFAR10 dataset, outperforming state-of-the-art (SOTA) Gaussian process (GP) methods. We are optimistic that NOVI possesses the potential to enhance the performance of deep Bayesian nonparametric models and could have significant implications for various practical applications.
Jian Xu 0021, Shian Du, Junmei Yang, Qianli Ma 0001, Delu Zeng
IEEE Trans. Neural Networks Learn. Syst.3
2025 Independently Reconfigurable Multiband All-Digital Transmitter Using SMASH Delta-Sigma Modulation
abstract
A multiband reconfigurable all-digital transmitter (ADT) employing a sturdy multistage noise-shaping (SMASH) delta-sigma modulation (DSM) is proposed in this article. TheN-stage SMASH DSM is presented digitally using unit signal transfer functions across cascaded stages, which supports arbitrary-stage extension for multiband transmission. Each stage is associated with a specific passband. By activating different stage numbers and adjusting only two parameters per stage, the numbers and frequencies of passbands are reconfigured simply and independently. Consequently, computational resources are significantly reduced, which is especially beneficial for four or more bands. Furthermore, the DSM outputs are directly extracted from individual stages to simplify the ADT architecture while the high efficiencies of subsequent power amplifiers (PAs) are maintained. By using the SMASH DSM, an ADT with multiband reconfigurable characteristics is proposed. With a resource-performance tradeoff, an ADT prototype based on a 2–2 SMASH DSM is implemented on a field programmable gate array (FPGA), which uses only 256 DSP48 slices. The transmissions are switched between single-band and dual-band configurations with carrier frequencies (CFs) independently tunable over 0–3.2 GHz. Measured adjacent channel leakage ratios (ACLRs) for 10- and 20-MHz dual-band signals are below −45 and −39.5 dBc, respectively, with error vector magnitudes (EVMs) below 1.6%. These results verify the agile multiband reconfiguration and excellent performance of the proposed SMASH DSM-based ADT.
Jie-Kai Xu, Yuan Chun Li, Wei-Heng Xue, Yun-Rui Feng, Junmei Yang, Quan Xue
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Superclass-aware visual feature disentangling for generalized zero-shot learning
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junmei Yang
Expert Syst. Appl.4
2024 Consistent representation joint adaptive adjustment for incremental zero-shot learning
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junmei Yang
Neurocomputing4
2024 Neural Ordinary Differential Equation Networks for Fintech Applications Using Internet of Things
abstract
The Internet-of-Things (IoT) technology is becoming increasingly pivotal in the financial services sector, with a growing number of algorithms being employed in high-frequency trading. High-frequency prediction in financial time series prediction presents a promising avenue of research. From convolutional neural networks to recurrent neural networks, deep learning have demonstrated exceptional capabilities in capturing the nonlinear characteristics of stock markets, thereby achieving high performance in stock index prediction. In this paper, we employ ODE-LSTM model for high-frequency price forecasting, predicting stock price data across various time scales, including 1-minute, 5-minutes, and 30-minutes frequencies. This approach introduces a novel concept, wherein the LSTM (Long Short-Term Memory) model is integrated with Neural ODE (Ordinary Differential Equations) to manage the hidden state and augment model interpretability. Over the course of 7 months, we achieved a 41.79% excess return on a simulated trading platform, with a daily average excess return of 0.30%, showcasing the commendable performance of our model and strategy.
Wei Chen 0165, Yican Liu, Junmei Yang, Delu Zeng, Zhiheng Zhou 0001
IEEE Internet Things J.4
2024 Generalized zero-shot action recognition through reservation-based gate and semantic-enhanced contrastive learning
Junyuan Shang, Chang Niu, Xiyuan Tao, Zhiheng Zhou 0001, Junmei Yang
Knowl. Based Syst.5
2024 DeepAR-Attention probabilistic prediction for stock price series
Wei Chen 0165, Zhiheng Zhou 0001, Junmei Yang, Delu Zeng
Neural Comput. Appl.4
2023 Taper Residual Dense Network for Audio Super-Resolution
Junmei Yang, Haosen Lin, Yiming Peng
ICANN (10)1
2022 Few-shot domain adaptation through compensation-guided progressive alignment and bias reduction
Junyuan Shang, Chang Niu, Junchu Huang, Zhiheng Zhou 0001, Junmei Yang
Appl. Intell.5
2022 Unbiased feature generating for generalized zero-shot learning
Chang Niu, Junyuan Shang, Junchu Huang, Junmei Yang, Yuting Song, Zhiheng Zhou 0001, Guoxu Zhou
J. Vis. Commun. Image Represent.4
2017 Joint Detection and Decoding for Polar Coded MIMO Systems
abstract
Generally, separate detection and decoding (SDD) scheme is usually adopted by multiple-input and multiple-output (MIMO) systems. In this paper, a novel approach which combines detection and decoding jointly using K-best detection and polar codes is proposed for the first time. Since the generation matrix of polar codes is triangular, polar codes could well adapt to the structure of K-best searching tree. Moreover, the property of polarization could reduce the latency of the proposed joint detection and decoding (JDD) scheme. Based on the joint optimization, the system model is given. For successive cancellation list (SCL) polar decoding, numerical results show that the performance of the proposed JDD is superior to the state-of-the-art SDD. At the frame error rate (FER) of 10-4, JDD outperforms SDD by approximately 2.5 dB for (256,128) polar coded 4×4 16-QAM MIMO system. Furthermore, for half rate polar codes, the proposed JDD could reduce 50% complexity compared to SDD. Results indicate that JDD shows superiorities in both performance and complexity. In addition, the corresponding hardware architectures are also given to demonstrate JDD's advantages and implementation feasibilities.
Yifei Shen 0003, Junmei Yang, Xiaohu You 0001, Chuan Zhang 0001
GLOBECOM2
2017 Algorithm and architecture for joint detection and decoding for MIMO with LDPC codes
abstract
With better spectral efficiency, multiple-input and multiple-output (MIMO) systems have drawn increasing attentions. Due to its near-optimal performance, K-best algorithm has been widely adopted for MIMO detection. To the best knowledge of the authors, this paper first proposes a joint detection and decoding (JDD) method for MIMO with low-density parity-check (LDPC) codes. By pruning the searching tree of K-best detection with LDPC coding constraint, the proposed JDD scheme benefits from both reduced tree-search complexity and improved performance compared to its uncoded MIMO counterpart. Numerical results of 16-QAM MIMO with (8, 2) LDPC code and 64-QAM MIMO with (18, 6) LDPC code have shown that, the proposed JDD scheme's performance is evidently superior over separated detection and decoding (SDD) scheme. More specifically, for the latter case with 12 antennas, JDD shows nearly 10 dB performance improvement than SDD when BER = 10-3. Hardware architecture and complexity analysis are also given in this paper to demonstrate JDD's advantages.
Shusen Jing, Junmei Yang, Zhongfeng Wang 0001, Xiaohu You 0001, Chuan Zhang 0001
ISCAS2
2016 Hardware Efficient and Low-Latency CA-SCL Decoder Based on Distributed Sorting
abstract
For polar codes, cyclic redundancy check (CRC)aided successive cancellation list (CA-SCL) decoder has attracted increasing attention from both academia and industry. In this paper, a hardware efficient and low-latency CA-SCL polar decoder based on distributed sorting is first proposed. For path metric (PM) sorting of each level, a distributed sorting (DS) algorithm is proposed to reduce the comparison complexity from (L2) to (L) (L denotes list size), together with the latency from kL2to kL (k is a coefficient independent of L). Employing folding technique, the N-bit folding polar decoder can be implemented based on the basic √N-bit polar decoder. In addition, pipelining technique is employed to refine the timing issue resulting from folding. The CRC is performed for 2L candidate paths serially to reduce hardware cost. According to demo of (1024, 512) code on Altera Stratix V FPGA, the proposed CA-SCL decoders with L = 2 and adjustable L = 2, 4 consume 9% and 50% board resources, respectively. Decoding latencies (in terms of clock cycles) are 2, 528 and 4, 064, respectively. For L = 2 and 4, we can achieve the frame error rate (FER) of 10-2at the signal noise ratio (SNR) of 2.36 dB and 2.06 dB, respectively. Compared with the floating point results, the performance degradation is negligible. Thus, the proposed design is suitable and adjustable for different real-life scenarios.
Xiao Liang 0005, Junmei Yang, Chuan Zhang 0001, Wenqing Song, Xiaohu You 0001
GLOBECOM2
2016 Efficient stochastic detector for large-scale MIMO
abstract
In this paper, a low-complexity stochastic belief propagation (BP) detector for large-scale MIMO is first proposed. Its efficient hardware architecture, with parallel pipeline, is presented in detail. Thanks to the stochastic approach, all arithmetic operations of the detector are implemented with simple logic structures. Several approaches which can potentially improve the detection performance are exploited. Simulation results have demonstrated that the stochastic BP detector can achieve similar detection performance compared with deterministic one for 32 × 32 MIMO system with 4-quadrature amplitude modulation (4-QAM). With the increase of antenna number, the detection performance improves at the linear expense of complexity and latency. Therefore, the proposed stochastic BP detector is suitable for large-scale MIMO system applications with good balance of detection performance and implementation complexity.
Junmei Yang, Chuan Zhang 0001, Shugong Xu, Xiaohu You 0001
ICASSP1
2016 Joint detection and decoding for MIMO systems with polar codes
abstract
As well known, the near-optimal K-best detection is popular in multiple-input and multiple-output (MIMO) systems. In this paper, we first propose the joint approaches of K-best detection and polar decoding. For joint detection and decoding (JDD) approach, both hard and soft decisions are considered. The simplified successive cancellation (SSC) decoding is exploited for hard decision, and the successive cancellation list (SCL) decoding is used as soft decision. The system setup for JDD is als o introduced, in which the modulation points across several channels are considered together. Simulation results have demonstrated the performance advantage of the JDD algorithms over the separated ones. For 1/2-rate polar codes, JDD schemes show 50% complexity reduction compared to the separated ones. Furthermore, by employing SSC hard decoding, the JDD algorithm is promising for high-throughput and low-complexity application s.
Junmei Yang, Chuan Zhang 0001, Wenqing Song, Shugong Xu, Xiaohu You 0001
ISCAS1
2016 Pipelined belief propagation polar decoders
abstract
Due to its inherent higher parallelism over successive cancellation (SC) polar decoder, belief propagation (BP) polar decoder becomes more favorable for high throughput applications. However, most existing BP decoders suffer from low utilization. In this paper, a new updating scheme, in which both left-to-right and right-to-left messages are considered identical, is first proposed for memory reduction. By revealing the similarity between BP polar decoder and fast Fourier transform (FFT) processor, both feed-forward and feed-back pipelined BP polar decoders are proposed along with detailed processing schedules. Implementation results have shown that both proposed pipelined BP decoders achieve more than 99.8% arithmetic logical units (ALUTs) reduction and 3.40% registers & block memory reduction, together with more than 7.45% speed-up, compared to the conventional fully parallel one. The proposed design approaches can be generalized with folding technique to achieve the required balance between area and speed flexibly.
Junmei Yang, Chuan Zhang 0001, Huayi Zhou 0002, Xiaohu You 0001
ISCAS1
2015 Pipelined implementations of polar encoder and feed-back part for SC polar decoder
abstract
In this paper, we first reveal the similarity of polar encoder and fast Fourier transform (FFT) processor. Based on this, both feed-forward and feed-back pipelined implementations of polar encoder are proposed. It is pointed out that the feedback part of SC polar decoder is nothing but a simplified version of polar encoder and therefore can be pipelined implemented also. Moreover, a general approach which uniformly constructs most pipelined polar encoders via folding transformation is proposed. Implementation results have shown that both proposed pipelined polar encoder architectures achieve more than 98.3% complexity reduction and more than 9.86% speed-up compared to the conventional implementation.
Chuan Zhang 0001, Junmei Yang, Xiaohu You 0001, Shugong Xu
ISCAS2