Lingwei Zhu

dblp:231/4574 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AnomalyFilter: Selective Denoising Diffusion Model for Time Series Anomaly Detection
Kohei Obata, Zheng Chen 0012, Yasuko Matsubara, Lingwei Zhu, Yasushi Sakurai
PAKDD (1)4
2025 q-exponential family for policy optimization
abstract
Policy optimization methods benefit from a simple and tractable policy parametrization, usually the Gaussian for continuous action spaces. In this paper, we consider a broader policy family that remains tractable: the $q$-exponential family. This family of policies is flexible, allowing the specification of both heavy-tailed policies ($q>1$) and light-tailed policies ($q<1$). This paper examines the interplay between $q$-exponential policies for several actor-critic algorithms conducted on both online and offline problems. We find that heavy-tailed policies are more effective in general and can consistently improve on Gaussian. In particular, we find the Student's t-distribution to be more stable than the Gaussian across settings and that a heavy-tailed $q$-Gaussian for Tsallis Advantage Weighted Actor-Critic consistently performs well in offline benchmark problems. In summary, we find that the Student's t policy a strong candidate for drop-in replacement to the Gaussian. Our code is available at \url{https://github.com/lingweizhu/qexp}.
Lingwei Zhu, Haseeb Shah, Han Wang 0066, Yukie Nagai, Martha White
ICLR1
2025 Fat-to-Thin Policy Optimization: Offline Reinforcement Learning with Sparse Policies
abstract
Sparse continuous policies are distributions that can choose some actions at random yet keep strictly zero probability for the other actions, which are radically different from the Gaussian. They have important real-world implications, e.g. in modeling safety-critical tasks like medicine. The combination of offline reinforcement learning and sparse policies provides a novel paradigm that enables learning completely from logged datasets a safety-aware sparse policy. However, sparse policies can cause difficulty with the existing offline algorithms which require evaluating actions that fall outside of the current support. In this paper, we propose the first offline policy optimization algorithm that tackles this challenge: Fat-to-Thin Policy Optimization (FtTPO). Specifically, we maintain a fat (heavy-tailed) proposal policy that effectively learns from the dataset and injects knowledge to a thin (sparse) policy, which is responsible for interacting with the environment. We instantiate FtTPO with the general $q$-Gaussian family that encompasses both heavy-tailed and sparse policies and verify that it performs favorably in a safety-critical treatment simulation and the standard MuJoCo suite. Our code is available at https://github.com/lingweizhu/fat2thin.
Lingwei Zhu, Han Wang 0066, Yukie Nagai
ICLR1
2023 Drugs Resistance Analysis from Scarce Health Records via Multi-task Graph Representation
Honglin Shu, Pei Gao, Lingwei Zhu, Zheng Chen 0012, Yasuko Matsubara, Yasushi Sakurai
ADMA (3)3
2023 General Munchausen Reinforcement Learning with Tsallis Kullback-Leibler Divergence
abstract
Many policy optimization approaches in reinforcement learning incorporate a Kullback-Leilbler (KL) divergence to the previous policy, to prevent the policy from changing too quickly. This idea was initially proposed in a seminal paper on Conservative Policy Iteration, with approximations given by algorithms like TRPO and Munchausen Value Iteration (MVI). We continue this line of work by investigating a generalized KL divergence---called the Tsallis KL divergence. Tsallis KL defined by the $q$-logarithm is a strict generalization, as $q = 1$ corresponds to the standard KL divergence; $q > 1$ provides a range of new options. We characterize the types of policies learned under the Tsallis KL, and motivate when $q >1$ could be beneficial. To obtain a practical algorithm that incorporates Tsallis KL regularization, we extend MVI, which is one of the simplest approaches to incorporate KL regularization. We show that this generalized MVI($q$) obtains significant improvements over the standard MVI($q = 1$) across 35 Atari games.
Lingwei Zhu, Zheng Chen 0012, Matthew Schlegel, Martha White
NeurIPS1
2023 A Two-View EEG Representation for Brain Cognition by Composite Temporal-Spatial Contrastive Learning
abstract
Electroencephalography (EEG) is a major tool for studying neurophysiological processes. Investigating reliable representations from highly noisy measurements is a pending challenge, however, the medically treasured and insufficient labeled data have driven this process away from a supervised learning manner. Recent works have turned their attention to self-supervised learning (SSL), putting the contrastive strategy on capturing the spatio-temporal characteristics of the neuronal events of interest. We argue that the temporal-spatial view is not the best choice for the SSL contrastive objective because there is a missing piece of the EEG representation that is usually ignored: dynamic fluctuations in brain neurons and the statistical learning of analog/artificial neural networks cannot handle the dynamic characteristics well. This paper proposes a novel two-view contrastive learning framework to refine EEG features from local-global and past-future views. An array of spiking neural networks is embedded to project spatio-temporal features onto the spike sequences to represent the dynamic fluctuation information of EEG. Experimenting with sleep stage classification and prediction of lethal epileptic seizures, we verify the proposal competes favorably against the state-of-the-art methods and offers high-quality features, that is, supervised learning on top of them observes a significant improvement in classification after only one training iteration.
Zheng Chen 0012, Lingwei Zhu, Haohui Jia, Takashi Matsubara 0001
SDM2
2023 Cautious policy programming: exploiting KL regularization for monotonic policy improvement in reinforcement learning
abstract
Abstract In this paper, we propose cautious policy programming (CPP), a novel value-based reinforcement learning (RL) algorithm that exploits the idea of monotonic policy improvement during learning. Based on the nature of entropy-regularized RL, we derive a new entropy-regularization-aware lower bound of policy improvement that depends on the expected policy advantage function but not on state-action-space-wise maximization as in prior work. CPP leverages this lower bound as a criterion for adjusting the degree of a policy update for alleviating policy oscillation. Different from similar algorithms that are mostly theory-oriented, we also propose a novel interpolation scheme that makes CPP better scale in high dimensional control problems. We demonstrate that the proposed algorithm can trade off performance and stability in both didactic classic control problems and challenging high-dimensional Atari games.
Lingwei Zhu, Takamitsu Matsubara
Mach. Learn.1
2022 Hierarchical Categorical Generative Modeling for Multi-omics Cancer Subtyping
abstract
Identifying a specific cancer subtype from a variety of candidates is vital for precise and effective treatment. However, cancer subtyping is highly non-trivial as a result of cancer heterogeneity. While significant efforts have been put into understanding the mechanism of cancer subtypes via studying the omics data, existing methods run the risk of presenting biased analyses resulted from overfitting the high-dimensional and scarce omics data. In this paper, we propose a novel generative model that directly models the cancer data distribution by which downstream tasks can circumvent the curse of overfitting and achieve better performance. Unlike conventional generative modeling schemes, the proposed method underlines hierarchical categorical latent spaces to extract global features and local details respectively from transcriptomics and genomics profiles, which is the first to be considered in the cancer subtyping literature. By extensive experiments we verify that the proposed architecture achieves more clearly separated subtypes, as well as medically significant insights into real subtyping.
Ziwei Yang 0002, Lingwei Zhu, Chen Li 0027, Zheng Chen 0012, Naoki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM2
2022 Multi-Tier Platform for Cognizing Massive Electroencephalogram
abstract
An end-to-end platform assembling multiple tiers is built for precisely cognizing brain activities. Being fed massive electroencephalogram (EEG) data, the time-frequency spectrograms are conventionally projected into the episode-wise feature matrices (seen as tier-1). A spiking neural network (SNN) based tier is designed to distill the principle information in terms of spike-streams from the rare features, which maintains the temporal implication in the nature of EEGs. The proposed tier-3 transposes time- and space-domain of spike patterns from the SNN; and feeds the transposed pattern-matrices into an artificial neural network (ANN, Transformer specifically) known as tier-4, where a special spanning topology is proposed to match the two-dimensional input form. In this manner, cognition such as classification is conducted with high accuracy. For proof-of-concept, the sleep stage scoring problem is demonstrated by introducing multiple EEG datasets with the largest comprising 42,560 hours recorded from 5,793 subjects. From experiment results, our platform achieves the general cognition overall accuracy of 87% by leveraging sole EEG, which is 2% superior to the state-of-the-art. Moreover, our developed multi-tier methodology offers visible and graphical interpretations of the temporal characteristics of EEG by identifying the critical episodes, which is demanded in neurodynamics but hardly appears in conventional cognition scenarios.
Zheng Chen 0012, Lingwei Zhu, Ziwei Yang 0002
IJCAI2
2022 Automated Cancer Subtyping via Vector Quantization Mutual Information Maximization
Zheng Chen 0012, Lingwei Zhu, Ziwei Yang 0002, Takashi Matsubara 0001
ECML/PKDD (1)2
2021 Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning
abstract
The recent boom in the literature on entropy-regularized reinforcement learning (RL) approaches reveals that Kullback-Leibler (KL) regularization brings advantages to RL algorithms by canceling out errors under mild assumptions. However, existing analyses focus on fixed regularization with a constant weighting coefficient and do not consider cases where the coefficient is allowed to change dynamically. In this paper, we study the dynamic coefficient scheme and present the first asymptotic error bound. Based on the dynamic coefficient error bound, we propose an effective scheme to tune the coefficient according to the magnitude of error in favor of more robust learning. Complementing this development, we propose a novel algorithm, Geometric Value Iteration (GVI), that features a dynamic error-aware KL coefficient design with the aim of mitigating the impact of errors on performance. Our experiments demonstrate that GVI can effectively exploit the trade-off between learning speed and robustness over uniform averaging of a constant KL coefficient. The combination of GVI and deep networks shows stable learning behavior even in the absence of a target network, where algorithms with a constant KL coefficient would greatly oscillate or even fail to converge.
Toshinori Kitamura, Lingwei Zhu, Takamitsu Matsubara
ACML2
2021 Cautious Actor-Critic
abstract
The oscillating performance of off-policy learning and persisting errors in the actor-critic(AC) setting call for algorithms that can conservatively learn to suit the stability-critical applications better. In this paper, we propose a novel off-policy AC algorithm cautious actor-critic (CAC). The name cautious comes from the doubly conservative nature that we exploit the classic policy interpolation from conservative policy iteration for the actor and the entropy-regularization of conservative value iteration for the critic. Our key observation is the entropy-regularized critic facilitates and simplifies the unwieldy interpolated actor update while still ensuring robust policy improvement. We compare CAC to state-of-the-art AC methods on a set of challenging continuous control problems and demonstrate thatCAC achieves comparable performance while significantly stabilizes learning.
Lingwei Zhu, Toshinori Kitamura, Takamitsu Matsubara
ACML1
2020 Dynamic Actor-Advisor Programming for Scalable Safe Reinforcement Learning
abstract
Real-world robots have complex strict constraints. Therefore, safe reinforcement learning algorithms that can simultaneously minimize the total cost and the risk of constraint violation are crucial. However, almost no algorithms exist that can scale to high-dimensional systems to the best of our knowledge. In this paper, we propose Dynamic Actor-Advisor Programming (DAAP), as an algorithm for sample-efficient and scalable safe reinforcement learning. DAAP employs two control policies, actor and advisor. They are updated to minimize total cost and risk of constraint violation intertwiningly and smoothly towards each other's direction by using the other as the baseline policy in the Kullback-Leibler divergence of Dynamic Policy Programming framework. We demonstrate the scalability and sample efficiency of DAAP through its application on simulated robot arm control tasks with performance comparisons to baselines.
Lingwei Zhu, Yunduan Cui, Takamitsu Matsubara
ICRA1