Litian Liang

dblp:301/7979 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 77% Motion planning and robot control · 11% Deep learning architectures and training · 5%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis
DNA language model
1.012026
CodonMoE: DNA language models for codon-dependent mRNA prediction · Bioinform. 2026
Machine learning › Reinforcement learning
imitation learning
0.912025
When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning? · AAAI 2025
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning
0.912025
When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning? · AAAI 2025
Machine learning › Reinforcement learning
model-based reinforcement learning
0.712023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Machine learning › Reinforcement learning
policy learning
0.712023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Robotics › Motion planning and robot control
trajectory optimization
0.712023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Bioinformatics and computational biology › survival analysis
survival prediction
0.712023
Causally-Aware Intraoperative Imputation for Overall Survival Time Prediction · CVPR 2023
Machine learning › Reinforcement learning
temporal difference learning
0.612022
Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks · ICML 2022
Machine learning › Reinforcement learning
value-based reinforcement learning
0.612022
Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks · ICML 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.312026
CodonMoE: DNA language models for codon-dependent mRNA prediction · Bioinform. 2026
Machine learning › Reinforcement learning › deep reinforcement learning
visual policy learning
0.312025
When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning? · AAAI 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.212023
Causally-Aware Intraoperative Imputation for Overall Survival Time Prediction · CVPR 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.212023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023

Methods — techniques the papers use, named apart from their topics

mixture of experts · 2.0adapter · 2.0causal graph · 1.3causal discovery · 1.3state-to-visual DAgger · 0.9empirical comparison · 0.9world model · 0.7variational bound · 0.7reparameterization · 0.7ensemble of deep networks · 0.6
YearPublicationVenuePosition
2026 CodonMoE: DNA language models for codon-dependent mRNA prediction
abstract
MOTIVATION: Genomic language models (gLMs) face a fundamental efficiency challenge: one must either maintain separate specialized models for each biological modality (DNA and RNA) or develop large multimodal architectures. Both approaches impose significant computational burdens-modality-specific models require redundant infrastructure despite inherent biological connections, while multi-modal architectures demand increased parameter counts and extensive cross-modality pretraining. RESULTS: To address this limitation, we introduce CodonMoE (Adaptive Mixture of Codon Reformative Experts), a lightweight adapter that transforms DNA language models into effective RNA analyzers without RNA-specific pretraining. Our theoretical analysis establishes CodonMoE as a universal approximator at the codon level, capable of mapping arbitrary functions from codon sequences to codon-dependent RNA properties given sufficient expert capacity. Across four RNA prediction tasks spanning stability, expression, and regulation, DNA models augmented with CodonMoE significantly outperform their unmodified counterparts, with the HyenaDNA+CodonMoE series achieving state-of-the-art results using 80% fewer parameters than specialized RNA models. By maintaining sub-quadratic complexity while achieving superior performance, our approach provides a principled path toward unifying genomic language modeling, leveraging more abundant DNA data and reducing computational overhead while preserving modality-specific performance advantages. AVAILABILITY AND IMPLEMENTATION: Source code for the method and to reproduce the results is available at https://github.com/Kingsford-Group/CodonMoE.
Shiyi Du, Litian Liang, Carl Kingsford
Bioinform.2
2025 When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning?
abstract
Learning policies from high-dimensional visual inputs, such as pixels and point clouds, is crucial in various applications. Visual reinforcement learning is a promising approach that directly trains policies from visual observations, although it faces challenges in sample efficiency and computational costs. This study conducts an empirical comparison of State-to-Visual DAgger — a two-stage framework that initially trains a state policy before adopting online imitation to learn a visual policy — and Visual RL across a diverse set of tasks. We evaluate both methods across 16 tasks from three benchmarks, focusing on their asymptotic performance, sample efficiency, and computational costs. Surprisingly, our findings reveal that State-to-Visual DAgger does not universally outperform Visual RL but shows significant advantages in challenging tasks, offering more consistent performance. In contrast, its benefits in sample efficiency are less pronounced, although it often reduces the overall wall-clock time required for training. Based on our findings, we provide recommendations for practitioners and hope that our results contribute valuable perspectives for future research in visual policy learning.
Tongzhou Mu, Zhaoyang Li 0007, Stanislaw Wiktor Strzelecki, Xiu Yuan, Yunchao Yao, Litian Liang, Hao Su 0001
AAAI6
2024 A Quantitative Approach for Evaluating Disease Focus and Interpretability of Deep Learning Models for Alzheimer's Disease Classification
abstract
Deep learning (DL) models have shown significant potential in Alzheimer’s Disease (AD) classification. However, understanding and interpreting these models remains challenging, which hinders the adoption of these models in clinical practice. Techniques such as saliency maps have been proven effective in providing visual and empirical clues about how these models work, but there still remains a gap in understanding which specific brain regions DL models focus on and whether these brain regions are pathologically associated with AD.To bridge such gap, in this study, we developed a quantitative disease-focusing strategy to first enhance the interpretability of DL models using saliency maps and brain segmentations; then we propose a disease-focus (DF) score that quantifies how much a DL model focuses on brain areas relevant to AD pathology based on clinically known MRI-based pathological regions of AD. Using this strategy, we compared several state-of-the-art DL models, including a baseline 3D ResNet model, a pretrained MedicalNet model, and a MedicalNet with data augmentation to classify patients with AD vs. cognitive normal patients using MRI data; then we evaluated these models in terms of their abilities to focus on disease-relevant regions. Our results show interesting disease-focusing patterns with different models, particularly characteristic patterns with the pretrained models and data augmentation, and also provide insight into their classification performance. These results suggest that the approach we developed for quantitatively assessing the abilities of DL models to focus on disease-relevant regions may help improve the interpretability of these models for AD classification and facilitate their adoption for AD diagnosis in clinical practice. The code is publicly available at https://github.com/Liang-lt/ADNI.
Thomas Yu Chow Tam, Litian Liang, Haohan Wang
BIBM2
2023 Causally-Aware Intraoperative Imputation for Overall Survival Time Prediction
abstract
Previous efforts in vision community are mostly made on learning good representations from visual patterns. Beyond this, this paper emphasizes the high-level ability of causal reasoning. We thus present a case study of solving the challenging task of Overall Survival (OS) time in primary liver cancers. Critically, the prediction of OS time at the early stage remains challenging, due to the unobvious image patterns of reflecting the OS. To this end, we propose a causal inference system by leveraging the intraoperative attributes and the correlation among them, as an intermediate supervision to bridge the gap between the images and the final OS. Particularly, we build a causal graph, and train the images to estimate the intraoperative attributes for final as prediction. We present a novel Causally-aware Intraoperative Imputation Model (CAWIM) that can sequentially predict each attribute using its parent nodes in the estimated causal graph. To determine the causal directions, we propose a splitting-voting mechanism, which votes for the direction for each pair of adjacent nodes among multiple predictions obtained via causal discovery from heterogeneity. The practicability and effectiveness of our method are demonstrated by the promising results on liver cancer dataset of 361 patients with long-term observations.
Xuelin Qian, Litian Liang, Lingjie Kong, Qiaole Dong, Jiejun Chen, Dingxia Liu, Xiuzhong Yao, Yanwei Fu 0001
CVPR3
2023 Reparameterized Policy Learning for Multimodal Trajectory Optimization
abstract
We investigate the challenge of parametrizing policies for reinforcement learning (RL) in high-dimensional continuous action spaces. Our objective is to develop a multimodal policy that overcomes limitations inherent in the commonly-used Gaussian parameterization. To achieve this, we propose a principled framework that models the continuous RL policy as a generative model of optimal trajectories. By conditioning the policy on a latent variable, we derive a novel variational bound as the optimization objective, which promotes exploration of the environment. We then present a practical model-based RL method, called Reparameterized Policy Gradient (RPG), which leverages the multimodal policy parameterization and learned world model to achieve strong exploration capabilities and high data efficiency. Empirical results demonstrate that our method can help agents evade local optima in tasks with dense rewards and solve challenging sparse-reward environments by incorporating an object-centric intrinsic reward. Our method consistently outperforms previous approaches across a range of tasks. Code and supplementary materials are available on the project page https://haosulab.github.io/RPG/
Zhiao Huang, Litian Liang, Zhan Ling, Chuang Gan 0001, Hao Su 0001
ICML2
2022 Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks
abstract
In temporal-difference reinforcement learning algorithms, variance in value estimation can cause instability and overestimation of the maximal target value. Many algorithms have been proposed to reduce overestimation, including several recent ensemble methods, however none have shown success in sample-efficient learning through addressing estimation variance as the root cause of overestimation. In this paper, we propose MeanQ, a simple ensemble method that estimates target values as ensemble means. Despite its simplicity, MeanQ shows remarkable sample efficiency in experiments on the Atari Learning Environment benchmark. Importantly, we find that an ensemble of size 5 sufficiently reduces estimation variance to obviate the lagging target network, eliminating it as a source of bias and further gaining sample efficiency. We justify intuitively and empirically the design choices in MeanQ, including the necessity of independent experience sampling. On a set of 26 benchmark Atari environments, MeanQ outperforms all tested baselines, including the best available baseline, SUNRISE, at 100K interaction steps in 16/26 environments, and by 68% on average. MeanQ also outperforms Rainbow DQN at 500K steps in 21/26 environments, and by 49% on average, and achieves average human-level performance using 200K ($\pm$100K) interaction steps. Our implementation is available at https://github.com/indylab/MeanQ.
Litian Liang, Yaosheng Xu, Stephen McAleer, Dailin Hu, Alexander Ihler, Pieter Abbeel, Roy Fox
ICML1