Jiachen Hu

dblp:239/5040 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 72% Transfer learning and domain adaptation · 21% Learning theory · 7%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 54% Quantum computing and quantum information · 46%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
regret minimization
1.322024
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret · ICML 2024
Near-Optimal Representation Learning for Linear Bandits and Linear RL · ICML 2021
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
1.222023
Provable Sim-to-real Transfer in Continuous Domain with Partial Observations · ICLR 2023
Understanding Domain Randomization for Sim-to-real Transfer · ICLR 2022
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.912025
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability · ICML 2025
Machine learning › Learning theory
online learning
0.912025
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability · ICML 2025
Machine learning › Reinforcement learning
strategic decision-making
0.912025
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability · ICML 2025
Algorithmic game theory and mechanism design › mechanism design
information asymmetry
0.912025
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability · ICML 2025
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.812024
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret · ICML 2024
Quantum computing and quantum information › quantum machine learning
quantum reinforcement learning
0.812024
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret · ICML 2024
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.712023
Provable Sim-to-real Transfer in Continuous Domain with Partial Observations · ICLR 2023
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization
0.612022
Understanding Domain Randomization for Sim-to-real Transfer · ICLR 2022
Machine learning › Reinforcement learning
exploration
0.612022
Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in Solver · ICLR 2022
Machine learning › Reinforcement learning › markov decision process › low-rank MDP
linear mixture MDP
0.612022
Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in Solver · ICLR 2022
Machine learning › Reinforcement learning
markov decision process
0.612022
Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in Solver · ICLR 2022
Machine learning › Reinforcement learning › exploration › exploration in markov decision processes
reward-free exploration
0.612022
Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in Solver · ICLR 2022
Machine learning › Reinforcement learning
constrained reinforcement learning
0.512021
Efficient Reinforcement Learning in Factored MDPs with Application to Constrained RL · ICLR 2021
Machine learning › Reinforcement learning
factored reinforcement learning
0.512021
Efficient Reinforcement Learning in Factored MDPs with Application to Constrained RL · ICLR 2021
Machine learning › Reinforcement learning
reinforcement learning theory
0.512021
Near-Optimal Representation Learning for Linear Bandits and Linear RL · ICML 2021
Machine learning › Reinforcement learning › exploration › efficient exploration
sample-efficient exploration
0.512021
Near-Optimal Representation Learning for Linear Bandits and Linear RL · ICML 2021
Machine learning › Reinforcement learning › bandit
bandit learning
0.412020
Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication · ICLR 2020
Machine learning › Reinforcement learning › multi-armed bandit › multi-agent bandit
distributed bandit learning
0.412020
Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication · ICLR 2020
Distributed systems › distributed machine learning
communication-efficient distributed learning
0.412020
Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication · ICLR 2020
Distributed systems
distributed coordination
0.412020
Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication · ICLR 2020
Machine learning › Reinforcement learning › markov decision process › finite markov decision processes
tabular markov decision process
0.212024
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret · ICML 2024
Machine learning › Reinforcement learning › bandit
linear bandits
0.112021
Near-Optimal Representation Learning for Linear Bandits and Linear RL · ICML 2021
Machine learning › Reinforcement learning
multi-armed bandit
0.112021
Near-Optimal Representation Learning for Linear Bandits and Linear RL · ICML 2021

Methods — techniques the papers use, named apart from their topics

sample-efficient algorithm · 1.7causal inference · 1.7value target regression · 1.5quantum estimation · 1.5lazy updating · 1.5UCRL · 1.5reinforcement learning · 0.6plug-in solver · 0.6optimism · 0.6domain randomization · 0.6regret analysis · 0.4communication compression · 0.4
YearPublicationVenuePosition
2025 The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability
abstract
Information asymmetry is a pervasive feature of multi-agent systems, especially evident in economics and social sciences. In these settings, agents tailor their actions based on private information to maximize their rewards. These strategic behaviors often introduce complexities due to confounding variables. Simultaneously, knowledge transportability poses another significant challenge, arising from the difficulties of conducting experiments in target environments. It requires transferring knowledge from environments where empirical data is more readily available. Against these backdrops, this paper explores a fundamental question in online learning: Can we employ non-i.i.d. actions to learn about confounders even when requiring knowledge transfer? We present a sample-efficient algorithm designed to accurately identify system dynamics under information asymmetry and to navigate the challenges of knowledge transfer effectively in reinforcement learning, framed within an online strategic interaction model. Our method provably achieves learning of an $\epsilon$-optimal policy with a tight sample complexity of $\tilde{O}(1/\epsilon^2)$.
Jiachen Hu, Rui Ai 0002, Han Zhong 0001, Xiaoyu Chen 0008, Liwei Wang 0001, Zhaoran Wang 0001, Zhuoran Yang
ICML1
2025 Guided Model-Based Policy Search Method for Fast Motor Learning of Robots With Learned Dynamics
abstract
Reinforcement learning recently has achieved impressive success in allowing robots to learn complex motor skills in simulation environments. However, most of these successes are difficult to transfer to physical robots since current algorithms require lots of practical training and complex sim-to-real transfer skills. To improve the learning efficiency and adaptability of physical robots, this article proposes a guided model-based policy search (GMBPS) algorithm inspired by a hypothetical model-free (MF) and model-based (MB) actor-critic brain implementation. This approach bridges the gap between MF and MB control processes, overcoming the suboptimality of MB methods and speeding up the learning rate of MF methods. Additionally, a one-step predictive control framework is proposed for minimizing the impact of delayed sensorimotor information in real-world tasks. This helps to accurately control the action cycle time and ensures the feasibility of MB planning for physical robots. The simulation and experimental results demonstrate that the proposed approach enables a 6-DOF UR5e robot arm to learn various reaching tasks in a few minutes with better policies and higher learning efficiency.Note to Practitioners—Reinforcement learning is becoming a popular framework that allows robots to learn complex motor skills without building analytical models of controlled plants. However, low learning efficiency severely limits its application in practical robots, where robots have to quickly adapt to dynamically changing environments in micro-data situations. To solve the inefficiency problem of physical robot learning from scratch, this paper proposes a MF and MB fusion control algorithm inspired by a hypothetical MF and MB actor-critic brain implementation. The motion decision process is modeled as an optimization problem with inequality constraints. The global MF value function is incorporated into the MB objective function, extending the short-term optimization into a long-term version to overcome the suboptimality of conventional MB methods. The MB policy is searched based on the quadratic penalty method with the guide of the MF policy, which helps improve the quality of policy at every decision-making step. Moreover, since the model dynamics is fitted by a probabilistic neural network, the proposed method is not only applicable to joint-driven robots but also provides a feasible solution for the control of various robotic systems with complex dynamics, such as soft robots and musculoskeletal robots.
Xiao Huang 0004, Xingfang Wang, Jiachen Hu, Hui Li 0047, Zhihong Jiang
IEEE Trans Autom. Sci. Eng.4
2024 ZeroSwap: Data-Driven Optimal Market Making in Decentralized Finance
Viraj Nadkarni, Jiachen Hu, Ranvir Rana, Chi Jin 0001, Sanjeev R. Kulkarni, Pramod Viswanath
FC (1)2
2024 Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret
abstract
While quantum reinforcement learning (RL) has attracted a surge of attention recently, its theoretical understanding is limited. In particular, it remains elusive how to design provably efficient quantum RL algorithms that can address the exploration-exploitation trade-off. To this end, we propose a novel UCRL-style algorithm that takes advantage of quantum computing for tabular Markov decision processes (MDPs) with $S$ states, $A$ actions, and horizon $H$, and establish an $\mathcal{O}(\mathrm{poly}(S, A, H, \log T))$ worst-case regret for it, where $T$ is the number of episodes. Furthermore, we extend our results to quantum RL with linear function approximation, which is capable of handling problems with large state spaces. Specifically, we develop a quantum algorithm based on value target regression (VTR) for linear mixture MDPs with $d$-dimensional linear representation and prove that it enjoys $\mathcal{O}(\mathrm{poly}(d, H, \log T))$ regret. Our algorithms are variants of UCRL/UCRL-VTR algorithms in classical RL, which also leverage a novel combination of lazy updating mechanisms and quantum estimation subroutines. This is the key to breaking the $\Omega(\sqrt{T})$-regret barrier in classical RL. To the best of our knowledge, this is the first work studying the online exploration in quantum RL with provable logarithmic worst-case regret.
Han Zhong 0001, Jiachen Hu, Yecheng Xue, Tongyang Li, Liwei Wang 0001
ICML2
2023 Provable Sim-to-real Transfer in Continuous Domain with Partial Observations
Jiachen Hu, Han Zhong 0001, Chi Jin 0001, Liwei Wang 0001
ICLR1
2022 Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in Solver
Xiaoyu Chen 0008, Jiachen Hu, Lin Yang 0011, Liwei Wang 0001
ICLR2
2022 Understanding Domain Randomization for Sim-to-real Transfer
Xiaoyu Chen 0008, Jiachen Hu, Chi Jin 0001, Lihong Li 0001, Liwei Wang 0001
ICLR2
2021 Efficient Reinforcement Learning in Factored MDPs with Application to Constrained RL
Xiaoyu Chen 0008, Jiachen Hu, Lihong Li 0001, Liwei Wang 0001
ICLR2
2021 Near-Optimal Representation Learning for Linear Bandits and Linear RL
abstract
This paper studies representation learning for multi-task linear bandits and multi-task episodic RL with linear value function approximation. We first consider the setting where we play $M$ linear bandits with dimension $d$ concurrently, and these bandits share a common $k$-dimensional linear representation so that $k\ll d$ and $k \ll M$. We propose a sample-efficient algorithm, MTLR-OFUL, which leverages the shared representation to achieve $\tilde{O}(M\sqrt{dkT} + d\sqrt{kMT} )$ regret, with $T$ being the number of total steps. Our regret significantly improves upon the baseline $\tilde{O}(Md\sqrt{T})$ achieved by solving each task independently. We further develop a lower bound that shows our regret is near-optimal when $d > M$. Furthermore, we extend the algorithm and analysis to multi-task episodic RL with linear value function approximation under low inherent Bellman error (Zanette et al., 2020a). To the best of our knowledge, this is the first theoretical result that characterize the benefits of multi-task representation learning for exploration in RL with function approximation.
Jiachen Hu, Xiaoyu Chen 0008, Chi Jin 0001, Lihong Li 0001, Liwei Wang 0001
ICML1
2020 Boundary-aware Segmentation Network Using Multi-Task Enhancement for Ultrasound Image
abstract
Complicated medical image analysis often requires a combination of disease classification, lesion detection and lesion segmentation. However, models designed for different tasks produce inconsistent or non-corresponding predictions and ignore the implicit connections between tasks. We propose a novel framework, which makes full use of the fact that segmentation and detection are mutually beneficial, boosts these three tasks in a unified framework. The proposed Information Enhancement Module uses classification information as a beneficial supplement to locate lesion quickly for segmentation. To further achieve fine segmentation with clear boundaries, we propose a Boundary-aware Loss, which dynamically adjusts supervised signal, so that our model pays more attention to boundary in later stages of training. Through experiments conducted on Thyroid Ultrasound dataset, we have demonstrated the good performance of the proposed method in joint segmentation and detection.
Jiachen Hu, Mei Yu 0004, Xi Wei 0002, Han Jiang 0004, Zhiqiang Liu 0002, Jie Gao 0008, Xuewei Li 0001
BIBM2
2020 Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication
Yuanhao Wang 0001, Jiachen Hu, Xiaoyu Chen 0008, Liwei Wang 0001
ICLR2
2020 Generative Adversarial Network Using Multi-modal Guidance for Ultrasound Images Inpainting
Jiachen Hu, Xi Wei 0002, Mei Yu 0004, Jie Gao 0008, Zhiqiang Liu 0002, Xuewei Li 0001
ICONIP (1)2