EDBT 2026 Demo / reviewers in the wild / expert
Haotian Fu
dblp:237/9681
· DBLP profile ↗
26ranked-venue papers
11as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective and efficient intracortical brain signal decoding with spiking neural networks
Haotian Fu, Peng Zhang 0106, Herui Zhang, Ziwei Wang 0009, Dongrui Wu |
Neurocomputing | 1 |
| 2026 | Spiking neural network for intra-cortical brain signal decoding
Haotian Fu, Herui Zhang, Peng Zhang 0106, Wei Li 0100, Dongrui Wu |
Knowl. Based Syst. | 2 |
| 2026 | Event-Based Motion Deblurring via Multi-Temporal Granularity FusionabstractConventional frame-based cameras inevitably produce blurry effects due to motion occurring during the exposure time. Event camera, a bio-inspired sensor offering continuous visual information could enhance the deblurring performance. Effectively utilizing the high-temporal-resolution event data is crucial for extracting precise motion information and enhancing deblurring performance. However, existing event-based image deblurring methods usually utilize voxel-based event representations, losing the fine-grained temporal details that are mathematically essential for fast motion deblurring. In this paper, we first introduce point cloud-based event representation into the image deblurring task and propose a Multi-Temporal Granularity Network (MTGNet). It combines the spatially dense but temporally coarse-grained voxel-based event representation and the temporally fine-grained but spatially sparse point cloud-based event. To seamlessly integrate such complementary representations, we design a Fine-grained Point Branch. An Aggregation and Mapping Module (AMM) is proposed to align the low-level point-based features with frame-based features and an Adaptive Feature Diffusion Module (AFDM) is designed to manage the resolution discrepancies between event data and image data by enriching the sparse point feature. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art approaches on both synthetic and real-world datasets. Our code is available at: https://github.com/xplin13/MTGNet. Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Yue Zhou 0010, Haotian Fu, Biao Pan, Bojun Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | ClearSight: Human Vision-Inspired Solutions for Event-Based Motion DeblurringabstractMotion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of event streams, we employ Spiking Neural Networks (SNNs) for motion feature extraction and Artificial Neural Networks (ANNs) for color information processing. Due to the non-uniform distribution and inherent redundancy of event data, existing cross-modal feature fusion methods exhibit certain limitations. Inspired by the visual attention mechanism in the human visual system, this study introduces a bioinspired dual-drive hybrid network (BDHNet). Specifically, the Neuron Configurator Module (NCM) is designed to dynamically adjusts neuron configurations based on cross-modal features, thereby focusing the spikes in blurry regions and adapting to varying blurry scenarios dynamically. Additionally, the Region of Blurry Attention Module (RBAM) is introduced to generate a blurry mask in an unsupervised manner, effectively extracting motion clues from the event features and guiding more accurate cross-modal feature fusion. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art methods on both synthetic and real-world datasets. Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Hongxiang Huang, Yue Zhou 0010, Haotian Fu, Bojun Cheng |
ICCV | 7 |
| 2025 | Knowledge Retention in Continual Model-Based Reinforcement LearningabstractWe propose DRAGO, a novel approach for continual model-based reinforcement learning aimed at improving the incremental development of world models across a sequence of tasks that differ in their reward functions but not the state space or dynamics. DRAGO comprises two key components: Synthetic Experience Rehearsal, which leverages generative models to create synthetic experiences from past tasks, allowing the agent to reinforce previously learned dynamics without storing data, and Regaining Memories Through Exploration, which introduces an intrinsic reward mechanism to guide the agent toward revisiting relevant states from prior tasks. Together, these components enable the agent to maintain a comprehensive and continually developing world model, facilitating more effective learning and adaptation across diverse environments. Empirical evaluations demonstrate that DRAGO is able to preserve knowledge across tasks, achieving superior performance in various continual learning scenarios. Haotian Fu, Yixiang Sun, Michael L. Littman, George Dimitri Konidaris |
ICML | 1 |
| 2025 | Learning Parameterized Skills from DemonstrationsabstractWe present DEPS, an end-to-end algorithm for discovering parameterized skills from expert demonstrations. Our method learns parameterized skill policies jointly with a meta-policy that selects the appropriate discrete skill and continuous parameters at each timestep. Using a combination of temporal variational inference and information-theoretic regularization methods, we address the challenge of degeneracy common in latent variable models, ensuring that the learned skills are temporally extended, semantically meaningful, and adaptable. We empirically show that learning parameterized skills from multitask expert demonstrations significantly improves generalization to unseen tasks. Our method outperforms multitask as well as skill learning baselines on both LIBERO and MetaWorld benchmarks. We also demonstrate that DEPS discovers interpretable parameterized skills, such as an object grasping skill whose continuous arguments define the grasp location. Vedant Gupta, Haotian Fu, Calvin Luo, Yiding Jiang, George Dimitri Konidaris |
NeurIPS | 2 |
| 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and RegressionabstractEvent cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks. Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A Simple and Effective Point-Based Network for Event Camera 6-DOFs Pose RelocalizationabstractEvent cameras exhibit remarkable attributes such as high dynamic range, asynchronicity, and low latency, making them highly suitable for vision tasks that involve highspeed motion in challenging lighting conditions. These cameras implicitly capture movement and depth information in events, making them appealing sensors for Camera Pose Relocalization (CPR) tasks. Nevertheless, existing CPR networks based on events neglect the pivotal finegrained temporal information in events, resulting in unsatisfactory performance. Moreover, the energy-efficient features are further compromised by the use of excessively complex models, hindering efficient deployment on edge devices. In this paper, we introduce PEPNet, a simple and effective point-based network designed to regress six degrees of freedom (6-DOFs) event camera poses. We rethink the relationship between the event camera and CPR tasks, leveraging the raw Point Cloud directly as network input to harness the high-temporal resolution and inherent sparsity of events. PEPNet is adept at abstracting the spatial and implicit temporal features through hierarchical structure and explicit temporal features by Attentive Bidirectional Long Short-Term Memory (A-Bi-LSTM). Byemploying a carefully crafted lightweight design, PEPNet delivers state-of-the-art (SOTA) performance on both indoor and outdoor datasets with meager computational resources. Specifically, PEPNet attains a significant 38% and 33% performance improvement on the random split IJRR and M3ED datasets, respectively. Moreover, the lightweight design version PEPNettinyaccomplishes results comparable to the SOTA while employing a mere 0.5% of the parameters. Jiadong Zhu, Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Bojun Cheng |
CVPR | 4 |
| 2024 | EPO: Hierarchical LLM Agents with Environment Preference OptimizationabstractLong-horizon decision-making tasks present significant challenges for LLM-based agents due to the need for extensive planning over multiple steps.In this paper, we propose a hierarchical framework that decomposes complex tasks into manageable subgoals, utilizing separate LLMs for subgoal prediction and lowlevel action generation.To address the challenge of creating training signals for unannotated datasets, we develop a reward model that leverages multimodal environment feedback to automatically generate reward signals.We introduce Environment Preference Optimization (EPO), a novel method that generates preference signals from the environment's feedback and uses them to train LLM-based agents.Extensive experiments on ALFRED demonstrate the state-of-the-art performance of our framework, achieving first place on the ALFRED public leaderboard and showcasing its potential to improve long-horizon decision-making in diverse environments. Haotian Fu, George Dimitri Konidaris |
EMNLP | 2 |
| 2024 | SpikePoint: An Efficient Point-based Spiking Neural Network for Event Cameras Action RecognitionabstractEvent cameras are bio-inspired sensors that respond to local changes in light intensity and feature low latency, high energy efficiency, and high dynamic range. Meanwhile, Spiking Neural Networks (SNNs) have gained significant attention due to their remarkable efficiency and fault tolerance. By synergistically harnessing the energy efficiency inherent in event cameras and the spike-based processing capabilities of SNNs, their integration could enable ultra-low-power application scenarios, such as action recognition tasks. However, existing approaches often entail converting asynchronous events into conventional frames, leading to additional data mapping efforts and a loss of sparsity, contradicting the design concept of SNNs and event cameras. To address this challenge, we propose SpikePoint, a novel end-to-end point-based SNN architecture. SpikePoint excels at processing sparse event cloud data, effectively extracting both global and local features through a singular-stage structure. Leveraging the surrogate training method, SpikePoint achieves high accuracy with few parameters and maintains low power consumption, specifically employing the identity mapping feature extractor on diverse datasets. SpikePoint achieves state-of-the-art (SOTA) performance on four event-based action recognition datasets using only 16 timesteps, surpassing other SNN methods. Moreover, it also achieves SOTA performance across all methods on three datasets, utilizing approximately 0.3 % of the parameters and 0.5 % of power consumption employed by artificial neural networks (ANNs). These results emphasize the significance of Point Cloud and pave the way for many ultra-low-power event-based data processing applications. Xiaopeng Lin, Haotian Fu, Bojun Cheng |
ICLR | 5 |
| 2024 | Language-guided Skill Learning with Temporal Variational InferenceabstractWe present an algorithm for skill discovery from expert demonstrations. The algorithm first utilizes Large Language Models (LLMs) to propose an initial segmentation of the trajectories. Following that, a hierarchical variational inference framework incorporates the LLM-generated segmentation information to discover reusable skills by merging trajectory segments. To further control the trade-off between compression and reusability, we introduce a novel auxiliary objective based on the Minimum Description Length principle that helps guide this skill discovery process. Our results demonstrate that agents equipped with our method are able to discover skills that help accelerate learning and outperform baseline skill learning approaches on new long-horizon tasks in BabyAI, a grid world navigation environment, as well as ALFRED, a household simulation environment. Haotian Fu, Pratyusha Sharma, Elias Stengel-Eskin, George Dimitri Konidaris, Nicolas Le Roux, Marc-Alexandre Côté, Xingdi Yuan |
ICML | 1 |
| 2024 | CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural NetworksabstractSpiking neural networks (SNNs) are promising brain-inspired energy-efficient models. Compared to conventional deep Artificial Neural Networks (ANNs), SNNs exhibit superior efficiency and capability to process temporal information. However, it remains a challenge to train SNNs due to their undifferentiable spiking mechanism. The surrogate gradients method is commonly used to train SNNs, but often comes with an accuracy disadvantage over ANNs counterpart. We link the degraded accuracy to the vanishing of gradient on the temporal dimension through the analytical and experimental study of the training process of Leaky Integrate-and-Fire (LIF) Neuron-based SNNs. Moreover, we propose the Complementary Leaky Integrate-and-Fire (CLIF) Neuron. CLIF creates extra paths to facilitate the backpropagation in computing temporal gradient while keeping binary output. CLIF is hyperparameter-free and features broad applicability. Extensive experiments on a variety of datasets demonstrate CLIF's clear performance advantage over other neuron models. Furthermore, the CLIF's performance even slightly surpasses superior ANNs with identical network structure and training conditions. The code is available at https://github.com/HuuYuLong/Complementary-LIF. Xiaopeng Lin, Haotian Fu, Zunchang Liu, Biao Pan, Bojun Cheng |
ICML | 4 |
| 2024 | Model-based Reinforcement Learning for Parameterized Action SpacesabstractWe propose a novel model-based reinforcement learning algorithm---Dynamics Learning and predictive control with Parameterized Actions (DLPA)---for Parameterized Action Markov Decision Processes (PAMDPs). The agent learns a parameterized-action-conditioned dynamics model and plans with a modified Model Predictive Path Integral control. We theoretically quantify the difference between the generated trajectory and the optimal trajectory during planning in terms of the value they achieved through the lens of Lipschitz Continuity. Our empirical results on several standard benchmarks show that our algorithm achieves superior sample efficiency and asymptotic performance than state-of-the-art PAMDP methods. Renhao Zhang, Haotian Fu, Yilin Miao, George Dimitri Konidaris |
ICML | 2 |
| 2024 | PipeCIM: A High-Throughput Computing-In-Memory Microprocessor With Nested Pipeline and RISC-V Extended InstructionsabstractThe large number of multiply accumulate (MAC) operations in Convolutional Neural Network (CNN) leads to substantial data migration and computation. Although computing-in-memory (CIM) proves to be a promising paradigm for MAC operations, high throughput CNN accelerator still confronts bottlenecks from: the low MAC utilization and the uncessary off-chip memory access. In this paper, we propose a high throughput CIM-based CNN accelerator PipeCIM with three hierarchies of pipelines: Intra-Macro, Near-Memory and Tile-Level. The Intra-Macro Pipeline parallelly executes data transfer and in-memory-computing (IMC) operations. The Near-Memory Pipeline alleviates memory access for pooling and data reshaping. The Tile-Level Pipeline establishes a layer-wise pipeline to further improve the throughput while reducing control complexity. PipeCIM introduces the nested scheme and a Unidirectional Divergent Connection Protocol (UDTCP) to simplify the control of data flow with the help of customized RISC-V instructions. To validate our design, PipeCIM was prototyped in 55 nm process node, achieving energy efficiency of 133.8 TOPS/W and peak throughput of 819 GOPS with a 16KB CIM array, which can accelerate VGG-16 to 128.56$\times$or Inception to 19.754$\times$compared to the baseline. Tingran Chen, Wenjia Wang 0011, Haotian Fu, Wente Yi, Bojun Cheng, He Zhang 0011, Biao Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | DS-CIM: A 40nm Asynchronous Dual-Spike Driven, MRAM Compute-In-Memory Macro for Spiking Neural NetworkabstractCompute-in-memory (CIM) based on emerging nonvolatile memory (eNVM) is an effective way to deploy neural networks to low-power edge devices for both storage and computation. NVMs such as ReRAM have been widely used in CIM. Meanwhile, MRAM has higher read and write cycles, lower device and cycle variation and a lower bit error rate, making it equally attractive for storage. However, the high read current and low on/off ratio result in large energy consumption in MRAM read limiting its large-scale application in CIM. The spiking neural network (SNN) represents the information as sparse spike sequences and facilitates hardware to achieve low-power computing by taking advantage of its spatial-temporal sparsity. To further increase the input sparsity of SNN and reduce the read energy consumption, this paper proposes ADC-free, dual-spike (DS) -CIM macro, a spiking MRAM CIM macro driven by asynchronous dual spikes. Compared to the conventional rate coding, our dual-spike coding method uses only 2 spikes to encode the information without losing accuracy. Moreover, the event-driven feature allows the macro to have sub-nW static power consumption. Our DS-CIM macro achieves comparable or higher accuracy while maintaining very low energy consumption. Specifically, it achieves accuracies of 96.99%, 82.87%, 90.00%, and 85.97% for digit classification, image classification, gesture recognition, and action recognition tasks, with energy consumption of only 8.07nJ, 71.26nJ, 729.3nJ, and 369.82nJ, respectively. These results emphasize the significance of DS-CIM and provide ideas for low-power inference on edge devices. Haotian Fu, Yulong Huang 0001, Tingran Chen, Chenyi Fu, Yue Zhou 0010, Shouzhong Peng, Zhirui Zong, Biao Pan, Bojun Cheng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Adaptive Processing for Video Streaming with Energy Constraint: A Multi-Agent Reinforcement Learning MethodabstractEdge computing is a highly promising technology that empowers mobile devices to offload video streaming tasks to edge servers, thereby improving the video stream analysis performance. However, most existing research on edge video streaming has failed to give adequate attention to the joint optimization of video streaming tasks with respect to dynamics, redundancy, and long-term energy constraints. To address this limitation, we propose a novel method based on a multi-agent reinforcement learning algorithm, which significantly enhances the performance of edge video stream analysis under long-term energy constraints. Specifically, our proposed method conducts video compression and offloading under long-term energy constraints to maximize the long-term rewards of video task processing. Experimental evaluations have demonstrated the convergence of the proposed method, which outperforms the baseline solutions, achieving higher long-term rewards. Haotian Fu, Shijing Yuan, Chentao Wu, Yuan Luo 0003, Jie Li 0002 |
GLOBECOM | 2 |
| 2023 | Performance Bounds for Model and Policy Transfer in Hidden-parameter MDPs
Haotian Fu, Jiayu Yao, Omer Gottesman, Finale Doshi-Velez, George Dimitri Konidaris |
ICLR | 1 |
| 2023 | Meta-learning Parameterized SkillsabstractWe propose a novel parameterized skill-learning algorithm that aims to learn transferable parameterized skills and synthesize them into a new action space that supports efficient learning in long-horizon tasks. We propose to leverage off-policy Meta-RL combined with a trajectory-centric smoothness term to learn a set of parameterized skills. Our agent can use these learned skills to construct a three-level hierarchical framework that models a Temporally-extended Parameterized Action Markov Decision Process. We empirically demonstrate that the proposed algorithms enable an agent to solve a set of highly difficult long-horizon (obstacle-course and robot manipulation) tasks. Haotian Fu, Shangqun Yu, Saket Tiwari, Michael L. Littman, George Dimitri Konidaris |
ICML | 1 |
| 2023 | TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasabstractEvent cameras have gained popularity in computer vision due to their data sparsity, high dynamic range, and low latency. As a bio-inspired sensor, event cameras generate sparse and asynchronous data, which is inherently incompatible with the traditional frame-based method. Alternatively, the point-based method can avoid additional modality transformation and naturally adapt to the sparsity of events. Still, it typically cannot reach a comparable accuracy as the frame-based method. We propose a lightweight and generalized point cloud network called TTPOINT which achieves competitive results even compared to the state-of-the-art (SOTA) frame-based method in action recognition tasks while only using 1.5 % of the computational resources. The model is adept at abstracting local and global geometry by hierarchy structure. By leveraging tensor-train compressed feature extractors, TTPOINT can be designed with minimal parameters and computational complexity. Additionally, we developed a straightforward downsampling algorithm to maintain the spatio-temporal feature. In the experiment, TTPOINT emerged as the SOTA method on three datasets while also attaining SOTA among point cloud methods on all five datasets. Moreover, by using the tensor-train decomposition method, the accuracy of the proposed TTPOINT is almost unaffected while compressing the parameter size by 55% in all five datasets. Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Renjing Xu, Bojun Cheng |
ACM Multimedia | 3 |
| 2023 | A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha |
Neural Networks | 15 |
| 2023 | In-Memory Computing Circuit Implementation of Complex-Valued Hopfield Neural Network for Efficient Portrait RestorationabstractComplex-valued neural networks have better optimization capabilities, stronger robustness, and richer characterization capabilities compared with real-valued neural networks, which has achieved good results in the field of portrait restoration. However, there is almost no circuit implementation of complex-valued neural networks. Based on this, this article proposes an in-memory computing circuit implementation of a complex-valued Hopfield neural network (CHNN) for the first time, which provides a highly accurate and efficient processing circuit for portrait restoration. First, a new memristive array is proposed, which can realize parallel complex-valued multiplication and complex-valued vector–matrix multiplication. On the basis, a CHNN circuit that can perform large-scale recursive computations is designed. Due to the characteristics of in-memory computation, the computation speed and robustness have been improved when realizing portrait restoration. Different portrait restoration scenarios can be realized based on the programmability of the memristive array. Pspice simulation results show that the recovery speed of CHNN can reach the level of 0.1 ms, and the accuracy can reach above 97.00%. Robustness analysis shows that the circuit can tolerate a certain degree of programming error and has strong anti-noise performance. Qinghui Hong, Haotian Fu, Yiyang Liu 0005, Jiliang Zhang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Model-based Lifelong Reinforcement Learning with Bayesian ExplorationabstractWe propose a model-based lifelong reinforcement-learning approach that estimates a hierarchical Bayesian posterior distilling the common structure shared across different tasks. The learned posterior combined with a sample-based Bayesian exploration procedure increases the sample efficiency of learning across a family of related tasks. We first derive an analysis of the relationship between the sample complexity and the initialization quality of the posterior in the finite MDP setting. We next scale the approach to continuous-state domains by introducing a Variational Bayesian Lifelong Reinforcement Learning algorithm that can be combined with recent model-based deep RL methods, and that exhibits backward transfer. Experimental results on several challenging domains show that our algorithms achieve both better forward and backward transfer performance than state-of-the-art lifelong RL methods. Haotian Fu, Shangqun Yu, Michael L. Littman, George Dimitri Konidaris |
NeurIPS | 1 |
| 2021 | Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive LearningabstractContext, the embedding of previous collected trajectories, is a powerful construct for Meta-Reinforcement Learning (Meta-RL) algorithms. By conditioning on an effective context, Meta-RL policies can easily generalize to new tasks within a few adaptation steps. We argue that improving the quality of context involves answering two questions: 1. How to train a compact and sufficient encoder that can embed the task-specific information contained in prior trajectories? 2. How to collect informative trajectories of which the corresponding context reflects the specification of tasks? To this end, we propose a novel Meta-RL framework called CCM (Contrastive learning augmented Context-based Meta-RL). We first focus on the contrastive nature behind different tasks and leverage it to train a compact and sufficient context encoder. Further, we train a separate exploration policy and theoretically derive a new information-gain-based objective which aims to collect informative trajectories in a few steps. Empirically, we evaluate our approaches on common benchmarks as well as several complex sparse-reward environments. The experimental results show that CCM outperforms state-of-the-art algorithms by addressing previously mentioned problems respectively. Haotian Fu, Hongyao Tang, Jianye Hao, Chen Chen 0077, Xidong Feng, Dong Li 0016, Wulong Liu |
AAAI | 1 |
| 2021 | Solving Non-Homogeneous Linear Ordinary Differential Equations Using Memristor-Capacitor CircuitabstractInhomogeneous linear ordinary differential equations (ODEs) and systems of ODEs can be solved in a variety of ways. However, hardware circuits that can perform the efficient analog computation to solve them are rarely in the literature. To address such problems, this paper proposes a general method of using a memristor-capacitor (M-C) circuit to solve inhomogeneous linear ODEs and systems of ODEs of any order in initial value problems. The M-C circuit can match the coefficients of the equations sought by adjusting the memristor resistance value according to the coefficient formula proposed in the paper, which has higher programmability. Then, some ODEs and systems of ODEs are given in the paper as examples to evaluate the proposed method. According to the comparison results based on MATLAB software simulation and the simulation based on OrCAD software, the designed M-C circuit has an effective improvement in speed and the accuracy exceeds 99.95% in software simulation. Based on practical verification, this paper gives the actual M-C circuit experiment based on PCB. Moreover, the proposed method can be used to quickly solve the object motion state in the spring mass damping system in actual engineering, and the accuracy can reach 99.98%. Haotian Fu, Qinghui Hong, Chunhua Wang 0001, Jingru Sun, Ya Li 0008 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | MGHRL: Meta Goal-Generation for Hierarchical Reinforcement Learning
Haotian Fu, Hongyao Tang, Jianye Hao, Wulong Liu, Chen Chen 0077 |
DAI | 1 |
| 2019 | Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action SpacesabstractDeep Reinforcement Learning (DRL) has been applied to address a variety of cooperative multi-agent problems with either discrete action spaces or continuous action spaces. However, to the best of our knowledge, no previous work has ever succeeded in applying DRL to multi-agent problems with discrete-continuous hybrid (or parameterized) action spaces which is very common in practice. Our work fills this gap by proposing two novel algorithms: Deep Multi-Agent Parameterized Q-Networks (Deep MAPQN) and Deep Multi-Agent Hierarchical Hybrid Q-Networks (Deep MAHHQN). We follow the centralized training but decentralized execution paradigm: different levels of communication between different agents are used to facilitate the training process, while each agent executes its policy independently based on local observations during execution. Our empirical results on several challenging tasks (simulated RoboCup Soccer and game Ghost Story) show that both Deep MAPQN and Deep MAHHQN are effective and significantly outperform existing independent deep parameterized Q-learning method. Haotian Fu, Hongyao Tang, Jianye Hao, Zihan Lei, Changjie Fan |
IJCAI | 1 |