Haowen Hou

dblp:248/9287 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 23% Video understanding and tracking · 20% Vision and language · 18%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.012026
PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
1.012026
PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.012026
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling · ACL (1) 2026
Computer vision › Video understanding and tracking
long video understanding
1.012026
VideoPro: Adaptive Program Reasoning for Long Video Understanding · ACL (1) 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling · ACL (1) 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
1.012026
PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026
Computer vision › 3D vision
ground reaction force estimation
0.912025
ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025
Computer vision › Video understanding and tracking › motion analysis
human motion analysis
0.912025
ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025
Machine learning › Generative modeling › motion generation
human motion imitation
0.912025
ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025
Robotics › Motion planning and robot control › robot dynamics
inverse dynamics
0.912025
ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding › human-centric scene understanding
human-scene interaction
0.812024
Revisit Human-Scene Interaction via Space Occupancy · ECCV (50) 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312026
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling · ACL (1) 2026
Computer vision › 3D vision
physical simulation
0.312025
ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025
Computer vision › 3D vision › 3d motion analysis
human pose and motion
0.212024
Revisit Human-Scene Interaction via Space Occupancy · ECCV (50) 2024

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.0reward smoothing · 1.0program synthesis · 1.0prioritized experience replay · 1.0momentum teacher · 1.0large language model · 1.0knowledge distillation · 1.0exponential moving average · 1.0physics simulation · 0.9motion imitation · 0.9
YearPublicationVenuePosition
2026 PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization
abstract
Reinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to maintain competitive performance. We propose PSPO (Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization), a lightweight yet effective enhancement to GRPO that improves training stability and sample efficiency through two complementary techniques. First, we introduce an experience-weighted reward smoothing mechanism, which uses exponential moving averages to track group-level reward statistics for each prompt. This enables more stable advantage estimation across training steps without storing entire trajectories, allowing the model to capture historical reward trends in a lightweight and memory-efficient manner. Second, we adopt a prompt-level prioritized sampling strategy, which is an online data selection method inspired by prioritized experience replay. It dynamically emphasizes higher-impact prompts based on their relative advantages, thereby improving data efficiency. Experiments on multiple mathematical reasoning benchmarks and models show that PSPO achieves comparable or better accuracy than GRPO, while significantly accelerating convergence, and maintaining low computational and memory overhead.
Ying He 0006, Haowen Hou, Ruichong Zhang, Nianbo Zeng, Yulin Peng, Jiongfeng Fang, F. Richard Yu
AAAI3
2026 Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
abstract
Interleaved reasoning paradigms enhance Multimodal Large Language Models (MLLMs) with visual feedback but are hindered by the prohibitive computational cost of re-encoding pixel-dense images. A promising alternative, latent visual reasoning, circumvents this bottleneck yet faces limitations: methods either fail to capture intermediate state evolution due to single-step, non-interleaved structures, or sacrifice precise perceptual modeling by over-compressing features. We introduce Interleaved Latent Visual Reasoning (ILVR), a framework that unifies dynamic state evolution with precise perceptual modeling. ILVR interleaves textual generation with latent visual representations that act as specific, evolving cues for subsequent reasoning. Specifically, we employ a self-supervision strategy where a momentum teacher model selectively distills relevant features from ground-truth intermediate images into sparse supervision targets. This adaptive selection mechanism guides the model to autonomously generate context-aware visual signals. Extensive experiments on multimodal reasoning benchmarks demonstrate that ILVR outperforms existing approaches, effectively bridging the gap between fine-grained perception and sequential multimodal reasoning. The code is available at https://github.com/XD111ds/ILVR.
Haowen Hou, Zhongyu Wei
ACL (1)5
2026 VideoPro: Adaptive Program Reasoning for Long Video Understanding
abstract
Chenglin Li, Feng Han, Yikun Wang, Ruilin Li, Shuai Dong, Haowen Hou, Haitao Li, Qianglong Chen, Feng Tao, Jingqi Tong, Yin Zhang, Jiaqi Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yikun Wang 0001, Haowen Hou, Qianglong Chen, Jingqi Tong, Yin Zhang 0006
ACL (1)6
2025 VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
abstract
Visual Language Models (VLMs) have rapidly progressed with the recent success of large language models. However, there have been few attempts to incorporate efficient linear Recurrent Neural Networks (RNNs) architectures into VLMs. In this study, we introduce VisualRWKV, the first application of a linear RNN model to multimodal learning tasks, leveraging the pre-trained RWKV language model. We propose a data-dependent recurrence and sandwich prompts to enhance our modeling capabilities, along with a 2D image scanning mechanism to enrich the processing of visual sequences. Extensive experiments demonstrate that VisualRWKV achieves competitive performance compared to Transformer-based models like LLaVA-1.5 on various benchmarks. Compared to LLaVA-1.5, VisualRWKV has a speed advantage of 3.98 times and can save 54% of GPU memory when reaching an inference length of 24K tokens. To facilitate further research and analysis, we have made the checkpoints and the associated code publicly accessible at the following GitHub repository: https://github.com/howard-hou/VisualRWKV.
Haowen Hou, Peigen Zeng, Fei Ma 0006, F. Richard Yu
COLING1
2025 ImDy: Human Inverse Dynamics from Imitated Observations
abstract
Inverse dynamics (ID), which aims at reproducing the driven torques from human kinematic observations, has been a critical tool for gait analysis. However, it is hindered from wider application to general motion due to its limited scalability. Conventional optimization-based ID requires expensive laboratory setups, restricting its availability. To alleviate this problem, we propose to exploit the recently progressive human motion imitation algorithms to learn human inverse dynamics in a data-driven manner. The key insight is that the human ID knowledge is implicitly possessed by motion imitators, though not directly applicable. In light of this, we devise an efficient data collection pipeline with state-of-the-art motion imitation algorithms and physics simulators, resulting in a large-scale human inverse dynamics benchmark as Imitated Dynamics (ImDy). ImDy contains over 150 hours of motion with joint torque and full-body ground reaction force data. With ImDy, we train a data-driven human inverse dynamics solver ImDyS(olver) in a fully supervised manner, which conducts ID and ground reaction force estimation simultaneously. Experiments on ImDy and real-world data demonstrate the impressive competency of ImDyS in human inverse dynamics and ground reaction force estimation. Moreover, the potential of ImDy(-S) as a fundamental motion analysis tool is exhibited with downstream applications. The project page is https://foruck.github.io/ImDy.
Xinpeng Liu 0002, Junxuan Liang, Zili Lin, Haowen Hou, Yong-Lu Li 0001, Cewu Lu
ICLR4
2025 RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
abstract
Existing Visual Language Modelsoften struggle with information loss and limited reasoning abilities when handling high-resolution web interfaces that combine complex visual, textual, and interactive elements. These challenges are particularly evident in tasks requiring webpage layout comprehension and multi-step interactive reasoning. To address these challenges, we propose RWKV-UI, a Visual Language Model based on the RWKV architecture, specifically designed to handle high-resolution UI images. During model training, we introduce layout detection as a visual prompt to help the model better understand the webpage layout structures. Additionally, we design a visual prompt based on the Chain-of-Thought(CoT) mechanism, which enhances the model’s ability to understand and reason about webpage content through reasoning chains. Experimental results show that RWKV-UI demonstrates significant performance improvements in high-resolution UI understanding and interactive reasoning tasks.
Haowen Hou
ICME2
2025 MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video Grounding
abstract
In untrimmed video tasks, identifying temporal boundaries in videos is crucial for temporal video grounding. With the emergence of multimodal large language models (MLLMs), recent studies have focused on endowing these models with the capability of temporal perception in untrimmed videos. To address the challenge, in this paper, we introduce a multimodal large language model named MLLM-TA with precise temporal perception to obtain temporal attention. Unlike the traditional MLLMs, answering temporal questions through one or two words related to temporal information, we leverage the text description proficiency of MLLMs to acquire video temporal attention with description. Specifically, we design a dual temporal-aware generative branches aimed at the visual space of the entire video and the textual space of global descriptions, simultaneously generating mutually supervised consistent temporal attention, thereby enhancing the video temporal perception capabilities of MLLMs. Finally, we evaluate our approach on both video grounding task and highlight detection task on three popular benchmarks, including Charades-STA, ActivityNet Captions and QVHighlights. The extensive results show that our MLLM-TA significantly outperforms previous approaches both on zero-shot and supervised setting, achieving state-of-the-art performance.
Yi Liu 0081, Haowen Hou, Fei Ma 0006, Shiguang Ni, F. Richard Yu
IEEE Signal Process. Lett.2
2024 Revisit Human-Scene Interaction via Space Occupancy
Xinpeng Liu 0002, Haowen Hou, Yanchao Yang 0001, Yong-Lu Li 0001, Cewu Lu
ECCV (50)2
2024 BagFormer: Better cross-modal retrieval via bag-wise interaction
Haowen Hou, Xiaopeng Yan, Yigeng Zhang
Eng. Appl. Artif. Intell.1