EDBT 2026 Demo / reviewers in the wild / expert
Haowen Hou
dblp:248/9287
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 23% Video understanding and tracking · 20% Vision and language · 18% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization |
1.0 | 1 | 2026 | PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.0 | 1 | 2026 | Interleaved Latent Visual Reasoning with Selective Perceptual Modeling · ACL (1) 2026 |
Computer vision › Video understanding and tracking
long video understanding |
1.0 | 1 | 2026 | VideoPro: Adaptive Program Reasoning for Long Video Understanding · ACL (1) 2026 |
Computer vision › Vision and language
multimodal reasoning |
1.0 | 1 | 2026 | Interleaved Latent Visual Reasoning with Selective Perceptual Modeling · ACL (1) 2026 |
Machine learning › Reinforcement learning
policy optimization |
1.0 | 1 | 2026 | PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
1.0 | 1 | 2026 | PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization · AAAI 2026 |
Computer vision › 3D vision
ground reaction force estimation |
0.9 | 1 | 2025 | ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025 |
Computer vision › Video understanding and tracking › motion analysis
human motion analysis |
0.9 | 1 | 2025 | ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025 |
Machine learning › Generative modeling › motion generation
human motion imitation |
0.9 | 1 | 2025 | ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025 |
Robotics › Motion planning and robot control › robot dynamics
inverse dynamics |
0.9 | 1 | 2025 | ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025 |
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding › human-centric scene understanding
human-scene interaction |
0.8 | 1 | 2024 | Revisit Human-Scene Interaction via Space Occupancy · ECCV (50) 2024 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2026 | Interleaved Latent Visual Reasoning with Selective Perceptual Modeling · ACL (1) 2026 |
Computer vision › 3D vision
physical simulation |
0.3 | 1 | 2025 | ImDy: Human Inverse Dynamics from Imitated Observations · ICLR 2025 |
Computer vision › 3D vision › 3d motion analysis
human pose and motion |
0.2 | 1 | 2024 | Revisit Human-Scene Interaction via Space Occupancy · ECCV (50) 2024 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 1.0reward smoothing · 1.0program synthesis · 1.0prioritized experience replay · 1.0momentum teacher · 1.0large language model · 1.0knowledge distillation · 1.0exponential moving average · 1.0physics simulation · 0.9motion imitation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy OptimizationabstractReinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to maintain competitive performance. We propose PSPO (Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization), a lightweight yet effective enhancement to GRPO that improves training stability and sample efficiency through two complementary techniques. First, we introduce an experience-weighted reward smoothing mechanism, which uses exponential moving averages to track group-level reward statistics for each prompt. This enables more stable advantage estimation across training steps without storing entire trajectories, allowing the model to capture historical reward trends in a lightweight and memory-efficient manner. Second, we adopt a prompt-level prioritized sampling strategy, which is an online data selection method inspired by prioritized experience replay. It dynamically emphasizes higher-impact prompts based on their relative advantages, thereby improving data efficiency. Experiments on multiple mathematical reasoning benchmarks and models show that PSPO achieves comparable or better accuracy than GRPO, while significantly accelerating convergence, and maintaining low computational and memory overhead. Ying He 0006, Haowen Hou, Ruichong Zhang, Nianbo Zeng, Yulin Peng, Jiongfeng Fang, F. Richard Yu |
AAAI | 3 |
| 2026 | Interleaved Latent Visual Reasoning with Selective Perceptual ModelingabstractInterleaved reasoning paradigms enhance Multimodal Large Language Models (MLLMs) with visual feedback but are hindered by the prohibitive computational cost of re-encoding pixel-dense images. A promising alternative, latent visual reasoning, circumvents this bottleneck yet faces limitations: methods either fail to capture intermediate state evolution due to single-step, non-interleaved structures, or sacrifice precise perceptual modeling by over-compressing features. We introduce Interleaved Latent Visual Reasoning (ILVR), a framework that unifies dynamic state evolution with precise perceptual modeling. ILVR interleaves textual generation with latent visual representations that act as specific, evolving cues for subsequent reasoning. Specifically, we employ a self-supervision strategy where a momentum teacher model selectively distills relevant features from ground-truth intermediate images into sparse supervision targets. This adaptive selection mechanism guides the model to autonomously generate context-aware visual signals. Extensive experiments on multimodal reasoning benchmarks demonstrate that ILVR outperforms existing approaches, effectively bridging the gap between fine-grained perception and sequential multimodal reasoning. The code is available at https://github.com/XD111ds/ILVR. Haowen Hou, Zhongyu Wei |
ACL (1) | 5 |
| 2026 | VideoPro: Adaptive Program Reasoning for Long Video UnderstandingabstractChenglin Li, Feng Han, Yikun Wang, Ruilin Li, Shuai Dong, Haowen Hou, Haitao Li, Qianglong Chen, Feng Tao, Jingqi Tong, Yin Zhang, Jiaqi Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yikun Wang 0001, Haowen Hou, Qianglong Chen, Jingqi Tong, Yin Zhang 0006 |
ACL (1) | 6 |
| 2025 | VisualRWKV: Exploring Recurrent Neural Networks for Visual Language ModelsabstractVisual Language Models (VLMs) have rapidly progressed with the recent success of large language models. However, there have been few attempts to incorporate efficient linear Recurrent Neural Networks (RNNs) architectures into VLMs. In this study, we introduce VisualRWKV, the first application of a linear RNN model to multimodal learning tasks, leveraging the pre-trained RWKV language model. We propose a data-dependent recurrence and sandwich prompts to enhance our modeling capabilities, along with a 2D image scanning mechanism to enrich the processing of visual sequences. Extensive experiments demonstrate that VisualRWKV achieves competitive performance compared to Transformer-based models like LLaVA-1.5 on various benchmarks. Compared to LLaVA-1.5, VisualRWKV has a speed advantage of 3.98 times and can save 54% of GPU memory when reaching an inference length of 24K tokens. To facilitate further research and analysis, we have made the checkpoints and the associated code publicly accessible at the following GitHub repository: https://github.com/howard-hou/VisualRWKV. Haowen Hou, Peigen Zeng, Fei Ma 0006, F. Richard Yu |
COLING | 1 |
| 2025 | ImDy: Human Inverse Dynamics from Imitated ObservationsabstractInverse dynamics (ID), which aims at reproducing the driven torques from human kinematic observations, has been a critical tool for gait analysis. However, it is hindered from wider application to general motion due to its limited scalability. Conventional optimization-based ID requires expensive laboratory setups, restricting its availability. To alleviate this problem, we propose to exploit the recently progressive human motion imitation algorithms to learn human inverse dynamics in a data-driven manner. The key insight is that the human ID knowledge is implicitly possessed by motion imitators, though not directly applicable. In light of this, we devise an efficient data collection pipeline with state-of-the-art motion imitation algorithms and physics simulators, resulting in a large-scale human inverse dynamics benchmark as Imitated Dynamics (ImDy). ImDy contains over 150 hours of motion with joint torque and full-body ground reaction force data. With ImDy, we train a data-driven human inverse dynamics solver ImDyS(olver) in a fully supervised manner, which conducts ID and ground reaction force estimation simultaneously. Experiments on ImDy and real-world data demonstrate the impressive competency of ImDyS in human inverse dynamics and ground reaction force estimation. Moreover, the potential of ImDy(-S) as a fundamental motion analysis tool is exhibited with downstream applications. The project page is https://foruck.github.io/ImDy. Xinpeng Liu 0002, Junxuan Liang, Zili Lin, Haowen Hou, Yong-Lu Li 0001, Cewu Lu |
ICLR | 4 |
| 2025 | RWKV-UI: UI Understanding with Enhanced Perception and ReasoningabstractExisting Visual Language Modelsoften struggle with information loss and limited reasoning abilities when handling high-resolution web interfaces that combine complex visual, textual, and interactive elements. These challenges are particularly evident in tasks requiring webpage layout comprehension and multi-step interactive reasoning. To address these challenges, we propose RWKV-UI, a Visual Language Model based on the RWKV architecture, specifically designed to handle high-resolution UI images. During model training, we introduce layout detection as a visual prompt to help the model better understand the webpage layout structures. Additionally, we design a visual prompt based on the Chain-of-Thought(CoT) mechanism, which enhances the model’s ability to understand and reason about webpage content through reasoning chains. Experimental results show that RWKV-UI demonstrates significant performance improvements in high-resolution UI understanding and interactive reasoning tasks. Haowen Hou |
ICME | 2 |
| 2025 | MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video GroundingabstractIn untrimmed video tasks, identifying temporal boundaries in videos is crucial for temporal video grounding. With the emergence of multimodal large language models (MLLMs), recent studies have focused on endowing these models with the capability of temporal perception in untrimmed videos. To address the challenge, in this paper, we introduce a multimodal large language model named MLLM-TA with precise temporal perception to obtain temporal attention. Unlike the traditional MLLMs, answering temporal questions through one or two words related to temporal information, we leverage the text description proficiency of MLLMs to acquire video temporal attention with description. Specifically, we design a dual temporal-aware generative branches aimed at the visual space of the entire video and the textual space of global descriptions, simultaneously generating mutually supervised consistent temporal attention, thereby enhancing the video temporal perception capabilities of MLLMs. Finally, we evaluate our approach on both video grounding task and highlight detection task on three popular benchmarks, including Charades-STA, ActivityNet Captions and QVHighlights. The extensive results show that our MLLM-TA significantly outperforms previous approaches both on zero-shot and supervised setting, achieving state-of-the-art performance. Yi Liu 0081, Haowen Hou, Fei Ma 0006, Shiguang Ni, F. Richard Yu |
IEEE Signal Process. Lett. | 2 |
| 2024 | Revisit Human-Scene Interaction via Space Occupancy
Xinpeng Liu 0002, Haowen Hou, Yanchao Yang 0001, Yong-Lu Li 0001, Cewu Lu |
ECCV (50) | 2 |
| 2024 | BagFormer: Better cross-modal retrieval via bag-wise interaction
Haowen Hou, Xiaopeng Yan, Yigeng Zhang |
Eng. Appl. Artif. Intell. | 1 |