EDBT 2026 Demo / reviewers in the wild / expert
Zhi Hou
dblp:128/0389
· DBLP profile ↗
13ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-2990-505XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 7 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action PolicyabstractVision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces. Although training data are collected from diverse camera perspectives, the models typically predict end-effector poses within the robot base coordinate frame, resulting in spatial inconsistencies. To mitigate this limitation, we introduce the Observation-Centric VLA (OC-VLA) framework, which grounds action predictions directly in the camera observation space. Leveraging the camera’s extrinsic calibration matrix, OC-VLA transforms end-effector poses from the robot base coordinate system into the camera coordinate system, thereby unifying prediction targets across heterogeneous viewpoints. This lightweight, plug-and-play strategy ensures robust alignment between perception and action, substantially improving model resilience to camera viewpoint variations. The proposed approach is readily compatible with existing VLA architectures, requiring no substantial modifications. Comprehensive evaluations on both simulated and real-world robotic manipulation tasks demonstrate that OC-VLA accelerates convergence, enhances task success rates, and improves cross-view generalization. Haonan Duan 0001, Haoran Hao 0003, Yu Qiao 0001, Jifeng Dai, Zhi Hou |
AAAI | 6 |
| 2025 | Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyabstractWhile recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or continuous actions constrains adaptability to heterogeneous action spaces. We present Dita, a scalable framework that leverages Transformer architectures to directly denoise continuous action sequences through a unified multimodal diffusion process. Departing from prior methods that condition denoising on fused embeddings via shallow networks, Dita employs in-context conditioning -- enabling fine-grained alignment between denoised actions and raw visual tokens from historical observations. This design explicitly models action deltas and environmental nuances. By scaling the diffusion action denoiser alongside the Transformer's scalability, Dita effectively integrates cross-embodiment datasets across diverse camera perspectives, observation scenes, tasks, and action spaces. Such synergy enhances robustness against various variances and facilitates the successful execution of long-horizon tasks. Evaluations across extensive benchmarks demonstrate state-of-the-art or comparative performance in simulation. Notably, Dita achieves robust real-world adaptation to environmental variances and complex long-horizon tasks through 10-shot finetuning, using only third-person camera inputs. The architecture establishes a versatile, lightweight and open-source baseline for generalist robot policy learning. Project Page: https://robodita.github.io. Zhi Hou, Yuwen Xiong, Haonan Duan 0001, Hengjun Pu, Ronglei Tong, Chengyang Zhao, Xizhou Zhu, Yu Qiao 0001, Jifeng Dai, Yuntao Chen |
ICCV | 1 |
| 2025 | Timeformer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust ReconstructionabstractDynamic scene reconstruction is a long-term challenge in 3D vision. Recent methods extend 3D Gaussian Splatting to dynamic scenes via additional deformation fields and apply explicit constraints like motion flow to guide the deformation. However, they learn motion changes from individual timestamps independently, making it challenging to reconstruct complex scenes, particularly when dealing with violent movement, extreme-shaped geometries, or reflective surfaces. To address the above issue, we design a plug-and-play module called TimeFormer to enable existing deformable 3D Gaussians reconstruction methods with the ability to implicitly model motion patterns from a learning perspective. Specifically, TimeFormer includes a Cross-Temporal Transformer Encoder, which adaptively learns the temporal relationships of deformable 3D Gaussians. Furthermore, we propose a two-stream optimization strategy that transfers the motion knowledge learned from TimeFormer to the base stream during the training phase. This allows us to remove TimeFormer during inference, thereby preserving the original rendering speed. Extensive experiments in the multi-view and monocular dynamic scenes validate qualitative and quantitative improvement brought by TimeFormer. Project Page: https://patrickddj.github.io/TimeFormer/ Dadong Jiang, Zhi Hou, Zhihui Ke, Xianghui Yang, Xiaobo Zhou 0003, Tie Qiu 0001 |
ICCV | 2 |
| 2025 | Learning to Explore Sample RelationshipsabstractDespite the great success achieved, deep learning technologies usually suffer from data scarcity issues in real-world applications, where existing methods mainly explore sample relationships in a vanilla way from the perspectives of either the input or the loss function. In this paper, we propose a batch transformer module, BatchFormerV1, to equip deep neural networks themselves with the abilities to explore sample relationships in a learnable way. Basically, the proposed method enables data collaboration, e.g., head-class samples will also contribute to the learning of tail classes. Considering that exploring instance-level relationships has very limited impacts on dense prediction, we generalize and refer to the proposed module as BatchFormerV2, which further enables exploring sample relationships for pixel-/patch-level dense representations. In addition, to address the train-test inconsistency where a mini-batch of data samples are neither necessary nor desirable during inference, we also devise a two-stream training pipeline, i.e., a shared model is first jointly optimized with and without BatchFormerV2 which is then removed during testing. The proposed module is plug-and-play without requiring any extra inference cost. Lastly, we evaluate the proposed method on over ten popular datasets, including 1) different data scarcity settings such as long-tailed recognition, zero-shot learning, domain generalization, and contrastive learning; and 2) different visual recognition tasks ranging from image classification to object detection and panoptic segmentation. Zhi Hou, Baosheng Yu, Yibing Zhan, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Large Point-to-Gaussian Model for Image-to-3D Generation
Longfei Lu, Huachen Gao, Tao Dai 0001, Yaohua Zha, Zhi Hou, Junta Wu, Shutao Xia |
ACM Multimedia | 5 |
| 2023 | AniPixel: Towards Animatable Pixel-Aligned Human AvatarabstractAlthough human reconstruction typically results in human-specific avatars, recent 3D scene reconstruction techniques utilizing pixel-aligned features show promise in generalizing to new scenes. Applying these techniques to human avatar reconstruction can result in a volumetric avatar with generalizability but limited animatability due to rendering only being possible for static representations. In this paper, we propose AniPixel, a novel animatable and generalizable human avatar reconstruction method that leverages pixel-aligned features for body geometry prediction and RGB color blending. Technically, to align the canonical space with the target space and the observation space, we propose a bidirectional neural skinning field based on skeleton-driven deformation to establish the target-to-canonical and canonical-to-observation correspondences. Then, we disentangle the canonical body geometry into a normalized neutral-sized body and a subject-specific residual for better generalizability. As the geometry and appearance are closely related, we introduce pixel-aligned features to facilitate the body geometry prediction and detailed surface normals to reinforce the RGB color blending. We also devise a pose-dependent and view direction-related shading module to represent the local illumination variance. Experiments show that AniPixel renders comparable novel views while delivering better novel pose animation results than state-of-the-art methods. Code will be released at AniPixel. https://github.com/loong8888/AniPixel. Jinlong Fan 0001, Jing Zhang 0037, Zhi Hou, Dacheng Tao |
ACM Multimedia | 3 |
| 2022 | BatchFormer: Learning to Explore Sample Relationships for Robust Representation LearningabstractDespite the success of deep neural networks, there are still many challenges in deep representation learning due to the data scarcity issues such as data imbalance, unseen distribution, and domain shift. To address the above-mentioned issues, a variety of methods have been devised to explore the sample relationships in a vanilla way (i.e., from the perspectives of either the input or the loss function), failing to explore the internal structure of deep neural networks for learning with sample relationships. Inspired by this, we propose to enable deep neural networks themselves with the ability to learn the sample relationships from each mini-batch. Specifically, we introduce a batch transformer module or BatchFormer, which is then applied into the batch dimension of each mini-batch to implicitly explore sample relationships during training. By doing this, the proposed method enables the collaboration of different samples, e.g., the head-class samples can also contribute to the learning of the tail classes for long-tailed recognition. Furthermore, to mitigate the gap between training and testing, we share the classifier between with or without the BatchFormer during training, which can thus be removed during testing. We perform extensive experiments on over ten datasets and the proposed method achieves significant improvements on different data scarcity applications without any bells and whistles, including the tasks of long-tailed recognition, compositional zero-shot learning, domain generalization, and contrastive learning. Code is made publicly available at https://github.com/zhihou7/BatchFormer. Zhi Hou, Baosheng Yu, Dacheng Tao |
CVPR | 1 |
| 2022 | Discovering Human-Object Interaction Concepts via Self-Compositional Learning
Zhi Hou, Baosheng Yu, Dacheng Tao |
ECCV (27) | 1 |
| 2022 | Load-Adaptive and Energy-Efficient Topology Control in LEO Mega-Constellation NetworksabstractThe Low-Earth-Orbit (LEO) mega-constellation networks, by providing low-latency and high-speed communications, are becoming indispensable infrastructures for the future six-generation (6G) architecture. Consequently, the topology, with thousands of satellites equipped with batteries of limited life, has to be adaptively controlled with high energy efficiency. However, existing work lacks the joint consideration of energy efficiency and load adaptation. In this paper, we first propose the line-of-sight condition to determine the candidate ISL set. Next, we model the energy consumption of the LEO mega-constellation networks. Along this direction, we formulate the Load-Adaptive and Energy-Efficient (LAEE) topology control problem in LEO mega-constellation networks and prove its NP-hardness. Finally, we propose the Amortized Energy based Topology Control (AETC) algorithm to solve the LAEE problem, with good adaptation to the fluctuating load and guarantees connectivities between any two satellites. Extensive simulation results demonstrate that the AETC algorithm outperforms related schemes in terms of energy consumption and results in good topology stability. Long Chen 0025, Feilong Tang 0001, Linghe Kong, Rui Li 0098, Zhi Hou, Jiacheng Liu 0001, Xu Li 0012, Song Guo 0001 |
GLOBECOM | 5 |
| 2021 | Affordance Transfer Learning for Human-Object Interaction DetectionabstractReasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we introduce an affordance transfer learning approach to jointly detect HOIs with novel object and recognize affordances. Specifically, HOI representations can be decoupled into a combination of affordance and object representations, making it possible to compose novel interactions by combining affordance representations and novel object representations from additional images, i.e. transferring the affordance to novel objects. With the proposed affordance transfer learning, the model is also capable of inferring the affordances of novel objects from known affordance representations. The proposed method can thus be used to 1) improve the performance of HOI detection, especially for the HOIs with unseen objects; and 2) infer the affordances of novel objects. Experimental results on two datasets, HICO-DET and HOI-COCO (from V-COCO), demonstrate significant improvements over recent state-of-the-art methods for HOI detection and object affordance detection. Code is available at https://github.com/zhihou7/HOI-CL. Zhi Hou, Baosheng Yu, Yu Qiao 0001, Xiaojiang Peng, Dacheng Tao |
CVPR | 1 |
| 2021 | Detecting Human-Object Interaction via Fabricated Compositional LearningabstractHuman-Object Interaction (HOI) detection, inferring the relationships between human and objects from images/videos, is a fundamental task for high-level scene understanding. However, HOI detection usually suffers from the open long-tailed nature of interactions with objects, while human has extremely powerful compositional perception ability to cognize rare or unseen HOI samples. Inspired by this, we devise a novel HOI compositional learning framework, termed as Fabricated Compositional Learning (FCL), to address the problem of open long-tailed HOI detection. Specifically, we introduce an object fabricator to generate effective object representations, and then combine verbs and fabricated objects to compose new HOI samples. With the proposed object fabricator, we are able to generate large-scale HOI samples for rare and unseen categories to alleviate the open long-tailed issues in HOI detection. Extensive experiments on the most popular HOI detection dataset, HICO-DET, demonstrate the effectiveness of the proposed method for imbalanced HOI detection and significantly improve the state-of-the-art performance on rare and unseen HOI categories. Code is available at https://github.com/zhihou7/HOI-CL. Zhi Hou, Baosheng Yu, Yu Qiao 0001, Xiaojiang Peng, Dacheng Tao |
CVPR | 1 |
| 2020 | Visual Compositional Learning for Human-Object Interaction Detection
Zhi Hou, Xiaojiang Peng, Yu Qiao 0001, Dacheng Tao |
ECCV (15) | 1 |
| 2019 | RTCRelief-F: an effective clustering and ordering-based ensemble pruning algorithm for facial expression recognition
Danyang Li 0004, Guihua Wen, Zhi Hou, Er-Yang Huan |
Knowl. Inf. Syst. | 3 |