VLDB 2026 Research / reviewers in the wild / expert
Liang Xu 0012
dblp:54/3420-12
· DBLP profile ↗
11ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-6441-4443ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu 0012, Chengqun Yang, Zili Lin, Fei Xu 0008, Congsheng Xu, Yiyi Zhang 0002, Jie Qin 0004, Xingdong Sheng, Yunhui Liu 0006, Xin Jin 0014, Yichao Yan, Wenjun Zeng 0001, Xiaokang Yang 0001 |
ICCV | 1 |
| 2024 | Inter-X: Towards Versatile Human-Human Interaction AnalysisabstractThe analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions, lack of hand gestures and fine- grained textual descriptions. To better perceive and generate human-human interactions, we propose Inter-X, a currently largest human-human interaction dataset with accurate body movements and diverse interaction patterns, together with detailed hand gestures. The dataset includes Liang Xu 0012, Xintao Lv, Yichao Yan, Xin Jin 0014, Shuwen Wu, Congsheng Xu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, Yunhui Liu 0006, Wenjun Zeng 0001, Xiaokang Yang 0001 |
CVPR | 1 |
| 2024 | ReGenNet: Towards Human Action-Reaction SynthesisabstractHumans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects, while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less explored. Human-human interactions can be regarded as asymmetric with actors and reactors in atomic interaction periods. In this paper, we compre-hensively analyze the asymmetric, dynamic, synchronous, and detailed nature of human-human interactions and propose the first multi-setting human action-reaction synthe-sis benchmark to generate human reactions conditioned on given human actions. To begin with, we propose to an-notate the actor-reactor order of the interaction sequences for the NTU120, InterHuman, and Chi3D datasets. Based on them, a diffusion-based generative model with a Trans-former decoder architecture called ReGenNet together with an explicit distance-based interaction loss is proposed to predict human reactions in an online manner, where the future states of actors are unavailable to reactors. Quantitative and qualitative results show that our method can gener-ate instant and plausible human reactions compared to the baselines, and can generalize to unseen actor motions and viewpoint changes. Liang Xu 0012, Yizhou Zhou, Yichao Yan, Xin Jin 0014, Wenhan Zhu, Fengyun Rao, Xiaokang Yang 0001, Wenjun Zeng 0001 |
CVPR | 1 |
| 2024 | HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects
Xintao Lv, Liang Xu 0012, Yichao Yan, Xin Jin 0014, Congsheng Xu, Shuwen Wu, Lincheng Li, Mengxiao Bi, Wenjun Zeng 0001, Xiaokang Yang 0001 |
ECCV (4) | 2 |
| 2023 | ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion GenerationabstractWe present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equipped with a Gaussian Process latent prior. Such a design combines the strong spatio-temporal representation capacity of Transformer, superiority in generative modeling of GAN, and inherent temporal correlations from the latent prior. Furthermore, ActFormer can be naturally extended to multi-person motions by alternately modeling temporal correlations and human interactions with Transformer encoders. To further facilitate research on multi-person motion generation, we introduce a new synthetic dataset of complex multi-person combat behaviors. Extensive experiments on NTU-13, NTU RGB+D 120, BABEL and the proposed combat dataset show that our method can adapt to various human motion representations and achieve superior performance over the state-of-the-art methods on both single-person and multi-person motion generation tasks, demonstrating a promising step towards a general human motion generator. The project website can be found at https://liangxuy.github.io/actformer/. Liang Xu 0012, Jing Su 0005, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001, Wei Wu 0021 |
ICCV | 1 |
| 2023 | HAKE: A Knowledge Engine Foundation for Human Activity UnderstandingabstractHuman activity understanding is of widespread interest in artificial intelligence and spans diverse applications like health care and behavior analysis. Although there have been advances with deep learning, it remains challenging. The object recognition-like solutions usually try to map pixels to semantics directly, but activity patterns are much different from object patterns, thus hindering another success. In this article, we propose a novel paradigm to reformulate this task in two-stage: first mapping pixels to an intermediate space spanned by atomic activity primitives, then programming detected primitives with interpretable logic rules to infer semantics. To afford a representative primitive space, we build a knowledge base including 26+ M primitive labels and logic rules from human priors or automatic discovering. Our framework, Human Activity Knowledge Engine (HAKE), exhibits superior generalization ability and performance upon canonical methods on challenging benchmarks. Code and data are available at http://hake-mvig.cn/. Yong-Lu Li 0001, Xinpeng Liu 0002, Yizhuo Li 0001, Zuoyu Qiu, Liang Xu 0012, Haoshu Fang, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Skeleton-Based Mutually Assisted Interacted Object Localization and Human Action RecognitionabstractSkeleton data carries valuable motion information and is widely explored in human action recognition. However, not only the motion information but also the interaction with the environment provides discriminative cues to recognize the action of persons. In this paper, we propose a joint learning framework for mutually assisted “interacted object localization” and “human action recognition” based on skeleton data. The two tasks are serialized together and collaborate to promote each other, where preliminary action type derived from skeleton alone helps improve interacted object localization, which in turn provides valuable cues for the final human action recognition. Besides, we explore the temporal consistency of interacted object as constraint to better localize the interacted object with the absence of ground-truth labels. Extensive experiments on the datasets of SYSU-3D, NTU60 RGB+D, Northwestern-UCLA and UAV-Human show that our method achieves the best or competitive performance with the state-of-the-art methods for human action recognition. Visualization results show that our method can also provide reasonable interacted object localization results. Liang Xu 0012, Cuiling Lan, Wenjun Zeng 0001, Cewu Lu |
IEEE Trans. Multim. | 1 |
| 2022 | Transferable Interactiveness Knowledge for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) Detection is an important problem to understand how humans interact with objects. In this paper, we explore Interactiveness Knowledge which indicates whether human and object interact with each other or not. We found that interactiveness knowledge can be learned across HOI datasets and alleviate the gap between diverse HOI category settings. Our core idea is to exploit an Interactiveness Network to learn the general interactiveness knowledge from multiple HOI datasets and perform Non-Interaction Suppression before HOI classification in inference. On account of the generalization of interactiveness, interactiveness network is a transferable knowledge learner and can be cooperated with any HOI detection models to achieve desirable results. We utilize the human instance and body part features together to learn the interactiveness in hierarchical paradigm, i.e., instance-level and body part-level interactivenesses. Thereafter, a consistency task is proposed to guide the learning and extract deeper interactive visual clues. We extensively evaluate the proposed method on HICO-DET, V-COCO, and a newly constructed HAKE-HOI dataset. With the learned interactiveness, our method outperforms state-of-the-art HOI detection methods, verifying its efficacy and flexibility. Code is available at https://github.com/DirtyHarryLYL/Transferable-Interactiveness-Network. Yong-Lu Li 0001, Xinpeng Liu 0002, Xijie Huang, Liang Xu 0012, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | PAL-Net: Predicate-Aware Learning Network for Visual Relationship RecognitionabstractVisual relationship recognition is essential for deeper scene understanding. It poses to recognize 〈subject-predicate-object〉 triplets between object pairs. Previous methods usually treat vastly different predicates equally and neglect the subtle differences between predicates. In this paper, we propose a novel and concise perspective called "predicate-aware learning network (PAL-Net)" for visual relationship recognition. "Predicate-aware" means that we take predicates as a condition in a task-driven manner. Our PAL-Net consists of two key modules: i) a predicate-guided regularization module designed to learn more differentiated representations for various predicates; ii) a predicate-aware contextual modeling module to integrate the efficacy of contextual objects for different predicates respectively. Extensive experiments on VRD and Visual Genome dataset yield remarkable performance gains, verifying the effectiveness of PAL-Net. Besides, PAL-Net also shows good applicability and achieves substantial improvement for human-object interaction detection. Liang Xu 0012, Yong-Lu Li 0001, Yan Hao, Cewu Lu |
ICME | 1 |
| 2020 | PaStaNet: Toward Human Activity Knowledge EngineabstractExisting image-based activity understanding methods mainly adopt direct mapping, i.e. from image to activity concepts, which may encounter performance bottleneck since the huge gap. In light of this, we propose a new path: infer human part states first and then reason out the activities based on part-level semantics. Human Body Part States (PaSta) are fine-grained action semantic tokens, e.g., which can compose the activities and help us step toward human activity knowledge engine. To fully utilize the power of PaSta, we build a large-scale knowledge base PaStaNet, which contains 7M+ PaSta annotations. And two corresponding models are proposed: first, we design a model named Activity2Vec to extract PaSta features, which aim to be general representations for various activities. Second, we use a PaSta-based Reasoning method to infer activities. Promoted by PaStaNet, our method achieves significant improvements, e.g. 6.4 and 13.9 mAP on full and one-shot sets of HICO in supervised learning, and 3.2 and 4.2 mAP on V-COCO and images-based AVA in transfer learning. Code and data are available at http://hake-mvig.cn/. Yong-Lu Li 0001, Liang Xu 0012, Xinpeng Liu 0002, Xijie Huang, Haoshu Fang, Ze Ma, Cewu Lu |
CVPR | 2 |
| 2019 | Transferable Interactiveness Knowledge for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) Detection is an important problem to understand how humans interact with objects. In this paper, we explore Interactiveness Knowledge which indicates whether human and object interact with each other or not. We found that interactiveness knowledge can be learned across HOI datasets, regardless of HOI category settings. Our core idea is to exploit an Interactiveness Network to learn the general interactiveness knowledge from multiple HOI datasets and perform Non-Interaction Suppression before HOI classification in inference. On account of the generalization of interactiveness, interactiveness network is a transferable knowledge learner and can be cooperated with any HOI detection models to achieve desirable results. We extensively evaluate the proposed method on HICO-DET and V-COCO datasets. Our framework outperforms state-of-the-art HOI detection results by a great margin, verifying its efficacy and flexibility. Code is available at https://github.com/DirtyHarryLYL/Transferable-Interactiveness-Network. Yong-Lu Li 0001, Xijie Huang, Liang Xu 0012, Ze Ma, Haoshu Fang, Cewu Lu |
CVPR | 4 |