EDBT 2026 Demo / reviewers in the wild / expert
Zhengxi Hu
dblp:276/2047
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0001-6119-4185ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DCCLA: Dense Cross Connections With Linear Attention for LiDAR-Based 3D Pedestrian DetectionabstractLiDAR-based 3D pedestrian detection has recently been extensively applied in autonomous driving and intelligent mobile robots. However, it remains a highly challenging perceptual task due to the sparsity of pedestrian point cloud data and the significant deformation of pedestrian body postures. To address these challenges, we propose a Dense Cross Connections network with Linear Attention (DCCLA), which mitigates the semantic discrepancy between the encoder and decoder of the network by integrating multiple 3D sparse convolutional layers within the skip connections. Furthermore, we enhance these connections by introducing cross-connections, thereby effectively promoting information interaction among various channels. To effectively retain crucial information while summarizing diverse pedestrian representations, we propose the Linear Self-Attention module for 3D point clouds (LSA3D), which significantly reduces model complexity. The experimental results demonstrate that our DCCLA achieves state-of-the-art Average Precision (AP) for the 3D pedestrian detection task on the JRDB large-scale dataset, outperforming the second-ranked method by 2.7% AP. Furthermore, our DCCLA enhances 1.6% mIoU over the benchmark method on the SemanticKITTI dataset. Therefore, our method achieves excellent performance through a cross-scale feature fusion strategy and linear attention that fully combines the advantages of convolution and transformer architectures. The project is publicly available athttps://github.com/jinzhengguang/DCCLA. Jinzheng Guang, Zhengxi Hu, Qianyi Zhang, Jingtai Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | UAGE: A Supervised Contrastive Method for Unconstrained Adaptive Gaze Estimation
Enfan Lan, Zhengxi Hu, Jingtai Liu |
ACCV (7) | 2 |
| 2024 | RPEA: A Residual Path Network with Efficient Attention for 3D pedestrian detection from LiDAR point clouds
Jinzheng Guang, Zhengxi Hu, Qianyi Zhang, Jingtai Liu |
Expert Syst. Appl. | 2 |
| 2024 | Hierarchical Context-Based Emotion Recognition With Scene GraphsabstractFor a better intention inference, we often try to figure out the emotional states of other people in social communications. Many studies on affective computing have been carried out to infer emotions through perceiving human states, i.e., facial expression and body posture. Such methods are skillful in a controlled environment. However, it often leads to misestimation due to the deficiency of effective inputs in unconstrained circumstances, that is, where context-aware emotion recognition appeared. We take inspiration from the advanced reasoning pattern of humans in perceived emotion recognition and propose the hierarchical context-based emotion recognition method with scene graphs. We propose to extract three contexts from the image, i.e., the entity context, the global context, and the scene context. The scene context contains abstract information about entity labels and their relationships. It is similar to the information processing of the human visual sensing mechanism. After that, these contexts are further fused to perform emotion recognition. We carried out a bunch of experiments on the widely used context-aware emotion datasets, i.e., CAER-S, EMOTIC, and BOdy Language Dataset (BoLD). We demonstrate that the hierarchical contexts can benefit emotion recognition by improving the accuracy of the SOTA score from 84.82% to 90.83% on CAER-S. The ablation experiments show that hierarchical contexts provide complementary information. Our method improves the F1 score of the SOTA result from 29.33% to 30.24% (C-F1) on EMOTIC. We also build the image-based emotion recognition task with BoLD-Img from BoLD and obtain a better emotion recognition score (ERS) score of 0.2153. Lei Zhou 0017, Zhengxi Hu, Jingtai Liu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | GFIE: A Dataset and Baseline for Gaze-Following from 2D to 3D in Indoor EnvironmentsabstractGaze-following is a kind of research that requires locating where the person in the scene is looking automatically under the topic of gaze estimation. It is an important clue for understanding human intention, such as identifying objects or regions of interest to humans. However, a survey of datasets used for gaze-following tasks reveals defects in the way they collect gaze point labels. Manual labeling may introduce subjective bias and is labor-intensive, while automatic labeling with an eye-tracking device would alter the person's appearance. In this work, we introduce GFIE, a novel dataset recorded by a gaze data collection system we developed. The system is constructed with two devices, an Azure Kinect and a laser rangefinder, which generate the laser spot to steer the subject's attention as they perform in front of the camera. And an algorithm is developed to locate laser spots in images for annotating 2D/3D gaze targets and removing ground truth introduced by the spots. The whole procedure of collecting gaze behavior allows us to obtain unbiased labels in unconstrained environments semi-automatically. We also propose a baseline method with stereo field-of-view (FoV) perception for establishing a 2D/3D gaze-following benchmark on the GFIE dataset. Project page: https://sites.google.com/view/gfie. Zhengxi Hu, Yuxue Yang, Xiaolin Zhai, Dingye Yang, Jingtai Liu |
CVPR | 1 |
| 2023 | Learning Group Residual Representation for Group Activity Prediction*abstractThe goal of group activity prediction is to infer the group activity involved multiple individuals before it is completely executed. Previous methods focused on capturing pair-wise relationships between individuals, but lacked the exploration of group-wise interactions which can provide global guidance from a macroscopic perspective. To further explore the group-wise interaction, we propose a Group Residual Module (GRM) which constructs a virtual leader node to summarize the group representation and designs a bidirectional message passing mechanism to build the bridge between group and individuals. To capture the spatial-temporal correlation jointly, we propose a Spatial-Temporal Group Residual Network composed of spatial GRMs and temporal GRMs. Different from existing methods that obtain additional information from the complete activity execution, temporal masks in the temporal GRMs are designed to enforce our network to excavate as much discriminative information as possible from the observed activity sequence. Moreover, experimental results show that our network achieves state-of-the-art performance on Volleyball Dataset and Collective Activity Dataset. Xiaolin Zhai, Zhengxi Hu, Dingye Yang, Jingtai Liu |
ICME | 2 |
| 2023 | The Human Gaze Helps Robots Run Bravely and Efficiently in CrowdsabstractIn human-aware navigation, the robot tacitly games with humans, balancing safety and efficiency according to human intentions. Poor balance or bad intent recognition causes the robot to stop conservatively or advance rashly, resulting in a deadlock or even a collision respectively. To address the issue, this paper proposes an improved limit cycle for collaboratively parameterizing human intentions and planning robot motions. The human-robot interaction is modeled as a dynamic chicken game with incomplete information, where the human gaze is introduced to depict the unique characteristics of each person, allowing the robot to approach with different safety margins. Our method is tested in challenging indoor scenarios and outperforms traditional methods in both safety and efficiency. We enable robots to utilize human wisdom to solve problems that cannot be solved on their own. The robot bravely goes through oncoming crowds by getting closer to people with higher attention on it and has the foresight to stably cross in front or behind people. Qianyi Zhang, Zhengxi Hu, Yinuo Song, Jiayi Pei, Jingtai Liu |
ICRA | 2 |
| 2023 | Advanced acoustic footstep-based person identification dataset and method using multimodal feature fusion
Xiaolin Zhai, Zhengxi Hu, Jingtai Liu |
Knowl. Based Syst. | 3 |
| 2022 | MGTR: End-to-End Mutual Gaze Detection with Transformer
Hang Guo 0002, Zhengxi Hu, Jingtai Liu |
ACCV (4) | 2 |
| 2022 | Social Aware Multi-modal Pedestrian Crossing Behavior Prediction
Xiaolin Zhai, Zhengxi Hu, Dingye Yang, Lei Zhou 0017, Jingtai Liu |
ACCV (4) | 2 |
| 2022 | Spatial Temporal Network for Image and Skeleton Based Group Activity Recognition
Xiaolin Zhai, Zhengxi Hu, Dingye Yang, Lei Zhou 0017, Jingtai Liu |
ACCV (4) | 2 |
| 2022 | Gaze Target Estimation Inspired by Interactive AttentionabstractAs an essential nonverbal cue, the human gaze reveals human intentions and plays a crucial role in human daily activities. Therefore, automatic detection of the person’s gaze target has drawn the interests of the computer vision community. This is useful not only for identifying whether children are attentive in class but also for locating items of interest to humans in retail settings. Existing gaze-following methods have only explored and exploited the scenes context and the head cues. Considering the significance of human-object interaction in understanding human intentions, we present the Visual-Spatial Graph and introduce a graph attention network to analyze the interaction probability between the human and elements in the scene. Then the interaction probability inferred from the visual-spatial information that is aggregated by the attention mechanism can be transformed into an interactive attention map that depicts the areas people care about. In addition, we construct a transformer as an encoder to integrate the features extracted by the scene and head pathways aiming to decode the gaze target. After introducing interactive attention, our proposed method achieves outstanding performance on two benchmarks: GazeFollow and VideoAttentionTarget. Our code is available athttps://github.com/nkuhzx/VSG-IA. Zhengxi Hu, Kunxu Zhao, Hang Guo 0002, Yuxue Yang, Jingtai Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |