EDBT 2026 Demo / reviewers in the wild / expert
Chi Xie 0001
dblp:84/2208-1
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-5808-1742ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bi-Level Keypoint Relation Helps Versatile and Occluded Human Pose EstimationabstractRecently, there has been significant progress in 2D pose estimation. However, accurately localizing limb keypoints and occluded keypoints is still challenging. To tackle these difficulties, prior in human body structure has been leveraged in previous studies. One approach involves localizing a challenging keypoint by utilizing its neighbor keypoint. A previous study successfully employed neighbor-joint spatial relation (SR), which transfers features from a neighbor keypoint to the target keypoint being predicted. Building upon this idea, our work extends the keypoint relation-based method by incorporating another level of keypoint relation, namely channel-wise feature relation. This additional feature relation (FR) module assists in selecting more suitable neighbor keypoint feature channels and enhances the effectiveness of SR. By combining FR and SR, we develop a simple and intuitive bi-level keypoint relation module that can be trained end-to-end with existing methods. Through comprehensive experimental results and ablation studies, we demonstrate the effectiveness of our approach. Shuang Liang 0001, Chi Xie 0001, Jiewen Wang, Gang Chu, Shuwei Yan |
FG | 2 |
| 2025 | Classifier Recalibration for Human-Object Interaction Detection
Shuwei Yan, Shuang Liang 0001, Kenan Ye, Baihua Liu, Chi Xie 0001, Shengjie Zhao 0001 |
ICIC (6) | 5 |
| 2025 | Enhancing Document Understanding with Group Position Embedding: A Novel Approach to Incorporate Layout InformationabstractRecent advancements in document understanding have been dominated by leveraging large language models (LLMs) and multimodal large models. However, enabling LLMs to comprehend complex document layouts and structural information often necessitates intricate network modifications or costly pre-training, limiting their practical applicability. In this paper, we introduce Group Position Embedding (GPE), a novel and efficient technique to enhance the layout understanding capabilities of LLMs without architectural changes or additional pre-training. GPE achieves this by strategically grouping the attention heads and feeding each group with distinct positional embeddings, effectively encoding layout information relevant to document comprehension. This simple yet powerful method allows for effective integration of layout information within the existing LLM framework. We evaluate GPE against several competitive baselines across five mainstream document tasks. We also introduce a challenging benchmark called BLADE, specifically designed to assess layout comprehension.
Extensive experiments on both established and BLADE benchmarks confirm the efficacy of GPE in significantly advancing the state-of-the-art in document understanding. Our code is available at https://github.com/antgroup/GroupPositionEmbedding.git Yuke Zhu, Chi Xie 0001, Zihua Xiong, Bo Zheng 0007, Sheng Guo 0005 |
ICLR | 4 |
| 2025 | RelationLMM: Large Multimodal Model as Open and Versatile Visual Relationship GeneralistabstractVisual relationships are crucial for visual perception and reasoning, and cover tasks like Scene Graph Generation, Human-Object Interaction, and object affordance. Despite significant efforts, this field still suffers from the following limitations: specialists for a specific task without considering similar ones, strict and complex task formulations with limited flexibility, and underexploited reasoning with language and knowledge. To solve these limitations, we seek to build a new framework, one model for all tasks, over Large Multimodal Models (LMMs). LMMs offer the potential of unifying tasks, flexible forms, and reasoning with language. However, they fail to handle visual relationship tasks well. We find the obstacles include the conflicts between different tasks and insufficient instance-level information. We solve these problems by reforming the data for LMMs, rather than architectures, considering their strong language-in language-out capability. We propose to disassemble tasks into simple and common sub-tasks, verbally estimate instance confidence, and augment instance diversity, all without additional modules. These strategies help us build a visual relationship generalist, RelationLMM, with a simple architecture. Exhaustive experiments demonstrate RelationLMM is strong, generalizable and flexible to different tasks, with one model and one suite of weight. Chi Xie 0001, Shuang Liang 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Causal Intervention for Panoptic Scene Graph GenerationabstractPanoptic Scene Graph Generation (PSG) is a computer vision task that involves recognizing objects and their relationships within a given image. However, due to the long-tail distribution of the dataset, current PSG methods face bias issues. Previous debiasing methods usually rely on resampling and reweighting, which are unable to address the issue fundamentally, and may lead to underfitting for the head categories to some extent. Causal intervention, as a method to identify confounders from a causal inference perspective, can serve as a suitable means of debiasing. In this work, we propose a novel debiasing method based on causal intervention, which treats the long-tail distribution prior of the dataset as a confounder and eliminates it using the backdoor criterion. An uncertainty estimation module is further employed to assist in determining hard samples. Experiments demonstrate a significant improvement in the competitiveness of our approach compared to the baseline. Moreover, our method exhibits notable enhancements in accuracy and generalization on tail categories. Shuang Liang 0001, Chi Xie 0001 |
ICME | 3 |
| 2024 | Sketch-based 3D shape retrieval via teacher-student learning
Shuang Liang 0001, Weidong Dai, Yiyang Cai, Chi Xie 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Scribble-based complementary graph reasoning network for weakly supervised salient object detection
Shuang Liang 0001, Zhiqi Yan, Chi Xie 0001, Hongming Zhu, Jiewen Wang |
Comput. Vis. Image Underst. | 3 |
| 2024 | Relation with Free Objects for Action RecognitionabstractRelevant objects are widely used for aiding human action recognition in still images. Such objects are founded by a dedicated and pre-trained object detector in all previous methods. Such methods have two drawbacks. First, training an object detector requires intensive data annotation. This is costly and sometimes unaffordable in practice. Second, the relation between objects and humans are not fully taken into account in training. This work proposes a systematic approach to address the two problems. We propose two novel network modules. The first is an object extraction module that automatically finds relevant objects for action recognition, without requiring annotations. Thus, it is free . The second is a human-object relation module that models the pairwise relation between humans and objects, and enhances their features. Both modules are trained in the action recognition network, end-to-end. Comprehensive experiments and ablation studies on three datasets for action recognition in still images demonstrate the effectiveness of the proposed approach. Our method yields state-of-the-art results. Specifically, on the HICO dataset, it achieves 44.9% mAP, which is 12% relative improvement over the previous best result. In addition, this work makes an observational contribution that it is no longer necessary to rely on a pre-trained object detector for this task. Relevant objects can be found via end-to-end learning with only action labels. This is encouraging for action recognition in the wild. Models and code will be released. Shuang Liang 0001, Wentao Ma 0004, Chi Xie 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Category Query Learning for Human-Object Interaction ClassificationabstractUnlike most previous HOI methods that focus on learning better human-object features, we propose a novel and complementary approach called category query learning. Such queries are explicitly associated to interaction categories, converted to image specific category representation via a transformer decoder, and learnt via an auxiliary image-level classification task. This idea is motivated by an earlier multi-label image classification method, but is for the first time applied for the challenging human-object interaction classification task. Our method is simple, general and effective. It is validated on three representative HOI baselines and achieves new state-of-the-art results on two benchmarks. Code will be available at https://github.com/charles-xie/CQL. Chi Xie 0001, Fangao Zeng, Yue Hu 0011, Shuang Liang 0001 |
CVPR | 1 |
| 2023 | Advancing Referring Expression Segmentation Beyond Single ImageabstractReferring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expression. However, in broader real-world scenarios, it is not always possible to determine if the described object exists in a specific image. Generally, a collection of images is available, some of which potentially contain the target objects. To this end, we propose a more realistic setting, named Group-wise Referring Expression Segmentation (GRES), which expands RES to a group of related images, allowing the described objects to exist in a subset of the input image group. To support this new setting, we introduce an elaborately compiled dataset named Grouped Referring Dataset (GRD), containing complete group-wise annotations of the target objects described by given expressions. Moreover, we also present a baseline method named Grouped Referring Segmenter (GRSer), which explicitly captures the language-vision and intra-group vision-vision interactions to achieve state-of-the-art results on the proposed GRES setting and related tasks, such as Co-Salient Object Detection and traditional RES. Our dataset and codes are publicly released in https://github.com/shikras/d-cube. Zhao Zhang 0018, Chi Xie 0001, Feng Zhu 0006, Rui Zhao 0001 |
ICCV | 3 |
| 2023 | Compositional Learning in Transformer-Based Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is an important part of understanding human activities and visual scenes. The long-tailed distribution of labeled instances is a primary challenge in HOI detection, promoting research in few-shot and zero-shot learning. Inspired by the combinatorial nature of HOI triplets, some existing approaches adopt the idea of compositional learning, in which object and action features are learned individually and re-composed as new training samples. However, these methods follow the CNN-based two-stage paradigm with limited feature extraction ability, and often rely on auxiliary information for better performance. Without introducing any additional information, we creatively propose a transformer-based framework for compositional HOI learning. Human-object pair representations and interaction representations are re-composed across different HOI instances, which involves richer contextual information and promotes the generalization of knowledge. Experiments show our simple but effective method achieves state-of-the-art performance, especially on rare HOI classes. Zikun Zhuang, Ruihao Qian, Chi Xie 0001, Shuang Liang 0001 |
ICME | 3 |
| 2023 | Described Object Detection: Liberating Object Detection with Flexible ExpressionsabstractDetecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to flexible language expressions for OVD and overcoming the limitation of REC only grounding the pre-existing object. We establish the research foundation for DOD by constructing a *Description Detection Dataset* ($D^3$). This dataset features flexible language expressions, whether short category names or long descriptions, and annotating all described objects on all images without omission. By evaluating previous SOTA methods on $D^3$, we find some troublemakers that fail current REC, OVD, and bi-functional methods. REC methods struggle with confidence scores, rejecting negative instances, and multi-target scenarios, while OVD methods face constraints with long and complex descriptions. Recent bi-functional methods also do not work well on DOD due to their separated training procedures and inference strategies for REC and OVD tasks. Building upon the aforementioned findings, we propose a baseline that largely improves REC methods by reconstructing the training data and introducing a binary classification sub-task, outperforming existing methods. Data and code are available at https://github.com/shikras/d-cube and related works are tracked in https://github.com/Charles-Xie/awesome-described-object-detection. Chi Xie 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001, Shuang Liang 0001 |
NeurIPS | 1 |
| 2023 | Temporal Dropout for Weakly Supervised Action LocalizationabstractWeakly supervised action localization is a challenging problem in video understanding and action recognition. Existing models usually formulate the training process as direct classification using video-level supervision. They tend to only locate the most discriminative parts of action instances and produce temporally incomplete detection results. A natural solution for this problem, the adversarial erasing strategy, is to remove such parts from training so that models can attend to complementary parts. Previous works do it in an offline and heuristic way. They adopt a multi-stage pipeline, where discriminative regions are determined and erased under the guidance of detection results from last stage. Such a pipeline can be both ineffective and inefficient, possibly hindering the overall performance. On the contrary, we combine adversarial erasing with dropout mechanism and propose a Temporal Dropout Module that learns where to remove in a data-driven and online manner. This plug-and-play module is trained without iterative stages, which not only simplifies the pipeline but also makes the regularization during training easier and more adaptive. Experiments show that the proposed method outperforms previous erasing-based methods by a large margin. More importantly, it achieves universal improvement when plugged into various direct classification methods and obtains state-of-the-art performance. Chi Xie 0001, Zikun Zhuang, Shengjie Zhao 0001, Shuang Liang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Joint relation based human pose estimation
Shuang Liang 0001, Gang Chu, Chi Xie 0001, Jiewen Wang |
Vis. Comput. | 3 |
| 2021 | Recurrent Graph Convolutional Autoencoder for Unsupervised Skeleton-Based Action RecognitionabstractSkeleton-based action recognition is a significant task in computer vision due to its robustness and wide application. Most unsupervised methods do not employ topological information of skeleton graphs, which actually ignore the spatial dependencies of action sequences. In this paper, we introduce a Recurrent Graph Convolutional Autoencoder (RGCA) for unsupervised action recognition from skeleton data. Our method explicitly exploits the spatial relationships among every frame’s joints while preserving the long-term temporal dynamics in whole sequences. Moreover, a Spatial Joints Attention Module is employed to measure the importance of joints in the input sequence automatically. We conduct experiments on three datasets (NTU RGB+D 60, NW-UCLA, and UWA3D) and exceed the state-of-the-art performance. Sev-Jue Zhao, Chi Xie 0001, Kenan Ye, Shuang Liang 0001 |
ICME | 3 |
| 2021 | Automatic Pose Quality Assessment for Adaptive Human Pose Refinement
Gang Chu, Chi Xie 0001, Shuang Liang 0001 |
MMM (1) | 2 |