EDBT 2026 Demo / reviewers in the wild / expert
Tao Pu 0002
dblp:137/9982-2
· DBLP profile ↗
17ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0001-9564-8288ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based SearchabstractRecent progress in robotics and embodied AI is largely driven by Large Multimodal Models (LMMs). However, a key challenge remains underexplored: how can we advance LMMs to discover tasks that assist humans in open-future scenarios, where human intentions are highly concurrent and dynamic. In this work, we formalize the problem of Human-centric Open-future Task Discovery (HOTD), focusing particularly on identifying tasks that reduce human effort across plausible futures. To facilitate this study, we propose HOTD-Bench, which features over 2K real-world videos, a semi-automated annotation pipeline, and a simulation-based protocol tailored for open-set future evaluation. Additionally, we propose the Collaborative Multi-Agent Search Tree (CMAST) framework, which decomposes complex reasoning through a multi-agent system and structures the reasoning process through a scalable search tree module. In our experiments, CMAST achieves the best performance on the HOTD-Bench, significantly surpassing existing LMMs. It also integrates well with existing LMMs, consistently improving performance. Xiaoxin Lin, Tao Pu 0002, Zhenlong Yuan, Guangrun Wang, Liang Lin 0004 |
AAAI | 3 |
| 2025 | DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label RecognitionabstractOpen-Vocabulary Multi-Label Recognition (OV-MLR) aims to identify multiple seen and unseen object categories within an image, requiring both precise intra-class localization to pinpoint objects and effective inter-class reasoning to model complex category dependencies. While Vision-Language Pre-training (VLP) models offer a strong open-vocabulary foundation, they often struggle with fine-grained localization under weak supervision and typically fail to explicitly leverage structured relational knowledge beyond basic semantics, limiting performance especially for unseen classes. To overcome these limitations, we propose the Dual Adaptive Refinement Transfer (DART) framework. DART enhances a frozen VLP backbone via two synergistic adaptive modules. For intra-class refinement, an Adaptive Refinement Module (ARM) refines patch features adaptively, coupled with a novel Weakly Supervised Patch Selecting (WPS) loss that enables discriminative localization using only image-level labels. Concurrently, for inter-class transfer, an Adaptive Transfer Module (ATM) leverages a Class Relationship Graph (CRG), constructed using structured knowledge mined from a Large Language Model (LLM), and employs graph attention network to adaptively transfer relational information between class representations. DART is the first framework, to our knowledge, to explicitly integrate external LLM-derived relational knowledge for adaptive inter-class transfer while simultaneously performing adaptive intra-class refinement under weak supervision for OV-MLR. Extensive experiments on challenging benchmarks demonstrate that our DART achieves new state-of-the-art performance, validating its effectiveness. Haijing Liu, Tao Pu 0002, Hefeng Wu, Keze Wang, Liang Lin 0004 |
ACM Multimedia | 2 |
| 2025 | Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal InterventionabstractEgocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for understanding egocentric human behavior. However, achieving such segmentation robustly is challenging due to ambiguities inherent in egocentric videos and biases present in training data. Consequently, existing methods often struggle, learning spurious correlations from skewed object-action pairings in datasets and fundamental visual confounding factors of the egocentric perspective, such as rapid motion and frequent occlusions. To address these limitations, we introduce Causal Ego-REferring Segmentation (CERES), a plug-in causal framework that adapts strong, pre-trained RVOS backbones to the egocentric domain. CERES implements dual-modal causal intervention: applying backdoor adjustment principles to counteract language representation biases learned from dataset statistics, and leveraging front-door adjustment concepts to address visual confounding by intelligently integrating semantic visual features with geometric depth information guided by causal principles, creating representations more robust to egocentric distortions. Extensive experiments demonstrate that CERES achieves state-of-the-art performance on Ego-RVOS benchmarks, highlighting the potential of applying causal reasoning to build more reliable models for broader egocentric video understanding. Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 0002, Keze Wang, Liang Lin 0004 |
NeurIPS | 4 |
| 2024 | Semantic Correlation Adaptation for Union-Set Multi-label Image Recognition
Tao Pu 0002, Dongyu Zhang 0002, Liang Lin 0004 |
ICPR (3) | 2 |
| 2024 | Dual-perspective semantic-aware representation blending for multi-label image recognition with partial labels
Tao Pu 0002, Tianshui Chen, Hefeng Wu, Yukai Shi, Zhijing Yang, Liang Lin 0004 |
Expert Syst. Appl. | 1 |
| 2024 | Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels
Tianshui Chen, Tao Pu 0002, Lingbo Liu, Yukai Shi, Zhijing Yang, Liang Lin 0004 |
Int. J. Comput. Vis. | 2 |
| 2024 | Dynamic Correlation Learning and Regularization for Multi-Label Confidence CalibrationabstractModern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confidence scores that necessitate calibration. While current confidence calibration techniques primarily address single-label scenarios, there is a lack of focus on more practical and generalizable multi-label contexts. This paper introduces the Multi-Label Confidence Calibration (MLCC) task, aiming to provide well-calibrated confidence scores in multi-label scenarios. Unlike single-label images, multi-label images contain multiple objects, leading to semantic confusion and further unreliability in confidence scores. Existing single-label calibration methods, based on label smoothing, fail to account for category correlations, which are crucial for addressing semantic confusion, thereby yielding sub-optimal performance. To overcome these limitations, we propose the Dynamic Correlation Learning and Regularization (DCLR) algorithm, which leverages multi-grained semantic correlations to better model semantic confusion for adaptive regularization. DCLR learns dynamic instance-level and prototype-level similarities specific to each category, using these to measure semantic correlations across different categories. With this understanding, we construct adaptive label vectors that assign higher values to categories with strong correlations, thereby facilitating more effective regularization. We establish an evaluation benchmark, re-implementing several advanced confidence calibration algorithms and applying them to leading multi-label recognition (MLR) models for fair comparison. Through extensive experiments, we demonstrate the superior performance of DCLR over existing methods in providing reliable confidence scores in multi-label scenarios. Tianshui Chen, Weihang Wang 0008, Tao Pu 0002, Jinghui Qin, Zhijing Yang, Jie Liu 0022, Liang Lin 0004 |
IEEE Trans. Image Process. | 3 |
| 2024 | Spatial-Temporal Knowledge-Embedded Transformer for Video Scene Graph GenerationabstractVideo scene graph generation (VidSGG) aims to identify objects in visual scenes and infer their relationships for a given video. It requires not only a comprehensive understanding of each object scattered on the whole scene but also a deep dive into their temporal motions and interactions. Inherently, object pairs and their relationships enjoy spatial co-occurrence correlations within each image and temporal consistency/transition correlations across different images, which can serve as prior knowledge to facilitate VidSGG model learning and inference. In this work, we propose a spatial-temporal knowledge-embedded transformer (STKET) that incorporates the prior spatial-temporal knowledge into the multi-head cross-attention mechanism to learn more representative relationship representations. Specifically, we first learn spatial co-occurrence and temporal transition correlations in a statistical manner. Then, we design spatial and temporal knowledge-embedded layers that introduce the multi-head cross-attention mechanism to fully explore the interaction between visual representation and the knowledge to generate spatial- and temporal-embedded representations, respectively. Finally, we aggregate these representations for each subject-object pair to predict the final semantic labels and their relationships. Extensive experiments show that STKET outperforms current competing algorithms by a large margin, e.g., improving the mR@50 by 8.1%, 4.7%, and 2.1% on different settings over current algorithms. Tao Pu 0002, Tianshui Chen, Hefeng Wu, Yongyi Lu, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2024 | Category-Adaptive Label Discovery and Noise Rejection for Multi-Label Recognition With Partial Positive LabelsabstractAs a cost-effective alternative to standard multi-label learning, the multi-label image recognition with partial positive labels (MLR-PPL) task attracts increasing attention, in which merely a portion of positive labels are given while the rest of positive labels and all negative labels are missing. To facilitate this task, we propose a novel framework that leverages semantic correlation among different images in a category-adaptive manner to complement unknown labels accurately. Specifically, the proposed framework consists of two complementary modules. 1) A category-adaptive label discovery (CALD) module is designed to measure the semantic similarity between positive samples and then complement unknown labels with high similarities. 2) A category-adaptive noise rejection (CANR) module is designed to compute the sample weights based on semantic similarities from different samples and discard noisy labels with low weights. Due to the various degrees of confidence calibration among different categories, searching appropriate thresholds for each category in the proposed framework is highly time-consuming. To avoid such a resource-intensive manual tuning, we introduce a category-adaptive threshold updating algorithm that introduces the category-specific positive and negative similarity to adjust the threshold adaptively. Extensive experiments on various benchmarks show that the proposed framework performs better than current state-of-the-art algorithms. Tao Pu 0002, Qianru Lao, Hefeng Wu, Tianshui Chen, Ling Tian, Jie Liu 0022, Liang Lin 0004 |
IEEE Trans. Multim. | 1 |
| 2023 | Video Scene Graph Generation with Spatial-Temporal KnowledgeabstractVarious video understanding tasks have been extensively explored in the multimedia community, among which the video scene graph generation (VidSGG) task is more challenging since it requires identifying objects in comprehensive scenes and deducing their relationships. Existing methods for this task generally aggregate object-level visual information from both spatial and temporal perspectives to better learn powerful relationship representations. However, these leading techniques merely implicitly model the spatial-temporal context, which may lead to ambiguous predicate predictions when visual relations vary frequently. In this work, I propose incorporating spatial-temporal knowledge into relation representation learning to effectively constrain the spatial prediction space within each image and sequential variation across temporal frames. To this end, I design a novel spatial-temporal knowledge-embedded transformer (STKET) that incorporates the prior spatial-temporal knowledge into the multi-head cross-attention mechanism to learn more representative relationship representations. Extensive experiments conducted on Action Genome demonstrate the effectiveness of the proposed STKET. Tao Pu 0002 |
ACM Multimedia | 1 |
| 2023 | Semantic representation and dependency learning for multi-label image recognition
Tao Pu 0002, Mingzhan Sun, Hefeng Wu, Tianshui Chen, Ling Tian, Liang Lin 0004 |
Neurocomputing | 1 |
| 2022 | Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial LabelsabstractTraining the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging and practical task. To address this task, current algorithms mainly depend on pre-training classification or similarity models to generate pseudo labels for the unknown labels. However, these algorithms depend on sufficient multi-label annotations to train the models, leading to poor performance especially with low known label proportion. In this work, we propose to blend category-specific representation across different images to transfer information of known labels to complement unknown labels, which can get rid of pre-training models and thus does not depend on sufficient annotations. To this end, we design a unified semantic-aware representation blending (SARB) framework that exploits instance-level and prototype-level semantic representation to complement unknown labels by two complementary modules: 1) an instance-level representation blending (ILRB) module blends the representations of the known labels in an image to the representations of the unknown labels in another image to complement these unknown labels. 2) a prototype-level representation blending (PLRB) module learns more stable representation prototypes for each category and blends the representation of unknown labels with the prototypes of corresponding labels to complement these labels. Extensive experiments on the MS-COCO, Visual Genome, Pascal VOC 2007 datasets show that the proposed SARB framework obtains superior performance over current leading competitors on all known label proportion settings, i.e., with the mAP improvement of 4.6%, 4.6%, 2.2% on these three datasets when the known label proportion is 10%. Codes are available at https://github.com/HCPLab-SYSU/HCP-MLR-PL. Tao Pu 0002, Tianshui Chen, Hefeng Wu, Liang Lin 0004 |
AAAI | 1 |
| 2022 | Structured Semantic Transfer for Multi-Label Recognition with Partial LabelsabstractMulti-label image recognition is a fundamental yet practical task because real-world images inherently possess multiple semantic labels. However, it is difficult to collect large-scale multi-label annotations due to the complexity of both the input images and output label spaces. To reduce the annotation cost, we propose a structured semantic transfer (SST) framework that enables training multi-label recognition models with partial labels, i.e., merely some labels are known while other labels are missing (also called unknown labels) per image. The framework consists of two complementary transfer modules that explore within-image and cross-image semantic correlations to transfer knowledge of known labels to generate pseudo labels for unknown labels. Specifically, an intra-image semantic transfer module learns image-specific label co-occurrence matrix and maps the known labels to complement unknown labels based on this matrix. Meanwhile, a cross-image transfer module learns category-specific feature similarities and helps complement unknown labels with high similarities. Finally, both known and generated labels are used to train the multi-label recognition models. Extensive experiments on the Microsoft COCO, Visual Genome and Pascal VOC datasets show that the proposed SST framework obtains superior performance over current state-of-the-art algorithms. Codes are available at https://github.com/HCPLab-SYSU/HCP-MLR-PL. Tianshui Chen, Tao Pu 0002, Hefeng Wu, Yuan Xie 0004, Liang Lin 0004 |
AAAI | 2 |
| 2022 | Content-Aware Hierarchical Representation Selection for Cross-View Geo-Localization
Zeng Lu, Tao Pu 0002, Tianshui Chen, Liang Lin 0004 |
ACCV (5) | 2 |
| 2022 | Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph LearningabstractFacial expression recognition (FER) has received significant attention in the past decade with witnessed progress, but data inconsistencies among different FER datasets greatly hinder the generalization ability of the models learned on one dataset to another. Recently, a series of cross-domain FER algorithms (CD-FERs) have been extensively developed to address this issue. Although each declares to achieve superior performance, comprehensive and fair comparisons are lacking due to inconsistent choices of the source/target datasets and feature extractors. In this work, we first propose to construct a unified CD-FER evaluation benchmark, in which we re-implement the well-performing CD-FER and recently published general domain adaptation algorithms and ensure that all these algorithms adopt the same source/target datasets and feature extractors for fair CD-FER evaluations. Based on the analysis, we find that most of the current state-of-the-art algorithms use adversarial learning mechanisms that aim to learn holistic domain-invariant features to mitigate domain shifts. However, these algorithms ignore local features, which are more transferable across different datasets and carry more detailed content for fine-grained adaptation. Therefore, we develop a novel adversarial graph representation adaptation (AGRA) framework that integrates graph representation propagation with adversarial learning to realize effective cross-domain holistic-local feature co-adaptation. Specifically, our framework first builds two graphs to correlate holistic and local regions within each domain and across different domains, respectively. Then, it extracts holistic-local features from the input image and uses learnable per-class statistical distributions to initialize the corresponding graph nodes. Finally, two stacked graph convolution networks (GCNs) are adopted to propagate holistic-local features within each domain to explore their interaction and across different domains for holistic-local feature co-adaptation. In this way, the AGRA framework can adaptively learn fine-grained domain-invariant features and thus facilitate cross-domain expression recognition. We conduct extensive and fair comparisons on the unified evaluation benchmark and show that the proposed AGRA framework outperforms previous state-of-the-art methods. Tianshui Chen, Tao Pu 0002, Hefeng Wu, Yuan Xie 0004, Lingbo Liu, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | AU-Expression Knowledge Constrained Representation Learning for Facial Expression RecognitionabstractRecognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Current deep-learning-based algorithms may achieve impressive performance in some lab-controlled environments, but they always fail to recognize the expressions accurately for the uncontrolled in-the-wild situation. Fortunately, facial action units (AU) describe subtle facial behaviors, and they can help distinguish uncertain and ambiguous expressions. In this work, we explore the correlations among the action units and facial expressions, and devise an AU-Expression Knowledge Constrained Representation Learning (AUE-CRL) framework to learn the AU representations without AU annotations and adaptively use representations to facilitate facial expression recognition. Specifically, it leverages AU-expression correlations to guide the learning of the AU classifiers, and thus it can obtain AU representations without incurring any AU annotations. Then, it introduces a knowledge-guided attention mechanism that mines useful AU representations under the constraint of AU-expression correlations. In this way, the framework can capture local discriminative and complementary features to enhance facial representation for facial expression recognition. We conduct experiments on the challenging uncontrolled datasets to demonstrate the superiority of the proposed framework over current state-of-the-art methods. Codes and trained models are available at https://github.com/HCPLab-SYSU/AUE-CRL. Tao Pu 0002, Tianshui Chen, Yuan Xie 0004, Hefeng Wu, Liang Lin 0004 |
ICRA | 1 |
| 2020 | Adversarial Graph Representation Adaptation for Cross-Domain Facial Expression RecognitionabstractData inconsistency and bias are inevitable among different facial expression recognition (FER) datasets due to subjective annotating process and different collecting conditions. Recent works resort to adversarial mechanisms that learn domain-invariant features to mitigate domain shift. However, most of these works focus on holistic feature adaptation, and they ignore local features that are more transferable across different datasets. Moreover, local features carry more detailed and discriminative content for expression recognition, and thus integrating local features may enable fine-grained adaptation. In this work, we propose a novel Adversarial Graph Representation Adaptation (AGRA) framework that unifies graph representation propagation with adversarial learning for cross-domain holistic-local feature co-adaptation. To achieve this, we first build a graph to correlate holistic and local regions within each domain and another graph to correlate these regions across different domains. Then, we learn the per-class statistical distribution of each domain and extract holistic-local features from the input image to initialize the corresponding graph nodes. Finally, we introduce two stacked graph convolution networks to propagate holistic-local feature within each domain to explore their interaction and across different domains for holistic-local feature co-adaptation. In this way, the AGRA framework can adaptively learn fine-grained domain-invariant features and thus facilitate cross-domain expression recognition. We conduct extensive and fair experiments on several popular benchmarks and show that the proposed AGRA framework achieves superior performance over previous state-of-the-art methods. Yuan Xie 0004, Tianshui Chen, Tao Pu 0002, Hefeng Wu, Liang Lin 0004 |
ACM Multimedia | 3 |