Hefeng Wu

dblp:126/4518 · DBLP profile ↗
← Back
52ranked-venue papers
7as first author
25since 2021 · last 2025
0000-0002-2132-6515ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
abstract
Long Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs).However, due to the inherently lengthy nature of long-text data, existing LCU benchmarks for LLMs often result in prohibitively high evaluation costs, like testing time and inference expenses.Through extensive experimentation, we discover that existing LCU benchmarks exhibit significant redundancy, which means the inefficiency in evaluation.In this paper, we propose a concise data compression method tailored for longtext data with sparse information characteristics.By pruning the well-known LCU benchmark LongBench, we create MiniLongBench.This benchmark includes only 237 test samples across six major task categories and 21 distinct tasks.Through empirical analysis of over 60 LLMs, MiniLongBench achieves an average evaluation cost reduced to only 4.5% of the original while maintaining an average rank correlation coefficient of 0.97 with Long-Bench results.Therefore, our MiniLongBench, as a low-cost benchmark, holds great potential to substantially drive future research into the LCU capabilities of LLMs.See Github for our code, data and tutorial.
Zhongzhan Huang, Guoming Ling, Shanshan Zhong, Hefeng Wu, Liang Lin 0004
ACL (1)4
2025 RoboPearls: Editable Video Simulation for Robot Manipulation
Tang Tao, Likui Zhang, Youpeng Wen, Kaidong Zhang, Jiawang Bian, Tianyi Yan, Kun Zhan, Peng Jia 0007, Hefeng Wu, Xiaodan Liang
ICCV10
2025 RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
abstract
Operating robots in open-ended scenarios with diverse tasks is a crucial research and application direction in robotics. While recent progress in natural language processing and large multimodal models has enhanced robots' ability to understand complex instructions, robot manipulation still faces the procedural skill dilemma and the declarative skill dilemma in open environments. Existing methods often compromise cognitive and executive capabilities. To address these challenges, in this paper, we propose RoBridge, a hierarchical intelligent architecture for general robotic manipulation. It consists of a high-level cognitive planner (HCP) based on a large-scale pre-trained vision-language model (VLM), an invariant operable representation (IOR) serving as a symbolic bridge, and a generalist embodied agent (GEA). RoBridge maintains the declarative skill of VLM and unleashes the procedural skill of reinforcement learning, effectively bridging the gap between cognition and execution. RoBridge demonstrates significant performance improvements over existing baselines, achieving a 75% success rate on new tasks and an 83% average success rate in sim-to-real generalization using only five real-world data samples per task. This work represents a significant step towards integrating cognitive reasoning with physical execution in robotic systems, offering a new paradigm for general robotic manipulation.
Kaidong Zhang, Rongtao Xu, Pengzhen Ren, Junfan Lin, Hefeng Wu, Xiaodan Liang
ICCV5
2025 DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
abstract
Open-Vocabulary Multi-Label Recognition (OV-MLR) aims to identify multiple seen and unseen object categories within an image, requiring both precise intra-class localization to pinpoint objects and effective inter-class reasoning to model complex category dependencies. While Vision-Language Pre-training (VLP) models offer a strong open-vocabulary foundation, they often struggle with fine-grained localization under weak supervision and typically fail to explicitly leverage structured relational knowledge beyond basic semantics, limiting performance especially for unseen classes. To overcome these limitations, we propose the Dual Adaptive Refinement Transfer (DART) framework. DART enhances a frozen VLP backbone via two synergistic adaptive modules. For intra-class refinement, an Adaptive Refinement Module (ARM) refines patch features adaptively, coupled with a novel Weakly Supervised Patch Selecting (WPS) loss that enables discriminative localization using only image-level labels. Concurrently, for inter-class transfer, an Adaptive Transfer Module (ATM) leverages a Class Relationship Graph (CRG), constructed using structured knowledge mined from a Large Language Model (LLM), and employs graph attention network to adaptively transfer relational information between class representations. DART is the first framework, to our knowledge, to explicitly integrate external LLM-derived relational knowledge for adaptive inter-class transfer while simultaneously performing adaptive intra-class refinement under weak supervision for OV-MLR. Extensive experiments on challenging benchmarks demonstrate that our DART achieves new state-of-the-art performance, validating its effectiveness.
Haijing Liu, Tao Pu 0002, Hefeng Wu, Keze Wang, Liang Lin 0004
ACM Multimedia3
2025 Diffusion-Driven 3D Gaussian Splatting for Occlusion-Free Egocentric Scene Reconstruction
abstract
Augmented reality and robotic navigation increasingly demand accurate and complete 3D scene reconstruction from egocentric viewpoints. However, dynamic occlusions—such as those caused by hand–object interactions—and frequent viewpoint changes often lead to persistent geometric incompleteness, hindering practical deployment. We present a diffusion-guided reconstruction framework that integrates 3D Gaussian Splatting (3DGS) with generative inpainting to achieve occlusion-free scene modeling. The method follows a two-stage pipeline: (1) Mask-guided Gaussian initialization constructs an occlusion-aware scene representation by explicitly excluding segmented occlusion regions; (2) Depth-aware diffusion refinement recovers missing structures through iterative cross-modal fusion of multi-view depth cues and pretrained semantic priors. This approach significantly improves scene completeness and visual fidelity over state-of-the-art methods. Extensive experiments demonstrate its robustness in occlusion-dense scenarios, especially in handling complex hand-induced occlusions, offering a practical solution for immersive augmented reality and robotic perception.
Roucheng Lai, Haijing Liu, Hefeng Wu
MMAsia3
2025 Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention
abstract
Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for understanding egocentric human behavior. However, achieving such segmentation robustly is challenging due to ambiguities inherent in egocentric videos and biases present in training data. Consequently, existing methods often struggle, learning spurious correlations from skewed object-action pairings in datasets and fundamental visual confounding factors of the egocentric perspective, such as rapid motion and frequent occlusions. To address these limitations, we introduce Causal Ego-REferring Segmentation (CERES), a plug-in causal framework that adapts strong, pre-trained RVOS backbones to the egocentric domain. CERES implements dual-modal causal intervention: applying backdoor adjustment principles to counteract language representation biases learned from dataset statistics, and leveraging front-door adjustment concepts to address visual confounding by intelligently integrating semantic visual features with geometric depth information guided by causal principles, creating representations more robust to egocentric distortions. Extensive experiments demonstrate that CERES achieves state-of-the-art performance on Ego-RVOS benchmarks, highlighting the potential of applying causal reasoning to build more reliable models for broader egocentric video understanding.
Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 0002, Keze Wang, Liang Lin 0004
NeurIPS3
2025 Improving Spatio-Temporal Awareness of Multimodal Large Language Models via Reinforcement Fine-Tuning
Hefeng Wu
PRCV (5)2
2025 SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
abstract
The class-agnostic counting (CAC) task has recently been proposed to solve the problem of counting all objects of an arbitrary class with several exemplars given in the input image. To address this challenging task, existing leading methods all resort to density map regression, which renders them impractical for downstream tasks that require object locations and restricts their ability to well explore the scale information of exemplars for supervision. Meanwhile, they generally model the interaction between the input image and the exemplars in an exemplar-by-exemplar way, which is inefficient and may not fully synthesize information from all exemplars. To address these limitations, we propose a novel localization-based CAC approach, termed Scale-modulated Query and Localization Network (SQLNet). It fully explores the scales of exemplars in both the query and localization stages and achieves effective counting by accurately locating each object and predicting its approximate size. Specifically, during the query stage, rich discriminative representations of the target class are acquired by the Hierarchical Exemplars Collaborative Enhancement (HECE) module from the few exemplars through multi-scale exemplar cooperation with equifrequent size prompt embedding. These representations are then fed into the Exemplars-Unified Query Correlation (EUQC) module to interact with the query features in a unified manner and produce the correlated query tensor. In the localization stage, the Scale-aware Multi-head Localization (SAML) module utilizes the query tensor to predict the confidence, location, and size of each potential object. Moreover, a scale-aware localization loss is introduced, which exploits flexible location associations and exemplar scales for supervision to optimize the model performance. Extensive experiments demonstrate that SQLNet outperforms state-of-the-art methods on popular CAC benchmarks, achieving excellent performance not only in counting accuracy but also in localization and bounding box generation.
Hefeng Wu, Yandong Chen 0002, Lingbo Liu, Tianshui Chen, Keze Wang, Liang Lin 0004
IEEE Trans. Image Process.1
2024 Dual-perspective semantic-aware representation blending for multi-label image recognition with partial labels
Tao Pu 0002, Tianshui Chen, Hefeng Wu, Yukai Shi, Zhijing Yang, Liang Lin 0004
Expert Syst. Appl.3
2024 Contrastive Transformer Learning With Proximity Data Generation for Text-Based Person Search
abstract
Given a descriptive text query, text-based person search (TBPS) aims to retrieve the best matched target person from an image gallery. Such a cross-modal retrieval task is quite challenging due to significant modality gap, fine-grained differences and insufficiency of annotated data. To better align the two modalities, most existing works focus on introducing sophisticated network structures and auxiliary tasks, which are complex and hard to implement. In this paper, we propose a simple yet effective dual Transformer model for text-based person search. By exploiting a hardness-aware contrastive learning strategy, our model achieves state-of-the-art performance without any special design for local feature alignment or side information. Moreover, we propose a proximity data generation (PDG) module to automatically produce more diverse data for cross-modal training. The PDG module first introduces an automatic generation algorithm based on a text-to-image diffusion model, which generates new text-image pair samples in the proximity space of original ones. Then it combines approximate text generation and feature-level mixup during training to further strengthen the data diversity. The PDG module can largely guarantee the reasonability of the generated samples that are directly used for training without any human inspection for noise rejection. It improves the performance of our model significantly, providing a feasible solution to the data insufficiency problem faced by such fine-grained visual-linguistic tasks. Extensive experiments on two popular datasets of the TBPS task (i.e., CUHK-PEDES and ICFG-PEDES) show that the proposed approach outperforms state-of-the-art approaches evidently, e.g., improving by 3.88%, 4.02%, 2.92% in terms of Top1, Top5, Top10 on CUHK-PEDES.
Hefeng Wu, Tianshui Chen, Zhiguang Chen 0001, Liang Lin 0004
IEEE Trans. Circuits Syst. Video Technol.1
2024 Spatial-Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation
abstract
Video scene graph generation (VidSGG) aims to identify objects in visual scenes and infer their relationships for a given video. It requires not only a comprehensive understanding of each object scattered on the whole scene but also a deep dive into their temporal motions and interactions. Inherently, object pairs and their relationships enjoy spatial co-occurrence correlations within each image and temporal consistency/transition correlations across different images, which can serve as prior knowledge to facilitate VidSGG model learning and inference. In this work, we propose a spatial-temporal knowledge-embedded transformer (STKET) that incorporates the prior spatial-temporal knowledge into the multi-head cross-attention mechanism to learn more representative relationship representations. Specifically, we first learn spatial co-occurrence and temporal transition correlations in a statistical manner. Then, we design spatial and temporal knowledge-embedded layers that introduce the multi-head cross-attention mechanism to fully explore the interaction between visual representation and the knowledge to generate spatial- and temporal-embedded representations, respectively. Finally, we aggregate these representations for each subject-object pair to predict the final semantic labels and their relationships. Extensive experiments show that STKET outperforms current competing algorithms by a large margin, e.g., improving the mR@50 by 8.1%, 4.7%, and 2.1% on different settings over current algorithms.
Tao Pu 0002, Tianshui Chen, Hefeng Wu, Yongyi Lu, Liang Lin 0004
IEEE Trans. Image Process.3
2024 Category-Adaptive Label Discovery and Noise Rejection for Multi-Label Recognition With Partial Positive Labels
abstract
As a cost-effective alternative to standard multi-label learning, the multi-label image recognition with partial positive labels (MLR-PPL) task attracts increasing attention, in which merely a portion of positive labels are given while the rest of positive labels and all negative labels are missing. To facilitate this task, we propose a novel framework that leverages semantic correlation among different images in a category-adaptive manner to complement unknown labels accurately. Specifically, the proposed framework consists of two complementary modules. 1) A category-adaptive label discovery (CALD) module is designed to measure the semantic similarity between positive samples and then complement unknown labels with high similarities. 2) A category-adaptive noise rejection (CANR) module is designed to compute the sample weights based on semantic similarities from different samples and discard noisy labels with low weights. Due to the various degrees of confidence calibration among different categories, searching appropriate thresholds for each category in the proposed framework is highly time-consuming. To avoid such a resource-intensive manual tuning, we introduce a category-adaptive threshold updating algorithm that introduces the category-specific positive and negative similarity to adjust the threshold adaptively. Extensive experiments on various benchmarks show that the proposed framework performs better than current state-of-the-art algorithms.
Tao Pu 0002, Qianru Lao, Hefeng Wu, Tianshui Chen, Ling Tian, Jie Liu 0022, Liang Lin 0004
IEEE Trans. Multim.3
2024 Improving Network Interpretability via Explanation Consistency Evaluation
abstract
While deep neural networks have achieved remarkable performance, they tend to lack transparency in prediction. The pursuit of greater interpretability in neural networks often results in a degradation of their original performance. Some works strive to improve both interpretability and performance, but they primarily depend on meticulously imposed conditions. In this paper, we propose a simple yet effective framework that acquires more explainable activation heatmaps and simultaneously increases the model performance, without the need for any extra supervision. Specifically, our concise framework introduces a new metric, i.e., explanation consistency, to reweight the training samples adaptively in model learning. The explanation consistency metric is utilized to measure the similarity between the model's visual explanations of the original samples and those of semantic-preserved adversarial samples, whose background regions are perturbed by using image adversarial attack techniques. Our framework then promotes the model learning by paying closer attention to those training samples with a high difference in explanations (i.e., low explanation consistency), for which the current model cannot provide robust interpretations. Comprehensive experimental results on various benchmarks demonstrate the superiority of our framework in multiple aspects, including higher recognition accuracy, greater data debiasing capability, stronger network robustness, and more precise localization ability on both regular networks and interpretable networks. We also provide extensive ablation studies and qualitative analyses to unveil the detailed contribution of each component.
Hefeng Wu, Keze Wang, Xianghuan He, Liang Lin 0004
IEEE Trans. Multim.1
2024 Dual-View Data Hallucination With Semantic Relation Guidance for Few-Shot Image Recognition
abstract
Learning to recognize novel concepts from just a few image samples is very challenging as the learned model is easily overfitted on the few data and results in poor generalizability. One promising but underexplored solution is to compensate for the novel classes by generating plausible samples. However, most existing works of this line exploit visual information only, rendering the generated data easy to be distracted by some challenging factors contained in the few available samples. Being aware of the semantic information in the textual modality that reflects human concepts, this work proposes a novel framework that exploits semantic relations to guide dual-view data hallucination for few-shot image recognition. The proposed framework enables generating more diverse and reasonable data samples for novel classes through effective information transfer from base classes. Specifically, an instance-view data hallucination module hallucinates each sample of a novel class to generate new data by employing local semantic correlated attention and global semantic feature fusion derived from base classes. Meanwhile, a prototype-view data hallucination module exploits semantic-aware measure to estimate the prototype of a novel class and the associated distribution from the few samples, which thereby harvests the prototype as a more stable sample and enables resampling a large number of samples. We conduct extensive experiments and comparisons with state-of-the-art methods on several popular few-shot benchmarks to verify the effectiveness of the proposed framework.
Hefeng Wu, Guangzhi Ye, Ziyang Zhou 0005, Ling Tian, Qing Wang 0018, Liang Lin 0004
IEEE Trans. Multim.1
2023 Multi-object Video Generation from Single Frame Layouts
abstract
In this paper, we study video synthesis with emphasis on simplifying the generation conditions. Most existing video synthesis models or datasets are designed to address complex motions of a single object, lacking the ability of comprehensively understanding the spatio-temporal relationships among multiple objects. Besides, current methods are usually conditioned on intricate annotations (e.g. video segmentations) to generate new videos, being fundamentally less practical. These motivate us to generate multi-object videos conditioning exclusively on object layouts from a single frame. To solve above challenges and inspired by recent research on image generation from layouts, we have proposed a novel video generative framework capable of synthesizing global scenes with local objects, via implicit neural representations and layout motion self-inference. Our framework is a non-trivial adaptation from image generation methods, and is new to this field. In addition, our model has been evaluated on two widely-used video recognition benchmarks, demonstrating effectiveness compared to the baseline model.
Hefeng Wu
ICME3
2023 Semantic representation and dependency learning for multi-label image recognition
Tao Pu 0002, Mingzhan Sun, Hefeng Wu, Tianshui Chen, Ling Tian, Liang Lin 0004
Neurocomputing3
2022 Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels
abstract
Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging and practical task. To address this task, current algorithms mainly depend on pre-training classification or similarity models to generate pseudo labels for the unknown labels. However, these algorithms depend on sufficient multi-label annotations to train the models, leading to poor performance especially with low known label proportion. In this work, we propose to blend category-specific representation across different images to transfer information of known labels to complement unknown labels, which can get rid of pre-training models and thus does not depend on sufficient annotations. To this end, we design a unified semantic-aware representation blending (SARB) framework that exploits instance-level and prototype-level semantic representation to complement unknown labels by two complementary modules: 1) an instance-level representation blending (ILRB) module blends the representations of the known labels in an image to the representations of the unknown labels in another image to complement these unknown labels. 2) a prototype-level representation blending (PLRB) module learns more stable representation prototypes for each category and blends the representation of unknown labels with the prototypes of corresponding labels to complement these labels. Extensive experiments on the MS-COCO, Visual Genome, Pascal VOC 2007 datasets show that the proposed SARB framework obtains superior performance over current leading competitors on all known label proportion settings, i.e., with the mAP improvement of 4.6%, 4.6%, 2.2% on these three datasets when the known label proportion is 10%. Codes are available at https://github.com/HCPLab-SYSU/HCP-MLR-PL.
Tao Pu 0002, Tianshui Chen, Hefeng Wu, Liang Lin 0004
AAAI3
2022 Structured Semantic Transfer for Multi-Label Recognition with Partial Labels
abstract
Multi-label image recognition is a fundamental yet practical task because real-world images inherently possess multiple semantic labels. However, it is difficult to collect large-scale multi-label annotations due to the complexity of both the input images and output label spaces. To reduce the annotation cost, we propose a structured semantic transfer (SST) framework that enables training multi-label recognition models with partial labels, i.e., merely some labels are known while other labels are missing (also called unknown labels) per image. The framework consists of two complementary transfer modules that explore within-image and cross-image semantic correlations to transfer knowledge of known labels to generate pseudo labels for unknown labels. Specifically, an intra-image semantic transfer module learns image-specific label co-occurrence matrix and maps the known labels to complement unknown labels based on this matrix. Meanwhile, a cross-image transfer module learns category-specific feature similarities and helps complement unknown labels with high similarities. Finally, both known and generated labels are used to train the multi-label recognition models. Extensive experiments on the Microsoft COCO, Visual Genome and Pascal VOC datasets show that the proposed SST framework obtains superior performance over current state-of-the-art algorithms. Codes are available at https://github.com/HCPLab-SYSU/HCP-MLR-PL.
Tianshui Chen, Tao Pu 0002, Hefeng Wu, Yuan Xie 0004, Liang Lin 0004
AAAI3
2022 Knowledge-Guided Multi-Label Few-Shot Learning for General Image Recognition
abstract
Recognizing multiple labels of an image is a practical yet challenging task, and remarkable progress has been achieved by searching for semantic regions and exploiting label dependencies. However, current works utilize RNN/LSTM to implicitly capture sequential region/label dependencies, which cannot fully explore mutual interactions among the semantic regions/labels and do not explicitly integrate label co-occurrences. In addition, these works require large amounts of training samples for each category, and they are unable to generalize to novel categories with limited samples. To address these issues, we propose a knowledge-guided graph routing (KGGR) framework, which unifies prior knowledge of statistical label correlations with deep neural networks. The framework exploits prior knowledge to guide adaptive information propagation among different categories to facilitate multi-label analysis and reduce the dependency of training samples. Specifically, it first builds a structured knowledge graph to correlate different labels based on statistical label co-occurrence. Then, it introduces the label semantics to guide learning semantic-specific features to initialize the graph, and it exploits a graph propagation network to explore graph node interactions, enabling learning contextualized image feature representations. Moreover, we initialize each graph node with the classifier weights for the corresponding label and apply another propagation network to transfer node messages through the graph. In this way, it can facilitate exploiting the information of correlated labels to help train better classifiers, especially for labels with limited training samples. We conduct extensive experiments on the traditional multi-label image recognition (MLR) and multi-label few-shot learning (ML-FSL) tasks and show that our KGGR framework outperforms the current state-of-the-art methods by sizable margins on the public benchmarks.
Tianshui Chen, Liang Lin 0004, Riquan Chen, Xiaolu Hui, Hefeng Wu
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning
abstract
Facial expression recognition (FER) has received significant attention in the past decade with witnessed progress, but data inconsistencies among different FER datasets greatly hinder the generalization ability of the models learned on one dataset to another. Recently, a series of cross-domain FER algorithms (CD-FERs) have been extensively developed to address this issue. Although each declares to achieve superior performance, comprehensive and fair comparisons are lacking due to inconsistent choices of the source/target datasets and feature extractors. In this work, we first propose to construct a unified CD-FER evaluation benchmark, in which we re-implement the well-performing CD-FER and recently published general domain adaptation algorithms and ensure that all these algorithms adopt the same source/target datasets and feature extractors for fair CD-FER evaluations. Based on the analysis, we find that most of the current state-of-the-art algorithms use adversarial learning mechanisms that aim to learn holistic domain-invariant features to mitigate domain shifts. However, these algorithms ignore local features, which are more transferable across different datasets and carry more detailed content for fine-grained adaptation. Therefore, we develop a novel adversarial graph representation adaptation (AGRA) framework that integrates graph representation propagation with adversarial learning to realize effective cross-domain holistic-local feature co-adaptation. Specifically, our framework first builds two graphs to correlate holistic and local regions within each domain and across different domains, respectively. Then, it extracts holistic-local features from the input image and uses learnable per-class statistical distributions to initialize the corresponding graph nodes. Finally, two stacked graph convolution networks (GCNs) are adopted to propagate holistic-local features within each domain to explore their interaction and across different domains for holistic-local feature co-adaptation. In this way, the AGRA framework can adaptively learn fine-grained domain-invariant features and thus facilitate cross-domain expression recognition. We conduct extensive and fair comparisons on the unified evaluation benchmark and show that the proposed AGRA framework outperforms previous state-of-the-art methods.
Tianshui Chen, Tao Pu 0002, Hefeng Wu, Yuan Xie 0004, Lingbo Liu, Liang Lin 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Physical-Virtual Collaboration Modeling for Intra- and Inter-Station Metro Ridership Prediction
abstract
Due to the widespread applications in real-world scenarios, metro ridership prediction is a crucial but challenging task in intelligent transportation systems. However, conventional methods either ignore the topological information of metro systems or directly learn on physical topology, and cannot fully explore the patterns of ridership evolution. To address this problem, we model a metro system as graphs with various topologies and propose a unified Physical-Virtual Collaboration Graph Network (PVCGN), which can effectively learn the complex ridership patterns from the tailor-designed graphs. Specifically, a physical graph is directly built based on the realistic topology of the studied metro system, while a similarity graph and a correlation graph are built with virtual topologies under the guidance of the inter-station passenger flow similarity and correlation. These complementary graphs are incorporated into a Graph Convolution Gated Recurrent Unit (GC-GRU) for spatial-temporal representation learning. Further, a Fully-Connected Gated Recurrent Unit (FC-GRU) is also applied to capture the global evolution tendency. Finally, we develop a Seq2Seq model with GC-GRU and FC-GRU to forecast the future metro ridership sequentially. Extensive experiments on two large-scale benchmarks (e.g., Shanghai Metro and Hangzhou Metro) well demonstrate the superiority of our PVCGN for station-level metro ridership prediction. Moreover, we apply the proposed PVCGN to address the online origin-destination (OD) ridership prediction and the experiment results show the universality of our method. Our code and benchmarks are available athttps://github.com/HCPLab-SYSU/PVCGN.
Lingbo Liu, Hefeng Wu, Jiajie Zhen, Guanbin Li, Liang Lin 0004
IEEE Trans. Intell. Transp. Syst.3
2021 Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting
abstract
Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we find that incorporating optical and thermal information can greatly help to recognize pedestrians. To promote future researches in this field, we introduce a large-scale RGBT Crowd Counting (RGBT-CC) benchmark, which contains 2,030 pairs of RGB-thermal images with 138,389 annotated people. Furthermore, to facilitate the multimodal crowd counting, we propose a cross-modal collaborative representation learning framework, which consists of multiple modality-specific branches, a modality-shared branch, and an Information Aggregation-Distribution Module (IADM) to capture the complementary information of different modalities fully. Specifically, our IADM incorporates two collaborative information transfers to dynamically enhance the modality-shared and modality-specific representations with a dual information propagation mechanism. Extensive experiments conducted on the RGBT-CC benchmark demonstrate the effectiveness of our framework for RGBT crowd counting. Moreover, the proposed approach is universal for multimodal crowd counting and is also capable to achieve superior performance on the ShanghaiTechRGBD [22] dataset. Finally, our source code and benchmark have been released at http://lingboliu.com/RGBT_Crowd_Counting.html.
Lingbo Liu, Hefeng Wu, Guanbin Li, Chenglong Li 0002, Liang Lin 0004
CVPR3
2021 AU-Expression Knowledge Constrained Representation Learning for Facial Expression Recognition
abstract
Recognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Current deep-learning-based algorithms may achieve impressive performance in some lab-controlled environments, but they always fail to recognize the expressions accurately for the uncontrolled in-the-wild situation. Fortunately, facial action units (AU) describe subtle facial behaviors, and they can help distinguish uncertain and ambiguous expressions. In this work, we explore the correlations among the action units and facial expressions, and devise an AU-Expression Knowledge Constrained Representation Learning (AUE-CRL) framework to learn the AU representations without AU annotations and adaptively use representations to facilitate facial expression recognition. Specifically, it leverages AU-expression correlations to guide the learning of the AU classifiers, and thus it can obtain AU representations without incurring any AU annotations. Then, it introduces a knowledge-guided attention mechanism that mines useful AU representations under the constraint of AU-expression correlations. In this way, the framework can capture local discriminative and complementary features to enhance facial representation for facial expression recognition. We conduct experiments on the challenging uncontrolled datasets to demonstrate the superiority of the proposed framework over current state-of-the-art methods. Codes and trained models are available at https://github.com/HCPLab-SYSU/AUE-CRL.
Tao Pu 0002, Tianshui Chen, Yuan Xie 0004, Hefeng Wu, Liang Lin 0004
ICRA4
2021 A survey of script learning
abstract
Script is the structured knowledge representation of prototypical real-life event sequences. Learning the commonsense knowledge inside the script can be helpful for machines in understanding natural language and drawing commonsensible inferences. Script learning is an interesting and promising research direction, in which a trained script learning system can process narrative texts to capture script knowledge and draw inferences. However, there are currently no survey articles on script learning, so we are providing this comprehensive survey to deeply investigate the standard framework and the major research topics on script learning. This research field contains three main topics: event representations, script learning models, and evaluation approaches. For each topic, we systematically summarize and categorize the existing script learning systems, and carefully analyze and compare the advantages and disadvantages of the representative systems. We also discuss the current state of the research and possible future directions.
Linbo Qiao, Jianming Zheng, Hefeng Wu, Dongsheng Li 0001, Xiangke Liao
Frontiers Inf. Technol. Electron. Eng.4
2021 Fine-Grained Image Captioning With Global-Local Discriminative Objective
abstract
Significant progress has been made in recent years in image captioning, an active topic in the fields of vision and language. However, existing methods tend to yield overly general captions and consist of some of the most frequent words/phrases, resulting in inaccurate and indistinguishable descriptions (see Fig. 1). This is primarily due to (i) the conservative characteristic of traditional training objectives that drives the model to generate correct but hardly discriminative captions for similar images and (ii) the uneven word distribution of the ground-truth captions, which encourages generating highly frequent words/phrases while suppressing the less frequent but more concrete ones. In this work, we propose a novel global-local discriminative objective that is formulated on top of a reference model to facilitate generating fine-grained descriptive captions. Specifically, from a global perspective, we design a novel global discriminative constraint that pulls the generated sentence to better discern the corresponding image from all others in the entire dataset. From the local perspective, a local discriminative constraint is proposed to increase attention such that it emphasizes the less frequent but more concrete words/phrases, thus facilitating the generation of captions that better describe the visual details of the given images. We evaluate the proposed method on the widely used MS-COCO dataset, where it outperforms the baseline methods by a sizable margin and achieves competitive performance over existing leading approaches. We also conduct self-retrieval experiments to demonstrate the discriminability of the proposed method.
Jie Wu 0030, Tianshui Chen, Hefeng Wu, Zhi Yang 0004, Guangchun Luo, Liang Lin 0004
IEEE Trans. Multim.3
2020 Knowledge Graph Transfer Network for Few-Shot Recognition
abstract
Few-shot learning aims to learn novel categories from very few samples given some base categories with sufficient training samples. The main challenge of this task is the novel categories are prone to dominated by color, texture, shape of the object or background context (namely specificity), which are distinct for the given few training samples but not common for the corresponding categories (see Figure 1). Fortunately, we find that transferring information of the correlated based categories can help learn the novel concepts and thus avoid the novel concept being dominated by the specificity. Besides, incorporating semantic correlations among different categories can effectively regularize this information transfer. In this work, we represent the semantic correlations in the form of structured knowledge graph and integrate this graph into deep neural networks to promote few-shot learning by a novel Knowledge Graph Transfer Network (KGTN). Specifically, by initializing each node with the classifier weight of the corresponding category, a propagation mechanism is learned to adaptively propagate node message through the graph to explore node interaction and transfer classifier information of the base categories to those of the novel ones. Extensive experiments on the ImageNet dataset show significant performance improvement compared with current leading competitors. Furthermore, we construct an ImageNet-6K dataset that covers larger scale categories, i.e, 6,000 categories, and experiments on this dataset further demonstrate the effectiveness of our proposed model.
Riquan Chen, Tianshui Chen, Xiaolu Hui, Hefeng Wu, Guanbin Li, Liang Lin 0004
AAAI4
2020 Exploiting Knowledge Embedded Soft Labels for Image Recognition
abstract
Objects from correlated classes usually share highly similar appearance while objects from uncorrelated classes are very different. Most of current image recognition works treat each class independently, which ignores these class correlations and inevitably leads to sub-optimal performance in many cases. Fortunately, object classes inherently form a hierarchy with different levels of abstraction and this hierarchy encodes rich correlations among different classes. In this work, we utilize a soft label vector that encodes the prior knowledge of class correlations as extra regularization to train the image classifiers. Specifically, for each class, instead of simply using a one-hot vector, we assign a high value to its correlated classes and assign small values to those uncorrelated ones, thus generating knowledge embedded soft labels. We conduct experiments on both general and fine-grained image recognition benchmarks and demonstrate its superiority compared with existing methods.
Lixian Yuan, Riquan Chen, Hefeng Wu, Tianshui Chen
ICPR3
2020 Adversarial Graph Representation Adaptation for Cross-Domain Facial Expression Recognition
abstract
Data inconsistency and bias are inevitable among different facial expression recognition (FER) datasets due to subjective annotating process and different collecting conditions. Recent works resort to adversarial mechanisms that learn domain-invariant features to mitigate domain shift. However, most of these works focus on holistic feature adaptation, and they ignore local features that are more transferable across different datasets. Moreover, local features carry more detailed and discriminative content for expression recognition, and thus integrating local features may enable fine-grained adaptation. In this work, we propose a novel Adversarial Graph Representation Adaptation (AGRA) framework that unifies graph representation propagation with adversarial learning for cross-domain holistic-local feature co-adaptation. To achieve this, we first build a graph to correlate holistic and local regions within each domain and another graph to correlate these regions across different domains. Then, we learn the per-class statistical distribution of each domain and extract holistic-local features from the input image to initialize the corresponding graph nodes. Finally, we introduce two stacked graph convolution networks to propagate holistic-local feature within each domain to explore their interaction and across different domains for holistic-local feature co-adaptation. In this way, the AGRA framework can adaptively learn fine-grained domain-invariant features and thus facilitate cross-domain expression recognition. We conduct extensive and fair experiments on several popular benchmarks and show that the proposed AGRA framework achieves superior performance over previous state-of-the-art methods.
Yuan Xie 0004, Tianshui Chen, Tao Pu 0002, Hefeng Wu, Liang Lin 0004
ACM Multimedia4
2020 Active Object Search
Jie Wu 0030, Tianshui Chen, Lishan Huang, Hefeng Wu, Guanbin Li, Ling Tian, Liang Lin 0004
ACM Multimedia4
2020 Efficient Crowd Counting via Structured Knowledge Transfer
abstract
Crowd counting is an application-oriented task and its inference efficiency is crucial for real-world applications. However, most previous works relied on heavy backbone networks and required prohibitive run-time consumption, which would seriously restrict their deployment scopes and cause poor scalability. To liberate these crowd counting models, we propose a novel Structured Knowledge Transfer (SKT) framework, which fully exploits the structured knowledge of a well-trained teacher network to generate a lightweight but still highly effective student network. Specifically, it is integrated with two complementary transfer modules, including an Intra-Layer Pattern Transfer which sequentially distills the knowledge embedded in layer-wise features of the teacher network to guide feature learning of the student network and an Inter-Layer Relation Transfer which densely distills the cross-layer correlation knowledge of the teacher to regularize the student's feature evolution. Consequently, our student network can derive the layer-wise and cross-layer knowledge from the teacher network to learn compact yet effective features. Extensive evaluations on three benchmarks well demonstrate the effectiveness of our SKT for extensive crowd counting models. In particular, only using around $6%$ of the parameters and computation cost of original models, our distilled VGG-based models obtain at least 6.5× speed-up on an Nvidia 1080 GPU and even achieve state-of-the-art performance. Our code and models are available at https://github.com/HCPLab-SYSU/SKT.
Lingbo Liu, Hefeng Wu, Tianshui Chen, Guanbin Li, Liang Lin 0004
ACM Multimedia3
2020 Multi-column point-CNN for sketch segmentation
Fei Wang 0056, Shujin Lin, Hefeng Wu, Tie Cai, Ruomei Wang 0001
Neurocomputing4
2020 Crowd counting via scale-communicative aggregation networks
Lixian Yuan, Zhilin Qiu, Lingbo Liu, Hefeng Wu, Tianshui Chen, Liang Lin 0004
Neurocomputing4
2019 ADCrowdNet: An Attention-Injective Deformable Convolutional Network for Crowd Understanding
abstract
We propose an attention-injective deformable convolutional network called ADCrowdNet for crowd understanding that can address the accuracy degradation problem of highly congested noisy scenes. ADCrowdNet contains two concatenated networks. An attention-aware network called Attention Map Generator (AMG) first detects crowd regions in images and computes the congestion degree of these regions. Based on detected crowd regions and congestion priors, a multi-scale deformable network called Density Map Estimator (DME) then generates high-quality density maps. With the attention-aware training scheme and multi-scale deformable convolutional scheme, the proposed ADCrowdNet achieves the capability of being more effective to capture the crowd features and more resistant to various noises. We have evaluated our method on four popular crowd counting datasets (ShanghaiTech, UCF_CC_50, WorldEXPO'10, and UCSD) and an extra vehicle counting dataset TRANCOS, and our approach beats existing state-of-the-art approaches on all of these datasets.
Yongchao Long, Changqing Zou, Qun Niu, Li Pan 0002, Hefeng Wu
CVPR6
2019 Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition
abstract
Recognizing multiple labels of images is a practical and challenging task, and significant progress has been made by searching semantic-aware regions and modeling label dependency. However, current methods cannot locate the semantic regions accurately due to the lack of part-level supervision or semantic guidance. Moreover, they cannot fully explore the mutual interactions among the semantic regions and do not explicitly model the label co-occurrence. To address these issues, we propose a Semantic-Specific Graph Representation Learning (SSGRL) framework that consists of two crucial modules: 1) a semantic decoupling module that incorporates category semantics to guide learning semantic-specific representations and 2) a semantic interaction module that correlates these representations with a graph built on the statistical label co-occurrence and explores their interactions via a graph propagation mechanism. Extensive experiments on public benchmarks show that our SSGRL framework outperforms current state-of-the-art methods by a sizable margin, e.g. with an mAP improvement of 2.5%, 2.6%, 6.7%, and 3.1% on the PASCAL VOC 2007 & 2012, Microsoft-COCO and Visual Genome benchmarks, respectively. Our codes and models are available at https://github.com/HCPLab-SYSU/SSGRL.
Tianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu, Liang Lin 0004
ICCV4
2019 SPFusionNet: Sketch Segmentation Using Multi-modal Data Fusion
abstract
The sketch segmentation problem remains largely unsolved because conventional methods are greatly challenged by the highly abstract appearances of freehand sketches and their numerous shape variations. In this work, we tackle such challenges by exploiting different modes of sketch data in a unified framework. Specifically, we propose a deep neural network SPFusionNet to capture the characteristic of sketch by fusing from its image and point set modes. The image modal component SketchNet learns hierarchically abstract ro-bust features and utilizes multi-level representations to produce pixel-wise feature maps, while the point set-modal component SPointNet captures local and global contexts of the sampled point set to produce point-wise feature maps. Then our framework aggregates these feature maps by a fusion network component to generate the sketch segmentation result. The extensive experimental evaluation and comparison with peer methods on our large SketchSeg dataset verify the effectiveness of the proposed framework.
Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Xiangjian He
ICME3
2019 Concrete Image Captioning by Integrating Content Sensitive and Global Discriminative Objective
abstract
Current methods for image captioning tend to generate sentences that are generally overly rigid and composed of some most frequent words/phrases, leading to inaccurate and indistinguishable descriptions. This is primarily due to the uneven word distribution of the ground truth captions that encourages to generate high frequent words/phrases while suppressing the less frequent but more concrete ones. In this work, we propose a new Content Sensitive and Global Discriminative objective, which is formulated as two constraints on top of a reference model to facilitate generating concrete and discriminative image captions. More specifically, the content sensitive constraint is designed to place greater focus on the less frequent and more concrete words/phrases, thus facilitating the generation of sentences that better describe visual details of the given images. To further improve the discriminability, the global discriminative constraint is designed to pull the generated sentence to better discern the corresponding image from others. We evaluate the proposed method on the widely used MS-COCO dataset, where it achieves superior performance over existing competing methods. We also conduct self-retrieval experiments to demonstrate the discriminability of the proposed method.
Jie Wu 0030, Tianshui Chen, Hefeng Wu, Zhi Yang 0004, Qing Wang 0018, Liang Lin 0004
ICME3
2019 Instance-aware representation learning and association for online multi-person tracking
Hefeng Wu, Yafei Hu, Keze Wang, Lin Nie
Pattern Recognit.1
2018 Learning deep similarity models with focus ranking for fabric image retrieval
Daiguo Deng, Ruomei Wang 0001, Hefeng Wu, Huayong He, Qi Li 0001
Image Vis. Comput.3
2018 PSI: A probabilistic semantic interpretable framework for fine-grained image ranking
abstract
Image Ranking is one of the key problems in information science research area. However, most current methods focus on increasing the performance, leaving the semantic gap problem, which refers to the learned ranking models are hard to be understood, remaining intact. Therefore, in this article, we aim at learning an interpretable ranking model to tackle the semantic gap in fine‐grained image ranking. We propose to combine attribute‐based representation and online passive‐aggressive (PA) learning based ranking models to achieve this goal. Besides, considering the highly localized instances in fine‐grained image ranking, we introduce a supervised constrained clustering method to gather class‐balanced training instances for local PA‐based models, and incorporate the learned local models into a unified probabilistic framework. Extensive experiments on the benchmark demonstrate that the proposed framework outperforms state‐of‐the‐art methods in terms of accuracy and speed.
Hefeng Wu, Shujin Lin, Zhuo Su 0001
J. Assoc. Inf. Sci. Technol.2
2018 Weak-structure-aware visual object tracking with bottom-up and top-down context exploration
Hefeng Wu, Hengzheng Zhu, Jin Zhan
Signal Process. Image Commun.3
2017 A Feature Preserved Mesh Subdivision Framework for Biomedical Mesh
abstract
As biomedical data in 3D space collected increasingly, there is a pressing need for efficient and accurate applications in the field of bioinformation analysis. For biomedical purpose, mesh subdivision techniques are commonly used to generate adaptive multi-resolution meshes for fast or accurate algorithms. However, current smoothing methods for each subdivision algorithm will moderate edge and vertex features from the original mesh. In this paper, we propose a feature preserved mesh subdivision framework, which generates a visually sensitive and a more precise result compared with commonly used subdivision methods, to preserve edge and vertex geometrical features of biomedical data.
Yongyi Gong, Hefeng Wu, Qi Li 0001
BIBE3
2017 Pedestrian Detection via Structure-Sensitive Deep Representation Learning
Deliang Huang, Shijia Huang, Hefeng Wu
ICIG (1)3
2017 Multi-view pairwise relationship learning for sketch based 3D shape retrieval
abstract
Recent progress in sketch-based 3D shape retrieval creates a novel and user-friendly way to explore massive 3D shapes on the Internet. However, current methods on this topic rely on designing invariant features for both sketches and 3D shapes, or complex matching strategies. Therefore, they suffer from problems like arbitrary drawings and inconsistent viewpoints. To tackle this problem, we propose a probabilistic framework based on Multi-View Pairwise Relationship (MVPR) learning. Our framework includes multiple views of 3D shapes as the intermediate layer between sketches and 3D shapes, and transforms the original retrieval problem into the form of inferring pairwise relationship between sketches and views. We accomplish pairwise relationship inference by a novel MVPR net, which can automatically predict and merge the pairwise relationships between a sketch and multiple views, thus freeing us from exhaustively selecting the best view of 3D shapes. We also propose to learn robust features for sketches and views via fine-tuning pre-trained networks. Extensive experiments on a large dataset demonstrate that the proposed method can outperform state-of-the-art methods significantly.
Hefeng Wu, Xiangjian He, Shujin Lin, Ruomei Wang 0001
ICME2
2017 A Data-Driven Approach for Sketch-Based 3D Shape Retrieval via Similar Drawing-Style Recommendation
abstract
Abstract Sketching is a simple and natural way of expression and communication for humans. For this reason, it gains increasing popularity in human computer interaction, with the emergence of multitouch tablets and styluses. In recent years, sketch‐based interactive methods are widely used in many retrieval systems. In particular, a variety of sketch‐based 3D model retrieval works have been presented. However, almost all of these works focus on directly matching sketches with the projection views of 3D models, and they suffer from the large differences between the sketch drawing and the views of 3D models, leading to unsatisfying retrieval results. Therefore, in this paper, during the matching procedure in the retrieval, we propose to match the sketch with each 3D model from historical users instead of projection views. Yet since the sketches between the current user and the historical users can have big difference, we also aim to handle users' personalized deviations and differences. To this end, we leverage recommendation algorithms to estimate the drawing style characteristic similarity between the current user and historical users. Experimental results on the Large Scale Sketch Track Benchmark(SHREC14LSSTB) demonstrate that our method outperforms several state‐of‐the‐art methods.
Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Fan Zhou 0001
Comput. Graph. Forum4
2017 Data-driven image completion for complex objects
Chengying Gao, Yanmei Luo, Hefeng Wu, Dong Wang 0041
Signal Process. Image Commun.3
2017 Distortion-Aware Correlation Tracking
abstract
Recently, correlation filter (CF)-based tracking methods have attracted considerable attention because of their high-speed performance. However, distortion, which refers to the phenomenon that the correlation outputs of CF-based trackers are distorted, remains a major obstacle for these methods. In this paper, we propose a distortion-aware correlation filter framework, which can detect distortions and recover from tracking failures. Our framework employs a simple yet effective feature termed normed correlation response to detect distortions. Meanwhile, we introduce a competition mechanism to handle distortions, in which we build a specialized graph to formulate and handle tracking under distortion as a maximum multi clique problem. Furthermore, a global-local context model is exploited to alleviate underlying distortions during the tracking process. Extensive experiments on the Online Tracking Benchmark show that our tracker can find the optimal target trajectory during the distortion period and retrieve the possibly missing target, consequently outperforms the state-of-the-art methods and improves the performance of CF-based trackers favorably.
Hefeng Wu, Huifang Zhang, Shujin Lin, Ruomei Wang 0001
IEEE Trans. Image Process.2
2016 Boosting Zero-Shot Image Classification via Pairwise Relationship Learning
Hefeng Wu, Shujin Lin, Ebroul Izquierdo
ACCV (1)2
2015 Cascaded probabilistic tracking with supervised dictionary learning
Jin Zhan, Hefeng Wu, Huifang Zhang
Signal Process. Image Commun.2
2015 Hierarchical Ensemble of Background Models for PTZ-Based Video Surveillance
abstract
In this paper, we study a novel hierarchical background model for intelligent video surveillance with the pan-tilt-zoom (PTZ) camera, and give rise to an integrated system consisting of three key components: background modeling, observed frame registration, and object tracking. First, we build the hierarchical background model by separating the full range of continuous focal lengths of a PTZ camera into several discrete levels and then partitioning the wide scene at each level into many partial fixed scenes. In this way, the wide scenes captured by a PTZ camera through rotation and zoom are represented by a hierarchical collection of partial fixed scenes. A new robust feature is presented for background modeling of each partial scene. Second, we locate the partial scenes corresponding to the observed frame in the hierarchical background model. Frame registration is then achieved by feature descriptor matching via fast approximate nearest neighbor search. Afterwards, foreground objects can be detected using background subtraction. Last, we configure the hierarchical background model into a framework to facilitate existing object tracking algorithms under the PTZ camera. Foreground extraction is used to assist tracking an object of interest. The tracking outputs are fed back to the PTZ controller for adjusting the camera properly so as to maintain the tracked object in the image plane. We apply our system on several challenging scenarios and achieve promising results.
Hefeng Wu
IEEE Trans. Cybern.2
2015 Robust tracking via discriminative sparse feature selection
Jin Zhan, Zhuo Su 0001, Hefeng Wu
Vis. Comput.3
2014 Weighted attentional blocks for probabilistic object tracking
Hefeng Wu, Guanbin Li
Vis. Comput.1
2012 Online boosted tracking with discriminative feature selection and scale adaptation
abstract
We track the object by separating it from the surrounding with an ensemble of boosted classifiers, which are trained in a discriminative feature space that is determined on the fly. Contour refinement and weight thresholding techniques are used to select good examples for training. While tracking, location calibration and scale adaptation are used to improve the tracker's performance. We update the ensemble of weak classifiers online to adapt to appearance changes, and use the positive occupancy ratio to detect occlusion. A center-surround discrepancy measure is presented to evaluate the discriminative power of the current feature space and to invoke re-initialization of feature selection and classifier training if necessary. Experiments on challenging video sequences demonstrate the effectiveness of the proposed approach.
Hefeng Wu, Guanbin Li, Zhuo Su 0001
ICIP1