Wei-Hong Li 0001

dblp:255/5590-1 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0003-4942-822XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2025 UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
abstract
Traditional spatiotemporal models generally rely on task-specific architectures, which limit their generalizability and scalability across diverse tasks due to domain-specific design requirements. In this paper, we introduce UniSTD, a unified Transformer-based framework for spatiotemporal modeling, which is inspired by advances in recent foundation models with the two-stage pretraining-then-adaption paradigm. Specifically, our work demonstrates that task-agnostic pretraining on 2D vision and vision-text datasets can build a generalizable model foundation for spatiotemporal learning, followed by specialized joint training on spatiotemporal datasets to enhance task-specific adaptability. To improve the learning capabilities across domains, our framework employs a rank-adaptive mixture-of-expert adaptation by using fractional interpolation to relax the discrete variables so that can be optimized in the continuous space. Additionally, we introduce a temporal module to incorporate temporal dynamics explicitly. We evaluate our approach on a large-scale dataset covering 10 tasks across 4 disciplines, demonstrating that a unified spatiotemporal model can achieve scalable, cross-task learning and support up to 10 tasks simultaneously within one model while reducing training costs in multi-domain applications. Code will be available at https://github.com/1hunters/UniSTD.
Xinzhu Ma, Encheng Su, Xiufeng Song, Xiaohong Liu 0001, Wei-Hong Li 0001, Lei Bai 0001, Wanli Ouyang, Xiangyu Yue 0001
CVPR6
2025 FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent Directions
Yilei Jiang, Wei-Hong Li 0001, Minghong Cai, Xiangyu Yue 0001
ICCV2
2024 Multi-task Learning with 3D-Aware Regularization
abstract
Deep neural networks have become the standard solution for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmentation thanks to their ability to capture complex correlations in high dimensional feature space across tasks. However, the cross-task correlations that are learned in the unstructured feature space can be extremely noisy and susceptible to overfitting, consequently hurting performance. We propose to address this problem by introducing a structured 3D-aware regularizer which interfaces multiple tasks through the projection of features extracted from an image encoder to a shared 3D feature space and decodes them into their task output space through differentiable rendering. We show that the proposed method is architecture agnostic and can be plugged into various prior multi-task backbones to improve their performance; as we evidence using standard benchmarks NYUv2 and PASCAL-Context.
Wei-Hong Li 0001, Steven McDonagh 0001, Ales Leonardis, Hakan Bilen
ICLR1
2024 Bifröst: 3D-Aware Image Compositing with Language Instructions
Kaixiong Gong, Wei-Hong Li 0001, Xili Dai, Tao Chen 0003, Xiangyu Yue 0001
NeurIPS3
2024 Universal Representations: A Unified Look at Multiple Task and Domain Learning
abstract
Abstract We propose a unified look at jointly learning multiple vision tasks and visual domains through universal representations, a single deep neural network. Learning multiple problems simultaneously involves minimizing a weighted sum of multiple loss functions with different magnitudes and characteristics and thus results in unbalanced state of one loss dominating the optimization and poor results compared to learning a separate model for each problem. To this end, we propose distilling knowledge of multiple task/domain-specific networks into a single deep neural network after aligning its representations with the task/domain-specific ones through small capacity adapters. We rigorously show that universal representations achieve state-of-the-art performances in learning of multiple dense prediction problems in NYU-v2 and Cityscapes, multiple image classification problems from diverse domains in Visual Decathlon Dataset and cross-domain few-shot learning in MetaDataset. Finally we also conduct multiple analysis through ablation and qualitative studies.
Wei-Hong Li 0001, Xialei Liu, Hakan Bilen
Int. J. Comput. Vis.1
2023 Learning Relation Models to Detect Important People in Still Images
abstract
Important people detection aims to identify the most important people (i.e., the people who play the main roles in scenes) in images, which is challenging since people's importance in images depends not only on their appearance but also on their interactions with others (i.e., relations among people) and their roles in the scene (i.e., relations between people and underlying events). In this work, we propose the People Relation Network (PRN) to solve this problem. PRN consists of three modules (i.e., the feature representation, relation and classification modules) to extract visual features, model relations and estimate people's importance, respectively. The relation module contains two submodules to model two types of relations, namely, the person-person relation submodule and the person-event relation submodule. The person-person relation submodule infers the relations among people from the interaction graph and the person-event relation submodule models the relations between people and events by considering the spatial correspondence between features. With the help of them, PRN can effectively distinguish important people from other individuals. Extensive experiments on the Multi-Scene Important People (MS) and NCAA Basketball Image (NCAA) datasets show that PRN achieves state-of-the-art performance and generalizes well when available data is limited.
Yukun Qiu, Fa-Ting Hong, Wei-Hong Li 0001, Wei-Shi Zheng 0001
IEEE Trans. Multim.3
2022 Cross-domain Few-shot Learning with Task-specific Adapters
abstract
In this paper, we look at the problem of cross-domain few-shot classification that aims to learn a classifier from previously unseen classes and domains withfew labeled samples. Recent approaches broadly solve this problem by pa-rameterizing their few-shot classifiers with task-agnostic and task-specific weights where the former is typically learned on a large training set and the latter is dynamically predicted through an auxiliary network conditioned on a small support set. In this work, we focus on the estimation of the latter, and propose to learn task-specific weights from scratch directly on a small support set, in contrast to dynamically estimating them. In particular, through systematic analysis, we show that task-specific weights through parametric adapters in matrix form with residual connections to multiple intermediate layers of a backbone network significantly improves the per-formance of the state-of-the-art models in the Meta-Dataset benchmark with minor additional cost.
Wei-Hong Li 0001, Xialei Liu, Hakan Bilen
CVPR1
2022 Learning Multiple Dense Prediction Tasks from Partially Annotated Data
abstract
Despite the recent advances in multi-task learning of dense prediction problems, most methods rely on expensive labelled datasets. In this paper, we present a label efficient approach and look at jointly learning of multiple dense prediction tasks on partially annotated data (i.e. not all the task labels are available for each image), which we call multi-task partially-supervised learning. We propose a multi-task training procedure that successfully leverages task relations to supervise its multi-task learning when data is partially annotated. In particular, we learn to map each task pair to a joint pairwise task-space which enables sharing information between them in a computationally efficient way through another network conditioned on task pairs, and avoids learning trivial cross-task relations by retaining high-level information about the input image. We rigorously demonstrate that our proposed method effectively exploits the images with unlabelled tasks and outperforms existing semi-supervised learning approaches and related methods on three standard benchmarks.
Wei-Hong Li 0001, Xialei Liu, Hakan Bilen
CVPR1
2021 Universal Representation Learning from Multiple Domains for Few-shot Classification
abstract
In this paper, we look at the problem of few-shot image classification that aims to learn a classifier for previously unseen classes and domains from few labeled samples. Recent methods use various adaptation strategies for aligning their visual representations to new domains or select the relevant ones from multiple domain-specific feature extractors. In this work, we present URL, which learns a single set of universal visual representations by distilling knowledge of multiple domain-specific networks after co-aligning their features with the help of adapters and centered kernel alignment. We show that the universal representations can be further refined for previously unseen domains by an efficient adaptation step in a similar spirit to distance learning methods. We rigorously evaluate our model in the recent Meta-Dataset benchmark and demonstrate that it significantly outperforms the previous methods while being more efficient.
Wei-Hong Li 0001, Xialei Liu, Hakan Bilen
ICCV1
2020 Learning to Detect Important People in Unlabelled Images for Semi-Supervised Important People Detection
abstract
Important people detection is to automatically detect the individuals who play the most important roles in a social event image, which requires the designed model to understand a high-level pattern. However, existing methods rely heavily on supervised learning using large quantities of annotated image samples, which are more costly to collect for important people detection than for individual entity recognition (i.e., object recognition). To overcome this problem, we propose learning important people detection on partially annotated images. Our approach iteratively learns to assign pseudo-labels to individuals in un-annotated images and learns to update the important people detection model based on data with both labels and pseudo-labels. To alleviate the pseudo-labelling imbalance problem, we introduce a ranking strategy for pseudo-label estimation, and also introduce two weighting strategies: one for weighting the confidence that individuals are important people to strengthen the learning on important people and the other for neglecting noisy unlabelled images (i.e., images without any important people). We have collected two large-scale datasets for evaluation. The extensive experimental results clearly confirm the efficacy of our method attained by leveraging unlabelled images for improving the performance of important people detection.
Fa-Ting Hong, Wei-Hong Li 0001, Wei-Shi Zheng 0001
CVPR2
2020 MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection
Fa-Ting Hong, Xuanteng Huang, Wei-Hong Li 0001, Wei-Shi Zheng 0001
ECCV (13)3
2019 Learning to Learn Relation for Important People Detection in Still Images
abstract
Humans can easily recognize the importance of people in social event images, and they always focus on the most important individuals. However, learning to learn the relation between people in an image, and inferring the most important person based on this relation, remains undeveloped. In this work, we propose a deep imPOrtance relatIon NeTwork (POINT) that combines both relation modeling and feature learning. In particular, we infer two types of interaction modules: the person-person interaction module that learns the interaction between people and the event-person interaction module that learns to describe how a person is involved in the event occurring in an image. We then estimate the importance relations among people from both interactions and encode the relation feature from the importance relations. In this way, POINT automatically learns several types of relation features in parallel, and we aggregate these relation features and the person's feature to form the importance feature for important people classification. Extensive experimental results show that our method is effective for important people detection and verify the efficacy of learning to learn relations for important people detection.
Wei-Hong Li 0001, Fa-Ting Hong, Wei-Shi Zheng 0001
CVPR1
2019 Towards Photo-Realistic Visible Watermark Removal with Conditional Generative Adversarial Networks
Xiang Li 0032, Chan Lu, Danni Cheng, Wei-Hong Li 0001, Mei Cao, Jiechao Ma, Wei-Shi Zheng 0001
ICIG (1)4
2019 One-pass person re-identification by sketch online discriminant analysis
Wei-Hong Li 0001, Allen Z. Zhong, Wei-Shi Zheng 0001
Pattern Recognit.1
2018 PersonRank: Detecting Important People in Images
abstract
Always, some individuals in images are more important/attractive than the others in some events such as presentation, basketball game or speech. However, it is challenging to ?nd important people among all individuals in an image directly based on their spatial or appearance information due to the existence of diverse variations of pose, action, appearance of persons an various changes of occasions. We overcome this challenge by constructing a multiple HyperInteraction Graph that treats each individual in an image as a node and inferring the most active node from the interactions estimated by using various types of cues. We model a pairwise interaction between people as an edge message communicated between nodes, resulting in a bidirectional pairwise-interaction graph. To enrich the person-person interaction estimation, we further introduce a unidirectional hyper-interaction graph that models the consensus of interactions between a focal person and any person in his/her local region around. Finally, we modify the PageRank algorithm to infer the activeness of people on the multiple Hybrid-Interaction Graph (HIG), the union of the pairwise-interaction and hyper-interaction graphs, and we call our algorithm the PersonRank. In order to provide publicable datasets for evaluation, we have contributed a new dataset called Multi-scene Important People Image Dataset and gathered a NCAA Basketball Image Dataset from sports game sequences. We have demonstrated that the proposed PersonRank outperforms related methods clearly and substantially. Our code and datasets are available at https://weihonglee.github.io/Projects/PersonRank.htm.
Wei-Hong Li 0001, Benchao Li, Wei-Shi Zheng 0001
FG1
2018 Large-Scale Visible Watermark Detection and Removal with Deep Convolutional Networks
Danni Cheng, Xiang Li 0032, Wei-Hong Li 0001, Chan Lu, Fake Li, Wei-Shi Zheng 0001
PRCV (3)3
2017 Correlation Based Identity Filter: An Efficient Framework for Person Search
Wei-Hong Li 0001, Yafang Mao, Ancong Wu, Wei-Shi Zheng 0001
ICIG (1)1
2016 Sketch metric learning
abstract
The main theme of this paper is to develop a systematic framework to learn a Mahalanobis distance metric based on matrix sketching. Within this framework, we present a novel sketch metric learning algorithm which sequentially sketches the received samples from training dataset and formulates a new kind of constraint for metric learning. This is in contrast to the traditional constraints that are only consisted of data from training dataset. In this paper, one training instance in the constraint is replaced by a pseudo center, which is generated during the sketching stage. Due to this change, our learning algorithm can focus on pushing every received sample to its corresponding similar pseudo center closer and pulling it far away from the dissimilar one. In addition, it can further achieve better performance of some kinds of time-varying process (e.g. on object tracking) than the compared related competitors. We demonstrate how to implement other methods in our algorithm framework and experiment to show that our method outperforms the competitors and relevant baselines on multiple datasets.
Yuting Mai, Wei-Hong Li 0001, Yongyi Tang, Xixi Bi, Wei-Shi Zheng 0001
IJCNN2