Conghui Hu

dblp:218/6215 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Rethink Cross-Modal Fusion in Weakly-Supervised Audio-Visual Video Parsing
abstract
Existing works on weakly-supervised audio-visual video parsing adopt hybrid attention network (HAN) as the multi-modal embedding to capture the cross-modal context. It embeds the audio and visual modalities with a shared network, where the cross-attention is performed at the input. However, such an early fusion method highly entangles the two non-fully correlated modalities and leads to sub-optimal performance in detecting single-modality events. To deal with this problem, we propose the messenger-guided mid-fusion transformer to reduce the uncorrelated cross-modal context in the fusion. The messengers condense the full cross-modal context into a compact representation to only preserve useful cross-modal information. Furthermore, due to the fact that microphones capture audio events from all directions, while cameras only record visual events within a restricted field of view, there is a more frequent occurrence of unaligned cross-modal context from audio for visual event predictions. We thus propose cross-audio prediction consistency to suppress the impact of irrelevant audio information on visual event prediction. Experiments consistently illustrate the superior performance of our framework compared to existing state-of-the-art methods.
Yating Xu, Conghui Hu, Gim Hee Lee
WACV2
2023 Sketch-based Video Object Segmentation: Benchmark and Analysis
Ruolin Yang 0001, Da Li 0001, Conghui Hu, Timothy M. Hospedales, Honggang Zhang 0002, Yi-Zhe Song
BMVC3
2023 Motion and Context-Aware Audio-Visual Conditioned Video Prediction
Yating Xu, Conghui Hu, Gim Hee Lee
BMVC2
2023 Unsupervised Feature Representation Learning for Domain-generalized Cross-domain Image Retrieval
abstract
Cross-domain image retrieval has been extensively studied due to its high practical value. In recently proposed unsupervised cross-domain image retrieval methods, efforts are taken to break the data annotation barrier. However, applicability of the model is still confined to domains seen during training. This limitation motivates us to present the first attempt at domain-generalized unsupervised crossdomain image retrieval (DG-UCDIR) aiming at facilitating image retrieval between any two unseen domains in an unsupervised way. To improve domain generalizability of the model, we thus propose a new two-stage domain augmentation technique for diversified training data generation. DG-UCDIR also shares all the challenges present in the unsupervised cross-domain image retrieval, where domain-agnostic and semantic-aware feature representations are supposed to be learned without external supervision. To accomplish this, we introduce a novel crossdomain contrastive learning strategy by utilizing phase image as a proxy to mitigate the domain gap. Extensive experiments are carried out using PACS and DomainNet dataset, and consistently illustrate the superior performance of our framework compared to existing state-ofthe-art methods. Our source code is available at https://github.com/conghui1002/DG-UCDIR.
Conghui Hu, Can Zhang 0007, Gim Hee Lee
ICCV1
2023 Generalized Few-Shot Point Cloud Segmentation Via Geometric Words
abstract
Existing fully-supervised point cloud segmentation methods suffer in the dynamic testing environment with emerging new classes. Few-shot point cloud segmentation algorithms address this problem by learning to adapt to new classes at the sacrifice of segmentation accuracy for the base classes, which severely impedes its practicality. This largely motivates us to present the first attempt at a more practical paradigm of generalized few-shot point cloud segmentation, which requires the model to generalize to new categories with only a few support point clouds and simultaneously retain the capability to segment base classes. We propose the geometric words to represent geometric components shared between the base and novel classes, and incorporate them into a novel geometric-aware semantic representation to facilitate better generalization to the new classes without forgetting the old ones. Moreover, we introduce geometric prototypes to guide the segmentation with geometric prior knowledge. Extensive experiments on S3DIS and ScanNet consistently illustrate the superior performance of our method over baseline methods. Our code is available at: https://github.com/Pixie8888/GFS-3DSeg_GWs.
Yating Xu, Conghui Hu, Na Zhao 0004, Gim Hee Lee
ICCV2
2022 Towards Unsupervised Sketch-based Image Retrieval
Conghui Hu, Yongxin Yang, Yunpeng Li 0004, Timothy M. Hospedales, Yi-Zhe Song
BMVC1
2022 Feature Representation Learning for Unsupervised Cross-Domain Image Retrieval
Conghui Hu, Gim Hee Lee
ECCV (37)1
2021 End-to-End Semi-supervised Learning for Differentiable Particle Filters
abstract
Recent advances in incorporating neural networks into particle filters provide the desired flexibility to apply particle filters in large-scale real-world applications. The dynamic and measurement models in this framework are learnable through the differentiable implementation of particle filters. Past efforts in optimising such models often require the knowledge of true states which can be expensive to obtain or even unavailable in practice. In this paper, in order to reduce the demand for annotated data, we present an end-to-end learning objective based upon the maximisation of a pseudo-likelihood function which can improve the estimation of states when large portion of true states are unknown. We assess performance of the proposed method in state estimation tasks in robotics with simulated and real-world datasets.
Xiongjie Chen, Georgios Papagiannis, Conghui Hu, Yunpeng Li 0001
ICRA4
2020 Sketch-a-Segmenter: Sketch-Based Photo Segmenter Generation
abstract
Given pixel-level annotated data, traditional photo segmentation techniques have achieved promising results. However, these photo segmentation models can only identify objects in categories for which data annotation and training have been carried out. This limitation has inspired recent work on few-shot and zero-shot learning for image segmentation. In this paper, we show the value of sketch for photo segmentation, in particular as a transferable representation to describe a concept to be segmented. We show, for the first time, that it is possible to generate a photo-segmentation model of a novel category using just a single sketch and furthermore exploit the unique fine-grained characteristics of sketch to produce more detailed segmentation. More specifically, we propose a sketch-based photo segmentation method that takes sketch as input and synthesizes the weights required for a neural network to segment the corresponding region of a given photo. Our framework can be applied at both the category-level and the instance-level, and fine-grained input sketches provide more accurate segmentation in the latter. This framework generalizes across categories via sketch and thus provides an alternative to zero-shot learning when segmenting a photo from a category without annotated training data. To investigate the instance-level relationship across sketch and photo, we create the SketchySeg dataset which contains segmentation annotations for photos corresponding to paired sketches in the Sketchy Dataset.
Conghui Hu, Da Li 0001, Yongxin Yang, Timothy M. Hospedales, Yi-Zhe Song
IEEE Trans. Image Process.1
2018 Sketch-a-Classifier: Sketch-Based Photo Classifier Generation
abstract
Contemporary deep learning techniques have made image recognition a reasonably reliable technology. However training effective photo classifiers typically takes numerous examples which limits image recognition's scalability and applicability to scenarios where images may not be available. This has motivated investigation into zero-shot learning, which addresses the issue via knowledge transfer from other modalities such as text. In this paper we investigate an alternative approach of synthesizing image classifiers: almost directly from a user's imagination, via freehand sketch. This approach doesn't require the category to be nameable or describable via attributes as per zero-shot learning. We achieve this via training a model regression network to map from free-hand sketch space to the space of photo classifiers. It turns out that this mapping can be learned in a category-agnostic way, allowing photo classifiers for new categories to be synthesized by user with no need for annotated training photos. We also demonstrate that this modality of classifier generation can also be used to enhance the granularity of an existing photo classifier, or as a complement to name-based zero-shot learning.
Conghui Hu, Da Li 0001, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales
CVPR1
2017 Now You See Me: Deep Face Hallucination for Unviewed Sketches
Conghui Hu, Da Li 0001, Yi-Zhe Song, Timothy M. Hospedales
BMVC1