EDBT 2026 Demo / reviewers in the wild / expert
Hu Wang 0005
dblp:62/2712-5
· DBLP profile ↗
25ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0003-1725-873XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PeaCap: Patch-Level Retrieval for Lightweight Retrieval-Augmented Image Captioning
Robin Viltoriano, Wei Zhang 0098, Hu Wang 0005, Mong Yuan Sim, Yanjun Shu |
SIGIR | 3 |
| 2026 | DuPLUS: Dual-Prompt Vision-Language Framework for Universal Medical Image Segmentation and PrognosisabstractDeep learning for medical imaging is hampered by task-specific models that lack generalizability and prognostic capabilities, while existing ’universal’ approaches suffer from simplistic conditioning and poor medical semantic understanding. To address these limitations, we introduce DuPLUS, a deep learning framework for efficient multimodal medical image analysis. DuPLUS introduces a novel vision-language framework that leverages hierarchical semantic prompts for fine-grained control over the analysis task, a capability absent in prior universal models. To enable extensibility to other medical tasks, it includes a hierarchical, text-controlled architecture driven by a unique dual-prompt mechanism. For segmentation, DuPLUS is able to generalize across three imaging modalities, ten different anatomically various medical datasets, encompassing more than 30 organs and tumor types. It outperforms the state-of-the-art task-specific and universal models on 8 out of 10 datasets. We demonstrate extensibility of its text-controlled architecture by seamless integration of electronic health record (EHR) data for prognosis prediction, and on a head and neck cancer dataset, DuPLUS achieved a Concordance Index (CI) of 0.69. Parameter-efficient fine-tuning enables rapid adaptation to new tasks and modalities from varying centers, establishing DuPLUS as a versatile and clinically relevant solution for medical image analysis. The code for this work is made available at: Code Numan Saeed, Tausifa Jan Saleem, Fadillah A. Maani, Muhammad Ridzuan, Hu Wang 0005, Mohammad Yaqub |
WACV | 5 |
| 2026 | Unpaired multi-modal multi-label learning for detecting endometriosis signsabstractEndometriosis is a widespread gynecological disorder causing severe pain and infertility, with diagnosis currently relying on slow, costly, and risky laparoscopy. This highlights the critical need for non-invasive imaging diagnostics using transvaginal ultrasound (TVUS) and magnetic resonance imaging (MRI). A key challenge is that patients typically receive only one scan modality in practice, despite TVUS and MRI offering differing diagnostic strengths for endometriosis signs like Pouch of Douglas (POD) obliteration and bowel nodules (BN). Previous work partially addressed this challenge by leveraging unpaired multi-modal data for detecting a single marker: Pouch of Douglas (POD) obliteration. However, this is restrictive because endometriosis signs, such as POD obliteration and bowel nodules (BN), often provide correlated diagnostic cues. Capturing these correlations is essential for accurate detection of endometriosis imaging signs, particularly when combined with multi-modal learning, as each modality offers complementary strengths for different signs. To overcome these limitations, we propose EndoFusion, a novel unpaired multi-modal, multi-label learning framework that enables the detection of POD and BN from TVUS and MRI. Our approach introduces three key innovations: (1) label-based pairing, mixup, and cross-modal feature exchange for robust single-modality inference; (2) Dynamic Mutual Knowledge Distillation (DMKD), which adaptively selects teachers using a worst-student-oriented strategy for effective cross-modal transfer; and (3) label correlations modeling with multi-head attention and a specialized loss to handle imbalance and boost accuracy. This design ensures that knowledge from the superior modality and from co-occurring signs is effectively transferred, mitigating modality-specific weaknesses and improving robustness in imaging sign detection. Experiments on our endometriosis dataset show that our method significantly outperforms all comparison methods, achieving an average AUC of 0.827 (95% CI: 0.790-0.861) when evaluated using single-modality inference. These results represent an initial proof-of-concept toward multi-modal, non-invasive assessment of selected endometriosis imaging signs from MRI and TVUS. Hu Wang 0005, Yutong Xie 0001, Minh-Son To, Steven Knox, Mathew Leonardi, George Condous, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
Artif. Intell. Medicine | 2 |
| 2025 | A Novel Perspective for Multi-Modal Multi-Label Skin Lesion ClassificationabstractThe efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and patient metadata) and addressing the challenges of multi-label classification. Current approaches tend to rely on limited multi-modal techniques and treat the multi-label problem as a multiple multi-class problem, overlooking issues related to imbalanced learning and multi-label correlation. This paper introduces the innovative Skin Lesion Classifier, utilizing a Multi-modal Multilabel TransFormer-based model (SkinM2Former). For multi-modal analysis, we introduce the Tri-Modal Cross-attention Transformer (TMCT) that fuses the three image and metadata modalities at various feature levels of a transformer encoder. For multi-label classification, we introduce a multi-head attention (MHA) module to learn multi-label correlations, complemented by an optimisation that handles multi-label and imbalanced learning problems. SkinM2Former achieves a mean average accuracy of 77.27% and a mean diagnostic accuracy of 77.85% on the public Derm7pt dataset, outperforming state-of-the-art (SOTA) methods. Yutong Xie 0001, Hu Wang 0005, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
WACV | 3 |
| 2024 | Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic SegmentationabstractAugmented Reality (AR) devices, emerging as prominent mobile interaction platforms, face challenges in user safety, particularly concerning oncoming vehicles. While some solutions leverage onboard camera arrays, these cameras often have limited field-of-view (FoV) with front or downward perspectives. Addressing this, we propose a new out-of-view semantic segmentation task and Segment Beyond View (SBV), a novel audio-visual semantic segmentation method. SBV supplements the visual modality, which miss the information beyond FoV, with the auditory information using a teacher-student distillation model (Omni2Ego). The model consists of a vision teacher utilising panoramic information, an auditory teacher with 8-channel audio, and an audio-visual student that takes views with limited FoV and binaural audio as input and produce semantic segmentation for objects outside FoV. SBV outperforms existing models in comparative evaluations and shows a consistent performance across varying FoV ranges and in monaural audio settings. Renjie Wu 0008, Hu Wang 0005, Feras Dayoub, Hsiang-Ting Chen |
AAAI | 2 |
| 2024 | Unraveling Instance Associations: A Closer Look for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate cross-modal alignment between sound and visual objects. Successful audio-visual learning requires two essential components: 1) a challenging dataset with high-quality pixel-level multiclass annotated images associated with audio files, and 2) a model that can establish strong links between audio information and its corresponding visual object. However, these requirements are only partially addressed by current methods, with training sets containing biased audio-visual data, and models that generalise poorly beyond this biased training set. In this work, we propose a new cost-effective strategy to build challenging and relatively unbiased high-quality audio-visual segmentation benchmarks. We also propose a new informative sample mining method for audio-visual supervised contrastive learning to leverage discriminative contrastive samples to enforce cross-modal understanding. We show empirical results that demonstrate the effectiveness of our benchmark. Furthermore, experiments conducted on existing AVS datasets and on our new benchmark show that our method achieves state-of-the-art (SOTA) segmentation accuracy11This work was supported by Australian Research Council through grant FT190100525. Yuanhong Chen, Yuyuan Liu, Hu Wang 0005, Fengbei Liu, Chong Wang 0012, Helen Frazer, Gustavo Carneiro 0001 |
CVPR | 3 |
| 2024 | CPM: Class-Conditional Prompting Machine for Audio-Visual Segmentation
Yuanhong Chen, Chong Wang 0012, Yuyuan Liu, Hu Wang 0005, Gustavo Carneiro 0001 |
ECCV (10) | 4 |
| 2024 | ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang 0005, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ECCV (1) | 3 |
| 2024 | Disentangling Specificity for Abstractive Multi-document Summarization
Congbo Ma, Wei Zhang 0098, Hu Wang 0005, Haojie Zhuang, Mingyu Guo 0001 |
IJCNN | 3 |
| 2024 | CORELOCKER: Neuron-level Usage ControlabstractThe growing complexity of deep neural network models in modern application domains necessitates a complex training process that involves extensive data, sophisticated design, and substantial computation. The trained model inherently encapsulates the intellectual property owned by the model developer (or the model owner). Consequently, safeguarding the model from unauthorized use by entities who obtain access to the model (or the model controllers), i.e., preserving the fundamental rights and proprietary interests of the model owner, has become a critical necessity.In this work, we propose CORELOCKER, employing the strategic extraction of a small subset of significant weights from the neural network. This subset serves as the access key to unlock the model’s complete capability. The extraction of the key can be customized to varying levels of utility that the model owner intends to release. Authorized users with the access key have full access to the model, while unauthorized users can have access to only part of its capability. We establish a formal foundation to underpin CORELOCKER, which provides crucial lower and upper bounds for the utility disparity between pre- and post-protected networks. We evaluate CORELOCKER using representative datasets such as Fashion-MNIST, CIFAR-10, and CIFAR-100, as well as real-world models including Vg-gNet, ResNet, and DenseNet. Our experimental results confirm its efficacy. We also demonstrate CORELOCKER’s resilience against advanced model restoration attacks based on fine-tuning and pruning. Zhongkui Ma, Xinguo Feng, Ruoxi Sun 0001, Hu Wang 0005, Minhui Xue 0001, Guangdong Bai |
SP | 5 |
| 2024 | ReFs: A hybrid pre-training paradigm for 3D medical image segmentation
Yutong Xie 0001, Lingqiao Liu, Hu Wang 0005, Yiwen Ye, Johan Verjans, Yong Xia 0001 |
Medical Image Anal. | 4 |
| 2024 | Counting Crowd by Weighing Counts: A Sequential Decision-Making PerspectiveabstractWe show that crowd counting can be formulated as a sequential decision-making (SDM) problem. Inspired by human counting, we evade one-step estimation mostly executed in existing counting models and decompose counting into sequential sub-decision problems. During implementation, a key insight is to interpret sequential counting as a physical process in reality-scale weighing. This analogy allows us to implement a novel "counting scale" termed LibraNet. Our idea is that, by placing a crowd image on the scale, LibraNet (agent) learns to place appropriate weights to match the count: at each step, one weight (action) is chosen from the weight box (the predefined action pool) conditioned on the image features and the placed weights (state) until the pointer (the agent output) informs balance. We investigate two forms of state definition and explore four types of LibraNet implementations under different learning paradigms, including deep Q-network (DQN), actor-critic (AC), imitation learning (IL), and mixed AC+IL. Experiments show that LibraNet indeed mimics scale weighing, that it outperforms or performs comparably against state-of-the-art approaches on five crowd counting benchmarks, that it can be used as a plug-in to improve off-the-shelf counting models, and particularly that it demonstrates remarkable cross-dataset generalization. Code and models are available at https://git.io/libranet. Hao Lu 0003, Liang Liu 0001, Hu Wang 0005, Zhiguo Cao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-Modal Learning with Missing Modality via Shared-Specific Feature ModellingabstractThe missing modality issue is critical but non-trivial to be solved by multi-modal models. Current methods aiming to handle the missing modality problem in multi-modal tasks, either deal with missing modalities only during evaluation or train separate models to handle specific missing modality settings. In addition, these models are designed for specific tasks, so for example, classification models are not easily adapted to segmentation tasks and vice versa. In this paper, we propose the Shared-Specific Feature Modelling (ShaSpec) method that is considerably simpler and more effective than competing approaches that address the issues above. ShaSpec is designed to take advantage of all available input modalities during training and evaluation by learning shared and specific features to better represent the input data. This is achieved from a strategy that relies on auxiliary tasks based on distribution alignment and domain classification, in addition to a residual feature fusion procedure. Also, the design simplicity of ShaSpec enables its easy adaptation to multiple tasks, such as classification and segmentation. Experiments are conducted on both medical image segmentation and computer vision classification, with results indicating that ShaSpec outperforms competing methods by a large margin. For instance, on BraTS2018, ShaSpec improves the SOTA by more than 3% for enhancing tumour, 5% for tumour core and 3% for whole tumour.11This work received funding from the Australian Government the through Medical Research Futures Fund: Primary Health Care Research Data Infrastructure Grant 2020 and from Endometriosis Australia. G.C. was supported by Australian Research Council through grant FT190100525. Hu Wang 0005, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
CVPR | 1 |
| 2023 | BoMD: Bag of Multi-label Descriptors for Noisy Chest X-ray ClassificationabstractDeep learning methods have shown outstanding classification accuracy in medical imaging problems, which is largely attributed to the availability of large-scale datasets manually annotated with clean labels. However, given the high cost of such manual annotation, new medical imaging classification problems may need to rely on machine-generated noisy labels extracted from radiology reports. Indeed, many Chest X-Ray (CXR) classifiers have been modelled from datasets with noisy labels, but their training procedure is in general not robust to noisy-label samples, leading to sub-optimal models. Furthermore, CXR datasets are mostly multi-label, so current multi-class noisy-label learning methods cannot be easily adapted. In this paper, we propose a new method designed for noisy multi-label CXR learning, which detects and smoothly re-labels noisy samples from the dataset to be used in the training of common multi-label classifiers. The proposed method optimises a bag of multi-label descriptors (BoMD) to promote their similarity with the semantic descriptors produced by language models from multi-label image annotations. Our experiments on noisy multi-label training sets and clean testing sets show that our model has state-of-the-art accuracy and robustness in many CXR multi-label classification benchmarks, including a new benchmark that we propose to systematically assess noisy multi-label methods. Code is available at https://github.com/cyh-0/BoMD. Yuanhong Chen, Fengbei Liu, Hu Wang 0005, Chong Wang 0012, Yuyuan Liu, Yu Tian 0001, Gustavo Carneiro 0001 |
ICCV | 3 |
| 2023 | Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
Hu Wang 0005, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
MICCAI (4) | 1 |
| 2023 | Data Hiding With Deep Learning: A Survey Unifying Digital Watermarking and SteganographyabstractThe advancement of secure communication and identity verification fields has significantly increased through the use of deep learning techniques for data hiding. By embedding information into a noise-tolerant signal, such as audio, video, or images, digital watermarking and steganography techniques can be used to protect sensitive intellectual property (IP) and enable confidential communication, ensuring that the information embedded is only accessible to authorized parties. This survey provides an overview of recent developments in deep learning techniques deployed for data hiding, categorized systematically according to model architectures and noise injection methods. In addition, potential future research directions that unite digital watermarking and steganography on software engineering to enhance security and mitigate risks are suggested and deliberated. This contribution furthers the creation of a more trustworthy digital world and advances responsible artificial intelligence (AI). Olivia Byrnes, Hu Wang 0005, Ruoxi Sun 0001, Congbo Ma, Huaming Chen, Qi Wu 0001, Minhui Xue 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Uncertainty-Aware Multi-modal Learning via Cross-Modal Random Network Prediction
Hu Wang 0005, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
ECCV (37) | 1 |
| 2022 | Multi-view Local Co-occurrence and Global Consistency Learning Improve Mammogram Classification Generalisation
Yuanhong Chen, Hu Wang 0005, Chong Wang 0012, Yu Tian 0001, Fengbei Liu, Yuyuan Liu, Michael Elliott, Davis J. McCarthy, Helen Frazer, Gustavo Carneiro 0001 |
MICCAI (3) | 2 |
| 2022 | M$^4$I: Multi-modal Models Membership InferenceabstractWith the development of machine learning techniques, the attention of research has been moved from single-modal learning to multi-modal learning, as real-world data exist in the form of different modalities. However, multi-modal models often carry more information than single-modal models and they are usually applied in sensitive scenarios, such as medical report generation or disease identification. Compared with the existing membership inference against machine learning classifiers, we focus on the problem that the input and output of the multi-modal models are in different modalities, such as image captioning. This work studies the privacy leakage of multi-modal models through the lens of membership inference attack, a process of determining whether a data record involves in the model training process or not. To achieve this, we propose Multi-modal Models Membership Inference (M$^4$I) with two attack methods to infer the membership status, named metric-based (MB) M$^4$I and feature-based (FB) M$^4$I, respectively. More specifically, MB M$^4$I adopts similarity metrics while attacking to infer target data membership. FB M$^4$I uses a pre-trained shadow multi-modal feature extractor to achieve the purpose of data inference attack by comparing the similarities from extracted input and output features. Extensive experimental results show that both attack methods can achieve strong performances. Respectively, 72.5% and 94.83% of attack success rates on average can be obtained under unrestricted scenarios. Moreover, we evaluate multiple defense mechanisms against our attacks. The source code of M$^4$I attacks is publicly available at https://github.com/MultimodalMI/Multimodal-membership-inference.git. Pingyi Hu, Ruoxi Sun 0001, Hu Wang 0005, Minhui Xue 0001 |
NeurIPS | 4 |
| 2022 | Incorporating Linguistic Knowledge for Abstractive Multi-document Summarization
Congbo Ma, Wei Zhang 0098, Hu Wang 0005, Mingyu Guo 0001 |
PACLIC | 3 |
| 2021 | Oriole: Thwarting Privacy Against Trustworthy Deep Learning Models
Liuqiao Chen, Hu Wang 0005, Benjamin Zi Hao Zhao, Minhui Xue 0001, Haifeng Qian |
ACISP | 2 |
| 2021 | Fully Quantized Image Super-Resolution NetworksabstractWith the rising popularity of intelligent mobile devices, it is of great practical significance to develop accurate, real-time and energy-efficient image Super-Resolution (SR) methods. A prevailing method for improving inference efficiency is model quantization, which allows for replacing the expensive floating-point operations with efficient bitwise arithmetic. To date, it is still challenging for quantized SR frameworks to deliver a feasible accuracy-efficiency trade-off. Here, we propose a Fully Quantized image Super-Resolution framework (FQSR) to jointly optimize efficiency and accuracy. In particular, we target obtaining end-to-end quantized models for all layers, especially including skip connections, which was rarely addressed in the literature of SR quantization. We further identify obstacles faced by low-bit SR networks and propose a novel method to counteract them accordingly. The difficulties are caused by 1) for SR task, due to the existence of skip connections, high-resolution feature maps would occupy a huge amount of memory spaces; 2) activation and weight distributions being vastly distinctive in different layers; 3) the inaccurate approximation of the quantization. We apply our quantization scheme on multiple mainstream super-resolution architectures, including SRResNet, SRGAN and EDSR. Experimental results show that our FQSR with low-bits quantization is able to achieve on par performance compared with the full-precision counterparts on five benchmark datasets and surpass the state-of-the-art quantized SR methods with significantly reduced computational cost and memory consumption. Code is available at https://git.io/JWxPp. Hu Wang 0005, Peng Chen 0037, Bohan Zhuang, Chunhua Shen |
ACM Multimedia | 1 |
| 2020 | Soft Expert Reward Learning for Vision-and-Language Navigation
Hu Wang 0005, Qi Wu 0001, Chunhua Shen |
ECCV (9) | 1 |
| 2020 | Unsupervised Representation Learning by Predicting Random DistancesabstractDeep neural networks have gained great success in a broad range of tasks due to its remarkable capability to learn semantically rich features from high-dimensional data. However, they often require large-scale labelled data to successfully learn such features, which significantly hinders their adaption in unsupervised learning tasks, such as anomaly detection and clustering, and limits their applications to critical domains where obtaining massive labelled data is prohibitively expensive. To enable unsupervised learning on those domains, in this work we propose to learn features without using any labelled data by training neural networks to predict data distances in a randomly projected space. Random mapping is a theoretically proven approach to obtain approximately preserved distances. To well predict these distances, the representation learner is optimised to learn genuine class structures that are implicitly embedded in the randomly projected space. Empirical results on 19 real-world datasets show that our learned representations substantially outperform a few state-of-the-art methods for both anomaly detection and clustering tasks. Code is available at: \url{https://git.io/RDP} Hu Wang 0005, Guansong Pang, Chunhua Shen, Congbo Ma |
IJCAI | 1 |
| 2019 | Multi-label Thoracic Disease Image Classification with Cross-Attention Networks
Congbo Ma, Hu Wang 0005, Steven C. H. Hoi |
MICCAI (6) | 2 |