EDBT 2026 Demo / reviewers in the wild / expert
Dadong Wang
dblp:06/2186
· DBLP profile ↗
43ranked-venue papers
3as first author
29since 2021 · last 2025
0000-0003-0409-2259ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentabstractAccurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous audio-visual correspondences—such as nearby visually similar but acoustically different objects and frequent shifts in objects’ sounding status. Consequently, they may struggle to reliably correlate audio and visual cues, leading to over-or under-segmentation. To address these limitations, we propose a novel framework with two primary components: an audio-guided modality alignment (AMA) module and an uncertainty estimation (UE) module. Instead of indiscriminately correlating audio-visual cues through a global attention mechanism, AMA performs audio-visual interactions within multiple groups and consolidates group features into compact representations based on their responsiveness to audio cues, effectively directing the model’s attention to audio-relevant areas. Leveraging contrastive learning, AMA further distinguishes sounding regions from silent areas by treating features with strong audio responses as positive samples and weaker responses as negatives. Additionally, UE integrates spatial and temporal information to identify high-uncertainty regions caused by frequent changes in sound state, reducing prediction errors by lowering confidence in these areas. Experimental results demonstrate that our approach achieves superior accuracy compared to existing state-of-the-art methods, particularly in challenging scenarios where traditional approaches struggle to maintain reliable segmentation. Chen Liu 0018, Peike Li, Dadong Wang, Lincheng Li, Xin Yu 0002 |
CVPR | 4 |
| 2025 | Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio SemanticsabstractSound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions, without fully addressing the inherent challenges posed by audio natures, i.e., (1) feature confusion due to the overlapping nature of audio signals, and (2) audiovisual matching difficulty from the varied sounds produced by the same object. To address these challenges, we propose Dynamic Derivation and Elimination (DDESeg): a novel audio-visual segmentation framework. Specifically, to mitigate feature confusion, DDESeg reconstructs the semantic content of the mixed audio signal by enriching the distinct semantic information of each individual source, deriving representations that preserve the unique characteristics of each sound. To reduce the matching difficulty, we introduce a discriminative feature learning module, which enhances the semantic distinctiveness of generated audio representations. Considering that not all derived audio representations directly correspond to visual features (e.g., off-screen sounds), we propose a dynamic elimination module to filter out non-matching elements. This module facilitates targeted interaction between sounding regions and relevant audio semantics. By scoring the interacted features, we identify and filter out irrelevant audio information, ensuring accurate audio-visual alignment. Comprehensive experiments demonstrate that our framework achieves superior performance in AVS datasets. Our code is here. Peike Li, Dadong Wang, Lincheng Li, Xin Yu 0002 |
CVPR | 4 |
| 2025 | Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion ModelabstractBitstream-corrupted video recovery aims to fill in realistic video content due to bitstream corruption during video storage or transmission. Most existing methods typically assume that the predefined masks of the corrupted regions are known in advance. However, manually annotating these masks is laborious and time-consuming, limiting the applicability of existing methods in real-world scenarios. Therefore, we expect to relax this assumption by defining a new blind video recovery setting where the recovery of corrupted regions does not rely on predefined masks. There are two significant challenges in this setting: (i) without predefined masks, how accurately can a model identify the regions requiring recovery? (ii) how to recover contents from extensive and irregular regions, especially when large portions of frames are severely degraded? To address these challenges, we introduce a Metadata-Guided Diffusion Model, dubbed M-GDM. To enable a diffusion model focusing on the corrupted regions, we leverage intrinsic video metadata as a corruption indicator and design a dual-stream metadata encoder. This encoder first embeds the motion vectors and frame types of a video separately and then merges them into a unified metadata representation. The metadata representation will interact with the corrupted latent feature through cross-attention mechanisms at each diffusion step. Meanwhile, to preserve the intact regions, we propose a prior-driven mask predictor that generates pseudo masks for the corrupted regions by leveraging the metadata prior and diffusion prior. These pseudo masks enable the separation and recombination of intact and recovered regions through hard masking. However, imperfections in pseudo mask predictions and hard masking processes often result in boundary artifacts. Thus, we introduce a post-refinement module that refines the hard-masked outputs, enhancing the consistency between intact and recovered regions. Extensive experiment results validate the effectiveness of our method and demonstrate its superiority in the blind video recovery task. Hu Zhang 0005, Dadong Wang, Xin Yu 0002 |
CVPR | 4 |
| 2025 | Jailbreaking the Non-Transferable Barrier via Test-Time Data DisguisingabstractNon-Transferable learning (NTL) has been proposed to protect model intellectual property (IP) by creating a "nontransferable barrier" to restrict generalization from authorized to unauthorized domains. Recently, well-designed attack, which restores the unauthorized-domain performance by fine-tuning NTL models on few authorized samples, highlights the security risks of NTL-based applications. However, such attack requires modifying model weights, thus being invalid in the black-box scenario. This raises a critical question: can we trust the security of NTL models deployed as black-box systems? In this work, we reveal the first loophole of black-box NTL models by proposing a novel attack method (dubbed as JailNTL) to jailbreak the non-transferable barrier through test-time data disguising. The main idea of JailNTL is to disguise unauthorized data so it can be identified as authorized by the NTL model, thereby bypassing the non-transferable barrier without modifying the NTL model weights. Specifically, JailNTL encourages unauthorized-domain disguising in two levels, including: (i) data-intrinsic disguising (DID) for eliminating domain discrepancy and preserving class-related content at the input-level, and (ii) model-guided disguising (MGD) for mitigating output-level statistics difference of the NTL model. Empirically, when attacking state-of-the-art (SOTA) NTL models in the black-box scenario, Jail-Ntl achieves an accuracy increase of up to 55.7% in the unauthorized domain by using only 1% authorized samples, largely exceeding existing SOTA white-box attacks. Code is released at https://github.com/tmllab/2025_CVPR_JailNTL. Yongli Xiang, Ziming Hong, Lina Yao 0001, Dadong Wang, Tongliang Liu |
CVPR | 4 |
| 2025 | Chain-of-Focus Prompting: Leveraging Sequential Visual Cues to Prompt Large Autoregressive Vision ModelsabstractIn-context learning (ICL) has revolutionized natural language processing by enabling models to adapt to diverse tasks with only a few illustrative examples. However, the exploration of ICL within the field of computer vision remains limited. Inspired by Chain-of-Thought (CoT) prompting in the language domain, we propose Chain-of-Focus (CoF) Prompting, which enhances vision models by enabling step-by-step visual comprehension. CoF Prompting addresses the challenges of absent logical structure in visual data by generating intermediate reasoning steps through visual saliency. Moreover, it provides a solution for creating tailored prompts from visual inputs by selecting contextually informative prompts based on query similarity and target richness. The significance of CoF prompting is demonstrated by the recent introduction of Large Autoregressive Vision Models (LAVMs), which predict downstream targets via in-context learning with pure visual inputs. By integrating intermediate reasoning steps into visual prompts and effectively selecting the informative ones, the LAVMs are capable of generating significantly better inferences. Extensive experiments on downstream visual understanding tasks validate the effectiveness of our proposed method for visual in-context learning. Jiyang Zheng, Jialiang Shen, Yu Yao 0005, Yang Yang 0002, Dadong Wang, Tongliang Liu |
ICLR | 6 |
| 2025 | UniMRG: Refining Medical Semantic Understanding Across Modalities via LLM-Orchestrated Synergistic Evolution
Hongyan Xu 0002, Arcot Sowmya, Ian Katz, Dadong Wang |
MICCAI (5) | 4 |
| 2025 | Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task PlanningabstractTo enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This requires a smart map that fuses accurate geometric structure with rich, human-understandable semantics. To address this, we introduce the 3D Queryable Scene Representation (3D QSR), a novel framework built on multimedia data that unifies three complementary 3D representations: (1) 3D-consistent novel view rendering and segmentation from panoptic reconstruction, (2) precise geometry from 3D point clouds, and (3) structured, scalable organization via 3D scene graphs. Built on an object-centric design, the framework integrates with large vision-language models to enable semantic queryability by linking multimodal object embeddings, and supporting object-level retrieval of geometric, visual, and semantic information. The retrieved data are then loaded into a robotic task planner for downstream execution. Xun Li 0004, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang 0005, Madhawa Perera, Ziwei Wang 0003, Ahalya Ravendran, Brandon J. Matthews, Matt Adcock, Dadong Wang, Jiajun Liu 0004 |
ACM Multimedia | 11 |
| 2025 | Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationabstractText-to-Audio-Video (T2AV) generation aims to produce temporally and semantically aligned visual and auditory content from natural language descriptions. While recent progress in text-to-audio and text-to-video models has improved generation quality within each modality, jointly modeling them remains challenging due to incomplete and asymmetric correspondence: audio often reflects only a subset of the visual scene, and vice versa. Naively enforcing full alignment introduces semantic noise and temporal mismatches. To address this, we propose a novel framework that performs selective cross-modal alignment through a learnable masking mechanism, enabling the model to isolate and align only the shared latent components relevant to both modalities. This mechanism is integrated into an adaptation module that interfaces with pretrained encoders and decoders from latent video and audio diffusion models, preserving their generative capacity with reduced training overhead. Theoretically, we show that our masked objective provably recovers the minimal set of shared latent variables across modalities. Empirically, our method achieves state-of-the-art performance on standard T2AV benchmarks, demonstrating significant improvements in audiovisual synchronization and semantic consistency. Jiyang Zheng, Siqi Pan, Yu Yao 0005, Zhaoqing Wang, Dadong Wang, Tongliang Liu |
NeurIPS | 5 |
| 2025 | Facial Expression Recognition with Controlled Privacy Preservation and Feature CompensationabstractFacial expression recognition (FER) systems raise significant privacy concerns due to the potential exposure of sensitive identity information. This paper presents a study on removing identity information while preserving FER capabilities. Drawing on the observation that lowfrequency components predominantly contain identity information and high-frequency components capture expression, we propose a novel two-stream framework that applies privacy enhancement to each component separately. We introduce a controlled privacy enhancement mechanism to optimize performance and a feature compensator to enhance task-relevant features without compromising privacy. Furthermore, we propose a novel privacy-utility trade-off, providing a quantifiable measure of privacy preservation efficacy in closed-set FER tasks. Extensive experiments on the benchmark CREMA-D dataset demonstrate that our framework achieves 78.84 % recognition accuracy with a privacy (facial identity) leakage ratio of only 2.01 %, highlighting its potential for secure and reliable video-based FER applications. We encourage the readers to visit the project page: https://fengxxu.github.io/ppfer/. David Ahmedt-Aristizabal, Lars Petersson, Dadong Wang, Xun Li 0004 |
WACV | 4 |
| 2025 | ESceme: Vision-and-Language Navigation with Episodic Scene MemoryabstractAbstract Vision-and-language navigation (VLN) simulates a visual agent that follows natural-language navigation instructions in real-world scenes. Existing approaches have made enormous progress in navigation in new environments, such as beam search, pre-exploration, and dynamic or hierarchical history encoding. To balance generalization and efficiency, we resort to memorizing visited scenarios apart from the ongoing route while navigating. In this work, we introduce a mechanism of Episodic Scene memory (ESceme) for VLN that wakes an agent’s memories of past visits when it enters the current scene. The episodic scene memory allows the agent to envision a bigger picture of the next prediction. This way, the agent learns to utilize dynamically updated information instead of merely adapting to the current observations. We provide a simple yet effective implementation of ESceme by enhancing the accessible views at each location and progressively completing the memory while navigating. We verify the superiority of ESceme on short-horizon (R2R), long-horizon (R4R), and vision-and-dialog (CVDN) VLN tasks. Our ESceme also wins first place on the CVDN leaderboard. Code is available: https://github.com/qizhust/esceme . Qi Zheng 0003, Daqing Liu, Jing Zhang 0037, Dadong Wang, Dacheng Tao |
Int. J. Comput. Vis. | 5 |
| 2024 | Benchmarking Audio Visual Segmentation for Long-Untrimmed VideosabstractExisting audio-visual segmentation datasets typically focus on short-trimmed videos with only one pixel-map annotation for a per-second video clip. In contrast, for untrimmed videos, the sound duration, start- and end-sounding time positions, and visual deformation of audible objects vary significantly. Therefore, we observed that current AVS models trained on trimmed videos might struggle to segment sounding objects in long videos. To investigate the feasibility of grounding audible objects in videos along both temporal and spatial dimensions, we introduce the Long-Untrimmed Audio-Visual Segmentation dataset (LU-AVS), which includes precise frame-level annotations of sounding emission times and provides exhaustive mask annotations for all frames. Considering that pixel-level annotations are difficult to achieve in some complex scenes, we also provide the bounding boxes to indicate the sounding regions. Specifically, LU-AVS contains 10M mask annotations across 6.6K videos, and 11M bounding box annotations across 7K videos. Compared with the existing datasets, LU-AVS videos are on average 4~8 times longer, with the silent duration being 3~15 times greater. Furthermore, we try our best to adapt some baseline models that were originally designed for audio-visual-relevant tasks to examine the challenges of our newly curated LU-AVS. Through comprehensive evaluation, we demonstrate the challenges of LU-AVS compared to the ones containing trimmed videos. Therefore, LU-AVS provides an ideal yet challenging platform for evaluating audio-visual segmentation and localization on untrimmed long videos. The dataset is publicly available at: https://yenanliu.github.io/LU-AVS/. Chen Liu 0028, Peike Li, Qingtao Yu, Hongwei Sheng, Dadong Wang, Lincheng Li, Xin Yu 0002 |
CVPR | 5 |
| 2024 | Enhancing Contrastive Learning for Ordinal Regression via Ordinal Content Preserved Data AugmentationabstractContrastive learning, while highly effective for a lot of tasks, shows limited improvement in ordinal regression. We find that the limitation comes from the predefined strong data augmentations employed in contrastive learning. Intuitively, for ordinal regression datasets, the discriminative information (ordinal content information) contained in instances is subtle. The strong augmentations can easily overshadow or diminish this ordinal content information. As a result, when contrastive learning is used to extract common features between weakly and strongly augmented images, the derived features often lack this essential ordinal content, rendering them less useful in training models for ordinal regression. To improve contrastive learning's utility for ordinal regression, we propose a novel augmentation method to replace the predefined strong argumentation based on the principle of minimal change. Our method is designed in a generative manner that can effectively generate images with different styles but contains desired ordinal content information. Extensive experiments validate the effectiveness of our proposed method, which serves as a plug-and-play solution and consistently improves the performance of existing state-of-the-art methods in ordinal regression tasks. Jiyang Zheng, Yu Yao 0005, Bo Han 0003, Dadong Wang, Tongliang Liu |
ICLR | 4 |
| 2024 | SCD-NAS: Towards Zero-Cost Training in Melanoma DiagnosisabstractDiagnosing melanoma remains challenging despite advances in Convolutional Neural Networks (CNNs) for skin cancer detection. Their application in clinical settings is often limited by differences between natural and clinical images. To address this, we introduce the Skin Cancer Detection Neural Architecture Search (SCD-NAS) framework. In our method, Large Language Model (LLM) is leveraged as a proxy, which helps SCD-NAS achieve cost-free training. Additionally, to maximize the benefits of various architectural design spaces, we introduce a Search Space Expansion (SSE) methodology. This effectively combines the merits of diverse architectural configurations, thereby enhancing model performance. We conducted experiments on the ISIC 2020, MedMNISTv2, CIFAR-10 and CIFAR-100 datasets. Our SCD-NAS-derived ResNet50 model achieved an Area Under the Curve (AUC) of 91.23% on the ISIC 2020 dataset, improving the baseline by 5.93%. It also exceeded the CIFAR-10 benchmark by 2.45% in accuracy. Hongyan Xu 0002, Xiu Su, Arcot Sowmya, Ian Katz, Dadong Wang |
ICME | 5 |
| 2024 | AMFP-net: Adaptive multi-scale feature pyramid network for diagnosis of pneumoconiosis from chest X-ray imagesabstractEarly detection of pneumoconiosis by routine health screening of workers in the mining industry is critical for preventing the progression of this incurable disease. Automated pneumoconiosis classification in chest X-ray images is challenging due to the low contrast of opacities, inter-class similarity, intra-class variation and the existence of artifacts. Compared to traditional methods, convolutional neural networks have shown significant improvement in pneumoconiosis classification tasks, however, accurate classification remains challenging due to mainly the inability to focus on semantically meaningful lesion opacities. Most existing networks focus on high level abstract information and ignore low level detailed object information. Different from natural images where an object occupies large space, the classification of pneumoconiosis depends on the density of small opacities inside the lung. To address this issue, we propose a novel two-stage adaptive multi-scale feature pyramid network called AMFP-Net for the diagnosis of pneumoconiosis from chest X-rays. The proposed model consists of 1) an adaptive multi-scale context block to extract rich contextual and discriminative information and 2) a weighted feature fusion module to effectively combine low level detailed and high level global semantic information. This two-stage network first segments the lungs to focus more on relevant regions by excluding irrelevant parts of the image, and then utilises the segmented lungs to classify pneumoconiosis into different categories. Extensive experiments on public and private datasets demonstrate that the proposed approach can outperform state-of-the-art methods for both segmentation and classification. Md. Shariful Alam, Dadong Wang, Arcot Sowmya |
Artif. Intell. Medicine | 2 |
| 2024 | Bypass network for semantics driven image paragraph captioningabstractImage paragraph captioning aims to describe a given image with a sequence of coherent sentences. Most existing methods model the coherence through the topic transition that dynamically infers a topic vector from preceding sentences. However, these methods still suffer from immediate or delayed repetitions in generated paragraphs because (i) the entanglement of syntax and semantics distracts the topic vector from attending pertinent visual regions; (ii) there are few constraints or rewards for learning long-range transitions. In this paper, we propose a bypass network that separately models semantics and linguistic syntax of preceding sentences. Specifically, the proposed model consists of two main modules, i.e. a topic transition module and a sentence generation module. The former takes previous semantic vectors as queries and applies attention mechanism on regional features to acquire the next topic vector, which reduces immediate repetition by eliminating linguistics. The latter decodes the topic vector and the preceding syntax state to produce the following sentence. To further reduce delayed repetition in generated paragraphs, we devise a replacement-based reward for the REINFORCE training. Comprehensive experiments on the widely used benchmark demonstrate the superiority of the proposed model over the state of the art for coherence while maintaining high accuracy. Qi Zheng 0003, Dadong Wang |
Comput. Vis. Image Underst. | 3 |
| 2024 | BAVS: Bootstrapping Audio-Visual Segmentation by Integrating Foundation KnowledgeabstractGiven an audio-visual pair, audio-visual segmentation (AVS) aims to locate sounding sources by predicting pixel-wise maps. Previous methods assume that each sound component in an audio signal always has a visual counterpart in the image. However, this assumption overlooks that off-screen sounds and background noise often contaminate the audio recordings in real-world scenarios. They impose significant challenges on building a consistent semantic mapping between audio and visual signals for AVS models and thus impede precise sound localization. In this work, we propose a two-stage bootstrapping audio-visual segmentation framework by incorporating multi-modal foundation knowledge$^{1}$In a nutshell, our BAVS is designed to eliminate the interference of background noise or off-screen sounds in segmentation by establishing the audio-visual correspondences in an explicit manner. In the first stage, we employ a segmentation model to localize potential sounding objects from visual data without being affected by contaminated audio signals. Meanwhile, we also utilize a foundation audio classification model to discern audio semantics. Considering the audio tags provided by the audio foundation model are noisy, associating object masks with audio tags is not trivial. Thus, in the second stage, we develop an audio-visual semantic integration strategy (AVIS) to localize the authentic-sounding objects. Here, we construct an audio-visual tree based on the hierarchical correspondence between sounds and object categories. We then examine the label concurrency between the localized objects and classified audio tags by tracing the audio-visual tree. With AVIS, we can effectively segment real-sounding objects. Extensive experiments demonstrate the superiority of our method on AVS datasets, particularly in scenarios involving background noise. Our project website ishttps://yenanliu.github.io/AVSS.github.io/. Chen Liu 0028, Peike Li, Hu Zhang 0005, Lincheng Li, Zi Huang, Dadong Wang, Xin Yu 0002 |
IEEE Trans. Multim. | 6 |
| 2023 | PADDLES: Phase-Amplitude Spectrum Disentangled Early Stopping for Learning with Noisy LabelsabstractConvolutional Neural Networks (CNNs) are powerful in learning patterns of different vision tasks, but they are sensitive to label noise and may overfit to noisy labels during training. The early stopping strategy averts updating CNNs during the early training phase and is widely employed in the presence of noisy labels. Motivated by biological findings that the amplitude spectrum (AS) and phase spectrum (PS) in the frequency domain play different roles in the animal’s vision system, we observe that PS, which captures more semantic information, can increase the robustness of CNNs to label noise, more so than AS can. We thus propose early stops at different times for AS and PS by disentangling the features of some layer(s) into AS and PS using Discrete Fourier Transform (DFT) during training. Our proposed Phase-AmplituDe DisentangLed Early Stopping (PADDLES) method is shown to be effective on both synthetic and real-world label-noise datasets. PADDLES out-performs other early stopping methods and obtains state-of-the-art performance. Huaxi Huang, Olivier Salvado, Thierry Rakotoarivelo, Dadong Wang, Tongliang Liu |
ICCV | 6 |
| 2023 | Detection of Basal Cell Carcinoma in Whole Slide Images
Hongyan Xu 0002, Dadong Wang, Arcot Sowmya, Ian Katz |
MICCAI (6) | 2 |
| 2023 | Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsabstractThe audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior arts are prone to segment a certain salient object in a video regardless of the audio information. This is because sounding objects are often the most salient ones in the AVS dataset. Thus, current AVS methods might fail to localize genuine sounding objects due to the dataset bias. In this work, we present an audio-visual instance-aware segmentation approach to overcome the dataset bias. In a nutshell, our method first localizes potential sounding objects in a video by an object segmentation network, and then associates the sounding object candidates with the given audio. We notice that an object could be a sounding object in one video but a silent one in another video. This would bring ambiguity in training our object segmentation network as only sounding objects have corresponding segmentation masks. We thus propose a silent object-aware segmentation objective to alleviate the ambiguity. Moreover, since the category information of audio is unknown, especially for multiple sounding sources, we propose to explore the audio-visual semantic correlation and then associate audio with potential objects. Specifically, we attend predicted audio category scores to potential instance masks and these scores will highlight corresponding sounding instances while suppressing inaudible ones. When we enforce the attended instance masks to resemble the ground-truth mask, we are able to establish audio-visual semantics correlation. Experimental results on the AVS benchmarks demonstrate that our method can effectively segment sounding objects without being biased to salient objects and also achieves state-of-the-art performance in both the single-source and multi-source scenarios. Chen Liu 0028, Peike Li, Xingqun Qi, Hu Zhang 0005, Lincheng Li, Dadong Wang, Xin Yu 0002 |
ACM Multimedia | 6 |
| 2023 | Subclass-Dominant Label Noise: A Counterexample for the Success of Early StoppingabstractIn this paper, we empirically investigate a previously overlooked and widespread type of label noise, subclass-dominant label noise (SDN). Our findings reveal that, during the early stages of training, deep neural networks can rapidly memorize mislabeled examples in SDN. This phenomenon poses challenges in effectively selecting confident examples using conventional early stopping techniques. To address this issue, we delve into the properties of SDN and observe that long-trained representations are superior at capturing the high-level semantics of mislabeled examples, leading to a clustering effect where similar examples are grouped together. Based on this observation, we propose a novel method called NoiseCluster that leverages the geometric structures of long-trained representations to identify and correct SDN. Our experiments demonstrate that NoiseCluster outperforms state-of-the-art baselines on both synthetic and real-world datasets, highlighting the importance of addressing SDN in learning with noisy labels. The code is available at https://github.com/tmllab/2023_NeurIPS_SDN. Yingbin Bai, Zhongyi Han, Erkun Yang, Jun Yu 0001, Bo Han 0003, Dadong Wang, Tongliang Liu |
NeurIPS | 6 |
| 2023 | A Multi-Scale Context Aware Attention Model for Medical Image SegmentationabstractMedical image segmentation is critical for efficient diagnosis of diseases and treatment planning. In recent years, convolutional neural networks (CNN)-based methods, particularly U-Net and its variants, have achieved remarkable results on medical image segmentation tasks. However, they do not always work consistently on images with complex structures and large variations in regions of interest (ROI). This could be due to the fixed geometric structure of the receptive fields used for feature extraction and repetitive down-sampling operations that lead to information loss. To overcome these problems, the standard U-Net architecture is modified in this work by replacing the convolution block with a dilated convolution block to extract multi-scale context features with varying sizes of receptive fields, and adding a dilated inception block between the encoder and decoder paths to alleviate the problem of information recession and the semantic gap between features. Furthermore, the input of each dilated convolution block is added to the output through a squeeze and excitation unit, which alleviates the vanishing gradient problem and improves overall feature representation by re-weighting the channel-wise feature responses. The original inception block is modified by reducing the size of the spatial filter and introducing dilated convolution to obtain a larger receptive field. The proposed network was validated on three challenging medical image segmentation tasks with varying size ROIs: lung segmentation on chest X-ray (CXR) images, skin lesion segmentation on dermoscopy images and nucleus segmentation on microscopy cell images. Improved performance compared to state-of-the-art techniques demonstrates the effectiveness and generalisability of the proposed Dilated Convolution and Inception blocks-based U-Net (DCI-UNet). Md. Shariful Alam, Dadong Wang, Qiyu Liao, Arcot Sowmya |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Data Agnostic Filter Gating For Efficient Deep NetworksabstractFilter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of the CNN’s loss function, thereby ignoring the variance of input batches to differences in sparse structure over the filters. In this paper, we propose a data-agnostic filter pruning method that uses an auxiliary network named Dagger module to induce pruning with the pre-trained weights as input. Besides, to help prune filters with a preset FLOPs constraint, we utilize an explicit FLOPs-aware regularisation mechanism to directly promote pruning filters toward the target FLOPs. Experimental results on CIFAR-10 and ImageNet datasets show that the proposed filter pruning method surpasses the state-of-the-art. Hongyan Xu 0002, Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002, Dadong Wang, Arcot Sowmya |
ICASSP | 9 |
| 2022 | Multi-scale alignment and Spatial ROI Module for COVID-19 DiagnosisabstractCoronavirus Disease 2019 (COVID-19) has spread globally and become a health crisis faced by humanity since first reported. Radiology imaging technologies such as computer tomography (CT) and chest X-ray imaging (CXR) are effective tools for diagnosing COVID-19. However, in CT and CXR images, the infected area occupies only a small part of the image. Some common deep learning methods that integrate large-scale receptive fields may cause the loss of image detail, resulting in the omission of the region of interest (ROI) in COVID-19 images and are therefore not suitable for further processing. To this end, we propose a deep spatial pyramid pooling (D-SPP) module to integrate contextual information over different resolutions, aiming to extract information under different scales of COVID-19 images effectively. Besides, we propose a COVID-19 infection detection (CID) module to draw attention to the lesion area and remove interference from irrelevant information. Extensive experiments on four CT and CXR datasets have shown that our method produces higher accuracy of detecting COVID-19 lesions in CT and CXR images. It can be used as a computer-aided diagnosis tool to help doctors effectively diagnose and screen for COVID-19. Hongyan Xu 0002, Dadong Wang, Arcot Sowmya |
IJCNN | 2 |
| 2022 | RSA: Reducing Semantic Shift from Aggressive Augmentations for Self-supervised LearningabstractMost recent self-supervised learning methods learn visual representation by contrasting different augmented views of images. Compared with supervised learning, more aggressive augmentations have been introduced to further improve the diversity of training pairs. However, aggressive augmentations may distort images' structures leading to a severe semantic shift problem that augmented views of the same image may not share the same semantics, thus degrading the transfer performance. To address this problem, we propose a new SSL paradigm, which counteracts the impact of semantic shift by balancing the role of weak and aggressively augmented pairs. Specifically, semantically inconsistent pairs are of minority, and we treat them as noisy pairs. Note that deep neural networks (DNNs) have a crucial memorization effect that DNNs tend to first memorize clean (majority) examples before overfitting to noisy (minority) examples. Therefore, we set a relatively large weight for aggressively augmented data pairs at the early learning stage. With the training going on, the model begins to overfit noisy pairs. Accordingly, we gradually reduce the weights of aggressively augmented pairs. In doing so, our method can better embrace aggressive augmentations and neutralize the semantic shift problem. Experiments show that our model achieves 73.1% top-1 accuracy on ImageNet-1K with ResNet-50 for 200 epochs, which is a 2.5% improvement over BYOL. Moreover, experiments also demonstrate that the learned representations can transfer well for various downstream tasks. Code is released at: https://github.com/tmllab/RSA. Yingbin Bai, Erkun Yang, Zhaoqing Wang, Bo Han 0003, Cheng Deng 0002, Dadong Wang, Tongliang Liu |
NeurIPS | 7 |
| 2022 | Category attention transfer for efficient fine-grained visual categorizationabstractFine-Grained Visual Categorization (FGVC) aims at distinguishing subordinate-level categories with subtle interclass differences. Although previous research shows the impressive effectiveness of the recurrent multi-attention models and the second-order feature encoding, they often require an enormous amount of both computation and memory space, making them inadequate for mobile applications. This paper proposed a Category Attention Transfer CNN (CAT-CNN) to address the efficiency issue in solving FGVC problems. We transfer part attention knowledge from a very large-scale FGVC network to a small but efficient network to significantly improve its presentation ability. Using the proposed CAT-CNN, the accuracy of the efficient networks, such as ShuffleNet, MobilieNet, and EfficientNet, can be improved by up to 5.7% on the CUB-2011-200 dataset without increasing computation complexity or memory cost. Our experiments show that the proposed CAT-CNN can be applied to multiple structures to enhance their performance. With a single efficient network structure and single inference, the proposed CAT-MobileNet-large-1.0 and the CAT-EfficientNet-b0 can achieve accuracies of 86.5% and 86.7%, respectively, on the CUB-2011-200 dataset, which is close to or better than the results from state-of-the-art methods using large scale networks and multiple inferences, and make FGVC feasible on mobile devices. Qiyu Liao, Dadong Wang, Min Xu 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Weakly Supervised RGB-D Salient Object Detection With Prediction Consistency Training and Active Scribble BoostingabstractRGB-D salient object detection (SOD) has attracted increasingly more attention as it shows more robust results in complex scenes compared with RGB SOD. However, state-of-the-art RGB-D SOD approaches heavily rely on a large amount of pixel-wise annotated data for training. Such densely labeled annotations are often labor-intensive and costly. To reduce the annotation burden, we investigate RGB-D SOD from a weakly supervised perspective. More specifically, we use annotator-friendly scribble annotations as supervision signals for model training. Since scribble annotations are much sparser compared to ground-truth masks, some critical object structure information might be neglected. To preserve such structure information, we explicitly exploit the complementary edge information from two modalities (i.e., RGB and depth). Specifically, we leverage the dual-modal edge guidance and introduce a new network architecture with a dual-edge detection module and a modality-aware feature fusion module. In order to use the useful information of unlabeled pixels, we introduce a prediction consistency training scheme by comparing the predictions of two networks optimized by different strategies. Moreover, we develop an active scribble boosting strategy to provide extra supervision signals with negligible annotation cost, leading to significant SOD performance improvement. Extensive experiments on seven benchmarks validate the superiority of our proposed method. Remarkably, the proposed method with scribble annotations achieves competitive performance in comparison to fully supervised state-of-the-art methods. Yunqiu Xu, Xin Yu 0002, Jing Zhang 0052, Linchao Zhu, Dadong Wang |
IEEE Trans. Image Process. | 5 |
| 2021 | Bidirectional Convolutional-LSTM based Network for lung segmentation of chest X-ray imagesabstractDeep Neural Networks (DNN)-based methods, particularly UNet, are considered as state-of-the-art for many medical imaging tasks. However, despite remarkable progress on segmenting the normal lung, performance of the UNet is unsatisfactory on challenging chest X-ray (CXR) images. This could be due to mainly two limiting factors: (1) skip connections that merge feature maps of similar size from encoding and decoding paths, and (2) loss of spatial information due to repetitive down-sampling operations. To overcome these problems, in this study, we propose a DNN-based new architecture that replaces the skip connections with a bidirectional convolutional-LSTM (BC-LSTM) module that allows exchange of more information between encoder and decoder paths and also capture spatiotemporal information. For further improvement, we add a multiple kernel pooling (MKP) block at the lowest level of UNet to encode more spatial information by different sized pooling operations. To evaluate the performance of our method, we use CXR images with different pulmonary diseases such as tuberculosis, pneumoconiosis, and Covid-19 from four public datasets as well as a private dataset and compare its performance with a standard UNet model. Results suggest that the proposed framework outperforms the UNet for all five datasets on lung segmentation, in terms of two evaluation metrics, namely Dice Coefficient (DC) and Jaccard Index (JI). Md. Sharitul Alam, Dadong Wang, Arcot Sowmya |
ICTAI | 2 |
| 2021 | Local-CycleGAN: a general end-to-end network for visual enhancement in complex deep-water environment
Xianhui Zong, Zhehan Chen, Dadong Wang |
Appl. Intell. | 3 |
| 2021 | PDANet: Pyramid density-aware attention based network for accurate crowd counting
Saeed Amirgholipour Kasmani, Wenjing Jia, Lei Liu 0036, Xiaochen Fan, Dadong Wang, Xiangjian He |
Neurocomputing | 5 |
| 2020 | Structural correlation filters combined with a Gaussian particle filter for hierarchical visual tracking
Manna Dai, Gao Xiao, Shuying Cheng, Dadong Wang, Xiangjian He |
Neurocomputing | 4 |
| 2019 | Design of multi-scale receptive field convolutional neural network for surface inspection of hot rolled steels
Di He 0005, Ke Xu 0005, Dadong Wang |
Image Vis. Comput. | 3 |
| 2019 | Object tracking in the presence of shaking motions
Manna Dai, Shuying Cheng, Xiangjian He, Dadong Wang |
Neural Comput. Appl. | 4 |
| 2018 | A-CCNN: Adaptive CCNN for Density Estimation and Crowd CountingabstractCrowd counting, for estimating the number of people in a crowd using vision-based computer techniques, has attracted much interest in the research community. Although many attempts have been reported, real-world problems, such as huge variation in subjects' sizes in images and serious occlusion among people, make it still a challenging problem. In this paper, we propose an Adaptive Counting Convolutional Neural Network (A-CCNN) and consider the scale variation of objects in a frame adaptively so as to improve the accuracy of counting. Our method takes advantages of contextual information to provide more accurate and adaptive density maps and crowd counting in a scene. Extensively experimental evaluation is conducted using different benchmark datasets for object-counting and shows that the proposed approach is effective and outperforms state-of-the-art approaches. Saeed Amirgholipour Kasmani, Xiangjian He, Wenjing Jia, Dadong Wang, Michelle Zeibots |
ICIP | 4 |
| 2016 | Automated Opal Grading by Imaging and Statistical LearningabstractQuantitative grading of opals is a challenging task even for skilled opal assessors. Current opal evaluation practices are highly subjective due to the complexities of opal assessment and the limitations of human visual observation. In this paper, we present a novel machine vision system for the automated grading of opals-the gemological digital analyzer (GDA). The grading is based on statistical machine learning with multiple characteristics extracted from opal images. The assessment workflow includes calibration, opal image capture, image analysis, and opal classification and grading. Experimental results show that the GDA-based grading is more consistent and objective compared with the manual evaluations conducted by the skilled opal assessors. Dadong Wang, Leanne Bischof, Ryan Lagerstrom, Volker Hilsenstein, Angus Hornabrook, Graham Hornabrook |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2015 | High-throughput and low-complexity binary arithmetic decoder based on logarithmic domainabstractThis paper proposes a low-complexity and high-throughput decoder (D_LBAC) based on Logarithmic Binary Arithmetic Coding (LBAC). The proposed D_LBAC has high throughput and low complexity. It does not use multiplication and division operations nor look up tables (LUTs). The proposed D_LBAC is a simple algorithm structure and only requires additions and shift operations. Experimental results show it can decode 3.5 symbols per cycle on average. The hardware implementation design described in this paper can achieve the high symbol processing capability and the low hardware costs. Quanhe Yu, Xiaozhen Zheng, Jianhua Zheng, Wei Yu 0007, Dadong Wang, Junyou Chen |
ICIP | 6 |
| 2015 | A new efficient bypass coding scheme based on logarithmic domainabstractThis paper proposes a new efficient bypass coding scheme (EBCS) based on Logarithmic Binary Arithmetic Coding (LBAC). The bypass coding model is used to encode a symbol which has equal probability (0.5). The percentage of the bypass coding model is about 25 in CABAC of H.265/HEVC. The proposed EBCS provides a hardware-efficient design that can significantly increase the processing speed, and it has a simple algorithm structure. Experimental results show that the EBCS can reduce bypass coding time by 60% roughly. For a hardware implementation in this paper, the overall processing speed is improved by about 96%, and the hardware cost is low. Quanhe Yu, Xiaozhen Zheng, Jianhua Zheng, Wei Yu 0007, Dadong Wang, Junyou Chen |
PCS | 6 |
| 2015 | Parallel multi-level 2D-DWT on CUDA GPUs and its application in ring artifact removalabstractSummary This paper presented two schemes of parallel 2D discrete wavelet transform (DWT) on Compute Unified Device Architecture graphics processing units. For the first scheme, the image and filter are transformed to spectral domain by using Fast Fourier Transformation (FFT), multiplied and then transformed back to space domain by using inverse FFT. For the second scheme, the image pixels are convolved directly with filters. Because there is no data relevance, the convolution for data points on different positions could be executed concurrently. To reduce data transfer, the boundary extension and down‐sampling are processed during data loading stage, and transposing is completed implicitly during data storage. A similar skill is adopted when parallelizing inverse 2D DWT. To further speed up the data access, the filter coefficients are stored in the constant memory. We have parallelized the 2D DWT for dozens of wavelet types and achieved a speedup factor of over 380 times compared with that of its CPU version. We applied the parallel 2D DWT in a ring artifact removal procedure; the executing speed was accelerated near 200 times compared with its CPU version. The experimental results showed that the proposed parallel 2D DWT on graphics processing units can significantly improve the performance for a wide variety of wavelet types and is promising for various applications. Copyright © 2015 John Wiley & Sons, Ltd. Leqing Zhu, Daxing Zhang, Dadong Wang, Huiyan Wang 0002, Xun Wang 0007 |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | Soft Cost Aggregation with Multi-resolution Fusion
Xiao Tan 0001, Changming Sun, Dadong Wang, Yi Guo 0001, Tuan D. Pham |
ECCV (5) | 3 |
| 2014 | An improved method for the removal of ring artifacts in synchrotron radiation images by using GPGPU computing with compute unified device architectureabstractSUMMARY Ring artifacts are a common problem in computed tomography, positron emission tomography, magnetic resonance imaging, and synchrotron radiation images. Before further processing the images such as segmentation and quantification, these artifacts have to be removed or suppressed. Otherwise, they may introduce additional errors for the segmentation and subsequent analysis. This paper proposes an improved ring artifact removal method based on biorthogonal wavelet transform, one‐dimensional fast Fourier transform, and Gaussian damping, which is implemented on general‐purpose computing on graphics processing unit with compute unified device architecture. The experimental results show that the proposed algorithms can be speed up several hundred times compared with the previous algorithms on CPU. The significant performance improvement makes the algorithms much more practical in processing large volume of images in real time. Copyright © 2013 John Wiley & Sons, Ltd. Leqing Zhu, Dadong Wang, Huiyan Wang 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Cloud Computing for High Performance Image Analysis on a National InfrastructureabstractCloud computing services offer highly reliable, scalable and efficient solutions with a large pool of easily accessible, virtualized resources. They are becoming an increasingly prevalent delivery model. We have developed a cloud-based image analysis toolbox to provide a wide user base with easy access to the software tools we have developed over the last decade. The toolbox is provided as a service on an Australian national cloud infrastructure. The design and implementation of the cloud-based service are presented, including its architecture, key components and some image analysis and visualization examples showing the capabilities of the service for biomedical image analysis. Dadong Wang, Tomasz Bednarz, Yulia Arzhaeva, John A. Taylor, Piotr Szul, Shiping Chen 0001, Neil Burdett, Alex Khassapov, Tim E. Gureyev |
CCGRID | 1 |
| 2013 | Cloud based Services for Biomedical Image Analysis
Dadong Wang, Tomasz Bednarz, Yulia Arzhaeva, Piotr Szul, Shiping Chen 0001, Neil Burdett, Alex Khassapov, Tim E. Gureyev, John A. Taylor |
CLOSER | 1 |
| 2013 | Decomposition of volume scattering, polarized light and chlorophyll fluorescence by in-situ polarization measurementabstractThe remotely sensed radiation of green vegetation canopy over the spectral region of 350-2500 nm typically mixes with three components of different optical properties, namely the leaf interior volume scattering, canopy surface polarized light and leaf internal emitted chlorophyll fluorescence (ChlF). They are tightly superimposed together but convey different information of vegetation. This study emphasizes the distinction of the three radiant fluxes above and disentangles them from the observed apparent radiance of Scindapsus aureus canopy by in-situ polarization measurements. Results demonstrate that the polarization measurement enables the quantitatively separation of the volume scattering, polarized light and ChlF. This study provides further understanding of light scattering properties of the vegetation canopy and particularly has the potential of allowing improvements of current reflectance-based vegetation models. Changping Huang, Dadong Wang, Taixia Wu, Qingxi Tong |
IGARSS | 3 |
| 2009 | Membrane boundary extraction using circular multiple paths
Changming Sun, Pascal Vallotton, Dadong Wang, Jamie Lopez, Yvonne Ng, David E. James |
Pattern Recognit. | 3 |