Bin Duan 0004

dblp:67/2554-4 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
0009-0003-3138-0923ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Automated Update of Android Deprecated API Usages With Large Language Models
abstract
Android apps rely on application programming interfaces (APIs) to access various functionalities of Android devices. These APIs however are regularly updated to incorporatenew features while the old APIs get deprecated. Even though the importance of updating deprecated API usages with the recommended replacement APIs has been widely recognized, it is non-trivial to update the deprecated API usages. Therefore, the usages of deprecated APIs linger in Android apps and cause compatibility issues in practice. This paper introduces GUPPY, an automated approach that utilizes large language models (LLMs) to update Android deprecated API usages. By employing carefully crafted Chain-of-Thoughts prompts, GUPPY leverages GPT-4, one of the most powerful LLMs, to update deprecated-API usages, ensuring compatibility in both the old and new API levels. Additionally, GUPPY uses GPT-4 to generate tests, identify incorrect updates, and refine the API usage through an iterative process until the tests pass or a specified limit is reached. Our evaluation, conducted on 360 benchmark API usages from 20 deprecated APIs and an additional 156 deprecated API usages from the latest API levels 33 and 34, demonstrates GUPPY’s advantages over the state-of-the-art techniques.
Tarek Mahmud, Bin Duan 0004, Meiru Che, Awatif Yasmin, Anne H. H. Ngu, Guowei Yang 0001
IEEE Trans. Software Eng.2
2025 MaskSAM: Auto-Prompt SAM with Mask Classification for Volumetric Medical Image Segmentation
Hao Tang 0005, Bin Duan 0004, Dawen Cai, Yan Yan 0002, Gady Agam
ICCV3
2025 Harnessing LLMs for Document-Guided Fuzzing of OpenCV Library
abstract
The combination of computer vision and artificial intelligence is fundamentally transforming a broad spectrum of industries by enabling machines to interpret and act upon visual data with high levels of accuracy. As the biggest and by far the most popular open-source computer vision library, OpenCV library provides an extensive suite of programming functions supporting real-time computer vision. Bugs in the OpenCV library can affect the downstream computer vision applications, and it is critical to ensure the reliability of the OpenCV library. This paper introduces VistaFuzz, a novel technique for harnessing large language models (LLMs) for document-guided fuzzing of the OpenCV library. Vistafuzz utilizes LLMs to parse API documentation and obtain standardized API information. Based on this standardized information, Vista Fuzz extracts constraints on individual input parameters and dependencies between these. Using these constraints and dependencies, VistaFuzz then generates new input values to systematically test each target API. We evaluate the effectiveness of Vistafuzz in testing 330 APIs in the OpenCV library, and the results show that Vistafuzz detected 17 new bugs, where 10 bugs have been confirmed, and 5 of these have been fixed.
Bin Duan 0004, Tarek Mahmud, Meiru Che, Yan Yan 0002, Naipeng Dong, Dong Seong Kim 0001, Guowei Yang 0001
ICSME1
2025 XAMT: Cross-Framework API Matching for Testing Deep Learning Libraries
abstract
Deep learning powers critical applications such as autonomous driving, healthcare, and finance, where the correctness of underlying libraries is essential. Bugs in widely used deep learning APIs can propagate to downstream systems, causing serious consequences. While existing fuzzing techniques detect bugs through intra-framework testing across hardware backends (CPU vs. GPU), they may miss bugs that manifest identically across backends and thus escape detection under these strategies. To address this problem, we propose XAMT, a cross-framework fuzzing method that tests deep learning libraries by matching and comparing functionally equivalent APIs across different frameworks. XAMT matches APIs using similarity-based rules based on names, descriptions, and parameter structures. It then aligns inputs and applies variance-guided differential testing to detect bugs. We evaluated XAMT on five popular frameworks, including PyTorch, TensorFlow, Keras, Chainer, and JAX. XAMT matched 839 APIs and identified 238 matched API groups, and detected 17 bugs, 12 of which have been confirmed. Our results show that XAMT uncovers bugs undetectable by intraframework testing, especially those that manifest consistently across backends. XAMT offers a complementary approach to existing methods and offers a new perspective on the testing of deep learning libraries.
Bin Duan 0004, Ruican Dong, Naipeng Dong, Dong Seong Kim 0001, Guowei Yang 0001
ISSRE1
2025 X-Field: A Physically Informed Representation for 3D X-ray Reconstruction
abstract
X-ray imaging is indispensable in medical diagnostics, yet its use is tightly regulated due to radiation exposure. Recent research borrows representations from the 3D reconstruction area to complete two tasks with reduced radiation dose: X-ray Novel View Synthesis (NVS) and Computed Tomography (CT) reconstruction. However, these representations fail to fully capture the penetration and attenuation properties of X-ray imaging as they originate from visible light imaging. In this paper, we introduce X-Field, a 3D representation informed in the physics of X-ray imaging. First, we employ homogeneous 3D ellipsoids with distinct attenuation coefficients to accurately model diverse materials within internal structures. Second, we introduce an efficient path-partitioning algorithm that resolves the intricate intersection of ellipsoids to compute cumulative attenuation along an X-ray path. We further propose a hybrid progressive initialization to refine the geometric accuracy of X-Field and incorporate material-based optimization to enhance model fitting along material boundaries. Experiments show that X-Field achieves superior visual fidelity on both real-world human organ and synthetic object datasets, outperforming state-of-the-art methods in X-ray NVS and CT Reconstruction. Our code is available on the project page: https://github.com/Brack-Wang/X-Field.
Jiachen Tao, Junyi Wu 0002, Haoxuan Wang 0002, Bin Duan 0004, Kai Wang 0036, Zongxin Yang, Yan Yan 0002
NeurIPS5
2024 Token Transformation Matters: Towards Faithful Post-Hoc Explanation for Vision Transformer
abstract
While Transformers have rapidly gained popularity in various computer vision applications, post-hoc explanations of their internal mechanisms remain largely unexplored. Vision Transformers extract visual information by representing image regions as transformed tokens and integrating them via attention weights. However, existing post-hoc explanation methods merely consider these attention weights, neglecting crucial information from the transformed tokens, which fails to accurately illustrate the rationales behind the models' predictions. To incorporate the influence of token transformation into interpretation, we propose TokenTM, a novel post-hoc explanation method that utilizes our introduced measurement of token transformation effects. Specifically, we quantify token transformation effects by measuring changes in token lengths and correlations in their directions pre- and post-transformation. Moreover, we develop initialization and aggregation rules to integrate both attention weights and token transformation effects across all layers, capturing holistic token contributions throughout the model. Experimental results on segmentation and perturbation tests demonstrate the superiority of our proposed TokenTM compared to state-of-the-art Vision Transformer explanation methods.
Junyi Wu 0002, Bin Duan 0004, Weitai Kang, Hao Tang 0005, Yan Yan 0002
CVPR2
2024 Adaptive Cross-Architecture Mutual Knowledge Distillation
abstract
Knowledge distillation (KD), which distills knowledge from complex networks (teacher) to lightweight (student) networks, has been actively studied recently. Despite previous studies have proposed several advanced KD losses or intricate training strategies, the core concept of KD proves ineffective if the student model is too weak to mimic the teacher's performance. In this study, we aim to narrow the performance discrepancy between Transformer-based teacher and student models by incorporating the inductive biases of several heterogeneous student models. To this end, we put forward a novel cross-architecture knowledge distillation approach called Adaptive Cross-architecture Mutual Knowledge Distillation (ACMKD), which tries to mitigate the performance gap issue using a multi-students mutual learning strategy. Specifically, we utilize three mainstream models associated with various inductive biases (CNN, INN, and Transformer) as the student models. In addition, we propose an effective attention similarity mechanism to facilitate the student models in mimicking specific portions of the teacher model. Drawing inspiration from the Cannikin Law, we devise a unique second-stage KD process that dynamically enables the weakest student model to learn from other stronger student models again. We validate our proposed methods on ImageNet and CIFAR100 datasets, and the results confirm that our ACMKD method significantly narrows the performance gap compared to other KD methods.
Jianyuan Ni, Hao Tang 0005, Yuzhang Shang, Bin Duan 0004, Yan Yan 0002
FG4
2024 Mining and Unifying Heterogeneous Contrastive Relations for Weakly-Supervised Actor-Action Segmentation
abstract
We introduce a novel weakly-supervised video actor-action segmentation (VAAS) framework, where only video-level tags are available. Previous VAAS methods follow a synthesize-and-refine scheme, i.e., they first synthesize the pseudo-segmentation and recursively refine the segmentation. However, this process requires significant time costs and heavily relies on the quality of the initial segmentation. Unlike existing works, our method hierarchically mines contrastive relations to supplement each other for learning a visually-plausible segmentation model. Specifically, three contrastive relations are abstracted from the pixel-level and frame-level, i.e., low-level edge-aware, class-activation map aware, and semantic tag-aware relations. Then, the discovered contrastive relations are unified into a universal objective for training the segmentation model, regardless of their heterogeneity. Moreover, we incorporate motion cues and unlabeled samples to increase the discriminative power and robustness of the segmentation model. Extensive experiments indicate that our proposed method produces reasonable segmentation.
Bin Duan 0004, Hao Tang 0005, Changchang Sun, Yan Yan 0002
WACV1
2023 MLP-GAN for Brain Vessel Image Segmentation
abstract
Brain vessel image segmentation can be used as a promising biomarker for better prevention and treatment of different diseases. One successful approach is to consider the segmentation as an image-to-image translation task and perform a conditional Generative Adversarial Network (cGAN) to learn a transformation between two distributions. In this paper, we present a novel multi-view approach, MLP-GAN, which splits a 3D volumetric brain vessel image into three different dimensional 2D images (i.e., sagittal, coronal, axial) and then feed them into three different 2D cGANs. The proposed MLP-GAN not only alleviates the memory issue which exists in the original 3D neural networks but also retains 3D spatial information. Specifically, we utilize U-Net as the backbone for our generator and redesign the pattern of skip connection integrated with the MLP-Mixer [1] which has attracted lots of attention recently. Our model obtains the ability to capture cross-patch information to learn global information with the MLP-Mixer. Extensive experiments are performed on the public brain vessel dataset [2] that show our MLP-GAN outperforms other state-of-the-art methods.
Hao Tang 0005, Bin Duan 0004, Dawen Cai, Yan Yan 0002
ICASSP3
2023 Towards Saner Deep Image Registration
abstract
With recent advances in computing hardware and surges of deep-learning architectures, learning-based deep image registration methods have surpassed their traditional counterparts, in terms of metric performance and inference time. However, these methods focus on improving performance measurements such as Dice, resulting in less attention given to model behaviors that are equally desirable for registrations, especially for medical imaging. This paper investigates these behaviors for popular learning-based deep registrations under a sanity-checking microscope. We find that most existing registrations suffer from low inverse consistency and nondiscrimination of identical pairs due to overly optimized image similarities. To rectify these behaviors, we propose a novel regularization-based sanity-enforcer method that imposes two sanity checks on the deep model to reduce its inverse consistency errors and increase its discriminative power simultaneously. Moreover, we derive a set of theoretical guarantees for our sanity-checked image registration method, with experimental results supporting our theoretical findings and their effectiveness in increasing the sanity of models without sacrificing any performance.
Bin Duan 0004, Ming Zhong 0008, Yan Yan 0002
ICCV1
2022 Learning Omnidirectional Flow in 360$^\circ $ Video via Siamese Representation
Keshav Bhandari, Bin Duan 0004, Gaowen Liu, Hugo Latapie, Ziliang Zong, Yan Yan 0002
ECCV (8)2
2022 Lipschitz Continuity Retained Binary Neural Network
Yuzhang Shang, Dan Xu 0002, Bin Duan 0004, Ziliang Zong, Liqiang Nie, Yan Yan 0002
ECCV (11)3
2022 Win The Lottery Ticket Via Fourier Analysis: Frequencies Guided Network Pruning
abstract
With the remarkable success of deep learning recently, efficient network compression algorithms are urgently demanded for releasing the potential computational power of edge devices, such as smartphones or tablets. However, optimal network pruning is a non-trivial task which mathematically is an NP-hard problem. Previous researchers explain training a pruned network as buying a lottery ticket. In this paper, we investigate the Magnitude-Based Pruning (MBP) scheme and analyze it from a novel perspective through Fourier analysis on the deep learning model to guide model designation. Besides explaining the generalization ability of MBP using Fourier transform, we also propose a novel two-stage pruning approach, where one stage is to obtain the topological structure of the pruned network and the other stage is to retrain the pruned network to recover the capacity using knowledge distillation from lower to higher on the frequency domain. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate the superiority of our novel Fourier analysis based MBP compared to other traditional MBP algorithms.
Yuzhang Shang, Bin Duan 0004, Ziliang Zong, Liqiang Nie, Yan Yan 0002
ICASSP2
2021 Lipschitz Continuity Guided Knowledge Distillation
abstract
Knowledge distillation has become one of the most important model compression techniques by distilling knowledge from larger teacher networks to smaller student ones. Although great success has been achieved by prior distillation methods via delicately designing various types of knowledge, they overlook the functional properties of neural networks, which makes the process of applying those techniques to new tasks unreliable and non-trivial. To alleviate such problem, in this paper, we initially leverage Lipschitz continuity to better represent the functional characteristic of neural networks and guide the knowledge distillation process. In particular, we propose a novel Lipschitz Continuity Guided Knowledge Distillation framework to faithfully distill knowledge by minimizing the distance between two neural networks’ Lipschitz constants, which enables teacher networks to better regularize student networks and improve the corresponding performance. We derive an explainable approximation algorithm with an explicit theoretical derivation to address the NP-hard problem of calculating the Lipschitz constant. Experimental results have shown that our method outperforms other benchmarks over several knowledge distillation tasks (e.g., classification, segmentation and object detection) on CIFAR-100, ImageNet, and PASCAL VOC datasets. Our code is available at https://github.com/42Shawn/LONDON/tree/master.
Yuzhang Shang, Bin Duan 0004, Ziliang Zong, Liqiang Nie, Yan Yan 0002
ICCV2
2021 Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention
abstract
The major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that the attention mechanism is beneficial to the fusion process. In this paper, we propose a novel joint attention mechanism with multi-modal fusion methods for audio-visual event localization. Particularly, we present a concise yet valid architecture that effectively learns representations from multiple modalities in a joint manner. Initially, visual features are combined with auditory features and then turned into joint representations. Next, we make use of the joint representations to attend to visual features and auditory features, respectively. With the help of this joint co-attention, new visual and auditory features are produced, and thus both features can enjoy the mutually improved benefits from each other. It is worth noting that the joint co-attention unit is recursive meaning that it can be performed multiple times for obtaining better joint representations progressively. Extensive experiments on the public AVE dataset have shown that the proposed method achieves significantly better results than the state-of-the-art methods.
Bin Duan 0004, Hao Tang 0005, Wei Wang 0108, Ziliang Zong, Guowei Yang 0001, Yan Yan 0002
WACV1
2020 Cascade Attention Guided Residue Learning GAN for Cross-Modal Translation
abstract
Since we were babies, we intuitively develop the ability to correlate the input from different cognitive sensors such as vision, audio, and text. However, in machine learning, this cross-modal learning is a nontrivial task because different modalities have no homogeneous properties. Previous works discover that there should be bridges among different modalities. From a neurology and psychology perspective, humans have the capacity to link one modality with another one, e.g., associating a picture of a bird with the only hearing of its singing and vice versa. Is it possible for machine learning algorithms to recover the scene given the audio signal? In this paper, we propose a novel Cascade Attention-Guided Residue GAN (CAR-GAN), aiming at reconstructing the scenes given the corresponding audio signals. Particularly, we present a residue module to mitigate the gap between different modalities progressively. Moreover, a cascade attention guided network with a novel classification loss function is designed to tackle the cross-modal learning task. Our model keeps consistency in the high-level semantic label domain and is able to balance two different modalities. The experimental results demonstrate that our model achieves the state-of-the-art cross-modal audio-visual generation on the challenging Sub-URMP dataset.
Bin Duan 0004, Wei Wang 0108, Hao Tang 0005, Hugo Latapie, Yan Yan 0002
ICPR1