Youshan Zhang

dblp:228/8382 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal
abstract
Purple flare, a diffuse chromatic aberration artifact commonly found around highlight areas, severely degrades the tone transition and color of the image. Existing traditional methods are based on hand-crafted features, which lack flexibility and rely entirely on fixed priors, while the scarcity of paired training data critically hampers deep learning. To address this issue, we propose a novel network built upon decoupled HSV Look-Up Tables (LUTs). The method aims to simplify color correction by adjusting the Hue (H), Saturation (S), and Value (V) components independently. This approach resolves the inherent color coupling problems in traditional methods. Our model adopts a two-stage architecture: First, a Chroma-Aware Spectral Tokenizer (CAST) converts the input image from RGB space to HSV space and independently encodes the Hue (H) and Value (V) channels into a set of semantic tokens describing the Purple flare status; second, the HSV-LUT module takes these tokens as input and dynamically generates independent correction curves (1D-LUTs) for the three channels H, S, and V. To effectively train and validate our model, we built the first large-scale purple flare dataset with diverse scenes. We also proposed new metrics and a loss function specifically designed for this task. Extensive experiments demonstrate that our model not only significantly outperforms existing methods in visual effects but also achieves state-of-the-art performance on all quantitative metrics.
Pu Wang 0008, Shuning Sun, Jialang Lu, Chen Wu 0006, Youshan Zhang, Chenggang Shan, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
AAAI6
2026 UHD image dehazing via anDehazeFormer with atmospheric-aware KV cache
Pu Wang 0008, Zhixuan Mao, Wenhao Li 0006, Liubing Hu, Dianjie Lu, Guijuan Zhang, Youshan Zhang, Zhuoran Zheng
Neurocomputing7
2025 AgentPolyp: Accurate Polyp Segmentation via Image Enhancement Agent
abstract
Captured polyp images often suffer from degradation, such as dim lighting, blur, and overexposure. Direct segmentation is prone to artifact diffusion, which significantly degrades the performance of downstream segmentation algorithms and leads to inaccurate boundary delineation. Addressing these varied degradations requires a dynamic, intelligent process that diagnoses and applies targeted corrections. We present AgentPolyp, a novel framework driven by an intelligent agent that integrates CLIP-based semantic guidance and dynamic image enhancement with a lightweight segmentation network. The agent adaptively selects reinforcement learning strategies to perform context-aware denoising, contrast adjustment, and artifact reduction. This selection process is continuously optimized through a feedback loop that includes quality assessment, ensuring the optimization and enhancement of downstream segmentation. This approach addresses degradation complexity and feature compatibility issues, offering a deployable solution for endoscopic polyp analysis.
Pu Wang 0008, Guangwei Gao, Youshan Zhang, Zhuoran Zheng
IEEE Signal Process. Lett.4
2024 SparrowVQE: Visual Question Explanation for Course Content Understanding
abstract
Visual Question Answering (VQA) research seeks to create AI systems to answer natural language questions in images, yet VQA methods often yield overly simplistic and short answers. This paper aims to advance the field by introducing Visual Question Explanation (VQE), which enhances the ability of VQA to provide detailed explanations rather than brief responses and address the need for more complex interaction with visual content. We first created an MLVQE dataset from a 14-week streamed video machine learning course, including 885 slide images, 110,407 words of transcripts, and 9,416 designed question-answer (QA) pairs. Next, we proposed a novel SparrowVQE, a small 3 billion parameters multimodal model. We trained our model with a three-stage training mechanism consisting of multimodal pre-training (slide images and transcripts feature alignment), instruction tuning (tuning the pre-trained model with transcripts and QA pairs), and domain fine-tuning (fine-tuning slide image and QA pairs). Eventually, our SparrowVQE can understand and connect visual information using the SigLIP model with transcripts using the Phi-2 language model with an MLP adapter. Experimental results demonstrate that our SparrowVQE achieves better performance in our developed MLVQE dataset and outperforms state-of-the-art methods in the other five benchmark VQA datasets. The source code is available at https://github.com/rrymn/SparrowVQE.
Jialu Li 0004, Manish Kumar Thota, Ruslan Gokhman, Radek Holik, Youshan Zhang
IEEE Big Data5
2024 Pink Guardian: A Gateway to Early Breast Cancer Detection
abstract
Breast cancer (BC) remains a pivotal concern in global health, and existing methods for early detection are often limited by accessibility and diagnostic accuracy. We present “Pink Guardian,” a mobile application that offers a significant advancement in the early detection of breast cancer, leveraging a proposed BC-Inception V3 architecture within TensorFlow Lite for real-time, precise classification of mammogram images. This sophisticated model has been integrated into a user-friendly interface, enabling users to easily obtain diagnostic predictions and corresponding confidence scores. Through rigorous experiments, BC-Inception-V3 emerged as the best model, showcasing superior training and test accuracies, thus signifying a break-through in clinical breast cancer screening that could be readily available on users' smartphones. Source code is available at https://github.com/kanchanmaurya95/PinkGuardian.git.
Kanchan Subhashchandra Maurya, Youshan Zhang
HSI2
2024 Vision Transformer Segmentation for Visual Bird Sound Denoising
Sahil Kumar, Jialu Li 0004, Youshan Zhang
INTERSPEECH3
2024 Complex Image-Generative Diffusion Transformer for Audio Denoising
Pu Wang 0008, Jialu Li 0004, Youshan Zhang
INTERSPEECH4
2024 Diffusion Gaussian Mixture Audio Denoise
Pu Wang 0008, Jialu Li 0004, Youshan Zhang
INTERSPEECH5
2023 Complex Image Generation SwinTransformer Network for Audio Denoising
abstract
Achieving high-performance audio denoising is still a challenging task in real-world applications.Existing time-frequency methods often ignore the quality of generated frequency domain images.This paper converts the audio denoising problem into an image generation task.We first develop a complex image generation SwinTransformer network to capture more information from the complex Fourier domain.We then impose structure similarity and detailed loss functions to generate highquality images and develop an SDR loss to minimize the difference between denoised and clean audios.Extensive experiments on two benchmark datasets demonstrate that our proposed model is better than state-of-the-art methods.
Youshan Zhang, Jialu Li 0004
INTERSPEECH1
2023 BirdSoundsDenoising: Deep Visual Audio Denoising for Bird Sounds
abstract
Audio denoising has been explored for decades using both traditional and deep learning-based methods. However, these methods are still limited to either manually added artificial noise or lower denoised audio quality. To overcome these challenges, we collect a large-scale natural noise bird sound dataset. We are the first to transfer the audio denoising problem into an image segmentation problem and propose a deep visual audio denoising (DVAD) model. With a total of 14,120 audio images, we develop an audio ImageMask tool and propose to use a few-shot generalization strategy to label these images. Extensive experimental results demonstrate that the proposed model achieves state-of-the-art performance. We also show that our method can be easily generalized to speech denoising, audio separation, audio enhancement, and noise estimation.
Youshan Zhang, Jialu Li 0004
WACV1
2021 Deep Least Squares Alignment for Unsupervised Domain Adaptation
Youshan Zhang, Brian D. Davison 0001
BMVC1
2021 Correlated Adversarial Joint Discrepancy Adaptation Network
abstract
Domain adaptation aims to mitigate the domain shift problem when transferring knowledge from one domain into another similar but different domain. However, most existing works rely on extracting marginal features without considering class labels. Moreover, some methods name their model as so-called unsupervised domain adaptation while tuning the parameters using the target domain label. To address these issues, we propose a novel approach called correlated adversarial joint discrepancy adaptation network (CAJNet), which minimizes the joint discrepancy of two domains and achieves competitive performance with tuning parameters using the correlated label. By training the joint features, we can align the marginal and conditional distributions between the two domains. In addition, we introduce a probability-based top-K correlated label (K-label), which is a powerful indicator of the target domain and effective metric to tune parameters to aid predictions. Extensive experiments on benchmark datasets demonstrate significant improvements in classification accuracy over the state of the art.
Youshan Zhang, Brian D. Davison 0001
CBMI1
2021 Enhanced Separable Disentanglement for Unsupervised Domain Adaptation
abstract
Domain adaptation aims to mitigate the domain gap when transferring knowledge from an existing labeled domain to a new domain. However, existing disentanglement-based methods do not fully consider separation between domain-invariant and domain-specific features, which means the domain-invariant features are not discriminative. The reconstructed features are also not sufficiently used during training. In this paper, we propose a novel enhanced separable disentanglement (ESD) model. We first employ a disentangler to distill domain-invariant and domain-specific features. Then, we apply feature separation enhancement processes to minimize contamination between domain-invariant and domain-specific features. Finally, our model reconstructs complete feature vectors, which are used for further disentanglement during the training phase. Extensive experiments from three benchmark datasets outperform state-of-the-art methods, especially on challenging cross-domain tasks.
Youshan Zhang, Brian D. Davison 0001
ICIP1
2021 Adversarial Reinforcement Learning for Unsupervised Domain Adaptation
abstract
Transferring knowledge from an existing labeled domain to a new domain often suffers from domain shift in which performance degrades because of differences between the domains. Domain adaptation has been a prominent method to mitigate such a problem. There have been many pre-trained neural networks for feature extraction. However, little work discusses how to select the best feature instances across different pre-trained models for both the source and target domain. We propose a novel approach to select features by employing reinforcement learning, which learns to select the most relevant features across two domains. Specifically, in this framework, we employ Q-learning to learn policies for an agent to make feature selection decisions by approximating the action-value function. After selecting the best features, we propose an adversarial distribution alignment learning to improve the prediction results. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods.
Youshan Zhang, Brian D. Davison 0001
WACV1
2021 Domain adaptation for object recognition using subspace sampling demons
Youshan Zhang, Brian D. Davison 0001
Multim. Tools Appl.1
2020 Bayesian Geodesic Regression on Riemannian Manifolds
Youshan Zhang
BMVC1
2020 Inference attacks on genomic privacy with an improved HMM and an RCNN model for unrelated individuals
Hongfa Ding, Youliang Tian, Changgen Peng, Youshan Zhang, Shuwen Xiang
Inf. Sci.4
2019 Transductive Learning Via Improved Geodesic Sampling
Youshan Zhang, Sihong Xie, Brian D. Davison 0001
BMVC1
2019 ShapeNet: Age-focused Landmark Shape Prediction with Regressive CNN
abstract
Deep neural networks are widely used in the segmentation and classification of medical images. However, little work has addressed the prediction of shapes based on population data over time as a regression problem. In this paper, we introduce a regressive convolutional neural network for landmark-based shape prediction. Unlike the conventional CNN model, the proposed network takes the input of a target age, and outputs the corresponding shape for that age. Experimental results demonstrate the effectiveness of the proposed ShapeNet to predict corpus callosum and mandible shapes with correct topology and accurate fitting that matches real-world scenarios. The proposed ShapeNet can predict the shape variation of high dimensional and nonlinear data, which is often critical to understanding the processes that change the shape of anatomy in biology and medical fields.
Youshan Zhang, Brian D. Davison 0001
CBMI1
2018 Corticospinal Tract (CST) Reconstruction Based on Fiber Orientation Distributions (FODs) Tractography
abstract
The Corticospinal Tract (CST) is a part of pyramidal tract (PT) and it can innervate the voluntary movement of skeletal muscle through spinal interneurons (the 4th layer of the Rexed gray board layers), and anterior horn motorneurons (which control trunk and proximal limb muscles). Spinal cord injury (SCI) is a highly disabling disease often caused by traffic accidents. The recovery of CST and the functional reconstruction of spinal anterior horn motor neurons play an essential role in the treatment of SCI. However, the localization and reconstruction of CST are still challenging issues, the accuracy of the geometric reconstruction can directly affect the results of the surgery. The main contribution of this paper is the reconstruction of the CST based on the fiber orientation distributions (FODs) tractography. Differing from tensor-based tractography in which the primary direction is a determined orientation, the direction of FODs tractography is determined by the probability. The spherical harmonics (SPHARM) can be used to approximate the efficiency of FODs tractography. We manually delineate the three ROIs (the posterior limb of the internal capsule, the cerebral peduncle, and the anterior pontine area) by the ITK-SNAP software, and use the pipeline software to reconstruct both the left and right sides of the CST fibers. Our results demonstrate that FOD-based tractography can show more and correct anatomical CST fiber bundles.
Youshan Zhang
BIBE1