Youxin Chen

dblp:152/8103 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PAT: Pruning-Aware Tuning for Large Language Models
abstract
Large language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet, traditional post-hoc pruning often leads to significant performance loss, with limited recovery from further fine-tuning due to reduced capacity. Since the model fine-tuning refines the general and chaotic knowledge in pre-trained models, we aim to incorporate structural pruning with the fine-tuning, and propose the Pruning-Aware Tuning (PAT) paradigm to eliminate model redundancy while preserving the model performance to the maximum extend. Specifically, we insert the innovative Hybrid Sparsification Modules (HSMs) between the Attention and FFN components to accordingly sparsify the upstream and downstream linear modules. The HSM comprises a lightweight operator and a globally shared trainable mask. The lightweight operator maintains a training overhead comparable to that of LoRA, while the trainable mask unifies the channels to be sparsified, ensuring structural pruning. Additionally, we propose the Identity Loss which decouples the transformation and scaling properties of the HSMs to enhance training robustness. Extensive experiments demonstrate that PAT excels in both performance and efficiency. For example, our Llama2-7b model with a 25% pruning ratio achieves 1.33x speedup while outperforming the LoRA-finetuned model by up to 1.26% in accuracy with a similar training cost.
Yijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang, Yuan Du
AAAI3
2025 UPR-Net: A Unified Pyramid Recurrent Network for Video Frame Interpolation
Xin Jin 0023, Longhai Wu, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm
Int. J. Comput. Vis.4
2025 O-PRESS: Boosting OCT axial resolution with Prior guidance, Recurrence, and Equivariant Self-Supervision
Kaiyan Li 0005, Jingyuan Yang 0015, Wenxuan Liang, Xingde Li, Lulu Chen, Chan Wu, Xiao Zhang 0059, Zhiyan Xu, Yueling Wang, Lihui Meng, Yue Zhang 0042, Youxin Chen, Shaohua Kevin Zhou
Medical Image Anal.13
2024 Dynamic Video Frame Interpolation with Integrated Difficulty Pre-Assessment
abstract
Video frame interpolation (VFI) has witnessed great progress in recent years. However, existing VFI models still struggle to achieve a good trade-off between accuracy and efficiency. Accurate VFI models typically rely on heavy compute to process all samples, ignoring the fact that easy samples with small motion or clear texture can be well addressed by a fast VFI model and do not require such heavy compute. In this paper, we present a dynamic VFI pipeline with integrated pre-assessment of interpolation difficulty. Specifically, it leverages a difficulty pre-assessment model to measure the difficulty level of interpolating input frames, and then dynamically selects an accurate or a fast VFI model for frame interpolation. Furthermore, we contribute a large-scale annotated dataset to train our VFI difficulty pre-assessment model. Extensive experiments show that our dynamic VFI pipeline can achieve an excellent trade-off between accuracy and efficiency, by feeding hard samples to accurate model, and passing easy samples through fast model.
Ban Chen, Xin Jin 0023, Youxin Chen, Longhai Wu, Jayoon Koo, Cheul-Hee Hahm
ICASSP3
2024 MITER: Medical Image-TExt joint adaptive pretRaining with multi-level contrastive learning
abstract
Recently multimodal medical pretraining models play a significant role in automatic medical image and text analysis that has wide social and economical impact in healthcare. Despite being able to be quickly transferred to downstream tasks, the models are greatly limited due to the fact that these models can only be pretrained with professional medical image-text datasets, which usually contain a very small number of samples. In this work We propose MITER (Medical Image-Text Joint adaptive Pretraining), a joint adaptive pretraining framework via multi-level contrastive learning to overcome this limitation by pretraining image and text models for medical domain and utilizing existing models pretrained on generic data, which contain enormous number of samples. MITER features two types of objectives to solve the problem. The first type is uni-modal objectives that pretrain the models with medical images and text separately on uni-modal tasks. The other type is a cross-modal objective that pretrains jointly, allowing the models to influence each other on cross-modal tasks. We also introduce a strategy to dynamically select hard negative samples during the training process for better performance. Experimental results over four medical tasks, image-report retrieval, multi-label image classification, visual question answering, and report generation, show that our MITER framework solves the limitation problem by greatly outperforming existing benchmark models on all the tasks. The source code of our framework is available online.2
Xiaochu Tang, Jing Xiao 0006, Youxin Chen, Xiu Li 0001, Qian Zhang 0018, Zheng Lu 0002
Expert Syst. Appl.5
2023 Supervised Domain Adaptation for Recognizing Retinal Diseases from Wide-Field Fundus Images
abstract
This paper addresses the emerging task of recognizing multiple retinal diseases from wide-field (WF) and ultra-wide-field (UWF) fundus images. For an effective use of existing large amount of labeled color fundus photo (CFP) data and the relatively small amount of WF and UWF data, we propose a supervised domain adaptation method named Cross-domain Collaborative Learning (CdCL). Inspired by the success of fixed-ratio based mixup in unsupervised domain adaptation, we re-purpose this strategy for the current task. Due to the intrinsic disparity between the field-of-view of CFP and WF/UWF images, a scale bias naturally exists in a mixup sample that the anatomic structure from a CFP image will be considerably larger than its WF/UWF counterpart. The CdCL method resolves the issue by Scale-bias Correction, which employs Transformers for producing scale-invariant features. As demonstrated by extensive experiments on multiple datasets covering both WF and UWF images, the proposed method compares favorably against a number of competitive baselines.
Qijie Wei, Jingyuan Yang 0004, Bo Wang 0011, Jinrui Wang, Jianchun Zhao, Niranchana Manivannan, Youxin Chen, Dayong Ding, Jing Zhou 0005, Xirong Li 0001
BIBM9
2023 A Unified Pyramid Recurrent Network for Video Frame Interpolation
abstract
Flow-guided synthesis provides a common framework for frame interpolation, where optical flow is estimated to guide the synthesis of intermediate frames between consecutive inputs. In this paper, we present UPR-Net, a novel Unified Pyramid Recurrent Network for frame interpolation. Cast in a flexible pyramid framework, UPR-Net exploits lightweight recurrent modules for both bi-directional flow estimation and intermediate frame synthesis. At each pyramid level, it leverages estimated bi-directional flow to generate forward-warped representations for frame synthesis; across pyramid levels, it enables iterative refinement for both optical flow and intermediate frame. In particular, we show that our iterative synthesis strategy can significantly improve the robustness of frame interpolation on large motion cases. Despite being extremely lightweight (1.7M parameters), our base version of UPR-Net achieves excellent performance on a large range of benchmarks. Code and trained models of our UPR-Net series are available at: https://github.com/srcn-iv1/UPR-Net.
Xin Jin 0023, Longhai Wu, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm
CVPR4
2023 Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation
abstract
Effectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a mixed way or devise separate modules for each type of information, which lead to representation ambiguity and low efficiency. In this paper, we propose a new module to explicitly extract motion and appearance information via a unified operation. Specifically, we rethink the information process in inter-frame attention and reuse its attention map for both appearance feature enhancement and motion information extraction. Furthermore, for efficient VFI, our proposed module could be seamlessly integrated into a hybrid CNN and Transformer architecture. This hybrid pipeline can alleviate the computational complexity of inter-frame attention as well as preserve detailed low-level structure information. Experimental results demonstrate that, for both fixed- and arbitrary-timestep interpolation, our method achieves state-of-the-art performance on various datasets. Meanwhile, our approach enjoys a lighter computation overhead over models with close performance. The source code and models are available at https://github.com/MCG-NJU/EMA-VFI.
Youxin Chen, Gangshan Wu, Limin Wang 0002
CVPR4
2023 Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective
abstract
Zijian Zhang, Chang Shu, Ya Xiao, Yuan Shen, Di Zhu, Youxin Chen, Jing Xiao, Jey Han Lau, Qian Zhang, Zheng Lu. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Ya Xiao 0006, Youxin Chen, Jing Xiao 0006, Jey Han Lau, Qian Zhang 0018, Zheng Lu 0002
EACL6
2023 Enhanced Bi-directional Motion Estimation for Video Frame Interpolation
abstract
We propose a simple yet effective algorithm for motion-based video frame interpolation. Existing motion-based interpolation methods typically rely on an off-the-shelf optical flow model or a U-Net based pyramid network for motion estimation, which either suffer from large model size or limited capacity in handling various challenging motion cases. In this work, we present a novel compact model to simultaneously estimate the bi-directional motions between input frames. It is designed by carefully adapting the ingredients (e.g., warping, correlation) in optical flow research for simultaneous bi-directional motion estimation within a flexible pyramid recurrent framework. Our motion estimator is extremely lightweight (15x smaller than PWC-Net), yet enables reliable handling of large and complex motion cases. Based on estimated bi-directional motions, we employ a synthesis network to fuse forward-warped representations and predict the intermediate frame. Our method achieves excellent performance on a broad range of frame interpolation benchmarks. Code and trained models are available at https://github.com/srcn-ivl/EBME.
Xin Jin 0023, Longhai Wu, Guotao Shen, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm
WACV4
2023 APP-Net: Auxiliary-Point-Based Push and Pull Operations for Efficient Point Cloud Recognition
abstract
Aggregating neighbor features is essential for point cloud neural network. In the existing work, each point in the cloud may inevitably be selected as the neighbors of multiple aggregation centers, as all centers will gather neighbor features from the whole point cloud independently. Thus, each point has to participate in the calculation repeatedly, generating redundant duplicates in the memory, leading to intensive computation costs and memory consumption. Meanwhile, to pursue higher accuracy, previous methods often rely on a complex local aggregator to extract fine geometric representation, further slowing down the processing pipeline. To address these issues, we propose a new local aggregator of linear complexity for point cloud analysis, coined as APP. Specifically, we introduce an auxiliary container as an anchor to exchange features between the source point and the aggregating center. Each source point pushes its feature to only one auxiliary container, and each center point pulls features from only one auxiliary container. This avoids the re-computation issue of each source point. To facilitate the learning of the local structure of point cloud, we use an online normal estimation module to provide explainable geometric information to enhance our APP modeling capability. Our built network is more efficient than all the previous baselines with a clear margin while still consuming a lower memory. Experiments on classification and semantic segmentation demonstrate that APP-Net reaches comparable accuracies to other networks. In the classification task, it can process more than 10,000 samples per second with less than 10GB of memory on a single GPU. We will release the code at https://github.com/MCG-NJU/ APP-Net.
Tao Lu 0005, Chunxu Liu, Youxin Chen, Gangshan Wu, Limin Wang 0002
IEEE Trans. Image Process.3
2022 ICAF: Iterative Contrastive Alignment Framework for Multimodal Abstractive Summarization
abstract
Integrating multimodal knowledge for abstractive summarization task is a work-in-progress research area, with present techniques inheriting fusion-then-generation paradigm. Due to semantic gaps between computer vision and natural language processing, current methods often treat multiple data points as separate objects and rely on attention mechanisms to search for connection in order to fuse together. In addition, missing awareness of cross-modal matching from many frameworks leads to performance reduction. To solve these two drawbacks, we propose an Iterative Contrastive Alignment Framework (ICAF) that uses recurrent alignment and contrast to capture the coherences between images and texts. Specifically, we design a recurrent alignment (RA) layer to gradually investigate fine-grained semantical relationships between image patches and text tokens. At each step during the encoding process, crossmodal contrastive losses are applied to directly optimize the embedding space. According to ROUGE, relevance scores, and human evaluation, our model outperforms the state-of-the-art baselines on MSMO dataset. Experiments on the applicability of our proposed framework and hyperparameters settings have been also conducted.
Youxin Chen, Jing Xiao 0006, Qian Zhang 0018, Zheng Lu 0002
IJCNN3
2022 Lesion Localization in OCT by Semi-Supervised Object Detection
abstract
Over 300 million people worldwide are affected by various retinal diseases. By noninvasive Optical Coherence Tomography (OCT) scans, a number of abnormal structural changes in the retina, namely retinal lesions, can be identified. Automated lesion localization in OCT is thus important for detecting retinal diseases at their early stage. To conquer the lack of manual annotation for deep supervised learning, this paper presents a first study on utilizing semi-supervised object detection (SSOD) for lesion localization in OCT images. To that end, we develop a taxonomy to provide a unified and structured viewpoint of the current SSOD methods, and consequently identify key modules in these methods. To evaluate the influence of these modules in the new task, we build OCT-SS, a new dataset consisting of over 1k expert-labeled OCT B-scan images and over 13k unlabeled B-scans. Extensive experiments on OCT-SS identify Unbiased Teacher (UnT) as the best current SSOD method for lesion localization. Moreover, we improve over this strong baseline, with mAP increased from 49.34 to 50.86.
Jianchun Zhao, Jingyuan Yang 0004, Weihong Yu, Youxin Chen, Xirong Li 0001
ICMR6
2022 Learning Two-Stream CNN for Multi-Modal Age-Related Macular Degeneration Categorization
abstract
This paper tackles automated categorization of Age-related Macular Degeneration (AMD), a common macular disease among people over 50. Previous research efforts mainly focus on AMD categorization with a single-modal input, let it be a color fundus photograph (CFP) or an OCT B-scan image. By contrast, we consider AMD categorization given a multi-modal input, a direction that is clinically meaningful yet mostly unexplored. Contrary to the prior art that takes a traditional approach of feature extraction plus classifier training that cannot be jointly optimized, we opt for end-to-end multi-modal Convolutional Neural Networks (MM-CNN). Our MM-CNN is instantiated by a two-stream CNN, with spatially-invariant fusion to combine information from the CFP and OCT streams. In order to visually interpret the contribution of the individual modalities to the final prediction, we extend the class activation mapping (CAM) technique to the multi-modal scenario. For effective training of MM-CNN, we develop two data augmentation methods. One is GAN-based CFP/OCT image synthesis, with our novel use of CAMs as conditional input of a high-resolution image-to-image translation GAN. The other method is Loose Pairing, which pairs a CFP image and an OCT image on the basis of their classes instead of eye identities. Experiments on a clinical dataset consisting of 1,094 CFP images and 1,289 OCT images acquired from 1,093 distinct eyes show that the proposed solution obtains better F1 and Accuracy than multiple baselines for multi-modal AMD categorization. Code and data are available at https://github.com/li-xirong/mmc-amd.
Weisen Wang, Xirong Li 0001, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Dayong Ding, Youxin Chen
IEEE J. Biomed. Health Informatics7
2021 Multi-Modal Multi-Instance Learning for Retinal Disease Recognition
abstract
This paper attacks an emerging challenge of multi-modal retinal disease recognition. Given a multi-modal case consisting of a color fundus photo (CFP) and an array of OCT B-scan images acquired during an eye examination, we aim to build a deep neural network that recognizes multiple vision-threatening diseases for the given case. As the diagnostic efficacy of CFP and OCT is disease-dependent, the network's ability of being both selective and interpretable is important. Moreover, as both data acquisition and manual labeling are extremely expensive in the medical domain, the network has to be relatively lightweight for learning from a limited set of labeled multi-modal samples. Prior art on retinal disease recognition focuses either on a single disease or on a single modality, leaving multi-modal fusion largely underexplored. We propose in this paper Multi-Modal Multi-Instance Learning (MM-MIL) for selectively fusing CFP and OCT modalities. Its lightweight architecture (as compared to current multi-head attention modules) makes it suited for learning from relatively small-sized datasets. For an effective use of MM-MIL, we propose to generate a pseudo sequence of CFPs by over sampling a given CFP. The benefits of this tactic include well balancing instances across modalities, increasing the resolution of the CFP input, and finding out regions of the CFP most relevant with respect to the final diagnosis. Extensive experiments on a real-world dataset consisting of 1,206 multi-modal cases from 1,193 eyes of 836 subjects demonstrate the viability of the proposed model.
Xirong Li 0001, Hailan Lin, Jianchun Zhao, Dayong Ding, Weihong Yu, Youxin Chen
ACM Multimedia8
2020 Learn to Segment Retinal Lesions and Beyond
abstract
Towards automated retinal screening, this paper makes an endeavor to simultaneously achieve pixel-level retinal lesion segmentation and image-level disease classification. Such a multi-task approach is crucial for accurate and clinically interpretable disease diagnosis. Prior art is insufficient due to three challenges, i.e., lesions lacking objective boundaries, clinical importance of lesions irrelevant to their size, and the lack of one-to-one correspondence between lesion and disease classes. This paper attacks the three challenges in the context of diabetic retinopathy (DR) grading. We propose Lesion-Net, a new variant of fully convolutional networks, with its expansive path redesigned to tackle the first challenge. A dual Dice loss that leverages both semantic segmentation and image classification losses is introduced to resolve the second challenge. Lastly, we build a multi-task network that employs Lesion-Net as a side-attention branch for both DR grading and result interpretation. A set of 12K fundus images is manually segmented by 45 ophthalmologists for 8 DR-related lesions, resulting in 290K manual segments in total. Extensive experiments on this large-scale dataset show that our proposed approach surpasses the prior art for multiple tasks including lesion segmentation, lesion classification and DR grading.
Qijie Wei, Xirong Li 0001, Weihong Yu, Yongpeng Zhang, Bojie Hu, Bin Mo, Di Gong, Dayong Ding, Youxin Chen
ICPR11
2019 Two-Stream CNN with Loose Pair Training for Multi-modal AMD Categorization
Weisen Wang, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Jingyuan Yang 0004, Zhikun Yang, Dayong Ding, Youxin Chen, Xirong Li 0001
MICCAI (1)10
2018 Laser Scar Detection in Fundus Images Using Convolutional Neural Networks
Qijie Wei, Xirong Li 0001, Dayong Ding, Weihong Yu, Youxin Chen
ACCV (4)6
2014 An MQDF-CNN Hybrid Model for Offline Handwritten Chinese Character Recognition
abstract
An MQDF-CNN hybrid model is presented for offline handwritten Chinese character recognition. The main idea behind MQDF-CNN hybrid model is that the significant difference on features and classification mechanisms between MQDF and CNN can complement each other. Linear confidence accumulation and multiplication confidence criteria are used for fusion outputs of MQDF and CNN. Experiments have been conducted on CASIA-HWDB1.1 and ICDAR2013 offline handwritten Chinese character recognition competition dataset. On both datasets, CNN beats MQDF by more than 1% of the accuracy, and the MQDF-CNN hybrid model has achieved the test accuracies of 92.03% and 94.44% respectively. The result on competition dataset is comparable to the state-of-the-art result though less training samples and only one CNN is used.
Xin Li 0144, Changsong Liu, Xiaoqing Ding, Youxin Chen
ICFHR5
2014 Enhanced Non-linear Features for On-line Handwriting Recognition Using Deep Learning
Minhua Wu, Zhenbo Luo, Youxin Chen
ICONIP (1)4