Jiawen Lin

dblp:199/0668 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation
abstract
3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in simple, single-object scenes, they suffer from severe performance degradation in complex, multi-object scenes that are common in real-world settings, hindering practical deployment. Existing methods face two key challenges in complex, multi-object scenes: inadequate parsing of implicit localization cues critical for disambiguating visually similar objects, and ineffective suppression of dynamic spatial interference from co-occurring objects, resulting in degraded grounding accuracy. To address these challenges, we propose PC-CrossDiff, a unified dual-task framework with a dual-level cross-modal differential attention architecture for 3DREC and 3DRES. Specifically, the framework introduces: (i) Point-Level Differential Attention (PLDA) modules that apply bidirectional differential attention between text and point clouds, adaptively extracting implicit localization cues via learnable weights to improve discriminative representation; (ii) Cluster-Level Differential Attention (CLDA) modules that establish a hierarchical attention mechanism to adaptively enhance localization-relevant spatial relationships while suppressing ambiguous or irrelevant spatial relations through a localization-aware differential attention block. To address the scale disparity and conflicting gradients in joint 3DREC–3DRES training, we propose L_DGTL, a unified loss function that explicitly reduces multi-task crosstalk and enables effective parameter sharing across tasks. Our method achieves state-of-the-art performance on the ScanRefer, NR3D, and SR3D benchmarks. Notably, on the Implicit subsets of ScanRefer, it improves the [email protected] score by +10.16% for the 3DREC task, highlighting its strong ability to parse implicit spatial cues.
Wenbin Tan 0001, Jiawen Lin, Fangyong Wang, Yuan Xie 0006, Yachao Zhang 0001, Yanyun Qu
AAAI2
2026 Instructing visual feature modeling with semantic guidance for 3D visual grounding
Yachao Zhang 0001, Shiran Bian, Jiahao Li 0003, Jiawen Lin, Fangyong Wang, Yuan Xie 0006, Yanyun Qu
Pattern Recognit.4
2025 Active Learning for Meibomian Gland Segmentation in Infrared Meibography Images
abstract
Meibomian Gland Dysfunction (MGD) is a key contributor to clinical Dry Eye Disease (DED), and accurate meibomian gland segmentation is essential for its intelligent diagnosis. However, the high cost of annotation remains a major barrier to improving segmentation quality. Active Learning (AL), which selects the most informative samples for annotation, offers an efficient solution. This study explores the application of AL to infrared meibography image segmentation and introduces a novel progressive hybrid sampling framework. Initial AL stages often rely on random sampling for initial labeled sets, risking unstable performance from uninformative data and cold starts. Additionally, meibography images always have specular reflections whose late introduction in training degrades future segmentation accuracy. To address these challenges, we incorporate prior knowledge of specular reflections into the initial labeled set construction, enabling the model to handle such artifacts from the beginning. We then implement a two-stage dynamic sampling strategy: entropy-based uncertainty sampling is used in the early iterations to maximize annotated data informativeness and rapidly boost model performance. However, as AL progresses, only focusing on uncertainty may overlook global data distribution and lead to redundant annotations. Hence, an adaptive threshold is introduced to monitor sample redundancy. When redundancy exceeds this threshold, a diversity sampling module is activated to improve model generalization. Experimental results on the public MGD-1K dataset show that our method achieves superior segmentation performance under the same annotation budget. Unlike mainstream methods that often fail to outperform random baselines, our approach consistently delivers accurate and efficient meibomian gland segmentation.
Kunfeng Lai, Dongqi Li, Ryan G. Benton, Glen M. Borchert, Jiawen Lin, Jingshan Huang
BIBM6
2025 Structure-Aware Unsupervised Enhancement of Low-Light Fundus Images
abstract
In real-world clinical settings, fundus images often suffer from quality degradation caused by uncontrolled imaging conditions, with underexposure being especially common. Such low-light images degrade visual quality and obscure critical retinal structures, hindering diagnosis and automated analysis. To address this issue, we propose a structure-aware unsupervised GAN-based enhancement method for low-light fundus images. The generator employs an attention mechanism that fuses estimated illumination with high-frequency components to enhance brightness while preserving fine anatomical details. A dualdiscriminator framework ensures both global consistency and local detail fidelity: the multi-scale global discriminator enforces structural coherence, and the local discriminator refines regional illumination to prevent over- or underexposure. Quantitative and qualitative results on the EyeQ dataset, along with downstream tasks such as vessel segmentation and diabetic retinopathy (DR) grading, demonstrate that our method significantly improves brightness and contrast while maintaining retinal structure, supporting reliable clinical and computational analysis.
Jiaqi Zheng 0018, Dongqi Li, Glen M. Borchert, Jiawen Lin, Jingshan Huang
BIBM4
2025 Time-Frequency Feature Enhancement Method for Moving Multiple Sound Source Localization in Noisy Environments
Qian Weng, Chenjie Zeng, Jiawen Lin, Yuanxun Kang
ICIC (17)4
2025 SDFCNet: A Spatial-Domain and Frequency-Domain Collaborative Network for Building Extraction in High-Resolution Remote Sensing Images
abstract
To address the low accuracy in building boundary extraction from high-resolution remote sensing images, this paper proposes a spatial-frequency collaborative building extraction network named SDFCNet, which integrates boundary information from both the frequency domain and the spatial domain. The High-Resolution Feature Processing Module (HRFPM) is designed to fuse multi-scale features within the encoder, which compensates for detail loss due to downsampling and provides richer edge information for subsequent extraction. Additionally, a High-Frequency Signal Processing Module (HFSMP) extracts edge features from high-resolution characteristics and fuses them with high-frequency signals from the original image, which enhances the precision and completeness of boundary extraction by constraining the decoder’s feature boundaries. Finally, a deep supervision strategy is introduced to provide auxiliary supervision for both high-frequency signals and decoder outputs. The experimental results on the INRIA and WHU datasets demonstrate that this approach outperforms mainstream building extraction networks in multiple evaluation metrics, offering enhanced accuracy and completeness in building boundary extraction.
Xiansheng Huang, Jiawen Lin, Cairen Jian, Qian Weng
ICIP2
2025 Semi-Supervised Infrared Meibomian Gland Segmentation with Intra-Patient Registration and Feature Supervision
abstract
Low-cost and high-precision infrared meibomian gland segmentation is an important basis for early diagnosis and monitoring of many ocular diseases in ophthalmic clinical practice. To address the issue of limited labeled data, we propose a novel semi-supervised meibomian gland segmentation approach. By leveraging the prior knowledge of the patient each image belongs to, intra-patient registration is taken to generate diverse and lifelike pseudo-labeled data. Contrastive learning strategy with reliable negative sample filtering is also introduced to resolve the insufficient supervision in the feature space. Experimental results on the private dataset demonstrate the success of the proposed approach, exhibiting superiority over the state-of-the-art methods. The Code can be found at SS-MGS.
Yushun Huang, Kunfeng Lai, Taichen Lai, Jiawen Lin
ICIP4
2025 SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
abstract
3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for real-world applications since eliminating scene-specific training requirements. However, existing zero-shot methods face challenges of spatial-limited reasoning due to reliance on single-view localization, and contextual omissions or detail degradation. To address these issues, we propose SeqVLM, a novel zero-shot 3DVG framework that leverages multi-view real-world scene images with spatial information for target object reasoning. Specifically, SeqVLM first generates 3D instance proposals via a 3D semantic segmentation network and refines them through semantic filtering, retaining only semantic-relevant candidates. A proposal-guided multi-view projection strategy then projects these candidate proposals onto real scene image sequences, preserving spatial relationships and contextual details in the conversion process of 3D point cloud to images. Furthermore, to mitigate VLM computational overload, we implement a dynamic scheduling mechanism that iteratively processes sequances-query prompts, leveraging VLM's cross-modal reasoning capabilities to identify textually specified objects. Experiments on the ScanRefer and Nr3D benchmarks demonstrate state-of-the-art performance, achieving [email protected] scores of 55.6% and 53.2%, surpassing previous zero-shot methods by 4.0% and 5.2%, respectively, which advance 3DVG toward greater generalization and real-world applicability.
Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan 0001, Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu
ACM Multimedia1
2025 A review of deep learning for fundus image enhancement
abstract
Fundus images are capable of accurately depicting the fundus structure of the examinee, providing crucial diagnostic evidence for systemic diseases like hypertension, enabling early detection, and facilitating timely preventive interventions. Nevertheless, fundus images often exhibit varying degrees of quality defects during acquisition. These imperfections not only undermine the accuracy of manual diagnosis but also pose significant challenges to automated analysis. Recently, deep learning-based image enhancement techniques have witnessed remarkable progress in addressing the issue of image quality defects. However, the direct application of these techniques to low-quality fundus images is still fraught with limitations. This article presents a comprehensive and systematic summary of deep learning techniques for fundus image enhancement, including a brief description and analysis of existing state-of-the-art approaches, datasets, and evaluation metrics. Comparisons between such methods are also illustrated. Finally, challenges and future research directions are discussed.
Jiawen Lin, Jiaqi Zheng 0018
Discov. Comput.1
2024 Artificial intelligence assisted recognition and diagnosis of Magnetically controlled capsule gastroscopy
abstract
To enhance the diagnostic efficiency and accuracy of gastric mucosal lesions in magnetically controlled capsule endoscopy (MCCE) images, a Jigsaw Guided Deep Feature Fusion (JG-DFF) artificial intelligence (AI)-assisted recognition model based on the ResNet-50 convolutional neural network (CNN) is presented in this paper. Based on a dataset consisting of 4,053 MCCE images retrospectively collected from Sun Yat-sen Memorial Hospital, we constructed a ResNet-50 AI recognition model, as well as a JG-DFF AI recognition model under the jigsaw guided deep feature fusion. Additionally, the diagnostic ability and run time of the JG-DFF model were compared against a digestive endoscopist using the same data. The results indicate that the JG-DFF model has the potential for clinical application in improving diagnostic efficiency and reducing false positive rates. The overall accuracy of the JG-DFF artificial intelligence assisted recognition model established in this study is 92.1%. Compared to the ResNet-50 model, recognition by the JG-DFF model was more consistent; Compared with the digestive endoscopist, observed specificity of the JG-DFF model in identifying positive lesions was similar (P>0.05). That said, the JG-DFF model was more sensitive in identifying certain classes (P1
Susu Chen, Chuyu Wei, Dongqi Li, Glen M. Borchert, Jiawen Lin, Jingshan Huang
BIBM6
2024 Scribbled-Supervised Meibomian Gland Segmentation via Perturbation and Conflict in Dual-Branch Network
Lingjie Lin, Kunfeng Lai, Yushun Huang, Jiawen Lin
PRCV (14)5
2024 BFRNet: Bimodal Fusion and Rectification Network for Remote Sensing Semantic Segmentation
Qian Weng, Yifeng Lin, Zengying Pan, Jiawen Lin, Gengwei Chen
PRCV (13)4
2023 Semi-supervised meibomian gland segmentation via mutual consistency constraints and uncertainty rectification
abstract
To solve the label scarcity of meibomian gland segmentation in infrared meibography images, a novel framework for semi-supervised meibomian gland segmentation is firstly presented in this paper. Extra mutual feature consistency constraint is added along with the cross pseudo supervision , guiding the model more robustness and discriminative. Meanwhile, cross uncertainty rectification is introduced to avoid noisy labels, further improving the pseudo supervision. Experimental results on an internal dataset reveals that our method yields significant performances using only 10% of the labeled data compared to the fully supervised segmentation, and outperforms the state-of-the art semi-supervised segmentation methods. Combination of mutual consistency regularization and cross uncertainty rectifi-cation guides model to distinguish glands from background well with limited labeled data.
Jiawen Lin, Lingjie Lin, Dongqi Li, Glen M. Borchert, Jingshan Huang
BIBM1
2020 Multi-attention based cross-domain beauty product image retrieval
Zhihui Wang 0001, Xing Liu 0005, Jiawen Lin, Caifei Yang
Sci. China Inf. Sci.3
2020 Retinal image quality assessment for diabetic retinopathy screening: A survey
Jiawen Lin, Lun Yu, Qian Weng, Xianghan Zheng
Multim. Tools Appl.1
2019 A saliency and Gaussian net model for retinal vessel segmentation
abstract
Retinal vessel segmentation is a significant problem in the analysis of fundus images. A novel deep learning structure called the Gaussian net (GNET) model combined with a saliency model is proposed for retinal vessel segmentation. A saliency image is used as the input of the GNET model replacing the original image. The GNET model adopts a bilaterally symmetrical structure. In the left structure, the first layer is upsampling and the other layers are max-pooling. In the right structure, the final layer is max-pooling and the other layers are upsampling. The proposed approach is evaluated using the DRIVE database. Experimental results indicate that the GNET model can obtain more precise features and subtle details than the UNET models. The proposed algorithm performs well in extracting vessel networks, and is more accurate than other deep learning methods. Retinal vessel segmentation can help extract vessel change characteristics and provide a basis for screening the cerebrovascular diseases.
Lanyan Xue, Jiawen Lin, Xinrong Cao, Shaohua Zheng, Lun Yu
Frontiers Inf. Technol. Electron. Eng.2
2017 Land-Use Classification via Extreme Learning Classifier Based on Deep Convolutional Features
abstract
One of the challenging issues in high-resolution remote sensing images is classifying land-use scenes with high quality and accuracy. An effective feature extractor and classifier can boost classification accuracy in scene classification. This letter proposes a deep-learning-based classification method, which combines convolutional neural networks (CNNs) and extreme learning machine (ELM) to improve classification performance. A pretrained CNN is initially used to learn deep and robust features. However, the generalization ability is finite and suboptimal, because the traditional CNN adopts fully connected layers as classifier. We use an ELM classifier with the CNN-learned features instead of the fully connected layers of CNN to obtain excellent results. The effectiveness of the proposed method is tested on the UC-Merced data set that has 2100 remotely sensed land-use-scene images with 21 categories. Experimental results show that the proposed CNN-ELM classification method achieves satisfactory results.
Qian Weng, Zhengyuan Mao, Jiawen Lin, Wenzhong Guo
IEEE Geosci. Remote. Sens. Lett.3
2016 Accumulative Energy-Based Seam Carving for Image Resizing
abstract
With the diversified development of the digital devices, such as computer, mobile phone, pad and television, how to resize an image or video to adapt to different display screens has been attracting more and more peoples' attention. Seam carving has been an important method for image resizing. If multiple removed or inserted seams are located within a certain region, it can lead to discontinuity image content. Besides, the salient objects tend to be destroyed if the energy function only contains the gradient information. Therefore, we propose an accumulative energy-based seam carving method for image resizing. When removing a certain seam, we distribute the energy of each pixel on the seam to its adjacent 8-connected pixels in order to avoid the extreme concentration of seams, especially within a texture region. In addition, we add the image saliency and the edge information into the energy function to reduce the distortion. Since the computational complexity of seam carving method is very high, we use parallel computing environment to achieve efficient computation. Experimental results show that compared with the existing methods, our method can both avoid the discontinuity of image content and distortions as well as better maintain the shape of the salient objects.
Yuzhen Niu, Jiawen Lin, Haifeng Zhang 0012
PDCAT3