Junyan Lin

dblp:318/8004 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
0009-0002-4028-2383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Real-Time Robotic Diffusion Policy Accelerator Exploiting Self- and Cross-Guided Modal Similarity
abstract
Diffusion Policy (DP) has demonstrated strong potential in robotic visuomotor control, offering robust generalization and seamless integration of multi-modal data. However, its complex model structure and increasing multi-modal inputs have brought latency and power challenges for edge resource-constrained robotic platforms. To address the above challenges, we identify the potential intra-model and inter-model redundancies in DP. We observe that DP relies on frequent multi-modal inputs such as images and text during execution. However, the demands of fine-grained robotic manipulation result in substantial intra-modal similarity across consecutive image frames, which, combined with inter-modal semantic redundancy between images and language, indicates that much of the input information is repetitive and potentially compressible. Yet prior works have not exploited these characteristics for targeted optimization. We therefore propose a hardware–software co-design accelerator. On the algorithmic side, we introduce self- and cross-guided modal compression, leveraging intra- and inter-modality similarity to reduce redundant computation within the key DP modules. On the hardware side, we design a tailored architecture that supports multiple operators with optimized sparse memory access, lightweight computation engines, and reconfigurable on-chip dataflow, substantially reducing energy cost. Experimental results demonstrate a 26× speedup over a high-performance GPU while consuming only 1.5 W, enabling low-power and real-time robotic control on edge robotic devices.
Boju Chen, Xiaoyu Feng, Junyan Lin, Huazhong Yang, Yongpan Liu
DATE3
2026 Efficient Token Compression for the Understanding and Generation Unified MLLMs
Junyan Lin, Jinming Liu 0001, Shengyang Zhao, Xin Jin 0014
ISCAS1
2026 Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
abstract
Classical visual coding and Multimodal Large Language Model (MLLM) token technology share the core objective - maximizing information fidelity while minimizing computational cost. Therefore, this paper reexamines MLLM token technology, including tokenization, token compression, and token reasoning, through the established principles of long-developed visual coding area. From this perspective, we (1) establish a unified formulation bridging token technology and visual coding, enabling a systematic, module-by-module comparative analysis; (2) synthesize bidirectional insights, exploring how visual coding principles can enhance MLLM token techniques' efficiency and robustness, and conversely, how token technology paradigms can inform the design of next-generation semantic visual codecs; (3) prospect for promising future research directions and critical unsolved challenges. In summary, this study presents the first comprehensive and structured technology comparison of MLLM token and visual coding, paving the way for more efficient multimodal models and more powerful visual codecs simultaneously.
Jinming Liu 0001, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Zhibo Chen 0001, Huan Wang 0014, Xin Jin 0014
ISCAS2
2026 An efficient constrained multi-objective evolutionary algorithm with a spatial discretization evaluation mechanism for unmanned aerial vehicle path planning
Zhiyuan Cai, Chaoda Peng, Junyan Lin, Yueting Xu, Haoyu Luo
Eng. Appl. Artif. Intell.4
2025 Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices
abstract
Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM.
Junyan Lin, Yingqi Fan, Hui Su, Jinlan Fu
CVPR1
2025 Multimodal Language Models See Better When They Look Shallower
abstract
Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT).This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis.While prior studies suggest that different ViT layers capture different types of information-shallower layers focusing on fine visual details and deeper layers aligning more closely with textual semantics, the impact of this variation on MLLM performance remains underexplored.We present the first comprehensive study of visual layer selection for MLLMs, analyzing representation similarity across ViT layers to establish shallow, middle, and deep layer groupings.Through extensive evaluation of MLLMs (1.4B-7B parameters) across 10 benchmarks encompassing 60+ tasks, we find that while deep layers excel in semantic-rich tasks like OCR, shallow and middle layers significantly outperform them on fine-grained visual tasks including counting, positioning, and object localization.Building on these insights, we propose a lightweight feature fusion method that strategically incorporates shallower layers, achieving consistent improvements over both single-layer and specialized fusion baselines.Our work offers the first principled study of visual layer selection in MLLMs, showing that MLLMs can often see better when they look shallower.
Junyan Lin, Xinghao Chen 0009, Jianfeng Dong, Xin Jin 0014, Hui Su, Jinlan Fu, Xiaoyu Shen 0001
EMNLP2
2025 Dynamic Cross-Modal Feature Interaction Network for Hyperspectral and LiDAR Data Classification
abstract
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data joint classification is a challenging task. Existing multisource remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on expert knowledge. To address these limitations, we propose a novel dynamic cross-modal feature interaction network (DCMNet), the first framework leveraging a dynamic routing mechanism for HSI and LiDAR classification. Specifically, our approach introduces three feature interaction blocks: bilinear spatial attention block (BSAB), bilinear channel attention block (BCAB), and integration convolutional block (ICB). These blocks are designed to effectively enhance spatial, spectral, and discriminative feature interactions. A multilayer routing space with routing gates is designed to determine optimal computational paths, enabling data-dependent feature fusion. Additionally, bilinear attention mechanisms are employed to enhance feature interactions in spatial and channel representations. Extensive experiments on three public HSI and LiDAR datasets demonstrate the superiority of DCMNet over the state-of-the-art methods. Our codes are available athttps://github.com/oucailab/DCMNet.
Junyan Lin, Feng Gao 0005, Lin Qi 0004, Junyu Dong, Qian Du 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimodal Large Language Models
abstract
In recent years, multimodal large language models (MLLMs) have garnered significant attention from both industry and academia.However, there is still considerable debate on constructing MLLM architectures, particularly regarding the selection of appropriate connectors for perception tasks of varying granularities.This paper systematically investigates the impact of connectors on MLLM performance.Specifically, we classify connectors into feature-preserving and featurecompressing types.Utilizing a unified classification standard, we categorize sub-tasks from three comprehensive benchmarks, MM-Bench, MME, and SEED-Bench, into three task types: coarse-grained perception, fine-grained perception, and reasoning, and evaluate the performance.Our findings reveal that featurepreserving connectors excel in fine-grained perception tasks due to their ability to retain detailed visual information.In contrast, featurecompressing connectors, while less effective in fine-grained perception tasks, offer significant speed advantages and perform comparably in coarse-grained perception and reasoning tasks.These insights are crucial for guiding MLLM architecture design and advancing the optimization of MLLM architectures.
Junyan Lin, Xiaoyu Shen 0001
EMNLP1
2024 Sparse Focus Network for Multi-Source Remote Sensing Data Classification
abstract
Multi-source remote sensing data classification has emerged as a prominent research topic with the advancement of various sensors. Existing multi-source data classification methods are susceptible to irrelevant information interference during multi-source feature extraction and fusion. To solve this issue, we propose a sparse focus network for multi-source data classification. Sparse attention is employed in Transformer block for HSI and SAR/LiDAR feature extraction, thereby the most useful self-attention values are maintained for better feature aggregation. Furthermore, cross-attention is used to enhance multi-source feature interactions, and further improves the efficiency of cross-modal feature fusion. Experimental results on the Berlin and Houston2018 datasets highlight the effectiveness of SF-Net, outperforming existing state-of-the-art methods.
Xuepeng Jin, Junyan Lin, Feng Gao 0005, Lin Qi 0004
IGARSS2
2024 Boosting Spatial-Spectral Masked Auto-Encoder Through Mining Redundant Spectra for HSI-SAR/LiDAR Classification
abstract
Although recent masked image modeling (MIM)-based HSI-LiDAR/SAR classification methods have gradually recognized the importance of the spectral information, they have not adequately addressed the redundancy among different spectra, resulting in information leakage during the pretraining stage. This issue directly impairs the representation ability of the model. To tackle the problem, we propose a new strategy, named Mining Redundant Spectra (MRS). Unlike randomly masking spectral bands, MRS selectively masks them by similarity to increase the reconstruction difficulty. Specifically, a random spectral band is chosen during pretraining, and the selected and highly similar bands are masked. Experimental results demonstrate that employing the MRS strategy during the pretraining stage effectively improves the accuracy of existing MIM-based methods on the Berlin and Houston 2018 datasets.
Junyan Lin, Xuepeng Jin, Feng Gao 0005, Junyu Dong, Hui Yu 0001
IGARSS1
2024 Hyperspectral Image Change Detection via Cross-Sample Slot Attention and Dual Gated Feed-Forward Network
Luyao Cheng, Junyan Lin, Feng Gao 0005
PRCV (13)2
2024 Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs
abstract
We present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability.
Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP3
2024 BO-SHAP-BLS: a novel machine learning framework for accurate forecasting of COVID-19 testing capabilities
Choujun Zhan, Lingfeng Miao, Junyan Lin, Minghao Tan, Kim Fung Tsang, Tianyong Hao, Hu Min, Xuejiao Zhao
Neural Comput. Appl.3
2023 Multi-Scale Transformer Network for Hyperspectral Image Denoising
abstract
Removing noise from hyperspectral images (HSIs) has been widely regarded as one of the most meaningful preprocessing tasks in remote sensing image interpretation. In this paper, we aim to extend the Transformer backbone to HSI denoising, and propose a Multi-scale Transformer Denoising Network (MTDNet). Specifically, we design a multi-head global attention module to alleviate the computational burden caused by self-attention. Furthermore, we propose a multi-scale feed-forward network in which three branches of multi-scale features are extracted through dilated convolution. It enriches the non-linear feature transformation in the Transformer block. Both the objective and subjective experiments on the ICVL dataset demonstrate the superiority of the proposed MTDNet over four closely related methods.
Shuai Hu, Yikun Hu 0002, Junyan Lin, Feng Gao 0005, Junyu Dong
IGARSS3
2023 Hyperspectral and SAR Image Classification via Recursive Feature Interactive Fusion Network
abstract
Most of existing mutli-source remote sensing data classification methods are based on convolutional neural networks. Recently, the emergence of Vision Transformer greatly challenges the dominance of CNN-based methods. The self-attention mechanism in Transformer and other dynamic networks imply that high-order feature interactions are beneficial to improve the feature representation and fusion. To explore the high-order feature interactions in multi-source image fusion, in this paper, we proposed a novel recursive feature interactive fusion network. It is composed of cross-shaped window self-attention encoder, and recursive feature interactive fusion. We use gated convolution recursively to mix multi-modal features and exploit their spatial relations. Experimental results on two datasets show that the proposed method achieves better performance than closely related methods.
Junyan Lin, Feng Gao 0005, Lin Qi 0004, Junyu Dong
IGARSS2
2023 Gated-Cross Aggregation Network for Hyperspectral and LiDAR Data Classification
abstract
Existing hyperspectral image (HSI) and LiDAR data joint classification methods commonly treat LiDAR data equally with HSI in the network. As a result, these methods may fail to effectively leverage the spectral features from HSI and elevation information from LiDAR. In this paper, we show that better cross-modal alignments can be achieved through an HSI encoder for jointly embedding elevation features from LiDAR during spectral feature encoding. To this end, we propose a Gated-Cross Aggregation Network (GCA-Net) to fully investigate the complementary clues hidden in multi-source data progressively. The beneficial spectral-elevation cues are then exploited by cross-attention feature fusion. Then, useful LiDAR features are integrated into the HSI features via an elevation gating module, which occurs at each stage of the network. Experimental results on the Houston 2013 dataset and Trento dataset reveal that the proposed GCA-Net achieves better performance than several closely related methods.
Xiaochen Shi, Junyan Lin, Yuan Rao 0001, Feng Gao 0005
IGARSS2
2023 Multiview Siamese Collaborative Network for Hyperspectral Image Unmixing
abstract
Existing hyperspectral image unmixing methods acquire information from single-view, and therefore can hardly make full use of diverse spectral information, and the learning of space and spectrum is often relatively independent and cannot be combined effectively. In order to solve the problem of insufficient feature representation caused by the single-view spectral information and the lack of close relationship between spatial and spectral learning, we introduce the idea of multiview data construction, which divides the spectral bands into different views for multiview learning. In addition, we proposed a spatial-spectral siamese network. Deep collaborative learning is used to construct an unmixing model by combining multiview representation and the siamese network. Experimental results on the Jasper Ridge and Urban datasets demonstrate the effectiveness of the proposed method.
Zimo Yang, Lin Qi 0004, Feng Gao 0005, Junyan Lin
IGARSS4
2023 SS-MAE: Spatial-Spectral Masked Autoencoder for Multisource Remote Sensing Image Classification
abstract
Masked image modeling (MIM) is a highly popular and effective self-supervised learning method for image understanding. Existing MIM-based methods mostly focus on spatial feature modeling, neglecting spectral feature modeling. Meanwhile, existing MIM-based methods use Transformer for feature extraction, some local or high-frequency information may get lost. To this end, we propose a spatial-spectral masked auto-encoder (SS-MAE) for HSI and LiDAR/SAR data joint classification. Specifically, SS-MAE consists of a spatial-wise branch and a spectral-wise branch. The spatial-wise branch masks random patches and reconstructs missing pixels, while the spectral-wise branch masks random spectral channels and reconstructs missing channels. Our SS-MAE fully exploits the spatial and spectral representations of the input data. Furthermore, to complement local features in the training stage, we add two lightweight CNNs for feature extraction. Both global and local features are taken into account for feature modeling. To demonstrate the effectiveness of the proposed SS-MAE, we conduct extensive experiments on three publicly available datasets. Extensive experiments on three multi-source datasets verify the superiority of our SS-MAE compared with several state-of-the-art baselines. The source codes are available at https://github.com/summitgao/SS-MAE.
Junyan Lin, Feng Gao 0005, Xiaochen Shi, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Application and Challenges of Blockchain in Heterogeneous Identity Trust
Zhaolei Zhang, Guishan Dong, Junyan Lin
BlockSys4