Junfeng Luo

dblp:71/10406 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0008-8315-123XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks
abstract
End-to-end backpropagation remains the dominant training paradigm in deep learning, yet it suffers from inherent drawbacks, including update locking, high GPU memory consumption, and limited biological plausibility. Supervised local learning alleviates these issues by dividing the network into multiple blocks and training each block independently with an auxiliary network. However, gradient isolation also weakens the influence of downstream representations on earlier blocks, often resulting in a clear accuracy gap to end-to-end training. We propose Momentum Auxiliary Network++ (MAN++), a scalable framework that improves supervised local learning via a lightweight parameter-space transfer between adjacent blocks. MAN++ employs the exponential moving average (EMA) of parameters from adjacent blocks to propagate contextual information across the network. To address feature mismatches arising from direct EMA parameter transfer, we introduce a learnable scaling bias, which compensates feature statistics mismatch and stabilizes the transfer. Extensive experiments on image classification, object detection, and semantic segmentation across multiple architectures illustrate that MAN++ achieves accuracy on par with end-to-end training while substantially reducing GPU memory usage. These results position MAN++ as a practical and effective alternative to conventional backpropagation, offering new insights into scalable supervised local learning for vision tasks.
Junhao Su, Hengyu Shi, Tianyang Han, Yurui Qiu, Junfeng Luo, Xiaoming Wei, Jialin Gao
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 InstructOCR: Instruction Boosting Scene Text Spotting
abstract
In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks.
Chen Duan, Qianyi Jiang, Pei Fu, Shengxi Li, Shan Guo, Junfeng Luo
AAAI8
2025 Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
abstract
Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visuallanguage alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQA-Mask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a trainingefficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten.
Tongkun Guan, Pei Fu, Chen Duan, Qianyi Jiang, Zhentao Guo, Shan Guo, Junfeng Luo, Wei Shen 0002, Xiaokang Yang 0001
CVPR8
2025 A Token-Level Text Image Foundation Model for Document Understanding
Tongkun Guan, Pei Fu, Zhengtao Guo, Wei Shen 0002, Tiezhu Yue, Chen Duan, Qianyi Jiang, Junfeng Luo, Xiaokang Yang 0001
ICCV11
2024 Text2Street: Controllable Text-to-Image Generation for Street Views
Songen Gu, Jinming Su, Yiting Duan, Xingyue Chen, Junfeng Luo
ICPR (6)5
2021 Rethinking BiSeNet for Real-Time Semantic Segmentation
abstract
BiSeNet [28], [27] has been proved to be a popular two-stream network for real-time segmentation. However, its principle of adding an extra path to encode spatial information is time-consuming, and the backbones borrowed from pretrained tasks, e.g., image classification, may be inefficient for image segmentation due to the deficiency of task-specific design. To handle these problems, we propose a novel and efficient structure named Short-Term Dense Concatenate network (STDC network) by removing structure redundancy. Specifically, we gradually reduce the dimension of feature maps and use the aggregation of them for image representation, which forms the basic module of STDC network. In the decoder, we propose a Detail Aggregation module by integrating the learning of spatial information into low-level layers in single-stream manner. Finally, the low-level features and deep features are fused to predict the final segmentation results. Extensive experiments on Cityscapes and CamVid dataset demonstrate the effectiveness of our method by achieving promising trade-off between segmentation accuracy and inference speed. On Cityscapes, we achieve 71.9% mIoU on the test set with a speed of 250.4 FPS on NVIDIA GTX 1080Ti, which is 45.2% faster than the latest methods, and achieve 76.8% mIoU with 97.0 FPS while inferring on higher resolution images. Code is available at https://github.com/MichaelFan01/STDC-Seg.
Mingyuan Fan 0002, Shenqi Lai, Junshi Huang, Xiaoming Wei, Zhenhua Chai, Junfeng Luo, Xiaolin Wei
CVPR6
2021 Structure Guided Lane Detection
abstract
Recently, lane detection has made great progress with the rapid development of deep neural networks and autonomous driving. However, there exist three mainly problems including characterizing lanes, modeling the structural relationship between scenes and lanes, and supporting more attributes (e.g., instance and type) of lanes. In this paper, we propose a novel structure guided framework to solve these problems simultaneously. In the framework, we first introduce a new lane representation to characterize each instance. Then a top-down vanishing point guided anchoring mechanism is proposed to produce intensive anchors, which efficiently capture various lanes. Next, multi-level structural constraints are used to improve the perception of lanes. In the process, pixel-level perception with binary segmentation is introduced to promote features around anchors and restore lane details from bottom up, a lane-level relation is put forward to model structures (i.e., parallel) around lanes, and an image-level attention is used to adaptively attend different regions of the image from the perspective of scenes. With the help of structural guidance, anchors are effectively classified and regressed to obtain precise locations and shapes. Extensive experiments on public benchmark datasets show that the proposed approach outperforms state-of-the-art methods with 117 FPS on a single GPU.
Jinming Su, Junfeng Luo, Xiaoming Wei, Xiaolin Wei
IJCAI4
2017 Effective Iris Recognition for Distant Images Using Log-Gabor Wavelet Based Contourlet Transform Features
Lasker Ershad Ali, Junfeng Luo, Jinwen Ma
ICIC (1)2
2015 Image segmentation with the competitive learning based MS model
abstract
In this paper, we propose a competitive learning approach to image segmentation by coupling the Mumford-Shah (MS) model and the Distance Sensitive Rival Penalized Competitive Learning (DSRPCL) mechanism, being denoted as the DBMS model. Actually, the DBMS model with the evolution of the level set function can get highly accurate segmentation of the image by automatically detecting the appropriate number of segmented regions and overcoming the problems of vacuum and overlap. It is demonstrated by experimental results on BSDS500 that our DBMS approach can obtain the state-of-the-art segmentation result under the evaluation of ODS index.
Junfeng Luo, Jinwen Ma
ICIP1