Di Sun 0001

dblp:60/11415-1 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-2793-7066ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 DSTFuse: Enhancing Degraded Images via Style Transfer for Visible and Infrared Image Fusion
Gang Pan 0002, Zhijie Sui, Yonglu Liu, Di Sun 0001
ICIC (17)4
2026 Lightweight Image Forgery Localization via Diffusion-Prior-Guided Dynamic Convolution Attention
Lei Zhou 0031, Gang Pan 0002, Di Sun 0001
ICIC (19)3
2026 Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
abstract
Grounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions. GMNER task exhibits two challenging attributes: 1) The tenuous correlation between images and text on social media contributes to a notable proportion of named entities being ungroundable. 2) There exists a distinction between coarse-grained noun phrases used in similar tasks (e.g., phrase localization) and fine-grained named entities. In this paper, we propose RiVEG, a unified framework that reformulates GMNERinto a joint modeling paradigm spanning MNER,VE, and VGperspectives by leveraging large language models (LLMs) as connecting bridges. This reformulation brings two benefits: 1) It enables us to optimize the MNER module for optimal MNER performance and eliminates the need to pre-extract region features using object detection methods, thus naturally addressing the two major limitations of existing GMNER methods. 2) The introduction of Entity Expansion Expression module and Visual Entailment (VE) module unifies Visual Grounding (VG) and Entity Grounding (EG). This endows the proposed framework with unlimited data and model scalability. Furthermore, to address the potential ambiguity stemming from the coarse-grained bounding box output in GMNER, we further construct the new Segmented Multimodal Named Entity Recognition (SMNER) task and corresponding Twitter-SMNER dataset aimed at generating fine-grained segmentation masks, and experimentally demonstrate the feasibility and effectiveness of using box prompt-based Segment Anything Model (SAM) to empower any GMNER model with the ability to accomplish the SMNER task. Extensive experiments demonstrate that RiVEG significantly outperforms SoTA methods on four datasets across the MNER, GMNER, and SMNER tasks. Datasets and Code will be released athttps://github.com/JinYuanLi0012/RiVEG.
Jianfei Yu, Di Sun 0001, Gang Pan 0002
IEEE Trans. Multim.6
2025 EGAP-YOLO: An Efficient Crack Detection Model Based on YOLO Architecture
Jianrong Li, Haifeng Fan, Di Sun 0001, Chuanlei Zhang, Yinglun Dong
ICIC (2)6
2025 MS-DETR: Multi-Scale and Attention-Enhanced Rust Detection for Bolts and Nuts in Transmission Lines
Di Sun 0001, Chaojie Yao, Haifeng Fan, Chuanlei Zhang
ICIC (11)1
2025 LSGNSF: A Graph-Based Time Series Anomaly Detection Algorithm
Chuanlei Zhang, Yinglun Dong, Jianrong Li, Haifeng Fan, Di Sun 0001
ICIC (20)6
2025 Formula Spotting Based on Synergy Perception and Representation Mining
abstract
Formula spotting aims to simultaneously detect and recognize formulas in documents, with broad applications in intelligent document parsing, mathematical reasoning, and more. Although existing methods that first detect and then recognize have achieved prominent results, they still suffer from semantic confusion from similar character structures, semantic loss from bounding box perturbation, and visual interference from non-formula regions. To address these issues, we propose a Synergy Perception and Representation Mining Network. This network facilitates explicit interaction between the RoI features of the detection module and the semantic features of the recognition module, leveraging additional visual priors to distinguish subtle differences in similar characters. Moreover, to better perceive the boundary character structure of formulas and filter out irrelevant visual interference, a Formula Representation Mining module is proposed. This module employs progressive attention mining to achieve a complementarity between semantic information and visual context without disrupting the linguistic priors of the formulas. Additionally, to enhance the efficiency of formula decoding, we propose a parallel mask, allowing the network to output multiple LaTeX tokens simultaneously in a single prediction step. To evaluate the effectiveness of our method in formula spotting, two novel datasets: Formula-7K and Exam-1K are established. To the best of our knowledge, they are the first formula spotting datasets. Experimental results on Formula-7K and Exam-1K validate the generality and effectiveness of the proposed method. Code is available at https://github.com/hongen123/SynRMFormer.
Gang Pan 0002, Hongen Liu, Di Sun 0001
ACM Multimedia3
2025 Image Retargeting based on Text Region Awareness
abstract
Image retargeting (IR) with text regions is a challenging yet underexplored task that focuses on resizing an image's aspect ratio while preserving both semantic objects and the legibility of textual content. This task introduces three primary challenges, which can be summarized as follows: (1) the distinct probability distributions between text and non-text regions in images; (2) the lack of dedicated mechanisms in existing IR methods for handling text regions, often leading to text distortion or blurring; (3) the absence of paired datasets specifically designed for IR tasks involving text regions. To tackle these challenges, we propose SSIR, a unified framework that reformulates IR as a joint Semantic Segmentation and Image Retargeting (SS-IR) task, leveraging an attention mechanism to bridge these components. Specifically, we first employ a semantic segmentation sub-network that extracts text region features using text segmentation techniques to improve text-awareness in retargeting tasks. Then, we integrate text features into the image's visual representation through an attention-driven module designed to preserve both textual and semantic content during retargeting. Finally, we address the absence of paired datasets with an unsupervised learning paradigm based on a Cycle-IR framework, which employs cyclic consistency reconstruction, enabling effective learning without the need for paired training data. Experimental results show that the SSIR algorithm effectively preserves text information and delivers high-quality visual retargeting results.
Gang Pan 0002, Meihua Liu, Lei Zhou 0031, Di Sun 0001
ACM Multimedia5
2025 AFFIR: Dual-Modal Attention Feature Fusion for Scene Text Image Retargeting
abstract
Image retargeting technique aims to adjust and reorganize the content of original images to fit different display sizes and visual requirements. Text elements frequently appear in real-world images and play a crucial role in conveying information. Existing algorithms often treat the image as a whole during retargeting, neglecting the unique features of textual content. This oversight results in missing textual information or distorted character structures, ultimately failing to effectively preserve the integrity of text regions, thereby affecting both the efficiency of information transmission and visual quality of the final image. To address the aforementioned issues, we start from the perception of textual content, which guides retargeted image generation through the fusion of attention features. Specifically, a Transformer-based model is employed for the image retargeting tasks in this study. Text and image features are extracted separately, accompanied by a dual-modal feature fusion strategy, which integrates text and image features through attention maps generated. The training process adopts a cyclic training strategy, where the retargeted results are fed back into the model in reverse. This approach is applicable to retargeting images of various sizes, ensuring that detailed information from both text and image content is accurately preserved. Extensive evaluations on benchmark datasets demonstrate that our method significantly outperforms existing techniques in maintaining both textual clarity and overall visual quality, making it a promising solution for advanced multimedia applications in computer science.
Gang Pan 0002, Liming Pan, Hongze Mi, Rongyu Xiong, Di Sun 0001
ACM Multimedia6
2025 AFAN: An Attention-Driven Forgery Adversarial Network for Blind Image Inpainting
abstract
Blind image inpainting is a challenging task aimed at reconstructing corrupted regions without relying on mask information. Due to the lack of mask priors, previous methods usually integrate a mask prediction network in the initial phase, followed by an inpainting backbone. However, this multi-stage generation process may result in feature misalignment. While recent end-to-end generative methods bypass the mask prediction step, they typically struggle with weak perception of contaminated regions and introduce structural distortions. This study presents a novel mask region perception strategy for blind image inpainting by combining adversarial training with forgery detection. To implement this strategy, we propose an attention-driven forgery adversarial network (AFAN), which leverages adaptive contextual attention (ACA) blocks for effective feature modulation. Specifically, within the generator, ACA employs self-attention to enhance content reconstruction by utilizing the rich contextual information of adjacent tokens. In the discriminator, ACA utilizes cross-attention with noise priors to guide adversarial learning for forgery detection. Moreover, we design a high-frequency omni-dimensional dynamic convolution (HODC) based on edge feature enhancement to improve detail representation. Extensive evaluations across multiple datasets demonstrate that the proposed AFAN model outperforms existing generative methods in blind image inpainting, particularly in terms of quality and texture fidelity.
Gang Pan 0002, Di Sun 0001, Jiawan Zhang
IEEE Trans. Multim.3
2024 Domain Adaptive Object Detection with Dehazing Module
Gang Pan 0002, Jingxin Li, Rufei Zhang, Sheng Shen 0013, Zhiliang Zeng, Di Sun 0001
ICIC (11)8
2024 Anomaly Detection of Transmission Line Large Metal Based on EGFPN-YOLO and UAVs
Gongcheng Shi, Jianrong Li, Yicong Li 0009, Di Sun 0001, Chuanlei Zhang
ICIC (5)7
2024 MulTIR: Deep Multi-Target Image Retargeting
Di Sun 0001, Chaojie Yao, Yijing Mei, Dufeng Chen, Gang Pan 0002
ICIC (7)1
2024 DGAP-YOLO: A Crack Detection Method Based on UAV Images and YOLO
Yunyi Li, Jianrong Li, Di Sun 0001, Chuanlei Zhang
ICIC (11)6
2024 Arbitrary Scale Texture Synthesis with Feature Map Swapping
Di Sun 0001, Yangde Lin, Sheng Shen 0013, Zhiliang Zeng, Shizhao Zhang
ICIC (7)1
2024 BiRGAN: Bi-directional Deep Image Retargeting
Di Sun 0001, Yunxiang Wang, Yijing Mei, Gang Pan 0002
ICIC (7)1
2024 Chinese Character Image Inpainting with Skeleton Extraction and Adversarial Learning
Di Sun 0001, Xiangyu Pan, Gang Pan 0002
ICIC (7)1
2024 Rust Detection Network for Transmission Line Based on UAV Inspection
Di Sun 0001, Chao Ren 0003, Chuanlei Zhang
ICIC (11)1
2024 Improved Real-Time Monitoring Lightweight Model for UAVs Based on YOLOv8
Chuanlei Zhang, Xingchen Zhao, Di Sun 0001, Guoyi Xu, Runjun Zhao
ICIC (11)3
2021 Deep Supervised Image Retargeting
abstract
Recent learning-based image retargeting methods have achieved significant improvement. However, two main is-sues remain in this challenging task: (i) it is difficult to build ground truth datasets for supervised learning; (ii) most methods are based on a certain operator, not suitable for various images with different target sizes. In this paper, for the first time, we address these issues by providing a deep supervised image retargeting solution. We introduce a new dataset1of 6, 576 pairs generated by multiple operators using Image Re-targeting Quality Assessment (IRQA) algorithm. We then develop a mult-operator image retargeting model named MR-GAN, which learns the deformation process of retargeted images using multiple methods and conducts retargeting operations in feature space. Experimental results validate the effectiveness as well as its superiority against state-of-the-art alternatives of the proposed approach.
Yijing Mei, Xiaojie Guo 0001, Di Sun 0001, Gang Pan 0002, Jiawan Zhang
ICME3
2021 Chinese Character Inpainting with Contextual Semantic Constraints
abstract
Chinese character inpainting is a challenging task where large missing regions have to be filled with both visually and semantic realistic contents. Existing methods generally produce pseudo or ambiguous characters due to lack of semantic information. Given the key observation that Chinese characters contain visually glyph representation and intrinsic contextual semantics, we tackle the challenge of similar Chinese characters by modeling the underlying regularities among glyph and semantic information. We propose a semantics enhanced generative framework for Chinese character inpainting, where a global semantic supervising module (GSSM) is introduced to constrain contextual semantics. In particular, sentence embedding is used to guide the encoding of continuous contextual characters. The method can not only generate realistic Chinese character, but also explicitly utilize context as reference during network training to eliminate ambiguity. The proposed method is evaluated on both handwritten and printed Chinese characters with various masks. The experiments show that the method successfully predicts missing character information without any mask input, and achieves significant sentence-level results benefiting from global semantic supervising in a wide variety of scenes.
Gang Pan 0002, Di Sun 0001, Jiawan Zhang
ACM Multimedia3
2018 Symmetry-Aware Face Completion with Generative Adversarial Networks
Jiawan Zhang, Rui Zhan, Di Sun 0001, Gang Pan 0002
ACCV (4)3
2018 Mural Sketch Generation via Style-aware Convolutional Neural Network
abstract
Sketch is one of the most important art expression forms for traditional Chinese painting. This paper presents a complete sketch generation framework for ancient mural paintings. First, we propose a deep learning network to perform mural-to-sketch prediction by combining meaningful convolutional features in a holistic manner. A dedicated mural database with fine-grained ground truth is built for network training and testing. Then we design a style-aware image fusion approach by detecting the specific feature region in a mural, from which the artistic style can be maximally preserved. Experimental results have demonstrated its validity in extracting style mural sketch. This work has the potential to provide a computer aided tool for artists and restorers to imitate and restore time-honored paintings.
Gang Pan 0002, Di Sun 0001, Rui Zhan, Jiawan Zhang
CGI2
2018 Mural2Sketch: A Combined Line Drawing Generation Method for Ancient Mural Painting
abstract
Line drawing is a unique drawing technique developed over millennia in China. Since ancient murals have line drawings of beautiful form and vast history, it is incredibly important to digitally curate these pieces. In this paper, we propose a line drawing generation method named Mural2Sketch (MS) for ancient mural paintings. MS first utilizes heuristic routing to detect the outer edge of a stroke, and then high frequency enhancement filtering is used to extract the information inside the stroke. A complete stroke is then generated by collaborative representation. MS is capable of outputting the result in vector form and producing different artistic styles. Experimental results show that our method is simple but effective. This research has the potential to support digital mural copying, mural protection, as well as related cultural research and application.
Di Sun 0001, Jiawan Zhang, Gang Pan 0002, Rui Zhan
ICME1
2012 StoryWizard: a framework for fast stylized story illustration
Jiawan Zhang, Yukun Hao, Liang Li 0039, Di Sun 0001
Vis. Comput.4
2011 Structure-Aware Image Completion with Texture Propagation
abstract
Structure-aware image completion has the reputation of keeping the salient structure of images, except for the case in texture propagation of images with distinguishing texture characteristics and large missing region. This inspirits us to design this new algorithm that could intensify texture-aware functions in the structure completion processing in an optimal manner. To avoid the occurrence of visually inconsistent results, we consider the filling of missing region as the energy minimization, using Belief Propagation (BP), of discrete Markov random field with integrated constraints of the costs of label, structure coherence and texture coherence. Moreover, wavelet multi-resolution image pyramid is adopted in our method, which not only enhances the global geometric structures and detailed texture features within large missing region, but as well speeds up the rate of convergence in the synthesis process.
Di Sun 0001, Yi Zhang 0070, Jiawan Zhang, Gang Pan 0002
ICIG1