Junsong Zhang

dblp:15/9171 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 NeurIPT: Foundation Model for Neural Interfaces
abstract
Electroencephalography (EEG) has wide-ranging applications, from clinical diagnosis to brain-computer interfaces (BCIs). With the increasing volume and variety of EEG data, there has been growing interest in establishing foundation models (FMs) to scale up and generalize neural decoding. Despite showing early potential, applying FMs to EEG remains challenging due to substantial inter-subject, inter-task, and inter-condition variability, as well as diverse electrode configurations across recording setups. To tackle these open challenges, we propose **NeurIPT**, a foundation model tailored for diverse EEG-based **Neur**al **I**nterfaces with a **P**re-trained **T**ransformer by capturing both homogeneous and heterogeneous spatio-temporal characteristics inherent in EEG signals. Temporally, we introduce Amplitude-Aware Masked Pretraining (AAMP), masking based on signal amplitude rather than random intervals, to learn robust representations across varying signal intensities beyond local interpolation. Moreover, this temporal representation is enhanced by a progressive Mixture-of-Experts (MoE) architecture, where specialized expert subnetworks are progressively introduced at deeper layers, adapting effectively to the diverse temporal characteristics of EEG signals. Spatially, NeurIPT leverages the 3D physical coordinates of electrodes, enabling effective transfer across varying EEG settings, and develops Intra-Inter Lobe Pooling (IILP) during fine-tuning to efficiently exploit regional brain features. Empirical evaluations across nine downstream BCI datasets, via fine-tuning and training from scratch, demonstrated NeurIPT consistently achieved state-of-the-art performance, highlighting its broad applicability and robust generalization. Our work pushes forward the state of FMs in EEG and offers insights into scalable and generalizable neural information processing systems.
Zitao Fang, Hongting Zhou, Shuyang Yu, Guodong Du 0002, Ashwaq Qasem, Jing Li 0034, Junsong Zhang, Sim Kuan Goh
NeurIPS9
2025 Conditional Font Generation With Content Pre-Train and Style Filter
abstract
Abstract Automatic font generation aims to streamline the design process by creating new fonts with minimal style references. This technology significantly reduces the manual labour and costs associated with traditional font design. Image‐to‐image translation has been the dominant approach, transforming font images from a source style to a target style using a few reference images. However, this framework struggles to fully decouple content from style, particularly when dealing with significant style shifts. Despite these limitations, image‐to‐image translation remains prevalent due to two main challenges faced by conditional generative models: (1) inability to handle unseen characters and (2) difficulty in providing precise content representations equivalent to the source font. Our approach tackles these issues by leveraging recent advancements in Chinese character representation research to pre‐train a robust content representation model. This model not only handles unseen characters but also generalizes to non‐existent ones, a capability absent in traditional image‐to‐image translation. We further propose a Transformer‐based Style Filter that not only accurately captures stylistic features from reference images but also handles any combination of them, fostering greater convenience for practical automated font generation applications. Additionally, we incorporate content loss with commonly used pixel‐ and perceptual‐level losses to refine the generated results from a comprehensive perspective. Extensive experiments validate the effectiveness of our method, particularly its ability to handle unseen characters, demonstrating significant performance gains over existing state‐of‐the‐art methods.
Yinfei Li, Xiaojun Qiao, Junsong Zhang
Comput. Graph. Forum4
2025 An Intelligent Surveillance Platform With Deep Tampered Video Detection in Secure Edge-Cloud Services
abstract
The increasing complexity of video tampering techniques poses a significant threat to the integrity and security of Internet of Multimedia Things (IoMT) ecosystems, particularly in resource‐constrained edge‐cloud infrastructures. This paper introduces Multiscale Gated Multihead Attention Depthwise Separable CNN (MGMA‐DSCNN), an advanced deep learning framework specifically optimized for real‐time tampered video detection in IoMT environments. By integrating lightweight convolutional neural networks (CNNs) with multihead attention mechanisms, MGMA‐DSCNN significantly enhances feature extraction while maintaining computational efficiency. Unlike conventional methods, this approach employs a multiscale attention mechanism to refine feature representations, effectively identifying deepfake manipulations, frame insertions, splicing, and adversarial forgeries across diverse multimedia streams. Extensive experiments on multiple forensic video datasets—including the HTVD dataset—demonstrate that MGMA‐DSCNN outperforms state‐of‐the‐art architectures such as VGGNet‐16, ResNet, and DenseNet, achieving an unprecedented detection accuracy of 98.1%. Furthermore, by leveraging edge‐cloud synergy, our framework optimally distributes computational loads, effectively reducing latency and energy consumption, making it highly suitable for real‐time security surveillance and forensic investigations. These advancements position MGMA‐DSCNN as a scalable, high‐performance solution for next‐generation intelligent video authentication, offering robust, low‐latency detection capabilities in dynamic and resource‐constrained IoMT environments.
Yuwen Shao, Qiuling Wang, Junsong Zhang, Haiying Tian
Int. J. Intell. Syst.3
2025 Improving Chinese character representation with formation tree
Xiaojun Qiao, Yinfei Li, Junsong Zhang
Neurocomputing5
2025 SGFormer: Spherical Geometry Transformer for 360° Depth Estimation
abstract
Panoramic distortion poses a significant challenge in 360° depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range dependencies to capture global structures, resulting in either unclear structure or insufficient local perception. In this paper, we propose a spherical geometry transformer, named SGFormer, to address the above issues, with an innovative step to integrate spherical geometric priors into vision transformers. To this end, we retarget the transformer decoder to a spherical prior decoder (termed SPDecoder), which endeavors to uphold the integrity of spherical structures during decoding. Concretely, we leverage bipolar reprojection, circular rotation, and curve local embedding to preserve the spherical characteristics of equidistortion, continuity, and surface distance, respectively. Furthermore, we present a query-based global conditional position embedding to compensate for spatial structure at varying resolutions. It not only boosts the global perception of spatial position but also sharpens the depth structure across different patches. Finally, we conduct extensive experiments on popular benchmarks, demonstrating our superiority over state-of-the-art solutions. Our code will be made publicly athttps://github.com/iuiuJaon/SGFormer.
Junsong Zhang, Zisong Chen, Chunyu Lin, Zhijie Shen, Lang Nie, Kang Liao, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 360 Layout Estimation via Orthogonal Planes Disentanglement and Multi-View Geometric Consistency Perception
abstract
Existing panoramic layout estimation solutions tend to recover room boundaries from a vertically compressed sequence, yielding imprecise results as the compression process often muddles the semantics between various planes. Besides, these data-driven approaches impose an urgent demand for massive data annotations, which are laborious and time-consuming. For the first problem, we propose an orthogonal plane disentanglement network (termed DOPNet) to distinguish ambiguous semantics. DOPNet consists of three modules that are integrated to deliver distortion-free, semantics-clean, and detail-sharp disentangled representations, which benefit the subsequent layout recovery. For the second problem, we present an unsupervised adaptation technique tailored for horizon-depth and ratio representations. Concretely, we introduce an optimization strategy for decision-level layout analysis and a 1D cost volume construction method for feature-level multi-view aggregation, both of which are designed to fully exploit the geometric consistency across multiple perspectives. The optimizer provides a reliable set of pseudo-labels for network training, while the 1D cost volume enriches each view with comprehensive scene information derived from other perspectives. Extensive experiments demonstrate that our solution outperforms other SoTA models on both monocular layout estimation and multi-view layout estimation tasks.
Zhijie Shen, Chunyu Lin, Junsong Zhang, Lang Nie, Kang Liao, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Color Hint-guided Ink Wash Painting Colorization with Ink Style Prediction Mechanism
abstract
We propose an end-to-end generative adversarial network that allows for controllable ink wash painting generation from sketches by specifying the colors via color hints. To the best of our knowledge, this is the first study for interactive Chinese ink wash painting colorization from sketches. To help our network understand the ink style and artistic conception, we introduced an ink style prediction mechanism for our discriminator, which enables the discriminator to accurately predict the style with the help of a pre-trained style encoder. We also designed our generator to receive multi-scale feature information from the feature pyramid network for detail reconstruction of ink wash painting. Experimental results and user study show that ink wash paintings generated by our network have higher realism and richer artistic conception than existing image generation methods.
Yao Zeng, Junsong Zhang
ACM Trans. Appl. Percept.4
2024 NRTR: Neuron Reconstruction With Transformer From 3D Optical Microscopy Images
abstract
The neuron reconstruction from raw Optical Microscopy (OM) image stacks is the basis of neuroscience. Manual annotation and semi-automatic neuron tracing algorithms are time-consuming and inefficient. Existing deep learning neuron reconstruction methods, although demonstrating exemplary performance, greatly demand complex rule-based components. Therefore, a crucial challenge is designing an end-to-end neuron reconstruction method that makes the overall framework simpler and model training easier. We propose a Neuron Reconstruction Transformer (NRTR) that, discarding the complex rule-based components, views neuron reconstruction as a direct set-prediction problem. To the best of our knowledge, NRTR is the first image-to-set deep learning model for end-to-end neuron reconstruction. The overall pipeline consists of the CNN backbone, Transformer encoder-decoder, and connectivity construction module. NRTR generates a point set representing neuron morphological characteristics for raw neuron images. The relationships among the points are established through connectivity construction. The point set is saved as a standard SWC file. In experiments using the BigNeuron and VISoR-40 datasets, NRTR achieves excellent neuron reconstruction results for comprehensive benchmarks and outperforms competitive baselines. Results of extensive experiments indicate that NRTR is effective at showing that neuron reconstruction is viewed as a set-prediction problem, which makes end-to-end model training available.
Rui Lang, Junsong Zhang
IEEE Trans. Medical Imaging4
2023 Adversarial Interactive Cartoon Sketch Colourization with Texture Constraint and Auxiliary Auto-Encoder
abstract
Abstract Colouring cartoon sketches can help children develop their intellect and inspire their artistic creativity. Unlike photo colourization or anime line art colourization, cartoon sketch colourization is challenging due to the scarcity of texture information and the irregularity of the line structure, which is mainly reflected in the phenomenon of colour‐bleeding artifacts in generated images. We propose a colourization approach for cartoon sketches, which takes both sketches and colour hints as inputs to produce impressive images. To solve the problem of colour‐bleeding artifacts, we propose a multi‐discriminator colourization framework that introduces a texture discriminator in the conditional generative adversarial network (cGAN). Then we combined this framework with a pre‐trained auxiliary auto‐encoder, where an auxiliary feature loss is designed to further improve colour quality, and a condition input is introduced to increase the generalization ability over hand‐drawn sketches. We present both quantitative and qualitative evaluations, which prove the effectiveness of our proposed method. We test our method on sketches of varying complexity and structure, then build an interactive programme based on our model for user study. Experimental results demonstrate that the method generates natural and consistent colour images in real time from sketches drawn by non‐professionals.
Shaoqiang Zhu, Yao Zeng, Junsong Zhang
Comput. Graph. Forum4
2023 Functional network: A novel framework for interpretability of deep neural networks
Ben Zhang 0002, Zhetong Dong, Junsong Zhang
Neurocomputing3
2022 Visual perception driven collage synthesis
abstract
A collage is a composite artwork made from the spatial layout of multiple pictures on a canvas, collected from the Internet or user photographs. Collages, usually made by skilled artists, involve a complex manual process, especially when searching for component pictures and adjusting their spatial layout to meet artistic requirements. In this paper, we present a visual perception driven method for automatically synthesizing visually pleasing collages. Unlike previous works, we focus on how to design a collage layout which not only provides easy access to the theme of the overall image, but also conforms to human visual perception. To achieve this goal, we formulate the generation of collages as a mapping problem: given a canvas image, first, compute a saliency map for it and a vector field for each sub-region of it. Second, using a divide-and-conquer strategy, generate a series of patch sets from the canvas image, where the salient map and the vector field are used to determine each patch’s size and direction respectively. Third, construct a Gestalt-based energy function to choose the most visually pleasing and orderly patch set as the final layout. Finally, using a semantic-color metric, map the picture set to the patch set to generate the final collage. Extensive experimental and user study results show that this method can generate visual pleasing collages.
Zuyi Yang, Qinghui Dai, Junsong Zhang
Comput. Vis. Media3
2022 UGSC-GAN: User-guided sketch colorization with deep convolution generative adversarial networks
abstract
Abstract Inspired by the creating process of human paintings, we propose a novel adversarial architecture for multiple sketch colorization which is a scribble‐based, automatic, and exemplar‐based colorization method. The proposed framework has two stages, namely, imitating stage and shading stage. In the imitating stage, to address the challenge of lack of texture in the sketch, we train a grayscale generation network to accomplish a mapping task, namely, generating a grayscale map with textured, grayscale, boundary information from the input sparse sketch. In the shading stage, the model can accurately colorize the objects in the gray image generated in the previous stage, and generate high‐quality colorized images. With the proposed model trained on our database, the experimental results show that our method can generate vivid colorized images and achieve a better performance than previous methods evaluated by FID metric.
Junsong Zhang, Shaoqiang Zhu, Kunxiang Liu
Comput. Animat. Virtual Worlds1
2022 Creating Word Paintings Jointly Considering Semantics, Attention, and Aesthetics
abstract
In this article, we present a content-aware method for generating a word painting. Word painting is a composite artwork made from the assemblage of words extracted from a given text, which carries similar semantics and visual features to a given source image. However, word painting, usually created by skilled artists, involves tedious manual processes, especially when generating streamlines and laying out text. Hence, we provide an easy method to create word paintings for users. How to design textural layout that simultaneously conveys the input image and enables easy access to the semantic theme is the key challenge to generating a visually pleasing word painting. To address this issue, given an image and its content-related text, we first decompose the input image into several regions and approximate each region with a smooth vector field. At the same time, by analyzing the input text, we extract some weighted keywords as the graphic elements. Then, to measure the likelihood of positions in the input image that attract the observers’ attention, we generate a saliency map with our trained visual attention model. Finally, jointly considering visual attention and aesthetic rules, we propose an energy-based optimization framework to arrange extracted keywords into the decomposed regions and synthesize a word painting. Experimental results and user studies show that this method is able to generate a fashionable and appealing word painting.
Junsong Zhang, Zuyi Yang, Linchengyu Jin, Zhitang Lu
ACM Trans. Appl. Percept.1
2022 Reconfiguration of the brain during aesthetic experience on Chinese calligraphy - Using brain complex networks
abstract
Chinese calligraphy, as a well-known performing art form, occupies an important role in the intangible cultural heritage of China. Previous studies focused on the psychophysiological benefits of Chinese calligraphy. Little attention has been paid to its aesthetic attributes and effectiveness on the cognitive process. To complement our understanding of Chinese calligraphy, this study investigated the aesthetic experience of Chinese cursive-style calligraphy using brain functional network analysis. Subjects stayed on the coach and rested for several minutes. Then, they were requested to appreciate artwork of cursive-style calligraphy. Results showed that (1) changes in functional connectivity between fronto-occipital, fronto-parietal, bilateral parietal, and central–occipital areas are prominent for calligraphy condition, (2) brain functional network showed an increased normalized cluster coefficient for calligraphy condition in alpha2 and gamma bands. These results demonstrate that the brain functional network undergoes a dynamic reconfiguration during the aesthetic experience of Chinese calligraphy. Providing evidence that the aesthetic experience of Chinese calligraphy has several similarities with western art while retaining its unique characters as an eastern traditional art form.
Xiaofei Jia, Changle Zhou, Junsong Zhang
Vis. Informatics4
2021 A Lightweight and Secure Anonymous User Authentication Protocol for Wireless Body Area Networks
abstract
The recent development of wireless body area network (WBAN) technology plays a significant role in the modern healthcare system for patient health monitoring. However, owing to the open nature of the wireless channel and the sensitivity of the transmitted messages, the data security and privacy threats in WBAN have been widely discussed and must be solved. In recent years, many authentication protocols had been proposed to provide security and privacy protection in WBANs. However, many of these schemes are not computationally efficient in the authentication process. Inspired by these studies, a lightweight and secure anonymous authentication protocol is presented to provide data security and privacy for WBANs. The proposed scheme adopts a random value and hash function to provide user anonymity. Besides, the proposed protocol can provide user authentication without a trusted third party, which makes the proposed scheme have no computational bottleneck in terms of architecture. Finally, the security and performance analyses demonstrate that the proposed scheme can meet security requirements with low computational and communication costs.
Junsong Zhang, Qikun Zhang, Xianling Lu, Yong Gan
Secur. Commun. Networks1
2021 A Novel Privacy-Preserving Authentication Protocol Using Bilinear Pairings for the VANET Environment
abstract
With the rapid development of communication and microelectronic technology, the vehicular ad hoc network (VANET) has received extensive attention. However, due to the open nature of wireless communication links, it will cause VANET to generate many network security issues such as data leakage, network hijacking, and eavesdropping. To solve the above problem, this paper proposes a new authentication protocol which uses bilinear pairings and temporary pseudonyms. The proposed authentication protocol can realize functions such as the identity authentication of the vehicle and the verification of the message sent by the vehicle. Moreover, the proposed authentication protocol is capable of preventing any party (peer vehicles, service providers, etc.) from tracking the vehicle. To improve the efficiency of message verification, this paper also presents a batch authentication method for the vehicle to verify all messages received within a certain period of time. Finally, through security and performance analysis, it is actually easy to find that the proposed authentication protocol can not only resist various security threats but also have good computing and communication performance in the VANET environment.
Junsong Zhang, Qikun Zhang, Xianling Lu, Yong Gan
Wirel. Commun. Mob. Comput.1
2020 Learning EEG topographical representation for classification via convolutional neural network
Meiyan Xu, Junfeng Yao, Zhihong Zhang 0001, Baorong Yang, Chunyan Li 0002, Junsong Zhang
Pattern Recognit.8
2019 Artistic Augmentation of Photographs with Droplets
Mohan Zhang, Kang Zhang 0001, Junsong Zhang
J. Comput. Sci. Technol.4
2017 Synthesizing Ornamental Typefaces
abstract
Abstract We present a method for creating ornamental typeface images. Ornamental typefaces are a composite artwork made from the assemblage of images that carry similar semantics to words. These appealing word‐art works often attract the attention of more people and convey more meaningful information than general typefaces. However, traditional ornamental typefaces are usually created by skilled artists, which involves tedious manual processes, especially when searching for appropriate materials and assembling them. Hence, we aim to provide an easy way to create ornamental typefaces for novices. How to combine users' design intentions with image semantic and shape information to obtain readable and appealing ornamental typefaces is the key challenge to generate ornamental typefaces. To address this problem, we first provide a scribble‐based interface for users to segment the input typeface into strokes according to their design concepts. To ensure the consistency of the image semantics and stroke shape, we then define a semantic‐shape similarity metric to select a set of suitable images. Finally, to beautify the typeface structure, an optional optimal strategy is investigated. Experimental results and user studies show that the proposed algorithm effectively generates attractive and readable ornamental typefaces.
Junsong Zhang, Weiyi Xiao, Zhenshan Luo
Comput. Graph. Forum1
2017 Computational Aesthetic Evaluation of Logos
abstract
Computational aesthetics has become an active research field in recent years, but there have been few attempts in computational aesthetic evaluation of logos. In this article, we restrict our study on black-and-white logos, which are professionally designed for name-brand companies with similar properties, and apply perceptual models of standard design principles in computational aesthetic evaluation of logos. We define a group of metrics to evaluate some aspects in design principles such as balance, contrast, and harmony of logos. We also collect human ratings of balance, contrast, harmony, and aesthetics of 60 logos from 60 volunteers. Statistical linear regression models are trained on this database using a supervised machine-learning method. Experimental results show that our model-evaluated balance, contrast, and harmony have highly significant correlation of over 0.87 with human evaluations on the same dimensions. Finally, we regress human-evaluated aesthetics scores on model-evaluated balance, contrast, and harmony. The resulted regression model of aesthetics can predict human judgments on perceived aesthetics with a high correlation of 0.85. Our work provides a machine-learning-based reference framework for quantitative aesthetic evaluation of graphic design patterns and also the research of exploring the relationship between aesthetic perceptions of human and computational evaluation of design principles extracted from graphic designs.
Kang Zhang 0001, Xianjun Sam Zheng, Junsong Zhang
ACM Trans. Appl. Percept.5
2017 Generating Multi-Destination Maps
abstract
Multi-destination maps are a kind of navigation maps aimed to guide visitors to multiple destinations within a region, which can be of great help to urban visitors. However, they have not been developed in the current online map service. To address this issue, we introduce a novel layout model designed especially for generating multi-destination maps, which considers the global and local layout of a multi-destination map. We model the layout problem as a graph drawing that satisfies a set of hard and soft constraints. In the global layout phase, we balance the scale factor between ROIs. In the local layout phase, we make all edges have good visibility and optimize the map layout to preserve the relative length and angle of roads. We also propose a perturbation-based optimization method to find an optimal layout in the complex solution space. The multi-destination maps generated by our system are potential feasible on the modern mobile devices and our result can show an overview and a detail view of the whole map at the same time. In addition, we perform a user study to evaluate the effectiveness of our method, and the results prove that the multi-destination maps achieve our goals well.
Junsong Zhang, Jiepeng Fan, Zhenshan Luo
IEEE Trans. Vis. Comput. Graph.1
2007 A Novel Method for Vectorizing Historical Documents of Chinese Calligraphy
abstract
We develop a novel method for feature point detection and employ it to generate outline font from historical document of Chinese calligraphy. The feature points at a character contour subdivide the contour into segments. Each segment can be then fitted by a parametric curve to obtain the outline font. Some experimental results are also presented in the paper.
Junsong Zhang
CAD/Graphics1