EDBT 2026 Demo / reviewers in the wild / expert
Junpeng Tan
dblp:273/3906
· DBLP profile ↗
22ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0001-9546-4331ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural clothing tryer: Customized virtual try-on via semantic enhancement and controlling diffusion model
Zhijing Yang, Yukai Shi, Junpeng Tan, Tianshui Chen, Liruo Zhong |
Expert Syst. Appl. | 6 |
| 2026 | Progressive multi-branch video style transfer network via confidence reweighted projection
Kunbo Han, Hongyan Yin, Junpeng Tan, Chong-zhi Gao, Chunmei Qing |
Neural Networks | 3 |
| 2025 | Enhancing fNIRS Signal Classification with Test-Time Training by Improved Spatiotemporal Feature ExtractionabstractFunctional Near-Infrared Spectroscopy (fNIRS) is a convenient brain imaging technology that is adaptable to various complex environments. It can detect the brain's responses to external stimuli across different contexts. However, existing fNIRS processing methods fail to mine the brain's sequential association responses in long-term signals over time. To better capture the hidden spatiotemporal correlations and long-range dependencies inherent in fNIRS signals, we propose fNIRS-TTT, a novel classification architecture leveraging Test-Time Training (TTT). Our framework introduces two core innovations: the Cross-Attention Embedding module (CAE) and the TTT Block. The CAE module combines the wide-kernel Patch Conv (capturing cross-channel spatial patterns) and the Channel Conv (processing temporal features within individual channels) together. Their outputs are fused via Cross-Attention to generate spatiotemporal tokens enriched with multi-scale spatial information for fNIRS. The TTT components dynamically adapt to the sequential nature of fNIRS data, effectively capturing long-range temporal context and mitigating overfitting compared to former models. KFold Cross-Validation (KFold-CV) and Leave-One-Subject Cross-Validation (LOSO-CV) are conducted on three open datasets, which demonstrate the superior performance of the proposed fNIRS-TTT compared to state-of-the-art models. Wanxiang Luo, Chunmei Qing, Junpeng Tan, Yihang Zou, Xiangmin Xu 0001 |
BIBM | 3 |
| 2025 | DP-Net: A 3D Dilated Projection Framework For Precise Fetal Brain Tissue SegmentationabstractPrecise segmentation of fetal brain tissues in MRI is essential for studying brain development and for the early diagnosis and treatment of neurological disorders. However, the complex and variable anatomy of the fetal brain, significant morphological changes at different gestational ages, and the low-quality MRI and inherent noise of fetal acquisition pose significant challenges. To address these, we propose a novel 3D Dilated Projection U-net segmentation framework, DP-Net, which incorporates large kernel convolutions, atrous convolution for receptive field expansion, and dual skip connections mechanism to enhance global semantic consistency. Specifically, we introduce a Dilated Projection Block (DPB) that leverages atrous convolution to capture global context across multiple anatomical regions without additional parameters. Furthermore, we propose a Dual Skip Connection (DSC) mechanism to maintain encoder-decoder global consistency by fusing low-level and projected high-level features, mitigating blind spots introduced by atrous convolution. Extensive experiments show that our method significantly outperforms state-of-the-art methods, demonstrating its robustness and effectiveness in addressing the challenges of fetal brain tissue segmentation. Junpeng Tan, Mingjin Chen, Chunmei Qing, Xin Zhang 0013, Xiangmin Xu 0001 |
ICIP | 1 |
| 2025 | Artistic Style Transfer via Fine-Grained Text Guidance and Contrastive Semantics SimilarityabstractDue to the development of text-image multimodal methods, text is used to guide the style transfer of images, which has attracted growing attention. Notably, The existing text-guided image style methods are limited to expressing specific artistic style through simple text. It can only accept coarse-grained text input such as “Van Gogh” and “White Cloud”, and cannot understand fine-grained text input such as “The Night Café by Vincent van Gogh”. To this end, this paper proposes a novel artistic style transfer network based on the fine-grained text guidance and the contrastive semantics similarity, named as TCStyler. It can accept images or texts as style guidance, which is more suitable for fine-grained content understanding stylization. In this network, to address the issue of text-image cross-modal discrepancy, the residual attention feature mapper (RAFM) is introduced to constrain the differences between different modalities in feature space. Then, the global cascading style-sharing module (GCSM) is proposed for performing content-style feature fusion and image-text modality fusion by adopting a global feature-sharing strategy. Furthermore, the contrastive semantics similarity loss is designed to address the problem of multimodal universality. Quantitative and visualization experiments demonstrate that our TCStyler can handle fine-grained artistic text inputs and maintain consistency in the style transfer results guided by different modalities. Chunmei Qing, Junpeng Tan, Jianxiu Jin, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Multi-view Panoramic Image Style Transfer with Multi-scale Attention and Global SharingabstractStyle transfer for panoramic images is a challenging task, due to the problems associated with its unique structure, including edge discontinuities, pole distortion, fuzzy details, and memory limitation. In this article, we propose a novel Multi-view Transformation network for Panorama Style Transfer (MuTPST). First, this architecture has a multi-view panoramic transformation mechanism, which includes a multi-view cubic projection and a multi-view equirectangular re-projection of panoramic images. This can address pole distortion and edge discontinuity by skillfully applying multiple types of projections and transformations. To capture different levels of context and structure in the stylization stage, we carefully design a multi-scale attention content encoder, which can coordinate the distribution of visual attention across space and channels. Besides, by the sharing of global style features in thumbnails and patches, MuTPST can process ultra-high-resolution panoramic images (e.g., 10,000 \(\times\) 5,000 pixels) with limited GPU memory. Extensive experiments illustrate that the proposed method outperforms the state-of-the-art with a discernible improvement in panoramic image style transfer. More results and interactive features can be found on https://weiyang001.github.io/MuTPST/ . Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Consistent Panoramic Video Style Transfer via Temporal-Spatial Cross Perception
Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
ICIC (6) | 3 |
| 2024 | Fetal MRI Reconstruction by Global Diffusion and Consistent Implicit Representation
Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Chaoxiang Yang, He Zhang 0023, Gang Li 0001, Xiangmin Xu 0001 |
MICCAI (7) | 1 |
| 2024 | Asymmetric low-rank double-level cooperation for scalable discrete cross-modal hashing
Junpeng Tan, Yinghong Zhou, Zhijing Yang, Feiping Nie 0001, Tianshui Chen |
Expert Syst. Appl. | 2 |
| 2024 | Unsupervised multi-perspective fusing semantic alignment for cross-modal hashing retrieval
Yongfeng Chen, Junpeng Tan, Zhijing Yang, Yukai Shi, Jinghui Qin |
Multim. Tools Appl. | 2 |
| 2024 | Discriminative latent semantics-preserving similarity embedding hashing for cross-modal retrieval
Yongfeng Chen, Junpeng Tan, Zhijing Yang, Yongqiang Cheng 0001 |
Neural Comput. Appl. | 2 |
| 2024 | MASANet: Multi-Aspect Semantic Auxiliary Network for Visual Sentiment AnalysisabstractRecently, multi-modal affective computing has demonstrated that introducing multi-modal information can enhance performance. However, multi-modal research faces significant challenges due to its high requirements regarding data acquisition, modal integrity, and feature alignment. The widespread use of multi-modal pre-training methods offers the possibility of aiding visual sentiment analysis by introducing cross-domain knowledge. This paper proposes a Multi-Aspect Semantic Auxiliary Network (MASANet) for visual sentiment analysis. Specifically, MASANet achieves modality expansion through cross-modal generation, making it possible to introduce cross-domain semantic assistance. Then, a cross-modal gating module and an adaptive modal fusion module are proposed for aspect-level and cross-modal interaction, respectively. In addition, a designed semantic polarity constraint loss is presented to improve sentiment multi-classification performance. Evaluations of eight widely-used affective image datasets demonstrate that our proposed method outperforms the state-of-the-art methods. Further ablation experiments and visualization results also confirm the effectiveness of the proposed method and its modules. Jinglun Cen, Chunmei Qing, Haochun Ou, Xiangmin Xu 0001, Junpeng Tan |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | DPHANet: Discriminative Parallel and Hierarchical Attention Network for Natural Language Video LocalizationabstractNatural Language Video Localization (NLVL) has recently attracted much attention because of its practical significance. However, the existing methods still face the following challenges: 1) When the models learn intra-modal semantic association, the temporal causal interaction information and contextual semantic discriminative information are ignored, resulting in the lack of intra-modal semantic context connection; 2) When learning fusion representations, existing cross-modal interaction modules lack hierarchical attention function to extract inter-modal similarity information and intra-modal self-correlation information, resulting in insufficient cross-modal information interaction; and 3) When the loss function is optimized, the existing models ignore the correlation of causal inference between the start and end boundaries, resulting in inaccurate start and end boundary calibrations. To conquer the above challenges, we proposed a novel NLVL model, called Discriminative Parallel and Hierarchical Attention Network (DPHANet). Specifically, we emphasized the importance of temporal causal interaction information and contextual semantic discriminative information and correspondingly proposed a Discriminative Parallel Attention Encoder (DPAE) module to infer and encode the above critical information. Besides, to overcome the shortcomings of the existing cross-modal interaction modules, we designed a Video-Query Hierarchical Attention (VQHA) module, which can perform cross-modal interaction and intra-modal self-correlation modeling in a hierarchical manner. Furthermore, a novel deviation loss function was proposed to capture the correlation of causal inference between the start and end boundaries and force the model to focus on the continuity and temporal causality in the video. Finally, extensive experiments on three benchmark datasets demonstrated the superiority of our proposed DPHANet model, which has achieved about 1.5% and 3.5% average performance improvement and about 2.5% and 7.5% maximum performance improvement on the Charades-STA and TACoS datasets respectively. Junpeng Tan, Zhijing Yang, Yongqiang Cheng 0001, Liang Lin 0004 |
IEEE Trans. Multim. | 2 |
| 2024 | Extensible Max-Min Collaborative Retention for Online Mini-Batch Learning Hash RetrievalabstractAlong with the concern of similarity measures in linear space, supervised online hash methods have been applied to the retrieval task. However, they ignored multi-dimensional space semantic mining and association characteristics will cause quantization errors of information hash code: 1) The similarity relation of discretized data needs to be considered in different spaces; 2) Latent semantic features need to be continuously embedded into hash code learning; 3) The correlation between the structure similarity and discrete hash matrices needs to be continuously optimized. To tackle these challenges, this paper proposes a novel Extensible Max-min Collaborative Retention Online Hash retrieval method based on mini-batch training data (EMCROH). It mainly includes the Max-min Bayesian Similarity Sparse Latent Hash module (MBSSLH), and the Repetition Collaborative Projection Learning module (RCPL). Specifically, MBSSLH is a max-min optimization model. Firstly, to explore the semantic similarity of multi-dimensional space, we propose a novel liner and nonlinear semantic similarity discrimination mechanism based on the log maximum likelihood similarity estimation with Euclidean space and minimize the input batch data features with a common projection matrix. Moreover, to further mine the potential semantic information of the discretization, we also propose a robust sparse discrete latent semantic information extraction submodule based on double latent factors. RCPL can extend the data externally using the repetition collaborative projection matrix with robustness regularization constraint. Finally, a novel max-min embedding iterative step is proposed to solve the batch discrete optimization problem based on Augmented Lagrange Multipliers (ALM) with Alternating Direction Minimization (ADM). Extensive experiments on several well-known large databases demonstrate that EMCROH outperforms the state-of-the-art hash methods. Code and datasets have been publicly available athttps://github.com/Tjeep-Tan/EMCROH. Junpeng Tan, Zhijing Yang, Yongyi Lu, Liang Lin 0004 |
IEEE Trans. Multim. | 1 |
| 2024 | Fourier Domain Robust Denoising Decomposition and Adaptive Patch MRI ReconstructionabstractThe sparsity of the Fourier transform domain has been applied to magnetic resonance imaging (MRI) reconstruction in k -space. Although unsupervised adaptive patch optimization methods have shown promise compared to data-driven-based supervised methods, the following challenges exist in MRI reconstruction: 1) in previous k -space MRI reconstruction tasks, MRI with noise interference in the acquisition process is rarely considered. 2) Differences in transform domains should be resolved to achieve the high-quality reconstruction of low undersampled MRI data. 3) Robust patch dictionary learning problems are usually nonconvex and NP-hard, and alternate minimization methods are often computationally expensive. In this article, we propose a method for Fourier domain robust denoising decomposition and adaptive patch MRI reconstruction (DDAPR). DDAPR is a two-step optimization method for MRI reconstruction in the presence of noise and low undersampled data. It includes the low-rank and sparse denoising reconstruction model (LSDRM) and the robust dictionary learning reconstruction model (RDLRM). In the first step, we propose LSDRM for different domains. For the optimization solution, the proximal gradient method is used to optimize LSDRM by singular value decomposition and soft threshold algorithms. In the second step, we propose RDLRM, which is an effective adaptive patch method by introducing a low-rank and sparse penalty adaptive patch dictionary and using a sparse rank-one matrix to approximate the undersampled data. Then, the block coordinate descent (BCD) method is used to optimize the variables. The BCD optimization process involves valid closed-form solutions. Extensive numerical experiments show that the proposed method has a better performance than previous methods in image reconstruction based on compressed sensing or deep learning. Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Xiangmin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Superpoint Transformer for 3D Scene Instance SegmentationabstractMost existing methods realize 3D instance segmentation by extending those models used for 3D object detection or 3D semantic segmentation. However, these non-straightforward methods suffer from two drawbacks: 1) Imprecise bounding boxes or unsatisfactory semantic predictions limit the performance of the overall 3D instance segmentation framework. 2) Existing method requires a time-consuming intermediate step of aggregation. To address these issues, this paper proposes a novel end-to-end 3D instance segmentation method based on Superpoint Transformer, named as SPFormer. It groups potential features from point clouds into superpoints, and directly predicts instances through query vectors without relying on the results of object detection or semantic segmentation. The key step in this framework is a novel query decoder with transformers that can capture the instance information through the superpoint cross-attention mechanism and generate the superpoint masks of the instances. Through bipartite matching based on superpoint masks, SPFormer can implement the network training without the intermediate aggregation step, which accelerates the network. Extensive experiments on ScanNetv2 and S3DIS benchmarks verify that our method is concise yet efficient. Notably, SPFormer exceeds compared state-of-the-art methods by 4.3% on ScanNetv2 hidden test set in terms of mAP and keeps fast inference speed (247ms per frame) simultaneously. Code is available at https://github.com/sunjiahao1999/SPFormer. Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
AAAI | 3 |
| 2023 | Multi-Scale Transformer Network for Saliency Prediction on 360-Degree ImagesabstractThe latest methods for saliency prediction on 360° images show that better results can be obtained using equirectangular (ERP) images as input. Due to the limitation of the receptive field, existing convolution-based networks cannot capture long-range information in complex 360° images. Although the transformer has the innate ability to capture long-range correlations with self-attention, large dataset requirement limit its application in saliency prediction of 360° images. In this paper, we present a novel Multi-scale Transformer framework for Saliency prediction on 360° images (MTSal360). The Multi-scale Transformer Module (MTM) is designed in the network to aggregate the contextual long-range information, which includes a Convolutional Positional Encoder (CPE) to enable the model could train and test on cubic and ERP format separately to address the insufficient data. Experiments on two public datasets illustrate that MTSal360 achieves better results over the state-of-the-art methods. Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
ICIP | 3 |
| 2023 | Cross-modal hash retrieval based on semantic multiple similarity learning and interactive projection matrix learning
Junpeng Tan, Zhijing Yang, Jielin Ye, Yongqiang Cheng 0001, Jinghui Qin, Yongfeng Chen |
Inf. Sci. | 1 |
| 2023 | Context-Based Adaptive Multimodal Fusion Network for Continuous Frame-Level Sentiment PredictionabstractRecently, video sentiment computing has become the focus of research because of its benefits in many applications such as digital marketing, education, healthcare, and so on. The difficulty of video sentiment prediction mainly lies in the regression accuracy of long-term sequences and how to integrate different modalities. In particular, different modalities may express different emotions. In order to maintain the continuity of long time-series sentiments and mitigate the multimodal conflicts, this paper proposes a novel Context-Based Adaptive Multimodal Fusion Network (CAMFNet) for consecutive frame-level sentiment prediction. A Context-based Transformer (CBT) module was specifically designed to embed clip features into continuous frame features, leveraging its capability to enhance the consistency of prediction results. Moreover, to resolve the multi-modal conflict between modalities, this paper proposed an Adaptive multimodal fusion (AMF) method based on the self-attention mechanism. It can dynamically determines the degree of shared semantics across modalities, enabling the model to flexibly adapt its fusion strategy. Through adaptive fusion of multimodal features, the AMF method effectively resolves potential conflicts arising from diverse modalities, ultimately enhancing the overall performance of the model. The proposed CAMFNet for consecutive frame-level sentiment prediction can ensure the continuity of long time-series sentiments. Extensive experiments illustrate the superiority of the proposed method especially in multimodal conflicts videos. Maochun Huang, Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | A Novel Robust Low-rank Multi-view Diversity Optimization Model with Adaptive-Weighting Based Manifold Learning
Junpeng Tan, Zhijing Yang, Jinchang Ren, Yongqiang Cheng 0001, Bingo Wing-Kuen Ling |
Pattern Recognit. | 1 |
| 2021 | SRAGL-AWCL: A two-step multi-view clustering via sparse representation and adaptive weighted cooperative learning
Junpeng Tan, Zhijing Yang, Yongqiang Cheng 0001, Jielin Ye |
Pattern Recognit. | 1 |
| 2021 | Unsupervised Multi-View Clustering by Squeezing Hybrid Knowledge From Cross View and Each ViewabstractMulti-view clustering methods have been a focus in recent years because of their superiority in clustering performance. However, typical traditional multi-view clustering algorithms still have shortcomings in some aspects, such as removal of redundant information, utilization of various views and fusion of multi-view features. In view of these problems, this paper proposes a new multi-view clustering method, low-rank subspace multi-view clustering based on adaptive graph regularization. We construct two new data matrix decomposition models into a unified optimization model. In this framework, we address the significance of the common knowledge shared by the cross view and the unique knowledge of each view by presenting new low-rank and sparse constraints on the sparse subspace matrix. To ensure that we achieve effective sparse representation and clustering performance on the original data matrix, adaptive graph regularization and unsupervised clustering constraints are also incorporated in the proposed model to preserve the internal structural features of the data. Finally, the proposed method is compared with several state-of-the-art algorithms. Experimental results for five widely used multi-view benchmarks show that our proposed algorithm surpasses other state-of-the-art methods by a clear margin. Junpeng Tan, Yukai Shi, Zhijing Yang, Caizhen Wen, Liang Lin 0004 |
IEEE Trans. Multim. | 1 |