Jingyu Wu

dblp:79/10401 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
18since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2027 Multi-scale fusion diffusion network for salient object detection in optical remote sensing images
Jingyu Wu, Fuming Sun, Mingyu Lu
Expert Syst. Appl.1
2026 Multi-modal cooperative fusion network for dual-stream RGB-D salient object detection
Jingyu Wu, Fuming Sun, Mingyu Lu
Image Vis. Comput.1
2026 Cross-Modal Fusion With Mixture-of-Experts for Efficient RGB-D Salient Object Detection
abstract
Current RGB-D salient object detection (SOD) models are plagued by issues, including excessive model parameters and high computational complexity. These drawbacks impede the model's efficient deployment and constrain the enhancement of model performance. This paper introduces an efficient and lightweight cross-modal feature cross-fusion network, termed CMFNet. In particular, we design an efficient model utilizing MobileViT as the dual-stream backbone network, thereby significantly reducing computational complexity while maintaining robust feature extraction capabilities. Firstly, we propose a Cross-Fusion Module (CFM) designed to integrate multi-scale semantic information from RGB and depth features dynamically. Furthermore, we design a Lightweight-Mixture-of-Experts Module (L-MoE) for multi-modal features, which enhances the representational capacity of fused features at various levels by employing a dynamic routing mechanism and balanced constraint strategy to allocate appropriate expert processing units to features at different scales. Additionally, we design a Multi-scale Feature Refinement Module (MSFR) that captures multi-scale context through a combination of channel-spatial attention mechanisms and depthwise separable convolution, and gradually eliminates feature distribution differences between modalities using a dual-path residual learning strategy. Abundant experimental findings verify that the proposed CMFNet outperforms the 25 existing State-of-the-art (SOTA) methods.
Jingyu Wu, Fuming Sun, Mingyu Lu
IEEE Trans. Multim.1
2025 AgileIR: Memory-Efficient Group Shifted Windows Attention for Lightweight Image Restoration
Hongyi Cai, Mohammad Mahdinur Rahman, Jingyu Wu, Wenzhen Dong
ICANN (2)3
2025 CFPFormer: Cross Feature-Pyramid Transformer Decoder for Medical Image Segmentation
abstract
Feature pyramids have been widely adopted in convolutional neural networks and transformers for tasks in medical image segmentation. However, existing models generally focus on the Encoder-side Transformer for feature extraction. We further explore the potential in improving the feature decoder with a well-designed architecture. We propose Cross Feature Pyramid Transformer decoder (CFPFormer), a novel decoder block that integrates feature pyramids and transformers. Even though transformer-like architecture impress with outstanding performance in segmentation, the concerns to reduce the redundancy and training costs still exist. Specifically, by leveraging patch embedding, cross-layer feature concatenation mechanisms, CFPFormer enhances feature extraction capabilities while complexity issue is mitigated by our Gaussian Attention. Benefiting from Transformer structure and U-shaped connections, our work is capable of capturing long-range dependencies and effectively up-sample feature maps. Experimental results are provided to evaluate CFPFormer on medical image segmentation datasets, demonstrating the efficacy and effectiveness. With a ResNet50 backbone, our method achieves 92.02% Dice Score, highlighting the efficacy of our methods. Notably, our VGG-based model outperformed baselines with more complex ViT and Swin Transformer backbone.
Hongyi Cai, Mohammad Mahdinur Rahman, Wenzhen Dong, Jingyu Wu
IJCNN4
2025 Brillouin temperature extraction technology based on RHN fusion Self-Attention mechanism
abstract
We propose a Brillouin temperature extraction method (Bi-SA-RHN) based on Recurrent Highway Network (RHN) fused with Self-Attention mechanism and Bi-LSTM to solve the problems of low measurement accuracy and poor real-time performance in Brillouin optical time-domain sensing systems. Traditional Lorenz curve fitting (LCF) has limitations, such as poor real-time performance, long computation time, and low accuracy.This study integrates Self-Attention mechanism into RHN, which can better capture global and local features, improve the accuracy of temperature measurement, and more effectively analyze complex fiber optic signals. The model incorporates Bi-LSTM layers, allowing it to capture features in both forward and backward directions while preserving feature integrity. This approach allows for faster and more accurate temperature extraction, improving the accuracy of temperature prediction. Compared to traditional recursive networks, the proposed Bi-SA-RHN more effectively leverages the temporal correlation features of time series, leading to improved prediction accuracy. We constructed a 20 kilometer Brillouin optical time-domain experimental system for data acquisition in the experiment, and conducted simulation training and testing to evaluate the performance of temperature extraction under different network parameters. The findings demonstrate that the Bi-SA-RHN maintains high measurement accuracy even in conditions of low signal-to-noise ratio and large scanning frequency step size. Both simulation and experimental results reveal that the algorithm achieves a minimum root mean square error of 0.27 °C. Furthermore, the performance of Bi-SA-RHN was compared with traditional methods such as Lorenz curve fitting (LCF) and deep neural networks (DNN). When the fiber temperature reached 80 °C, Bi-SA-RHN maintained good extraction accuracy, whereas DNN and LCF showed poor performance.
Jingyu Wu, Zhihui Sun, Shaodong Jiang, Faxiang Zhang, Changfeng Li
IJCNN1
2025 VAEnvGen: A Real-Time Virtual Agent Environment Generation System Based on Large Language Models
abstract
Environment plays an important role in non-verbal communication for human-virtual agent interaction. Existing research explores the influence of an agent’s appearance and attributes to enhance human-virtual agent communication. However, there is no common practice for dynamically adjusting the surrounding environments of the virtual agent. In this paper, we introduce a real-time virtual agent environment generation system (VAEnvGen), which contributes to the field by enhancing users’ content perception and improving task performance through dynamic environment adjustment. The system dynamically analyzes both the appropriate communication environment and filters the key information according to the current context. Leveraging Large Language Models, it generates a pseudo-3D background space to create an engaging atmosphere and a dynamic foreground content space for vivid key information display, thereby significantly enhancing content perception. For widespread adoption and flexibility, VAEnvGen is developed as a web application. We further evaluate the impact of VAEnvGen on content perception, user attention, and subjective satisfaction through a mixed-design user study with 50 participants. Quantitative and qualitative results reveal significant improvements in content perception, task completion time, and user satisfaction when using VAEnvGen. The system effectively redistributes user attention from subtitles and the virtual agent itself to the dynamically generated background and key foreground information, leading to a more immersive and less fatiguing user experience.
Jingyu Wu, Pengchen Chen, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun
Int. J. Hum. Comput. Interact.1
2025 Accurate, Secure, and Efficient Semi-Constrained Navigation Over Encrypted City Maps
abstract
Navigation services enable users to find the shortest path from a starting point$S$to a destination$D$, reducing time, gas, and traffic congestion. Still, navigation users risk the exposure of their sensitive location data. Our motivation arises from how users can accurately, securely, and efficiently navigate from$S$to$D$while passing through$k$unordered stops, i.e., midway locations with a non-fixed visiting order. In this work, we formally define Semi-Constrained Navigation (SCN) and present a novel scheme Hermes to achieve accurate, secure, and efficient SCN. Specifically, we propose a divide-and-conquer approach to strike a good balance between accuracy and efficiency. It recursively depth-first-searches the whole area (a navigation tree) and invokes five carefully-crafted strategies stop-by-stop to compute three subpaths in three sequential subareas. We construct a path-distance oracle to encrypt the road graph and securely implement the strategies by using homomorphic encryption and garble circuits. We formally prove the security in the random oracle model and analyze the search complexity to be less than$O(k^{2})$. We experiment over a real-world city map and compare with six baselines. Results show that path search with$k=4$among$N=1000$intersections requires 5.58 seconds with a 3.2% distance deviation rate and an 82.5% path similarity.
Meng Li 0006, Yifei Chen 0005, Jingyu Wu, Zijian Zhang 0001, Jialing He, Liehuang Zhu, Mauro Conti, Xiaodong Lin 0001
IEEE Trans. Dependable Secur. Comput.4
2024 Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
abstract
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on various visual-language understanding and generation tasks. However, MLLMs occasionally generate content inconsistent with the given images, which is known as "hallucination". Prior works primarily center on evaluating hallucination using standard, unperturbed benchmarks, which overlook the prevalent occurrence of perturbed inputs in real-world scenarios-such as image cropping or blurring-that are critical for a comprehensive assessment of MLLMs' hallucination. In this paper, to bridge this gap, we propose Hallu-PI, the first benchmark designed to evaluate Hallucination in MLLMs within Perturbed Inputs. Specifically, Hallu-PI consists of seven perturbed scenarios, containing 1,260 perturbed images from 11 object types. Each image is accompanied by detailed annotations, which include fine-grained hallucination types, such as existence, attribute, and relation. We equip these annotations with a rich set of questions, making Hallu-PI suitable for both discriminative and generative tasks. Extensive experiments on 12 mainstream MLLMs, such as GPT-4V and Gemini-Pro Vision, demonstrate that these models exhibit significant hallucinations on Hallu-PI, which is not observed in unperturbed scenarios. Furthermore, our research reveals a severe bias in MLLMs' ability to handle different types of hallucinations. We also design two baselines specifically for perturbed scenarios, namely Perturbed-Reminder and Perturbed-ICL. We hope that our study will bring researchers' attention to the limitations of MLLMs when dealing with perturbed inputs, and spur further investigations to address this issue. Our code and datasets are publicly available at https://github.com/NJUNLP/Hallu-PI.
Peng Ding 0001, Jingyu Wu, Jun Kuang, Dan Ma 0008, Xuezhi Cao, Shi Chen 0005, Jiajun Chen 0001, Shujian Huang
ACM Multimedia2
2024 A Novel zk-SNARKs Method for Cross-chain Transactions in Multi-chain System
Pengcheng Xia 0004, Jingyu Wu, Yiyang Ni 0001, Jun Li 0004
TrustCom2
2024 Convergence to Equilibrium of No-Regret Dynamics in Congestion Games
Volkan Cevher, Wei Chen 0020, Leello Tadesse Dadi, Jing Dong 0008, Ioannis Panageas, Stratis Skoulakis, Luca Viano, Baoxiang Wang 0001, Siwei Wang 0002, Jingyu Wu
WINE10
2024 CNAMD Corpus: A Chinese Natural Audiovisual Multimodal Database of Conversations for Social Interactive Agents
abstract
Impressive progress has been made in developing companion Socially Interactive Agents (SIAs) that provide companionship and reduce loneliness. However, recent works focus on analyzing multimodal feedback in Answer part but ignore Question part. Furthermore, research on SIAs is primarily based on English, which poses a challenge for Chinese SIAs because of cultural differences between English and Chinese. Therefore, we introduce a Chinese Natural Audiovisual Multimodal Database (CNAMD) corpus, the first and largest freely available Chinese multimodal database for multi-person interaction, containing 48 hours of videos and annotations across eight modalities. Using CNAMD, we analyze the characteristics of vocal-verbal, audio, behavioral, and multimodal combinations during questioning, test the performance of six baselines on three tasks, and propose improvements for processing daily Chinese data. The present findings will help designers consider Chinese customs and language when designing Chinese SIAs, making them more suitable for the Chinese cultural context and users.
Jingyu Wu, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun, Hongzeng Zhang, Yanxu Li
Int. J. Hum. Comput. Interact.1
2024 Unsupervised image style transformation of generative adversarial networks based on cyclic consistency
Jingyu Wu, Fuming Sun, Mingyu Lu
Multim. Syst.1
2023 Preserving Structural Consistency in Arbitrary Artist and Artwork Style Transfer
abstract
Deep generative models are effective in style transfer. Previous methods learn one or several specific artist-style from a collection of artworks. These methods not only homogenize the artist-style of different artworks of the same artist but also lack generalization for the unseen artists. To solve these challenges, we propose a double-style transferring module (DSTM). It extracts different artist-style and artwork-style from different artworks (even untrained) and preserves the intrinsic diversity between different artworks of the same artist. DSTM swaps the two styles in the adversarial training and encourages realistic image generation given arbitrary style combinations. However, learning style from single artwork can often cause over-adaption to it, resulting in the introduction of structural features of style image. We further propose an edge enhancing module (EEM) which derives edge information from multi-scale and multi-level features to enhance structural consistency. We broadly evaluate our method across six large-scale benchmark datasets. Empirical results show that our method achieves arbitrary artist-style and artwork-style extraction from a single artwork, and effectively avoids introducing the style image’s structural features. Our method improves the state-of-the-art deception rate from 58.9% to 67.2% and the average FID from 48.74 to 42.83.
Jingyu Wu, Lefan Hou, Zejian Li, Jun Liao 0001, Li Liu 0001, Lingyun Sun
AAAI1
2023 Cultural Self-Adaptive Multimodal Gesture Generation Based on Multiple Culture Gesture Dataset
abstract
Co-speech gesture generation is essential for multimodal chatbots and agents. Previous research extensively studies the relationship between text, audio, and gesture. Meanwhile, to enhance cross-culture communication, culture-specific gestures are crucial for chatbots to learn cultural differences and incorporate cultural cues. However, culture-specific gesture generation faces two challenges: lack of large-scale, high-quality gesture datasets that include diverse cultural groups, and lack of generalization across different cultures. Therefore, in this paper, we first introduce a Multiple Culture Gesture Dataset (MCGD), the largest freely available gesture dataset to date. It consists of ten different cultures, over 200 speakers, and 10,000 segmented sequences. We further propose a Cultural Self-adaptive Gesture Generation Network (CSGN) that takes multimodal relationships into consideration while generating gestures using a cascade architecture and learnable dynamic weight. The CSGN adaptively generates gestures with different cultural characteristics without the need to retrain a new network. It extracts cultural features from the multimodal inputs or a cultural style embedding space with a designated culture. We broadly evaluate our method across four large-scale benchmark datasets. Empirical results show that our method achieves multiple cultural gesture generation and improves comprehensiveness of multimodal inputs. Our method improves the state-of-the-art average FGD from 53.7 to 48.0 and culture deception rate (CDR) from 33.63% to 39.87%.
Jingyu Wu, Shi Chen 0005, Shuyu Gan, Chang-yuan Yang, Lingyun Sun
ACM Multimedia1
2022 Aggregate interactive learning for RGB-D salient object detection
Jingyu Wu, Fuming Sun, Fasheng Wang
Expert Syst. Appl.1
2022 Detection and Biomass Estimation of Phaeocystis globosa Blooms off Southern China From UAV-Based Hyperspectral Measurements
abstract
Phaeocystis globosa (P. globosa) is a unique causative species of harmful algal blooms, which can form gelatinous colonies. We, for the first time, used unmanned aerial vehicle (UAV) measurements to identify P. globosa blooms and to quantify the biomass. Based on in situ measured remote sensing reflectance (${R_{\mathrm{ rs}}}$), it is found that, for P. globosa blooms, the maximum of the second-derivative (${d\lambda ^{2}{R}_{\mathrm{ rs}}}$) of${R_{\mathrm{ rs}}(\lambda)}$in the 460–480-nm domain is beyond 466 nm. An analysis of the absorption properties from algal cultures suggested that this feature comes from the absorption of chlorophyll${c_{3}}$(Chl$-/{c_{3}}$) around 466 nm, a prominent feature of P. globosa. This position of$ {d\lambda ^{2}{R}_{\mathrm{ rs}}}$maximum was, thus, selected as the criterion for P. globosa identification. The spatial extent of P. globosa blooms in two bays off southern China was then mapped by applying the criterion to UAV-measured${R_{\mathrm{ rs}}}$. Twelve out of 16 UAV and in situ match-up stations were consistently identified as dominated by P. globosa, indicating the accuracy of 75%. Furthermore, using localized empirical models, chlorophyll a (Chl$-/{a}$) concentration and colony numbers of P. globosa were estimated from UAV-derived${R_{\mathrm{ rs}}}$, where P. globosa colonies were found in a range of ~3–37 gel matrix/L, indicating the occurrence of weak to moderate P. globosa blooms during the surveys. The promising results suggest a high potential for detection and quantification of P. globosa blooms in near-shore bays or harbors using UAV-based hyperspectral remote sensing, where conventional ocean color satellite remote sensing runs into difficulties.
Xue Li 0027, Shaoling Shang, ZhongPing Lee, Gong Lin, Yongnian Zhang, Jingyu Wu, Zhenjun Kang, Xiangxu Liu
IEEE Trans. Geosci. Remote. Sens.6
2021 Image Synthesis from Layout with Locality-Aware Mask Adaption
abstract
This paper is concerned with synthesizing images conditioned on a layout (a set of bounding boxes with object categories). Existing works construct a layout-mask-image pipeline. Object masks are generated separately and mapped to bounding boxes to form a whole semantic segmentation mask (layout-to-mask), with which a new image is generated (mask-to-image). However, overlapped boxes in layouts result in overlapped object masks, which reduces the mask clarity and causes confusion in image generation. We hypothesize the importance of generating clean and semantically clear semantic masks. The hypothesis is supported by the finding that the performance of state-of-the-art LostGAN decreases when input masks are tainted. Motivated by this hypothesis, we propose Locality-Aware Mask Adaption (LAMA) module to adapt overlapped or nearby object masks in the generation. Experimental results show our proposed model with LAMA outperforms existing approaches regarding visual fidelity and alignment with input layouts. On COCO-stuff in 256×256, our method improves the state-of-the-art FID score from 41.65 to 31.12 and the SceneFID from 22.00 to 18.64.
Zejian Li, Jingyu Wu, Immanuel Koh, Lingyun Sun
ICCV2