Yuyu Jia

dblp:301/6390 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reasoning via Implicit Self-supervised Emergence for Instruction Segmentation
abstract
We challenge the assumption that complex instruction-guided segmentation tasks necessitate equally complex and explicit supervision. This paper introduces RISE (Reasoning via Implicit Self-supervised Emergence), a framework that learns intricate compositional reasoning, spanning spatial relations to world knowledge, without a single ground-truth mask. To achieve this, RISE employs reinforcement learning with GRPO guided by a single, strikingly simple reward: the semantic alignment score between the textual instruction and the predicted image region. Our primary discovery is the implicit emergence of a high-quality chain-of-thought process from this minimalist signal. Within a structured format, the model autonomously learns to understand instructions by accessing its latent knowledge, inferring spatial relationships—capabilities inherent in its architecture but unlocked by our simple objective. Remarkably, our emergent reasoning yields highly competitive results: RISE achieves 58.7 gIoU on the ReasonSeg benchmark, on par with methods using geometric rewards. Furthermore, we show extreme data efficiency: a variant trained on only 2,000 ImageNet-label pairs establishes a new state-of-the-art for annotation-free referring segmentation with 79.6 cIoU on RefCOCO.
Lichang Yang, Yuyu Jia, Junyu Gao 0001, Weiping Ni, Junzheng Wu, Qi Wang 0009
AAAI3
2026 Enhancing streamflow forecasting using an LSTM hybrid model with lightweight frequency-domain feature learning
Yubo Jia, Xiaoling Su, Te Zhang, Haijiang Wu, Yuyu Jia
Expert Syst. Appl.5
2025 High-Order Anchor Graph-Based Clustering for Efficient Structured Proximity Matrix Learning
Zihua Zhao, Yuyu Jia, Fangyuan Xie, Rong Wang 0001
ADMA (4)2
2025 Oscillation Suppression of Acoustic Trapping: A Disturbance Observer-based Approach
abstract
Acoustic tweezers have been a valuable tool across various fields, from nano-microfabrication to biology. Their unique characteristics enable three-dimensional particle manipulation, where acoustic trapping serves as a fundamental requirement. However, traditional methods struggle to maintain steady particle positioning due to nonlinear forces and complex dynamic coupling effects. As a result, particle oscillations are inevitable and cannot be effectively compensated by predesigned acoustic trapping. To address these challenges, this study introduces a novel visual feedback control approach that dynamically adjusts the acoustic field distribution to mitigate oscillations along the z-axis of the acoustic trapping. A binocular microscopic vision system is employed for precise particle localization, while a disturbance observer estimates the effects of strong nonlinearity and uncertainties of the acoustic trapping. The proposed methodology is validated through simulations and experiments, demonstrating a significant reduction in z-axis oscillations from 1.33× wavelength to within 0.03× wavelength. This advancement marks a step forward in achieving precise and complex acoustic manipulation using traveling-wave acoustic tweezers.
Yuyu Jia, Yizhou Gong, Zhenhuan Sun, Yalin Shi, Yang Wang 0063, Song Liu 0003
IROS1
2025 Dual-View Classifier Evolution for Generalized Remote Sensing Few-Shot Segmentation
abstract
Advancements in few-shot segmentation (FSS) for remote sensing images have significantly improved the ability to binarization parse novel classes using only a few supports. Generalized few-shot segmentation (GFSS), a challenging and practical task, has recently attracted research attention. It involves recognizing base and novel classes while segmenting multiple categories in a query. Most GFSS methods adopt a two-stage approach: base classifier training and novel classifier registering. However, they encounter two key challenges: the data scale disparity between base and novel classes and significant intraclass variation in remote sensing images. In this article, we present a dual-view classifier evolution (DiCE) method. Our approach utilizes the well-trained base classifier to allocate attention within the novel classifier, effectively addressing the disparities between the two. Simultaneously, it fosters context-driven interactions between the query and the classifier, tailoring sample-specific classifiers to mitigate intraclass variations. Furthermore, we propose a binocular hybrid training (BHT) mechanism that integrates normal base training with episodic training, endowing the model with the ability to adapt to few-shot tasks. Extensive experiments on the iSAID-$5^{i}$dataset demonstrate the superior performance of DiCE.
Yuyu Jia, Junyu Gao 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2025 Generalized Few-Shot Semantic Segmentation for Remote Sensing Images
abstract
Few-shot segmentation (FSS) techniques enhance pixel-level interpretation of unseen classes while reducing reliance on extensive labeled data. However, FSS still faces significant limitations in practical applications: it is restricted to segmenting novel classes and relies on manually constructed support-query pairs during inference. We are the first to introduce the generalized few-shot segmentation (GFSS) task to remote sensing analysis. It enables simultaneous segmentation of base and novel classes without manual prior interventions. The most intuitive construction is to extend a pretrained base classifier with a novel classifier. Nevertheless, since the latter is aggregated from a limited number of supports while the former is trained on abundant data, this disparity inevitably introduces a base class bias, leading to suboptimal segmentation results. This article proposes a background-aware self-mining prototype learning (BSPL) strategy to address the issues above. Specifically, we design a dynamic prototype update mechanism during training to enhance the model’s adaptability in few-shot scenarios and thereby mitigate the base class bias. Considering the intraclass variation and complex background elements in remote sensing images, we customize segmentation guidance for each query through background-aware self-mining, achieving more precise segmentation performance. Compared to peer algorithms, extensive experiments demonstrate that BSPL achieves the best overall segmentation performance for both base and novel classes, indicating its significant practicality.
Yuyu Jia, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2025 Entity-Guided Attention Twisting Network for Referring Remote Sensing Image Segmentation
abstract
Referring Remote Sensing Image Segmentation (RRSIS) aims to establish pixel-level interpretation of specific regions queried by textual expressions, bridging textual semantics and intelligent analysis of remote sensing imagery. In contrast to natural scenarios, the intricate backgrounds in remote sensing scenarios result in low target-background contrast, often leading to semantic dispersion in segmented regions. Furthermore, conventional cross-attention-based referring image segmentation (RIS) methods struggle to bridge the modal gap, hindering fine-grained alignment between linguistic descriptions and geographical features. To overcome these challenges, we present a pioneering Entity-Guided Attention Twisting Network (Enti-TwistNet) for RRSIS. Our framework first introduces a SAM-inspired Entity Guidance (SEG) module that extracts spatially constrained entity prompts through a self-reasoning mask generation mechanism, constructing a comprehensive entity-visual-text tri-modal information cube. Subsequently, during cross-modal interaction, we propose a Dual-phase Attention-Twisting (DAT) mechanism: (1) initially sequential channel-wise scanning to facilitate cross-modal semantic propagation; (2) Subsequently, twist attention to the spatial dimension, integrating entity guidance to enhance the representation of irregular geographic boundaries. Extensive experiments on two widely used benchmarks, RefSegRS and RRSIS-D, demonstrate that Enti-TwistNet achieves significant performance improvements over existing state-of-the-art models.
Yuyu Jia, Junyu Gao 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2025 A Semantic-Guided Framework for Few-Shot Remote Sensing Object Detection
abstract
Few-Shot Object Detection (FSOD) aims to recognize novel class targets using limited annotated data. Conventional approaches rely on extensive base class training, followed by fine-tuning where few instances from both base and novel classes are sampled for each category. Although they demonstrate remarkable performance in natural image domains, the specificity of remote sensing scenarios poses two critical challenges for FSOD: 1) The morphological differences between remote sensing images and natural images are significant, leading to a loss of structural priors in the Region Proposal Network (RPN). This makes it difficult for structural priors pretrained on natural images to generalize to remote sensing images, especially for novel class with scarce data; 2) Differences in imaging conditions lead to appearance variations among similar objects, leading to sparse visual features are insufficient to represent the common semantic structure of the entire class. To solve problems above, we introduce an innovative framework named ST-FSOD. Primarily, we introduce the SA-RPN module, which leverages efficient pixel association capability to generate high-quality foreground object proposals. Subsequently, through a text guiding learner module (TGL), we use textual labels of each category to generate image-agnostic text-guided prototypes. The enhanced text prototypes are fused with visual features to complement the sparse visual features. Extensive experiments conducted on the DIOR, NWPU VHR-10 and RSOD benchmarks demonstrate that the proposed method consistently surpasses strong baselines and achieves superior performance compared to previous state-of-the-art (SOTA) approaches. Our project will be open-sourced soon on https://github.com/wdcjhyy/ST-FSOD.
Chenchen Sun, Yuyu Jia, Qiang Li 0042, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.2
2025 Embedding Generalized Semantic Knowledge Into Few-Shot Remote Sensing Segmentation
abstract
Few-shot segmentation (FSS) for remote sensing (RS) imagery leverages supporting information from limited annotated samples to achieve query segmentation of novel classes. Previous efforts are dedicated to mining segmentation-guiding visual cues from a constrained set of support samples. However, they still struggle to address the pronounced intra-class differences in RS images, as sparse visual cues make it challenging to establish robust class-specific representations. In this article, we propose a holistic semantic embedding (HSE) approach that effectively harnesses general semantic knowledge, i.e., class description (CD) embeddings. Instead of the naive combination of CD embeddings and visual features for segmentation decoding, we investigate embedding the general semantic knowledge during the feature extraction stage. Specifically, in HSE, a spatial dense interaction (SDI) module allows the interaction of visual support features with CD embeddings along the spatial dimension via self-attention. Furthermore, a global content modulation (GCM) module efficiently augments the global information of the target category in both support and query features, thanks to the transformative fusion of visual features and CD embeddings. These two components holistically synergize CD embeddings and visual cues, constructing a robust class-specific representation. Through extensive experiments on the standard FSS benchmark, the proposed HSE approach demonstrates superior performance compared to peer work, setting a new state-of-the-art.
Qi Wang 0009, Yuyu Jia, Wei Huang 0068, Junyu Gao 0001, Qiang Li 0042
IEEE Trans. Geosci. Remote. Sens.2
2025 Like Humans to Few-Shot Learning Through Knowledge Permeation of Visual and Language
abstract
Few-shot learning aims to generalize the recognizer from seen categories to an entirely novel scenario. With only a few support samples, several advanced methods initially introduce class names as prior knowledge for identifying novel classes. However, obstacles still impede achieving a comprehensive understanding of how to harness the mutual advantages of visual and textual knowledge. In this paper, we set out to fill this gap via a coherent Bidirectional Knowledge Permeation strategy called BiKop, which is grounded in human intuition: a class name description offers a moregeneralrepresentation, whereas an image captures thespecificityof individuals. BiKop primarily establishes a hierarchical joint general-specific representation through bidirectional knowledge permeation. On the other hand, considering the bias of joint representation towards the base set, we disentangle base-class-relevant semantics during training, thereby alleviating the suppression of potential novel-class-relevant information. Experiments on four challenging benchmarks demonstrate the remarkable superiority of BiKop, particularly outperforming previous methods by a substantial margin in the 1-shot setting (improving the accuracy by 7.58% onminiImageNet).
Yuyu Jia, Junyu Gao 0001, Qiang Li 0042, Qi Wang 0009
IEEE Trans. Multim.1
2023 Noncontact Particle Manipulation on Water Surface with Ultrasonic Phased Array System and Microscopic Vision
abstract
Noncontact particle manipulation (NPM) shows great application potential than its conventional counterpart particularly in terms of non-invasiveness, and thus has significantly extended robotic manipulation capacity into bio- medical engineering, material science, etc. As NPM by means of electric, magnetic, and optical field has successfully demonstrated powerful strength in both academia and industry, NPM boosted by acoustic field, however, still faces staggering challenges. It is indeed in the very recent years that controllable dynamic airborne or waterborne acoustic field modulation technology emerged in academia. In this paper, we report our latest research regarding dexterous and dynamic noncontact micro-particle manipulation on water surface effected by acoustic field in terms of automated trapping, closed-loop positioning, and real-time motion planning, which can be applied to scenarios such as parallel 3D printing, cell assembly, etc. The main contribution of this work is we demonstrated the feasibility of objective-oriented and fully automated acoustic manipulation of micro-particle in precision scale based on robotic approach in 2D plane. Experiment results showed that the repetitive positioning accuracy can reach as high as 16 μm, which is essentially the pixel scale factor.
Yexin Zhang, Jiaqi Li 0029, Yuyu Jia, Teng Li 0017, Yang Wang 0063, David C. Jeong, Hu Su, Song Liu 0003
ICRA3
2023 Exploring Hard Samples in Multiview for Few-Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification is of high practical value in real situations where data are scarce and annotated costly. The few-shot learner needs to identify new categories with limited examples, and the core issue of this assignment is how to prompt the model to learn transferable knowledge from a large-scale base dataset. Although current approaches based on transfer learning or meta-learning have achieved significant performance on this task, there are still two problems to be addressed: (i) as an essential characteristic of remote sensing images, spatial rotation insensitivity surprisingly remains largely unexplored; (ii) the high distribution uncertainty of hard samples reduces the discriminative power of the model decision boundary. Stimulated by these, we propose a corresponding end-to-end framework termed a Hard Sample Learning (HSL) and Multi-view Integration (MI) Network (HSL-MINet). First, the MI module contains a pretext task introduced to guide the knowledge transfer, and a multiview-attention mechanism used to extract correlational information across different rotation views of images. Second, aiming at increasing the discrimination of the model decision boundary, the HSL module is designed to evaluate and select hard samples via a class-wise adaptive threshold strategy, and then decrease the uncertainty of their feature distributions by a devised triplet loss. Extensive evaluations on NWPU-RESISC45, WHU-RS19, and UCM datasets show that the effectiveness of our HSL-MINet surpasses the former state-of-the-art approaches.
Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2023 Holistic Mutual Representation Enhancement for Few-Shot Remote Sensing Segmentation
abstract
Few-shot segmentation endeavors to utilize a minimal amount of annotated samples (support) to guide the segmentation of unseen objects (query). Previous techniques primarily employ asupport-to-queryparadigm, neglecting to sufficiently leverage the mutual representation between query and support images, which leaves models suffering from intra-class variations and background interference in remote sensing images. This paper proposes a Holistic Mutual Representation Enhancement (HMRE) method to bridge these gaps. First, a Dual Activation (DA) module is devised to establish information symmetry between the two branches and forms the foundation for mutual representation enhancement. Subsequently, the holistic mutual enhancement is jointly constructed by the Global Semantic (GS) and Spatial Dense (SD) mutual enhancement modules. In the prediction stage for segmentation, we integrate the enhanced mutual representation into the Mutual-Fusion Decoder to activate the homologous object regions bidirectionally. To expedite the replication of investigation in this task, we further create a corresponding benchmark Flood-3i. The whole dataset is attainable at https://drive.google.com/drive/folders/1FMAKf2sszoFKjq0UrUmSLnJDbwQSpfxR. Extensive experiments on two benchmarks iSAID-5i and Flood-3i demonstrate the superiority of our proposed method, which also sets a new state-of-the-art.
Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2021 Simultaneous Precision Assembly of Multiple Objects through Coordinated Micro-robot Manipulation
abstract
Simultaneous assembly of multiple objects is a key technology to form solid connections among objects to get compact structures in precision assembly and micro-assembly. Dramatically different from traditional assembly of two objects, the interaction among multiple objects is more complicated on analysis and control. During simultaneous assembly of multiple objects, there are multiple mutually effected contact surfaces, and multiple force sensors are needed to perceive the interaction status. In this paper, a coordinated micro-robot manipulation strategy is proposed for simultaneous assembly problem, which is based on microscopic vision and force information. Taking simultaneous assembly of three objects as an instance, the proposed method is well articulated, including calibration of assembly system, force analysis for each contacting surface, and insertion control strategy for assembly process. The proposed method is applicable also to case with more objects. Experiment results demonstrate effectiveness of the proposed method.
Song Liu 0003, Yuyu Jia, Youfu Li 0001, Yao Guo 0002, Haojian Lu
ICRA2
2021 Wireless Energy Transfer in Extra-Large Massive MIMO Rician Channels
abstract
In application scenarios such as Internet of Things, a large number of energy receivers (ERs) exist and line-of-sight (LOS) propagation could be common. Considering this, we investigate wireless energy transfer (WET) in extra-large massive MIMO Rician channels. We derive analytical expressions of the received net energy for different schemes, including 1) training-based WET, where the ER sends beacon signal for channel training and the energy transmitter (ET) uses the channel estimate for energy beamforming, 2) LOS beamforming, where the ET transmits to the LOS direction of the ER, and 3) energy harvesting, which allows an ER to harvest the training energy from the other ERs. We derive a path loss threshold for switching between training and LOS beamforming-based WET. We further show that the WET scheme selection of one ER is not affected by the other ERs, and the energy harvested from training is minimal in practice. With these insights, we propose an algorithm for the multi-ER scenario, which minimizes the power consumption by iteratively updating the WET scheme selection and power allocation for all ERs. Simulations show that the proposed algorithm achieves near-optimal performance as compared to exhaustive searching, while with much lower implementation complexity.
Jue Wang 0006, Ye Li 0004, Yuyu Jia, Jun Zhang 0023, Shi Jin 0002, Tony Q. S. Quek
IEEE Trans. Wirel. Commun.3