Yuqi Ji

dblp:278/1746 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 26% Transfer learning and domain adaptation · 26% Representation and self-supervised learning · 26%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-set object detection
1.012026
Toward Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining · IEEE Trans. Image Process. 2026
Machine learning › Representation and self-supervised learning
prototype learning
1.012026
Toward Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining · IEEE Trans. Image Process. 2026
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
1.012026
Toward Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining · IEEE Trans. Image Process. 2026
Computer vision › Vision and language
multimodal fusion
0.612022
Modality Eigen-Encodings Are Keys to Open Modality Informative Containers · ACM Multimedia 2022
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.212022
Modality Eigen-Encodings Are Keys to Open Modality Informative Containers · ACM Multimedia 2022
Computer vision › Vision and language
visual question answering
0.212022
Modality Eigen-Encodings Are Keys to Open Modality Informative Containers · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

memory bank · 1.0knowledge mining · 1.0clustering · 1.0multi-scale alignment · 0.6modality eigen-encoding · 0.6
YearPublicationVenuePosition
2026 DINO-PCB: Two-stage vision foundation model pretraining and distillation for real-time circuit-board defect detection
Junjie Ke, Lihuo He, Jing Zhang 0037, Yuqi Ji, Hui Chen 0013, Jie Li 0001, Sicheng Zhao, Guiguang Ding, Xinbo Gao 0001
Pattern Recognit.5
2026 Toward Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining
abstract
Existing object detection methods struggle to generalize across increasingly data domains while simultaneously adapting to the emergence of novel categories. To tackle this challenge, adaptive open-set object detection (AOOD) has been introduced, which employs supervised training on base categories within the source domain while enabling unsupervised adaptation to both base and novel categories in the target domain. However, existing AOOD approaches are still hindered by several limitations, including insufficient cross-domain feature representation, inter-category ambiguity in novel classes, and inherent feature bias toward the source domain. To overcome these issues, this paper proposes a category-level collaboration knowledge mining strategy designed to comprehensively exploit both inter-class and intra-class feature relationships across domains. Specifically, a clustering-based memory bank (CMB) is initially constructed to aggregate class prototype features, class auxiliary features, and intra-class disparity features, thereby embedding rich category-level knowledge into a unified memory structure. The CMB is iteratively updated through unsupervised clustering, which facilitates the modeling of intra-category relationships and enhances its capacity for cross-domain knowledge representation. Subsequently, a base-to-novel selection metric (BNSM) is designed to identify features corresponding to novel categories within the source domain by regulating the relationships between the novel categories and each base category. The selected features are then leveraged to initialize the object detector for the classification of novel categories. Finally, an adaptive feature assignment (AFA) strategy is introduced to transfer the learned category-level knowledge to the target domain, enabling the assignment of category labels to features. The memory bank is updated asynchronously with these assigned features to mitigate source domain bias. Extensive experiments conducted on diverse domain datasets demonstrate that the proposed method consistently outperforms state-of-the-art AOOD approaches, achieving performance gains of 1.1 to 5.5 mAP. Code is available at https://github.com/Jandsome/CCKM.
Yuqi Ji, Junjie Ke, Lihuo He, Lizhi Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 A Two-Stage Payload Dynamic Parameter Identification Method for Interactive Industrial Robots With Large Components
abstract
Taking human-robot collaborative assembly as an example, the methods based on contact forces can improve the assembly efficiency of industrial robots with large components in industrial manufacturing. However, due to the large size, high payload, assembly accuracy and dynamic changes in grip position, accurately estimating the contact forces between the payload and the operator becomes challenging when handling these large components. In this paper, a two-stage method is proposed for payload dynamic parameter identification. The parameter identification equation in the sensor coordinate system is initially established. Furthermore, the identification model of recursive restricted total least squares (RRTLS) based on total least squares (TLS) is constructed to achieve low-consumption online identification. According to the assembly requirements and payload characteristics, the posture coordinate system is designed for safety, including the feasible workspace for the robot. Subsequently, the static identification postures and dynamic excitation trajectory are planned to obtain static values and dynamic inertial parameters. In the end, a high-payload human-robot collaborative assembly system is built to validate the proposed method. Experimental results show that compared with the existing methods, the proposed approach can effectively identify and compensate the payload, leading to more accurate external force sensing.
Mingxuan Liu 0003, Pengcheng Li 0020, Jinjun Duan, Lunqian Liu, Ye Shen, Yuqi Ji
IEEE Trans Autom. Sci. Eng.7
2023 Interactive Real-Time Monitoring and Information Traceability for Complex Aircraft Assembly Field Based on Digital Twin
abstract
High-quality aircraft assembly is critical in aircraft manufacturing. To meet increasing quality demands, the aircraft assembly process needs to be monitored in real time. In recent years, digital twins have attracted increasing interest in different applications. However, due to the heterogeneous data and multiple protocols in the assembly field, the large amount of data needs to satisfy high throughput, low latency, and low storage cost requirements without interrupting the assembly process, which introduces challenges to real-time monitoring systems. In this article, a real-time monitoring system based on the digital twin that provides interactive services for users based on multisource, multidimensional data, and sophisticated analyses is designed. Moreover, we introduce a time series database (TSDB) for dynamic data storage and SQLServer for static data storage. Time series data can be regarded as a standard type of organized data that are convenient for hierarchical representations and constructing multidimensional data cubes for information traceability. The storage cost is significantly reduced because the compression rate reaches 3.8%. Furthermore, the powerful and comprehensive analysis engine of the TSDB ensures that the monitoring system has real-time responses of less than 4 ms. In addition, a flexible and scalable system that collects multidimensional data without affecting the existing workflow is proposed. We perform various experiments to verify the feasibility and effectiveness of assembling large-scale aircraft components. This study provides a solid data foundation for accelerating the pace of intelligent aircraft manufacturing.
Wei Liu 0037, Yang Zhang 0011, Changyong Gao, Yuqi Ji
IEEE Trans. Ind. Informatics7
2022 Modality Eigen-Encodings Are Keys to Open Modality Informative Containers
abstract
Vision-Language fusion relies heavily on precise cross-modal information synergy. Nevertheless, modality divergence makes mutual description with the other modality extremely difficult. Despite various attempts to tap into semantic unity in vision and language, most existing approaches utilize modality-specific features via the high-dimensionality tensors as the smallest unit of information, limiting the interactivity of multi-modal fine-grained fusion. Furthermore, in previous works, cross-modal interaction is commonly depicted by the similarity between semantically insufficient global features. Differently, we propose a novel scheme for multi-modal fusion named Vision Language Interaction (VLI). To represent more fine-grained and flexible information of the modality, we consider high-dimensional features as containers of modality-specific information, while homogeneous semantic information between heterogeneous modalities is the key stored in the containers. We first construct information containers via multi-scale alignment and then utilize modality eigen-encodings to take out the homogeneous semantics on the vector level. Finally, we iteratively embed the eigen-encodings of one modality into the eigen-encodings of the other modality to perform cross-modal semantic interaction. After embeddings interaction, vision and language information can break the existing representation bottleneck through the representation level of granularity never achieved in previous work. Extensive experimental results on vision-language tasks validate the effectiveness of VLI. On the three benchmarks of Referring Expression Comprehension (REC), Referring Expression Segmentation (RES), and Visual Question Answering (VQA), VLI significantly outperforms the existing state-of-the-art methods.
Yuqi Ji
ACM Multimedia2