Ya Jing

dblp:55/3653 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-4179-8210ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Fine-Grained Alignment Supervision Matters in Vision-and-Language Navigation
abstract
The Vision-and-Language Navigation (VLN) task involves an agent navigating within 3D indoor environments based on provided instructions. Achieving cross-modal alignment presents one of the most critical challenges in VLN, as the predicted trajectory needs to precisely align with the given instruction. This paper focuses on addressing cross-modal alignment in VLN from a fine-grained perspective. Firstly, to address the issue of weak cross-modal alignment supervision arising from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset called Landmark-RxR. This dataset aims to offer precise, fine-grained supervision for VLN. Secondly, in order to comprehensively demonstrate the potential and advantage of the fine-grained data from Landmark-RxR, we explore the core components of the training process that depend on the characteristics of the training data. These components include data augmentation, training paradigm, reward shaping, and navigation loss design. Leveraging our fine-grained data, we carefully design methods for handling them and introduce a novel evaluation mechanism. The experimental results demonstrate that the fine-grained data can effectively improve the agent's cross-modal alignment ability.
Keji He, Yan Huang 0008, Ya Jing, Qi Wu 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Implicit Chain-of-Thought Reasoning via Task-Aware Latent Motion for transferable VLA
Ya Jing, Xinghang Li, Lifang Wu
Pattern Recognit.1
2026 Disco: Disentangled identity-action extraction and spatiotemporal context modeling for LLM-based identity-aware basketball video captioning
Zeyu Xi, Ya Jing, Haoying Sun, Lifang Wu
Pattern Recognit.2
2024 Vision-Language Foundation Models as Effective Robot Imitators
abstract
Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on robotics data. To this end, we derive a simple and novel vision-language manipulation framework, dubbed RoboFlamingo, built upon the open-source VLMs, OpenFlamingo. Unlike prior works, RoboFlamingo utilizes pre-trained VLMs for single-step vision-language comprehension, models sequential history information with an explicit policy head, and is slightly fine-tuned by imitation learning only on language-conditioned manipulation datasets. Such a decomposition provides RoboFlamingo the flexibility for open-loop control and deployment on low-performance platforms. By exceeding the state-of-the-art performance with a large margin on the tested benchmark, we show RoboFlamingo can be an effective and competitive alternative to adapt VLMs to robot control. Our extensive experimental results also reveal several interesting conclusions regarding the behavior of different pre-trained VLMs on manipulation tasks. We believe RoboFlamingo has the potential to be a cost-effective and easy-to-use solution for robotics manipulation, empowering everyone with the ability to fine-tune their own robotics policy. Our code will be made public upon acceptance.
Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Chilam Cheang, Ya Jing, Weinan Zhang 0001, Huaping Liu 0001, Tao Kong
ICLR8
2024 Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
abstract
Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative pre-training. We introduce GR-1, a GPT-style model designed for multi-task language-conditioned visual robot manipulation. GR-1 takes as inputs a language instruction, a sequence of observation images, and a sequence of robot states. It predicts robot actions as well as future images in an end-to-end manner. Thanks to a flexible design, GR-1 can be seamlessly finetuned on robot data after pre-trained on a large-scale video dataset. We perform extensive experiments on the challenging CALVIN benchmark and a real robot. On CALVIN benchmark, our method outperforms state-of-the-art baseline methods and improves the success rate from 88.9% to 94.9%. In the setting of zero-shot unseen scene generalization, GR-1 improves the success rate from 53.3% to 85.4%. In real robot experiments, GR-1 also outperforms baseline methods and shows strong potentials in generalization to unseen scenes and objects. We provide inaugural evidence that a unified GPT-style transformer, augmented with large-scale video generative pre-training, exhibits remarkable generalization to multi-task visual robot manipulation. Project page: https://GR1-Manipulation.github.io
Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Tao Kong
ICLR2
2024 Memory-Adaptive Vision-and-Language Navigation
Keji He, Ya Jing, Yan Huang 0008, Zhihe Lu, Dong An 0002, Liang Wang 0001
Pattern Recognit.2
2023 Learning to Explore Informative Trajectories and Samples for Embodied Perception
abstract
We are witnessing significant progress on perception models, specifically those trained on large-scale internet images. However, efficiently generalizing these perception models to unseen embodied tasks is insufficiently studied, which will help various relevant applications (e.g., home robots). Unlike static perception methods trained on pre-collected images, the embodied agent can move around in the environment and obtain images of objects from any viewpoints. Therefore, efficiently learning the exploration policy and collection method to gather informative training samples is the key to this task. To do this, we first build a 3D semantic distribution map to train the exploration policy self-supervised by introducing the semantic distribution disagreement and the semantic distribution uncertainty rewards. Note that the map is generated from multi-view observations and can weaken the impact of misidentification from an unfamiliar viewpoint. Our agent is then encouraged to explore the objects with different semantic distributions across viewpoints, or uncertain semantic distributions. With the explored informative trajectories, we propose to select hard samples on trajectories based on the semantic distribution uncertainty to reduce unnecessary observations that can be correctly identified. Experiments show that the perception model fine-tuned with our method outperforms the baselines trained with other exploration policies. Further, we demonstrate the robustness of our method in real-robot experiments.
Ya Jing, Tao Kong
ICRA1
2023 Exploring Visual Pre-training for Robot Manipulation: Datasets, Models and Methods
abstract
Visual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipes of visual pre-training for robot manipulation tasks are yet to be built. In this paper, we thoroughly investigate the effects of visual pre-training strategies on robot manipulation tasks from three fundamental perspectives: pre-training datasets, model architectures and training methods. Several significant experimental findings are provided that are beneficial for robot learning. Further, we propose a visual pre-training scheme for robot manipulation termed Vi-PRoM, which combines self-supervised learning and supervised learning. Concretely, the former employs contrastive learning to acquire underlying patterns from large-scale unlabeled data, while the latter aims learning visual semantics and temporal dynamics. Extensive experiments on robot manipulations in various simulation environments and the real robot demonstrate the superiority of the proposed scheme. Videos and more details can be found on https://explore-pretrain-robot.github.io.
Ya Jing, Xuelin Zhu, Xingbin Liu, Qie Sima, Taozheng Yang, Yunhai Feng, Tao Kong
IROS1
2023 MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation
abstract
In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate actions in an end-to-end manner, they often suffer from insufficient action accuracy and robustness against noise. On the other hand, classical control-based methods can enhance system robustness, but at the cost of extensive parameter tuning. To address these challenges, we present MOMA-Force, a visual-force imitation method that seamlessly combines representation learning for perception, imitation learning for complex motion generation, and admittance whole-body control for system robustness and controllability. MOMA-Force enables a mobile manipulator to learn multiple complex contact-rich tasks with high success rates and small contact forces. In a real household setting, our method outperforms baseline methods in terms of task success rates. Moreover, our method achieves smaller contact forces and smaller force variances compared to baseline methods without force imitation. Overall, we offer a promising approach for efficient and robust mobile manipulation in the real world. Videos and more details can be found on https://visual-force-imitation.github.io.
Taozheng Yang, Ya Jing, Jiafeng Xu, Kuankuan Sima, Guangzeng Chen, Qie Sima, Tao Kong
IROS2
2022 Towards Unifying Reference Expression Generation and Comprehension
abstract
Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a promising way to improve both. However, the problem of distinct inputs, as well as building connections between them in a single model, brings challenges to the design and training of the joint model. To address the problems, we propose a unified model for REG and REC, named UniRef. It unifies these two tasks with the carefully-designed Image-Region-Text Fusion layer (IRTF), which fuses the image, region and text via the image cross-attention and region cross-attention. Additionally, IRTF could generate pseudo input regions for the REC task to enable a uniform way for sharing the identical representation space across the REC and REG. We further propose Vision-conditioned Masked Language Modeling (VMLM) and Text-Conditioned Region Prediction (TRP) to pre-train UniRef model on multi-granular corpora. The VMLM and TRP are directly related to REG and REC, respectively, but could help each other. We conduct extensive experiments on three benchmark datasets, RefCOCO, RefCOCO+ and RefCOCOg. Experimental results show that our model outperforms previous state-of-the-art methods on both REG and REC.
Duo Zheng, Tao Kong, Ya Jing, Jiaan Wang, Xiaojie Wang 0006
EMNLP3
2021 Locate Then Segment: A Strong Pipeline for Referring Image Segmentation
abstract
Referring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature interaction mechanism to fuse the visual-linguistic features to directly generate the final segmentation mask without explicitly modeling the localization information of the referent instances. To tackle these problems, we view this task from another perspective by decoupling it into a "Locate-Then-Segment" (LTS) scheme. Given a language expression, people generally first perform attention to the corresponding target image regions, then generate a fine segmentation mask about the object based on its context. The LTS first extracts and fuses both visual and textual features to get a cross-modal representation, then applies a cross-model interaction on the visual-textual features to locate the referred object with position prior, and finally generates the segmentation result with a light-weight segmentation network. Our LTS is simple but surprisingly effective. On three popular benchmark datasets, the LTS outperforms all the previous state-of-the-arts methods by a large margin (e.g., +3.2% on RefCOCO+ and +3.4% on RefCOCOg). In addition, our model is more interpretable with explicitly locating the object, which is also proved by visualization experiments. We believe this framework is promising to serve as a strong baseline for referring image segmentation.
Ya Jing, Tao Kong, Wei Wang 0115, Liang Wang 0001, Lei Li 0005, Tieniu Tan
CVPR1
2021 Learning Aligned Image-Text Representations Using Graph Attentive Relational Network
abstract
Image-text matching aims to measure the similarities between images and textual descriptions, which has made great progress recently. The key to this cross-modal matching task is to build the latent semantic alignment between visual objects and words. Due to the widespread variations of sentence structures, it is very difficult to learn the latent semantic alignment using only global cross-modal features. Many previous methods attempt to learn the aligned image-text representations by the attention mechanism but generally ignore the relationships within textual descriptions which determine whether the words belong to the same visual object. In this paper, we propose a graph attentive relational network (GARN) to learn the aligned image-text representations by modeling the relationships between noun phrases in a text for the identity-aware image-text matching. In the GARN, we first decompose images and texts into regions and noun phrases, respectively. Then a skip graph neural network (skip-GNN) is proposed to learn effective textual representations which are a mixture of textual features and relational features. Finally, a graph attention network is further proposed to obtain the probabilities that the noun phrases belong to the image regions by modeling the relationships between noun phrases. We perform extensive experiments on the CUHK Person Description dataset (CUHK-PEDES), Caltech-UCSD Birds dataset (CUB), Oxford-102 Flowers dataset and Flickr30K dataset to verify the effectiveness of each component in our model. Experimental results show that our approach achieves the state-of-the-art results on these four benchmark datasets.
Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
IEEE Trans. Image Process.1
2020 Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search
abstract
Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric.
Ya Jing, Chenyang Si, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
AAAI1
2020 Cross-Modal Cross-Domain Moment Alignment Network for Person Search
abstract
Text-based person search has drawn increasing attention due to its wide applications in video surveillance. However, most of the existing models depend heavily on paired image-text data, which is very expensive to acquire. Moreover, they always face huge performance drop when directly exploiting them to new domains. To overcome this problem, we make the first attempt to adapt the model to new target domains in the absence of pairwise labels, which combines the challenges from both cross-modal (text-based) person search and cross-domain person search. Specially, we propose a moment alignment network (MAN) to solve the cross-modal cross-domain person search task in this paper. The idea is to learn three effective moment alignments including domain alignment (DA), cross-modal alignment (CA) and exemplar alignment (EA), which together can learn domain-invariant and semantic aligned cross-modal representations to improve model generalization. Extensive experiments are conducted on CUHK Person Description dataset (CUHK-PEDES) and Richly Annotated Pedestrian dataset (RAP). Experimental results show that our proposed model achieves the state-of-the-art performances on five transfer tasks.
Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
CVPR1
2020 Relational graph neural network for situation recognition
Ya Jing, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
Pattern Recognit.1
2020 Skeleton-based action recognition with hierarchical spatial reasoning and temporal stack learning network
Chenyang Si, Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
Pattern Recognit.2
2018 Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning
Chenyang Si, Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ECCV (1)2
2005 Iterative receivers for space-frequency block-coded OFDM systems with error propagation suppression
abstract
This paper considers the design of iterative receivers for space-frequency block-coded orthogonal frequency division multiplexing (SFBC-OFDM) systems in unknown wireless dispersive fading channel with outer channel coding. The iterative joint channel estimation and symbol detection algorithm is derived. A simple decision-error judgment criterion is proposed and an improved training scheme is applied in the system to suppress the error propagation. Both theoretical analysis and computer simulation show that this algorithm gives better BER performance and saves the overhead of OFDM systems
Ya Jing, Ming Chen 0001, Shixin Cheng, Haifeng Wang 0002
PIMRC1
2005 Subspace-based noise variance and SNR estimation for OFDM systems [mobile radio applications]
abstract
Noise variance and hence signal to noise ratio (SNR) estimates are very important for the channel quality control in communication systems. Noting that in mobile communications the multipath time delays are slowly varying in time, in this paper we derive a subspace-based estimation method for orthogonal frequency division multiplexing (OFDM) systems, which is based on an eigenvector decomposition of the estimated channel correlation matrix. Simulation results show that the proposed estimator can obtain accurate real time measurements of the noise variance and SNR after an observation interval of about 20 OFDM symbols for various fading channels.
Xiaodong Xu 0001, Ya Jing, Xiaohu You 0001
WCNC2