Jingru Yi

dblp:156/1720 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0001-9648-9389ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
abstract
With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens of minutes has become both feasible and increasingly prevalent. However, the inherently redundant nature of video sequences poses significant challenges for contemporary state-of-the-art models. These challenges stem from two primary aspects: 1) efficiently incorporating a larger number of frames within memory constraints, and 2) extracting discriminative information from the vast volume of input data. In this paper, we introduce a novel end-to-end schema for long-form video understanding, which includes an information-density-based adaptive video sampler (AVS) and an autoencoder-based spatiotemporal video compressor (SVC) integrated with a multimodal large language model (MLLM). Our proposed system offers two major advantages: it adaptively and effectively captures essential information from video sequences of varying durations, and it achieves high compression rates while preserving crucial discriminative information. The proposed framework demonstrates promising performance across various benchmarks, excelling in both long-form video understanding tasks and standard video understanding benchmarks. These results underscore the versatility and efficacy of our approach, particularly in managing the complexities of prolonged video sequences.
Jue Wang 0010, Jingru Yi, Zhaowei Cai, Xinyu Li 0003, Hao Yang 0043, Davide Modolo
WACV4
2024 Augment the Pairs: Semantics-Preserving Image-Caption Pair Augmentation for Grounding-Based Vision and Language Models
abstract
Grounding-based vision and language models have been successfully applied to low-level vision tasks, aiming to precisely locate objects referred in captions. The effectiveness of grounding representation learning heavily relies on the scale of the training dataset. Despite being a useful data enrichment strategy, data augmentation has received minimal attention in existing vision and language tasks as augmentation for image-caption pairs is non-trivial. In this study, we propose a robust phrase grounding model trained with text-conditioned and text-unconditioned data augmentations. Specifically, we apply text-conditioned color jittering and horizontal flipping to ensure semantic consistency between images and captions. To guarantee image-caption correspondence in the training samples, we modify the captions according to pre-defined keywords when applying horizontal flipping. Additionally, inspired by recent masked signal reconstruction, we propose to use pixel-level masking as a novel form of data augmentation. While we demonstrate our data augmentation method with MDETR framework, the proposed approach is applicable to common grounding-based vision and language tasks with other frameworks. Finally, we show that image encoder pretrained on large-scale image and language datasets (such as CLIP) can further improve the results. Through extensive experiments on three commonly applied datasets: Flickr30k, referring expressions and GQA, our method demonstrates advanced performance over the state-of-the-arts with various metrics. Code can be found in https://github.com/amzn/augment-the-pairs-wacv2024.
Jingru Yi, Burak Uzkent, Oana Ignat, Zili Li 0014, Amanmeet Garg, Linda Liu
WACV1
2023 Dynamic Inference with Grounding Based Vision and Language Models
abstract
Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks. On the other hand, there exists a large amount of computational redundancy in these large models which skips their run-time efficiency. To address this problem, we propose dynamic inference for grounding based vision and language models conditioned on the input image-text pair. We first design an approach to dynamically skip multihead self-attention and feed forward network layers across two backbones and multimodal network. Additionally, we propose dynamic token pruning and fusion for two backbones. In particular, we remove redundant tokens at different levels of the backbones and fuse the image tokens with the language tokens in an adaptive manner. To learn policies for dynamic inference, we train agents using reinforcement learning. In this direction, we replace the CNN backbone in a recent grounding-based vision and language model, MDETR, with a vision transformer and call it ViTMDETR. Then, we apply our dynamic inference method to ViTMDETR, called D-ViTDMETR, and perform experiments on image-language tasks. Our results show that we can improve the run-time efficiency of the state-of-the-art models MDETR and GLIP by up to ~ 50% on Referring Expression Comprehension and Segmentation, and VQA with only maximum ~ 0.3% accuracy drop.
Burak Uzkent, Amanmeet Garg, Keval Doshi, Jingru Yi, Mohamed Omar
CVPR5
2022 Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images
Qilong Zhangli, Jingru Yi, Di Liu 0003, Xiaoxiao He, Zhaoyang Xia, Ligong Han, Yunhe Gao, Song Wen 0001, Haiming Tang, He Wang 0016, Mu Zhou, Dimitris N. Metaxas
MICCAI (4)2
2021 Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors
abstract
Oriented object detection in aerial images is a challenging task as the objects in aerial images are displayed in arbitrary directions and are usually densely packed. Cur-rent oriented object detection methods mainly rely on two-stage anchor-based detectors. However, the anchor-based detectors typically suffer from a severe imbalance issue be-tween the positive and negative anchor boxes. To address this issue, in this work we extend the horizontal keypoint-based object detector to the oriented object detection task. In particular, we first detect the center keypoints of the objects, based on which we then regress the box boundary-aware vectors (BBAVectors) to capture the oriented bounding boxes. The box boundary-aware vectors are distributed in the four quadrants of a Cartesian coordinate system for all arbitrarily oriented objects. To relieve the difficulty of learning the vectors in the corner cases, we further classify the oriented bounding boxes into horizontal and rotational bounding boxes. In the experiment, we show that learning the box boundary-aware vectors is superior to directly predicting the width, height, and angle of an oriented bounding box, as adopted in the baseline method. Besides, the proposed method competes favorably with state-of-the-art methods. Code is available at https://github.com/yijingru/BBAVectors-Oriented-Object-Detection.
Jingru Yi, Pengxiang Wu, Bo Liu 0005, Qiaoying Huang, Dimitris N. Metaxas
WACV1
2021 Dynamic MRI reconstruction with end-to-end motion-guided network
Qiaoying Huang, Yikun Xian, Dong Yang 0005, Jingru Yi, Pengxiang Wu, Dimitris N. Metaxas
Medical Image Anal.5
2021 Object-Guided Instance Segmentation With Auxiliary Feature Refinement for Biological Images
abstract
Instance segmentation is of great importance for many biological applications, such as study of neural cell interactions, plant phenotyping, and quantitatively measuring how cells react to drug treatment. In this paper, we propose a novel box-based instance segmentation method. Box-based instance segmentation methods capture objects via bounding boxes and then perform individual segmentation within each bounding box region. However, existing methods can hardly differentiate the target from its neighboring objects within the same bounding box region due to their similar textures and low-contrast boundaries. To deal with this problem, in this paper, we propose an object-guided instance segmentation method. Our method first detects the center points of the objects, from which the bounding box parameters are then predicted. To perform segmentation, an object-guided coarse-to-fine segmentation branch is built along with the detection branch. The segmentation branch reuses the object features as guidance to separate target object from the neighboring ones within the same bounding box region. To further improve the segmentation quality, we design an auxiliary feature refinement module that densely samples and refines point-wise features in the boundary regions. Experimental results on three biological image datasets demonstrate the advantages of our method. The code will be available at https://github.com/yijingru/ObjGuided-Instance-Segmentation.
Jingru Yi, Pengxiang Wu, Bo Liu 0005, Qiaoying Huang, Lianyi Han, Wei Fan 0001, Daniel J. Hoeppner, Dimitris N. Metaxas
IEEE Trans. Medical Imaging1
2020 Object-Guided Instance Segmentation for Biological Images
abstract
Instance segmentation of biological images is essential for studying object behaviors and properties. The challenges, such as clustering, occlusion, and adhesion problems of the objects, make instance segmentation a non-trivial task. Current box-free instance segmentation methods typically rely on local pixel-level information. Due to a lack of global object view, these methods are prone to over- or under-segmentation. On the contrary, the box-based instance segmentation methods incorporate object detection into the segmentation, performing better in identifying the individual instances. In this paper, we propose a new box-based instance segmentation method. Mainly, we locate the object bounding boxes from their center points. The object features are subsequently reused in the segmentation branch as a guide to separate the clustered instances within an RoI patch. Along with the instance normalization, the model is able to recover the target object distribution and suppress the distribution of neighboring attached objects. Consequently, the proposed model performs excellently in segmenting the clustered objects while retaining the target object details. The proposed method achieves state-of-the-art performances on three biological datasets: cell nuclei, plant phenotyping dataset, and neural cells.
Jingru Yi, Pengxiang Wu, Bo Liu 0005, Daniel J. Hoeppner, Dimitris N. Metaxas, Lianyi Han, Wei Fan 0001
AAAI1
2020 Weakly Supervised Deep Nuclei Segmentation Using Partial Points Annotation in Histopathology Images
abstract
Nuclei segmentation is a fundamental task in histopathology image analysis. Typically, such segmentation tasks require significant effort to manually generate accurate pixel-wise annotations for fully supervised training. To alleviate such tedious and manual effort, in this paper we propose a novel weakly supervised segmentation framework based on partial points annotation, i.e., only a small portion of nuclei locations in each image are labeled. The framework consists of two learning stages. In the first stage, we design a semi-supervised strategy to learn a detection model from partially labeled nuclei locations. Specifically, an extended Gaussian mask is designed to train an initial model with partially labeled data. Then, self-training with background propagation is proposed to make use of the unlabeled regions to boost nuclei detection and suppress false positives. In the second stage, a segmentation model is trained from the detected nuclei locations in a weakly-supervised fashion. Two types of coarse labels with complementary information are derived from the detected points and are then utilized to train a deep neural network. The fully-connected conditional random field loss is utilized in training to further refine the model without introducing extra computational complexity during inference. The proposed method is extensively evaluated on two nuclei segmentation datasets. The experimental results demonstrate that our method can achieve competitive performance compared to the fully supervised counterpart and the state-of-the-art methods while requiring significantly less annotation effort.
Pengxiang Wu, Qiaoying Huang, Jingru Yi, Zhennan Yan, Kang Li 0004, Gregory M. Riedlinger, Subhajyoti De, Shaoting Zhang 0001, Dimitris N. Metaxas
IEEE Trans. Medical Imaging4
2019 Point Cloud Processing via Recurrent Set Encoding
abstract
We present a new permutation-invariant network for 3D point cloud processing. Our network is composed of a recurrent set encoder and a convolutional feature aggregator. Given an unordered point set, the encoder firstly partitions its ambient space into parallel beams. Points within each beam are then modeled as a sequence and encoded into subregional geometric features by a shared recurrent neural network (RNN). The spatial layout of the beams is regular, and this allows the beam features to be further fed into an efficient 2D convolutional neural network (CNN) for hierarchical feature aggregation. Our network is effective at spatial feature learning, and competes favorably with the state-of-the-arts (SOTAs) on a number of benchmarks. Meanwhile, it is significantly more efficient compared to the SOTAs.
Pengxiang Wu, Chao Chen 0012, Jingru Yi, Dimitris N. Metaxas
AAAI3
2019 Multi-scale Cell Instance Segmentation with Keypoint Graph Based Bounding Boxes
Jingru Yi, Pengxiang Wu, Qiaoying Huang, Bo Liu 0005, Daniel J. Hoeppner, Dimitris N. Metaxas
MICCAI (1)1
2019 ASSD: Attentive single shot multibox detector
Jingru Yi, Pengxiang Wu, Dimitris N. Metaxas
Comput. Vis. Image Underst.1
2019 Attentive neural cell instance segmentation
Jingru Yi, Pengxiang Wu, Menglin Jiang, Qiaoying Huang, Daniel J. Hoeppner, Dimitris N. Metaxas
Medical Image Anal.1
2017 Towards large-scale MR thigh image analysis via an integrated quantification framework
Chaowei Tan, Kang Li 0004, Zhennan Yan, Jingru Yi, Pengxiang Wu, Hui Jing Yu, Klaus Engelke, Dimitris N. Metaxas
Neurocomputing4