Dong Jing

dblp:206/3646 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 32% Learning paradigms · 18% Representation and self-supervised learning · 14%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 85% Image and video processing · 15%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
multi-task learning
1.012026
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026
Machine learning › Optimization for machine learning › multi-task optimization
task balancing
1.012026
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026
Machine learning › Learning paradigms › multi-task learning
task weighting
1.012026
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning
1.012026
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval · ICCV 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval · AAAI 2025
Computer vision › Vision and language
vision-language model
0.912025
Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval · AAAI 2025
Information retrieval › image retrieval
composed image retrieval
0.912025
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval · ICCV 2025
Information retrieval
cross-modal retrieval
0.912025
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval · ICCV 2025
Information retrieval › image retrieval › composed image retrieval
zero-shot composed image retrieval
0.912025
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval · ICCV 2025
Multimedia analysis and retrieval › image retrieval
composed image retrieval
0.912025
Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval · AAAI 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding · NeurIPS 2024
Machine learning › Representation and self-supervised learning › contrastive learning › dense contrastive learning
region-level contrastive learning
0.812024
FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding · NeurIPS 2024
Computer vision › Vision and language
vision-language pretraining
0.812024
FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding · NeurIPS 2024
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
light field saliency detection
0.512021
Occlusion-aware Bi-directional Guided Network for Light Field Salient Object Detection · ACM Multimedia 2021
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.512021
Occlusion-aware Bi-directional Guided Network for Light Field Salient Object Detection · ACM Multimedia 2021
Information retrieval
image retrieval
0.312025
Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval · AAAI 2025
Information retrieval › query understanding
intent-aware retrieval
0.312025
Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval · AAAI 2025
Computer vision › Segmentation and scene understanding
dense prediction
0.212024
FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

prompt learning · 2.6instruction tuning · 2.6vision-language pre-training · 1.7multi-scale reasoning · 1.7chain-of-thought · 1.7validation-performance-based task balancing · 1.0bi-directional guiding flow · 1.0vision-language pretraining · 0.9self-distillation · 0.8region-text pairs · 0.8contrastive learning · 0.8occlusion extraction module · 0.5epipolar plane image · 0.5
YearPublicationVenuePosition
2026 Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty
abstract
Visual instruction tuning is a key training stage of large multimodal models. However, when learning multiple visual tasks simultaneously, this approach often results in suboptimal and imbalanced overall performance due to latent knowledge conflicts across tasks. To mitigate this issue, we propose a novel Adaptive Task Balancing approach tailored for visual instruction tuning (VisATB). Specifically, we measure two critical dimensions for visual task balancing based on validation performance: (1) Inter-Task Contribution, the mechanism where learning one task enhances the performance on others owing to shared knowledge across tasks, and (2) Intra-Task Difficulty, which denotes the inherent learning difficulty of a single task. Furthermore, we propose prioritizing three categories of tasks with greater weight: those that offer substantial contributions to others, those that receive minimal contributions from others, and those that present high learning difficulties. Among these three task weighting strategies, the first and third focus on improving overall performance, and the second targets the mitigation of performance imbalance. Extensive experiments on three benchmarks demonstrate that our VisATB approach consistently achieves superior and more balanced overall performance in visual instruction tuning. The data, code, and models are available at https://github.com/YanqiDai/VisATB.
Yanqi Dai, Zebin You, Dong Jing, Xiangxiang Chu, Zhiwu Lu 0001
WWW4
2025 Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion strategies for addressing the task. However, these methods typically fail to simultaneously meet two key requirements of CIR: comprehensively extracting visual information and faithfully following the user intent. In this work, we propose CIR-LVLM, a novel framework that leverages the large vision-language model (LVLM) as the powerful user intent-aware encoder to better meet these requirements. Our motivation is to explore the advanced reasoning and instruction-following capabilities of LVLM for accurately understanding and responding the user intent. Furthermore, we design a novel hybrid intent instruction module to provide explicit intent guidance at two levels: (1) The task prompt clarifies the task requirement and assists the model in discerning user intent at the task level. (2) The instance-specific soft prompt, which is adaptively selected from the learnable prompt pool, enables the model to better comprehend the user intent at the instance level compared to a universal prompt for all instances. CIR-LVLM achieves state-of-the-art performance across three prominent benchmarks with acceptable inference efficiency. We believe this study provides fundamental insights into CIR-related fields.
Zelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei, Zhiwu Lu 0001
AAAI2
2025 CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
Zelong Sun, Dong Jing, Zhiwu Lu 0001
ICCV2
2024 FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding
abstract
Contrastive Language-Image Pre-training (CLIP) achieves impressive performance on tasks like image classification and image-text retrieval by learning on large-scale image-text datasets. However, CLIP struggles with dense prediction tasks due to the poor grasp of the fine-grained details. Although existing works pay attention to this issue, they achieve limited improvements and usually sacrifice the important visual-semantic consistency. To overcome these limitations, we propose FineCLIP, which keeps the global contrastive learning to preserve the visual-semantic consistency and further enhances the fine-grained understanding through two innovations: 1) A real-time self-distillation scheme that facilitates the transfer of representation capability from global to local features. 2) A semantically-rich regional contrastive learning paradigm with generated region-text pairs, boosting the local representation capabilities with abundant fine-grained knowledge. Both cooperate to fully leverage diverse semantics and multi-grained complementary information. To validate the superiority of our FineCLIP and the rationality of each design, we conduct extensive experiments on challenging dense prediction and image-level tasks. All the observations demonstrate the effectiveness of FineCLIP.
Dong Jing, Xiaolong He 0003, Yutian Luo, Nanyi Fei, Guoxing Yang, Huiwen Zhao, Zhiwu Lu 0001
NeurIPS1
2024 A Multicloud Collaborative Data Security Sharing Scheme With Blockchain Indexing in Industrial Internet Environments
abstract
Data sharing in the field of industrial Internet provides a powerful impetus for the vigorous development of industrial Internet by integrating management and flexible configuration of resources. However, there are single point of failure and data leakage problems in the current industrial Internet data sharing and storage process due to the adoption of single-cloud storage mode. This paper proposes a multi-cloud collaborative data security sharing scheme with blockchain indexing. The blockchain-assisted multi-cloud storage model is constructed to enable indirect retrieval of industrial data through a federation chain. Enterprises on the alliance chain cannot directly access the industrial data stored in their respective private clouds for each other. To address the lack of trustworthiness of ciphertext retrieval, a tree-structured indexing method supporting keyword ciphertext retrieval is designed to realize efficient keyword token search. To prevent unfair trading behavior of data sharing participants, a fair trading contract is designed to overcome the disadvantage of dishonest execution of protocols by cloud servers. The security analysis shows that the scheme satisfies IND-CKA security. Simulation results show that the tree indexing approach constructed in this paper has less time overhead compared to the traditional sequential indexing and other robust technique. And the efficiency advantage is more prominent when the number of keywords is huge.
Dong Jing, Jingyu Feng
IEEE Internet Things J.3
2021 Occlusion-aware Bi-directional Guided Network for Light Field Salient Object Detection
abstract
Existing light field based works utilize either views or focal stacks for saliency detection. However, since depth information exists implicitly in adjacent views or different focal slices, it is difficult to exploit scene depth information from both. By comparison, Epipolar Plane Images (EPIs) provide explicit accurate scene depth and occlusion information by projected pixel lines. Due to the fact that the depth of an object is often continuous, the distribution of occlusion edges concentrates more on object boundaries compared with traditional color edges, which is more beneficial for improving accuracy and completeness of saliency detection. In this paper, we propose a learning-based network to exploit occlusion features from EPIs and integrate high-level features from the central view for accurate salient object detection. Specifically, a novel Occlusion Extraction Module is proposed to extract occlusion boundary features from horizontal and vertical EPIs. In order to naturally combine occlusion features in EPIs and high-level features in central view, we design a concise Bi-directional Guiding Flow based on cascaded decoders. The flow leverages generated salient edge predictions and salient object predictions to refine features in mutual encoding processes. Experimental results demonstrate that our approach achieves state-of-the-art performance in both segmentation accuracy and edge clarity.
Dong Jing, Shuo Zhang 0003, Runmin Cong, Youfang Lin
ACM Multimedia1