Jia Wang 0020

dblp:58/6299-20 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-0998-251XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 DiffusionREC: Diffusion Model with Adaptive Condition for Referring Expression Comprehension
abstract
The objective of referring expression comprehension (REC) is to accurately identify the object in an image described by a given expression. Existing REC methods, including transformer-based and graph-based approaches among others, have shown robust performance in REC tasks. In this study, we present a groundbreaking framework named DiffusionREC for REC task. This framework reimagines the REC task as a text guided bounding box denoising diffusion process, through which noisy bounding boxes are refined and distilled to pinpoint the target box. Throughout the training process, the bounding box of the target object diffuses from its ground-truth position towards a random distribution. Simultaneously, a filtering-based object decoder is introduced to reverse this diffusion of noise, conditional on the provided expression, the result from previous denoised step and the interaction between the expression and the image. At the inference stage, we begin by randomly generating a collection of boxes. Subsequently, the filtering-based object decoder is iteratively employed to refine and prune these bounding boxes, taking into account the conditions on the given expression, the results from the previous denoised step, and the interaction between the expression and the image. Extensive experiments conducted on six datasets demonstrate that DiffusionREC outperforms previous REC methods, yielding superior performances.
Jingcheng Ke, Wai Keung Wong, Jia Wang 0020, Mu Li 0005, Lunke Fei, Jie Wen 0001
AAAI3
2025 SMPV: Social Media Prediction for Videos
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng
ACM Multimedia5
2025 Graph-based referring expression comprehension with expression-guided selective filtering and noun-oriented reasoning
Jingcheng Ke, Qi Zhang 0059, Jia Wang 0020, Hongqing Ding, Jie Wen 0001
Pattern Recognit.3
2025 Graph-Based Group Division Network for Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims at locating the target object described by an expression. We observe that most of the graph-based REC methods only focus on establishing relations between all objects in an image and the given expression during the graph construction while ignoring the relationships between objects in the same category. As a result, these methods are sub-optimal in locating the target object described by the expression, particularly when the target object is surrounded by objects of similar categories. Meanwhile, during reasoning, numerous irrelevant objects are considered for expression, which will introduce significant harmful noise. To address these issues, this paper proposes a new graph-based group division network (GBGDN). Different from the existing works, our work partitions the constructed graphs into several sub-graphs based on the categories of objects and expressions. In each sub-graph, the common visual features of objects will be strengthened through a feature enhancement strategy. Subsequently, the enhanced sub-graphs and expressions undergo joint processing via a filtering-based reasoning module designed to reduce the influence of unrelated nodes in each sub-graph, facilitating more accurate reasoning and matching. Experimental results across various datasets, including RefCOCO /+/g, Flickr30K Entities, RefClef, and Ref-reasoning, showcase the superiority of our proposed method over existing approaches. Most importantly, our method does not need pre-training.
Jingcheng Ke, Jia Wang 0020, Wai Keung Wong, Anne Toomey, Jie Wen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Multi-Perspective Cross-Modal Object Encoding for Referring Expression Comprehension
abstract
Referring expression comprehension (REC) is a crucial task in understanding how a given text description identifies a target object within an image. Existing two-stage REC methods have demonstrated strong performance due to their rational framework design. However, during the encoding of object candidates in an image, most two-stage methods rely exclusively on features extracted from pre-trained detectors, often neglecting the contextual relationships between an object and its neighboring elements. This limitation hinders the full capture of contextual and relational information, reducing the discriminative power of object representations and negatively impacting subsequent processing. In this paper, we propose two novel plug-and-adapt modules: expression-guided label representation module (ELR) and cross-modal calibrated semantic module (CCS), designed to enhance two-stage REC methods. Specifically, the ELR module connects the noun phases of expression to the categorical labels of object candidates in the image, ensuring effective alignment between them. Guided by these connections, a CCS module is introduced to represent each object candidate by integrating its features with those of neighboring candidates from multiple perspectives. This preserves the intrinsic information of each candidate while incorporating relational cues from other objects, enabling more precise embeddings and effective downstream processing in two-stage REC methods. Extensive experiments on six datasets demonstrate the importance of incorporating prior statistical knowledge, and detailed analysis shows that the proposed modules strengthen the alignment between image and text. As a result, our method achieves competitive performance and is compatible with most two-stage methods in the REC task. The code is available on Github: https://github.com/freedom6927/ELR_CCS.git.
Jingcheng Ke, Jie Wen 0001, Huiting Wang, Wen-Huang Cheng, Jia Wang 0020
IEEE Trans. Image Process.5
2024 SMP Challenge Summary: Social Media Prediction Challenge
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng
ACM Multimedia5
2024 CLIPREC: Graph-Based Domain Adaptive Network for Zero-Shot Referring Expression Comprehension
abstract
Referring expression comprehension (REC) is a cross-modal matching task that aims to localize the target object in an image specified by a text description. Most existing approaches for this task focus on identifying only objects whose categories are covered by training data. This restricts their generalization to unseen categories and practical usage. To address this issue, we propose a domain adaptive network called CLIPREC for zero-shot REC, which integrates the Contrastive Language-Image Pretraining (CLIP) model for graph-based REC. The proposed CLIPREC is composed of a graph collaborative attention module with two directed graphs: one for objects in an image and the other for their corresponding categorical labels. To carry out zero-shot REC, we leverage the strong common image-text feature space from the CLIP model to correlate the two graphs. Furthermore, a multilayer perceptron is introduced to enable feature alignment so that the CLIP model is adapted to the expression representation from the language parser, resulting in effective reasoning from expressions involving both seen and unseen object categories. Extensive experimental and ablation results on several widely-adopted benchmarks show that the proposed approach performs favorably against state-of-the-art approaches for zero-shot REC.
Jingcheng Ke, Jia Wang 0020, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin
IEEE Trans. Multim.2
2024 Language-guided Residual Graph Attention Network and Data Augmentation for Visual Grounding
abstract
Visual grounding is an essential task in understanding the semantic relationship between the given text description and the target object in an image. Due to the innate complexity of language and the rich semantic context of the image, it is still a challenging problem to infer the underlying relationship and to perform reasoning between the objects in an image and the given expression. Although existing visual grounding methods have achieved promising progress, cross-modal mapping across different domains for the task is still not well handled, especially when the expressions are complex and long. To address the issue, we propose a language-guided residual graph attention network for visual grounding (LRGAT-VG), which enables us to apply deeper graph convolution layers with the assistance of residual connections between them. This allows us to better handle long and complex expressions than other graph-based methods. Furthermore, we perform a Language-guided Data Augmentation (LGDA), which is based on copy-paste operations on pairs of source and target images to increase the diversity of training data while maintaining the relationship between the objects in the image and the expression. With extensive experiments on three visual grounding benchmarks, including RefCOCO, RefCOCO+, and RefCOCOg, LRGAT-VG with LGDA achieves competitive performance with other state-of-the-art graph network-based referring expression approaches and demonstrates its effectiveness.
Jia Wang 0020, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ACM Trans. Multim. Comput. Commun. Appl.1
2023 SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge
abstract
Social Media Popularity Prediction (SMPP) is a crucial task that involves automatically predicting future popularity values of online posts, leveraging vast amounts of multimodal data available on social media platforms. Studying and investigating social media popularity becomes central to various online applications and requires novel methods of comprehensive analysis, multimodal comprehension, and accurate prediction.
Bo Wu 0018, Peiye Liu, Wen-Huang Cheng, Bei Liu 0001, Zhaoyang Zeng, Jia Wang 0020, Qiushi Huang, Jiebo Luo 0001
ACM Multimedia6
2023 Referring Expression Comprehension Via Enhanced Cross-modal Graph Attention Networks
abstract
Referring expression comprehension aims to localize a specific object in an image according to a given language description. It is still challenging to comprehend and mitigate the gap between various types of information in the visual and textual domains. Generally, it needs to extract the salient features from a given expression and match the features of expression to an image. One challenge in referring expression comprehension is the number of region proposals generated by object detection methods is far more than the number of entities in the corresponding language description. Remarkably, the candidate regions without described by the expression will bring a severe impact on referring expression comprehension. To tackle this problem, we first propose a novel Enhanced Cross-modal Graph Attention Networks (ECMGANs) that boosts the matching between the expression and the entity position of an image. Then, an effective strategy named Graph Node Erase (GNE) is proposed to assist ECMGANs in eliminating the effect of irrelevant objects on the target object. Experiments on three public referring expression comprehension datasets show unambiguously that our ECMGANs framework achieves better performance than other state-of-the-art methods. Moreover, GNE is able to obtain higher accuracies of visual-expression matching effectively.
Jia Wang 0020, Jingcheng Ke, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Residual Graph Attention Network and Expression-Respect Data Augmentation Aided Visual Grounding
abstract
Visual grounding aims to localize a target object in an image based on a given text description. Due to the innate complexity of language, it is still a challenging problem to perform reasoning of complex expressions and to infer the underlying relationship between the expression and the object in an image. To address these issues, we propose a residual graph attention network for visual grounding. The proposed approach first builds an expression-guided relation graph and then performs multi-step reasoning followed by matching the target object. It allows performing better visual grounding with complex expressions by using deeper layers than other graph network approaches. Moreover, to increase the diversity of training data, we perform an expression-respect data augmentation based on copy-paste operations to pairs of source and target images. The proposed approach achieves better performance with extensive experiments than other state-of-the-art graph network-based approaches and demonstrates its effectiveness.
Jia Wang 0020, Hung-Yi Wu, Jun-Cheng Chen, Hong-Han Shuai, Wen-Huang Cheng
ICIP1