Jingcheng Ke

dblp:239/0105 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0002-2262-6261ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DiffusionREC: Diffusion Model with Adaptive Condition for Referring Expression Comprehension
abstract
The objective of referring expression comprehension (REC) is to accurately identify the object in an image described by a given expression. Existing REC methods, including transformer-based and graph-based approaches among others, have shown robust performance in REC tasks. In this study, we present a groundbreaking framework named DiffusionREC for REC task. This framework reimagines the REC task as a text guided bounding box denoising diffusion process, through which noisy bounding boxes are refined and distilled to pinpoint the target box. Throughout the training process, the bounding box of the target object diffuses from its ground-truth position towards a random distribution. Simultaneously, a filtering-based object decoder is introduced to reverse this diffusion of noise, conditional on the provided expression, the result from previous denoised step and the interaction between the expression and the image. At the inference stage, we begin by randomly generating a collection of boxes. Subsequently, the filtering-based object decoder is iteratively employed to refine and prune these bounding boxes, taking into account the conditions on the given expression, the results from the previous denoised step, and the interaction between the expression and the image. Extensive experiments conducted on six datasets demonstrate that DiffusionREC outperforms previous REC methods, yielding superior performances.
Jingcheng Ke, Wai Keung Wong, Jia Wang 0020, Mu Li 0005, Lunke Fei, Jie Wen 0001
AAAI1
2025 Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and Comprehension
abstract
Referring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. Moreover, REG methods frequently struggle to bridge the visual and textual domains due to the limited capacity, leading to low-quality and restricted diversity in expression generation. To address these issues, we propose a novel VIsion-guided Expression Diffusion Model (VIE-DM) for the REG task, where diverse synonymous expressions adhering to both image and text contexts of the target object are generated to augment REC datasets. VIE-DM consists of a vision-text condition (VTC) module and a transformer decoder. Our VTC and token selection design effectively addresses the feature discrepancy problem prevalent in existing REG methods. This enables us to generate high-quality, diverse synonymous expressions that can serve as augmented data for REC model learning. Extensive experiments on five datasets demonstrate the high quality and large diversity of our generated expressions. Furthermore, the augmented image-expression pairs consistently enhance the performance of existing REC models, achieving state-of-the-art results.
Jingcheng Ke, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin
ICLR1
2025 Graph-based referring expression comprehension with expression-guided selective filtering and noun-oriented reasoning
Jingcheng Ke, Qi Zhang 0059, Jia Wang 0020, Hongqing Ding, Jie Wen 0001
Pattern Recognit.1
2025 Graph-Based Group Division Network for Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims at locating the target object described by an expression. We observe that most of the graph-based REC methods only focus on establishing relations between all objects in an image and the given expression during the graph construction while ignoring the relationships between objects in the same category. As a result, these methods are sub-optimal in locating the target object described by the expression, particularly when the target object is surrounded by objects of similar categories. Meanwhile, during reasoning, numerous irrelevant objects are considered for expression, which will introduce significant harmful noise. To address these issues, this paper proposes a new graph-based group division network (GBGDN). Different from the existing works, our work partitions the constructed graphs into several sub-graphs based on the categories of objects and expressions. In each sub-graph, the common visual features of objects will be strengthened through a feature enhancement strategy. Subsequently, the enhanced sub-graphs and expressions undergo joint processing via a filtering-based reasoning module designed to reduce the influence of unrelated nodes in each sub-graph, facilitating more accurate reasoning and matching. Experimental results across various datasets, including RefCOCO /+/g, Flickr30K Entities, RefClef, and Ref-reasoning, showcase the superiority of our proposed method over existing approaches. Most importantly, our method does not need pre-training.
Jingcheng Ke, Jia Wang 0020, Wai Keung Wong, Anne Toomey, Jie Wen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Deep Multi-View Clustering With Meta Information Compression
abstract
Multi-view clustering typically leverages the consistency and complementarity among views to partition different samples. However, existing deep learning-based methods often face the dilemma between selecting complementary information and capturing essential details: 1) Capturing complementary semantics among views may introduce label-irrelevant redundant information. 2) Only extracting consistent semantic information will cause information loss, hindering the clarity in downstream tasks. To address these issues, we propose a novel method from the perspective of meta-learning to learn clustering-friendly representations with minimal redundancy. Specifically, we train an information compressor to guide the model in describing the original samples as compact as possible with minimal information, thus learning the key semantics with minimized redundancy. Meta-learning bi-level optimization promotes the nested optimization of feature embedding and information compressor. Meanwhile, a semantic puzzle mechanism complements the semantic fragments by exploiting the relationships between low-level features, resulting in a consensus representation with strong discriminative power. We conducted extensive experiments on datasets with various sizes to validate the effectiveness of our model, demonstrating significant performance improvements over several state-of-the-art methods.
Jinrong Cui, Bang Liufu, Chongjie Dong, Jingcheng Ke, Jie Wen 0001
IEEE Trans. Image Process.4
2025 Multi-Perspective Cross-Modal Object Encoding for Referring Expression Comprehension
abstract
Referring expression comprehension (REC) is a crucial task in understanding how a given text description identifies a target object within an image. Existing two-stage REC methods have demonstrated strong performance due to their rational framework design. However, during the encoding of object candidates in an image, most two-stage methods rely exclusively on features extracted from pre-trained detectors, often neglecting the contextual relationships between an object and its neighboring elements. This limitation hinders the full capture of contextual and relational information, reducing the discriminative power of object representations and negatively impacting subsequent processing. In this paper, we propose two novel plug-and-adapt modules: expression-guided label representation module (ELR) and cross-modal calibrated semantic module (CCS), designed to enhance two-stage REC methods. Specifically, the ELR module connects the noun phases of expression to the categorical labels of object candidates in the image, ensuring effective alignment between them. Guided by these connections, a CCS module is introduced to represent each object candidate by integrating its features with those of neighboring candidates from multiple perspectives. This preserves the intrinsic information of each candidate while incorporating relational cues from other objects, enabling more precise embeddings and effective downstream processing in two-stage REC methods. Extensive experiments on six datasets demonstrate the importance of incorporating prior statistical knowledge, and detailed analysis shows that the proposed modules strengthen the alignment between image and text. As a result, our method achieves competitive performance and is compatible with most two-stage methods in the REC task. The code is available on Github: https://github.com/freedom6927/ELR_CCS.git.
Jingcheng Ke, Jie Wen 0001, Huiting Wang, Wen-Huang Cheng, Jia Wang 0020
IEEE Trans. Image Process.1
2025 Make Graph-Based Referring Expression Comprehension Great Again Through Expression-Guided Dynamic Gating and Regression
abstract
One common belief is that with complex models and pre-training on large-scale datasets, transformer-based methods for referring expression comprehension (REC) perform much better than existing graph-based methods. We observe that since most graph-based methods adopt an off-the-shelf detector to locate candidate objects (i.e., regions detected by the object detector), they face two challenges that result in subpar performance: (1) the presence of significant noise caused by numerous irrelevant objects during reasoning, and (2) inaccurate localization outcomes attributed to the provided detector. To address these issues, we introduce a plug-and-adapt module guided by sub-expressions, called dynamic gate constraint (DGC), which can adaptively disable irrelevant proposals and their connections in graphs during reasoning. We further introduce an expression-guided regression strategy (EGR) to refine location prediction. Extensive experimental results on the RefCOCO, RefCOCO+, RefCOCOg, Flickr30 K, RefClef, and Ref-reasoning datasets demonstrate the effectiveness of the DGC module and the EGR strategy in consistently boosting the performances of various graph-based REC methods. Without any pretaining, the proposed graph-based method achieves better performance than the state-of-the-art (SOTA) transformer-based methods.
Jingcheng Ke, Dele Wang, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin
IEEE Trans. Multim.1
2024 CLIPREC: Graph-Based Domain Adaptive Network for Zero-Shot Referring Expression Comprehension
abstract
Referring expression comprehension (REC) is a cross-modal matching task that aims to localize the target object in an image specified by a text description. Most existing approaches for this task focus on identifying only objects whose categories are covered by training data. This restricts their generalization to unseen categories and practical usage. To address this issue, we propose a domain adaptive network called CLIPREC for zero-shot REC, which integrates the Contrastive Language-Image Pretraining (CLIP) model for graph-based REC. The proposed CLIPREC is composed of a graph collaborative attention module with two directed graphs: one for objects in an image and the other for their corresponding categorical labels. To carry out zero-shot REC, we leverage the strong common image-text feature space from the CLIP model to correlate the two graphs. Furthermore, a multilayer perceptron is introduced to enable feature alignment so that the CLIP model is adapted to the expression representation from the language parser, resulting in effective reasoning from expressions involving both seen and unseen object categories. Extensive experimental and ablation results on several widely-adopted benchmarks show that the proposed approach performs favorably against state-of-the-art approaches for zero-shot REC.
Jingcheng Ke, Jia Wang 0020, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin
IEEE Trans. Multim.1
2023 Referring Expression Comprehension Via Enhanced Cross-modal Graph Attention Networks
abstract
Referring expression comprehension aims to localize a specific object in an image according to a given language description. It is still challenging to comprehend and mitigate the gap between various types of information in the visual and textual domains. Generally, it needs to extract the salient features from a given expression and match the features of expression to an image. One challenge in referring expression comprehension is the number of region proposals generated by object detection methods is far more than the number of entities in the corresponding language description. Remarkably, the candidate regions without described by the expression will bring a severe impact on referring expression comprehension. To tackle this problem, we first propose a novel Enhanced Cross-modal Graph Attention Networks (ECMGANs) that boosts the matching between the expression and the entity position of an image. Then, an effective strategy named Graph Node Erase (GNE) is proposed to assist ECMGANs in eliminating the effect of irrelevant objects on the target object. Experiments on three public referring expression comprehension datasets show unambiguously that our ECMGANs framework achieves better performance than other state-of-the-art methods. Moreover, GNE is able to obtain higher accuracies of visual-expression matching effectively.
Jia Wang 0020, Jingcheng Ke, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ACM Trans. Multim. Comput. Commun. Appl.2
2019 A novel grouped sparse representation for face recognition
abstract
Grouped sparse representation classification methods (GSRCMs) have been attracted much attention by scholars, especially in face recognition. However, pervious literatures of GSRCMs only fuse the scores from different groups to classification the test sample, not consider relationships of the groups. Moreover, in real-world application, many methods of face recognition cannot obtain satisfied recognition accuracies because of the variation of poses, illuminations and facial representations of face image. In order to overcome above-mentioned bottlenecks, in this paper, we proposed a novel grouped fusion-based method in face recognition. The proposed method uses the axis-symmetrical property of face to designs a framework and perform it on original training set to generate a kind of virtual samples. The virtual samples are able to reflect the possible change of face images. Meanwhile, to consider the relationship of different groups and strengthen the representation capability of test sample, the proposed method exploits a novel weighted fusion approach to classify the test sample. Experimental results on five face databases demonstrate that our method is reasonable and can obtain higher recognition rate than the other 11 state-of-the-art methods.
Jingcheng Ke, Shigang Liu, Zengguo Sun
Multim. Tools Appl.1