Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mengjun Cheng

dblp:302/3451 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-3271-7589ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 50% Vision and language · 22% Knowledge representation and reasoning · 14%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding › image segmentation
boundary-aware segmentation
0.912025
Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation · IEEE Trans. Multim. 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.912025
Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation · IEEE Trans. Multim. 2025
Computer vision › Segmentation and scene understanding › medical image segmentation
polyp segmentation
0.912025
Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation · IEEE Trans. Multim. 2025
Natural language and speech › Information extraction and text analysis
document understanding
0.812024
Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents · ECCV (45) 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction › multimodal information extraction
visual information extraction
0.812024
Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents · ECCV (45) 2024
Computer vision › Vision and language
cross-modal retrieval
0.612022
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval · CVPR 2022
Computer vision › Vision and language › multimodal understanding
scene text understanding
0.612022
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval · CVPR 2022

Methods — techniques the papers use, named apart from their topics

top-down architecture · 0.9encoder-decoder · 0.9textual grounding · 0.8open-vocabulary learning · 0.8transformer aggregation · 0.6dual contrastive learning · 0.6
YearPublicationVenuePosition
2025 Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation
abstract
The diagnosis of colon polyps is important for the prevention of colorectal cancer. Polyp segmentation, however, is still a challenging problem given that recent medical computer-aided equipment suffers from situations of polyp variations in terms of size, color, texture, and poor illuminations in endoscopy videos. These obstacles hinder the prediction of polyp boundaries. Inspired by the observation that the values of pixels on the border region change more sharply than others, we propose the oriented-derivative (OD) representation to capture the relationship between pixels and the boundary region given distance and orientation. To adaptively use the proposed representation in arbitrary frameworks, we design plug-in modules to learn the representation and aggregate features to improve the accuracy of boundary predictions in the polyp segmentation task, which can be implemented in frameworks including the encoder-decoder and top-down architectures. Extensive experimental results show the improvement from the proposed oriented-derivative representation for the polyp segmentation task and the extendibility of our proposed modules in different architectures. Our methods achieved an improvement ranging from 0.3% to 2.5% (mDice) compared with the baseline on five publicly available datasets, includingKvasir, CVC-ClinicDB, EndoScene, CVC-ColonDB, andETIS.
Mengjun Cheng, Xiawu Zheng, Rongrong Ji, Jie Chen 0001
IEEE Trans. Multim.2
2024 Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents
Mengjun Cheng, Chengquan Zhang, Chang Liu 0047, Xiawu Zheng, Rongrong Ji, Jie Chen 0001
ECCV (45)1
2023 ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai
ICDAR (2)10
2022 ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
abstract
Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@ 1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework.
Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Jie Chen 0001, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
CVPR1
2021 Learnable Oriented-Derivative Network for Polyp Segmentation
Mengjun Cheng, Zishang Kong, Guoli Song, Yonghong Tian 0001, Yongsheng Liang 0001, Jie Chen 0001
MICCAI (1)1