EDBT 2026 Demo / reviewers in the wild / expert
Mengjun Cheng
dblp:302/3451
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-3271-7589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Segmentation and scene understanding · 50% Vision and language · 22% Knowledge representation and reasoning · 14% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding › image segmentation
boundary-aware segmentation |
0.9 | 1 | 2025 | Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation · IEEE Trans. Multim. 2025 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.9 | 1 | 2025 | Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation · IEEE Trans. Multim. 2025 |
Computer vision › Segmentation and scene understanding › medical image segmentation
polyp segmentation |
0.9 | 1 | 2025 | Oriented-Derivative Representation for Boundary-Aware Polyp Segmentation · IEEE Trans. Multim. 2025 |
Natural language and speech › Information extraction and text analysis
document understanding |
0.8 | 1 | 2024 | Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents · ECCV (45) 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction › multimodal information extraction
visual information extraction |
0.8 | 1 | 2024 | Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents · ECCV (45) 2024 |
Computer vision › Vision and language
cross-modal retrieval |
0.6 | 1 | 2022 | ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval · CVPR 2022 |
Computer vision › Vision and language › multimodal understanding
scene text understanding |
0.6 | 1 | 2022 | ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
top-down architecture · 0.9encoder-decoder · 0.9textual grounding · 0.8open-vocabulary learning · 0.8transformer aggregation · 0.6dual contrastive learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Oriented-Derivative Representation for Boundary-Aware Polyp SegmentationabstractThe diagnosis of colon polyps is important for the prevention of colorectal cancer. Polyp segmentation, however, is still a challenging problem given that recent medical computer-aided equipment suffers from situations of polyp variations in terms of size, color, texture, and poor illuminations in endoscopy videos. These obstacles hinder the prediction of polyp boundaries. Inspired by the observation that the values of pixels on the border region change more sharply than others, we propose the oriented-derivative (OD) representation to capture the relationship between pixels and the boundary region given distance and orientation. To adaptively use the proposed representation in arbitrary frameworks, we design plug-in modules to learn the representation and aggregate features to improve the accuracy of boundary predictions in the polyp segmentation task, which can be implemented in frameworks including the encoder-decoder and top-down architectures. Extensive experimental results show the improvement from the proposed oriented-derivative representation for the polyp segmentation task and the extendibility of our proposed modules in different architectures. Our methods achieved an improvement ranging from 0.3% to 2.5% (mDice) compared with the baseline on five publicly available datasets, includingKvasir, CVC-ClinicDB, EndoScene, CVC-ColonDB, andETIS. Mengjun Cheng, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents
Mengjun Cheng, Chengquan Zhang, Chang Liu 0047, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
ECCV (45) | 1 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 10 |
| 2022 | ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalabstractVisual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@ 1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework. Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Jie Chen 0001, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 1 |
| 2021 | Learnable Oriented-Derivative Network for Polyp Segmentation
Mengjun Cheng, Zishang Kong, Guoli Song, Yonghong Tian 0001, Yongsheng Liang 0001, Jie Chen 0001 |
MICCAI (1) | 1 |