EDBT 2026 Demo / reviewers in the wild / expert
Yukang Liang
dblp:344/3496
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › visual question answering
text-based visual question answering |
0.7 | 1 | 2023 | Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA · AAAI 2023 |
Computer vision › Vision and language
visual question answering |
0.7 | 1 | 2023 | Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA · AAAI 2023 |
Computer vision › Vision and language
cross-modal alignment |
0.2 | 1 | 2023 | Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA · AAAI 2023 |
Computer vision › Vision and language › cross-modal alignment
visual-semantic alignment |
0.2 | 1 | 2023 | Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
region proposal network · 0.7pre-trained language model · 0.7language refinement network · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Summarizing Like Human: Edit-Based Text Summarization with Keywords
Yukang Liang, Junliang Guo, Yongxin Zhu 0003, Linli Xu 0002 |
ICANN (7) | 1 |
| 2023 | Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQAabstractIn this paper, we propose a novel multi-modal framework for Scene Text Visual Question Answering (STVQA), which requires models to read scene text in images for question answering. Apart from text or visual objects, which could exist independently, scene text naturally links text and visual modalities together by conveying linguistic semantics while being a visual object in an image simultaneously. Different to conventional STVQA models which take the linguistic semantics and visual semantics in scene text as two separate features, in this paper, we propose a paradigm of "Locate Then Generate" (LTG), which explicitly unifies this two semantics with the spatial bounding box as a bridge connecting them. Specifically, at first, LTG locates the region in an image that may contain the answer words with an answer location module (ALM) consisting of a region proposal network and a language refinement network, both of which can transform to each other with one-to-one mapping via the scene text bounding box. Next, given the answer words selected by ALM, LTG generates a readable answer sequence with an answer generation module (AGM) based on a pre-trained language model. As a benefit of the explicit alignment of the visual and linguistic semantics, even without any scene text based pre-training tasks, LTG can boost the absolute accuracy by +6.06% and +6.92% on the TextVQA dataset and the ST-VQA dataset respectively, compared with a non-pre-training baseline. We further demonstrate that LTG effectively unifies visual and text modalities through the spatial bounding box connection, which is underappreciated in previous methods. Yongxin Zhu 0003, Yukang Liang, Xin Li 0118, Hao Liu 0003, Changcun Bao, Linli Xu 0002 |
AAAI | 3 |
| 2023 | End-to-End Word-Level Pronunciation Assessment with MASK Pre-training
Yukang Liang, Kaitao Song, Shaoguang Mao, Huiqiang Jiang, Luna Qiu, Yuqing Yang 0001, Dongsheng Li 0002, Lili Qiu |
INTERSPEECH | 1 |