VLDB 2026 Research / reviewers in the wild / expert
Zhangzikang Li
dblp:337/2174
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 77% Segmentation and scene understanding · 18% Video understanding and tracking · 5% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
0.7 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Computer vision › Vision and language › image captioning › structured image captioning
scene graph-based captioning |
0.7 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Computer vision › Vision and language › video-text retrieval
text-to-video retrieval |
0.7 | 1 | 2023 | Learning Trajectory-Word Alignments for Video-Language Tasks · ICCV 2023 |
Computer vision › Vision and language › video-language understanding
video-text tasks |
0.7 | 1 | 2023 | Learning Trajectory-Word Alignments for Video-Language Tasks · ICCV 2023 |
Computer vision › Segmentation and scene understanding › scene graph
visual scene graph |
0.7 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Computer vision › Vision and language
caption generation |
0.2 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Computer vision › Video understanding and tracking
video question answering |
0.2 | 1 | 2023 | Learning Trajectory-Word Alignments for Video-Language Tasks · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.7trajectory-to-word attention · 0.7scene graph transformation · 0.7hierarchical frame selector · 0.7contrastive learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Transforming Visual Scene Graphs to Image CaptionsabstractXu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu, Qinghao Ye, Chenliang Li, Songfang Huang, Fei Huang, Zhangzikang Li, Yu Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xu Yang 0021, Jiawei Peng 0001, Zihua Wang, Haiyang Xu 0001, Qinghao Ye, Chenliang Li 0003, Songfang Huang, Fei Huang 0002, Zhangzikang Li, Yu Zhang 0004 |
ACL (1) | 9 |
| 2023 | Learning Trajectory-Word Alignments for Video-Language TasksabstractIn a video, an object usually appears as the trajectory, i.e., it spans over a few spatial but longer temporal patches, that contains abundant spatiotemporal contexts. However, modern Video-Language BERTs (VDL-BERTs) neglect this trajectory characteristic that they usually follow image-language BERTs (IL-BERTs) to deploy the patch-to-word (P2W) attention that may over-exploit trivial spatial contexts and neglect significant temporal contexts. To amend this, we propose a novel TW-BERT to learn Trajectory-Word alignment by a newly designed trajectory-to-word (T2W) attention for solving video-language tasks. Moreover, previous VDL-BERTs usually uniformly sample a few frames into the model while different trajectories have diverse graininess, i.e., some trajectories span longer frames and some span shorter, and using a few frames will lose certain useful temporal contexts. However, simply sampling more frames will also make pre-training infeasible due to the largely increased training burdens. To alleviate the problem, during the fine-tuning stage, we insert a novel Hierarchical Frame-Selector (HFS) module into the video encoder. HFS gradually selects the suitable frames conditioned on the text context for the later cross-modal encoder to learn better trajectory-word alignments. By the proposed T2W attention and HFS, our TW-BERT achieves SOTA performances on text-to-video retrieval tasks, and comparable performances on video question-answering tasks with some VDL-BERTs trained on much more data. The code will be available in the supplementary material. Xu Yang 0021, Zhangzikang Li, Haiyang Xu 0001, Hanwang Zhang, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Yu Zhang 0004, Fei Huang 0002, Songfang Huang |
ICCV | 2 |