Ziyao Chen

dblp:185/7872 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0009-0014-5406ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning paradigms · 48% Vision and language · 31% Video understanding and tracking · 21%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 76% Geometric modeling and processing · 24%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms › continual learning › catastrophic forgetting
catastrophic forgetting mitigation
1.012026
KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning · AAAI 2026
Machine learning › Learning paradigms
continual learning
1.012026
KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning · AAAI 2026
Computer vision › Video understanding and tracking › temporal localization
temporal event localization
0.912025
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning · AAAI 2025
Multimedia analysis and retrieval › video captioning
dense video captioning
0.912025
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning · AAAI 2025
Multimedia analysis and retrieval
video understanding
0.812024
Short Video Ordering via Position Decoding and Successor Prediction · SIGIR 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning
0.312026
KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning · AAAI 2026
Geometric modeling and processing
multi-view geometry
0.212016
Novel Coplanar Line-Points Invariants for Robust Line Matching Across Views · ECCV (8) 2016

Methods — techniques the papers use, named apart from their topics

mask generation · 1.7dual-mode captioning · 1.7complementary masking · 1.7orthogonal subspace regularization · 1.0mixture of experts · 1.0knowledge subspace decomposition · 1.0successor prediction · 0.8position decoding · 0.8pairwise ordering · 0.8listwise ordering · 0.8coplanar line-point invariants · 0.2
YearPublicationVenuePosition
2026 KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning
abstract
Multimodal Large Language Models (MLLMs) employing the Mixture-of-Experts (MoE) structure exhibit encouraging results in visual language tasks. However, they struggle with catastrophic forgetting due to a lack of effective collaboration among experts and negative transfer across tasks. This happens because the router typically employed in MoE for managing expert assignments is inadequate when there are significant shifts in data distribution across various tasks. A drop in the effectiveness of earlier tasks is caused by negative transfer, which occurs due to conflicts in shared knowledge between tasks, disturbing the knowledge already acquired. To address these issues, we propose the Knowledge Space Synergy Framework in Mixture of Experts (KSS-MoE) for Continual Visual Instruction Tuning (CVIT). It dynamically combines the knowledge subspaces of experts to improve the integration of fine-grained complementary knowledge and collaborative abilities of experts, thus addressing the limitations of the basic router. Furthermore, we introduce a general expert that maintains orthogonal subspaces for shared knowledge, enabling effective cross-task knowledge utilization while reducing negative transfer. Extensive experiments conducted on eight CVIT tasks confirm the excellence of KSS-MoE, showcasing its top-tier performance.
Lingyun Song, Ziyao Chen, Kang Pan, Xiaolin Han 0002, Xinbiao Gan, Yudai Pan, Xiaofan Sun, Xuequn Shang 0001
AAAI2
2025 Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
abstract
Weakly-Supervised Dense Video Captioning (WSDVC) aims to localize and describe all events of interest in a video without requiring annotations of event boundaries. This setting poses a great challenge in accurately locating the temporal location of event, as the relevant supervision is unavailable. Existing methods rely on explicit alignment constraints between event locations and captions, which involve complex event proposal procedures during both training and inference. To tackle this problem, we propose a novel implicit location-caption alignment paradigm by complementary masking, which simplifies the complex event proposal and localization process while maintaining effectiveness. Specifically, our model comprises two components: a dual-mode video captioning module and a mask generation module. The dual-mode video captioning module captures global event information and generates descriptive captions, while the mask generation module generates differentiable positive and negative masks for localizing the events. These masks enable the implicit alignment of event locations and captions by ensuring that captions generated from positively and negatively masked videos are complementary, thereby forming a complete video description. In this way, even under weak supervision, the event location and event caption can be aligned implicitly. Extensive experiments on the public datasets demonstrate that our method outperforms existing weakly-supervised methods and achieves competitive results compared to fully-supervised methods.
Shiping Ge, Zhiwei Jiang 0001, Yafeng Yin 0002, Liu Qin, Ziyao Chen, Qing Gu 0001
AAAI6
2024 Short Video Ordering via Position Decoding and Successor Prediction
abstract
Short video collection is an easy way for users to consume coherent content on various online short video platforms, such as TikTok, YouTube, Douyin, and WeChat Channel. These collections cover a wide range of content, including online courses, TV series, movies, and cartoons. However, short video creators occasionally publish videos in a disorganized manner due to various reasons, such as revisions, secondary creations, deletions, and reissues, which often result in a poor browsing experience for users. Therefore, accurately reordering videos within a collection based on their content coherence is a vital task that can enhance user experience and presents an intriguing research problem in the field of video narrative reasoning. In this work, we curate a dedicated multimodal dataset for this Short Video Ordering (SVO) task and present the performance of some benchmark methods on the dataset. In addition, we further propose an advanced SVO framework with the aid of position decoding and successor prediction. The proposed framework combines both pairwise and listwise ordering paradigms, which can get rid of the issues from both quadratic growth and cascading conflict in the pairwise paradigm, and improve the performance of existing listwise methods. Extensive experiments demonstrate that our method achieves the best performance on our open SVO dataset, and each component of the framework contributes to the final performance. Both the SVO dataset and code will be released at https://github.com/ShipingGe/SVO.
Shiping Ge, Zhiwei Jiang 0001, Yafeng Yin 0002, Ziyao Chen, Qing Gu 0001
SIGIR5
2016 Novel Coplanar Line-Points Invariants for Robust Line Matching Across Views
Qi Jia 0001, Xinkai Gao, Xin Fan 0001, Zhongxuan Luo, Ziyao Chen
ECCV (8)6