EDBT 2026 Demo / reviewers in the wild / expert
Shiyu Xie
dblp:53/8391
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 3Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E2ETMPN: An End-to-End Template Matching Prediction Network for Intra CodingabstractIntra coding plays a critical role in video compression, yet conventional directional intra prediction mainly relies on boundary pixels and fails to capture complex textures and long-range spatial correlations. Template Matching Prediction (TMP) alleviates this issue by searching similar templates in reconstructed regions without motion-vector signaling; however, its MSE/MAE-based matching and simple averaging fusion are sensitive to noise and often unreliable for complex content. This paper proposes E2ETMPN, an end-to-end template matching prediction network for intra coding. E2ETMPN jointly models template matching and multi-candidate fusion under a unified loss function. It consists of two components: (1) a CNN-based multiscale matching sub-network that searches candidate predictions from reconstructed regions at different spatial scales; and (2) a Transformer-based fusion sub-network that adaptively fuses multiple candidates with template information using attention mechanisms to generate the final prediction. Experimental results demonstrate that integrating E2ETMPN into the reference encoder achieves consistent BD-rate reductions over conventional TMP and standard intra prediction methods on standard test sequences. Notably, E2ETMPN yields larger gains on screen content and scenes with rich non-local repetitive structures, validating the effectiveness of end-to-end learning for template matching prediction in intra coding. Qijun Wang, Shiyu Xie |
DCC | 2 |
| 2026 | Compressed Video Stream Learning for Video-Text RetrievalabstractVideo-Text Retrieval (VTR) aims to align video content with natural language descriptions and is a fundamental task in multi-modal understanding. Most existing methods model videos as uniformly sampled RGB frames, which overlooks rich temporal cues, especially motion dynamics encoded in videos. We propose Compressed Video Stream Learning for Video-Text Retrieval (CVSVTR), a framework that directly exploits information from compressed video streams to enhance retrieval performance without fully decoding videos. Specifically, CVSVTR decodes only the I-frames of each GOP and extracts appearance features using a CLIP-based encoder. Meanwhile, motion vectors and residuals are parsed from the compressed bitstream and processed by a lightweight P-frame Feature Generation (PFG) module to construct motion-aware representations for P-frames. A spatial-channel attention mechanism is further introduced to adaptively fuse appearance features with compressed-domain motion cues, compensating for temporal information missed by uniform frame sampling. Extensive experiments on the MSR-VTT and MSVD benchmarks demonstrate that CVSVTR consistently outperforms existing video-text retrieval methods across multiple evaluation metrics, validating the effectiveness of leveraging compressed video streams for efficient and accurate temporal modeling in VTR. Qijun Wang, Shiyu Xie, Xuguang Liu |
DCC | 2 |
| 2024 | Effective semi-supervised graph clustering with pairwise constraints
Shiyu Xie, Hui Yang 0005, Feiping Nie 0001 |
Inf. Sci. | 2 |
| 2022 | A novel method for optimizing spectral rotation embedding K-means with coordinate descent
Jianyong Zhu, Bingxia Feng, Shiyu Xie, Hui Yang 0005, Feiping Nie 0001 |
Inf. Sci. | 4 |
| 2022 | FGC_SS: Fast Graph Clustering Method by Joint Spectral Embedding and Improved Spectral Rotation
Jianyong Zhu, Shiyu Xie, Hui Yang 0005, Feiping Nie 0001 |
Inf. Sci. | 3 |