EDBT 2026 Demo / reviewers in the wild / expert
Siru Zhong
dblp:359/5794
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0001-4465-8686ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OccamVTS: Distilling Vision Models to 1% Parameters for Time Series ForecastingabstractTime series forecasting is fundamental to diverse applications, with recent approaches leverage large vision models (LVMs) to capture temporal patterns through visual representations. We reveal that while vision models enhance forecasting performance, 99% of their parameters are unnecessary for time series tasks. Through cross-modal analysis, we find that time series align with low-level textural features but not high-level semantics, which can impair forecasting accuracy. We propose OccamVTS, a knowledge distillation framework that extracts only the essential 1% of predictive information from LVMs into lightweight networks. Using pre-trained LVMs as privileged teachers, OccamVTS employs pyramid-style feature alignment combined with correlation and feature distillation to transfer beneficial patterns while filtering out semantic noise. Counterintuitively, this aggressive parameter reduction improves accuracy by eliminating overfitting to irrelevant visual features while preserving essential temporal patterns. Extensive experiments across multiple benchmark datasets demonstrate that OccamVTS consistently achieves state-of-the-art performance with only 1% of the original parameters, particularly excelling in few-shot and zero-shot scenarios. Sisuo Lyu, Siru Zhong, Weilin Ruan, Qingxiang Liu 0004, Qingsong Wen, Hui Xiong 0001, Yuxuan Liang 0002 |
AAAI | 2 |
| 2025 | UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator PredictionabstractUrban socioeconomic indicator prediction aims to infer various metrics related to sustainable development in diverse urban landscapes using data-driven methods. However, prevalent pretrained models, particularly those reliant on satellite imagery, face dual challenges. Firstly, concentrating solely on macro-level patterns from satellite data may introduce bias, lacking nuanced details at micro levels, such as architectural details at a place. Secondly, the text generated by the precursor work UrbanCLIP, which fully utilizes the extensive knowledge of LLMs, frequently exhibits issues such as hallucination and homogenization, resulting in a lack of reliable quality. In response to these issues, we devise a novel framework entitled UrbanVLP based on Vision-Language Pretraining. Our UrbanVLP seamlessly integrates multi-granularity information from both macro (satellite) and micro (street-view) levels, overcoming the limitations of prior pretrained models. Moreover, it introduces automatic text generation and calibration, providing a robust guarantee for producing high-quality text descriptions of urban imagery. Rigorous experiments conducted across six socioeconomic indicator prediction tasks underscore its superior performance. Xixuan Hao, Wei Chen 0070, Siru Zhong, Kun Wang 0056, Qingsong Wen, Yuxuan Liang 0002 |
AAAI | 4 |
| 2025 | AirRadar: Inferring Nationwide Air Quality in China with Deep Neural NetworksabstractMonitoring real-time air quality is essential for safeguarding public health and fostering social progress. However, the widespread deployment of air quality monitoring stations is constrained by their significant costs. To address this limitation, we introduce AirRadar, a deep neural network designed to accurately infer real-time air quality in locations lacking monitoring stations by utilizing data from existing ones. By leveraging learnable mask tokens, AirRadar reconstructs air quality features in unmonitored regions. Specifically, it operates in two stages: first capturing spatial correlations and then adjusting for distribution shifts. We validate AirRadar’s efficacy using a year-long dataset from 1,085 monitoring stations across China, demonstrating its superiority over multiple baselines, even with varying degrees of unobserved data. Qiongyan Wang, Yutong Xia, Siru Zhong, Weichuang Li, Shifen Cheng, Junbo Zhang 0004, Yuxuan Liang 0002 |
AAAI | 3 |
| 2025 | Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series ForecastingabstractRecent advancements in time series forecasting have explored augmenting models with text or vision modalities to improve accuracy. While text provides contextual understanding, it often lacks fine-grained temporal details. Conversely, vision captures intricate temporal patterns but lacks semantic context, limiting the complementary potential of these modalities. To address this, we propose Time-VLM, a novel multimodal framework that leverages pre-trained Vision-Language Models (VLMs) to bridge temporal, visual, and textual modalities for enhanced forecasting. Our framework comprises three key components: (1) a Retrieval-Augmented Learner, which extracts enriched temporal features through memory bank interactions; (2) a Vision-Augmented Learner, which encodes time series as informative images; and (3) a Text-Augmented Learner, which generates contextual textual descriptions. These components collaborate with frozen pre-trained VLMs to produce multimodal embeddings, which are then fused with temporal features for final prediction. Extensive experiments demonstrate that Time-VLM achieves superior performance, particularly in few-shot and zero-shot scenarios, thereby establishing a new direction for multimodal time series forecasting. Code is available at https://github.com/CityMind-Lab/ICML25-TimeVLM. Siru Zhong, Weilin Ruan, Ming Jin 0005, Qingsong Wen, Yuxuan Liang 0002 |
ICML | 1 |
| 2025 | Fine-grained Urban Heat Island Effect Forecasting: A Context-aware Thermodynamic Modeling FrameworkabstractClimate change and rapid urbanization have led to the Urban Heat Island (UHI) effect, resulting in higher temperatures in metropolitan areas and negatively impacting urban communities. Accurate UHI forecasting is crucial for identifying high-risk periods and locations, especially in cities with vulnerable populations. Current methods are limited by data granularity and inadequate modeling of regional thermodynamics, which affects both accuracy and spatio-temporal granularity. In this paper, we propose DeepUHI, a data-driven context-aware framework for modeling local thermodynamics based on the heat equation, alongside the SeoulTemp dataset, the first multi-modal dataset for UHI effect predictions at the street level. Our framework utilizes a heat decomposition method to represent urban thermodynamics through thermodynamic cycles and thermal flows, effectively integrating urban environmental data. Extensive experiments show that our framework improves accuracy in UHI effect prediction and warning tasks, outperforming leading models. We have integrated DeepUHI into our SeoUHI platform to provide hourly street-level UHI forecasting for Seoul. The code, platform, and dataset are accessible at https://github.com/CityMind-Lab/DeepUHI. Xingchen Zou, Weilin Ruan, Siru Zhong, Yuehong Hu, Yuxuan Liang 0002 |
KDD (2) | 3 |
| 2025 | Towards Multi-Scenario Forecasting of Building Electricity Loads with Multimodal Data
Yongzheng Liu, Siru Zhong, Gefeng Luo, Weilin Ruan, Yuxuan Liang 0002 |
ACM Multimedia | 2 |
| 2025 | Multimodal Learning for Spatio-Temporal Data MiningabstractSpatio-temporal data mining (STDM) has become crucial in multimedia, driven by the surge of multimodal data from remote sensing, IoT sensors, social media, surveillance systems, mobile devices, and crowdsourced platforms. Traditional single-modal methods, though successful, struggle to capture real-world complexity. Integrating multiple modalities yields richer, more accurate insights, boosting spatio-temporal analysis. This half-day tutorial, MM4ST: Multimodal Learning for STDM, offers a comprehensive overview, covering STDM fundamentals, challenges in aligning and fusing heterogeneous data, advanced multimodal modeling techniques, and emerging research directions. Attendees will acquire practical knowledge to develop scalable and robust spatio-temporal mining solutions. All materials will be publicly available online. Siru Zhong, Xixuan Hao, Hao Miao 0001, Yan Zhao 0008, Qingsong Wen, Roger Zimmermann, Yuxuan Liang 0002 |
ACM Multimedia | 1 |
| 2025 | Learning to Factorize Spatio-Temporal Foundation ModelsabstractSpatio-Temporal Foundation Models (STFMs) promise zero/few-shot generalization across various datasets, yet joint spatio-temporal pretraining is computationally prohibitive and struggles with domain-specific spatial correlations. To this end, we introduce FactoST, a factorized STFM that decouples universal temporal pretraining from spatio-temporal adaptation. The first stage pretrains a space-agnostic backbone with multi-frequency reconstruction and domain-aware prompting, capturing cross-domain temporal regularities at low computational cost. The second stage freezes or further fine-tunes the backbone and attaches an adapter that fuses spatial metadata, sparsifies interactions, and aligns domains with continual memory replay. Extensive forecasting experiments reveal that, in few-shot setting, FactoST reduces MAE by up to 46.4% versus UniST, uses 46.2% fewer parameters, and achieves 68% faster inference than OpenCity, while remaining competitive with expert models. We believe this factorized view offers a practical and scalable path toward truly universal STFMs. The code will be released upon notification. Siru Zhong, Junjie Qiu, Yangyu Wu 0004, Xingchen Zou, Zhongwen Rao, Bin Yang 0002, Chenjuan Guo, Yuxuan Liang 0002 |
NeurIPS | 1 |
| 2025 | Cross Space and Time: A Spatio-Temporal Unitized Model for Traffic Flow ForecastingabstractPredicting spatio-temporal traffic flow presents significant challenges due to complex interactions between spatial and temporal factors. Existing approaches often address these dimensions in isolation, neglecting their critical interdependencies. In this paper, we introduce theSpatio-TemporalUnitizedModel (STUM), a unified framework designed to capture both spatial and temporal dependencies while addressing spatio-temporal heterogeneity through techniques such as distribution alignment and feature fusion. It also ensures both predictive accuracy and computational efficiency. Central to STUM is the Adaptive Spatio-temporal Unitized Cell (ASTUC), which utilizes low-rank matrices to seamlessly store, update, and interact with space, time, as well as their correlations. Our framework is also modular, allowing it to integrate with various spatio-temporal graph neural networks through components such as backbone models, feature extractors, residual fusion blocks, and the predictor to collectively enhance forecasting outcomes. Experimental results across multiple real-world datasets demonstrate that STUM consistently improves prediction performance with minimal computational cost. These findings are further supported by hyperparameter optimization, ablation studies, and result visualization. We provide our source code for reproducibility athttps://github.com/RWLinno/STUM Weilin Ruan, Wenzhuo Wang, Siru Zhong, Wei Chen 0070, Li Liu 0001, Yuxuan Liang 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Spatio-Temporal Field Neural Networks for Air Quality Inference
Yutong Feng, Qiongyan Wang, Yutong Xia, Siru Zhong, Yuxuan Liang 0002 |
IJCAI | 5 |
| 2024 | Predicting Carpark Availability in Singapore with Cross-Domain Data: A New Dataset and A Data-Driven Approach
Huaiwu Zhang, Yutong Xia, Siru Zhong, Kun Wang 0042, Zekun Tong, Qingsong Wen, Roger Zimmermann, Yuxuan Liang 0002 |
IJCAI | 3 |
| 2024 | UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain AdaptationabstractUrbanization challenges underscore the necessity for effective satellite image-text retrieval methods to swiftly access specific information enriched with geographic semantics for urban applications. However, existing methods often overlook significant domain gaps across diverse urban landscapes, primarily focusing on enhancing retrieval performance within single domains. To tackle this issue, we present UrbanCross, a new framework for cross-domain satellite image-text retrieval. UrbanCross leverages a high-quality, cross-domain dataset enriched with extensive geo-tags from three countries to highlight domain diversity. It employs the Large Multimodal Model (LMM) for textual refinement and the Segment Anything Model (SAM) for visual augmentation, achieving fine-grained alignment of images, segments and texts, yielding a 10% improvement in retrieval performance. Additionally, UrbanCross incorporates an adaptive curriculum-based source sampler and a weighted adversarial cross-domain fine-tuning module, progressively enhancing adaptability across various domains. Extensive experiments confirm UrbanCross's superior efficiency in retrieval and adaptation to new urban environments, demonstrating an average performance increase of 15% over its version without domain adaptation mechanisms, effectively bridging the domain gap. Our code and dataset are publicly accessible at https://github.com/siruzhong/UrbanCross. Siru Zhong, Xixuan Hao, Ying Zhang 0047, Yangqiu Song, Yuxuan Liang 0002 |
ACM Multimedia | 1 |
| 2024 | UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebabstractUrban region profiling from web-sourced data is of utmost importance for urban computing. We are witnessing a blossom of LLMs for various fields, especially in multi-modal data research such as vision-language learning, where text modality serves as a supplement for images. As textual modality has rarely been introduced into modality combinations in urban region profiling, we aim to answer two fundamental questions: i) Can text modality enhance urban region profiling? ii) and if so, in what ways and which aspects? To answer the questions, we leverage the power of Large Language Models (LLMs) and introduce the first-ever LLM-enhanced framework that integrates the knowledge of text modality into urban imagery, named LLM-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining (UrbanCLIP ). Specifically, it first generates a detailed textual description for each satellite image by Image-to-Text LLMs. Then, the model is trained on image-text pairs, seamlessly unifying language supervision for urban visual representation learning, jointly with contrastive loss and language modeling loss. Results on urban indicator prediction in four major metropolises show its superior performance, with an average improvement of 6.1% on R2 compared to the state-of-the-art methods. Our code and dataset are available at https://github.com/StupidBuluchacha/UrbanCLIP. Haomin Wen, Siru Zhong, Wei Chen 0070, Qingsong Wen, Roger Zimmermann, Yuxuan Liang 0002 |
WWW | 3 |