VLDB 2026 Research / reviewers in the wild / expert
Lei Bai 0001
dblp:119/1223-1
· DBLP profile ↗
147ranked-venue papers
8as first author
135since 2021 · last 2026
0000-0003-3378-7201ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 106 · 4 first-author · 99 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 1 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 17 since 2021Databases, data management, data science and information retrieval · 15 · 3 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynWeather: Weather Observation Data Synthesis Across Multiple Regions and Variables via a General Diffusion TransformerabstractWith the advancement of meteorological instruments, abundant data has become available. However, due to instruments’ intrinsic limitations such as environmental sensitivity and orbital constraints, raw data often suffer from temporal or spatial gaps, making it urgent to leverage data synthesis techniques to fill in missing information. Current approaches are typically focus on single-variable, single-region tasks and primarily rely on deterministic modeling. This limits unified synthesis across variables and regions, overlooks cross-variable complementarity and often leads to over-smoothed results. To address above challenges, we introduce SynWeather, the first dataset designed for Unified Multi-region and Multi-variable Weather Observation Data Synthesis. SynWeather covers four representative regions: the Continental United States, Europe, East Asia, and Tropical Cyclone regions, as well as provides high-resolution observations of key weather variables, including Composite Radar Reflectivity, Hourly Precipitation, Visible Light, and Microwave Brightness Temperature. In addition, we introduce SynWeatherDiff, a general and probabilistic weather synthesis model built upon the Diffusion Transformer framework to address the over-smoothed problem. Experiments on the SynWeather dataset demonstrate the effectiveness of our network compared with both task-specific and general models. Moreover, SynWeatherDiff is able to generate results that are both fine-grained and accurate in high-value regions. Through the dataset and baseline model, we aim to advance meteorological downstream tasks and promote the development of general models for weather variable synthesis. Kaiyi Xu, Junchao Gong, Zhiwang Zhou, Zhangrui Li, Yuandong Pu, Ben Fei, Fenghua Ling, Lei Bai 0001 |
AAAI | 10 |
| 2026 | The Avengers: A Routing Recipe for Collective Intelligence in Language ModelsabstractProprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, data efficiency, and values of its sole parameter---the number of clusters. Hao Li 0069, Linyao Chen, Qiaosheng Zhang 0002, Peng Ye 0006, Shi Feng 0001, Xinrun Wang, Xu Jia 0012, Lei Bai 0001, Shuyue Hu |
AAAI | 10 |
| 2026 | FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge FlowabstractYusong Hu, Runmin Ma, Yue Fan, Jinxin Shi, Zongsheng Cao, Yuhao Zhou, Jiakang Yuan, Shuaiyu Zhang, Shiyang Feng, Xiangchao Yan, Shufei Zhang, Wenlong Zhang, Lei Bai, Bo Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yusong Hu, Runmin Ma, Jinxin Shi, Zongsheng Cao, Yuhao Zhou 0005, Jiakang Yuan, Shuaiyu Zhang, Shiyang Feng, Xiangchao Yan, Shufei Zhang, Lei Bai 0001, Bo Zhang 0069 |
ACL (1) | 13 |
| 2026 | Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningabstractZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Yifan Zhou, Qiang He, Xiangyuan Xue, Heng Zhou, Yutao Fan, Zhong-Zhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang, Zhenfei Yin, Philip Torr, Lei Bai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Xiangyuan Xue, Yutao Fan, Zhongzhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang 0007, Zhenfei Yin, Philip Torr 0001, Lei Bai 0001 |
ACL (1) | 17 |
| 2026 | A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven EnhancementabstractShengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shengji Tang, Jianjian Cao, Weihao Lin 0002, Jiale Hong, Bo Zhang 0069, Shuyue Hu, Lei Bai 0001, Tao Chen 0003, Wanli Ouyang, Peng Ye 0006 |
ACL (1) | 7 |
| 2026 | R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement LearningabstractYiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lihao Wang, Lei Bai, Lijun Wu, Wei-Ying Ma, Hao Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. YiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lei Bai 0001, Lijun Wu 0003, Wei-Ying Ma, Hao Zhou 0012 |
ACL (1) | 7 |
| 2026 | MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint EmbeddingsabstractYiqun Zhang, Hao Li, Zihan Wang, Shi Feng, Xiaocui Yang, Daling Wang, Bo Zhang, Lei Bai, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Li 0069, Shi Feng 0001, Xiaocui Yang, Daling Wang, Bo Zhang 0069, Lei Bai 0001, Shuyue Hu |
ACL (1) | 8 |
| 2026 | Nature-Inspired Population-Based Evolution of Large Language ModelsabstractYiqun Zhang, Peng Ye, Xiaocui Yang, Shi Feng, Shufei Zhang, Lei Bai, Wanli Ouyang, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peng Ye 0006, Xiaocui Yang, Shi Feng 0001, Shufei Zhang, Lei Bai 0001, Wanli Ouyang, Shuyue Hu |
ACL (1) | 6 |
| 2026 | MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMsabstractXiangyu Zhao, Wanghan Xu, Bo Liu, Yuhao Zhou, Fenghua Ling, Ben Fei, Xiaoyu Yue, Lei Bai, Wenlong Zhang, Xiao-Ming Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wanghan Xu, Yuhao Zhou 0005, Fenghua Ling, Ben Fei, Xiaoyu Yue, Lei Bai 0001 |
ACL (1) | 8 |
| 2026 | Modeling Point-to-Point Dependency for High-Dimensional Long-Term Series Forecasting
Xinyu Li 0014, Kexi Chen, Ying Zheng 0004, Zhiyi Yao, Yi Xie 0003, Jihan Dai, Lei Bai 0001, Jin Zhao 0001, Jiajie Shen, Yunqi Cai, Hong Lu 0001, Xin Wang 0002 |
WWW | 7 |
| 2026 | Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
Chongjun Tu, Peng Ye 0006, Dongzhan Zhou, Lei Bai 0001, Gang Yu 0002, Tao Chen 0003, Wanli Ouyang |
Int. J. Comput. Vis. | 4 |
| 2026 | $\beta $-DARTS++: Bi-Level Regularization for Proxy-Robust Differentiable Architecture SearchabstractNeural Architecture Search (NAS) has attracted increasing attention in recent years because of its capability to design neural networks automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for search efficiency. However, they still suffer from three main issues, that are, the weak stability due to the performance collapse, the poor generalization ability of the searched architectures, and the inferior robustness to different kinds of proxies (i.e., computationally reduced search configurations). To solve the search stability and searched architecture's generalization problems, a simple-but-effective regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process (referred as $\beta$β-DARTS). Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from being too large, thereby ensuring fair competition among architecture parameters and making the supernet less sensitive to the impact of input on the operation set. In-depth theoretical analyses on how it works and why it works are provided, and comprehensive experiments on a variety of search spaces and datasets validate that Beta-Decay regularization can help to stabilize the searching process and make the searched network more transferable across different datasets. To address the proxy robustness problem, we first benchmark differentiable NAS methods under a wide range of proxy data, proxy channels, proxy layers, and proxy epochs, since the robustness of NAS under different kinds of proxies has not been explored before. We then conclude some interesting findings and find that $\beta$β-DARTS always achieves the best result among all compared NAS methods under almost all proxy settings. We further introduce the novel flooding regularization to the weight optimization of $\beta$β-DARTS (termed as Bi-level regularization), and experimentally and theoretically verify its effectiveness for improving the proxy robustness of differentiable NAS. Peng Ye 0006, Tong He 0001, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | A renaissance of explicit motion information mining from transformers for action recognition
Peiqin Zhuang, Lei Bai 0001, Yichao Wu, Ding Liang, Luping Zhou, Yali Wang 0001, Wanli Ouyang |
Pattern Recognit. | 2 |
| 2026 | SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference OptimizationabstractPrecipitation nowcasting based on radar echoes plays a crucial role in monitoring extreme weather and supporting disaster prevention. Although deep learning approaches have achieved significant progress, they still face notable limitations. For example, deterministic models tend to produce over-smoothed predictions, which struggle to capture extreme events and fine-scale precipitation patterns. Probabilistic generative models, due to their inherent randomness, often show fluctuating performance across different metrics and rarely achieve consistently optimal results. Furthermore, precipitation nowcasting is typically evaluated using multiple metrics, some of which are inherently conflicting. For instance, there is often a trade-off between the Critical Success Index (CSI) and the False Alarm Ratio (FAR), making it challenging for existing models to deliver forecasts that perform well on both metrics simultaneously. To address these challenges, we introduce preference optimization into precipitation nowcasting for the first time, motivated by the success of reinforcement learning from human feedback in large language models. Specifically, we propose SynCast, which lever-ages the two-stage post-training framework of Diffusion Sequential Preference Optimization (Diffusion-SPO) to progressively align conflicting metrics. In the first stage, the framework focuses on reducing FAR to deliver clean and high-precision predictions. Building on this foundation, the second stage further optimizes CSI under strict FAR constraints, thereby achieving synergistic improvements across these conflicting metrics. Experiments on three radar precipitation datasets demonstrate that SynCast reduces FAR while improving CSI, and achieves performance comparable to state-of-the-art methods. Furthermore, we verify that the post-training framework of Diffusion-SPO is compatible with multiple diffusion models for precipitation nowcasting, demonstrating its generalizability. The code for SynCast is available at https://github.com/Dtdtxuky/SynCast. Kaiyi Xu, Junchao Gong, Ben Fei, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | TRACK: Temporal Decoupled Kriging for Inductive Spatio-Temporal GraphabstractThe deployment of sensors enables data-driven urban management, but necessitates inductive spatio-temporal kriging to infer unmonitored areas. Existing methods impute these unknown observations by smoothing temporal features based on spatial dependencies, overlooking the decoupling ofinherent propertiesanddynamic correlationsin message passing. In particular, the inherent properties reveal non-transitive signals, and current coupled aggregation leads to inaccurate results. To this end, we proposeTempoRAl deCoupledKriging, named TRACK, to decouple two factors with the help of node-specific inherency. Specifically, we first construct a node-specific profile to represent its inherency including geographical and periodic features, which is subsequently transformed into decoupling prompts. Secondly, the coupled temporal features are separated through querying each prompt embedding, facilitating precise temporal aggregation for inherent properties and spatial aggregation for dynamic correlations. Finally, a multi-task training strategy is further adopted to mimic the inductive scenarios during testing. We evaluate TRACK on four real-world datasets spanning urban traffic and air quality prediction tasks. TRACK achieves state-of-the-art performance, with average improvements of 3.10% in MAE and 4.45% in RMSE over strong baselines. Moreover, we further demonstrated its robust generalization in a challenging cross-city inductive setting. Code is available athttps://github.com/JeremyChou28/TRACK. Jianping Zhou 0004, Weida Wang, Bin Lu 0005, Guanjie Zheng, Lei Bai 0001, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical ConstraintsabstractTropical cyclone (TC) intensity forecasting is crucial for early disaster warning and emergency decision-making. Numerous researchers have explored deep-learning methods to address computational and post-processing issues in operational forecasting. Regrettably, they exhibit subpar long-term forecasting capabilities. We use two strategies to enhance long-term forecasting. (1) By enhancing the matching between TC intensity and spatial information, we can improve long-term forecasting performance. (2) Incorporating physical knowledge and physical constraints can help mitigate the accumulation of forecasting errors. To achieve the above strategies, we propose the VQLTI framework. VQLTI transfers the TC intensity information to a discrete latent space while retaining the spatial information differences, using large-scale spatial meteorological data as conditions. Furthermore, we leverage the forecast from the weather prediction model FengWu to provide additional physical knowledge for VQLTI. Additionally, we calculate the potential intensity (PI) to impose physical constraints on the latent variables. In the global long-term TC intensity forecasting, VQLTI achieves state-of-the-art results for the 24h to 120h, with the MSW (Maximum Sustained Wind) forecast error reduced by 35.65%-42.51% compared to ECMWF-IFS. Lei Liu 0029, Tao Han 0002, Bin Li 0025, Lei Bai 0001 |
AAAI | 6 |
| 2025 | SURVEYFORGE : On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey WritingabstractSurvey paper plays a crucial role in scientific research, especially given the rapid growth of research publications. Recently, researchers have begun using LLMs to automate survey generation for better efficiency. However, the quality gap between LLM-generated surveys and those written by human remains significant, particularly in terms of outline quality and citation accuracy. To close these gaps, we introduce SURVEYFORGE, which first generates the outline by analyzing the logical structure of human-written outlines and referring to the retrieved domain-related articles. Subsequently, leveraging high-quality papers retrieved from memory by our scholar navigation agent, SURVEYFORGE can automatically generate and refine the content of the generated article. Moreover, to achieve a comprehensive evaluation, we construct SurveyBench, which includes 100 human-written survey papers for win-rate comparison and assesses AI-generated survey papers across three dimensions: reference, outline, and content quality. Experiments demonstrate that SURVEYFORGEcan outperform previous works such as AutoSurvey. Xiangchao Yan, Shiyang Feng, Jiakang Yuan, Renqiu Xia, Bin Wang 0065, Lei Bai 0001, Bo Zhang 0069 |
ACL (1) | 6 |
| 2025 | Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and FeedbackabstractJiakang Yuan, Xiangchao Yan, Bo Zhang, Tao Chen, Botian Shi, Wanli Ouyang, Yu Qiao, Lei Bai, Bowen Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiakang Yuan, Xiangchao Yan, Bo Zhang 0069, Tao Chen 0003, Botian Shi, Wanli Ouyang, Yu Qiao 0001, Lei Bai 0001, Bowen Zhou 0002 |
ACL (1) | 8 |
| 2025 | UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesabstractTraditional spatiotemporal models generally rely on task-specific architectures, which limit their generalizability and scalability across diverse tasks due to domain-specific design requirements. In this paper, we introduce UniSTD, a unified Transformer-based framework for spatiotemporal modeling, which is inspired by advances in recent foundation models with the two-stage pretraining-then-adaption paradigm. Specifically, our work demonstrates that task-agnostic pretraining on 2D vision and vision-text datasets can build a generalizable model foundation for spatiotemporal learning, followed by specialized joint training on spatiotemporal datasets to enhance task-specific adaptability. To improve the learning capabilities across domains, our framework employs a rank-adaptive mixture-of-expert adaptation by using fractional interpolation to relax the discrete variables so that can be optimized in the continuous space. Additionally, we introduce a temporal module to incorporate temporal dynamics explicitly. We evaluate our approach on a large-scale dataset covering 10 tasks across 4 disciplines, demonstrating that a unified spatiotemporal model can achieve scalable, cross-task learning and support up to 10 tasks simultaneously within one model while reducing training costs in multi-domain applications. Code will be available at https://github.com/1hunters/UniSTD. Xinzhu Ma, Encheng Su, Xiufeng Song, Xiaohong Liu 0001, Wei-Hong Li 0001, Lei Bai 0001, Wanli Ouyang, Xiangyu Yue 0001 |
CVPR | 7 |
| 2025 | Satellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary ResolutionabstractAccurate acquisition of surface meteorological conditions at arbitrary locations holds significant importance for weather forecasting and climate simulation. Meteorological states derived from satellite observations are often provided in the form of low-resolution grid fields. If spatial interpolation is applied directly to obtain meteorological states for specific locations, there will often be significant discrepancies compared to actual observations. Existing downscaling methods for acquiring meteorological state information at higher resolutions commonly overlook the correlation with satellite observations. To bridge the gap, we propose Satellite-observations Guided Diffusion Model (SGD), a conditional diffusion model pre-trained on ERA5 reanalysis data with satellite observations (GridSat) as conditions, which is employed for sampling downscaled meteorological states through a zero-shot guided sampling strategy and patch-based methods. During the training process, we propose to fuse the information from GridSat satellite observations into ERA5 maps via the attention mechanism, enabling SGD to generate atmospheric states that align more accurately with actual conditions. In the sampling, we employed optimizable convolutional kernels to simulate the upscale process, thereby generating high-resolution ERA5 maps using low-resolution ERA5 maps as well as observations from weather stations as guidance. Moreover, our devised patch-based method promotes SGD to generate meteorological states at arbitrary resolutions. Experiments demonstrate SGD fulfills accurate meteorological states downscaling to 6.25km. The code is available at https://github.com/Tusiwei/SGD Siwei Tu, Ben Fei, Weidong Yang 0001, Fenghua Ling, Hao Chen 0045, Kun Chen 0004, Hang Fan, Wanli Ouyang, Lei Bai 0001 |
CVPR | 10 |
| 2025 | IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion PriorabstractVariation of Arctic sea ice has significant impacts on polar ecosystems, transporting routes, coastal communities, and global climate. Tracing the change of sea ice at a finer scale is paramount for both operational applications and scientific studies. Recent pan-Arctic sea ice forecasting methods that leverage advances in artificial intelligence have made promising progress over numerical models. However, forecasting sea ice at higher resolutions is still under-explored. To bridge the gap, we propose a two-module cooperative deep learning framework, IceDiff, to forecast sea ice concentration at finer scales. IceDiff first leverages a vision transformer to generate coarse yet superior forecasting results over previous methods at a regular 25 km grid. This high-quality sea ice forecasting can be utilized as reliable guidance for the next module. Subsequently, an unconditional diffusion model pre-trained on low-resolution sea ice concentration maps is utilized for sampling down-scaled sea ice forecasting via a zero-shot guided sampling strategy and a patch-based method. For the first time, IceDiff demonstrates sea ice forecasting with a 6.25 km resolution. IceDiff extends the boundary of existing sea ice forecasting models and more importantly, its capability to generate high-resolution sea ice concentration data is vital for pragmatic usages and research. Code is available at https://github.com/EtronTech/IceDiff. Siwei Tu, Weidong Yang 0001, Ben Fei, Shuhao Li 0001, Keyi Liu, Yeqi Luo, Lipeng Ma, Lei Bai 0001 |
CVPR | 9 |
| 2025 | ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI SystemsabstractMuch previous AI research has focused on developing monolithic models to maximize their intelligence, with the primary goal of enhancing performance on specific tasks. In contrast, this work attempts to study using LLM-based agents to design collaborative AI systems autonomously. To explore this problem, we first introduce ComfyBench to evaluate agents’s ability to design collaborative AI systems in ComfyUI. ComfyBench is a comprehensive benchmark comprising 200 diverse tasks covering various instruction-following generation challenges, along with detailed annotations for 3,205 nodes and 20 workflows. Based on ComfyBench, we further develop ComfyAgent, a novel framework that empowers LLM-based agents to autonomously design collaborative AI systems by generating workflows. ComfyAgent is based on two core concepts. First, it represents workflows with code, which can be reversibly converted into workflows and executed as collaborative systems by the interpreter. Second, it constructs a multi-agent system that cooperates to learn from existing workflows and generate new workflows for a given task. While experimental results demonstrate that ComfyAgent achieves a comparable resolve rate to o1-preview and significantly surpasses other agents on ComfyBench, ComfyAgent has resolved only 15% of creative tasks. LLM-based agents still have a long way to go in autonomously designing collaborative AI systems. Progress with ComfyBench is paving the way for more intelligent and autonomous collaborative AI systems. Our code is available at: https://github.com/xxyQwQ/ComfyBench. Xiangyuan Xue, Zidong Wang 0004, Wanli Ouyang, Lei Bai 0001 |
CVPR | 6 |
| 2025 | Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized RoutingabstractBalancing performance and efficiency is a central challenge in large language model (LLM) advancement. GPT-5 addresses this with test-time routing, dynamically assigning queries to either an efficient or a high-capacity model during inference. In this work, we present Avengers-Pro, a test-time routing framework that ensembles LLMs of varying capacities and efficiencies, providing a unified solution for all performance-efficiency tradeoffs. The Avengers-Pro embeds and clusters incoming queries, then routes each to the most suitable model based on a performance-efficiency score. Across 6 challenging benchmarks and 8 leading models—including GPT-5-medium, Gemini-2.5-pro, and Claude-opus-4.1—Avengers-Pro achieves state-of-the-art results: by varying a performance-efficiency trade-off parameter, it can surpass the strongest single model (GPT-5-medium) by +7% in average accuracy. Moreover, it can match the average accuracy of the strongest single model at 27% lower cost, and reach ∼ 90% of that performance at 63% lower cost. Last but not least, it achieves a Pareto frontier, consistently yielding the highest accuracy for any given cost, and the lowest cost for any given accuracy, among all single models. Code is available at https://github.com/ZhangYiqun018/AvengersPro. Hao Li 0069, Jianhao Chen 0001, Hangfan Zhang, Peng Ye 0006, Lei Bai 0001, Shuyue Hu |
DAI | 6 |
| 2025 | ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning TasksabstractMulti-agent systems have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving.However, current MAS frameworks are limited by poor flexibility and scalability, with underdeveloped optimization strategies.To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process.The core of ReSo is the proposed Collaborative Reward Model, which can provide fine-grained reward signals for MAS cooperation for optimization.We also introduce an automated data synthesis framework for generating MAS benchmarks, without human annotations.Experimentally, ReSo matches or outperforms existing methods.ReSo achieves 33.7% and 32.3% accuracy on Math-MAS and SciBench-MAS SciBench, while other methods completely fail.The code and data are available at Reso. Hejia Geng, Xiangyuan Xue, Yiran Qin, Zhiyong Wang 0001, Zhenfei Yin, Lei Bai 0001 |
EMNLP | 8 |
| 2025 | DiffSR: Learning Radar Reflectivity Synthesis via Diffusion Model from Satellite ObservationsabstractWeather radar data synthesis can fill in data for areas where ground observations are missing. Existing methods often employ reconstruction-based approaches with MSE loss to reconstruct radar data from satellite observation. However, such methods lead to over-smoothing, which hinders the generation of high-frequency details or high-value observation areas associated with convective weather. To address this issue, we propose a two-stage diffusion-based method called DiffSR. We first pretrain a reconstruction model on global-scale data to obtain radar estimation and then synthesize radar reflectivity by combining radar estimation results with satellite data as conditions for the diffusion model. Extensive experiments show that our method achieves state-of-the-art (SOTA) results, demonstrating the ability to generate high-frequency details and high-value areas. Zhiwang Zhou, Hao Chen 0045, Lei Bai 0001 |
ICASSP | 7 |
| 2025 | Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive ModelabstractAccurate forecasting of tropical cyclone (TC) intensity is crucial for formulating disaster risk reduction strategies. Current methods predominantly rely on limited spatiotemporal information from ERA5 data and neglect the causal relationships between these physical variables, failing to fully capture the spatial and temporal patterns required for intensity forecasting. To address this issue, we propose a Multi-modal multi-Scale Causal AutoRegressive model (MSCAR), which is the first model that combines causal relationships with large-scale multimodal data for global TC intensity autoregressive forecasting. Furthermore, given the current absence of a TC dataset that offers a wide range of spatial variables, we present the Satellite and ERA5-based Tropical Cyclone Dataset (SETCD), which stands as the longest and most comprehensive global dataset related to TCs. Experiments on the dataset show that MSCAR outperforms the state-of-the-art methods, achieving maximum reductions in global and regional forecast errors of 9.52% and 6.74%, respectively. The code and dataset are publicly available at https://github.com/1457756434/MSCAR.git. Lei Liu 0029, Tao Han 0002, Bin Li 0025, Lei Bai 0001 |
ICASSP | 6 |
| 2025 | VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather ForecastingabstractThis paper presents Variables-Adaptive Mixture of Experts (VA-MoE), a novel framework for incremental weather forecasting that dynamically adapts to evolving spatiotemporal patterns in real-time data. Traditional weather prediction models often struggle with exorbitant computational expenditure and the need to continuously update forecasts as new observations arrive. VA-MoE addresses these challenges by leveraging a hybrid architecture of experts, where each expert specializes in capturing distinct sub-patterns of atmospheric variables (e.g., temperature, humidity, wind speed). Moreover, the proposed method employs a variable-adaptive gating mechanism to dynamically select and combine relevant experts based on the input context, enabling efficient knowledge distillation and parameter sharing. This design significantly reduces computational overhead while maintaining high forecast accuracy. Experiments on real-world ERA5 dataset demonstrate that VA-MoE performs comparable against state-of-the-art models in both short-term (e.g., 1–3 days) and long-term (e.g., 5 days) forecasting tasks, with only about 25\% of trainable parameters and 50\% of the initial training data. Hao Chen 0045, Tao Han 0002, Song Guo 0001, Jie Zhang 0076, Yonghan Dong, Yue Yu 0001, Lei Bai 0001 |
ICCV | 7 |
| 2025 | InfGen: A Resolution-Agnostic Paradigm for Scalable Image SynthesisabstractArbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution, causing 4K image generation delays over 100 seconds. To solve this, we explore the second generation upon the latent diffusion models, where the fixed latent generated by diffusion models is regarded as the content representation and we propose to decode arbitrary resolution images with a compact generated latent using a one-step generator. Thus, we present the \textbf{InfGen}, replacing the VAE decoder with the new generator, for generating images at any resolution from a fixed-size latent without retraining the diffusion models, which simplifies the process, reducing computational complexity and can be applied to any model using the same latent space. Experiments show InfGen is capable of improving many models into the arbitrary high-resolution era while cutting 4K image generation time to under 10 seconds. Tao Han 0002, Wanghan Xu, Junchao Gong, Xiaoyu Yue, Song Guo 0001, Luping Zhou, Lei Bai 0001 |
ICCV | 7 |
| 2025 | Chimera: Improving Generalist Model with Domain-Specific ExpertsabstractRecent advancements in Large Multi-modal Models (LMMs) underscore the importance of scaling by increasing image-text paired data, achieving impressive performance on general tasks. Despite their effectiveness in broad applications, generalist models are primarily trained on web-scale datasets dominated by natural images, resulting in the sacrifice of specialized capabilities for domain-specific tasks that require extensive domain prior knowledge. Moreover, directly integrating expert models tailored for specific domains is challenging due to the representational gap and imbalanced optimization between the generalist model and experts. To address these challenges, we introduce Chimera, a scalable and low-cost multi-modal pipeline designed to boost the ability of existing LMMs with domain-specific experts. Specifically, we design a progressive training strategy to integrate features from expert models into the input of a generalist LMM. To address the imbalanced optimization caused by the well-aligned general visual encoder, we introduce a novel Generalist-Specialist Collaboration Masking (GSCM) mechanism. This results in a versatile model that excels across the chart, table, math, and document domains, achieving state-of-the-art performance on multi-modal reasoning and visual content extraction tasks, both of which are challenging tasks for assessing existing LMMs. Tianshuo Peng, Mingsheng Li, Jiakang Yuan, Hongbin Zhou, Renqiu Xia, Renrui Zhang, Lei Bai 0001, Song Mao, Bin Wang 0065, Aojun Zhou, Botian Shi, Tao Chen 0003, Bo Zhang 0069, Xiangyu Yue 0001 |
ICCV | 7 |
| 2025 | RoboFactory: Exploring Embodied Agent Collaboration with Compositional ConstraintsabstractDesigning effective embodied multi-agent systems is critical for solving complex real-world tasks across domains. Due to the complexity of multi-agent embodied systems, existing methods fail to automatically generate safe and efficient training data for such systems. To this end, we propose the concept of compositional constraints for embodied multi-agent systems, addressing the challenges arising from collaboration among embodied agents. We design various interfaces tailored to different types of constraints, enabling seamless interaction with the physical world. Leveraging compositional constraints and specifically designed interfaces, we develop an automated data collection framework for embodied multi-agent systems and introduce the first benchmark for embodied multi-agent manipulation, RoboFactory. Based on RoboFactory benchmark, we adapt and evaluate the method of imitation learning and analyzed its performance in different difficulty agent tasks. Furthermore, we explore the architectures and training strategies for multi-agent imitation learning, aiming to build safe and efficient embodied multi-agent systems. Yiran Qin, Xiufeng Song, Zhenfei Yin, Xiaohong Liu 0001, Xihui Liu, Ruimao Zhang, Lei Bai 0001 |
ICCV | 8 |
| 2025 | VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorabstractVideo diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation. Xindi Yang, Baolu Li 0001, Zhenfei Yin, Lei Bai 0001, Liqian Ma, Zhiyong Wang 0001, Jianfei Cai 0001, Tien-Tsin Wong, Huchuan Lu, Xu Jia 0012 |
ICCV | 5 |
| 2025 | DIFFODE: Neural ODE with Differentiable Hidden State for Irregular Time Series AnalysisabstractIrregular time series analysis is increasingly essential in data management due to the proliferation of complex data irregularly sampled by real-world systems. Traditional time series models, including RNN-based models and transformer variants, face significant challenges in generalizing to continuous-time paradigms, which are essential for capturing the ongoing dynamics of irregular time series. Neural Ordinary Differential Equations (NODEs) assume a continuous latent dynamic and provide an elegant framework for irregular time series analysis, yet they suffer from limitations like fragmented latent processes and the inability to fully exploit interdependencies among observations. To address these challenges, we propose a novel Differentiable hidden state enhanced neural ODE framework, termed DIFFODE, designed to effectively model irregular time series. Concretely, we introduce an attention-based differential hidden state that maps irregular observations into a continuous hidden state space, enabling the extraction of latent dynamics while preserving temporal continuity. Leveraging the theory of generalized inverses, DIFFODE innovatively derives ODEs to describe hidden state dynamics. Furthermore, we incorporate the Hoyer metric into our framework to enhance its capacity to capture subtle yet critical temporal shifts, significantly improving the accuracy of time series modeling. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of DIFFODE across three key tasks, including irregular time series classification, interpolation, and extrapolation. Yudong Zhang 0005, Xu Wang 0029, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
ICDE | 6 |
| 2025 | PostCast: Generalizable Postprocessing for Precipitation Nowcasting via Unsupervised Blurriness ModelingabstractPrecipitation nowcasting plays a pivotal role in socioeconomic sectors, especially in severe convective weather warnings. Although notable progress has been achieved by approaches mining the spatiotemporal correlations with deep learning, these methods still suffer severe blurriness as the lead time increases, which hampers accurate predictions for extreme precipitation. To alleviate blurriness, researchers explore generative methods conditioned on blurry predictions. However, the pairs of blurry predictions and corresponding ground truth need to be given in advance, making the training pipeline cumbersome and limiting the generality of generative models within blurry modes that appear in training data. By rethinking the blurriness in precipitation nowcasting as a blur kernel acting on predictions, we propose an unsupervised postprocessing method to eliminate the blurriness without the requirement of training with the pairs of blurry predictions and corresponding ground truth. Specifically, we utilize blurry predictions to guide the generation process of a pre-trained unconditional denoising diffusion probabilistic model (DDPM) to obtain high-fidelity predictions with eliminated blurriness. A zero-shot blur kernel estimation mechanism and an auto-scale denoise guidance strategy are introduced to adapt the unconditional DDPM to any blurriness modes varying from datasets and lead times in precipitation nowcasting. Extensive experiments are conducted on 7 precipitation radar datasets, demonstrating the generality and superiority of our method. Junchao Gong, Siwei Tu, Weidong Yang 0001, Ben Fei, Kun Chen 0004, Xiaokang Yang 0001, Wanli Ouyang, Lei Bai 0001 |
ICLR | 9 |
| 2025 | TimeKAN: KAN-based Frequency Decomposition Learning Architecture for Long-term Time Series ForecastingabstractReal-world time series often have multiple frequency components that are intertwined with each other, making accurate time series forecasting challenging. Decomposing the mixed frequency components into multiple single frequency components is a natural choice. However, the information density of patterns varies across different frequencies, and employing a uniform modeling approach for different frequency components can lead to inaccurate characterization. To address this challenges, inspired by the flexibility of the recent Kolmogorov-Arnold Network (KAN), we propose a KAN-based Frequency Decomposition Learning architecture (TimeKAN) to address the complex forecasting challenges caused by multiple frequency mixtures. Specifically, TimeKAN mainly consists of three components: Cascaded Frequency Decomposition (CFD) blocks, Multi-order KAN Representation Learning (M-KAN) blocks and Frequency Mixing blocks. CFD blocks adopt a bottom-up cascading approach to obtain series representations for each frequency band. Benefiting from the high flexibility of KAN, we design a novel M-KAN block to learn and represent specific temporal patterns within each frequency band. Finally, Frequency Mixing blocks is used to recombine the frequency bands into the original format. Extensive experimental results across multiple real-world time series datasets demonstrate that TimeKAN achieves state-of-the-art performance as an extremely lightweight architecture. Code is available at https://github.com/huangst21/TimeKAN. Songtao Huang, Zhen Zhao 0001, Can Li 0014, Lei Bai 0001 |
ICLR | 4 |
| 2025 | VAE-Var: Variational Autoencoder-Enhanced Variational Methods for Data Assimilation in MeteorologyabstractData assimilation (DA) is an essential statistical technique for generating accurate estimates of a physical system's states by combining prior model predictions with observational data, especially in the realm of weather forecasting. Effectively modeling the prior distribution while adapting to diverse observational sources presents significant challenges for both traditional and neural network-based DA algorithms. This paper introduces VAE-Var, a novel neural network-based data assimilation algorithm aimed at 1) enhancing accuracy by capturing the non-Gaussian characteristics of the conditional background distribution $p(\mathbf{x}|\mathbf{x}_b)$, and 2) efficiently assimilating real-world observational data. VAE-Var utilizes a variational autoencoder to learn the background error distribution, with its decoder creating a variational cost function to optimize the analysis states. The advantages of VAE-Var include: 1) it maintains the framework of traditional variational assimilation, enabling it to accommodate various observation operators, particularly irregular observations; 2) it lessens the dependence on expert knowledge for constructing the background distribution, allowing for improved modeling of non-Gaussian structures; and 3) experimental results indicate that, when applied to the FengWu weather forecasting model, VAE-Var outperforms DiffDA and two traditional algorithms (interpolation and 3DVar) in terms of assimilation accuracy in sparse observational contexts, and is capable of assimilating real-world GDAS prepbufr observations over a year. Qilong Jia, Kun Chen 0004, Lei Bai 0001, Wei Xue 0003 |
ICLR | 4 |
| 2025 | WeatherGFM: Learning a Weather Generalist Foundation Model via In-context LearningabstractThe Earth's weather system involves intricate weather data modalities and diverse weather understanding tasks, which hold significant value to human life.
Existing data-driven models focus on single weather understanding tasks (e.g., weather forecasting).
While these models have achieved promising results, they fail to tackle various complex tasks within a single and unified model.
Moreover, the paradigm that relies on limited real observations for a single scenario hinders the model's performance upper bound.
Inspired by the in-context learning paradigm from visual foundation models and large language models, in this paper, we introduce the first generalist weather generalist foundation model (WeatherGFM) to address weather understanding tasks in a unified manner.
Specifically, we first unify the representation and definition for diverse weather understanding tasks.
Subsequently, we design weather prompt formats to handle different weather data modalities, including single, multiple, and temporal modalities.
Finally, we adopt a visual prompting question-answering paradigm for the training of unified weather understanding tasks.
Extensive experiments indicate that our WeatherGFM can effectively handle up to 12 weather understanding tasks, including weather forecasting, super-resolution, weather image translation, and post-processing. Our method also showcases generalization ability on unseen tasks. The source code is available at https://github.com/xiangyu-mm/WeatherGFM. Zhiwang Zhou, Junchao Gong, Hao Chen 0045, Ben Fei, Wanli Ouyang, Lei Bai 0001 |
ICLR | 12 |
| 2025 | WorldSimBench: Towards Video Generation Models as World SimulatorsabstractRecent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing benchmarks are unable to effectively evaluate higher-capability, highly embodied predictive models from an embodied perspective. In this work, we classify the functionalities of predictive models into a hierarchy and take the first step in evaluating World Simulators by proposing a dual evaluation framework called WorldSimBench. WorldSimBench includes Explicit Perceptual Evaluation and Implicit Manipulative Evaluation, encompassing human preference assessments from the visual perspective and action-level evaluations in embodied tasks, covering three representative embodied scenarios: Open-Ended Embodied Environment, Autonomous, Driving, and Robot Manipulation. In the Explicit Perceptual Evaluation, we introduce the HF-Embodied Dataset, a video assessment dataset based on fine-grained human feedback, which we use to train a Human Preference Evaluator that aligns with human perception and explicitly assesses the visual fidelity of World Simulater. In the Implicit Manipulative Evaluation, we assess the video-action consistency of World Simulators by evaluating whether the generated situation-aware video can be accurately translated into the correct control signals in dynamic environments. Our comprehensive evaluation offers key insights that can drive further innovation in video generation models, positioning World Simulators as a pivotal advancement toward embodied artificial intelligence. Yiran Qin, Zhelun Shi, Jiwen Yu, Enshen Zhou, Zhenfei Yin, Xihui Liu, Lu Sheng, Lei Bai 0001, Ruimao Zhang |
ICML | 11 |
| 2025 | Multi-agent Architecture Search via Agentic SupernetabstractLarge Language Model (LLM)-empowered multi-agent systems extend the cognitive boundaries of individual agents through disciplined collaboration and interaction, while constructing these systems often requires labor-intensive manual designs. Despite the availability of methods to automate the design of agentic workflows, they typically seek to identify a static, complex, one-size-fits-all system, which, however, fails to dynamically allocate inference resources based on the difficulty and domain of each query. To address this challenge, we shift away from the pursuit of a monolithic agentic system, instead optimizing the \textbf{agentic supernet}, a probabilistic and continuous distribution of agentic architectures. We introduce \textbf{MaAS}, an automated framework that samples query-dependent agentic systems from the supernet, delivering high-quality solutions and tailored resource allocation (\textit{e.g.}, LLM calls, tool calls, token cost). Comprehensive evaluation across six benchmarks demonstrates that MaAS \textbf{(I)} requires only $6\\sim45\\%$ of the inference costs of existing handcrafted or automated multi-agent systems, \textbf{(II)} surpasses them by $0.54\\%\sim11.82\\%$, and \textbf{(III)} enjoys superior cross-dataset and cross-LLM-backbone transferability. Guibin Zhang, Luyang Niu, Junfeng Fang, Kun Wang 0056, Lei Bai 0001, Xiang Wang 0010 |
ICML | 5 |
| 2025 | DAWP: A framework for global observation forecasting via Data Assimilation and Weather Prediction in satellite observation spaceabstractWeather prediction is a critical task for human society, where impressive progress has been made by training artificial intelligence weather prediction (AIWP) methods with reanalysis data.
However, reliance on reanalysis data limits the AIWPs with shortcomings, including data assimilation biases and temporal discrepancies.
To liberate AIWPs from the reanalysis data, observation forecasting emerges as a transformative paradigm for weather prediction.
One of the key challenges in observation forecasting is learning spatiotemporal dynamics across disparate measurement systems with irregular high-resolution observation data, which constrains the design and prediction of AIWPs.
To this end, we propose our DAWP as an innovative framework to enable AIWPs to operate in a complete observation space by initialization with an artificial intelligence data assimilation (AIDA) module.
Specifically, our AIDA module applies a mask multi-modality autoencoder (MMAE) for assimilating irregular satellite observation tokens encoded by mask ViT-VAEs.
For AIWP, we introduce a spatiotemporal decoupling transformer with cross-regional boundary conditioning (CBC), learning the dynamics in observation space, to enable sub-image-based global observation forecasting.
Comprehensive experiments demonstrate that AIDA initialization significantly improves the roll-out and efficiency of AIWP.
Additionally, we show that DAWP holds promising potential to be applied in global precipitation forecasting. Junchao Gong, Ben Fei, Fenghua Ling, Kun Chen 0004, Wanghan Xu, Weidong Yang 0001, Xiaokang Yang 0001, Lei Bai 0001 |
NeurIPS | 10 |
| 2025 | RadarQA: Multi-modal Quality Analysis of Weather Radar ForecastsabstractQuality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evolution. With the rapid development of Multi-modal Large Language Models (MLLMs), these models become potential tools to overcome the above challenges. In this work, we introduce an MLLM-based weather forecast analysis method, RadarQA, integrating key physical attributes with detailed assessment reports. We introduce a novel and comprehensive task paradigm for multi-modal quality analysis, encompassing both single frame and sequence, under both rating and assessment scenarios. To support training and benchmarking, we design a hybrid annotation pipeline that combines human expert labeling with automated heuristics. With such an annotation method, we construct RQA-70K, a large-scale dataset with varying difficulty levels for radar forecast quality evaluation. We further design a multi-stage training strategy that iteratively improves model performance at each stage. Extensive experiments show that RadarQA outperforms existing general MLLMs across all evaluation settings, highlighting its potential for advancing quality analysis in weather prediction. Zhiyuan You, Junchao Gong, Couhua Liu, Xiaoyu Yue, Peiqin Zhuang, Lei Bai 0001 |
NeurIPS | 8 |
| 2025 | VIKI‑R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement LearningabstractCoordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large language models (LLMs) for multi-agent planning, a few have begun to explore vision-language models (VLMs) for visual reasoning. However, these VLM-based approaches remain limited in their support for diverse embodiment types. In this work, we introduce VIKI-Bench, the first hierarchical benchmark tailored for embodied multi-agent cooperation, featuring three structured levels: agent activation, task planning, and trajectory perception. VIKI-Bench includes diverse robot embodiments, multi-view visual observations, and structured supervision signals to evaluate reasoning grounded in visual inputs. To demonstrate the utility of VIKI-Bench, we propose VIKI-R, a two-stage framework that fine-tunes a pretrained vision-language model (VLM) using Chain-of-Thought annotated demonstrations, followed by reinforcement learning under multi-level reward signals. Our extensive experiments show that VIKI-R significantly outperforms baselines method across all task levels. Furthermore, we show that reinforcement learning enables the emergence of compositional cooperation patterns among heterogeneous agents. Together, VIKI-Bench and VIKI-R offer a unified testbed and method for advancing multi-agent, visual-driven cooperation in embodied AI systems. Xiufeng Song, Yiran Qin, Jie Yang 0009, Xiaohong Liu 0001, Philip Torr 0001, Lei Bai 0001, Zhenfei Yin |
NeurIPS | 8 |
| 2025 | LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied AgentsabstractScientific embodied agents play a crucial role in modern laboratories by automating complex experimental workflows.Compared to typical household environments, laboratory settings impose significantly higher demands on perception of physical-chemical transformations and long-horizon planning, making them an ideal testbed for advancing embodied intelligence.However, its development has been long hampered by the lack of suitable simulator and benchmarks.In this paper, we address this gap by introducing LabUtopia, a comprehensive simulation and benchmarking suite designed to facilitate the development of generalizable, reasoning-capable embodied agents in laboratory settings. Specifically, it integrates i) LabSim, a high-fidelity simulator supporting multi-physics and chemically meaningful interactions; ii) LabScene, a scalable procedural generator for diverse scientific scenes; and iii) LabBench, a hierarchical benchmark spanning five levels of complexity from atomic actions to long-horizon mobile manipulation. LabUtopia supports 30 distinct tasks and includes more than 200 scene and instrument assets, enabling large-scale training and principled evaluation in high-complexity environments.We demonstrate that LabUtopia offers a powerful platform for advancing the integration of perception, planning, and control in scientific-purpose agents and provides a rigorous testbed for exploring the practical capabilities and generalization limits of embodied intelligence in future research. Project web page: https://rui-li023.github.io/labutopia-site/ Rui Li 0054, Wenxi Qu, Jinouwen Zhang, Zhenfei Yin, Sha Zhang 0002, Xuantuo Huang, Jiangmiao Pang, Wanli Ouyang, Lei Bai 0001, Wangmeng Zuo, Ling-Yu Duan, Dongzhan Zhou, Shixiang Tang |
NeurIPS | 12 |
| 2025 | Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable RefinementabstractUltra-High Definition (UHD) image restoration struggles to balance computational efficiency and detail retention.
While Variational Autoencoders (VAEs) offer improved efficiency by operating in the latent space, with the Gaussian variational constraint, this compression preserves semantics but sacrifices critical high-frequency attributes specific to degradation and thus compromises reconstruction fidelity.
% This compromises reconstruction fidelity, even when global semantics are preserved.
Consequently, a VAE redesign is imperative to foster a robust semantic representation conducive to generalization and perceptual quality, while simultaneously enabling effective high-frequency information processing crucial for reconstruction fidelity.
To address this, we propose \textit{Latent Harmony}, a two-stage framework that reinvigorates VAEs for UHD restoration by concurrently regularizing the latent space and enforcing high-frequency-aware reconstruction constraints.
Specifically, Stage One introduces the LH-VAE, which fortifies its latent representation through visual semantic constraints and progressive degradation perturbation for enhanced semantics robustness; meanwhile, it incorporates latent equivariance to bolster its high-frequency reconstruction capabilities.
Then, Stage Two facilitates joint training of this refined VAE with a dedicated restoration model.
This stage integrates High-Frequency Low-Rank Adaptation (HF-LoRA), featuring two distinct modules: an encoder LoRA, guided by a fidelity-oriented high-frequency alignment loss, tailored for the precise extraction of authentic details from degradation-sensitive high-frequency components; and a decoder LoRA, driven by a perception-oriented loss, designed to synthesize perceptually superior textures. These LoRA modules are meticulously trained via alternating optimization with selective gradient propagation to preserve the integrity of the pre-trained latent structure. This methodology culminates in a flexible fidelity-perception trade-off at inference, managed by an adjustable parameter
$\alpha$.
Extensive experiments demonstrate that \textit{Latent Harmony} effectively balances perceptual and reconstructive objectives with efficiency, achieving superior restoration performance across diverse UHD and standard-resolution scenarios. Yidi Liu, Xueyang Fu, Jie Huang 0017, Jie Xiao 0002, Dong Li 0055, Lei Bai 0001, Zhengjun Zha |
NeurIPS | 7 |
| 2025 | Retro-R1: LLM-based Agentic RetrosynthesisabstractRetrosynthetic planning is a fundamental task in chemical discovery. Due to the vast combinatorial search space, identifying viable synthetic routes remains a significant challenge--even for expert chemists. Recent advances in Large Language Models (LLMs), particularly equipped with reinforcement learning, have demonstrated strong human-like reasoning and planning abilities, especially in mathematics and code problem solving. This raises a natural question: Can the reasoning capabilities of LLMs be harnessed to develop an AI chemist capable of learning effective policies for multi-step retrosynthesis? In this study, we introduce Retro-R1, a novel LLM-based retrosynthesis agent trained via reinforcement learning to design molecular synthesis pathways. Unlike prior approaches, which typically rely on single-turn, question-answering formats, Retro-R1 interacts dynamically with plug-in single-step retrosynthesis tools and learns from environmental feedback. Experimental results show that Retro-R1 achieves a 55.79\% pass@1 success rate, surpassing the previous state of the art by 8.95\%. Notably, Retro-R1 demonstrates strong generalization to out-of-domain test cases, where existing methods tend to fail despite their high in-domain performance. Our work marks a significant step toward equipping LLMs with advanced, chemist-like reasoning abilities, highlighting the promise of reinforcement learning for enabling data-efficient, generalizable, and sophisticated scientific problem-solving in LLM-based agents. Jiangtao Feng, Hongli Yu, Yuxuan Song 0002, Shufei Zhang, Lei Bai 0001, Wei-Ying Ma, Hao Zhou 0012 |
NeurIPS | 7 |
| 2025 | Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple PreferencesabstractData assimilation (DA) aims to estimate the full state of a dynamical system by combining partial and noisy observations with a prior model forecast, commonly referred to as the background. In atmospheric applications, this problem is fundamentally ill-posed due to the sparsity of observations relative to the high-dimensional state space. Traditional methods address this challenge by simplifying background priors to regularize the solution, which are empirical and require continual tuning for application. Inspired by alignment techniques in text-to-image diffusion models, we propose Align-DA, which formulates DA as a generative process and uses reward signals to guide background priors—replacing manual tuning with data-driven alignment. Specifically, we train a score-based model in the latent space to approximate the background-conditioned prior, and align it using three complementary reward signals for DA: (1) assimilation accuracy, (2) forecast skill initialized from the assimilated state, and (3) physical adherence of the analysis fields. Experiments with multiple reward signals demonstrate consistent improvements in analysis quality across different evaluation metrics and observation-guidance strategies. These results show that preference alignment, implemented as a soft constraint, can automatically adapt complex background priors tailored to DA, offering a promising new direction for advancing the field. Jing-An Sun, Hang Fan, Junchao Gong, Ben Fei, Kun Chen 0004, Fenghua Ling, Wanghan Xu, Pierre Gentine, Lei Bai 0001 |
NeurIPS | 11 |
| 2025 | scMRDR: A scalable and flexible framework for unpaired single-cell multi-omics data integrationabstractAdvances in single-cell sequencing have enabled high-resolution profiling of diverse molecular modalities, while integrating unpaired multi-omics single-cell data remains challenging. Existing approaches either rely on pair information or prior correspondences, or require computing a global pairwise coupling matrix, limiting their scalability and flexibility. In this paper, we introduce a scalable and flexible generative framework called single-cell Multi-omics Regularized Disentangled Representations (scMRDR) for unpaired multi-omics integration. Specifically, we disentangle each cell’s latent representations into modality-shared and modality-specific components using a well-designed $\beta$-VAE architecture, which are augmented with isometric regularization to preserve intra-omics biological heterogeneity, adversarial objective to encourage cross-modal alignment, and masked reconstruction loss strategy to address the issue of missing features across modalities. Our method achieves excellent performance on benchmark datasets in terms of batch correction, modality alignment, and biological signal preservation. Crucially, it scales effectively to large-scale datasets and supports integration of more than two omics, offering a powerful and flexible solution for large-scale multi-omics data integration and downstream biological discovery. Jianle Sun, Chaoqi Liang, Peng Zheng 0004, Lei Bai 0001, Wanli Ouyang, Hongliang Yan, Peng Ye 0006 |
NeurIPS | 5 |
| 2025 | Native-Resolution Image SynthesisabstractWe introduce native-resolution image synthesis, a novel paradigm in generative modeling capable of synthesizing images at arbitrary resolutions and aspect ratios. This approach overcomes the limitations of standard fixed-resolution, square-image methods by inherently handling variable-length visual tokens—a core challenge for conventional techniques. To this end, we propose the Native-resolution diffusion Transformer (NiT), an architecture that explicitly models varying resolutions and aspect ratios within its denoising process. Unconstrained by fixed formats, NiT learns intrinsic visual distributions from images encompassing a wide range of resolutions and aspect ratios. Notably, a single NiT model simultaneously achieves the state-of-the-art performance on both ImageNet-256x256 and 512x512 benchmarks. Surprisingly, akin to the robust zero-shot capabilities seen in advanced Large Language Models, NiT, pretrained solely on ImageNet, demonstrates excellent zero-shot generalization performance. It successfully generates high-fidelity images at previously unseen high resolutions (e.g., 1024x1024, 1536x1536) and diverse aspect ratios (e.g., 16:9,3:1, 4:3), as shown in Figure 1. These findings indicate the significant potential of native-resolution modeling as a bridge between visual generative modeling and advanced LLM methodologies. Zidong Wang 0004, Lei Bai 0001, Xiangyu Yue 0001, Wanli Ouyang |
NeurIPS | 2 |
| 2025 | GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion PoliciesabstractDespite significant advances in robotic policy generation, effective coordination in embodied multi-agent systems remains a fundamental challenge—particularly in scenarios where agents must balance individual perspectives with global environmental awareness.
Existing approaches often struggle to balance fine-grained local control with comprehensive scene understanding, resulting in limited scalability and compromised collaboration quality.
In this paper, we present GauDP, a novel Gaussian-image synergistic representation that facilitates scalable, perception-aware imitation learning in multi-agent collaborative systems.
Specifically, GauDP reconstructs a globally consistent 3D Gaussian field from local-view RGB images, allowing all agents to dynamically query task-relevant features from a shared scene representation.
This design facilitates both fine-grained control and globally coherent behavior without requiring additional sensing modalities.
We evaluate GauDP on the RoboFactory benchmark, which includes diverse multi-arm manipulation tasks.
Our method achieves superior performance over existing image-based methods and approaches the effectiveness of point-cloud-driven methods, while maintaining strong scalability as the number of agents increases.
Extensive ablations and visualizations further demonstrate the robustness and efficiency of our unified local-global perception framework for multi-agent embodied learning. Yiran Qin, Jiahua Ma, Zhanglin Peng, Lei Bai 0001, Ruimao Zhang |
NeurIPS | 6 |
| 2025 | Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionabstractWith the rise of the fine-tuned–pretrained paradigm, storing numerous fine-tuned models for multi-tasking creates significant storage overhead.
Delta compression alleviates this by storing only the pretrained model and the highly compressed delta weights (the differences between fine-tuned and pretrained model weights).
However, existing methods fail to maintain both high compression and performance, and often rely on data.
To address these challenges, we propose UltraDelta, the first data-free delta compression pipeline that achieves both ultra-high compression and strong performance.
UltraDelta is designed to minimize redundancy, maximize information, and stabilize performance across inter-layer, intra-layer, and global dimensions, using three key components:
(1) Variance-Based Mixed Sparsity Allocation assigns sparsity based on variance, giving lower sparsity to high-variance layers to preserve inter-layer information.
(2) Distribution-Aware Compression applies uniform quantization and then groups parameters by value, followed by group-wise pruning, to better preserve intra-layer distribution.
(3) Trace-Norm-Guided Rescaling uses the trace norm of delta weights to estimate a global rescaling factor, improving model stability under higher compression.
Extensive experiments across
(a) large language models (fine-tuned on LLaMA-2 7B and 13B) with up to 50$\times$ compression,
(b) general NLP models (RoBERTa-base, T5-base) with up to 224$\times$ compression,
(c) vision models (ViT-B/32, ViT-L/14) with up to 132$\times$ compression, and
(d) multi-modal models (BEiT-3) with 18$\times$ compression,
demonstrate that UltraDelta consistently outperforms existing methods, especially under ultra-high compression.
Code is available at https://github.com/xiaohuiwang000/UltraDelta. Peng Ye 0006, Chenyu Huang 0001, Shenghe Zheng, Bo Zhang 0069, Lei Bai 0001, Wanli Ouyang, Tao Chen 0003 |
NeurIPS | 6 |
| 2025 | BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning DatasetabstractIn this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community. Zhiheng Xi, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 11 |
| 2025 | Learning Urban Climate Dynamics via Physics-Guided Urban Surface-Atmosphere InteractionsabstractUrban warming differs markedly from regional background trends, highlighting the unique behavior of urban climates and the challenges they present. Accurately predicting local urban climate necessitates modeling the interactions between urban surfaces and atmospheric forcing. Although off-the-shelf machine learning (ML) algorithms offer considerable accuracy for climate prediction, they often function as black boxes, learning data mappings rather than capturing physical evolution. As a result, they struggle to capture key land-atmosphere interactions and may produce physically inconsistent predictions. To address these limitations, we propose UCformer, a novel multi-task, physics-guided Transformer architecture designed to emulate nonlinear urban climate processes. UCformer jointly estimates 2-m air temperature $\(T\)$, specific humidity $\(q\)$, and dew point temperature $\(t\)$ in urban areas, while embedding domain and physical priors into its learning structure. Experimental results demonstrate that incorporating domain and physical knowledge leads to significant improvements in emulation accuracy and generalizability under future urban climate scenarios. Further analysis reveals that learning shared correlations across cities enables the model to capture transferable urban surface–atmosphere interaction patterns, resulting in improved accuracy in urban climate emulation. Finally, UCformer shows strong potential to fit real-world data: when fine-tuned with limited observational data, it achieves competitive performance in estimating urban heat fluxes compared to a physics-based model. Jiyang Xia, Fenghua Ling, Zhenhui Jessie Li, David Topping, Lei Bai 0001, Zhonghua Zheng |
NeurIPS | 7 |
| 2025 | LoRA-EnVar: Parameter-Efficient Hybrid Ensemble Variational Assimilation for Weather ForecastingabstractAccurate estimation of background error (i.e., forecast error) distribution is critical for effective data assimilation (DA) in numerical weather prediction (NWP). In state-of-the-art operational DA systems, it is common to account for the temporal evolution of background errors by employing hybrid methods, which blend a static climatological covariance with a flow-dependent ensemble-derived component. While effective to some extent, these methods typically assume Gaussian-distributed errors and rely heavily on hand-crafted covariance structures and domain expertise, limiting their ability to capture the complex, non-Gaussian nature of atmospheric dynamics. In this work, we propose LoRA-EnVar, a novel hybrid ensemble variational DA algorithm that integrates low-rank adaptation (LoRA) into a deep generative modeling framework. We first learn a climatological background error distribution using a variational autoencoder (VAE) trained on historical data. To incorporate flow-dependent uncertainty, we introduce LoRA modules that efficiently adapt the learned distribution in response to flow-dependent ensemble perturbations. Our approach supports online finetuning, enabling dynamic updates of the background error distribution without catastrophic forgetting. We validate LoRA-EnVar in high-resolution assimilation settings using the FengWu forecast model and simulated observations from ERA5 reanalysis. Experimental results show that LoRA-EnVar significantly improves assimilation accuracy over models assuming static background error distribution and achieves comparable or better performance than full finetuning while reducing the number of trainable parameters by three orders of magnitude. This demonstrates the potential of parameter-efficient adaptation for scalable, non-Gaussian DA in operational meteorology. Hang Fan, Kun Chen 0004, Ben Fei, Wei Xue 0003, Lei Bai 0001 |
NeurIPS | 7 |
| 2025 | SIFusion: A Unified Fusion Framework for Multi-granularity Arctic Sea Ice ForecastingabstractArctic sea ice performs a vital role in global climate and has paramount impacts on both polar ecosystems and coastal communities. In the last few years, multiple deep learning based pan-Arctic sea ice concentration (SIC) forecasting methods have emerged and showcased superior performance over physics-based dynamical models. However, previous methods forecast SIC at a fixed temporal granularity, e.g. sub-seasonal or seasonal, thus only leveraging inter-granularity information and overlooking the plentiful inter-granularity correlations. SIC at various temporal granularities exhibits cumulative effects and are naturally consistent, with short-term fluctuations potentially impacting long-term trends and long-term trends provides effective hints for facilitating short-term forecasts in Arctic sea ice. Therefore, in this study, we propose to cultivate temporal multi-granularity that naturally derived from Arctic sea ice reanalysis data and provide a unified perspective for modeling SIC via our Sea Ice Fusion framework. SIFusion is delicately designed to leverage both intra-granularity and inter-granularity information for capturing granularity-consistent representations that promote forecasting skills. Our extensive experiments show that SIFusion outperforms off-the-shelf deep learning models for their specific temporal granularity. Weidong Yang 0001, Keyi Liu, Yeqi Luo, Ben Fei, Lei Bai 0001 |
NeurIPS | 7 |
| 2025 | Understand Before You Generate: Self-Guided Training for Autoregressive Image GenerationabstractRecent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenges. In this work, we present the first systematic investigation into the mechanisms of applying the next-token prediction paradigm to the visual domain. We identify three key properties that hinder the learning of high-level visual semantics: local and conditional dependence, inter-step semantic inconsistency, and spatial invariance deficiency. We show that these issues can be effectively addressed by introducing self-supervised objectives during training, leading to a novel training framework, Self-guided Training for AutoRegressive models (ST-AR). Without relying on pre-trained representation models, ST-AR significantly enhances the image understanding ability of autoregressive models and leads to improved generation quality. Specifically, ST-AR brings approximately 42% FID improvement for LlamaGen-L and 49% FID improvement for LlamaGen-XL, while maintaining the same sampling strategy. Xiaoyu Yue, Zidong Wang 0004, Xihui Liu, Wanli Ouyang, Lei Bai 0001, Luping Zhou |
NeurIPS | 7 |
| 2025 | Actial: Activate Spatial Reasoning Ability of Multimodal Large Language ModelsabstractRecent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding. Xiaoyu Zhan, Wenxuan Huang 0001, Xinyu Fu 0009, Changfeng Ma, Shaosheng Cao, Bohan Jia, Shaohui Lin, Zhenfei Yin, Lei Bai 0001, Wanli Ouyang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001 |
NeurIPS | 10 |
| 2025 | Scaling Physical Reasoning with the PHYSICS DatasetabstractLarge Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding, received limited academic and industrial attention. This paper introduces PHYSICS, a dataset containing 16,568 high-quality physics problems spanning subjects and difficulty levels, to facilitate this issue. Specifically, PHYSICS is curated with exercises from over 100 textbooks through a carefully designed pipeline for quality control. It covers five major physics domains: Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. It also spans a wide range of difficulty levels, from high school to graduate-level physics courses. To utilize the data for improving and evaluating the model's physical reasoning capabilities, we split the dataset into training and test sets, and provide reasoning paths generated by powerful reasoning models for the training data to facilitate model training. In addition, for the evaluation part, we find that existing evaluation frameworks exhibit biases in aspects such as units, simplification, and precision in physics domain. To balance efficiency and accuracy, we introduce a Rule+Model evaluation framework tailored to physics problems. Our evaluations on current state-of-the-art open-source and proprietary models highlight the limitations of current models in handling physics-related tasks. We hope that our dataset and evaluation methodology will jointly advance the development of LLMs in the field of physics. The code and data can be found at: https://github.com/Zhengsh123/PHYSICS. Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Ning Ding 0002, Yu Cheng 0001, Shuyue Hu, Lei Bai 0001, Dongzhan Zhou, Ganqu Cui, Peng Ye 0006 |
NeurIPS | 9 |
| 2025 | NUTS: Eddy-Robust Reconstruction of Surface Ocean Nutrients via Two-Scale ModelingabstractReconstructing ocean surface nutrients from sparse observations is critical for understanding long-term biogeochemical cycles. Most prior work focuses on reconstructing atmospheric fields and treats the reconstruction problem as image inpainting, assuming smooth, single-scale dynamics. In contrast, nutrient transport follows advection–diffusion dynamics under nonstationary, multiscale ocean flow. This mismatch leads to instability, as small errors in unresolved eddies can propagate through time and distort nutrient predictions.
To address this, we introduce NUTS, a two-scale reconstruction model that decouples large-scale transport and mesoscale variability. The homogenized solver captures stable, coarse-scale advection under filtered flow. A refinement module then restores mesoscale detail conditioned on the residual eddy field.
NUTS is stable, interpretable, and robust to mesoscale perturbations, with theoretical guarantees from homogenization theory. NUTS outperforms all data-driven baselines in global reconstruction and achieves site-wise accuracy comparable to numerical models. On real observations, NUTS reduces NRMSE by 79.9% for phosphate and 19.3% for nitrate over the best baseline. Ablation studies validate the effectiveness of each module. Shiyu Liang, Chaofan Sun, Lei Bai 0001, Enhui Liao |
NeurIPS | 5 |
| 2025 | Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and ReasoningabstractScientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this discovery process in realistic workflows. However, current scientific benchmarks mostly focus on evaluating the knowledge understanding capabilities of MLLMs, leading to an inadequate assessment of their perception and reasoning abilities. To address this gap, we present the Scientists’ First Exam (SFE) benchmark, designed to evaluate the scientific cognitive capacities of MLLMs through three interconnected levels: scientific signal perception, scientific attribute understanding, scientific comparative reasoning. Specifically, SFE comprises 830 expert-verified VQA pairs across three question types, spanning 66 multimodal tasks across five high-value disciplines. Extensive experiments reveal that current state-of-the-art GPT-o3 and InternVL-3 achieve only 34.08% and 26.52% on SFE, highlighting significant room for MLLMs to improve in scientific realms. We hope the insights obtained in SFE will facilitate further developments in AI-enhanced scientific discoveries. Yuhao Zhou 0005, Ruoyao Xiao, Qiantai Feng, Zijie Guo, Yuejin Yang, Wenxuan Huang 0001, Dan Si, Xiuqi Yao, Jia Bu, Haiwen Huang, Tianfan Fu, Shixiang Tang, Ben Fei, Dongzhan Zhou, Fenghua Ling, Yan Lu 0001, Chenhui Li 0001, Guanjie Zheng, Lei Bai 0001 |
NeurIPS | 27 |
| 2025 | Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-IdentificationabstractRecently, person re-identification (ReID) has witnessed fast development due to its broad practical applications and proposed various settings, e.g., traditional ReID, clothes-changing ReID, and visible-infrared ReID. However, current studies primarily focus on single specific tasks, which limits model applicability in real-world scenarios. This paper aims to address this issue by introducing a novel instruct-ReID task that unifies 6 existing ReID tasks in one model and retrieves images based on provided visual or textual instructions. Instruct-ReID is the first exploration of a general ReID setting, where 6 existing ReID tasks can be viewed as special cases by assigning different instructions. To facilitate research in this new instruct-ReID task, we propose a large-scale OmniReID++ benchmark equipped with diverse data and comprehensive evaluation methods, e.g., task-specific and task-free evaluation settings. In the task-specific evaluation setting, gallery sets are categorized according to specific ReID tasks. We propose a novel baseline model, IRM, with an adaptive triplet loss to handle various retrieval tasks within a unified framework. For task-free evaluation setting, where target person images are retrieved from task-agnostic gallery sets, we further propose a new method called IRM++ with novel memory bank-assisted learning. Extensive evaluations of IRM and IRM++ on OmniReID++ benchmark demonstrate the superiority of our proposed methods, achieving state-of-the-art performance on 10 test sets. Weizhen He, Yiheng Deng, Yunfeng Yan, Feng Zhu 0006, Yizhou Wang 0007, Lei Bai 0001, Qingsong Xie, Rui Zhao 0001, Donglian Qi, Wanli Ouyang, Shixiang Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Hulk: A Universal Knowledge Translator for Human-Centric TasksabstractHuman-centric perception tasks, e.g., pedestrian detection, skeleton-based action recognition, and pose estimation, have wide industrial applications, such as metaverse and sports analysis. There is a recent surge to develop human-centric foundation models that can benefit a broad range of human-centric perception tasks. While many human-centric foundation models have achieved success, they did not explore 3D and vision-language tasks for human-centric and required task-specific finetuning. These limitations restrict their application to more downstream tasks and situations. To tackle these problems, we present Hulk, the first multimodal human-centric generalist model, capable of addressing 2D vision, 3D vision, skeleton-based, and vision-language tasks without task-specific finetuning. The key to achieving this is condensing various task-specific heads into two general heads, one for discrete representations, e.g., languages, and the other for continuous representations, e.g., location coordinates. The outputs of two heads can be further stacked into four distinct input and output modalities. This uniform representation enables Hulk to treat diverse human-centric tasks as modality translation, integrating knowledge across a wide range of tasks. Comprehensive evaluations of Hulk on 12 benchmarks covering 8 human-centric tasks demonstrate the superiority of our proposed method, achieving state-of-the-art performance in 11 benchmarks. Yizhou Wang 0007, Weizhen He, Xun Guo 0001, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Jian Wu 0001, Tong He 0001, Wanli Ouyang, Shixiang Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Stimulative Training++: Go Beyond the Performance Limits of Residual NetworksabstractResidual networks have shown great success and become indispensable in recent deep neural network models. In this work, we aim to re-investigate the training process of residual networks from a novel perspective of loafing, and further propose a new training scheme as well as three improved strategies for boosting residual networks beyond their performance limits. Previous research has suggested that residual networks can be considered as ensembles of shallow networks, which implies that the final performance of a residual network is influenced by a group of subnetworks. Furthermore, we identify a previously overlooked problem, where subnetworks within a residual network are prone to exert less effort when working as part of a group compared to working alone. We define this problem as network loafing. Since network loafing may inevitably cause the sub-par performance of the residual network, we propose a novel training scheme called stimulative training, which randomly samples a residual subnetwork and calculates the KL divergence loss between the sampled subnetwork and the given residual network for extra supervision. In order to unleash the potential of stimulative training, we further propose three simple-yet-effective strategies, including a novel KL- loss that only aligns the network logits direction, random smaller inputs for subnetworks, and inter-stage sampling rules. Comprehensive experiments and analysis verify the effectiveness of stimulative training as well as its three improved strategies. For example, the proposed method can boost the performance of ResNet50 on ImageNet to 80.5% Top1 accuracy without using any extra data, model, trick, or changing the structure. With only uniform augment, the performance can be further improved to 81.0% Top1 accuracy, better than the best training recipes provided by Timm library and PyTorch official version. We also verify its superiority on various typical models, datasets, and tasks and give some theoretical analysis. As such, we advocate utilizing the proposed method as a general and next-generation technology to train residual networks. Peng Ye 0006, Tong He 0001, Shengji Tang, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Impact of VAEformer Compression Algorithm Precision Loss on the Tropospheric Delays for Microwave Remote SensingabstractRay-tracing through numerical weather models (NWMs) is one of the most accurate methods for determining slant tropospheric delays (STDs) in microwave remote sensing. However, the massive data volumes of high-resolution NWMs create substantial I/O operations, limiting large-scale ray-tracing on general hardware. This constraint has historically necessitated parameterized tropospheric delay models, which are disseminated as standardized products (e.g., zenith delays with mapping functions and horizontal gradients). Recently, the AI-driven VAE-former algorithm revolutionized NWM compression, achieving >470:1 ratios by compressing 37 pressure level, 0.25°×0.25° ERA5 data into files smaller than surface-only VMF3 products (1°×1° resolution). This breakthrough challenges the conventional reliance on parameterized models as the sole practical solution. We quantified discrepancies in tropospheric delay parameters between original ERA5 and VAEformer-compressed CRA5 data across 2022, evaluating compression fidelity on global grids and against in-situ zenith tropospheric delay (ZTD) estimates. Results show global average precision loss from compression is10 mm). Our findings demonstrate CRA5 as a reliable ERA5 substitute, with compression-induced inaccuracies being negligible for most microwave-based remote sensing applications. This work underscores that parameterized delay modeling is no longer the exclusive pathway, enabling efficient local computation of high-precision STDs without through mapping functions and gradients. Junsheng Ding, Cancan Xu, Wu Chen 0001, Junping Chen, Yize Zhang, Lei Bai 0001, Tao Han 0002, Yuhao Xiong |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Topographic Informed Kolmogorov-Arnold Neural Interpolator for Downscaling and Correcting Meteorological Fields From In Situ ObservationsabstractObtaining accurate weather forecasts at station locations is a critical challenge due to systematic biases arising from the mismatch between multi-scale, continuous atmospheric characteristic and their discrete, gridded representations. Previous works have primarily focused on modeling gridded meteorological data, inherently neglecting the off-grid, continuous nature of atmospheric states and leaving such biases unresolved. To address this, we propose theKolmogorov–Arnold Neural Interpolator(KANI), a novel framework that redefines meteorological field representation as continuous neural functions derived from discretized grids. Grounded in the Kolmogorov–Arnold theorem, KANI captures the inherent continuity of atmospheric states and leverages sparse in-situ observations to correct these biases systematically. Furthermore, KANI introduces an innovativezero-shotdownscaling capability, guided by high-resolution topographic textures without requiring high-resolution meteorological fields for supervision. Experimental results across three sub-regions of the continental United States indicate that KANI achieves an accuracy improvement of 40.28% for temperature and 67.41% for wind speed, highlighting its significant improvement over traditional interpolation methods. This enables continuous neural representation of meteorological variables through neural networks, transcending the limitations of conventional grid-based representations. Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | KAN-Enhanced Transformer for Wind Profile Retrieval From Lidar SpectraabstractAccurate detection of wind profiles within the troposphere is essential for atmospheric dynamics research and extreme weather forecasting. Coherent Doppler wind lidar (CDWL) is widely regarded as the most suitable technique for high spatial and temporal resolution wind profile detection. However, since coherent detection relies heavily on the concentration of aerosol particles, which cause Mie scattering, the received backscattering lidar signal exhibits significantly low intensity at high altitudes. As a result, conventional methods, such as spectral centroid estimator, often fail to produce credible and accurate wind retrieval results in these regions. To address this issue, we introduce a deep learning model for lidar wind profile retrieval, which integrates Transformer architecture with the Kolmogorov-Arnold network. Our model is trained solely on targets derived from the traditional wind retrieval algorithm and utilizes radiosonde measurements as the ground truth for test results evaluation. Experimental results demonstrate that our model not only extends the maximum vertical wind profile detection range but also produces more accurate results, exhibiting a level of precision that surpasses the labeled targets. This phenomenon, which we refer to assuper-accuracy, is explored by investigating the potential underlying factors that contribute to this intriguing occurrence. In addition, we compare the performance of our method with other state-of-the-art (SOTA) models, highlighting its superior effectiveness and capability in high-resolution wind retrieval. In the future, our proposed model could potentially be extended to other remote sensing spectrum signal processing applications (e.g., microwave Doppler lidar), thereby establishing a benchmark for future research and driving the advancement of deep learning models in this domain. Chong Wang 0022, Hao Chen 0045, Mingjiao Jia, Xiang Shang, Luoyuan Qu, Guoliang Shentu, Yanyu Lu, Yanfeng Huo, Lei Bai 0001, Xianghui Xue, Xiankang Dou |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | VegeDiff: Latent Diffusion Model for Geospatial Vegetation ForecastingabstractIn the context of global climate change and frequent extreme weather events, forecasting future geospatial vegetation states under these conditions is of significant importance. The vegetation change process is influenced by the complex interplay between dynamic meteorological variables and static environmental variables, leading to high levels of uncertainty. Existing deterministic methods are inadequate in addressing this uncertainty and fail to accurately model the impact of these variables on vegetation, resulting in blurry and inaccurate forecasting results. To address these issues, VegeDiff is proposed for the geospatial vegetation forecasting task. To our best knowledge, VegeDiff is the first to employ a diffusion model to probabilistically capture the uncertainties in vegetation change processes, enabling the generation of clear and accurate future vegetation states. VegeDiff also separately models the global impact of dynamic meteorological variables and the local effects of static environmental variables, thus accurately modeling the impact of these variables. Extensive experiments on geospatial vegetation forecasting tasks demonstrate the effectiveness of VegeDiff. By capturing the uncertainties in vegetation changes and modeling the complex influence of relevant variables, VegeDiff outperforms existing deterministic methods, providing clear and accurate forecasting results of future vegetation states. Interestingly, this study demonstrate the potential of VegeDiff in applications of forecasting future vegetation states from multiple aspects and exploring the impact of meteorological variables on vegetation dynamics. The code of this work will be available at https://github.com/walking-shadow/Official VegeDiff. Sijie Zhao, Hao Chen 0045, Xueliang Zhang 0002, Pengfeng Xiao, Lei Bai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Online Test-Time Adaptation of Spatial-Temporal Traffic Flow ForecastingabstractAccurate spatial-temporal traffic flow forecasting is crucial in aiding traffic managers in implementing control measures and assisting drivers in selecting optimal travel routes. Traditional deep-learning based methods for traffic flow forecasting typically rely on historical data to train their models, which are then used to make predictions on future data. However, the performance of the trained model usually degrades due to the temporal drift between the historical and future data. To make the model trained on historical data better adapt to future data in a fully online manner, this paper conducts the first study of the online test-time adaptation techniques for spatial-temporal traffic flow forecasting problems. To this end, we propose anAdaptiveDoubleCorrection bySeriesDecomposition (ADCSD) method, which first decomposes the output of the trained model into seasonal and trend-cyclical parts and then corrects them by two separate modules during the testing phase using the latest observed dataentry by entry. In the proposed ADCSD method, instead of fine-tuning the whole trained model during the testing phase, a lite network is attached after the trained model, and only the lite network is fine-tuned in the testing process each time a data entry is observed. Moreover, to satisfy that different time series variables may have different levels of temporal drift, two adaptive vectors are adopted to provide different weights for different time series variables. Extensive experiments on four real-world traffic flow forecasting datasets demonstrate the effectiveness of the proposed ADCSD method. The code is available athttps://github.com/Pengxin-Guo/ADCSD Pengxin Guo 0001, Pengrong Jin, Ziyue Li 0002, Lei Bai 0001, Yu Zhang 0006 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | VisionTraj: A Noise-Robust Trajectory Recovery Framework Based on Large-Scale Camera NetworkabstractTrajectory recovery from snapshots captured by a city-wide multi-camera network facilitates urban mobility sensing and road network optimization. State-of-the-art solutions for such vision-based schemes typically rely on predefined rules or unsupervised iterative feedback, but they struggle with multiple challenges, such as the lack of open-source datasets for training the entire pipeline and the vulnerability to noise in visual inputs. In response to the dilemma, this paper proposes VisionTraj, the first learning-based model that reconstructs vehicle trajectories from snapshots recorded by road network cameras. Along with this, we present two well-designed vision-trajectory datasets that provide extensive trajectory data and corresponding visual snapshots, enabling the extraction of supervised vision-trajectory interactions. After the data creation, based on the results from the off-the-shelf multi-modal vehicle clustering, we first re-formulate the trajectory recovery problem as a generative task and introduce the canonical Transformer as the autoregressive backbone. Next, to identify clustering noise (i.e., false positives) based on the snapshots’ spatiotemporal dependencies, a graph convolutional neural network-based soft-denoising module is built upon the fine- and coarse-grained clusters. Additionally, we leverage strong semantic information extracted from the tracklet to provide detailed insights into the vehicle’s entry and exit behaviors during trajectory recovery. The denoising and tracklet components can also serve as plug-and-play modules to enhance baselines. Experimental results on the two hand-crafted datasets show that the proposed VisionTraj achieves a maximum improvement of +11.5% against the sub-best model. Furthermore, we explore potential downstream applications, and our model continues to outperform its peers. The code and data are available herehttps://github.com/bonaldli/VisionTraj Zhishuai Li, Ziyue Li 0002, Xiaoru Hu, Guoqing Du, Yunhao Nie, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Enhancing Mapless Trajectory Prediction Through Knowledge DistillationabstractScene information plays a crucial role in trajectory forecasting systems for autonomous driving by providing semantic clues and constraints on potential future paths of traffic agents. Prevalent trajectory prediction techniques often take High-Definition maps (HD maps) as part of the inputs to provide scene knowledge. Although HD maps offer accurate road information, they may suffer from the high cost of annotation or restrictions of law, which limits their widespread use. Therefore, it is crucial for trajectory prediction methods to generate reliable prediction results in mapless scenarios. In this paper, we tackle the problem of improving the consistency of predicted trajectories and the scene road topology when map information is unavailable during the test phase. To achieve this, we propose a universal knowledge distillation (KD) framework. This KD framework trains a map-based teacher network on samples with annotated HD maps and subsequently transfers the knowledge to a student mapless predictor through a two-fold knowledge distillation process. Experimental results show that our method stably improves prediction performance in test-time mapless situations on many widely used trajectory prediction baselines, and achieves state-of-the-art mapless prediction performances. Qualitative visualization results demonstrate that our approach helps infer unseen map information. Our solution is generalizable for common trajectory prediction networks and datasets. Pu Zhang 0001, Lei Bai 0001, Jianru Xue |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image GenerationabstractA 360-degree (omni-directional) image provides an all-encompassing spherical view of a scene. Recently, there has been an increasing interest in synthesising 360-degree images from conventional narrow field of view (NFoV) images captured by digital cameras and smartphones, for providing immersive experiences in various scenarios such as virtual reality. Yet, existing methods typically fall short in synthesizing intricate visual details or ensure the generated images align consistently with user-provided prompts. In this study, autoregressive omni-aware generative network (AOG-Net) is proposed for 360-degree image generation by outpainting an incomplete 360-degree image progressively with NFoV and text guidances joinly or individually. This autoregressive scheme not only allows for deriving finer-grained and text-consistent patterns by dynamically generating and adjusting the process but also offers users greater flexibility to edit their conditions throughout the generation process. A global-local conditioning mechanism is devised to comprehensively formulate the outpainting guidance in each autoregressive step. Text guidances, omni-visual cues, NFoV inputs and omni-geometry are encoded and further formulated with cross-attention based transformers into a global stream and a local stream into a conditioned generative backbone model. As AOG-Net is compatible to leverage large-scale models for the conditional encoder and the generative prior, it enables the generation to use extensive open-vocabulary text guidances. Comprehensive experiments on two commonly used 360-degree image datasets for both indoor and outdoor settings demonstrate the state-of-the-art performance of our proposed method. Our code is available at https://github.com/zhuqiangLu/AOG-NET-360. Zhuqiang Lu, Kun Hu 0008, Lei Bai 0001, Zhiyong Wang 0001 |
AAAI | 4 |
| 2024 | Towards Dynamic Spatial-Temporal Graph Learning: A Decoupled PerspectiveabstractWith the progress of urban transportation systems, a significant amount of high-quality traffic data is continuously collected through streaming manners, which has propelled the prosperity of the field of spatial-temporal graph prediction. In this paper, rather than solely focusing on designing powerful models for static graphs, we shift our focus to spatial-temporal graph prediction in the dynamic scenario, which involves a continuously expanding and evolving underlying graph. To address inherent challenges, a decoupled learning framework (DLF) is proposed in this paper, which consists of a spatial-temporal graph learning network (DSTG) with a specialized decoupling training strategy. Incorporating inductive biases of time-series structures, DSTG can interpret time dependencies into latent trend and seasonal terms. To enable prompt adaptation to the evolving distribution of the dynamic graph, our decoupling training strategy is devised to iteratively update these two types of patterns. Specifically, for learning seasonal patterns, we conduct thorough training for the model using a long time series (e.g., three months of data). To enhance the learning ability of the model, we also introduce the masked auto-encoding mechanism. During this period, we frequently update trend patterns to expand new information from dynamic graphs. Considering both effectiveness and efficiency, we develop a subnet sampling strategy to select a few representative nodes for fine-tuning the weights of the model. These sampled nodes cover unseen patterns and previously learned patterns. Experiments on dynamic spatial-temporal graph datasets further demonstrate the competitive performance, superior efficiency, and strong scalability of the proposed framework. Binwu Wang, Pengkun Wang 0001, Yudong Zhang 0005, Xu Wang 0029, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
AAAI | 6 |
| 2024 | MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsabstractGenerating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion directly from textual action descriptions, they often support only a single modality of the control signal, which limits their application in the real digital human industry. This paper presents a Motion General-Purpose generaTor (MotionGPT) that can use multimodal control signals, e.g., text and single-frame poses, for generating consecutive human motions by treating multimodal signals as special input tokens in large language models (LLMs). Specifically, we first quantize multimodal control signals into discrete codes and then formulate them in a unified prompt instruction to ask the LLMs to generate the motion answer. Our MotionGPT demonstrates a unified human motion generation model with multimodal control signals by tuning a mere 0.4% of LLM parameters. To the best of our knowledge, MotionGPT is the first method to generate human motion by multimodal control signals, which we hope can shed light on this new direction. Visit our webpage at https://qiqiapink.github.io/MotionGPT/. Bin Liu 0016, Shixiang Tang, Yan Lu 0001, Lu Chen 0001, Lei Bai 0001, Qi Chu 0001, Nenghai Yu, Wanli Ouyang |
AAAI | 7 |
| 2024 | Non-Neighbors Also Matter to Kriging: A New Contrastive-Prototypical LearningabstractKriging aims to estimate the attributes of unseen geo-locations from observations in the spatial vicinity or physical connections. Existing works assume that neighbors’ information offers the basis for estimating the unobserved target while ignoring non-neighbors. However, neighbors could also be quite different or even misleading, and the non-neighbors could still offer constructive information. To this end, we propose "Contrastive-Prototypical" self-supervised learning for Kriging (KCP): (1) The neighboring contrastive module coarsely pushes neighbors together and non-neighbors apart. (2) In parallel, the prototypical module identifies similar representations via exchanged prediction, such that it refines the misleading neighbors and recycles the useful non-neighbors from the neighboring contrast component. As a result, not all the neighbors and some of the non-neighbors will be used to infer the target. (3) To learn general and robust representations, we design an adaptive augmentation module that encourages data diversity. Theoretical bound is derived for the proposed augmentation. Extensive experiments on real-world datasets demonstrate the superior performance of KCP compared to its peers with 6% improvements and exceptional transferability and robustness. Zhishuai Li, Yunhao Nie, Ziyue Li 0002, Lei Bai 0001, Rui Zhao 0001 |
AISTATS | 4 |
| 2024 | Instruct-ReID: A Multi-Purpose Person Re-Identification Task with InstructionsabstractHuman intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which limits the applications in the real world. This paper strives to resolve this problem by proposing a new instruct-ReID task that requires the model to retrieve images according to the given image or language instructions. Our instruct-ReID is a more general ReID setting, where existing 6 ReID tasks can be viewed as special cases by designing different instructions. We propose a large-scale OmniReID benchmark and an adaptive triplet loss as a baseline method to facilitate research in this new setting. Experimental results show that the proposed multi-purpose ReID model, trained on our OmniReID benchmark without finetuning, can improve +0.5%, +0.6%, +7.7% mAP on Market1501, MSMT17, CUHK03 for traditional ReID, +6.4%, +7.1%, +11.2% mAP on PRCC, VC-Clothes, LTCC for clothes-changing ReID, +11.7% mAP on COCAS+ real2 for clothes template based clothes-changing ReID when using only RGB images, +24.9% mAP on COCAS+ real2 for our newly defined language-instructed ReID, +4.3% on LLCM for visible-infrared ReID, +2.6% on CUHK-PEDES for text-to-image ReID. The datasets, the model, and code are available at https://github.com/hwz-zju/Instruct-ReID. Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang 0007, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Wanli Ouyang, Donglian Qi, Yunfeng Yan |
CVPR | 7 |
| 2024 | PredBench: Benchmarking Spatio-Temporal Prediction Across Diverse Disciplines
Zidong Wang 0004, Tong He 0001, Xihui Liu, Wanli Ouyang, Lei Bai 0001 |
ECCV (58) | 7 |
| 2024 | Q-Refine: A Perceptual Quality Refiner for AI-Generated ImageabstractWith the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited optimization capabilities for low-quality AIGIs but also brought negative optimization to high-quality AIGIs. To address this issue, a quality-award refiner named Q-Refine is proposed. Based on the preference of the Human Visual System (HVS), Q-Refine uses the Image Quality Assessment (IQA) metric to guide the refining process for the first time, and modify images of different qualities through three adaptive pipelines. Experimental data shows that for mainstream T2I models, Q-Refine can perform effective optimization to AIGIs of different qualities. It can be a general refiner to optimize AIGIs from both fidelity and aesthetic quality levels, thus expanding the application of the T2I generation models. The code is released on https://github.com/Q-Future/Q-Refine. Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
ICME | 6 |
| 2024 | CasCast: Skillful High-resolution Precipitation Nowcasting via Cascaded ModellingabstractPrecipitation nowcasting based on radar data plays a crucial role in extreme weather prediction and has broad implications for disaster management. Despite progresses have been made based on deep learning, two key challenges of precipitation nowcasting are not well-solved: (i) the modeling of complex precipitation system evolutions with different scales, and (ii) accurate forecasts for extreme precipitation. In this work, we propose CasCast, a cascaded framework composed of a deterministic and a probabilistic part to decouple the predictions for mesoscale precipitation distributions and small-scale patterns. Then, we explore training the cascaded framework at the high resolution and conducting the probabilistic modeling in a low dimensional latent space with a frame-wise-guided diffusion transformer for enhancing the optimization of extreme events while reducing computational costs. Extensive experiments on three benchmark radar precipitation datasets show that CasCast achieves competitive performance. Especially, CasCast significantly surpasses the baseline (up to +91.8%) for regional extreme-precipitation nowcasting. Junchao Gong, Lei Bai 0001, Peng Ye 0006, Wanghan Xu, Xiaokang Yang 0001, Wanli Ouyang |
ICML | 2 |
| 2024 | FiT: Flexible Vision Transformer for Diffusion ModelabstractIn the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically designed for generating images with unrestricted resolutions and aspect ratios. Unlike traditional methods that perceive images as static-resolution grids, FiT conceptualizes images as sequences of dynamically-sized tokens. This perspective enables a flexible training strategy that effortlessly adapts to diverse aspect ratios during both training and inference phases, thus promoting resolution generalization and eliminating biases induced by image cropping. Enhanced by a meticulously adjusted network structure and the integration of training-free extrapolation techniques, FiT exhibits remarkable flexibility in resolution extrapolation generation. Comprehensive experiments demonstrate the exceptional performance of FiT across a broad range of resolutions. Repository available at https://github.com/whlzy/FiT. Zidong Wang 0004, Chengyue Wu, Xihui Liu, Wanli Ouyang, Lei Bai 0001 |
ICML | 7 |
| 2024 | Towards a Self-contained Data-driven Global Weather Forecasting FrameworkabstractData-driven weather forecasting models are advancing rapidly, yet they rely on initial states (i.e., analysis states) typically produced by traditional data assimilation algorithms. Four-dimensional variational assimilation (4DVar) is one of the most widely adopted data assimilation algorithms in numerical weather prediction centers; it is accurate but computationally expensive. In this paper, we aim to couple the AI forecasting model, FengWu, with 4DVar to build a self-contained data-driven global weather forecasting framework, FengWu-4DVar. To achieve this, we propose an *AI-embedded* 4DVar algorithm that includes three components: (1) a 4DVar objective function embedded with the FengWu forecasting model and its error representation to enhance efficiency and accuracy; (2) a spherical-harmonic-transform-based (SHT-based) approximation strategy for capturing the horizontal correlation of background error; and (3) an auto-differentiation (AD) scheme for determining the optimal analysis fields. Experimental results show that under the ERA5 simulated observational data with varying proportions and noise levels, FengWu-4DVar can generate accurate analysis fields; remarkably, it has achieved stable self-contained global weather forecasts for an entire year for the first time, demonstrating its potential for real-world applications. Additionally, our framework is approximately 100 times faster than the traditional 4DVar algorithm under similar experimental conditions, highlighting its significant computational efficiency. Lei Bai 0001, Wei Xue 0003, Hao Chen 0045, Kun Chen 0004, Tao Han 0002, Wanli Ouyang |
ICML | 2 |
| 2024 | PAPS-OVQA: Projection-Aware Patch Sampling for Omnidirectional Video Quality AssessmentabstractIn immersive multimedia systems, the perceptual quality model of omnidirectional video is indispensable. However, to cope with its resolution that is several times higher than ordinary video, the existing omnidirectional video quality assessment (OVQA) models require extremely high computational complexity and usually need to transcode the projection into a certain format. Therefore, to assess the perceptual quality of omnidirectional video effectively, we propose Projection-Aware Patch Sampling (PAPS)-OVQA to process its three common projection formats simultaneously while resizing high-resolution video into patches sampled from uniform grids and finally apply Fragment Attention Network (FANet) to perform quality regression. As a result, we avoid the overhead computational cost of projection transcoding and reduce the complexity of the quality model greatly. Experimental data show that PAPS-OVQA guarantees good performance while retaining high efficiency under different projection formats. Chunyi Li 0001, Haoning Wu 0001, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ISCAS | 5 |
| 2024 | G-Refine: A General Quality Refiner for Text-to-Image Generation
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Tengchuan Kou, Chaofeng Chen, Lei Bai 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 7 |
| 2024 | FNP: Fourier Neural Processes for Arbitrary-Resolution Data AssimilationabstractData assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their significant advantages over traditional techniques in terms of computational consumption. However, existing AI-based data assimilation methods can only handle observations with a specific resolution, lacking the compatibility and generalization ability to assimilate observations with other resolutions. Considering that complex real-world observations often have different resolutions, we propose the Fourier Neural Processes (FNP) for arbitrary-resolution data assimilation in this paper. Leveraging the efficiency of the designed modules and flexible structure of neural processes, FNP achieves state-of-the-art results in assimilating observations with varying resolutions, and also exhibits increasing advantages over the counterparts as the resolution and the amount of observations increase. Moreover, our FNP trained on a fixed resolution can directly handle the assimilation of observations with out-of-distribution resolutions and the observational information reconstruction task without additional fine-tuning, demonstrating its excellent generalization ability across data resolutions as well as across tasks. Code is available at https://github.com/OpenEarthLab/FNP. Kun Chen 0004, Peng Ye 0006, Hao Chen 0045, Tao Han 0002, Wanli Ouyang, Tao Chen 0003, Lei Bai 0001 |
NeurIPS | 8 |
| 2024 | On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionabstractLarge numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identify diffusion-generated content with a diverse range of semantics. To advance the field of video forensics, we propose an innovative algorithm named Multi-Modal Detection(MM-Det) for detecting diffusion-generated videos. MM-Det utilizes the profound perceptual and comprehensive abilities of Large Multi-modal Models (LMMs) by generating a Multi-Modal Forgery Representation (MMFR) from LMM's multi-modal space, enhancing its ability to detect unseen forgery content. Besides, MM-Det leverages an In-and-Across Frame Attention (IAFA) mechanism for feature augmentation in the spatio-temporal domain. A dynamic fusion strategy helps refine forgery representations for the fusion. Moreover, we construct a comprehensive diffusion video dataset, called Diffusion Video Forensics (DVF), across a wide range of forgery videos. MM-Det achieves state-of-the-art performance in DVF, demonstrating the effectiveness of our algorithm. Both source code and DVF are available at https://github.com/SparkleXFantasy/MM-Det. Xiufeng Song, Jiache Zhang, Lei Bai 0001, Xiaoming Liu 0002, Guangtao Zhai, Xiaohong Liu 0001 |
NeurIPS | 5 |
| 2024 | Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid ModelingabstractData-driven artificial intelligence (AI) models have made significant advancements in weather forecasting, particularly in medium-range and nowcasting. However, most data-driven weather forecasting models are black-box systems that focus on learning data mapping rather than fine-grained physical evolution in the time dimension. Consequently, the limitations in the temporal scale of datasets prevent these models from forecasting at finer time scales. This paper proposes a physics-AI hybrid model (i.e., WeatherGFT) which generalizes weather forecasts to finer-grained temporal scales beyond training dataset. Specifically, we employ a carefully designed PDE kernel to simulate physical evolution on a small time scale (e.g., 300 seconds) and use a parallel neural networks with a learnable router for bias correction. Furthermore, we introduce a lead time-aware training framework to promote the generalization of the model at different lead times. The weight analysis of physics-AI modules indicates that physics conducts major evolution while AI performs corrections adaptively. Extensive experiments show that WeatherGFT trained on an hourly dataset, effectively generalizes forecasts across multiple time scales, including 30-minute, which is even smaller than the dataset's temporal resolution. Wanghan Xu, Fenghua Ling, Tao Han 0002, Hao Chen 0045, Wanli Ouyang, Lei Bai 0001 |
NeurIPS | 7 |
| 2024 | Hierarchical Diffusion Autoencoders and Disentangled Image ManipulationabstractDiffusion models have attained impressive visual quality for image synthesis. However, how to probe and manipulate the latent space of diffusion models has not been extensively explored. Prior work diffusion autoencoders encode the semantic representations with a single latent code, neglecting the low-level details and leading to entangled representations. To mitigate those limitations, we propose Hierarchical Diffusion Autoencoders (HDAE) that exploits the coarse-to-fine feature hierarchy for the latent space of diffusion models. Our HDAE converges 2+ times faster and encodes richer and more comprehensive coarse-to-fine representations of images. The hierarchical latent space inherently disentangles different semantic levels of features. Furthermore, we propose a truncated feature based approach for disentangled image manipulation. We demonstrate the effectiveness of our proposed HDAE with extensive experiments and applications on image reconstruction, style mixing, controllable interpolation, image editing, and multi-modal semantic image synthesis. The code will be released upon acceptance. Chengyue Wu, Yaohui Wang 0001, Lei Bai 0001, Yu Qiao 0001, Xihui Liu |
WACV | 5 |
| 2024 | Improving multiple sclerosis lesion segmentation across clinical sites: A federated learning approach with noise-resilient trainingabstractAccurately measuring the evolution of Multiple Sclerosis (MS) with magnetic resonance imaging (MRI) critically informs understanding of disease progression and helps to direct therapeutic strategy. Deep learning models have shown promise for automatically segmenting MS lesions, but the scarcity of accurately annotated data hinders progress in this area. Obtaining sufficient data from a single clinical site is challenging and does not address the heterogeneous need for model robustness. Conversely, the collection of data from multiple sites introduces data privacy concerns and potential label noise due to varying annotation standards. To address this dilemma, we explore the use of the federated learning framework while considering label noise. Our approach enables collaboration among multiple clinical sites without compromising data privacy under a federated learning paradigm that incorporates a noise-robust training strategy based on label correction. Specifically, we introduce a Decoupled Hard Label Correction (DHLC) strategy that considers the imbalanced distribution and fuzzy boundaries of MS lesions, enabling the correction of false annotations based on prediction confidence. We also introduce a Centrally Enhanced Label Correction (CELC) strategy, which leverages the aggregated central model as a correction teacher for all sites, enhancing the reliability of the correction process. Extensive experiments conducted on two multi-site datasets demonstrate the effectiveness and robustness of our proposed methods, indicating their potential for clinical applications in multi-site collaborations to train better deep learning models with lower cost in data collection and annotation. Lei Bai 0001, Dongang Wang, Hengrui Wang, Michael Barnett 0006, Mariano Cabezas, Tom Weidong Cai, Fernando Calamante, Kain Kyle, Dongnan Liu, Linda Ly, Aria Nguyen, Chun-Chien Shieh, Ryan Sullivan, Geng Zhan, Wanli Ouyang, Chenyu Wang 0001 |
Artif. Intell. Medicine | 1 |
| 2024 | Towards Frame Rate Agnostic Multi-object Tracking
Lei Bai 0001, Yongqiang Yao, Fengwei Yu, Wanli Ouyang |
Int. J. Comput. Vis. | 2 |
| 2024 | HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
Sha Zhang 0002, Jiajun Deng, Lei Bai 0001, Houqiang Li, Wanli Ouyang, Yanyong Zhang |
Int. J. Comput. Vis. | 3 |
| 2024 | Forecasting of Tropospheric Delay Using AI Foundation Models in Support of Microwave Remote SensingabstractAccurate tropospheric delay forecasts are imperative for microwave-based remote sensing techniques, playing a pivotal role in early warning and forecasting of natural disasters such as tsunamis, heavy rains, and hurricanes. Nevertheless, conventional methods for forecasting tropospheric delays entail substantial computational resources and high network transmission speeds, thereby restricting their real-time applicability in remote sensing operations. In this study, we introduce a novel approach to derive forecasted tropospheric delays using artificial intelligence (AI) weather forecast foundation models (FMs), exemplified by Huawei Cloud Pangu-Weather, Google DeepMind GraphCast, and Shanghai AI Lab FengWu. We assess the accuracy of these forecasts on a global scale employing fifth-generation ECMWF atmospheric re-analysis of the global climate (ERA5) (European Centre for Medium-Range Weather Forecasts (ECMWF) Reanalysis v5), ground-based Global Navigation Satellite System (GNSS), and in situ radiosonde (RS) measurements as reference data. Our results show that the FM-based scheme outperforms traditional methods in both forecast accuracy and length, with the ability to provide high-accuracy tropospheric delay parameters locally for 15-day forecasts at any location within minutes. Furthermore, the FM scheme still maintains accuracy better than empirical models when forecasting up to ten days in advance. This research demonstrates the potential of AI weather forecast FMs in delivering high-precision tropospheric delay medium-range forecasts and improvements for real-time remote sensing applications. Junsheng Ding, Xiaolong Mi, Wu Chen 0001, Junping Chen, Yize Zhang, Joseph L. Awange, Benedikt Soja, Lei Bai 0001, Yuanfan Deng |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2024 | Deriving Accurate Surface Meteorological States at Arbitrary Locations via Observation-Guided Continuous Neural Field ModelingabstractAccurately retrieving surface meteorological states at arbitrary locations is of great application significance in weather forecasting and climate modeling. Since meteorological variables are typically provided as coarse-resolution gridded fields, common methods that obtain the states at a specific location directly through spatial interpolation can lead to significant accuracy deviations compared to actual observations. Traditional downscaling, the process of obtaining fixed-scale high-resolution meteorological fields from low-resolution inputs, has been proposed as a way to indirectly improve the accuracy of retrieving states at arbitrary locations by providing more detailed subgrid-scale information. However, for arbitrary locations at the station scale, their states are influenced by subgrid information, resulting in systematic biases between the downscaled results after interpolation and the actual observations at specific station locations. To address this issue, in this article, we propose a new task called station-scale downscaling, which aims to directly derive accurate meteorological states at any given station location from a coarse-resolution meteorological field. To achieve this, we propose a new downscaling model based on hypernetwork architecture, namely, HyperDS, which efficiently integrates the multiscale observational information to guide the continuous neural field modeling of the meteorological variables, enabling accurate sampling of the states at any target location. Through extensive experiments, our proposed method outperforms other specially designed baseline models on multiple surface variables. Notably, the mean squared error (mse) for wind speed and surface pressure improved by 67% and 19.5% compared with other methods, respectively. Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Keyan Chen 0001, Wanli Ouyang, Zhengxia Zou, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MambaDS: Near-Surface Meteorological Field Downscaling With Topography Constrained Selective State-Space ModelingabstractIn an era of frequent extreme weather and global warming, obtaining precise, fine-grained near-surface weather forecasts is increasingly essential for human activities. Downscaling (DS), a crucial task in meteorological forecasting and remote sensing, enables the reconstruction of high-resolution meteorological states for target regions from global-scale forecast results. Previous downscaling methods, inspired by convolutional neural network (CNN) and Transformer-based super-resolution (SR) models, lacked tailored designs for meteorology and encountered structural limitations. Notably, they failed to efficiently integrate topography, a crucial prior to the downscaling process. In this article, we address these limitations by pioneering the selective state-space model (SSM) into the meteorological field downscaling and propose a novel model called MambaDS. This model retains the advantages of Mamba in long-range dependency modeling and linear computational complexity while enhancing the learning ability of multivariate correlation. In addition, by designing an efficient topography constraint layer, this prior information can be used more efficiently than ever before. Through extensive experiments in both China mainland and the continental United States (CONUS), we validated that our proposed MambaDS achieves state-of-the-art (SOTA) results in three different types of meteorological field downscaling settings. Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Wanli Ouyang, Zhengxia Zou, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | RS-Mamba for Large Remote Sensing Image Dense PredictionabstractContext modeling is critical for remote sensing image dense prediction tasks. Nowadays, the growing size of very-high-resolution (VHR) remote sensing images poses challenges in effectively modeling context. While transformer-based models possess global modeling capabilities, they encounter computational challenges when applied to large VHR images due to their quadratic complexity. The conventional practice of cropping large images into smaller patches results in a notable loss of contextual information. To address these issues, we propose the remote sensing Mamba (RSM) for dense prediction tasks in large VHR remote sensing images. RSM is specifically designed to capture the global context of remote sensing images with linear complexity, facilitating the effective processing of large VHR images. Considering that the land covers in remote sensing images are distributed in arbitrary spatial directions due to characteristics of remote sensing over-head imaging, the RSM incorporates an omnidirectional selective scan module (OSSM) to globally model the context of images in multiple directions, capturing large spatial features from various directions. We designed simple yet effective models based on RSM, achieving state-of-the-art performance on dense prediction tasks in VHR remote sensing images without fancy training strategies. Extensive experiments on semantic segmentation (SS) and change detection (CD) tasks across various land covers demonstrate the effectiveness of the proposed RSM. Leveraging the linear complexity and global modeling capabilities, RSM achieves better efficiency and accuracy than transformer-based models on large remote sensing images. Interestingly, we also demonstrated that our model generally performs better with a larger image size on dense prediction tasks. Sijie Zhao, Hao Chen 0045, Xueliang Zhang 0002, Pengfeng Xiao, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Self-Supervised Feature Learning for Appliance Recognition in Nonintrusive Load MonitoringabstractNonintrusive load monitoring (NILM) can monitor the operating state and energy consumption of electric appliances in a nonintrusive manner and provides a promising approach to improving electricity usage efficiency for residential and commercial buildings. Although machine learning (ML) methods are powerful and have significantly advanced the developments of NILM, they request a sizable amount of labeled data for model training. However, getting operational data of each electrical appliance in real life is challenging, so the requirements for labeled data limit the NLIM's practicality. To tackle this challenge, a novel multilayer momentum contrast (MLMoCo) learning mechanism is proposed for self-supervised feature representation learning. With only unlabeled aggregate load data, the proposed MLMoCo contrasts the augmented versions of the same sample (“positives”) with instances extracted from other samples (“negatives”). To maintain a dictionary with enough negative samples to be compared with the input, a momentum encoder is adopted to momentum update the parameters rather than by backpropagation during training. An event-based data augmentation method is also proposed to obtain the distinct but strongly related positive pairs for self-supervised feature learning. The experimental comparisons, including different state-of-the-art techniques and various downstream tasks with real-world datasets, demonstrate the remarkable performance gains of the proposed approach through learning from the unlabeled data, which could significantly increase the practicality of the NILM. Yinyan Liu, Lei Bai 0001, Jin Ma 0001, Wei Wang 0011, Wanli Ouyang |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A General Scenario-Agnostic Reinforcement Learning for Traffic Signal ControlabstractReinforcement learning (RL) can automatically learn a better policy through a trial-and-error paradigm and has been adopted to revolutionize and optimize traditional traffic signal control systems that are usually based on handcrafted methods. However, most existing RL-based models are either based on a single scenario or multiple independent scenarios, where each scenario has a separate simulation environment with predefined road network topology and traffic signal settings. These models implement training and testing in the same scenario, thus being strictly tied up with the specific setting and sacrificing model generalization heavily. While a few recent models could be trained by multiple scenarios, they require a huge amount of manual labor to label the intersection structure, hindering the model’s generalization. In this work, we aim at ageneralframework that could eliminate heavy labeling and model a variety of scenariossimultaneously. To this end, we propose a general Scenario-Agnostic (GESA) reinforcement learning framework for traffic signal control with: (1) A general plug-in module to map all different intersections into a unified structure, freeing us from the heavy manual labor to specify the structure of intersections; (2) A unified state and action space design to keep the model input and output consistently structured; (3) A large-scale co-training with multiple scenarios, leading to a generic traffic signal control algorithm. GESA can automatically handle various structured intersections from various cities without human labeling, and it co-trains a generalist agent to control traffic signals for multiple cities together, which also demonstrates superior transferability in zero-shot settings. In experiments, we demonstrate our algorithm as the first one that can be co-trained with seven different scenarios without manual annotation and gets13.27%higher rewards than baselines. When dealing with a new scenario, our model can still achieve9.39%higher rewards. The code, scenarios, and demos are available https://github.com/bonaldli/GESA. Haoyuan Jiang, Ziyue Li 0002, Zhishuai Li, Lei Bai 0001, Hangyu Mao, Wolfgang Ketter, Rui Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Adaptive and Interactive Multi-Level Spatio-Temporal Network for Traffic ForecastingabstractTraffic forecasting is a challenging research topic due to the complex spatial and temporal dependencies among different roads. Though great efforts have been made on traffic forecasting, existing works still have the following shortcomings: i) Most methods only directly perform on the original road network topology which cannot accommodate the diverse traffic patterns and multi-granularity traffic forecasting requirements driven by the natural multi-level urban structure and layout, ii) The existing studies based on the spatio-temporal multi-granularity perspective ignore the interactions between the fine-grained information and coarse-grained information, resulting in the spatio-temporal correlation under multi-granularity inaccurately modeled. To solve the problems, we propose an Adaptive and Interactive Multi-level Spatio-Temporal network (AIMST) for traffic forecasting. Specifically, we first devise a learnable adaptive hierarchical clustering method to automatically generate more coarse-grained graphs from the initial road networks and the traffic data. Then, the spatio-temporal graph convolutional networks are executed on the constructed hierarchical traffic graph of each level correspondingly to capture the spatio-temporal patterns. Furthermore, a multi-level bidirectional interaction module is designed to emphasize the multi-grained interaction patterns among different levels. Extensive experiments on two real-world traffic datasets demonstrate that our framework is superior to several state-of-the-art baselines. Yudong Zhang 0005, Pengkun Wang 0001, Binwu Wang, Xu Wang 0029, Zhe Zhao 0008, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Similarity- and Quality-Guided Relation Learning for Joint Detection and TrackingabstractJoint detection and tracking, which solves two fundamental vision challenges in a unified manner, is a challenging topic in computer vision. In this area, the proper use of spatial-temporal information in videos can help reduce local defects and improve the quality of feature representations. Although modeling low-level (usually pixel-wise) spatial-temporal information has been studied, instance-level spatial-temporal correlations (i.e., relations between semantic regions in which instances have occurred) have not been fully exploited. In comparison, modeling instance-level correlation is a more flexible and reasonable way to enhance feature representations. However, we have found that conventional instance-level relation learning that works for the separate tasks of detection or tracking is not effective in joint tasks in which a variety of scenarios may be presented. To try to resolve this problem, in this study, we effectively exploited instance-level spatial-temporal semantic information for joint detection and tracking via a joint relation learning pipeline with a novel relation learning mechanism called Similarity- and Quality-Guided Attention (SQGA). Specifically, we added task-specific SQGA relation modules before the corresponding task prediction heads to refine the instance feature representation using features of other reference instances in the neighboring frames; these features are aggregated on the basis of relational affinities. In particular, in SQGA, relational affinities were factorized to similarity and quality terms so that fine-grained supervision rules could be applied. Then we added task-specific attention losses for each SQGA relation module, resulting in a better feature aggregation for the corresponding task. Quantitative experiments based on several challenging multi-object tracking benchmarks showed that our approach was more effective than the baselines and provided competitive results compared with recent state-of-the-art methods. Lei Bai 0001, Yongqiang Yao, Weihao Gan, Wei Wu 0021, Wanli Ouyang |
IEEE Trans. Multim. | 2 |
| 2024 | Relation-Aware Distribution Representation Network for Person Clustering With Multiple ModalitiesabstractPerson clustering with multi-modal clues, including faces, bodies, and voices, is critical for various tasks, such as movie parsing and identity-based movie editing. Related methods such as multi-view clustering mainly project multi-modal features into a joint feature space. However, multi-modal clue features are usually rather weakly correlated due to the semantic gap from the modality-specific uniqueness. As a result, these methods are not suitable for person clustering. In this paper, we propose aRelation-AwareDistribution representation Network (RAD-Net) to generate adistribution representationfor multi-modal clues. The distribution representation of a clue is a vector consisting of the relation between this clue and all other clues from all modalities, thus beingmodality agnosticand good for person clustering. Accordingly, we introduce a graph-based method to construct distribution representation and employ a cyclic update policy to refine distribution representation progressively. Our method achieves substantial improvements of+6%and+8.2%in F-score on the Video Person-Clustering Dataset (VPCD) and VoxCeleb2 multi-view clustering dataset, respectively. Codes will be released athttps://github.com/bonaldli/RADNet. Kaijian Liu, Shixiang Tang, Ziyue Li 0002, Zhishuai Li, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Multi-Scale Control Signal-Aware Transformer for Motion Synthesis without PhaseabstractSynthesizing controllable motion for a character using deep learning has been a promising approach due to its potential to learn a compact model without laborious feature engineering. To produce dynamic motion from weak control signals such as desired paths, existing methods often require auxiliary information such as phases for alleviating motion ambiguity, which limits their generalisation capability. As past poses often contain useful auxiliary hints, in this paper, we propose a task-agnostic deep learning method, namely Multi-scale Control Signal-aware Transformer (MCS-T), with an attention based encoder-decoder architecture to discover the auxiliary information implicitly for synthesizing controllable motion without explicitly requiring auxiliary information such as phase. Specifically, an encoder is devised to adaptively formulate the motion patterns of a character's past poses with multi-scale skeletons, and a decoder driven by control signals to further synthesize and predict the character's state by paying context-specialised attention to the encoded past motion patterns. As a result, it helps alleviate the issues of low responsiveness and slow transition which often happen in conventional methods not using auxiliary information. Both qualitative and quantitative experimental results on an existing biped locomotion dataset, which involves diverse types of motion transitions, demonstrate the effectiveness of our method. In particular, MCS-T is able to successfully generate motions comparable to those generated by the methods using auxiliary information. Lintao Wang 0002, Kun Hu 0008, Lei Bai 0001, Wanli Ouyang, Zhiyong Wang 0001 |
AAAI | 3 |
| 2023 | UniHCP: A Unified Model for Human-Centric PerceptionsabstractHuman-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have their own relevant semantic aspect to focus on, they also share the same underlying semantic structure of the human body. However, few works have attempted to exploit such homogeneity and design a general-propose model for human-centric tasks. In this work, we revisit a broad range of human-centric tasks and unify them in a minimalist manner. We propose UniHCP, a Unified Model for Human-Centric Perceptions, which unifies a wide range of human-centric tasks in a simplified end-to-end manner with the plain vision transformer architecture. With large-scale joint training on 33 human-centric datasets, UniHCP can outperform strong baselines on several in-domain and downstream tasks by direct evaluation. When adapted to a specific task, UniHCP achieves new SOTAs on a wide range of human-centric tasks, e.g., 69.8 mIoU on CIHP for human parsing, 86.18 mA on PA100K for attribute prediction, 90.3 mAP on Market1501 for ReID, and 85.8 JI on CrowdHuman for pedestrian detection, performing better than specialized models tailored for each task. The code and pretrained model are available at https://github.com/OpenGVLab/UniHCP. Yuanzheng Ci, Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Fengwei Yu, Donglian Qi, Wanli Ouyang |
CVPR | 5 |
| 2023 | HumanBench: Towards General Human-Centric Perception with Projector Assisted PretrainingabstractHuman-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench. Shixiang Tang, Qingsong Xie, Meilin Chen, Yizhou Wang 0007, Yuanzheng Ci, Lei Bai 0001, Feng Zhu 0006, Haiyang Yang, Rui Zhao 0001, Wanli Ouyang |
CVPR | 7 |
| 2023 | FEND: A Future Enhanced Distribution-Aware Contrastive Learning Framework for Long-Tail Trajectory PredictionabstractPredicting the future trajectories of the traffic agents is a gordian technique in autonomous driving. However, trajectory prediction suffers from data imbalance in the prevalent datasets, and the tailed data is often more complicated and safety-critical. In this paper, we focus on dealing with the long-tail phenomenon in trajectory prediction. Previous methods dealing with long-tail data did not take into account the variety of motion patterns in the tailed data. In this paper, we put forward a future enhanced contrastive learning framework to recognize tail trajectory patterns and form a feature space with separate pattern clusters. Furthermore, a distribution aware hyper predictor is brought up to better utilize the shaped feature space. Our method is a model-agnostic framework and can be plugged into many well-known baselines. Experimental results show that our framework outperforms the state-of-the-art long-tail prediction method on tailed samples by 9.5% on ADE and 8.5% on FDE, while maintaining or slightly improving the averaged performance. Our method also surpasses many long-tail techniques on trajectory prediction task. Pu Zhang 0001, Lei Bai 0001, Jianru Xue |
CVPR | 3 |
| 2023 | Long-Tailed Time Series Classification via Feature Space Rebalancing
Pengkun Wang 0001, Xu Wang 0029, Binwu Wang, Yudong Zhang 0005, Lei Bai 0001, Yang Wang 0015 |
DASFAA (1) | 5 |
| 2023 | A Knowledge-Driven Memory System for Traffic Flow Prediction
Binwu Wang, Yudong Zhang 0005, Pengkun Wang 0001, Xu Wang 0029, Lei Bai 0001, Yang Wang 0015 |
DASFAA (4) | 5 |
| 2023 | Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text RetrievalabstractVideo-text retrieval has greatly benefited from the massive web video in recent years, while the performance is still limited to the weak supervision from the uncurated data. In this work, we propose to leverage the well-represented information of each original modality and exploit complementary information in two views of the same video, i.e., video clips and captions, by using one view to obtain positive samples with the neighboring samples of the other. Respecting the hierarchical organization of real-world data, we further design a hierarchical cross-modal pre-training method (HCP) to learn good representations in the common embedding space. We evaluate the pre-trained model on three downstream tasks, i.e. text-to-video retrieval, action step localization and video question answering and our method outperforms previous works under the same setting. Yuhan Xiang, Kaijian Liu, Shixiang Tang, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Xianming Lin |
ICASSP | 4 |
| 2023 | STEERER: Resolving Scale Variations for Counting and Localization via Selective Inheritance LearningabstractScale variation is a deep-rooted problem in object counting, which has not been effectively addressed by existing scale-aware algorithms. An important factor is that they typically involve cooperative learning across multi-resolutions, which could be suboptimal for learning the most discriminative features from each scale. In this paper, we propose a novel method termed STEERER (SelecTivE inhERitance lEaRning) that addresses the issue of scale variations in object counting. STEERER selects the most suitable scale for patch objects to boost feature extraction and only inherits discriminative features from lower to higher resolution progressively. The main insights of STEERER are a dedicated Feature Selection and Inheritance Adaptor (FSIA), which selectively forwards scale-customized features at each scale, and a Masked Selection and Inheritance Loss (MSIL) that helps to achieve high-quality density maps across all scales. Our experimental results on nine datasets with counting and localization tasks demonstrate the unprecedented scale generalization ability of STEERER. Code is available at https://github.com/taohan10200/STEERER. Tao Han 0002, Lei Bai 0001, Lingbo Liu, Wanli Ouyang |
ICCV | 2 |
| 2023 | Cycle-consistent Masked AutoEncoder for Unsupervised Domain Generalization
Haiyang Yang, Shixiang Tang, Feng Zhu 0006, Yizhou Wang 0007, Meilin Chen, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ICLR | 7 |
| 2023 | GReTo: Remedying dynamic graph topology-task discordance via target homophily
Zhengyang Zhou, Qihe Huang, Gengyu Lin, Kuo Yang 0002, Lei Bai 0001, Yang Wang 0015 |
ICLR | 5 |
| 2023 | MM-DAG: Multi-task DAG Learning for Multi-modal Data - with Application for Traffic Congestion AnalysisabstractThis paper proposes to learn Multi-task, Multi-modal Direct Acyclic Graphs (MM-DAGs), which are commonly observed in complex systems, e.g., traffic, manufacturing, and weather systems, whose variables are multi-modal with scalars, vectors, and functions. This paper takes the traffic congestion analysis as a concrete case, where a traffic intersection is usually regarded as a DAG. In a road network of multiple intersections, different intersections can only have someoverlapping and distinct variables observed. For example, a signalized intersection has traffic light-related variables, whereas unsignalized ones do not. This encourages the multi-task design: with each DAG as a task, the MM-DAG tries to learn the multiple DAGs jointly so that their consensus and consistency are maximized. To this end, we innovatively propose a multi-modal regression for linear causal relationship description of different variables. Then we develop a novel Causality Difference (CD) measure and its differentiable approximator. Compared with existing SOTA measures, CD can penalize the causal structural difference among DAGs with distinct nodes and can better consider the uncertainty of causal orders. We rigidly prove our design's topological interpretation and consistency properties. We conduct thorough simulations and one case study to show the effectiveness of our MM-DAG. The code is available under https://github.com/Lantian72/MM-DAG. Ziyue Li 0002, Zhishuai Li, Lei Bai 0001, Man Li 0003, Fugee Tsung, Wolfgang Ketter, Rui Zhao 0001, Chen Zhang 0007 |
KDD | 4 |
| 2023 | Pattern Expansion and Consolidation on Evolving Graphs for Continual Traffic PredictionabstractRecently, spatiotemporal graph convolutional networks are becoming popular in the field of traffic flow prediction and significantly improve prediction accuracy. However, the majority of existing traffic flow prediction models are tailored to static traffic networks and fail to model the continuous evolution and expansion of traffic networks. In this work, we move to investigate the challenge of traffic flow prediction on an expanding traffic network. And we propose an efficient and effective continual learning framework to achieve continuous traffic flow prediction without the access to historical graph data, namely Pattern Expansion and Consolidation based on Pattern Matching based (PECPM). Specifically, we first design a pattern bank based on pattern matching to store representative patterns of the road network. With the expansion of the road network, the model configured with such a bank module can achieve continuous traffic prediction by effectively managing patterns stored in the bank. The core idea is to continuously update new patterns while consolidating learned ones. Specifically, we design a pattern expansion mechanism that can detect evolved and new patterns from the updated network, then these unknown patterns are expanded into the pattern bank to adapt to the updated road network. Additionally, we propose a pattern consolidation mechanism that includes both a bank preservation mechanism and a pattern traceability mechanism. This can effectively consolidate the learned patterns in the bank without requiring access to detailed historical graph data. We construct experiments on real-world traffic datasets to demonstrate the competitive performance, superior efficiency, and strong generalization ability of PECPM. Binwu Wang, Yudong Zhang 0005, Xu Wang 0029, Pengkun Wang 0001, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
KDD | 6 |
| 2023 | LargeST: A Benchmark Dataset for Large-Scale Traffic ForecastingabstractRoad traffic forecasting plays a critical role in smart city initiatives and has experienced significant advancements thanks to the power of deep learning in capturing non-linear patterns of traffic data. However, the promising results achieved on current public datasets may not be applicable to practical scenarios due to limitations within these datasets. First, the limited sizes of them may not reflect the real-world scale of traffic networks. Second, the temporal coverage of these datasets is typically short, posing hurdles in studying long-term patterns and acquiring sufficient samples for training deep models. Third, these datasets often lack adequate metadata for sensors, which compromises the reliability and interpretability of the data. To mitigate these limitations, we introduce the LargeST benchmark dataset. It encompasses a total number of 8,600 sensors in California with a 5-year time coverage and includes comprehensive metadata. Using LargeST, we perform in-depth data analysis to extract data insights, benchmark well-known baselines in terms of their performance and efficiency, and identify challenges as well as opportunities for future research. We release the datasets and baseline implementations at: https://github.com/liuxu77/LargeST. Xu Liu 0014, Yutong Xia, Yuxuan Liang 0002, Junfeng Hu 0001, Yiwei Wang 0001, Lei Bai 0001, Chao Huang 0001, Zhenguang Liu, Bryan Hooi, Roger Zimmermann |
NeurIPS | 6 |
| 2023 | Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated ImagesabstractPhotos serve as a way for humans to record what they experience in their daily lives, and they are often regarded as trustworthy sources of information. However, there is a growing concern that the advancement of artificial intelligence (AI) technology may produce fake photos, which can create confusion and diminish trust in photographs. This study aims to comprehensively evaluate agents for distinguishing state-of-the-art AI-generated visual content. Our study benchmarks both human capability and cutting-edge fake image detection AI algorithms, using a newly collected large-scale fake image dataset Fake2M. In our human perception evaluation, titled HPBench, we discovered that humans struggle significantly to distinguish real photos from AI-generated ones, with a misclassification rate of 38.7\%. Along with this, we conduct the model capability of AI-Generated images detection evaluation MPBench and the top-performing model from MPBench achieves a 13\% failure rate under the same setting used in the human evaluation.We hope that our study can raise awareness of the potential risks of AI-generated images and facilitate further research to prevent the spread of false information. More information can refer to https://github.com/Inf-imagine/Sentry. Lei Bai 0001, Jingjing Qu, Chengyue Wu, Xihui Liu, Wanli Ouyang |
NeurIPS | 3 |
| 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and BenchmarkabstractLarge language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interaction with the world extends beyond only text as a modality, and other modalities such as vision are also crucial. Recent works on multi-modal large language models, such as GPT-4V and Bard, have demonstrated their effectiveness in handling visual modalities. However, the transparency of these works is limited and insufficient to support academic research. To the best of our knowledge, we present one of the very first open-source endeavors in the field, LAMM, encompassing a Language-Assisted Multi-Modal instruction tuning dataset, framework, and benchmark. Our aim is to establish LAMM as a growing ecosystem for training and evaluating MLLMs, with a specific focus on facilitating AI agents capable of bridging the gap between ideas and execution, thereby enabling seamless human-AI interaction. Our main contribution is three-fold: 1) We present a comprehensive dataset and benchmark, which cover a wide range of vision tasks for 2D and 3D vision. Extensive experiments validate the effectiveness of our dataset and benchmark. 2) We outline the detailed methodology of constructing multi-modal instruction tuning datasets and benchmarks for MLLMs, enabling rapid scaling and extension of MLLM research to diverse domains, tasks, and modalities. 3) We provide a primary but potential MLLM training framework optimized for modality extension. We also provide baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained within 24 A100 GPU hours, framework supports training with V100 and RTX3090 is available thanks to the open-source society. Codes and data are now available at https://openlamm.github.io. Zhenfei Yin, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang 0001, Lu Sheng, Lei Bai 0001, Wanli Ouyang |
NeurIPS | 10 |
| 2023 | Online Metro Origin-Destination Prediction via Heterogeneous Information AggregationabstractMetro origin-destination prediction is a crucial yet challenging time-series analysis task in intelligent transportation systems, which aims to accurately forecast two specific types of cross-station ridership, i.e., Origin-Destination (OD) one and Destination-Origin (DO) one. However, complete OD matrices of previous time intervals can not be obtained immediately in online metro systems, and conventional methods only used limited information to forecast the future OD and DO ridership separately. In this work, we proposed a novel neural network module termed Heterogeneous Information Aggregation Machine (HIAM), which fully exploits heterogeneous information of historical data (e.g., incomplete OD matrices, unfinished order vectors, and DO matrices) to jointly learn the evolutionary patterns of OD and DO ridership. Specifically, an OD modeling branch estimates the potential destinations of unfinished orders explicitly to complement the information of incomplete OD matrices, while a DO modeling branch takes DO matrices as input to capture the spatial-temporal distribution of DO ridership. Moreover, a Dual Information Transformer is introduced to propagate the mutual information among OD features and DO features for modeling the OD-DO causality and correlation. Based on the proposed HIAM, we develop a unified Seq2Seq network to forecast the future OD and DO ridership simultaneously. Extensive experiments conducted on two large-scale benchmarks demonstrate the effectiveness of our method for online metro origin-destination prediction. Our code is resealed at https://github.com/HCPLab-SYSU/HIAM. Lingbo Liu, Yuying Zhu 0007, Guanbin Li, Lei Bai 0001, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Towards Trajectory Forecasting From DetectionabstractTrajectory forecasting for traffic participants (e.g., vehicles) is critical for autonomous platforms to make safe plans. Currently, most trajectory forecasting methods assume that object trajectories have been extracted and directly develop trajectory predictors based on the ground truth trajectories. However, this assumption does not hold in practical situations. Trajectories obtained from object detection and tracking are inevitably noisy, which could cause serious forecasting errors to predictors built on ground truth trajectories. In this paper, we propose to predict trajectories directly based on detection results without relying on explicitly formed trajectories. Different from traditional methods which encode the motion cues of an agent based on its clearly defined trajectory, we extract the motion information only based on the affinity cues among detection results, in which an affinity-aware state update mechanism is designed to manage the state information. In addition, considering that there could be multiple plausible matching candidates, we aggregate the states of them. These designs take the uncertainty of association into account which relax the undesirable effect of noisy trajectory obtained from data association and improve the robustness of the predictor. Extensive experiments validate the effectiveness of our method and its generalization ability to different detectors or forecasting schemes. Pu Zhang 0001, Lei Bai 0001, Jianwu Fang, Jianru Xue, Nanning Zheng 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | SiamSampler: Video-Guided Sampling for Siamese Visual TrackingabstractWe present SiamSampler, the first to our knowledge investigating video sampling in visual object tracking. We observe that the random sampling applied in Siamese-based trackers cannot focus on important data or ensure data diversity, hindering the effective training of networks. This paper proposes the Video-Guided Sampling Strategy to solve the problems in random sampling from both inter and intra-video levels. At the inter-video level, we propose Modified Gaussian Sampling Strategy (MGSS) to automatically assign higher sampling probabilities to longer and more difficult videos and reduce the sampling probabilities of shorter and easier videos. At the intra-video level, the Farthest Image Pair Sampling Strategy (FPSS) is proposed to increase the diversity of training data. Extensive experiments on general benchmarks demonstrate the effectiveness of our method. Compared with the baseline model, our method improves tracking performance on five datasets, without affecting the testing speed. Peixia Li, Lei Bai 0001, Lei Qiao 0004, Bo Li 0114, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Cure-GNN: A Robust Curvature-Enhanced Graph Neural Network Against Adversarial AttacksabstractGraph neural networks (GNNs) are a specialized type of deep learning models on graphs by learning aggregations over neighbor nodes. However, recent studies reveal that the performance of GNNs are severely deteriorated by injecting adversarial examples. Hence, improving the robustness of GNNs is of significant importance. Prior works are devoted to reducing the influence of direct adversaries which are adversarial attacks by positioning a node's one-hop neighbors, yet these approaches are limited in protecting GNNs from indirect adversarial attacks within a node's multi-hop neighbors. In this work, we approach this problem from a new angle by exploring the graph Ricci curvature, which can characterize the relationships of both direct and indirect links from any two nodes’ neighborhoods in the Riemannian space. We first investigate the distinguishable properties of adversarial attacks with graph Ricci curvature distribution. Then, a novel defense framework called Cure-GNN is proposed to detect and mitigate adversarial effects. Cure-GNN discerns the distinction between adversarial edges and normal edges via computing curvature, and merges it into the node features reconstructed by a residual learning framework. Extensive experiments over real-world datasets on node classification task demonstrate the efficacy of Cure-GNN and achieves superiority to the state-of-the-arts without incurring high complexity. Yang Xiao 0014, Zhuolin Xing, Alex X. Liu, Lei Bai 0001, Qingqi Pei, Lina Yao 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Knowledge Expansion and Consolidation for Continual Traffic Prediction With Expanding GraphsabstractAccurate traffic prediction plays a vital role in intelligent transport managements and applications. However, in the vast majority of existing works, the focus is mainly on modeling spatiotemporal correlations in static traffic networks. Thus, the continuous expansion and evolution of traffic networks are ignored. In this work, we study the problem of traffic prediction with expanding road network structures under the continual learning paradigm. Considering the model prediction performance, efficiency, and data accessibility, a SpatioTemporal Knowledge Expansion and Consolidation (STKEC) framework is proposed. This framework contains an influence-based knowledge expansion strategy to help the spatiotemporal learning model integrate new spatiotemporal traffic patterns and a memory-augmented knowledge consolidation mechanism to preserve the learned spatiotemporal patterns without accessing the data in previous graphs. Extensive experiments are conducted on a large-scale dataset and verify the superior performance of STKEC in continual traffic prediction. Binwu Wang, Yudong Zhang 0005, Pengkun Wang 0001, Xu Wang 0029, Lei Bai 0001, Yang Wang 0015 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | SepFusion: Finding Optimal Fusion Structures for Visual Sound SeparationabstractMultiple modalities can provide rich semantic information; and exploiting such information will normally lead to better performance compared with the single-modality counterpart. However, it is not easy to devise an effective cross-modal fusion structure due to the variations of feature dimensions and semantics, especially when the inputs even come from different sensors, as in the field of audio-visual learning. In this work, we propose SepFusion, a novel framework that can smoothly produce optimal fusion structures for visual-sound separation. The framework is composed of two components, namely the model generator and the evaluator. To construct the generator, we devise a lightweight architecture space that can adapt to different input modalities. In this way, we can easily obtain audio-visual fusion structures according to our demands. For the evaluator, we adopt the idea of neural architecture search to select superior networks effectively. This automatic process can significantly save human efforts while achieving competitive performances. Moreover, since our SepFusion provides a series of strong models, we can utilize the model family for broader applications, such as further promoting performance via model assembly, or providing suitable architectures for the separation of certain instrument classes. These potential applications further enhance the competitiveness of our approach. Dongzhan Zhou, Xinchi Zhou, Di Hu 0001, Hang Zhou 0009, Lei Bai 0001, Ziwei Liu 0002, Wanli Ouyang |
AAAI | 5 |
| 2022 | Jointly Contrastive Representation Learning on Road Network and TrajectoryabstractRoad network and trajectory representation learning are essential for traffic systems since the learned representation can be directly used in various downstream tasks (e.g., traffic speed inference, travel time estimation). However, most existing methods only contrast within the same scale, i.e., treating road network and trajectory separately, which ignores valuable inter-relations. In this paper, we aim to propose a unified framework that jointly learns the road network and trajectory representations end-to-end. We design domain-specific augmentations for road-road contrast and trajectory-trajectory contrast separately, i.e., road segment with its contextual neighbors and trajectory with its detour replaced and dropped alternatives, respectively. On top of that, we further introduce the road-trajectory cross-scale contrast to bridge the two scales by maximizing the total mutual information. Unlike the existing cross-scale contrastive learning methods on graphs that only contrast a graph and its belonging nodes, the contrast between road segment and trajectory is elaborately tailored via novel positive sampling and adaptive weighting strategies. We conduct prudent experiments based on two real-world datasets with four downstream tasks, demonstrating improved performance and effectiveness. Zhenyu Mao, Ziyue Li 0002, Dedong Li, Lei Bai 0001, Rui Zhao 0001 |
CIKM | 4 |
| 2022 | DR.VIC: Decomposition and Reasoning for Video Individual CountingabstractPedestrian counting is a fundamental tool for under-standing pedestrian patterns and crowd flow analysis. Existing works (e.g., image-level pedestrian counting, cross-line crowd counting et al.) either only focus on the image-level counting or are constrained to the manual annotation of lines. In this work, we propose to conduct the pedes-trian counting from a new perspective - Video Individual Counting (VIC), which counts the total number of individual pedestrians in the given video (a person is only counted once). Instead of relying on the Multiple Object Tracking (MOT) techniques, we propose to solve the problem by decomposing all pedestrians into the initial pedestrians who existed in the first frame and the new pedestrians with separate identities in each following frame. Then, an end-to-end Decomposition and Reasoning Network (DRNet) is designed to predict the initial pedestrian count with the density estimation method and reason the new pedestrian's count of each frame with the differentiable optimal transport. Extensive experiments are conducted on two datasets with congested pedestrians and diverse scenes, demonstrating the effectiveness of our method over baselines with great superiority in counting the individual pedestrians. Code: https://github.com/taohan10200/DRNet. Tao Han 0002, Lei Bai 0001, Junyu Gao 0001, Qi Wang 0009, Wanli Ouyang |
CVPR | 2 |
| 2022 | Revisiting the Transferability of Supervised Pretraining: an MLP PerspectiveabstractThe pretrain-finetune paradigm is a classical pipeline in visual learning. Recent progress on unsupervised pretraining methods shows superior transfer performance to their supervised counterparts. This paper revisits this phenomenon and sheds new light on understanding the transferability gap between unsupervised and supervised pretraining from a multilayer perceptron (MLP) perspective. While previous works [6], [8], [17] focus on the effectiveness of MLP on unsupervised image classification where pretraining and evaluation are conducted on the same dataset, we reveal that the MLP projector is also the key factor to better transferability of unsupervised pretraining methods than supervised pretraining methods. Based on this observation, we attempt to close the transferability gap between supervised and unsupervised pretraining by adding an MLP projector before the classifier in supervised pretraining. Our analysis indicates that the MLP projector can help retain intra-class variation of visual features, decrease the feature distribution distance between pretraining and evaluation datasets, and reduce feature redundancy. Extensive experiments on public benchmarks demonstrate that the added MLP projector significantly boosts the transferability of supervised pretraining, e.g. +7.2% top-1 accuracy on the concept generalization task, +5.8% top-1 accuracy for linear evaluation on 12 -domain classification tasks, and +0.8% AP on COCO object detection task, making supervised pretraining comparable or even better than unsupervised pretraining. Yizhou Wang 0007, Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Donglian Qi, Wanli Ouyang |
CVPR | 4 |
| 2022 | Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking
Peixia Li, Lei Bai 0001, Lei Qiao 0004, Qiuhong Shen, Bo Li 0114, Weihao Gan, Wei Wu 0021, Wanli Ouyang |
ECCV (22) | 3 |
| 2022 | Fast-MoCo: Boost Momentum-Based Contrastive Learning with Combinatorial Patches
Yuanzheng Ci, Chen Lin 0003, Lei Bai 0001, Wanli Ouyang |
ECCV (26) | 3 |
| 2022 | Unifying Visual Contrastive Learning for Object Recognition from a Graph Perspective
Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Chenyu Wang 0001, Wanli Ouyang |
ECCV (26) | 3 |
| 2022 | Relative Contrastive Loss for Unsupervised Representation Learning
Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ECCV (27) | 3 |
| 2022 | Domain Invariant Masked Autoencoders for Self-supervised Learning from Multi-domains
Haiyang Yang, Shixiang Tang, Meilin Chen, Yizhou Wang 0007, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ECCV (31) | 6 |
| 2022 | Countering Modal Redundancy and Heterogeneity: A Self-Correcting Multimodal FusionabstractFusing multimodal heterogeneous data plays a vital role in recognition and prediction tasks in various fields, e.g., action recognition and traffic accident forecast. Yet, there remain some key challenges, such as heterogeneous feature interaction and feature redundancies, that significantly affect the performance of multimodal fusion. To tackle these challenges, we first devise a Unified Feature Interaction Module (UFIM) in which a novel orthogonal attention component is designed to obtain fine-grained inter-modal interaction information among heterogeneous features. Then, we propose a novel Self-Correcting Transformer Module (SCTM) which employs a modified transformer to obtain the one-to-many correlation information between the current modal feature and the merged features of other modalities to alleviate the redundancy problem. Extensive experiments on four cross-domain tasks demonstrate the effectiveness and generalization ability of our proposed method. Pengkun Wang 0001, Xu Wang 0029, Binwu Wang, Yudong Zhang 0005, Lei Bai 0001, Yang Wang 0015 |
ICDM | 5 |
| 2022 | Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningabstractUnsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to produce object priors, \emph{e.g.,} selective search, which separates the prior generation and detector learning and leads to sub-optimal solutions. In this work, we propose a novel object detection pretraining framework that could generate object priors and learn detectors jointly by generating accurate object priors from the model itself. Specifically, region priors are extracted by attention maps from the encoder, which highlights foregrounds. Instance priors are the selected high-quality output bounding boxes of the detection decoder. By assuming objects as instances in the foreground, we can generate object priors with both region and instance priors. Moreover, our object priors are jointly refined along with the detector optimization. With better object priors as supervision, the model could achieve better detection capability, which in turn promotes the object priors generation. Our method improves the competitive approaches by \textbf{+1.3 AP}, \textbf{+1.7 AP} in 1\% and 10\% COCO low-data regimes object detection. Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Feng Zhu 0006, Haiyang Yang, Lei Bai 0001, Rui Zhao 0001, Yunfeng Yan, Donglian Qi, Wanli Ouyang |
NeurIPS | 6 |
| 2022 | Face to purchase: Predicting consumer choices with structured facial and behavioral traits embedding
Zhe Liu 0023, Xianzhi Wang 0001, Lina Yao 0001, Jake An, Lei Bai 0001, Ee-Peng Lim |
Knowl. Based Syst. | 6 |
| 2022 | Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection for Autonomous DrivingabstractThe strong demand of autonomous driving in the industry has led to vigorous interest in 3D object detection and resulted in many excellent 3D object detection algorithms. However, the vast majority of algorithms only model single-frame data, ignoring the temporal clue in video sequence. In this work, we propose a new transformer, called Temporal-Channel Transformer (TCTR), to model the temporal-channel domain and spatial-wise relationships for video object detecting from Lidar data. As the special design of this transformer, the information encoded in the encoder is different from that in the decoder. The encoder encodes temporal-channel information of multiple frames while the decoder decodes the spatial-wise information for the current frame in a voxel-wise manner. Specifically, the temporal-channel encoder of the transformer is designed to encode the information of different channels and frames by utilizing the correlation among features from different channels and frames. On the other hand, the spatial decoder of the transformer decodes the information for each location of the current frame. Before conducting the object detection with detection head, a gate mechanism is further deployed for re-calibrating the features of current frame, which filters out the object-irrelevant information by repetitively refining the representation of target frame along with the up-sampling process. Experimental results reveal that TCTR achieves the state-of-the-art performance in grid voxel-based 3D object detection on the nuScenes benchmark. Zhenxun Yuan, Xiao Song 0002, Lei Bai 0001, Zhe Wang 0006, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Action Recognition With Motion Diversification and Dynamic SelectionabstractMotion modeling is crucial in modern action recognition methods. As motion dynamics like moving tempos and action amplitude may vary a lot in different video clips, it poses great challenge on adaptively covering proper motion information. To address this issue, we introduce a Motion Diversification and Selection (MoDS) module to generate diversified spatio-temporal motion features and then select the suitable motion representation dynamically for categorizing the input video. To be specific, we first propose a spatio-temporal motion generation (StMG) module to construct a bank of diversified motion features with varying spatial neighborhood and time range. Then, a dynamic motion selection (DMS) module is leveraged to choose the most discriminative motion feature both spatially and temporally from the feature bank. As a result, our proposed method can make full use of the diversified spatio-temporal motion information, while maintaining computational efficiency at the inference stage. Extensive experiments on five widely-used benchmarks, demonstrate the effectiveness of the method and we achieve state-of-the-art performance on Something-Something V1 & V2 that are of large motion variation. Peiqin Zhuang, Luping Zhou, Lei Bai 0001, Ding Liang, Zhiyong Wang 0001, Yali Wang 0001, Wanli Ouyang |
IEEE Trans. Image Process. | 5 |
| 2022 | Graph Neural Network for Robust Public Transit Demand PredictionabstractUnderstanding and forecasting mobility patterns and travel demand are fundamental and critical to efficient transport infrastructure planning and service operation. However, most existing studies focused on deterministic demand estimation/prediction/analytics. Differently, this study provides confidence interval based demand forecasting, which can help transport planning and operation authorities to better accommodate demand uncertainty/variability. The proposed Origin-Destination (OD) demand prediction approach well captures and utilizes the correlations among spatial and temporal information. In particular, the proposed Probabilistic Graph Convolution Model (PGCM) consists of two components: (i) a prediction module based on Graph Convolution Network and combined with the gated mechanism to predict OD demand by utilizing spatio-temporal relations; (ii) a Bayesian-based approximation module to measure the confidence interval of demand prediction by evaluating the graph-based model uncertainty. We use a large-scale real-world public transit dataset from the Greater Sydney area to test and evaluate the proposed approach. The experimental results demonstrate that the proposed method is capable of capturing the spatial-temporal correlations for more robust demand prediction against several established tools in the literature. Can Li 0014, Lei Bai 0001, Wei Liu 0101, Lina Yao 0001, S. Travis Waller |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Mutual CRF-GNN for Few-Shot LearningabstractGraph-neural-networks (GNN) is a rising trend for fewshot learning. A critical component in GNN is the affinity. Typically, affinity in GNN is mainly computed in the feature space, e.g., pairwise features, and does not take fully advantage of semantic labels associated to these features. In this paper, we propose a novel Mutual CRF-GNN (MCGN). In this MCGN, the labels and features of support data are used by the CRF for inferring GNN affinities in a principled and probabilistic way. Specifically, we construct a Conditional Random Field (CRF) conditioned on labels and features of support data to infer a affinity in the label space. Such affinity is fed to the GNN as the node-wise affinity. GNN and CRF mutually contributes to each other in MCGN. For GNN, CRF provides valuable affinity information. For CRF, GNN provides better features for inferring affinity. Experimental results show that our approach outperforms stateof-the-arts on datasets miniImageNet, tieredImageNet, and CIFAR-FS on both 5-way 1-shot and 5-way 5-shot settings. Shixiang Tang, Dapeng Chen, Lei Bai 0001, Kaijian Liu, Yixiao Ge, Wanli Ouyang |
CVPR | 3 |
| 2021 | GLiT: Neural Architecture Search for Global and Local Image TransformerabstractWe introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and thus could be sub-optimal when directly used for image recognition. In order to improve the visual representation ability for transformers, we propose a new search space and searching algorithm. Specifically, we introduce a locality module that models the local correlations in images explicitly with fewer computational cost. With the locality module, our search space is defined to let the search algorithm freely trade off between global and local information as well as optimizing the low-level design choice in each module. To tackle the problem caused by huge search space, a hierarchical neural architecture search method is proposed to search the optimal vision transformer from two levels separately with the evolutionary algorithm. Extensive experiments on the ImageNet dataset demonstrate that our method can find more discriminative and efficient trans-former variants than the ResNet family (e.g., ResNet101) and the baseline ViT for image classification. The source codes are available at https://github.com/bychen515/GLiT. Peixia Li, Chuming Li, Baopu Li, Lei Bai 0001, Chen Lin 0003, Ming Sun 0008, Wanli Ouyang |
ICCV | 5 |
| 2021 | Graph-Based 3D Multi-Person Pose Estimation Using Multi-View ImagesabstractThis paper studies the task of estimating the 3D human poses of multiple persons from multiple calibrated camera views. Following the top-down paradigm, we decompose the task into two stages, i.e. person localization and pose estimation. Both stages are processed in coarse-to-fine manners. And we propose three task-specific graph neural networks for effective message passing. For 3D person localization, we first use Multi-view Matching Graph Module (MMG) to learn the cross-view association and recover coarse human proposals. The Center Refinement Graph Module (CRG) further refines the results via flexible point-based prediction. For 3D pose estimation, the Pose Regression Graph Module (PRG) learns both the multi-view geometry and structural relations between human joints. Our approach achieves state-of-the-art performance on CMU Panoptic and Shelf datasets with significantly lower computation complexity. Size Wu, Sheng Jin 0007, Wentao Liu 0002, Lei Bai 0001, Chen Qian 0006, Dong Liu 0002, Wanli Ouyang |
ICCV | 4 |
| 2021 | Deep spatial-temporal sequence modeling for multi-step passenger demand prediction
Lei Bai 0001, Lina Yao 0001, Xianzhi Wang 0001, Can Li 0014, Xiang Zhang 0012 |
Future Gener. Comput. Syst. | 1 |
| 2020 | Knowledge Adaption for Demand Prediction based on Multi-task Memory Neural NetworkabstractAccurate demand forecasting of different public transport modes (e.g., buses and light rails) is essential for public service operation. However, the development level of various modes often varies significantly, which makes it hard to predict the demand of the modes with insufficient knowledge and sparse station distribution (i.e., station-sparse mode). Intuitively, different public transit modes may exhibit shared demand patterns temporally and spatially in a city. As such, we propose to enhance the demand prediction of station-sparse modes with the data from station-intensive mode and design a Memory-Augmented Multi-task Re current Network (MATURE) to derive the transferable demand patterns from each mode and boost the prediction of station-sparse modes through adapting the relevant patterns from the station-intensive mode. Specifically, MATURE comprises three components: 1) a memory-augmented recurrent network for strengthening the ability to capture the long-short term information and storing temporal knowledge of each transit mode; 2) a knowledge adaption module to adapt the relevant knowledge from a station-intensive source to station-sparse sources; 3) a multi-task learning framework to incorporate all the information and forecast the demand of multiple modes jointly. The experimental results on a real-world dataset covering four public transport modes demonstrate that our model can promote the demand forecasting performance for the station-sparse modes. Can Li 0014, Lei Bai 0001, Wei Liu 0101, Lina Yao 0001, S. Travis Waller |
CIKM | 2 |
| 2020 | Are You A Risk Taker? Adversarial Learning of Asymmetric Cross-Domain Alignment for Risk Tolerance PredictionabstractMost current studies on survey analysis and risk tolerance modelling lack professional knowledge and domain-specific models. Given the effectiveness of generative adversarial learning in cross-domain information, we design an Asymmetric cross-Domain Generative Adversarial Network (ADGAN) for domain scale inequality. ADGAN utilizes the information-sufficient domain to provide extra information to improve the representation learning on the information-insufficient domain via domain alignment. We provide data analysis and user model on two data sources: Consumer Consumption Information and Survey Information. We further test ADGAN on a real-world dataset with view embedding structures and show ADGAN can better deal with the class imbalance and unqualified data space than state-of-the-art, demonstrating the effectiveness of leveraging asymmetrical domain information. Zhe Liu 0023, Lina Yao 0001, Xianzhi Wang 0001, Lei Bai 0001, Jake An |
IJCNN | 4 |
| 2020 | Spectrum-Guided Adversarial Disparity LearningabstractIt has been a significant challenge to portray intraclass disparity precisely in the area of activity recognition, as it requires a robust representation of the correlation between subject-specific variation for each activity class. In this work, we propose a novel end-to-end knowledge directed adversarial learning framework, which portrays the class-conditioned intraclass disparity using two competitive encoding distributions and learns the purified latent codes by denoising learned disparity. Furthermore, the domain knowledge is incorporated in an unsupervised manner to guide the optimization and further boosts the performance. The experiments on four HAR benchmark datasets demonstrate the robustness and generalization of our proposed methods over a set of state-of-the-art. We further prove the effectiveness of automatic domain knowledge incorporation in performance enhancement. Zhe Liu 0023, Lina Yao 0001, Lei Bai 0001, Xianzhi Wang 0001, Can Wang 0004 |
KDD | 3 |
| 2020 | Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingabstractModeling complex spatial and temporal correlations in the correlated time series data is indispensable for understanding the traffic dynamics and predicting the future status of an evolving traffic system. Recent works focus on designing complicated graph neural network architectures to capture shared patterns with the help of pre-defined graphs. In this paper, we argue that learning node-specific patterns is essential for traffic forecasting while pre-defined graph is avoidable. To this end, we propose two adaptive modules for enhancing Graph Convolutional Network (GCN) with new capabilities: 1) a Node Adaptive Parameter Learning (NAPL) module to capture node-specific patterns; 2) a Data Adaptive Graph Generation (DAGG) module to infer the inter-dependencies among different traffic series automatically. We further propose an Adaptive Graph Convolutional Recurrent Network (AGCRN) to capture fine-grained spatial and temporal correlations in traffic series automatically based on the two modules and recurrent networks. Our experiments on two real-world traffic datasets show AGCRN outperforms state-of-the-art by a significant margin without pre-defined graphs about spatial connections. Lei Bai 0001, Lina Yao 0001, Can Li 0014, Xianzhi Wang 0001, Can Wang 0004 |
NeurIPS | 1 |
| 2020 | Prototype Similarity Learning for Activity RecognitionabstractHuman Activity Recognition (HAR) plays an irreplaceable role in various applications such as security, gaming, and assisted living. Recent studies introduce deep learning to mitigate the manual feature extraction (i.e., data representation) efforts and achieve high accuracy. However, there are still challenges in learning accurate representations for sensory data due to the weakness of representation modules and the subject variances. We propose a scheme called Distance-based HAR from Ensembled spatial-temporal Representations (DHARER) to address above challenges. The idea behind DHARER is straightforward—the same activities should have similar representations. We first learn representations of the input sensory segments and latent prototype representations of each class, using a Convolution Neural Network (CNN)-based dual-stream representation module; then the learned representations are projected to activity types by measuring their similarity to the learned prototypes. We have conducted extensive experiments under a strict subject-independent setting on three large-scale datasets to evaluate the proposed scheme, and our experimental results demonstrate superior performance of DHARER to several state-of-the-art methods. Lei Bai 0001, Lina Yao 0001, Xianzhi Wang 0001, Salil S. Kanhere, Yang Xiao 0014 |
PAKDD (1) | 1 |
| 2020 | Mobility Irregularity Detection with Smart Transit Card Data
Xuesong Wang 0002, Lina Yao 0001, Wei Liu 0101, Can Li 0014, Lei Bai 0001, S. Travis Waller |
PAKDD (1) | 5 |
| 2020 | An enhanced probabilistic fairness-aware group recommendation by incorporating social activenessabstractCompared with individual recommendation, recommending services to a group of users is more complicated because of various users' preference should be considered and introduces new challenging such as fairness, which has never been well studied in current works. In this paper, we propose a novel recommendation scheme called PFGR, which combines a probabilistic model with coalition game strategy, to ensure the accuracy and fairness between groups of users. Given a group of users and a set of services, PFGR models a generative process for service selection in light of several observations: 1) each group is related with several topics; 2) users' decisions on the service selection depends on their expertise, the opinions of members they are familiar with, and group influence; 3) each group contains active users and inactive user, whose activeness contributes to the existence of group. PFGR first estimates the preference of each user on a candidate service via combining user's expertise, inherent connection, and group influence. Then, it determines a group's decision on a service by aggregating the preference of group members using adaptive weights. Finally, PFGR considers users' activeness and employs a strategy based on coalition game to produce a ranked list which is fair to each group member as much as possible. Experimental results on three real-world datasets validate that PFGR can achieve higher Hit Rate and Average Reciprocal Hit Rank than state-of-the-art approaches, which indicates that PFGR attains both the precision and fairness of recommendation. Yang Xiao 0014, Qingqi Pei, Lina Yao 0001, Shui Yu 0001, Lei Bai 0001, Xianzhi Wang 0001 |
J. Netw. Comput. Appl. | 5 |
| 2019 | Spatio-Temporal Graph Convolutional and Recurrent Networks for Citywide Passenger Demand PredictionabstractOnline ride-sharing platforms have become a critical part of the urban transportation system. Accurately recommending hotspots to drivers in such platforms is essential to help drivers find passengers and improve users' experience, which calls for efficient passenger demand prediction strategy. However, predicting multi-step passenger demand is challenging due to its high dynamicity, complex dependencies along spatial and temporal dimensions, and sensitivity to external factors (meteorological data and time meta). We propose an end-to-end deep learning framework to address the above problems. Our model comprises three components in pipeline: 1) a cascade graph convolutional recurrent neural network to accurately extract the spatial-temporal correlations within citywide historical passenger demand data; 2) two multi-layer LSTM networks to represent the external meteorological data and time meta, respectively; 3) an encoder-decoder module to fuse the above two parts and decode the representation to predict over multi-steps into the future. The experimental results on three real-world datasets demonstrate that our model can achieve accurate prediction and outperform the most discriminative state-of-the-art methods. Lei Bai 0001, Lina Yao 0001, Salil S. Kanhere, Xianzhi Wang 0001, Wei Liu 0101, Zheng Yang 0002 |
CIKM | 1 |
| 2019 | Passenger Demographic Attributes Prediction for Human-Centered Public Transport
Can Li 0014, Lei Bai 0001, Wei Liu 0101, Lina Yao 0001, S. Travis Waller |
ICONIP (4) | 2 |
| 2019 | STG2Seq: Spatial-Temporal Graph to Sequence Model for Multi-step Passenger Demand ForecastingabstractMulti-step passenger demand forecasting is a crucial task in on-demand vehicle sharing services. However, predicting passenger demand is generally challenging due to the nonlinear and dynamic spatial-temporal dependencies. In this work, we propose to model multi-step citywide passenger demand prediction based on a graph and use a hierarchical graph convolutional structure to capture both spatial and temporal correlations simultaneously. Our model consists of three parts: 1) a long-term encoder to encode historical passenger demands; 2) a short-term encoder to derive the next-step prediction for generating multi-step prediction; 3) an attention-based output module to model the dynamic temporal and channel-wise information. Experiments on three real-world datasets show that our model consistently outperforms many baseline methods and state-of-the-art models. Lei Bai 0001, Lina Yao 0001, Salil S. Kanhere, Xianzhi Wang 0001, Quan Z. Sheng |
IJCAI | 1 |
| 2019 | Passenger Demand Forecasting with Multi-Task Convolutional Recurrent Neural Networks
Lei Bai 0001, Lina Yao 0001, Salil S. Kanhere, Zheng Yang 0002, Jing Chu, Xianzhi Wang 0001 |
PAKDD (2) | 1 |
| 2018 | Automatic Device Classification from Network Traffic Streams of Internet of ThingsabstractWith the widespread adoption of Internet of Things (IoT), billions of everyday objects are being connected to the Internet. Effective management of these devices to support reliable, secure and high quality applications becomes challenging due to the scale. As one of the key cornerstones of IoT device management, automatic cross-device classification aims to identify the semantic type of a device by analyzing its network traffic. It has the potential to underpin a broad range of novel features such as enhanced security (by imposing the appropriate rules for constraining the communications of certain types of devices) or context-awareness (by the utilization and interoperability of IoT devices and their high-level semantics) of IoT applications. We propose an automatic IoT device classification method to identify new and unseen devices. The method uses the rich information carried by the traffic flows of IoT networks to characterize the attributes of various devices. We first specify a set of discriminating features from raw network traffic flows, and then propose a LSTM-CNN cascade model to automatically identify the semantic type of a device. Our experimental results using a real-world IoT dataset demonstrate that our proposed method is capable of delivering satisfactory performance. We also present interesting insights and discuss the potential extensions and applications. Lei Bai 0001, Lina Yao 0001, Salil S. Kanhere, Xianzhi Wang 0001, Zheng Yang 0002 |
LCN | 1 |