Wei Wang 0353

dblp:35/7092-353 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SSTODE: Ocean-Atmosphere Physics-Informed Neural ODEs for Sea Surface Temperature Prediction
abstract
Sea Surface Temperature (SST) is crucial for understanding upper-ocean thermal dynamics and ocean-atmosphere interactions, which have profound economic and social impacts. While data-driven models show promise in SST prediction, their black-box nature often limits interpretability and overlooks key physical processes. Recently, physics-informed neural networks have been gaining momentum but struggle with complex ocean-atmosphere dynamics due to 1) inadequate characterization of seawater movement (e.g., coastal upwelling) and 2) insufficient integration of external SST drivers (e.g., turbulent heat fluxes). To address these challenges, we propose SSTODE, a physics-informed Neural Ordinary Differential Equations (Neural ODEs) framework for SST prediction. First, we derive ODEs from fluid transport principles, incorporating both advection and diffusion to model ocean spatiotemporal dynamics. Through variational optimization, we recover a latent velocity field that explicitly governs the temporal dynamics of SST. Building upon ODE, we introduce an Energy Exchanges Integrator (EEI)-inspired by ocean heat budget equations-to account for external forcing factors. Thus, the variations in the components of these factors provide deeper insights into SST dynamics. Extensive experiments demonstrate that SSTODE achieves state-of-the-art performances in global and regional SST forecasting benchmarks. Furthermore, SSTODE visually reveals the impact of advection dynamics, thermal diffusion patterns, and diurnal heating-cooling cycles on SST evolution. These findings demonstrate the model's interpretability and physical consistency.
Wei Wang 0353, Yi Wang 0013
AAAI2
2026 M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference
abstract
Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.
Chuxiong Sun, Qirui Ji, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Wei Wang 0353
AAAI7
2026 ConstructAI: From Real-Time Safety Insight to Skill Growth in Deployed Construction AI Systems
abstract
Ensuring safety in power grid construction remains a critical yet challenging task, as existing monitoring approaches often lack scalability, timeliness, and adaptability to diverse on-site conditions. To address these limitations, we present ConstructAI, a deployed AI-driven safety management system that integrates multi-source image and video acquisition devices with advanced multimodal large model reasoning. The system combines text, image, and video prompts through an efficient workflow powered by LLaMA3 and Meta SAM2 backbones, enhanced with LoRA and adaptor modules for multimodal fusion. Once deployed, ConstructAI continuously processes real-time construction footage to identify violations, assess risk levels, and generate standardized rectification requirements. The deployment has demonstrated measurable benefits across multiple sites, including a >70% increase in violation rectification rates, reduction of average rectification delays from hours to minutes, and a 45% decline in repeat violations. Beyond technical gains, ConstructAI has delivered significant business impacts, such as reduced safety incidents, improved compliance with national regulations, and higher operational efficiency. By enabling proactive risk management and structured safety feedback loops, our system exemplifies how innovative use of AI can translate into tangible improvements for industrial safety. The lessons learned from deployment highlight the importance of balancing algorithmic advances with practical integration into organizational workflows.
Wei Wang 0353, Lee Kong Tiong, Yi Wang 0013
AAAI2
2025 Enhancing Vision-Language Models with Morphological and Taxonomic Knowledge: Towards Coral Recognition for Ocean Health
abstract
Coral reefs play a crucial role in marine ecosystems, offering a nutrient-rich environment and safe shelter for numerous marine species. Automated coral image recognition aids in monitoring ocean health at a scale without experts' manual effort. Recently, large vision-language models like CLIP have greatly enhanced zero-shot and low-shot classification capabilities for various visual tasks. However, these models struggle with fine-grained coral-related tasks due to a lack of specific knowledge. To bridge this gap, we compile a fine-grained coral image dataset consisting of 16,659 images with taxonomy labels (from Kingdom to Species), accompanied by morphology-specific text descriptions for each species. Based on the dataset, we propose CORAL-Adapter, integrating two complementary kinds of coral-specific knowledge (biological taxonomy and coral morphology) with general knowledge learned by CLIP. CORAL-Adapter is a simple yet powerful extension of CLIP with only a few parameter updates and can be used as a plug-and-play module with various CLIP-based methods. We show improvements in accuracy across diverse coral recognition tasks, e.g., recognizing corals unseen during training that are prone to bleaching or originate from different oceans.
Hongyong Han, Wei Wang 0353, Yi Wang 0013
AAAI2
2025 XCotton: Advancing AI-Enabled Hardware/Software Integrated System for Foreign Fiber Cleaning
abstract
Cotton is a critical agricultural product and industrial raw material, playing a key role in the national economies and people's living conditions, particularly in developing countries. However, cotton picking and processing often result in the contamination with various foreign fibers, such as hair, hemp rope, plastic film, and polypropylene rope. These contaminants are difficult to remove during textile processing and tend to break into small fragments, significantly reducing the quality of cotton products and negatively impacting the cotton industry. In this paper, we present an AI-enabled hardware-software integrated system--XCotton, for identifying and removing foreign fibers. Our system has been deployed in actual cotton production environments in the multiple regions in China, Central Asia, and Africa. XCotton achieves a cleaning efficiency of 1000kg/h, representing a 43% improvement, with only 14 kWh energy consumption (63% less). Moreover, XCotton brings significant business values to its manufacturer and clients. XCotton not only enhances the quality of cotton products but also contributes to the value-adding and upgrading of the cotton industry in developing regions, supporting economic growth and improving living conditions.
Wei Wang 0353, Yusheng Peng, Yi Wang 0013
AAAI2
2025 Hierarchical Cognitive Graph Autoencoder for Multi-Agent Reinforcement Learning
Wei Wang 0353, Chuxiong Sun, Yi Wang 0013
CogSci2
2025 FenGePad-A Tangible Multi-Prompt Interactive Framework for Deep Dune Segmentation
abstract
Interactive segmentation has become critical for efficiently delineating dune boundaries from remote sensing landform images, enabling geographers to iteratively refine model predictions through minimal user guidance. However, geographers report two major challenges when working with existing tools: (1) handling segmentation around ambiguous dune boundaries forces geographers into dense, repetitive clicking, making the interaction tedious and reducing annotation efficiency; (2) conventional desktop-based annotation platforms mainly support sequential, isolated interactions, hindering the smooth, co-located collaboration necessary for dealing with difficult cases. We thus propose FenGePad, a tangible collaborative interactive segmentation framework. It supports flexible prompt types–clicks, polylines, and scribbles–designed to accommodate geographers’ diverse annotation preferences and improve annotation efficiency. To enhance model robustness and generalization, we introduce prompt generation strategies that simulate realistic annotation behaviors of geographers during training. Finally, we instantiate a tablet-based application supporting FenGePad’s tangible annotation and collaboration. Comprehensive experiments demonstrate that FenGePad achieves competitive segmentation performance while effectively improving annotation quality and collaborative efficiency. Our results demonstrate the promise of tangible interactive frameworks for applying deep learning in geographic research.
Haonan Kang, Zifeng Wu, Wei Wang 0353, Eerdun Hasi, Yi Wang 0013
ECAI4
2025 PatchST: A Patch Spatial-Temporal Network for Large-Scale Traffic Forecasting
abstract
Traffic prediction plays a critical role in mitigating congestion, optimizing traffic flow, and enhancing urban mobility. Accurate predictions contribute directly to improved safety, reduced travel times, and a more efficient transportation system. For example, when a surge in traffic is anticipated, traffic signal timings can be adjusted in real-time to prevent gridlock and minimize accidents. However, the dynamic and complex nature of traffic patterns poses significant challenges for accurate forecasting. Recently, deep learning techniques like Spatial-Temporal Graph Neural Networks (STGNNs) have been employed to address these challenges. While these methods have shown promise, they often struggle with high computational complexity and difficulty in explicitly capturing contextual dependencies. In this paper, we introduce the Patch Spatial-Temporal Network (PatchST), an efficient network design tailored for time series forecasting. Our approach offers three primary contributions: 1) a quadratic reduction in computational and memory overhead for attention maps, particularly for large datasets; 2) enhanced forecasting accuracy through the incorporation of more extensive historical context; and 3) improved capability to capture both long-term and short-term spatial dependencies. Comprehensive experiments on well-established traffic forecasting datasets demonstrate that our model achieves state-of-the-art performance in both effectiveness and efficiency.
Jinrun Li, Wei Wang 0353, Yi Wang 0013
ICASSP3
2025 CSD: Weather forecasting with graph neural network based on cross-scale diffusivity
abstract
Automated weather stations play a pivotal role in fine-grained weather forecasting, due to their cost-effectiveness and global deployment potential. Data-driven methods, particularly deep learning techniques, have emerged as potent tools for precise weather forecasting. However, capturing meaningful relationships among these decentralized stations, especially on global scale, presents a formidable challenge. Many transformer-based approaches, commonly used in data-driven forecasting, rely on the conventional point-wise or the series-wise attention mechanism to construct these relationships between weather stations. Unfortunately, this mechanism either brings about high computational complexity or falls short in explicitly capturing contextual dependencies, as it operates on keys and values derived from the same series of data. In this paper, we propose a novel approach that can adaptively establish cross-scale diffusivity from an energy-constrained diffusion perspective. Additionally, our model can adaptively learn two types of graphs: the static spatial graph and the dynamic temporal graph. Experimental results demonstrate that our method can achieve state-of-the-art performance.
Jinrun Li, Wei Wang 0353, Yi Wang 0013
ICASSP3
2025 TIDE-Net: A Physics-Based Graph Model for Predicting Tropical Cyclone Impacts on Estuarine Systems
abstract
Estuarine systems, located at the interface of land, ocean, and atmosphere, are vital to global ecosystems and economies due to their rich exchanges among multiple environments. Tropical cyclones cause significant fluctuations in salinity and other environmental parameters within estuaries, impacting their resilience. Accurately predicting these changes aids in decision-making to protect these ecosystems. While graph-based deep learning models have shown promise, they often fail to capture the complex interdependencies between observation stations. To address this, we propose an energy-constrained diffusion feature representation that synergizes with ocean dynamics. Using data from 24 observation points in the Yangtze River Estuary during three tropical cyclones, our method effectively predicts and interprets dynamic environmental changes, offering robust support for protecting estuarine systems.
Wei Wang 0353, Yi Wang 0013
ICASSP2
2025 ZeroPose: Leveraging Diffusion Models and Large Language Models for Advanced Multi-Hypothesis 3D Construction Workers' Pose Estimation
abstract
ZeroPose is a zero-shot, conditional diffusion-based model for 3D pose estimation from monocular images, addressing challenges in dynamic, crowded construction sites. Unlike traditional methods, it generates diverse 3D poses, handling ambiguities and joint occlusions without the need for extensive pre-labeled datasets. By integrating Large Language Models (LLMs) like LLama, ZeroPose aligns pose predictions with common sense, improving safety by identifying potential hazards and unsafe behaviors. Experimental results show that ZeroPose outperforms existing methods, offering flexibility for both monocular and multi-camera environments, and contributes to safer, healthier workplaces.
Wei Wang 0353, Yi Wang 0013
ICME2
2025 CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
abstract
Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the need for domain expertise. Visual Question Answering (VQA), powered by Large Vision-Language Models (LVLMs), has great potential in user-friendly interaction with coral reef images. However, applying VQA to coral imagery demands a dedicated dataset that addresses two key challenges: domain-specific annotations and multidimensional questions. In this work, we introduce CoralVQA, the first large-scale VQA dataset for coral reef analysis. It contains 12,805 real-world coral images from 67 coral genera collected from 3 oceans, along with 277,653 question-answer pairs that comprehensively assess ecological and health-related conditions. To construct this dataset, we develop a semi-automatic data construction pipeline in collaboration with marine biologists to ensure both scalability and professional-grade data quality. CoralVQA presents novel challenges and provides a comprehensive benchmark for studying vision-language reasoning in the context of coral reef images. By evaluating several state-of-the-art LVLMs, we reveal key limitations and opportunities. These insights form a foundation for future LVLM development, with a particular emphasis on supporting coral conservation efforts.
Hongyong Han, Wei Wang 0353, Yi Wang 0013
NeurIPS2
2025 Uncovering Non-native Speakers' Experiences in Global Software Development Teams - - a Bourdieusian Perspective
Yi Wang 0013, Yang Yue 0003, Wei Wang 0353
Comput. Support. Cooperative Work.3
2024 DCV2I: A Practical Approach for Supporting Geographers' Visual Interpretation in Dune Segmentation with Deep Vision Models
abstract
Visual interpretation is extremely important in human geography as the primary technique for geographers to use photograph data in identifying, classifying, and quantifying geographic and topological objects or regions. However, it is also time-consuming and requires overwhelming manual effort from professional geographers. This paper describes our interdisciplinary team's efforts in integrating computer vision models with geographers' visual image interpretation process to reduce their workload in interpreting images. Focusing on the dune segmentation task, we proposed an approach featuring a deep dune segmentation model to identify dunes and label their ranges in an automated way. By developing a tool to connect our model with ArcGIS, one of the most popular workbenches for visual interpretation, geographers can further refine the automatically-generated dune segmentation on images without learning any CV or deep learning techniques. Our approach thus realized a non-invasive change to geographers' visual interpretation routines, reducing their manual efforts while incurring minimal interruptions to their work routines and tools they are familiar with. Deployment with a leading Chinese geography research institution demonstrated the potential of our approach in supporting geographers in researching and solving drylands desertification.
Anqi Lu, Zifeng Wu, Wei Wang 0353, Eerdun Hasi, Yi Wang 0013
AAAI4
2024 The Dynamics of Cooperation with Commitment in A Population of Heterogeneous Preferences-An ABM Study
Wei Wang 0353, Luzhan Yuan, Yi Wang 0013
CogSci1
2024 Characterizing Developers' Behaviors in LLM -Supported Software Development
abstract
The emergence of large language models (LLMs) represented by ChatGPT has profoundly influenced the conventional software development process. Nevertheless, there is little research on how developers interact with LLMs. To investigate the interaction between developers and LLMs during software development, we conducted a user study with 56 participants, who were randomly assigned into two groups to perform different types of software development tasks. The first task was solving two simple coding puzzles, and the second was to fix two real-world bugs from a small-scale open source projects. We captured the full screen histories of all participants to construct a Markov activity transition model, depicting transitions among prevalent development activities such as coding, writing prompt and debugging. By characterizing developer behavior of the two groups of participants in the task, our study contributes empirical insights into developer behavior while interacting with LLMs. Such knowledge could guide the development of LLM -specific capabilities in supporting software development tasks, and offer valuable insights for developers aiming to utilize LLMs for more efficient problem-solving in their development practices.
Wei Wang 0353, Huilong Ning, Shuo Qian, Yi Wang 0013
COMPSAC1
2024 CR-Cross: Cross Domain Coral Recognitions with Reject Options For Coral Conservation
abstract
Although coral reefs are special and vital marine ecosystems, massive coral degradation began to occur due to the increase in global temperatures and the intensification of human industrial activities. Coral reef protection requires accurate coral recognition because it is the foundation for learning the distribution, disease, and growth of coral reefs, hereby informing the proper ways for further action. Recently, CNNs have been applied in automated coral image classification. These classifier models, however, are difficult to be generalized from the trained coral images in a marine region (source domain) to the coral images in a different marine region (target domain) since the corals have significant within-species morphological variability among the different geographic location domains. In this paper, a novel coral recognition algorithm is introduced via knowledge transfer across domains and its advantages lie in the following aspects. (1) It simultaneously transfers corals’ texture and structure features across domains thus providing useful knowledge to assist the coral recognition tasks in the target marine domain. (2) To overcome the difficulty that the confusing coral images (e.g., bleached corals) are prone to be misclassified and transfer useless or even negative information, our algorithm is equipped with the reject option for the confusing corals while adapting. These corals can be sent to an expert or a more expensive but accurate system, resulting in strengthened transferability and reliability. Furthermore, we develop a new cross-domain coral image dataset to enhance coral research. Without the label information from the target marine region, our method significantly reduces the distribution gap and domain shift among the different marine regions. In addition, CR-Cross goes a step further in tackling the challenges of missing coral data, maximizing the utilization of available coral datasets, and enhancing the reusability of both coral data and coral recognition models. A series of empirical studies show that our method remarkably outperforms a broad range of baselines.
Hongyong Han, Wei Wang 0353, Yi Wang 0013
COMPASS2
2024 FenGe-An Interactive Framework for Improving the Utility of Deep Dune Segmentation in Geographical Tasks
abstract
Segmenting dunes from remote sensing landforms images with deep vision models is promising by freeing geographers from manual visual interpretation tasks, making them more concentrated on the essential tasks in solving desertification challenges. However, geographers have reported that automated segmentation results may be not satisfactory though achieving high accuracy, implying there are potential gaps between pixel-level metrics and the utility in downstream geographic tasks. Therefore, pixel-wise metrics may be not proper in evaluating the deep dune segmentation in the geography domain, arising the necessity to develop domain-specific, human-centered measurements for deep dune segmentation. This paper first proposes a novel measurement based on geographers’ subjective judgments, which allows the evaluation of the alignment between deep dune segmentation models and geographical utility. We design an interactive framework integrating multiagent reinforcement learning (MARL) with geographers’ domain knowledge to improve models’ utility in the domain of geography. Our extensive experiments show that (1) our framework enables the interactive domain knowledge integration in the model-building process, and thus (2) the dune segmentation model better aligns with geographical utility, which ultimately improves the effectiveness of dune segmentation. We have deployed the framework with a number of geographers to support their various tasks including dune segmentation as a component. The results demonstrate our framework’s capabilities.
Anqi Lu, Zifeng Wu, Wei Wang 0353, Eerdun Hasi, Yi Wang 0013
ECAI4
2024 MA-ST3D: Motion Associated Self-Training for Unsupervised Domain Adaptation on 3D Object Detection
abstract
Recently, unsupervised domain adaptation (UDA) for 3D object detectors has increasingly garnered attention as a method to eliminate the prohibitive costs associated with generating extensive 3D annotations, which are crucial for effective model training. Self-training (ST) has emerged as a simple and effective technique for UDA. The major issue involved in ST-UDA for 3D object detection is refining the imprecise predictions caused by domain shift and generating accurate pseudo labels as supervisory signals. This study presents a novel ST-UDA framework to generate high-quality pseudo labels by associating predictions of 3D point cloud sequences during ego-motion according to spatial and temporal consistency, named motion-associated self-training for 3D object detection (MA-ST3D). MA-ST3D maintains a global-local pathway (GLP) architecture to generate high-quality pseudo-labels by leveraging both intra-frame and inter-frame consistencies along the spatial dimension of the LiDAR's ego-motion. It also equips two memory modules for both global and local pathways, called global memory and local memory, to suppress the temporal fluctuation of pseudo-labels during self-training iterations. In addition, a motion-aware loss is introduced to impose discriminated regulations on pseudo labels with different motion statuses, which mitigates the harmful spread of false positive pseudo labels. Finally, our method is evaluated on three representative domain adaptation tasks on authoritative 3D benchmark datasets (i.e. Waymo, Kitti, and nuScenes). MA-ST3D achieved SOTA performance on all evaluated UDA settings and even surpassed the weakly supervised DA methods on the Kitti and NuScenes object detection benchmark.
Chi Zhang 0060, Wei Wang 0353, Zhaoxiang Zhang 0001
IEEE Trans. Image Process.3
2023 OPRADI: Applying Security Game to Fight Drive under the Influence in Real-World
abstract
Driving under the influence (DUI) is one of the main causes of traffic accidents, often leading to severe life and property losses. Setting up sobriety checkpoints on certain roads is the most commonly used practice to identify DUI-drivers in many countries worldwide. However, setting up checkpoints according to the police's experiences may not be effective for ignoring the strategic interactions between the police and DUI-drivers, particularly when inspecting resources are limited. To remedy this situation, we adapt the classic Stackelberg security game (SSG) to a new SSG-DUI game to describe the strategic interactions in catching DUI-drivers. SSG-DUI features drivers' bounded rationality and social knowledge sharing among them, thus realizing improved real-world fidelity. With SSG-DUI, we propose OPRADI, a systematic approach for advising better strategies in setting up checkpoints. We perform extensive experiments to evaluate it in both simulated environments and real-world contexts, in collaborating with a Chinese city's police bureau. The results reveal its effectiveness in improving police's real-world operations, thus having significant practical potentials.
Luzhan Yuan, Wei Wang 0353, Yi Wang 0013
AAAI2
2023 Correntropy-Induced Wasserstein GCN: Learning Graph Embedding via Domain Adaptation
abstract
Graph embedding aims at learning vertex representations in a low-dimensional space by distilling information from a complex-structured graph. Recent efforts in graph embedding have been devoted to generalizing the representations from the trained graph in a source domain to the new graph in a different target domain based on information transfer. However, when the graphs are contaminated by unpredictable and complex noise in practice, this transfer problem is quite challenging because of the need to extract helpful knowledge from the source graph and to reliably transfer knowledge to the target graph. This paper puts forward a two-step correntropy-induced Wasserstein GCN (graph convolutional network, or CW-GCN for short) architecture to facilitate the robustness in cross-graph embedding. In the first step, CW-GCN originally investigates correntropy-induced loss in GCN, which places bounded and smooth losses on the noisy nodes with incorrect edges or attributes. Consequently, helpful information are extracted only from clean nodes in the source graph. In the second step, a novel Wasserstein distance is introduced to measure the difference in marginal distributions between graphs, avoiding the negative influence of noise. Afterwards, CW-GCN maps the target graph to the same embedding space as the source graph by minimizing the Wasserstein distance, and thus the knowledge preserved in the first step is expected to be reliably transferred to assist the target graph analysis tasks. Extensive experiments demonstrate the significant superiority of CW-GCN over state-of-the-art methods in different noisy environments.
Wei Wang 0353, Hongyong Han, Chi Zhang 0060
IEEE Trans. Image Process.1
2022 Achieving Consensus to Learn an Efficient and Robust Communication via Reinforcement Learning
Wei Qing, Zhaofeng He 0001, Junge Zhang, Luzhan Yuan, Wei Wang 0353
CogSci6
2021 Learning Robust Feature Transformation for Domain Adaptation
Wei Wang 0353, Hao Wang 0005, Zhi-Yong Ran
Pattern Recognit.1