Yi Wang 0013

dblp:67/6649-13 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0003-1321-4035ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 SSTODE: Ocean-Atmosphere Physics-Informed Neural ODEs for Sea Surface Temperature Prediction
abstract
Sea Surface Temperature (SST) is crucial for understanding upper-ocean thermal dynamics and ocean-atmosphere interactions, which have profound economic and social impacts. While data-driven models show promise in SST prediction, their black-box nature often limits interpretability and overlooks key physical processes. Recently, physics-informed neural networks have been gaining momentum but struggle with complex ocean-atmosphere dynamics due to 1) inadequate characterization of seawater movement (e.g., coastal upwelling) and 2) insufficient integration of external SST drivers (e.g., turbulent heat fluxes). To address these challenges, we propose SSTODE, a physics-informed Neural Ordinary Differential Equations (Neural ODEs) framework for SST prediction. First, we derive ODEs from fluid transport principles, incorporating both advection and diffusion to model ocean spatiotemporal dynamics. Through variational optimization, we recover a latent velocity field that explicitly governs the temporal dynamics of SST. Building upon ODE, we introduce an Energy Exchanges Integrator (EEI)-inspired by ocean heat budget equations-to account for external forcing factors. Thus, the variations in the components of these factors provide deeper insights into SST dynamics. Extensive experiments demonstrate that SSTODE achieves state-of-the-art performances in global and regional SST forecasting benchmarks. Furthermore, SSTODE visually reveals the impact of advection dynamics, thermal diffusion patterns, and diurnal heating-cooling cycles on SST evolution. These findings demonstrate the model's interpretability and physical consistency.
Wei Wang 0353, Yi Wang 0013
AAAI4
2026 ConstructAI: From Real-Time Safety Insight to Skill Growth in Deployed Construction AI Systems
abstract
Ensuring safety in power grid construction remains a critical yet challenging task, as existing monitoring approaches often lack scalability, timeliness, and adaptability to diverse on-site conditions. To address these limitations, we present ConstructAI, a deployed AI-driven safety management system that integrates multi-source image and video acquisition devices with advanced multimodal large model reasoning. The system combines text, image, and video prompts through an efficient workflow powered by LLaMA3 and Meta SAM2 backbones, enhanced with LoRA and adaptor modules for multimodal fusion. Once deployed, ConstructAI continuously processes real-time construction footage to identify violations, assess risk levels, and generate standardized rectification requirements. The deployment has demonstrated measurable benefits across multiple sites, including a >70% increase in violation rectification rates, reduction of average rectification delays from hours to minutes, and a 45% decline in repeat violations. Beyond technical gains, ConstructAI has delivered significant business impacts, such as reduced safety incidents, improved compliance with national regulations, and higher operational efficiency. By enabling proactive risk management and structured safety feedback loops, our system exemplifies how innovative use of AI can translate into tangible improvements for industrial safety. The lessons learned from deployment highlight the importance of balancing algorithmic advances with practical integration into organizational workflows.
Wei Wang 0353, Lee Kong Tiong, Yi Wang 0013
AAAI6
2025 Enhancing Vision-Language Models with Morphological and Taxonomic Knowledge: Towards Coral Recognition for Ocean Health
abstract
Coral reefs play a crucial role in marine ecosystems, offering a nutrient-rich environment and safe shelter for numerous marine species. Automated coral image recognition aids in monitoring ocean health at a scale without experts' manual effort. Recently, large vision-language models like CLIP have greatly enhanced zero-shot and low-shot classification capabilities for various visual tasks. However, these models struggle with fine-grained coral-related tasks due to a lack of specific knowledge. To bridge this gap, we compile a fine-grained coral image dataset consisting of 16,659 images with taxonomy labels (from Kingdom to Species), accompanied by morphology-specific text descriptions for each species. Based on the dataset, we propose CORAL-Adapter, integrating two complementary kinds of coral-specific knowledge (biological taxonomy and coral morphology) with general knowledge learned by CLIP. CORAL-Adapter is a simple yet powerful extension of CLIP with only a few parameter updates and can be used as a plug-and-play module with various CLIP-based methods. We show improvements in accuracy across diverse coral recognition tasks, e.g., recognizing corals unseen during training that are prone to bleaching or originate from different oceans.
Hongyong Han, Wei Wang 0353, Yi Wang 0013
AAAI5
2025 XCotton: Advancing AI-Enabled Hardware/Software Integrated System for Foreign Fiber Cleaning
abstract
Cotton is a critical agricultural product and industrial raw material, playing a key role in the national economies and people's living conditions, particularly in developing countries. However, cotton picking and processing often result in the contamination with various foreign fibers, such as hair, hemp rope, plastic film, and polypropylene rope. These contaminants are difficult to remove during textile processing and tend to break into small fragments, significantly reducing the quality of cotton products and negatively impacting the cotton industry. In this paper, we present an AI-enabled hardware-software integrated system--XCotton, for identifying and removing foreign fibers. Our system has been deployed in actual cotton production environments in the multiple regions in China, Central Asia, and Africa. XCotton achieves a cleaning efficiency of 1000kg/h, representing a 43% improvement, with only 14 kWh energy consumption (63% less). Moreover, XCotton brings significant business values to its manufacturer and clients. XCotton not only enhances the quality of cotton products but also contributes to the value-adding and upgrading of the cotton industry in developing regions, supporting economic growth and improving living conditions.
Wei Wang 0353, Yusheng Peng, Yi Wang 0013
AAAI4
2025 Hierarchical Cognitive Graph Autoencoder for Multi-Agent Reinforcement Learning
Wei Wang 0353, Chuxiong Sun, Yi Wang 0013
CogSci4
2025 FenGePad-A Tangible Multi-Prompt Interactive Framework for Deep Dune Segmentation
abstract
Interactive segmentation has become critical for efficiently delineating dune boundaries from remote sensing landform images, enabling geographers to iteratively refine model predictions through minimal user guidance. However, geographers report two major challenges when working with existing tools: (1) handling segmentation around ambiguous dune boundaries forces geographers into dense, repetitive clicking, making the interaction tedious and reducing annotation efficiency; (2) conventional desktop-based annotation platforms mainly support sequential, isolated interactions, hindering the smooth, co-located collaboration necessary for dealing with difficult cases. We thus propose FenGePad, a tangible collaborative interactive segmentation framework. It supports flexible prompt types–clicks, polylines, and scribbles–designed to accommodate geographers’ diverse annotation preferences and improve annotation efficiency. To enhance model robustness and generalization, we introduce prompt generation strategies that simulate realistic annotation behaviors of geographers during training. Finally, we instantiate a tablet-based application supporting FenGePad’s tangible annotation and collaboration. Comprehensive experiments demonstrate that FenGePad achieves competitive segmentation performance while effectively improving annotation quality and collaborative efficiency. Our results demonstrate the promise of tangible interactive frameworks for applying deep learning in geographic research.
Haonan Kang, Zifeng Wu, Wei Wang 0353, Eerdun Hasi, Yi Wang 0013
ECAI8
2025 PatchST: A Patch Spatial-Temporal Network for Large-Scale Traffic Forecasting
abstract
Traffic prediction plays a critical role in mitigating congestion, optimizing traffic flow, and enhancing urban mobility. Accurate predictions contribute directly to improved safety, reduced travel times, and a more efficient transportation system. For example, when a surge in traffic is anticipated, traffic signal timings can be adjusted in real-time to prevent gridlock and minimize accidents. However, the dynamic and complex nature of traffic patterns poses significant challenges for accurate forecasting. Recently, deep learning techniques like Spatial-Temporal Graph Neural Networks (STGNNs) have been employed to address these challenges. While these methods have shown promise, they often struggle with high computational complexity and difficulty in explicitly capturing contextual dependencies. In this paper, we introduce the Patch Spatial-Temporal Network (PatchST), an efficient network design tailored for time series forecasting. Our approach offers three primary contributions: 1) a quadratic reduction in computational and memory overhead for attention maps, particularly for large datasets; 2) enhanced forecasting accuracy through the incorporation of more extensive historical context; and 3) improved capability to capture both long-term and short-term spatial dependencies. Comprehensive experiments on well-established traffic forecasting datasets demonstrate that our model achieves state-of-the-art performance in both effectiveness and efficiency.
Jinrun Li, Wei Wang 0353, Yi Wang 0013
ICASSP4
2025 CSD: Weather forecasting with graph neural network based on cross-scale diffusivity
abstract
Automated weather stations play a pivotal role in fine-grained weather forecasting, due to their cost-effectiveness and global deployment potential. Data-driven methods, particularly deep learning techniques, have emerged as potent tools for precise weather forecasting. However, capturing meaningful relationships among these decentralized stations, especially on global scale, presents a formidable challenge. Many transformer-based approaches, commonly used in data-driven forecasting, rely on the conventional point-wise or the series-wise attention mechanism to construct these relationships between weather stations. Unfortunately, this mechanism either brings about high computational complexity or falls short in explicitly capturing contextual dependencies, as it operates on keys and values derived from the same series of data. In this paper, we propose a novel approach that can adaptively establish cross-scale diffusivity from an energy-constrained diffusion perspective. Additionally, our model can adaptively learn two types of graphs: the static spatial graph and the dynamic temporal graph. Experimental results demonstrate that our method can achieve state-of-the-art performance.
Jinrun Li, Wei Wang 0353, Yi Wang 0013
ICASSP4
2025 TIDE-Net: A Physics-Based Graph Model for Predicting Tropical Cyclone Impacts on Estuarine Systems
abstract
Estuarine systems, located at the interface of land, ocean, and atmosphere, are vital to global ecosystems and economies due to their rich exchanges among multiple environments. Tropical cyclones cause significant fluctuations in salinity and other environmental parameters within estuaries, impacting their resilience. Accurately predicting these changes aids in decision-making to protect these ecosystems. While graph-based deep learning models have shown promise, they often fail to capture the complex interdependencies between observation stations. To address this, we propose an energy-constrained diffusion feature representation that synergizes with ocean dynamics. Using data from 24 observation points in the Yangtze River Estuary during three tropical cyclones, our method effectively predicts and interprets dynamic environmental changes, offering robust support for protecting estuarine systems.
Wei Wang 0353, Yi Wang 0013
ICASSP3
2025 ZeroPose: Leveraging Diffusion Models and Large Language Models for Advanced Multi-Hypothesis 3D Construction Workers' Pose Estimation
abstract
ZeroPose is a zero-shot, conditional diffusion-based model for 3D pose estimation from monocular images, addressing challenges in dynamic, crowded construction sites. Unlike traditional methods, it generates diverse 3D poses, handling ambiguities and joint occlusions without the need for extensive pre-labeled datasets. By integrating Large Language Models (LLMs) like LLama, ZeroPose aligns pose predictions with common sense, improving safety by identifying potential hazards and unsafe behaviors. Experimental results show that ZeroPose outperforms existing methods, offering flexibility for both monocular and multi-camera environments, and contributes to safer, healthier workplaces.
Wei Wang 0353, Yi Wang 0013
ICME3
2025 CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
abstract
Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the need for domain expertise. Visual Question Answering (VQA), powered by Large Vision-Language Models (LVLMs), has great potential in user-friendly interaction with coral reef images. However, applying VQA to coral imagery demands a dedicated dataset that addresses two key challenges: domain-specific annotations and multidimensional questions. In this work, we introduce CoralVQA, the first large-scale VQA dataset for coral reef analysis. It contains 12,805 real-world coral images from 67 coral genera collected from 3 oceans, along with 277,653 question-answer pairs that comprehensively assess ecological and health-related conditions. To construct this dataset, we develop a semi-automatic data construction pipeline in collaboration with marine biologists to ensure both scalability and professional-grade data quality. CoralVQA presents novel challenges and provides a comprehensive benchmark for studying vision-language reasoning in the context of coral reef images. By evaluating several state-of-the-art LVLMs, we reveal key limitations and opportunities. These insights form a foundation for future LVLM development, with a particular emphasis on supporting coral conservation efforts.
Hongyong Han, Wei Wang 0353, Yi Wang 0013
NeurIPS5
2025 Uncovering Non-native Speakers' Experiences in Global Software Development Teams - - a Bourdieusian Perspective
Yi Wang 0013, Yang Yue 0003, Wei Wang 0353
Comput. Support. Cooperative Work.1
2025 Making Software Development More Diverse and Inclusive: Key Themes, Challenges, and Future Directions
abstract
Introduction : Digital products increasingly reshape industries, influencing human behavior and decision-making. However, the software development teams developing these systems often lack diversity, which may lead to designs that overlook the needs, equal treatment or safety of diverse user groups. These risks highlight the need for fostering diversity and inclusion in software development to create safer, more equitable technology. Method : This research is based on insights from an academic meeting in June 2023 involving 23 software engineering researchers and practitioners. We used the collaborative discussion method 1-2-4-ALL as a systematic research approach and identified six themes around the theme “challenges and opportunities to improve Software Developer Diversity and Inclusion (SDDI).” We identified benefits, harms, and future research directions for the four main themes. Then, we discuss the remaining two themes, AI & SDDI and AI & Computer Science education, which have a cross-cutting effect on the other themes. Results : This research explores the key challenges and research opportunities for promoting SDDI, providing a roadmap to guide both researchers and practitioners. We underline that research around SDDI requires a constant focus on maximizing benefits while minimizing harms, especially to vulnerable groups. As a research community, we must strike this balance in a responsible way.
Sonja Hyrynsalmi, Sebastian Baltes, Chris Brown 0001, Rafael Prikladnicki, Gema Rodríguez-Pérez, Alexander Serebrenik, Jocelyn Simmonds, Bianca Trinkenreich, Yi Wang 0013, Grischa Liebel
ACM Trans. Softw. Eng. Methodol.9
2024 DCV2I: A Practical Approach for Supporting Geographers' Visual Interpretation in Dune Segmentation with Deep Vision Models
abstract
Visual interpretation is extremely important in human geography as the primary technique for geographers to use photograph data in identifying, classifying, and quantifying geographic and topological objects or regions. However, it is also time-consuming and requires overwhelming manual effort from professional geographers. This paper describes our interdisciplinary team's efforts in integrating computer vision models with geographers' visual image interpretation process to reduce their workload in interpreting images. Focusing on the dune segmentation task, we proposed an approach featuring a deep dune segmentation model to identify dunes and label their ranges in an automated way. By developing a tool to connect our model with ArcGIS, one of the most popular workbenches for visual interpretation, geographers can further refine the automatically-generated dune segmentation on images without learning any CV or deep learning techniques. Our approach thus realized a non-invasive change to geographers' visual interpretation routines, reducing their manual efforts while incurring minimal interruptions to their work routines and tools they are familiar with. Deployment with a leading Chinese geography research institution demonstrated the potential of our approach in supporting geographers in researching and solving drylands desertification.
Anqi Lu, Zifeng Wu, Wei Wang 0353, Eerdun Hasi, Yi Wang 0013
AAAI6
2024 The Dynamics of Cooperation with Commitment in A Population of Heterogeneous Preferences-An ABM Study
Wei Wang 0353, Luzhan Yuan, Yi Wang 0013
CogSci5
2024 Characterizing Developers' Behaviors in LLM -Supported Software Development
abstract
The emergence of large language models (LLMs) represented by ChatGPT has profoundly influenced the conventional software development process. Nevertheless, there is little research on how developers interact with LLMs. To investigate the interaction between developers and LLMs during software development, we conducted a user study with 56 participants, who were randomly assigned into two groups to perform different types of software development tasks. The first task was solving two simple coding puzzles, and the second was to fix two real-world bugs from a small-scale open source projects. We captured the full screen histories of all participants to construct a Markov activity transition model, depicting transitions among prevalent development activities such as coding, writing prompt and debugging. By characterizing developer behavior of the two groups of participants in the task, our study contributes empirical insights into developer behavior while interacting with LLMs. Such knowledge could guide the development of LLM -specific capabilities in supporting software development tasks, and offer valuable insights for developers aiming to utilize LLMs for more efficient problem-solving in their development practices.
Wei Wang 0353, Huilong Ning, Shuo Qian, Yi Wang 0013
COMPSAC5
2024 CR-Cross: Cross Domain Coral Recognitions with Reject Options For Coral Conservation
abstract
Although coral reefs are special and vital marine ecosystems, massive coral degradation began to occur due to the increase in global temperatures and the intensification of human industrial activities. Coral reef protection requires accurate coral recognition because it is the foundation for learning the distribution, disease, and growth of coral reefs, hereby informing the proper ways for further action. Recently, CNNs have been applied in automated coral image classification. These classifier models, however, are difficult to be generalized from the trained coral images in a marine region (source domain) to the coral images in a different marine region (target domain) since the corals have significant within-species morphological variability among the different geographic location domains. In this paper, a novel coral recognition algorithm is introduced via knowledge transfer across domains and its advantages lie in the following aspects. (1) It simultaneously transfers corals’ texture and structure features across domains thus providing useful knowledge to assist the coral recognition tasks in the target marine domain. (2) To overcome the difficulty that the confusing coral images (e.g., bleached corals) are prone to be misclassified and transfer useless or even negative information, our algorithm is equipped with the reject option for the confusing corals while adapting. These corals can be sent to an expert or a more expensive but accurate system, resulting in strengthened transferability and reliability. Furthermore, we develop a new cross-domain coral image dataset to enhance coral research. Without the label information from the target marine region, our method significantly reduces the distribution gap and domain shift among the different marine regions. In addition, CR-Cross goes a step further in tackling the challenges of missing coral data, maximizing the utilization of available coral datasets, and enhancing the reusability of both coral data and coral recognition models. A series of empirical studies show that our method remarkably outperforms a broad range of baselines.
Hongyong Han, Wei Wang 0353, Yi Wang 0013
COMPASS5
2024 FenGe-An Interactive Framework for Improving the Utility of Deep Dune Segmentation in Geographical Tasks
abstract
Segmenting dunes from remote sensing landforms images with deep vision models is promising by freeing geographers from manual visual interpretation tasks, making them more concentrated on the essential tasks in solving desertification challenges. However, geographers have reported that automated segmentation results may be not satisfactory though achieving high accuracy, implying there are potential gaps between pixel-level metrics and the utility in downstream geographic tasks. Therefore, pixel-wise metrics may be not proper in evaluating the deep dune segmentation in the geography domain, arising the necessity to develop domain-specific, human-centered measurements for deep dune segmentation. This paper first proposes a novel measurement based on geographers’ subjective judgments, which allows the evaluation of the alignment between deep dune segmentation models and geographical utility. We design an interactive framework integrating multiagent reinforcement learning (MARL) with geographers’ domain knowledge to improve models’ utility in the domain of geography. Our extensive experiments show that (1) our framework enables the interactive domain knowledge integration in the model-building process, and thus (2) the dune segmentation model better aligns with geographical utility, which ultimately improves the effectiveness of dune segmentation. We have deployed the framework with a number of geographers to support their various tasks including dune segmentation as a component. The results demonstrate our framework’s capabilities.
Anqi Lu, Zifeng Wu, Wei Wang 0353, Eerdun Hasi, Yi Wang 0013
ECAI7
2024 Android Malware Family Labeling: Perspectives from the Industry
abstract
Labeling and classifying Android malware is important for identifying new threats, triaging security incidents, and demystifying evasion techniques. To automate the malware classification pipeline, state-of-the-art tools such as AVClass and Euphony unify raw labels from commercial antivirus vendors (i.e., VirusTotal) to produce family labels. These tools are widely used for automatic malware classification in both academic research and industry practice. However, they face significant limitations in real-world industrial scenarios with numerous and dynamically changing samples. For example, our industrial practices revealed that VirusTotal's results change over time, leading to temporal inconsistencies in family labeling results that rely on label unification, which can severely impact a company's security posture. Despite this, such issues and challenges remain understudied. In this paper, we present the first systematic measurement study of existing automatic Android malware family labeling systems from various aspects, including label dynamics, consistency, reliability, and etc. Based on a large-scale dataset, we validate that the labeling results of these systems do evolve with time, and such evolution can introduce bias into many previous studies on performance assessments. We also reveal substantial divergence in labeling decisions across different systems when given the same input. Besides, we identify a disclosure priority among families in these systems' labeling processes, which could threaten the industry by allowing malicious actors to exploit these discrepancies. Our findings could benefit both researchers and industry practitioners for further refinement of automatic malware family labeling systems, contributing to their practical applications.
Liu Wang 0002, Haoyu Wang 0001, Tao Zhang 0001, Haitao Xu 0002, Guozhu Meng, Peiming Gao, Yi Wang 0013
ASE8
2024 IdeoRate: Towards a Semi-automated Assessment Methodology for OSS Ideologies
abstract
Open source software (OSS) development, as any social movement, is driven by its ideologies, namely OSS ideologies [6]. Understanding OSS ideologies could provide significant insights into open source development, since OSS ideologies determine and influence the dynamics and outcomes of open source development [1, 2]. Assessing OSS ideologies within open source projects could bring various benefits to OSS practitioners, e.g., the owners and the maintainers of open source projects could identify important ideological elements that were previously ignored, and could improve open source development accordingly. Therefore, with such an assessment of OSS ideologies, institutional and individual stakeholders who are interested in open source development could make informed decisions when interacting with open source projects.
Yang Yue 0003, Yi Wang 0013, David F. Redmiles
ASE2
2024 Characterizing Developers' Linguistic Behaviors in Open Source Development across Their Social Statuses
abstract
Open Source Software (OSS) development has attracted numerous developers. As a typical complex sociotechnical system, an OSS project often forms a hierarchical social structure where a few developers are elite while the rest are non-elite. Differences in social status may result in distinct language use behaviors in interpersonal communication. Characterizing such behaviors is critical for supporting efficient and effective communication among developers with different social statuses. This study empirically compared elite and non-elite developers' language behaviors in their communication. We compiled a corpus of - 216,000 discourses collected from 20 large projects on GitHub. We investigated the linguistic differences in three aspects, namely, linguistic styles and characters, main concerns, and sentence patterns. Our findings reveal that elite and non-elite developers showed different linguistic patterns and had different concerns in their discourses. Their discourses also reflect the variation of the main focuses in the development process. Furthermore, elite and non-elite developers exhibited noticeable patterns in their linguistic behaviors in accordance with their roles and corresponding divisions of labor in the production process, no matter which semantic contexts. These findings provide implications for supporting communication that crosses social statuses in OSS development.
Yisi Han, Zhendong Wang 0003, Yang Feng 0003, Yi Wang 0013
Proc. ACM Hum. Comput. Interact.5
2023 OPRADI: Applying Security Game to Fight Drive under the Influence in Real-World
abstract
Driving under the influence (DUI) is one of the main causes of traffic accidents, often leading to severe life and property losses. Setting up sobriety checkpoints on certain roads is the most commonly used practice to identify DUI-drivers in many countries worldwide. However, setting up checkpoints according to the police's experiences may not be effective for ignoring the strategic interactions between the police and DUI-drivers, particularly when inspecting resources are limited. To remedy this situation, we adapt the classic Stackelberg security game (SSG) to a new SSG-DUI game to describe the strategic interactions in catching DUI-drivers. SSG-DUI features drivers' bounded rationality and social knowledge sharing among them, thus realizing improved real-world fidelity. With SSG-DUI, we propose OPRADI, a systematic approach for advising better strategies in setting up checkpoints. We perform extensive experiments to evaluate it in both simulated environments and real-world contexts, in collaborating with a Chinese city's police bureau. The results reveal its effectiveness in improving police's real-world operations, thus having significant practical potentials.
Luzhan Yuan, Wei Wang 0353, Yi Wang 0013
AAAI4
2023 Cross-status communication and project outcomes in OSS development
Yisi Han, Zhendong Wang 0003, Yang Feng 0003, Yi Wang 0013
Empir. Softw. Eng.5
2023 Off to a Good Start: Dynamic Contribution Patterns and Technical Success in an OSS Newcomer's Early Career
abstract
Attracting and retaining newcomers are critical aspects for OSS projects, as such projects rely on newcomers’ sustainable contributions. Considerable effort has been made to help newcomers by identifying and overcoming the barriers during the onboarding process. However, most newcomers eventually fail and drop out of their projects even after successful onboarding. Meanwhile, it has been long known that individuals’ early career stages profoundly impact their long-term career success. However, newcomers’ early careers are less investigated in SE research. In this paper, we sought to develop an empirical understanding of the relationships between newcomers’ dynamic contribution patterns in their early careers and their technical success. To achieve this goal, we compiled a dataset of newcomers’ contribution data from 54 large OSS projects under three different ecosystems and analyzed it with time series analysis and other statistical analysis techniques. Our analyses yield rich findings. The correlations between several contribution patterns and technical success were identified. In general, being consistent and persistent in newcomers’ early careers is positively associated with their technical success. While these correlations generally hold in all three ecosystems, we observed some differences in detailed contribution patterns correlated with technical success across ecosystems. In addition, we performed a case study to investigate whether another type of contributions, i.e., documentation contribution, could potentially have positive correlations with newcomers’ technical success. We discussed the implications and summarized practical recommendations to OSS newcomers. The insights gained from this work demonstrated the necessity of extending the focus of research and practice to newcomers’ early careers and hence shed light on future research in this direction.
Yang Yue 0003, Yi Wang 0013, David F. Redmiles
IEEE Trans. Software Eng.2
2021 IIAG: a data-driven and theory-inspired approach for advising how to interact with new remote collaborators in OSS teams
Yi Wang 0013, David F. Redmiles
Autom. Softw. Eng.1
2021 Living in a City, Living a Rural Life: Understanding Second Generation Mingongs' Experiences with Technologies in China
abstract
Rural-urban migrants (mingongs) provide crucial labor for China’s economic growth and global supply chains. Today, second generation mingongs who have spent most of their lives in cities have grown up. However, we know little about if their experiences with technologies are similar to their “urban-native” peers. This study reports on a qualitative study in a community in Beijing. We found a new type of “rurality”: second generation mingongs’ experiences with technologies differed from their urban-native peers in nearly every aspect, but exhibited similarities with their peers in rural areas. Taking nostalgia and memory as theoretical lenses, we demonstrate that such a “rurality” could be a coping mechanism for mingongs’ identity struggles. Our work contributed to HCI and CSCW literature by identifying the existence of a new type of “rurality.” That is, although residing in the city for nearly almost all of their lives, these second generation mingongs experiences greatly differed from their urban-native peers while exhibiting certain similarities with people living in rural areas.
Yi Wang 0013
ACM Trans. Comput. Hum. Interact.1
2020 Reducing implicit gender biases in software development: does intergroup contact theory work?
abstract
The software development profession suffers from severe gender biases, which could be explicit and implicit. However, SE literature has not systematically explored and evaluated the methods for reducing gender biases, especially for implicit gender biases. This paper reports on a field experiment to examine whether the intergroup contact theory could reduce implicit gender biases in software development. In the field experiment, 280 undergraduate students taking a project-centric introductory software engineering course were assigned to 70 teams with different contact configurations. We measured and compared their explicit and implicit gender biases before and after contacts in their teams. The study yields a rich set of findings. First, we confirmed the positive effects of intergroup contact theory in reducing gender biases, particularly the implicit gender biases in both general and SE-specific contexts. We further revealed that such effects were subjected to different contact configurations. The intergroup contact theory's effects were maximized in teams where the number of females is greater than or equal to the number of males. When the female is the minority group in a team, contacts among members contribute to reducing male members' implicit gender biases but fail to result in the same scale of effects on female members' implicit gender biases. The findings provide insights into using intergroup contact theory in reducing implicit gender biases in software development contexts.
Yi Wang 0013, Min Zhang 0002
ESEC/SIGSOFT FSE1
2020 Quality assessment of crowdsourced test cases
Yuan Zhao 0010, Yang Feng 0003, Yi Wang 0013, Chunrong Fang, Zhenyu Chen 0001
Sci. China Inf. Sci.3
2020 Unveiling Elite Developers' Activities in Open Source Projects
abstract
Open source developers, particularly the elite developers who own the administrative privileges for a project, maintain a diverse portfolio of contributing activities. They not only commit source code but also exert significant efforts on other communicative, organizational, and supportive activities. However, almost all prior research focuses on specific activities and fails to analyze elite developers’ activities in a comprehensive way. To bridge this gap, we conduct an empirical study with fine-grained event data from 20 large open source projects hosted on G IT H UB . We investigate elite developers’ contributing activities and their impacts on project outcomes. Our analyses reveal three key findings: (1) elite developers participate in a variety of activities, of which technical contributions (e.g., coding) only account for a small proportion; (2) as the project grows, elite developers tend to put more effort into supportive and communicative activities and less effort into coding; and (3) elite developers’ efforts in nontechnical activities are negatively correlated with the project’s outcomes in terms of productivity and quality in general, except for a positive correlation with the bug fix rate (a quality indicator). These results provide an integrated view of elite developers’ activities and can inform an individual’s decision making about effort allocation, which could lead to improved project outcomes. The results also provide implications for supporting these elite developers.
Zhendong Wang 0003, Yang Feng 0003, Yi Wang 0013, James A. Jones, David F. Redmiles
ACM Trans. Softw. Eng. Methodol.3
2019 KupC: A Formal Tool for Modeling and Verifying Dynamic Updating of C Programs
abstract
Dynamic Software Updating (DSU) is a useful technique for updating running software without incurring any downtime. Its correctness must be guaranteed because updating a running software is a complicated and safety-critical process. In this paper, we present a formal tool called KupC for modeling and verifying dynamic updating of C programs. The tool is built on $$\mathbb {K}$$ –a formal semantic framework for programming languages. We formalize a patch-based dynamic updating mechanism in $$\mathbb {K}$$ based on the formal executable operational semantics of C. The formalization automatically yields an interpreter and several verification tools, which can be used to formally analyze the correctness of dynamic updating for C programs. To our knowledge, KupC is the first formal tool for code-level verification of dynamic software updating.
Jiaqi Qian, Min Zhang 0002, Yi Wang 0013, Kazuhiro Ogata 0001
FASE3
2019 Country stereotypes, initial trust, and cooperation in global software development teams
abstract
People have to expose to collaborators from different countries in global software engineering (GSE) teams. They often rely on their perceptions of the country stereotypes to form their initial beliefs of the foreign collaborators and to make decisions on how to work with them. In this article, we employ the Stereotype Content Model (SCM) to investigate how explicit and implicit country stereotypes influence people' trust towards foreign collaborators and their decisions on cooperative behaviors. We conduct an empirical study with 92 professional software engineers with GSE experience. The results show that both SCM's explicit and implicit warmth, as well as the explicit competency, have significant impacts on the GSE team members' trust and cooperative behaviors in their initial interactions with unfamiliar foreign collaborators. Our findings indicate that Globally distributed collaboration practitioners may still need to overcome the over-reliance on country-of-origin cues when making attributions on unfamiliar foreign collaborators.
Yi Wang 0013, Min Zhang 0002
ICGSE1
2019 Collaboration in global software development: an investigation on research trends and evolution
abstract
Global software development (GSD) done by geographically distributed teams of developers is one of the most common ways of developing software nowadays. Though GSD has various benefits, it also introduces challenges that have led to a plethora of research. This paper analyzes research papers published in top software engineering venues in recent years (2009-2018) focusing on team collaboration in order to understand the trend in GSD research. Out of 4,292 papers published in these venues, we found 33 papers that focused on team collaboration in the context of GSD. We study the kinds of data used in these papers and classify them into primary data (i.e., interview and observation data) and secondary data (i.e., repository and communication data) and found that interview data is the dominant type of data in these papers. We also found that the strength of evidence presented in most papers tends to be moderate.
Yang Yue 0003, Iftekhar Ahmed 0001, Yi Wang 0013, David F. Redmiles
ICGSE3
2019 Emotions Extracted from Text vs. True Emotions-An Empirical Evaluation in SE Context
abstract
Emotion awareness research in SE context has been growing in recent years. Currently, researchers often rely on textual communication records to extract emotion states using natural language processing techniques. However, how well these extracted emotion states reflect people's real emotions has not been thoroughly investigated. In this paper, we report a multi-level, longitudinal empirical study with 82 individual members in 27 project teams. We collected their self-reported retrospective emotion states on a weekly basis during their year-long projects and also extracted corresponding emotions from the textual communication records. We then model and compare the dynamics of these two types of emotions using multiple statistical and time series analysis methods. Our analyses yield a rich set of findings. The most important one is that the dynamics of emotions extracted using text-based algorithms often do not well reflect the dynamics of self-reported retrospective emotions. Besides, the extracted emotions match self-reported retrospective emotions better at the team-level. Our results also suggest that individual personalities and the team's emotion display norms significantly impact the match/mismatch. Our results should warn the research community about the limitations and challenges of applying text-based emotion recognition tools in SE research.
Yi Wang 0013
ASE1
2017 Characterizing Developer Behavior in Cloud Based IDEs
abstract
Background: Cloud based integrated development environments (IDEs) are rapidly gaining popularity for its native support and potential to accelerate DevOps. However, there is little research of how developers behave when interacting with these environments. Aims: To develop empirical knowledge about how developers behave when interacting with cloud based IDEs to deal with programming tasks at various difficulty levels. Method: We conducted a user study using a cloud based IDE, JazzHub. We collected and coded session trace data, self-reported effort and frustration levels, and screen recordings. Results: We built a Markov activity transition model that describes the transitions among common development activities such as coding, debugging, and searching for information. It also captures extended interactions with remote resources. We correlated activity transition with different code growth trajectories. Conclusion: The findings are an early step toward realizing the potential for enhanced interactions in cloud based IDEs. Our study provides empirical evidence that may inspire the future evolution of cloud based IDE designs and features.
Yi Wang 0013
ESEM1
2017 Using Collaborative Online Drawing to Build Up Distributed Teams
abstract
Building up effective teams over a distance is a challenging but common problem in global software engineering. We propose an approach to help build up teams through collaborative online drawing. Our goal is to evaluate how drawing, as one activity that can facilitate expression of personal affective status, can benefit distributed teams. Preliminary results indicate positive effects of collaborative online drawing in increasing team cohesion and positive emotions. We discuss our approach and preliminary results in terms of design implications for future collaborative drawing systems for building up distributed teams.
Mengyao Zhao, Yi Wang 0013, David F. Redmiles
ICGSE2
2016 The Diffusion of Trust and Cooperation in Teams with Individuals¿ Variations on Baseline Trust
abstract
Baseline trust, which refers to the personality aspect of trust and varies with different individuals, is essential for understanding the development of trust and cooperation in a team. At the same time, informal, non-work-related conversations (aka, cheap talk) have positive influences on the diffusion of trust and cooperation in global software engineering (GSE) practice. This paper seeks to develop an understanding of the influences of individuals' baseline trust on the diffusion of trust and cooperation, in the presence of cheap talk over the Internet. We employ a novel approach, designing a virtual experiment that integrates abstract agent-based modeling and simulation with realistic, empirical network structures and baseline trust data from two large open source projects (Lucene and Google Chromium OS). The results highlight the significant impact of baseline trust on the diffusion of trust and cooperation, for instance, the emergence of non-traditional diffusion trajectories. The results also demonstrate that proper seeding strategies can improve the effectiveness and efficiency of diffusion of trust and cooperation.
Yi Wang 0013, David F. Redmiles
CSCW1
2016 Exploring Trust and Cooperation Development with Agent-Based Simulation in A Pseudo Scale-free Network
abstract
Globally distributed collaboration requires cooperation and trust among team members. Current research suggests that informal, non-work related communication plays a positive role in developing cooperation and trust. However, the way in which teams connect, i.e. via a social network, greatly influences cooperation and trust development. The study described in this paper employs agent-based modeling and simulation to investigate the cooperation and trust development with the presence of informal, non-work-related communication in networked teams. Leveraging game theory, we present a model of how an individual makes strategic decisions when interacting with her social network neighbors. The results of simulation on a pseudo scale-free network reveal the conditions under which informal communication has an impact, how different network degree distributions affect efficient trust and cooperation development, and how it is possible to "seed" trust and cooperation development amongst individuals in specific network positions. This study is the first to use agent-based modeling and simulation to examine the relationships between scale-free networks' topological features (degree distribution), cooperation and trust development, and informal communication.
Yi Wang 0013, David F. Redmiles
GROUP1
2016 Cheap talk, cooperation, and trust in global software engineering - An evolutionary game theory model with empirical support
Yi Wang 0013, David F. Redmiles
Empir. Softw. Eng.1
2015 Strengthening collaborative groups through art-mediated self-expression
abstract
Self-expression and interpersonal sharing of emotion have been shown to strengthen groups. However, how to accomplish such interpersonal sharing in public settings is a challenge. In a pilot study of a prototype system, we sought to facilitate public self-expression and sharing of affective information. We followed five design principles around the concept of art-mediated self-expression and created a collective doodling installation. The pilot trial demonstrated positive results around engagement of end users. As we reflected on the results from this trial, we found implications for building trust and collaboration in teams.
Mengyao Zhao, Yi Wang 0013, David F. Redmiles
VL/HCC2
2013 Globally distributed system developers: their trust expectations and processes
abstract
Trust remains a challenge in globally distributed development teams. In order to investigate how trust plays out in this context, we conducted a qualitative study of 5 multi-national IT organizations. We interviewed 58 individuals across 10 countries and made two principal findings. First, study participants described trust in terms of their expectations of their colleagues. These expectations fell into one of three dimensions: that socially correct behavior will persist, that team members possess technical competency, and that individuals will demonstrate concern for others. Second, our study participants described trust as a dynamic process, with phases including formation, dissolution, adjustment and restoration. We provide new insights into these dimensions and phases of trust within distributed teams which extend existing literature. Our study also provides guidelines on effective practices within distributed teams in addition to providing implications for the extension of software engineering and collaboration tools.
Ban Al-Ani, Matthew J. Bietz, Yi Wang 0013, Erik H. Trainer, Benjamin Koehne, Sabrina Marczak, David F. Redmiles, Rafael Prikladnicki
CSCW3
2012 Distributed Developers and the Non-use of Web 2.0 Technologies: A Proclivity Model
abstract
We sought to understand the role that Web 2.0 technologies play in supporting the development of trust in globally distributed development teams. We found the use of Web 2.0 technologies to be minimal, with less than 25% of our participants reporting using them and many reporting the disadvantages of adopting them. In response, we sought to understand the factors that led to the use and non-use of these technologies in distributed development teams. We adopted a mix of qualitative and quantitative methods to analyze data collected from 61 interviewees representing all common roles in systems development. We discovered six factors that influenced the use and non-use of Web 2.0 technology. We present a proclivity model to frame our findings as well as our conclusions about the interrelationships between the results of our qualitative and quantitative analyses. We also present implications for the design of collaboration tools, which could lead to greater support and usage by distributed developers.
Ban Al-Ani, Yi Wang 0013, Sabrina Marczak, Erik H. Trainer, David F. Redmiles
ICGSE2
2010 Penalty policies in professional software development practice: a multi-method field study
abstract
Organizational Punishment/Penalty is a pervasive phenomenon in many professional organizations. In some software development organizations, punishment measures have been adopted in an attempt to improve software developers' performance, reduce the software defects, and hence ensure software quality. It is unclear whether these measures are effective. This article presents the results of a multi-method field study that analyzes software engineers' perception towards penalty policies in relation to software quality in a software development process. The results were generated via both qualitative and quantitative methods. Through interviews, we collected the individuals' perception towards the penalty policy. By extracting data in a software configuration management system, we identified several patterns of defects change. We found that while a penalty mechanism does help to reduce software defects in daily coding activity, it fails in achieving programmers' maximum work potential. Meanwhile, experienced software programmers require less time to adapt to penalty policies and benefit from exist of less experienced developers. Some additional findings and implications are also discussed.
Yi Wang 0013, Min Zhang 0002
ICSE (2)1
2007 Specifying Pointcuts in AspectJ
abstract
Program verification is a promising approach to improving program quality. To formally verify aspect- oriented programs, we have to find a way to formally specify programs written in aspect-oriented languages. Pipa is a BISL tailored to AspectJ for specifying AspectJ programs. However, Pipa has not provided specification method for pointcuts in AspectJ programs. Based on the exist work of Pipa, and related issues, this paper proposes an approach to specifying pointcuts using purity conception in JML. This paper also provides several examples to illustrate our pointcut specification approach.
Yi Wang 0013, Jianjun Zhao 0001
COMPSAC (2)1