Shengyu Zhao

dblp:225/5446 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 OpenDigger: A Practical Framework for Assessing Community Health and Sustainability in Open Source Collaboration Platforms
abstract
The rapid development and widespread adoption of open source software, facilitated and accelerated by the web, have fostered a vibrant ecosystem for collaborative development and innovation. GitHub, a leading platform for collaborative software development, currently hosts more than 100 million registered users, creating a substantial ecosystem for examining open source community behaviors. Existing tools for measuring open source communities primarily focus on metrics such as issue response time, pull request response time, or incremental stars to provide insights into community activity. However, these tools are limited in their ability to assess the influence of communities from the perspective of collaboration networks. Moreover, current data collection solutions offer fixed functionalities and lack the flexibility to support multi-source, fine-grained, and customizable data acquisition, which is essential for comprehensive analysis of Open Source Ecosystems (OSEs). In this paper, we present OpenDigger, a framework for multi-dimensional assessment of collaboration activities in OSEs. To enable scalable, modular, and continuous acquisition of OSE data, we developed OpenCrawler, a one-line service providing customizable, fine-grained control over data collection. Using the collected data, OpenDigger computes 20 statistical and 2 network-based metrics, and our empirical analysis further verifies their effectiveness in enabling a comprehensive assessment of trends in OSEs. By continuously collecting logs from GitHub and Gitee, OpenDigger has now accumulated over 9 billion records. Our framework has already been deployed across multiple industrial environments, including Alibaba Group, Ant Group, Apache Foundation, and Mulan Open Source Community.
Wei Wang 0033, Fanyu Han, Shengyu Zhao, Xuan Zhou 0001, Weining Qian, Aoying Zhou, Xiaoya Xia, Moming Duan
WWW3
2026 Are cloud providers exploiting open-source? An exploratory study of Redis license change
abstract
Context: On March 20, 2024, Redis Inc. changed the Redis project’s license from the permissive Berkeley Software Distribution (BSD) license to a dual licensing model under the Redis Source Available License (RSALv2) and the Server Side Public License (SSPLv1). The official rationale for this change was to restrict cloud service providers from offering Redis as a managed service without contributing back. The license transition drew widespread attention within the open source community and raised concerns from some developers and organizations about its implications for project governance, contribution dynamics, and long-term sustainability. Objective: Analyze the contribution distribution within the Redis repository by examining the behavior of developers with different motivations, evaluate the changes that occurred during the license change period, and assess whether cloud providers are exploiting the open-source project. Method: This study categorizes developers’ motivations based on collaboration behavior data. By leveraging developer profile information, contributors are categorized into three motivation types: company-driven, community-driven, and communication-driven, enabling an analysis of trends across different contributor groups. The study primarily relies on project evaluation metrics from the Community Health Analytics Open Source Software (CHAOSS) community, along with collaboration metrics, to examine changes within the Redis repository. Results: Since 2017, cloud providers have consistently contributed to the Redis open-source project. During the license change period, there was a notable decrease in contributor participation and influence, particularly among company-driven and community-driven developers. Several core contributors transitioned to the newly established Valkey project. Conclusion: Following the license change by Redis Inc., the community experienced a certain degree of fragmentation, with major cloud providers migrating to the Valkey fork. Cloud providers have recognized the importance of the community and are willing to invest resources into the open-source projects they participate in, ensuring better collaboration and alignment with upstream development for their cloud services.
Fanyu Han, Shengyu Zhao, Xiaoya Xia, Wei Wang 0033
Inf. Softw. Technol.2
2026 OpenRank: A centrality algorithm for high-dimensional heterogeneous networks in open source collaboration
Fanyu Han, Shengyu Zhao, Wei Wang 0033, Jiaheng Peng, Xiaoya Xia
Inf. Sci.2
2024 OSGraph: A Data Visualization Insight Platform for Open Source Community
Wenrui Huang, Xiaoya Xia, Aoying Zhou, Xuan Zhou 0001, Wei Wang 0011, Shengyu Zhao, Sikang Bian
DASFAA (7)6
2024 HyperCRX: A Browser Extension for Insights into GitHub Projects and Developers
abstract
We present HyperCRX, a browser extension designed to enhance the open source experience by providing insights into GitHub projects and developers. Our tool seamlessly integrates with GitHub, offering interactive features that reveal the maintenance status of projects, the activity level of developers, and the connections between projects or developers. To ensure optimal performance, we have developed a special feature loading mechanism that supports the smooth operation of the existing 14 features. HyperCRX is now available on Google Chrome and Microsoft Edge, boasting a significant user base of 1,000 users on each browser. Our extension caters to a diverse range of users, including software engineers, project maintainers, open source newcomers, and company employers. With HyperCRX, users can make informed decisions and gain valuable insights to facilitate their engagement in the open source community. The source code is available on GitHub at https://github.com/hypertrons/hypertrons-crx, and a video screencast demonstrating the features can be found at https://youtu.be/_zm3FfpnZ28.
Yenan Tang, Shengyu Zhao, Xiaoya Xia, Fenglin Bi, Wei Wang 0033
ICPC2
2024 OpenGalaxy: An interactive exploration platform for a visualized GitHub Full Domain collaboration network
abstract
In this work, we introduce OpenGalaxy - an interactive exploration platform tool for a visualized GitHub Full Domain collaboration network based on 3D force-oriented layouts. We first collected Github domain-wide log data, built a developer-repository heterogeneous collaboration network, calculated both the Influence value for each repository and the activity value for each developer, finally performed visualization of the collaboration network and set up an 3D game-like interaction patterns.
Shengyu Zhao, Yenan Tang, Xiaoya Xia, Wei Wang 0033
ICPC2
2023 Understanding the Archived Projects on GitHub
abstract
Open source software (OSS) on GitHub is experiencing continued growth and a rapid rise in the creation of projects. As OSS evolved, numerous projects faced decline, with some inevitably degrading into unmaintained status and getting archived by the owners. Understanding the deprecated projects helps deepen the perception of OSS maintenance and evolution. This paper describes the study of 361 popular GitHub projects that have been archived. By reading repository READMEs and sending surveys to project maintainers, we found the software repositories were archived due to being transferred, evolved, or unmaintained. We provide a set of 16 reasons and 10 practices to describe why and how these projects were archived. We further reveal the OSS development activity lifecycle patterns by fitting their commit history curves. We also confirmed the importance of the bus factor risk that would influence the sustainability of open source projects. Archiving a software repository is an explicit indication of repository deprecation. By providing the reason (the why), the strategy (the how), and the lifecycle patterns (the what) of the archived projects, we bring implications and practices that promote the health and sustainability of open source projects and the ecosystem overall.
Xiaoya Xia, Shengyu Zhao, Zehua Lou, Wei Wang 0033, Fenglin Bi
SANER2
2022 Exploring Activity and Contributors on GitHub: Who, What, When, and Where
abstract
Apart from being a code hosting platform, GitHub is the place where large-scale open collaborations and contributions happen. Every minute, thousands of developers are submitting code, having discussions of issues or pull requests, with all user behaviors recorded in the GitHub Event Stream (GES). Exploration of the activities in the GES could help understand who is active, the way they work, the time when they are active and even their location. To this end, a large-scale analysis was initially performed based on the 0.86 billion event records generated in 2020. We extracted 902K active contributors out of 14 million GitHub accounts by observing their activity distribution, then explored their behavior distribution, active time in the day and week, and estimated time zone distributions on the basis of their circadian activity rhythm. To go deeper, a case study of 79 projects in CNCF and contrast analyses of different project maturity levels were conducted. Our results showed that from a macro perspective, bots are increasingly more active and can serve numerous projects. Contributors work on weekdays, and are globally more inclined toward the daytime working hours in the Americas and Europe. The time zone distribution also reveals that UTC+2 and UTC-4 have the most active contributors. A critical discovery was the validation and quantification of a high bus factor risk exists in the OSS ecosystem. Whether from a large group point of view or within specific projects, a rather small group of OSS contributors (less than 20%) undertook the majority of the work. The GES can provide a wealth of information about open source software (OSS). Our findings provide insights into global GitHub collaboration behaviors and may be of help for researchers and practitioners to further understand modern OSS ecosystem.
Xiaoya Xia, Zhenjie Weng, Wei Wang 0033, Shengyu Zhao
APSEC4
2022 PVNAS: 3D Neural Architecture Search With Point-Voxel Convolution
abstract
3D neural networks are widely used in real-world applications (e.g., AR/VR headsets, self-driving cars). They are required to be fast and accurate; however, limited hardware resources on edge devices make these requirements rather challenging. Previous work processes 3D data using either voxel-based or point-based neural networks, but both types of 3D models are not hardware-efficient due to the large memory footprint and random memory access. In this paper, we study 3D deep learning from the efficiency perspective. We first systematically analyze the bottlenecks of previous 3D methods. We then combine the best from point-based and voxel-based models together and propose a novel hardware-efficient 3D primitive, Point-Voxel Convolution (PVConv). We further enhance this primitive with the sparse convolution to make it more effective in processing large (outdoor) scenes. Based on our designed 3D primitive, we introduce 3D Neural Architecture Search (3D-NAS) to explore the best 3D network architecture given a resource constraint. We evaluate our proposed method on six representative benchmark datasets, achieving state-of-the-art performance with 1.8-23.7× measured speedup. Furthermore, our method has been deployed to the autonomous racing vehicle of MIT Driverless, achieving larger detection range, higher accuracy and lower latency.
Haotian Tang, Shengyu Zhao, Kevin Shao, Song Han 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Large Scale Image Completion via Co-Modulated Generative Adversarial Networks
Shengyu Zhao, Jonathan Cui, Yilun Sheng, Eric I-Chao Chang, Yan Xu 0001
ICLR1
2020 MaskFlownet: Asymmetric Feature Matching With Learnable Occlusion Mask
abstract
Feature warping is a core technique in optical flow estimation; however, the ambiguity caused by occluded areas during warping is a major problem that remains unsolved. In this paper, we propose an asymmetric occlusion-aware feature matching module, which can learn a rough occlusion mask that filters useless (occluded) areas immediately after feature warping without any explicit supervision. The proposed module can be easily integrated into end-to-end network architectures and enjoys performance gains while introducing negligible computational cost. The learned occlusion mask can be further fed into a subsequent network cascade with dual feature pyramids with which we achieve state-of-the-art performance. At the time of submission, our method, called MaskFlownet, surpasses all published optical flow methods on the MPI Sintel, KITTI 2012 and 2015 benchmarks. Code is available at https://github.com/microsoft/MaskFlownet.
Shengyu Zhao, Yilun Sheng, Eric I-Chao Chang, Yan Xu 0001
CVPR1
2020 Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution
Haotian Tang, Shengyu Zhao, Yujun Lin 0001, Ji Lin 0002, Hanrui Wang 0002, Song Han 0003
ECCV (28)3
2020 Differentiable Augmentation for Data-Efficient GAN Training
abstract
The performance of generative adversarial networks (GANs) heavily deteriorates given a limited amount of training data. This is mainly because the discriminatorsis memorizing the exact training set. To combat it, we propose Differentiable Augmentation (DiffAugment), a simple method that improves the data efficiency of GANs by imposing various types of differentiable augmentations on both real and fake samples. Previous attempts to directly augment the training data manipulate the distribution of real images, yielding little benefit; DiffAugment enables us to adopt the differentiable augmentation for the generated samples, effectively stabilizes training, and leads to better convergence. Experiments demonstrate consistent gains of our method over a variety of GAN architectures and loss functions for both unconditional and class-conditional generation. With DiffAugment, we achieve astate-of-the-art FID of 6.80 with an IS of 100.8 on ImageNet 128×128 and 2-4× reductions of FID given 1,000 images on FFHQ and LSUN. Furthermore, with only 20% training data, we can match the top performance on CIFAR-10 and CIFAR-100. Finally, our method can generate high-fidelity images using only 100 images without pre-training, while being on par with existing transfer learning algorithms. Code is available at https://github.com/mit-han-lab/data-efficient-gans.
Shengyu Zhao, Ji Lin 0002, Jun-Yan Zhu, Song Han 0003
NeurIPS1
2020 Characterization of Group-strategyproof Mechanisms for Facility Location in Strictly Convex Space
abstract
We characterize the class of group-strategyproof mechanisms for the single facility location game in any unconstrained strictly convex space. A mechanism is group-strategyproof,if no group of agents can misreport so that all its members are strictlybetter off. A strictly convex space is a normed vector space where |x+y|<2 holds for any pair of different unit vectors x ≠ y, e.g., any Lp space with p∈ (1,∞).
Pingzhong Tang, Dingli Yu, Shengyu Zhao
EC3
2020 Unsupervised 3D End-to-End Medical Image Registration With Volume Tweening Network
abstract
3D medical image registration is of great clinical importance. However, supervised learning methods require a large amount of accurately annotated corresponding control points (or morphing), which are very difficult to obtain. Unsupervised learning methods ease the burden of manual annotation by exploiting unlabeled data without supervision. In this article, we propose a new unsupervised learning method using convolutional neural networks under an end-to-end framework, Volume Tweening Network (VTN), for 3D medical image registration. We propose three innovative technical components: (1) An end-to-end cascading scheme that resolves large displacement; (2) An efficient integration of affine registration network; and (3) An additional invertibility loss that encourages backward consistency. Experiments demonstrate that our algorithm is 880x faster (or 3.3x faster without GPU acceleration) than traditional optimization-based methods and achieves state-of-the-art performance in medical image registration.
Shengyu Zhao, Ting Fung Lau, Ji Luo 0002, Eric I-Chao Chang, Yan Xu 0001
IEEE J. Biomed. Health Informatics1
2020 ANHIR: Automatic Non-Rigid Histological Image Registration Challenge
abstract
Automatic Non-rigid Histological Image Registration (ANHIR) challenge was organized to compare the performance of image registration algorithms on several kinds of microscopy histology images in a fair and independent manner. We have assembled 8 datasets, containing 355 images with 18 different stains, resulting in 481 image pairs to be registered. Registration accuracy was evaluated using manually placed landmarks. In total, 256 teams registered for the challenge, 10 submitted the results, and 6 participated in the workshop. Here, we present the results of 7 well-performing methods from the challenge together with 6 well-known existing methods. The best methods used coarse but robust initial alignment, followed by non-rigid registration, used multiresolution, and were carefully tuned for the data at hand. They outperformed off-the-shelf methods, mostly by being more robust. The best methods could successfully register over 98% of all landmarks and their mean landmark registration accuracy (TRE) was 0.44% of the image diagonal. The challenge remains open to submissions and all images are available for download.
Jirí Borovec, Jan Kybic, Ignacio Arganda-Carreras, Dmitry V. Sorokin, Gloria Bueno García, Alexander V. Khvostikov, Spyridon Bakas, Eric I-Chao Chang, Stefan Heldmann, Kimmo Kartasalo, Leena Latonen, Johannes Lotz 0002, Michelle Noga, Sarthak Pati, Kumaradevan Punithakumar, Pekka Ruusuvuori, Andrzej Skalski, Nazanin Tahmasebi, Masi Valkonen, Ludovic Venet, Nick Weiss, Marek Wodzinski, Yan Xu 0001, Paul A. Yushkevich, Shengyu Zhao, Arrate Muñoz-Barrutia
IEEE Trans. Medical Imaging28
2019 Recursive Cascaded Networks for Unsupervised Medical Image Registration
abstract
We present recursive cascaded networks, a general architecture that enables learning deep cascades, for deformable image registration. The proposed architecture is simple in design and can be built on any base network. The moving image is warped successively by each cascade and finally aligned to the fixed image; this procedure is recursive in a way that every cascade learns to perform a progressive deformation for the current warped image. The entire system is end-to-end and jointly trained in an unsupervised manner. In addition, enabled by the recursive architecture, one cascade can be iteratively applied for multiple times during testing, which approaches a better fit between each of the image pairs. We evaluate our method on 3D medical images, where deformable registration is most commonly applied. We demonstrate that recursive cascaded networks achieve consistent, significant gains and outperform state-of-the-art methods. The performance reveals an increasing trend as long as more cascades are trained, while the limit is not observed.
Shengyu Zhao, Eric I-Chao Chang, Yan Xu 0001
ICCV1