VLDB 2026 Research / reviewers in the wild / expert
Wenyi Zhao
dblp:38/6375 · also Wen-Yi Zhao
· DBLP profile ↗
76ranked-venue papers
34as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 17 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 21 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 11 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identification of cancer mini-drivers by deciphering selective landscape in the cancer genomeabstractCancer development is driven by somatic evolution and clonal selection. However, traditional selective pressure analysis methods have treated all sites within a gene equally, such a gene-level model oversimplifies the complexity of cancer evolution. In this study, we introduced CN/CS-calculator, a novel site-specific method that can capture selective pressures acting across different gene sites. By deciphering the interplay between the selection pattern and the function of a gene in oncogenesis, CN/CS-calculator uncovers a unique class of mini-driver genes, which exhibit weak positive selection, with certain critical sites providing context-dependent promoter effects on the fitness of cancer subclones while others are constrained by evolutionary conservation. Our method emphasizes the importance of site-specific analysis in uncovering how subtle evolutionary forces shape cancer biology. The refined understanding offers new insights into the mechanisms of cancer heterogeneity and molecular evolution, with potential implications for advancing therapeutic strategies and prognostic assessments. Xunuo Zhu, Wenyi Zhao, Jingqi Zhou, Binbin Zhou 0005, Zhan Zhou, Xun Gu 0002 |
Briefings Bioinform. | 2 |
| 2026 | Multi-scale feature enhancement network for object detection in severe foggy weather
Yingjun Wang, Yingjian Wang 0002, Peixian Zhuang, Wenyi Zhao, Haoxiang Lu, Weidong Zhang 0007 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | CFSPNet : Cross-domain feature synergy and perception for robust underwater object detection
Dechuan Kong, Yandi Zhang, Wenyi Zhao, Deguang Li, Weidong Zhang 0007 |
Expert Syst. Appl. | 4 |
| 2026 | URDNet: Unsupervised retinex decomposition network for low-light image enhancement
Xingyun Gao, Wenyi Zhao, Deguang Li, Zheng Liang 0001, Weidong Zhang 0007 |
Inf. Sci. | 2 |
| 2026 | FGDNet: Frequency-domain guided degradation-aware network for object detection in adverse weather
Yingjun Wang, Deguang Li, Zheng Liang 0001, Wenyi Zhao, Weidong Zhang 0007 |
Inf. Sci. | 6 |
| 2026 | Mutual masked image consistency and feature adversarial training for semi-supervised medical image segmentation
Wei Li 0243, Linye Ma, Wenyi Zhao |
Knowl. Based Syst. | 3 |
| 2026 | TD loss: Taylor expansion of Dice loss for robust medical image segmentation
Weijin Xu, Wenyi Zhao |
Medical Image Anal. | 4 |
| 2026 | Semi-supervised medical image segmentation via pseudo-labeling refinement and dual-adaptive adjustment schemes
Wenyi Zhao, Wentao Liu 0004, Chuanbo Qin |
Pattern Recognit. | 2 |
| 2026 | Underwater image color correction via retinex reflectance-layer color transfer
Kangle Qiao, Deguang Li, Ling Zhou 0003, Songlin Jin, Wenyi Zhao |
Pattern Recognit. Lett. | 5 |
| 2026 | AMDC: Attenuation map-guided dual-color space for underwater image color correction
Shilong Sun 0003, Baiqiang Yu, Ling Zhou 0003, Junpeng Xu, Wenyi Zhao, Weidong Zhang 0007 |
Pattern Recognit. Lett. | 5 |
| 2026 | Cross-scale coupled attention network for underwater image enhancement
Gaoli Zhao, Kefei Zhang 0001, Song Han 0007, Wenyi Zhao, Weidong Zhang 0007 |
Pattern Recognit. Lett. | 5 |
| 2026 | Underwater image color correction via global-local collaborative strategy
Ling Zhou 0003, Baiqiang Yu, Wenyi Zhao, Weidong Zhang 0007 |
Pattern Recognit. Lett. | 4 |
| 2026 | TLVNet: Triple Latent Variational Attention Network for underwater image enhancement
Gaoli Zhao, Junping Song, Haoxiang Lu, Wenyi Zhao, Zheng Liang 0001, Weidong Zhang 0007 |
Signal Process. Image Commun. | 6 |
| 2026 | A comprehensive review of low-light image enhancement methods
Ling Zhou 0003, Kaijie Jin, Songlin Jin, Zheng Liang 0001, Wenyi Zhao, Weidong Zhang 0007 |
Signal Process. Image Commun. | 6 |
| 2026 | MFPD: Mamba-Driven Feature Pyramid Decoding for Underwater Object DetectionabstractUnderwater object detection suffers from limited long-range dependency modeling, fine-grained feature representation, and noise suppression, resulting in blurred boundaries, frequent missed detections, and reduced robustness. To address these challenges, we propose the Mamba-Driven Feature Pyramid Decoding framework, which employs a parallel Feature Pyramid Network and Path Aggregation Network collaborative pathway to enhance semantic and geometric features. A lightweight Mamba Block models long-range dependencies, while an Adaptive Sparse Self-Attention module highlights discriminative targets and suppresses noise. Together, these components improve feature representation and robustness. Experiments on two publicly available underwater datasets demonstrate that MFPD significantly outperforms existing methods, validating its effectiveness in complex underwater environments. The code is publicly available at:https://github.com/YitengGuo/MFPD Yiteng Guo, Junpeng Xu, Wenyi Zhao, Weidong Zhang 0007 |
IEEE Signal Process. Lett. | 4 |
| 2026 | ES-DETR: Edge-Guided State-Space DETR for Foggy Remote Sensing Object Detection
Qiang Zhang 0011, Zheng Liang 0001, Wenyi Zhao, Weidong Zhang 0007 |
IEEE Signal Process. Lett. | 4 |
| 2026 | DFFR: DETR With Foreground-Guided Feature Refinement Network for End-to-End Underwater Object Detection
Gaoli Zhao, Kefei Zhang 0001, Zheng Liang 0001, Wenyi Zhao, Weidong Zhang 0007 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | Underwater Image Enhancement via Intelligent Optimized Multi-Exposure Image FusionabstractUnderwater images often suffer from visual degradation due to varying light absorption at different wavelengths and scattering from suspended particles. To tackle these issues, we present an intelligent optimized multi-exposure image fusion method called IMIF. Specifically, we propose an adaptive color transfer strategy that employs a colorless reference image to correct the color distortion issue by transferring the mean and standard deviation of the reference image to adjust a color-balanced image. Subsequently, we introduce a particle swarm optimization algorithm that intelligently selects the optimal set of exposure image sequences by employing information entropy and edge intensity of the image as fitness metrics. Meanwhile, we leverage a guided filtering strategy to decompose the exposure image sequences into basic and detailed layers, taking into account the exposure characteristics of each layer to generate corresponding weight maps. Finally, we employ a multi-exposure fusion strategy to adaptively fuse the exposed image sequences with weight maps, producing an enhanced result. Extensive experiments conducted on three datasets demonstrate that our IMIF method outperforms state-of-the-art (SOTA) methods in both qualitative and quantitative evaluations. Additionally, the enhanced results produced by our proposed IMIF method significantly improve the accuracy of object detection and keypoint detection. The is available at https://www.researchgate.net/publication/403951386_2026-IMIF. Weidong Zhang 0007, Baiqiang Yu, Wenyi Zhao, Zheng Liang 0001, Peixian Zhuang, Keran Zhu |
IEEE Trans. Image Process. | 3 |
| 2025 | Refined self-supervised learning via trident network and cross-hybrid optimization
Wenyi Zhao, Wei Li 0243, Yongqin Tian, Yingjian Wang 0002, Weidong Zhang 0007 |
Expert Syst. Appl. | 1 |
| 2025 | CIDNet: Cross-Scale Interference Mining Detection Network for underwater object detection
Gaoli Zhao, Kefei Zhang 0001, Liangzhi Wang, Wenyi Zhao, Weidong Zhang 0007 |
Knowl. Based Syst. | 4 |
| 2025 | Underwater Optical Image Contrast Enhancement via Color Channel MatchingabstractDue to the complex physical environment underwater, underwater captured images often suffer from issues such as color distortion, low contrast, and loss of texture details. To address this issue, we propose a color channel matching (CCM) method for underwater optical image contrast enhancement, called CCM. Specifically, we first convert the raw image into a grayscale image and employ histogram matching techniques to make the brightness distribution of the image more uniform, thereby reducing brightness variations caused by environmental factors. Then, we transform the matched image into the hue-saturation-intensity (HSI) color space and optimize the HSI channels separately. During this process, we decouple the intensity information from the color information to avoid interference during enhancement, which employs adaptive histogram equalization on the intensity channel to improve contrast and detailed representation further. Finally, we fuse the processed intensity channel with the optimized hue and saturation channels to obtain the final contrast-enhanced image. Extensive qualitative and quantitative experimental results demonstrate that the proposed method exhibits strong robustness and generalization capabilities in enhancing the contrast of underwater images. Xiaoguo Chen, Shilong Sun 0003, Yaqi Gao, Wenyi Zhao, Weidong Zhang 0007 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | S3H: Long-tailed classification via spatial constraint sampling, scalable network, and hybrid task
Wenyi Zhao, Wei Li 0243, Yongqin Tian, Enwen Hu, Wentao Liu 0004, Weidong Zhang 0007 |
Neural Networks | 1 |
| 2025 | Leveraging multi-level regularization for efficient Domain Adaptation of Black-box Predictors
Wei Li 0243, Wenyi Zhao, Xipeng Pan |
Pattern Recognit. | 2 |
| 2025 | RDANet: Retinex decomposition attention network for low-light image enhancement
Xingyun Gao, Weibo Zhang, Peixian Zhuang, Wenyi Zhao, Weidong Zhang 0007 |
Pattern Recognit. Lett. | 4 |
| 2025 | HSTNet: Hybrid Supervision-Driven Two-Stream Collaborative Network for Hyperspectral Wheat Variety ClassificationabstractHyperspectral remote sensing plays an important role in agricultural monitoring, and fine-grained wheat variety classification is essential for advancing smart agriculture. However, progress is limited by the scarcity of high-quality spectral samples and the challenges associated with collecting large-scale hyperspectral data. To cope with these issues, we design a hybrid supervised-driven two-stream collaborative network (HSTNet), which consists of a semi-supervised conditional generative adversarial network for data augmentation (SCGAN) and a supervised two-stream discriminative network (STDNet) for wheat classification. In SCGAN, the generator constructs a mapping relationship between input noise and real wheat hyperspectral samples to generate fake wheat hyperspectral samples that are highly matched with the distribution of the real sample, and the discriminator with multilayer perceptions utilizes discriminative learning to identify real and fake samples. In STDNet, it collaboratively extracts the spectral, spatial and texture features of wheat hyperspectral images employing the dual-stream branch structure of 3DCNN and 2DCNN. Subsequently, it utilizes the Fast Fourier Transform and the cross-attention mechanism to refine and fuse these features to improve their capability of feature expression. Noteworthy, the individual design effectively improves the classification results of wheat varieties via collaborative optimization among modules. Besides, we built a mixed wheat hyperspectral dataset (MWHD) with 4800 samples of 20 wheat varieties. Extensive experiments on our constructed MWHD dataset demonstrate that the proposed HSTNet outperforms state-of-the-art methods in wheat variety classification. The code is publicly available at: https://github.com/bakam412/HSTNet. Ling Zhou 0003, Shuoguo Cui, Qiang Zhang 0011, Wenyi Zhao, Zheng Liang 0001, Weidong Zhang 0007 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Constructing Balanced Training Samples: A New Perspective on Long-Tailed ClassificationabstractThe most significant characteristic of long-tailed classification is that severe sample imbalance causes the model to be biased towards the head category. While the long-tailed distribution of multimedia dataset remains a constant, we can enhance the acquisition of balanced training samples and corresponding features during the learning process. This paper innovatively designs a sample provider to construct balanced training samples to enhance the acquisition of comprehensive features, and proposes a Siamese-based parameter-sharing framework to handle data with long-tailed distributions. Specifically, one branch of the Siamese network is introduced to classify samples with conventional random cropping sampling, another branch integrates the advantages of constructed balanced samples and hybrid optimization to capture the balanced features to identify more precise category boundaries. This combination not only facilitates the learning of long-tailed distribution but also strengthens the model's extraction of balanced features through the incorporation of contrastive learning. Most significantly, extensive experiments on CIFAR10-LT, CIFAR100-LT, ImageNet-LT and iNaturalist 2018 datasets demonstrate our model not only achieves superior performance but also retains the benefits of end-to-end training. Specifically, our method achieves 60.7% accuracy on ImageNet-LT with an end-to-end ResNeXt-50 backbone. Wenyi Zhao, Wei Li 0243, Lu Yang 0006, Zhenhao Liang, Enwen Hu, Weidong Zhang 0007 |
IEEE Trans. Multim. | 1 |
| 2024 | Unified multi-color-model-learning-based deep support vector machine for underwater image classification
Weidong Zhang 0007, Baiqiang Yu, Guohou Li, Peixian Zhuang, Zheng Liang 0001, Wenyi Zhao |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | CATNet: Cascaded attention transformer network for marine species image classification
Weidong Zhang 0007, Gongchao Chen, Peixian Zhuang, Wenyi Zhao, Ling Zhou 0003 |
Expert Syst. Appl. | 4 |
| 2024 | Autonomous driving system: A comprehensive survey
Wenyi Zhao, Zhenghong Wang, Feng Zhang 0007, Wenxiang Zheng, Wanke Cao, Jinrui Nan, Yubo Lian, Andrew F. Burke |
Expert Syst. Appl. | 2 |
| 2024 | S4: Self-supervised learning with sparse-dense sampling
Yongqin Tian, Weidong Zhang 0007, Peixian Zhuang, Xiwang Xie, Wenyi Zhao |
Knowl. Based Syst. | 7 |
| 2024 | ZAP: Underwater Image Color Correction via Zero Approximation PrincipleabstractUnderwater images widely endure severe color distortion because of the absorption and scattering of the water medium. We present a zero-approximation principle for underwater image color correction, called ZAP, to tackle this issue. Specifically, we first present the channel’s a and b pixel values to subtract the corresponding channels’ average pixel values so that the histograms corresponding to their pixel values are symmetric about the zero within the CIELab model. Afterward, we utilize the standard deviation of the channels’ a and b compensated pixel values to adjust the channels’ dynamic range and correct the image color distortion. Broad qualitative and quantitative experiments prove the practicability of ZAP for correcting color distortion in underwater images. Baiqiang Yu, Weidong Zhang 0007, Wenqiang Yu, Peixian Zhuang, Wenyi Zhao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Diversity matters: Cross-head mutual mean-teaching for semi-supervised medical image segmentation
Wei Li 0243, Ruifeng Bian, Wenyi Zhao, Weijin Xu |
Medical Image Anal. | 3 |
| 2024 | DIAS: A dataset and benchmark for intracranial artery segmentation in DSA sequences
Wentao Liu 0004, Tong Tian, Lemeng Wang, Weijin Xu, Wenyi Zhao, Xipeng Pan, Yiming Deng, Xin Wang 0121, Ruisheng Su |
Medical Image Anal. | 7 |
| 2024 | CVANet: Cascaded visual attention network for single image super-resolution
Weidong Zhang 0007, Wenyi Zhao, Jia Li 0019, Peixian Zhuang, Hai-Han Sun, Chongyi Li |
Neural Networks | 2 |
| 2024 | Local Reference Feature Transfer (LRFT): A simple pre-processing step for image enhancement
Ling Zhou 0003, Weidong Zhang 0007, Yuchao Zheng 0001, Jianping Wang 0004, Wenyi Zhao |
Pattern Recognit. Lett. | 5 |
| 2024 | Underwater Image Enhancement via Weighted Wavelet Visual Perception FusionabstractUnderwater images typically suffer from various quality degradation issues due to the scattering and absorption of light, but these degraded-quality underwater images are unbeneficial for analysis and applications. To effectively solve these quality degradation issues, an underwater image enhancement method via weighted wavelet visual perception fusion is introduced, called WWPF. Concretely, we first present an attenuation-map-guided color correction strategy to correct the color distortion of an underwater image. Subsequently, we employ the maximum information entropy optimized global contrast strategy to the color-corrected image to obtain a global contrast-enhanced image. Meanwhile, we apply a fast integration optimized local contrast strategy to the color-corrected image to get a local contrast-enhanced image. To exploit the complementary of the global contrast-enhanced image and the local contrast-enhanced image, we introduce a weighted wavelet visual perception fusion strategy to obtain a high-quality underwater image by fusing the high-frequency and low-frequency components of images at different scales. Our extensive experiments on three benchmarks validate that our WWPF outperforms the state-of-the-art methods in qualitative and quantitative. Besides, the underwater images processed by our WWPF also benefit practical underwater applications. The code is availablehttps://github.com/Li-Chongyi/WWPF_code. Weidong Zhang 0007, Ling Zhou 0003, Peixian Zhuang, Guohou Li, Xipeng Pan, Wenyi Zhao, Chongyi Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Learning What and Where to Learn: A New Perspective on Self-Supervised LearningabstractSelf-supervised learning (SSL) has demonstrated its power in generalized model acquisition by leveraging the discriminative semantic and explicit positional information of unlabeled datasets. Unfortunately, mainstream contrastive learning-based methods excessive focus on semantic information and ignore the position is also the carrier of image content, resulting in inadequate data utilization and extensive computational consumption. To address these issues, we present an efficient SSL framework, learning What and Where to learn (W2SSL), to aggregate semantic and position features. Concretely, we devise a spatially-coupled sampling manner to process images through pre-defined rules, which integrates the advantage of semantic (What) and positional (Where) features into framework to enrich the diversity of feature representation capabilities and improve data utilization. Besides, a spectrum of latent vectors is obtained by mapping the positional features, which implicitly explores the relationship between these vectors. Whereafter, the corresponding discriminative and contrastive optimization objectives are seamlessly embedded in the framework via a cascade paradigm to explore semantic and positional features. The proposed W2SSL is verified on different types of datasets, which demonstrates that it still outperforms state-of-the-art SSL methods even with half the computational consumption. Code will be available at https://github.com/WilyZhao8/W2SSL. Wenyi Zhao, Lu Yang 0006, Weidong Zhang 0007, Yongqin Tian, Wenhe Jia, Wei Li 0243, Mu Yang, Xipeng Pan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | MATTE: a pipeline of transcriptome module alignment for anti-noise phenotype-gene-related analysisabstractA phenotype may be associated with multiple genes that interact with each other in the form of a gene module or network. How to identify these relationships is one important aspect of comparative transcriptomics. However, it is still a challenge to align gene modules associated with different phenotypes. Although several studies attempted to address this issue in different aspects, a general framework is still needed. In this study, we introduce Module Alignment of TranscripTomE (MATTE), a novel approach to analyze transcriptomics data and identify differences in a modular manner. MATTE assumes that gene interactions modulate a phenotype and models phenotype differences as gene location changes. Specifically, we first represented genes by a relative differential expression to reduce the influence of noise in omics data. Meanwhile, clustering and aligning are combined to depict gene differences in a modular way robustly. The results show that MATTE outperformed state-of-the-art methods in identifying differentially expressed genes under noise in gene expression. In particular, MATTE could also deal with single-cell ribonucleic acid-seq data to extract the best cell-type marker genes compared to other methods. Additionally, we demonstrate how MATTE supports the discovery of biologically significant genes and modules, and facilitates downstream analyses to gain insight into breast cancer. The source code of MATTE and case analysis are available at https://github.com/zjupgx/MATTE. Guoxing Cai, Wenyi Zhao, Zhan Zhou, Xun Gu 0002 |
Briefings Bioinform. | 2 |
| 2023 | TIVE: A toolbox for identifying video instance segmentation errors
Wenhe Jia, Lu Yang 0006, Zilong Jia, Wenyi Zhao, Qing Song 0006 |
Neurocomputing | 4 |
| 2023 | Global-and-Local sampling for efficient hybrid task self-supervised learning
Wenyi Zhao, Lingqiao Li |
Knowl. Based Syst. | 1 |
| 2023 | BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachabstractCompiler optimization plays an increasingly important role to boost the performance of machine learning models for data processing and management. With increasingly complex data, the dynamic tensor shape phenomenon emerges for ML models. However, existing ML compilers either can only handle static shape models or expose a series of performance problems for both operator fusion optimization and code generation in dynamic shape scenes. This paper tackles the main challenges of dynamic shape optimization: the fusion optimization without shape value, and code generation supporting arbitrary shapes. To tackle the fundamental challenge of the absence of shape values, it systematically abstracts and excavates the shape information and designs a cross-level symbolic shape representation. With the insight that what fusion optimization relies upon is tensor shape relationships between adjacent operators rather than exact shape values, it proposes the dynamic shape fusion approach based on shape information propagation. To generate code that adapts to arbitrary shapes efficiently, it proposes a compile-time and runtime combined code generation approach. Finally, it presents a complete optimization pipeline for dynamic shape models and implements an industrial-grade ML compiler, named BladeDISC. The extensive evaluation demonstrates that BladeDISC outperforms PyTorch, TorchScript, TVM, ONNX Runtime, XLA, Torch Inductor (dynamic shape), and TensorRT by up to 6.95×, 6.25×, 4.08×, 2.04×, 2.06×, 7.92×, and 4.16× (3.54×, 3.12×, 1.95×, 1.47×, 1.24×, 2.93×, and 1.46× on average) in terms of end-to-end inference speedup on the A10 and T4 GPU, respectively. BladeDISC's source code is publicly available at https://github.com/alibaba/BladeDISC. Zhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu 0004, Wenyi Zhao, Tianyou Guo, Xiafei Qiu, Minmin Sun, Feng Zhang 0007, Xiaoyong Du 0001, Jidong Zhai, Wei Lin 0016 |
Proc. ACM Manag. Data | 5 |
| 2023 | Embedding Global Contrastive and Local Location in Self-Supervised LearningabstractSelf-supervised representation learning (SSL) typically suffers from inadequate data utilization and feature-specificity due to the suboptimal sampling strategy and the monotonous optimization method. Existing contrastive-based methods alleviate these issues through exceedingly long training time and large batch size, resulting in non-negligible computational consumption and memory usage. In this paper, we present an efficient self-supervised framework, called GLNet. The key insights of this work are the novel sampling and ensemble learning strategies embedded in the self-supervised framework. We first propose a location-based sampling strategy to integrate the complementary advantages of semantic and spatial characteristics. Whereafter, a Siamese network with momentum update is introduced to generate representative vectors, which are used to optimize the feature extractor. Finally, we particularly embed global contrastive and local location tasks in the framework, which aims to leverage the complementarity between the high-level semantic features and low-level texture features. Such complementarity is significant for mitigating the feature-specificity and improving the generalizability, thus effectively improving the performance of downstream tasks. Extensive experiments on representative benchmark datasets demonstrate that GLNet performs favorably against the state-of-the-art SSL methods. Specifically, GLNet improves MoCo-v3 by 2.4% accuracy on ImageNet dataset, while improves 2% accuracy and consumes only 75% training time on the ImageNet-100 dataset. In addition, GLNet is appealing in its compatibility with popular SSL frameworks. Code is available at GLNet. Wenyi Zhao, Chongyi Li, Weidong Zhang 0007, Lu Yang 0006, Peixian Zhuang, Lingqiao Li, Kefeng Fan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesabstractThis work reveals that memory-intensive computation is a rising performance-critical factor in recent machine learning models. Due to a unique set of new challenges, existing ML optimizing compilers cannot perform efficient fusion under complex two-level dependencies combined with just-in-time demand. They face the dilemma of either performing costly fusion due to heavy redundant computation, or skipping fusion which results in massive number of kernels. Furthermore, they often suffer from low parallelism due to the lack of support for real-world production workloads with irregular tensor shapes. To address these rising challenges, we propose AStitch, a machine learning optimizing compiler that opens a new multi-dimensional optimization space for memory-intensive ML computations. It systematically abstracts four operator-stitching schemes while considering multi-dimensional optimization objectives, tackles complex computation graph dependencies with novel hierarchical data reuse, and efficiently processes various tensor shapes via adaptive thread mapping. Finally, AStitch provides just-in-time support incorporating our proposed optimizations for both ML training and inference. Although AStitch serves as a stand-alone compiler engine that is portable to any version of TensorFlow, its basic ideas can be generally applied to other ML frameworks and optimization compilers. Experimental results show that AStitch can achieve an average of 1.84x speedup (up to 2.73x) over the state-of-the-art Google's XLA solution across five production workloads. We also deploy AStitch onto a production cluster for ML workloads with thousands of GPUs. The system has been in operation for more than 10 months and saves about 20,000 GPU hours for 70,000 tasks per week. Zhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long, Kai Zhu 0004, Feiwen Zhu, Wenyi Zhao, Jun Yang 0052, Jidong Zhai, Shuaiwen Song, Wei Lin 0016 |
ASPLOS | 7 |
| 2022 | MODIG: integrating multi-omics and multi-dimensional gene network for cancer driver gene identification based on graph attention network modelabstractMOTIVATION: Identifying genes that play a causal role in cancer evolution remains one of the biggest challenges in cancer biology. With the accumulation of high-throughput multi-omics data over decades, it becomes a great challenge to effectively integrate these data into the identification of cancer driver genes. RESULTS: Here, we propose MODIG, a graph attention network (GAT)-based framework to identify cancer driver genes by combining multi-omics pan-cancer data (mutations, copy number variants, gene expression and methylation levels) with multi-dimensional gene networks. First, we established diverse types of gene relationship maps based on protein-protein interactions, gene sequence similarity, KEGG pathway co-occurrence, gene co-expression patterns and gene ontology. Then, we constructed a multi-dimensional gene network consisting of approximately 20 000 genes as nodes and five types of gene associations as multiplex edges. We applied a GAT to model within-dimension interactions to generate a gene representation for each dimension based on this graph. Moreover, we introduced a joint learning module to fuse multiple dimension-specific representations to generate general gene representations. Finally, we used the obtained gene representation to perform a semi-supervised driver gene identification task. The experiment results show that MODIG outperforms the baseline models in terms of area under precision-recall curves and area under the receiver operating characteristic curves. AVAILABILITY AND IMPLEMENTATION: The MODIG program is available at https://github.com/zjupgx/modig. The code and data underlying this article are also available on Zenodo, at https://doi.org/10.5281/zenodo.7057241. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wenyi Zhao, Xun Gu 0002, Jian Wu 0001, Zhan Zhou |
Bioinform. | 1 |
| 2022 | LESSL: Can LEGO sampling and collaborative optimization contribute to self-supervised learning?
Wenyi Zhao, Weidong Zhang 0007, Xipeng Pan, Peixian Zhuang, Xiwang Xie, Lingqiao Li |
Inf. Sci. | 1 |
| 2021 | CanDriS: posterior profiling of cancer-driving sites based on two-component evolutionary modelabstractCurrent cancer genomics databases have accumulated millions of somatic mutations that remain to be further explored. Due to the over-excess mutations unrelated to cancer, the great challenge is to identify somatic mutations that are cancer-driven. Under the notion that carcinogenesis is a form of somatic-cell evolution, we developed a two-component mixture model: while the ground component corresponds to passenger mutations, the rapidly evolving component corresponds to driver mutations. Then, we implemented an empirical Bayesian procedure to calculate the posterior probability of a site being cancer-driven. Based on these, we developed a software CanDriS (Cancer Driver Sites) to profile the potential cancer-driving sites for thousands of tumor samples from the Cancer Genome Atlas and International Cancer Genome Consortium across tumor types and pan-cancer level. As a result, we identified that approximately 1% of the sites have posterior probabilities larger than 0.90 and listed potential cancer-wide and cancer-specific driver mutations. By comprehensively profiling all potential cancer-driving sites, CanDriS greatly enhances our ability to refine our knowledge of the genetic basis of cancer and might guide clinical medication in the upcoming era of precision medicine. The results were displayed in a database CandrisDB (http://biopharm.zju.edu.cn/candrisdb/). Wenyi Zhao, Jingcheng Wu, Guoxing Cai, Jeffrey Haltom, Weijia Su, Michael J. Dong, Jian Wu 0001, Zhan Zhou, Xun Gu 0002 |
Briefings Bioinform. | 1 |
| 2021 | Region- and Pixel-Level Multi-Focus Image Fusion through Convolutional Neural Networks
Wenyi Zhao, Xipeng Pan |
Mob. Networks Appl. | 1 |
| 2021 | S2-aware network for visual recognition
Wenyi Zhao, Xipeng Pan, Lingqiao Li |
Signal Process. Image Commun. | 1 |
| 2020 | Predicting and reining in application-level slowdown on spatial multitasking GPUs
Mengze Wei, Wenyi Zhao, Quan Chen 0002, Jingwen Leng, Chao Li 0009, Wenli Zheng, Minyi Guo |
J. Parallel Distributed Comput. | 2 |
| 2020 | Multi-domain modeling of atrial fibrillation detection with twin attentional convolutional long short-term memory neural networks
Yanrui Jin, Chengjin Qin, Wenyi Zhao, Chengliang Liu 0001 |
Knowl. Based Syst. | 4 |
| 2019 | Themis: Predicting and Reining in Application-Level Slowdown on Spatial Multitasking GPUsabstractPredicting performance degradation of a GPU application when it is co-located with other applications on a spatial multitasking GPU without prior application knowledge is essential in public Clouds. Prior work mainly targets CPU co-location, and is inaccurate and/or inefficient for predicting performance of applications at co-location on spatial multitasking GPUs. Our investigation shows that hardware event statistics caused by co-located applications, which can be collected with negligible overhead, strongly correlate with their slowdowns. Based on this observation, we present Themis, an online slowdown predictor that can precisely and efficiently predict application slowdown without prior application knowledge. We first train a precise slowdown model offline using hardware event statistics collected from representative co-locations. When new applications co-run, Themis collects event statistics and predicts their slowdowns simultaneously. Our evaluation shows that Themis has negligible runtime overhead and can precisely predict application-level slowdown with prediction error smaller than 9.5%. Based on Themis, we also implement an SM allocation engine to rein in application slowdown at co-location. Case studies show that the engine successfully enforces fair sharing and QoS. Wenyi Zhao, Quan Chen 0002, Jingwen Leng, Chao Li 0009, Wenli Zheng, Li Li 0012, Minyi Guo |
IPDPS | 1 |
| 2015 | Surface Oriented Traverse for robust instance detection in RGB-DabstractWe address the problem of robust instance detection in RGB-D image in the presence of noisy data, cluttering, partial occlusion and large pose variation. We extract contour points from the depth image, construct a Surface Oriented Traverse (SOT) feature for each contour point and further classify it as either belonging or not belonging to the instance of interest. Starting from each contour point, its SOT feature is constructed by traversing and uniformly sampling along an oriented geodesic path on the object surface. After classification, all contour points vote for an instance-specific saliency map, from which the instance of interest is finally localized. Compared with the holistic template-based and learning-based methods, our method inherits advantages of the feature-based methods in dealing with cluttering, partial occlusion, and large pose variation. Furthermore, our method does not require accurate 3D models or high quality laser scan data as input and takes noisy data from commodity 3D sensors. Experimental results on the public RGB-D Object Dataset and our FindMe RGB-D Dataset demonstrate the effectiveness and robustness of our proposed instance detection algorithm. Ruizhe Wang 0002, Gérard G. Medioni, Wenyi Zhao |
IROS | 3 |
| 2007 | Image Restoration Under Significant Additive NoiseabstractThe task of deblurring, a form of image restoration, is to recover an image from its blurred version. Whereas most existing methods assume a small amount of additive noise, image restoration under significant additive noise remains an interesting research problem. We describe two techniques to improve the noise handling characteristics of a recently proposed variational framework for semi-blind image deblurring that is based on joint segmentation and deblurring. One technique uses a structure tensor as a robust edge-indicating function. The other uses nonlocal image averaging to suppress noise. We report promising results with these techniques for the case of a known blur kernel Wenyi Zhao, Art Pope |
IEEE Signal Process. Lett. | 1 |
| 2006 | Efficient Scene-based Nonuniformity Correction and EnhancementabstractWe propose a unified framework for scene-based nonuniformity correction (NUC) and enhancement that is required for the FPA-like (focal plane array) sensors to remove fixed-pattern noise and to enhance the image quality. In contrast to existing scene-based NUC methods, the new framework allows us to process image sequences under severe and structured nonuniformity efficiently and to obtain high quality images. We achieve this goal by applying an efficient registration-based method that is bootstrapped by statistical scene-based NUC methods. Specifically, we initialize the whole NUC-enhancement process by applying statistical methods in order to obtain images with quality just enough for image registration. To obtain high-quality images, we integrate the NUC process with super-resolution techniques to reduce noise and enhance resolution. This is achieved by adopting a new imaging model that includes linear NUC model and image blurring and subsampling. Experiments with real data demonstrate the efficacy of the proposed framework. Wenyi Zhao |
ICIP | 1 |
| 2006 | Flexible Image Blending for Image Mosaicing with Reduced ArtifactsabstractImage mosaicing involves geometric alignment among video frames and image compositing or blending. For dynamic mosaicing, image mosaics are constructed dynamically along with incoming video frames. Consequently, dynamic mosaicing demands efficient operations for both alignment and blending in order to achieve real-time performance. In this paper, we focus on efficient image blending methods that create good-quality image mosaics from any number of overlapping frames. One of the driving forces for efficient image processing is the huge market of mobile devices such as cell phones, PDAs that have image sensors and processors. In particular, we show that it is possible to have efficient sequential implementations of blending methods that simultaneously involve all accumulated video frames. The choices of image blending include traditional averaging, overlapping and flexible ones that take into consideration temporal order of video frames and user control inputs. In addition, we show that artifacts due to mis-alignment, image intensity difference can be significantly reduced by efficiently applying weighting functions when blending video frames. These weighting functions are based on pixel locations in a frame, view perspective and temporal order of this frame. One interesting application of flexible blending is to visualize moving objects on a mosaiced stationary background. Finally, to correct for significant exposure difference in video frames, we propose a pyramid extension based on intensity matching of aligned images at the coarsest resolution. Our experiments with real image sequences demonstrate the advantages of the proposed methods. Wenyi Zhao |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2005 | Alignment of Continuous Video onto 3D Point CloudsabstractWe propose a general framework for aligning continuous (oblique) video onto 3D sensor data. We align a point cloud computed from the video onto the point cloud directly obtained from a 3D sensor. This is in contrast to existing techniques where the 2D images are aligned to a 3D model derived from the 3D sensor data. Using point clouds enables the alignment for scenes full of objects that are difficult to model; for example, trees. To compute 3D point clouds from video, motion stereo is used along with a state-of-the-art algorithm for camera pose estimation. Our experiments with real data demonstrate the advantages of the proposed registration algorithm for texturing models in large-scale semiurban environments. The capability to align video before a 3D model is built from the 3D sensor data offers new practical opportunities for 3D modeling. We introduce a novel modeling-through-registration approach that fuses 3D information from both the 3D sensor and the video. Initial experiments with real data illustrate the potential of the proposed approach. Wenyi Zhao, David Nistér, Steven C. Hsu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Alignment of Continuous Video onto 3D Point Clouds
Wenyi Zhao, David Nistér, Steven C. Hsu |
CVPR (2) | 1 |
| 2004 | Motion-based spatial-temporal image repairingabstractIn this paper, we present a motion-based method for image repairing where images with large missing regions can be restored. The key idea of this method is to reconstruct motion fields for regions of missing data by exploring both temporal and spatial information. After the full motion fields are reconstructed, missing regions can be repaired by the replacement of good corresponding regions from other images. To handle general scenes, we employ dense optical flow as our motion model. To compute flow without correspondence (due to missing pixels), we propose an effective multi-resolution, multi-frame flow method. We demonstrate the efficacy of our method using real images. Wenyi Zhao |
ICIP | 1 |
| 2004 | Super-resolution with significant illumination change
Wenyi Zhao |
ICIP | 1 |
| 2002 | Is Super-Resolution with Optical Flow Feasible?
Wenyi Zhao, Harpreet Sawhney |
ECCV (1) | 1 |
| 2001 | Video Georegistration: Algorithm and Quantitative EvaluationabstractAn algorithm is presented for video georegistration, with a particular concern for aerial video, i.e., video captured from an airborne platform. The algorithm's input is a video stream with telemetry (camera model specification sufficient to define an initial estimate of the view) and geodetically calibrated reference imagery (coaligned digital orthoimage and elevation map). The output is a spatial registration of the video to the reference so that it inherits the available geodetic coordinates. The video is processed in a continuous fashion to yield a corresponding stream of georegistered results. Quantitative results of evaluating the developed approach with real world aerial video also are presented. The results suggest that the developed approach may provide valuable input to the analysis and interpretation of aerial video. Richard P. Wildes, David J. Hirvonen, Steven C. Hsu, Rakesh Kumar 0001, W. Brian Lehman, Bogdan Matei, Wenyi Zhao |
ICCV | 7 |
| 2001 | Video Enhancement by Scintillation RemovalabstractAs a comprehensive problem, video enhancement includes many subset problems, including super-resolution, noise reduction, increase of dynamic range, to name a few. In this paper, we focus on a special subset problem -- scintillation removal. Scintillation is one type of image distortions caused by the propagation of visible light through unstable atmosphere. One commonly known source for disturbing atmosphere is heat. We propose an image processing method to remove scintillation in video. This method explores the temporal information redundancy in video and is based on image registration and image fusion. We examine the scintillation problem in detail and explain why we take an image registration plus fusion approach. Finally we demonstrate successful application of scintillation removal on real video clips. Wenyi Zhao, Luca Bogoni, Michael W. Hansen |
ICME | 1 |
| 2001 | Symmetric Shape-from-Shading Using Self-ratio Image
Wenyi Zhao, Rama Chellappa |
Int. J. Comput. Vis. | 1 |
| 2000 | Illumination-Insensitive Face Recognition Using Symmetric Shape-from-ShadingabstractSensitivity to variations in illumination is a fundamental and challenging problem in face recognition. In this paper, we describe a new method based on symmetric shape-from-shading (SSFS) to develop a face recognition system that is robust to changes in illumination. The basic idea of this approach is to use the SSFS algorithm as a tool to obtain a prototype image which is illumination-normalized. It has been shown that the SSFS algorithm has a unique point-wise solution. But it is still difficult to recover accurate shape information given a single real face image with complex shape and varying albedo. In stead, we utilize the fact that all faces share a similar shape making the direct computation of the prototype image from a given face image feasible. Finally, to demonstrate the efficacy of our method, we have applied it to several publicly available face databases. Wenyi Zhao, Rama Chellappa |
CVPR | 1 |
| 2000 | SFS Based View Synthesis for Robust Face RecognitionabstractSensitivity to variations in pose is a challenging problem in face recognition using appearance-based methods. More specifically, the appearance of a face changes dramatically when viewing and/or lighting directions change. Various approaches have been proposed to solve this difficult problem. They can be broadly divided into three classes: (1) multiple image-based methods where multiple images of various poses per person are available; (2) hybrid methods where multiple example images are available during learning but only one database image per person is available during recognition; and (3) single image-based methods where no example-based learning is carried out. We present a method that comes under class 3. This method, based on shape-from-shading (SFS), improves the performance of a face recognition system in handling variations due to pose and illumination via image synthesis. Wenyi Zhao, Rama Chellappa |
FG | 1 |
| 2000 | Robust Image Based Face RecognitionabstractIn face recognition literature, 2D image based approaches are possibly the most promising ones. However, the 2D images/patterns can change dramatically in practice. We first study the performance degradation due to 2D distortions and illumination variations on the input images. We then propose several methods to improve the system performance. Finally experiments are carried out using FERET and other face databases to demonstrate the improvement of one particular system-the subspace LDA system. Wenyi Zhao, Rama Chellappa |
ICIP | 1 |
| 2000 | 3D Model Enhanced Face RecognitionabstractA personal identification system based on the analysis of frontal or profile images of the face has applications in human-computer interfaces, access control, surveillance etc. Among many practical face recognition schemes, image based approaches are possibly the most promising ones. However, the 2D images/patterns of 3D face objects can change dramatically due to lighting and viewing variations. In this paper, we propose a generic 3D model to enhance existing face recognition systems. More specifically, we use a 3D model to synthesize the so-called prototype image from a given image acquired under different lighting and viewing conditions. Wenyi Zhao, Rama Chellappa |
ICIP | 1 |
| 2000 | Performance Perturbation Analysis of Eigen-SystemsabstractEigenproblem has wide range of applications in engineering disciplines, including compression and recognition. When building practical systems for these applications, we are interested in actual system performance. The actual system performance usually differs from the ideal system performance. One way to quantify the actual performance is to measure the ideal performance and the performance degradation/perturbation. In this paper, we present a system performance perturbation analysis particular to eigen-systems based on the matrix perturbation theory. Such performance analysis can provide guidance for designing practically good systems. Wenyi Zhao |
ICPR | 1 |
| 2000 | Discriminant Component Analysis for Face RecognitionabstractWe propose using a feature extraction scheme, discriminant component analysis, for face recognition. This scheme decomposes a signal into orthogonal bases such that for each base there is an eigenvalue representing the discriminatory power of projection in that direction. The bases and eigenvalues are obtained by iteratively applying Fisher's linear discriminant analysis (LDA). We illustrate the motivation of this scheme and show how it can be used to construct new distance metrics for the purpose of enhanced classification. Finally, good performance for face recognition on a dataset of 738 gallery images and 115 probe images is obtained using new distance metrics. Wenyi Zhao |
ICPR | 1 |
| 2000 | A reliable descriptor for face objects in visual content
Wenyi Zhao, Dinkar Bhat, Nagaraj Nandhakumar, Rama Chellappa |
Signal Process. Image Commun. | 1 |
| 1998 | Empirical Performance Analysis of Linear Discriminant ClassifiersabstractIn face recognition literature, holistic template matching systems and geometrical local feature based systems have been pursued. In the holistic approach, PCA (Principal Component Analysis) and LDA (Linear Discriminant Analysis) are popular ones. More recently, the combination of PCA and LDA has been proposed as a superior alternative over pure PCA and LDA. In this paper, we illustrate the rationales behind these methods and the pros and cons of applying them to pattern classification task. A theoretical performance analysis of LDA suggests applying LDA over the principal components from the original signal space or the subspace. The improved performance of this combined approach is demonstrated through experiments conducted on both simulated data and real data. Wenyi Zhao, Rama Chellappa, Nagaraj Nandhakumar |
CVPR | 1 |
| 1998 | Face Similarity Space as Perceived by Humans and Artificial Systems
Peter Kalocsai, Wenyi Zhao, Egor Elagin |
FG | 2 |
| 1998 | Discriminant Analysis of Principal Components for Face Recognition
Wenyi Zhao, Rama Chellappa, Arvind Krishnaswamy |
FG | 1 |
| 1998 | Linear discriminant analysis of MPF for face recognitionabstractIn face recognition literature, major approaches based on holistic templates and geometrical local features have been taken. Both approaches have certain advantages and disadvantages. We explore a method which integrates the above two approaches. Among many specific systems, we select LDA (linear discriminant analysis) and MPF (matching pursuit filter) as the representative from the first type approach and the second type approach respectively. We treat MPF as the feature representation of the original input and LDA as the pattern classifier. We compare the performances of MPF system, LDA system and the hybrid LDA-MPF system for face recognition. Wenyi Zhao, Nagaraj Nandhakumar |
ICPR | 1 |
| 1997 | Model-based interpretation of stereo imagery of textured surfaces
Wenyi Zhao, Nagaraj Nandhakumar, Philip W. Smith |
Mach. Vis. Appl. | 1 |
| 1996 | Effects of camera alignment errors on stereoscopic depth estimates
Wenyi Zhao, Nagaraj Nandhakumar |
Pattern Recognit. | 1 |