Yun Guo

dblp:58/8480 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorArtificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Security and privacy · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Image and video processing · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image restoration
image deraining
1.422024
Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2024
From Sky to the Ground: A Large-scale Benchmark and Simple Baseline Towards Real Rain Removal · ICCV 2023
Image and video processing
image restoration
1.422024
Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2024
From Sky to the Ground: A Large-scale Benchmark and Simple Baseline Towards Real Rain Removal · ICCV 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Representation and self-supervised learning › contrastive learning
self-supervised contrastive learning
0.812024
Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Software testing
combinatorial testing
0.212016
Applying combinatorial test data generation to big data applications · ASE 2016
Software testing
test input generation
0.212016
Applying combinatorial test data generation to big data applications · ASE 2016

Methods — techniques the papers use, named apart from their topics

asymmetric contrastive loss · 1.5nonlocal self-similarity · 0.8non-local self-similarity · 0.8transformer · 0.7self-attention · 0.7low-rank tensor recovery · 0.7cross-layer attention · 0.7combinatorial testing · 0.5
YearPublicationVenuePosition
2024 Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-Similarity
abstract
Most existing learning-based deraining methods are supervisedly trained on synthetic rainy-clean pairs. The domain gap between the synthetic and real rain makes them less generalized to complex real rainy scenes. Moreover, the existing methods mainly utilize the property of the image or rain layers independently, while few of them have considered their mutually exclusive relationship. To solve above dilemma, we explore the intrinsic intra-similarity within each layer and inter-exclusiveness between two layers and propose an unsupervised non-local contrastive learning (NLCL) deraining method. The non-local self-similarity image patches as the positives are tightly pulled together and rain patches as the negatives are remarkably pushed away, and vice versa. On one hand, the intrinsic self-similarity knowledge within positive/negative samples of each layer benefits us to discover more compact representation; on the other hand, the mutually exclusive property between the two layers enriches the discriminative decomposition. Thus, the internal self-similarity within each layer (similarity) and the external exclusive relationship of the two layers (dissimilarity) serving as a generic image prior jointly facilitate us to unsupervisedly differentiate the rain from clean image. We further discover that the intrinsic dimension of the non-local image patches is generally higher than that of the rain patches. This insight motivates us to design an asymmetric contrastive loss that precisely models the compactness discrepancy of the two layers, thereby improving the discriminative decomposition. In addition, recognizing the limited quality of existing real rain datasets, which are often small-scale or obtained from the internet, we collect a large-scale real dataset under various rainy weathers that contains high-resolution rainy images. Extensive experiments conducted on different real rainy datasets demonstrate that the proposed method obtains state-of-the-art performance in real deraining.
Yi Chang 0002, Yun Guo, Yuntong Ye, Changfeng Yu, Lin Zhu 0012, Xi-Le Zhao, Luxin Yan, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 From Sky to the Ground: A Large-scale Benchmark and Simple Baseline Towards Real Rain Removal
abstract
Learning-based image deraining methods have made great progress. However, the lack of large-scale high-quality paired training samples is the main bottleneck to hamper the real image deraining (RID). To address this dilemma and advance RID, we construct a Large-scale High-quality Paired real rain benchmark (LHP-Rain), including 3000 video sequences with 1 million high-resolution (1920*1080) frame pairs. The advantages of the proposed dataset over the existing ones are three-fold: rain with higher-diversity and larger-scale, image with higher-resolution and higher-quality ground-truth. Specifically, the real rains in LHP-Rain not only contain the classical rain streak/veiling/occlusion in the sky, but also the splashing on the ground overlooked by deraining community. Moreover, we propose a novel robust low-rank tensor recovery model to generate the GT with better separating the static background from the dynamic rain. In addition, we design a simple transformer-based single image deraining baseline, which simultaneously utilize the self-attention and cross-layer attention within the image and rain layer with discriminative feature representation. Extensive experiments verify the superiority of the proposed dataset and deraining method over state-of-the-art.
Yun Guo, Xueyao Xiao, Yi Chang 0002, Shumin Deng, Luxin Yan
ICCV1
2023 Mutual Query Network for Multi-Modal Product Image Segmentation
abstract
Product image segmentation is vital in e-commerce. Most existing methods extract the product image foreground only based on the visual modality, making it difficult to distinguish irrelevant products. As product titles contain abundant appearance information and provide complementary cues for product image segmentation, we propose a mutual query network to segment products based on both visual and linguistic modalities. First, we design a language query vision module to obtain the response of language description in image areas, thus aligning the visual and linguistic representations across modalities. Then, a vision query language module utilizes the correlation between visual and linguistic modalities to filter the product title and effectively suppress the content irrelevant to the vision in the title. To promote the research in this field, we also construct a Multi-Modal Product Segmentation dataset (MMPS), which contains 30,000 images and corresponding titles. The proposed method significantly outperforms the state-of-the-art methods on MMPS.
Yun Guo, Xiancong Ren, Jingjing Lv, Xin Zhu 0008, Zhangang Lin, Jingping Shao
ICME1
2020 A Movie Recommendation System Based on Differential Privacy Protection
abstract
In the past decades, the ever-increasing popularity of the Internet has led to an explosive growth of information, which has consequently led to the emergence of recommendation systems. A series of cloud-based encryption measures have been adopted in the current recommendation systems to protect users’ privacy. However, there are still many other privacy attacks on the local devices. Therefore, this paper studies the encryption interference of applying a differential privacy protection scheme on the data in the user’s local devices under the assumption of an untrusted server. A dynamic privacy budget allocation method is proposed based on a localized differential privacy protection scheme while taking the specific application scene of movie recommendation into consideration. What is more, an improved user-based collaborative filtering algorithm, which adopts a matrix-based similarity calculation method instead of the traditional vector-based method when computing the user similarity, is proposed. Finally, it was proved by experimental results that the differential privacy-based movie recommendation system (DP-MRE) proposed in this paper could not only protect the privacy of users but also ensure the accuracy of recommendations.
Min Li 0045, Yingming Zeng, Yun Guo
Secur. Commun. Networks4
2019 Exoneration-based fault localization for SQL predicates
Yun Guo, Nan Li 0008, A. Jefferson Offutt, Amihai Motro
J. Syst. Softw.1
2018 An Integrative Analysis of Time-varying Regulatory Networks From High-dimensional Data
abstract
Directed networks have been widely used to describe many biological processes and functions. Understanding the structure of biological networks, especially regulatory networks, could help discover the mechanisms underlying important biological processes and pathogenesis of diseases. Most network inference methods assume the network structure is time-invariant or stationary. However, in some processes, the network structure is non-stationary or time-varying. The stationary network inference methods might not be able to directly used to reconstruct time-varying networks. Some non-stationary network learning methods have been proposed to infer the networks, but, the inferred networks are not regulatory networks which require activation and inhibition information. This work proposes an integrative approach, which combines the changepoint estimation, weighted network learning and searching, and model checking technique, to reconstruct time varying regulatory networks from high-dimensional time series data. We illustrate this approach to study the structure changes of Drosophila's regulatory networks in its life cycle.
Yun Guo, Haijun Gong
IEEE BigData2
2018 Examine Manipulated Datasets with Topology Data Analysis: A Case Study
Yun Guo, Daniel Sun 0004, Guoqiang Li 0001, Shiping Chen 0001
ICICS1
2018 Automatically Repairing SQL Faults
abstract
SQL is the standard database language, yet SQL statements can be complex and expensive to debug by hand. Automatic program repair techniques have the potential to reduce cost significantly. A previous attempt to repair SQL faults automatically used a decision tree (DT) algorithm that succeeded in some cases, but also generated many patches that passed the automated tests but that were not acceptable to the engineers. This paper proposes a novel fault localization and repair technique to repair faulty SQL statements. It targets faults in two common SQL constructs, JOIN and WHERE. It identifies the fault location and type precisely, and then creates a patch to fix the fault. We implemented this technique in a tool, and evaluated it on five medium to large-scale databases using 825 faulty queries with various complexity and faulty types. Experimental results showed that this technique can identify and repair JOIN faults when the DT approach is infeasible, and repair WHERE faults at about the same rate as the DT approach. Moreover, patches generated by our approach are more acceptable to engineers, and the tool is much faster.
Yun Guo, Nan Li 0008, A. Jefferson Offutt, Amihai Motro
QRS1
2017 Localizing and Fixing Faults in SQL Predicates
Yun Guo
ICST1
2017 Localizing Faults in SQL Predicates
abstract
Fault localization techniques have been applied to database and data-centric applications that use SQL or SQL-based languages. However, existing techniques can only identify the SQL statements that have faults, but not determine the precise location of the faults within SQL statements. Since SQL statements can be rather complex, programmers are still left with a difficult repair chore. We propose a novel fault localization method to localize multiple types of faults in SQL predicates, that is based on row-based dynamic slicing and delta debugging. Our method was implemented in a tool called ALTAR, and experiments were performed on two publicly available databases. Our method can be compared with existing fault localization techniques when these are applied to "drill-down" in SQL statements. The results showed that ALTAR can discover more types of faults. Moreover, for the type of faults discovered by current methods, ALTAR is more precise.
Yun Guo, Amihai Motro, Nan Li 0008
ICST1
2016 Applying combinatorial test data generation to big data applications
abstract
Big data applications (e.g., Extract, Transform, and Load (ETL) applications) are designed to handle great volumes of data. However, processing such great volumes of data is time-consuming. There is a need to construct small yet effective test data sets during agile development of big data applications.
Nan Li 0008, Yu Lei 0001, Haider Riaz Khan, Jingshu Liu, Yun Guo
ASE5
2015 A Scalable Big Data Test Framework
abstract
This paper identifies three problems when testing software that uses Hadoop-based big data techniques. First, processing big data takes a long time. Second, big data is transferred and transformed among many services. Do we need to validate the data at every transition point? Third, how should we validate the transferred and transformed data? We are developing a novel big data test framework to address these problems. The test framework generates a small and representative data set from an original large data set using input space partition testing. Using this data set for development and testing would not hinder the continuous integration and delivery when using agile processes. The test framework also accesses and validates data at various transition points when data is transferred and transformed.
Nan Li 0008, Anthony Escalona, Yun Guo, A. Jefferson Offutt
ICST3
2015 Robust Dynamic Background Model with Adaptive Region Based on T2FS and GMM
abstract
For many tracking and surveillance applications, Gaussian mixture model (GMM) provides an effective mean to segment the foreground from background. Though, because of insufficient and noisy data in complex dynamic scenes, the estimated parameters of the GMM, which are based on the assumption that the pixel process meets multi-modal Gaussian distribution, may not accurately reflect the underlying distribution of the observations. And the existing block-based GMM (BGMM) method may be able to segment only rough foreground objects with time-consuming calculations. To solve these difficulties, this paper proposes to use type-2 fuzzy sets (T2FSs) to handle GMM’s uncertain parameters (T2GMM). Furthermore, this paper also introduces a novel representation of contextual spatial information including the color, edge and texture features for each block which is faster and almost lossless (T2BGMM). Experimental results demonstrate the efficiency of the proposed methods.
Yun Guo, Yi Ji 0001, Jutao Zhang, Shengrong Gong, Chunping Liu
KSEM1
2014 Autonomy in collaborative manufacturing networks
abstract
A collaborative manufacturing network is an alliance of business entities (mostly manufacturers and suppliers) who collaborate on the production of complex products. An essential element of this type of collaboration is that the participating entities are allowed to retain some measure of autonomy
Yun Guo, Amihai Motro
CollaborateCom1
2014 Dynamic analysis and modeling of Forest above-ground biomass
abstract
Estimating forest above-ground biomass (AGB) and monitoring its variation are relevant for sustainable forest management, monitoring global change, carbon accounting, particularly for the Qilian Mountains (QMs), a water resource protection zone. In this work, the results of above-ground biomass (AGB) estimates from Landsat Thematic Mapper 5 (TM) images and field data from the fragmented landscape of the upper reaches of the Heihe River Basin (HRB), located in the Qilian Mountains of Gansu province in northwest China, are presented. An optimized k-Nearest Neighbor (k-NN) method was determined by varying both the mathematical formulation of the algorithm and remote sensing data input which resulted in 3,000 different model configurations. Following the sun-canopy-sensor plus C (SCS+C) topographic correction, performance of the optimized k-NN method was satisfied (R2=0.59, RMSE=24.92 ton/ha) which indicated that the optimized k-NN is capable of operational applications of forest AGB estimates in regions where only a few inventory data are available. Afterwards, the calibrated BIOME-BGC was applied to simulate the carbon fluxes over QMs forests with satisfactory accuracy. Finally, the dynamic analysis and modeling of forest AGB was conducted based on the remotely sensed estimation of forest AGB and the annual forest AGB increment from the ecological process model.
Xin Tian 0005, Zengyuan Li, Yun Guo, Erxue Chen, Zhongbo Su, Christiaan van der Tol, Feilong Ling
IGARSS3
2014 Comparison of estimating forest above-ground biomass over montane area by two non-parametric methods
abstract
Forest biomass reflects the ecological succession and human disturbance of the forest, and can fully embody the quality of forest ecosystem environment. The Qilian Mountain forest reserve at upper reaches of the Heihe River Basin was selected for the study. Landsat Thematic Mapper 5 (TM) images were selected as the source data, which were rectified by SCS + C terrain radiometric correction. Forest above-ground biomass was estimated using k-nearest neighbor (k-NN) method and support vector regression (SVR) method, respectively. The results show that spectral information of remote sensing image was recovered by the sun-canopy-sensor plus the C (SCS+C) terrain correction which can effectively improve the estimation accuracy of the models regardless of k-NN or SVR. The optimal k-NN method (R2=0.54, RMSE=26.62ton/ha) performs better than the optimal SVR method (R2=0.51, RMSE=27.45ton/ha).
Yun Guo, Xin Tian 0005, Zengyuan Li, Feilong Ling, Erxue Chen
IGARSS1
2014 Simulation of carbon flux of forest ecosystem by Biome-BGC and MODIS-PSN models
abstract
An approach was used to incorporate the forest carbon flux for Qilian Mountains by ecological-process-based model (Biome-BGC), and remote-sensing-based model (MODIS-PSN). The calibration phase, aiming at setting the ecophysiological parameters to effectively simulate the daily GPP behavior of the Qilian Mountains, was proceeded by adjusting the 8 day GPP outputs obtained from Biome-BGC using the optimized MODIS-PSN algorithm and the observations. The results showed that the optimized MODIS-PSN could describe the GPP behavior faithfully comparing to the eddy covariance-observed GPPs, with R2=0.77, RMSE=6.219gC/m2/8d. After validation, the calibrated Biome-BGC has been proved to estimate the performances of daily GPP behavior effectively comparing to the eddy covariance-observed GPPs, especially in summer and winter (with R2= 0.76, RMSE= 1.3115gC/m2/d), which illustrated that the combination of Biome-BGC and optimized MODIS-PSN could express the carbon fluxes well over the Qilian Mountains.
Zengyuan Li, Xin Tian 0005, Erxue Chen, Wangfei Zhang, Yun Guo
IGARSS6
2012 The SOAVE Platform: A Service-Oriented Architecture for Virtual Enterprises
Amihai Motro, Yun Guo
PRO-VE2
2010 A Novel IRC Botnet Detection Method Based on Packet Size Sequence
abstract
Botnets have become a serious threat to Internet and are often deployed to control a large pool of zombies and perform notorious activities such as DDoS, information theft and spam sending. In this paper, a new method is developed for detecting IRC botnets by analyzing the characteristic of packet size sequence of the TCP conversation between IRC zombies and their command and control (C&C) servers. In comparison with IRC chat, the TCP conversations within IRC botnets show a nature of approximate periodicity defined as quasi-periodicity in this paper. A simple yet effective detection method is presented to detect IRC botnets by measuring the quasi-periodicity degree and packet average size of IRC conversations based on ukkonen algorithm. We evaluated our method using real-world IRC botnet traces captured from honeynet. The results show that our method can detect real-world IRC botnets from IRC traffic with high accuracy and has a low false positive rate.
Xiaobo Ma 0001, Xiaohong Guan, Yun Guo
ICC5
2007 An Algorithm for Learning Principal Curves with Principal Component Analysis and Back-Propagation Network
abstract
A new algorithm for learning principal curves with definite mathematical representations is proposed based on combining the principal component analysis (PCA) and back-propagation (BP) network. The algorithm successfully turns an unsupervised learning problem into a supervised one by projecting a data set to its first component line and identifying the relation between the data points and their corresponding projection indices with BP network. This algorithm has been proved distinctly superior to the HS algorithm.
Yihuai Wang, Yun Guo, Y. C. Fu, Zu-yi Shen
ISDA2