VLDB 2026 Research / reviewers in the wild / expert
Yao Sun 0004
dblp:62/6846-4
· DBLP profile ↗
27ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-3429-3476ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11Theory of computation · 9 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Simultaneous Retrieval Algorithm of Water Cloud Optical and Microphysical Properties by High-Spectral-Resolution LidarabstractThe uncertainty of water cloud feedback on radiative forcing is one of the largest obstacles to producing confident projections of the global climate. Sufficient measurements of water clouds are crucial to addressing this issue. However, existing techniques based on remote sensing or in situ instruments face limitations in data capacity attributed to the short lifetime, high temporal variability, and complex vertical structure of water clouds. In this study, taking advantage of a dual-field-of-view (dual-FOV) high-spectral-resolution lidar (HSRL), we developed a novel algorithm to obtain diurnal simultaneous profiles of water cloud optical and microphysical properties with high temporal-spatial resolution. This technique does not rely on the widely used subadiabatic assumption about the vertical structure of water clouds. The retrieval algorithm, validated by simulations and cloud radar measurements, was applied to field experiment data collected at the Beijing and Hangzhou sites in China. The relationship functions between water cloud properties are presented to enhance our understanding of the underlying processes. Furthermore, the vertical distributions of retrieved properties are compared to the subadiabatic assumption. The dual-FOV HSRL technique enables comprehensive observations, enhancing our understanding of water clouds and providing significant insights into the interactions among clouds, aerosols, precipitation, and radiation. Kai Zhang 0062, Lingyun Wu, Daniel Rosenfeld, Detlef Müller, Chengcai Li, Chuanfeng Zhao, Eduardo Landulfo, Cristofer Jimenez, Shuaibo Wang, Xianzhe Hu, Xiaotao Li, Yao Sun 0004, Xueping Wan, Wentai Chen, Jing Li 0052, Yudi Zhou, Zhiji Deng, Zhewei Fu, Weilin Pan, Dong Liu 0020 |
IEEE Trans. Geosci. Remote. Sens. | 13 |
| 2022 | Fine-Grained Human-Centric Tracklet Segmentation with Single Frame SupervisionabstractIn this paper, we target at the Fine-grAined human-Centric Tracklet Segmentation (FACTS) problem, where 12 human parts, e.g., face, pants, left-leg, are segmented. To reduce the heavy and tedious labeling efforts, FACTS requires only one labeled frame per video during training. The small size of human parts and the labeling scarcity makes FACTS very challenging. Considering adjacent frames of videos are continuous and human usually do not change clothes in a short time, we explicitly consider the pixel-level and frame-level context in the proposed Temporal Context segmentation Network (TCNet). On the one hand, optical flow is on-line calculated to propagate the pixel-level segmentation results to neighboring frames. On the other hand, frame-level classification likelihood vectors are also propagated to nearby frames. By fully exploiting the pixel-level and frame-level context, TCNet indirectly uses the large amount of unlabeled frames during training and produces smooth segmentation results during inference. Experimental results on four video datasets show the superiority of TCNet over the state-of-the-arts. The newly annotated datasets can be downloaded via http://liusi-group.com/projects/FACTS for the further studies. Si Liu 0001, Guanghui Ren, Yao Sun 0004, Jinqiao Wang, Changhu Wang, Bo Li 0006, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | On the efficiency of solving Boolean polynomial systems with the characteristic set method
Zhenyu Huang 0004, Yao Sun 0004, Dongdai Lin |
J. Symb. Comput. | 2 |
| 2021 | Algorithms for computing greatest common divisors of parametric multivariate polynomials
Deepak Kapur, Michael B. Monagan, Yao Sun 0004, Dingkang Wang |
J. Symb. Comput. | 4 |
| 2019 | Building Detail-Sensitive Semantic Segmentation Networks With Polynomial PoolingabstractSemantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classification model, invariance to spatial perturbation resulting from the lose of detail-sensitivity prevents segmentation networks from achieving high performance. The use of standard poolings is one of the key factors for this invariance. The most common standard poolings are max and average pooling. Max pooling can increase both the invariance to spatial perturbations and the non-linearity of the networks. Average pooling, on the other hand, is sensitive to spatial perturbations, but is a linear function. For semantic segmentation, we prefer both the preservation of detailed cues within a local feature region and non-linearity that increases a network's functional complexity. In this work, we propose a polynomial pooling (P-pooling) function that finds an intermediate form between max and average pooling to provide an optimally balanced and self-adjusted pooling strategy for semantic segmentation. The P-pooling is differentiable and can be applied into a variety of pre-trained networks. Extensive studies on the PASCAL VOC, Cityscapes and ADE20k datasets demonstrate the superiority of P-pooling over other poolings. Experiments on various network architectures and state-of-the-art training strategies also show that models with P-pooling layers consistently outperform those directly fine-tuned using pre-trained classification models. Zhen Wei 0001, Jingyi Zhang 0005, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Yi Zhou 0007, Si Liu 0001, Yao Sun 0004, Ling Shao 0001 |
CVPR | 8 |
| 2019 | Preimage Attacks on Round-Reduced Keccak-224/256 via an Allocating Approach
Ting Li 0023, Yao Sun 0004 |
EUROCRYPT (3) | 2 |
| 2019 | GPS: Group People Segmentation with Detailed Part InferenceabstractNoticeable progress has been witnessed in general object detection, semantic segmentation and instance segmentation, while parsing a group of people is still a challenging task for human-centric visual understanding due to severe occlusion and various poses. In this paper, we present a new large-scale dataset named “GPS (Group People Segmentation)” to boost academical study and technology development. GPS contains 14000 elaborately annotated images with 20 fine-grained semantic category labels related to human, divided into two sub-datasets corresponding to indoor and outdoor scenes involving various poses, occlusion and background. We further propose a novel GPSNet for group people segmentation. GPSNet consists of a new “Adjusted RoI Align” module to adjust position of detected person and align RoI features, such that the network does not need to fit various positions of each person. A fusion of global and local features is also employed to refine parsing results. Compared with baseline methods, GPSNet achieves the best performance on GPS Dataset. Yue Liao, Si Liu 0001, Tianrui Hui, Chen Gao 0005, Yao Sun 0004, Bo Li 0006 |
ICME | 5 |
| 2019 | Accurate Facial Image Parsing at Real-Time SpeedabstractIn this paper, we propose a design scheme for deep learning networks in face parsing task with promising accuracy and real-time inference speed. By analyzing the differences between general image parsing task and face parsing task, we first revisit the structure of traditional FCN and make improvements to adapt to the unique properties of face parsing task. Especially, the concept of Normalized Receptive Field is proposed to give more insights on designing the network. Then a novel loss function called Statistical Contextual Loss is introduced, which integrates richer contextual information and regularizes features during training. For further model acceleration, we propose a semi-supervised distillation scheme that effectively transfers the learned knowledge to a lighter network. Extensive experiments on LFW and Helen dataset demonstrate the significant superiority of the new design scheme on both efficacy and efficiency. Zhen Wei 0001, Si Liu 0001, Yao Sun 0004 |
IEEE Trans. Image Process. | 3 |
| 2018 | Cross-Domain Human Parsing via Adversarial Feature and Label AdaptationabstractHuman parsing has been extensively studied recently due to its wide applications in many important scenarios. Mainstream fashion parsing models (i.e., parsers) focus on parsing the high-resolution and clean images. However, directly applying the parsers trained on benchmarks of high-quality samples to a particular application scenario in the wild, e.g., a canteen, airport or workplace, often gives non-satisfactory performance due to domain shift. In this paper, we explore a new and challenging cross-domain human parsing problem: taking the benchmark dataset with extensive pixel-wise labeling as the source domain, how to obtain a satisfactory parser on a new target domain without requiring any additional manual labeling? To this end, we propose a novel and efficient cross-domain human parsing model to bridge the cross-domain differences in terms of visual appearance and environment conditions and fully exploit commonalities across domains. Our proposed model explicitly learns a feature compensation network, which is specialized for mitigating the cross-domain differences. A discriminative feature adversarial network is introduced to supervise the feature compensation to effectively reduces the discrepancy between feature distributions of two domains. Besides, our proposed model also introduces a structured label adversarial network to guide the parsing results of the target domain to follow the high-order relationships of the structured labels shared across domains. The proposed framework is end-to-end trainable, practical and scalable in real applications. Extensive experiments are conducted where LIP dataset is the source domain and 4 different datasets including surveillance videos, movies and runway shows without any annotations, are evaluated as target domains. The results consistently confirm data efficiency and performance advantages of the proposed method for the challenging cross-domain human parsing problem. Si Liu 0001, Yao Sun 0004, Defa Zhu, Guanghui Ren, Jiashi Feng, Jizhong Han |
AAAI | 2 |
| 2018 | An Efficient Algorithm for Computing Parametric Multivariate Polynomial GCDabstractA new efficient algorithm for computing a parametric greatest common divisor (GCD) of parametric multivariate polynomials over k[u][x] is presented. The algorithm is based on a well-known simple insight that the GCD of two multivariate polynomials (non-parametric as well as parametric) can be extracted using the generator of the quotient ideal of a polynomial with respect to the second polynomial. And, further, this generator can be obtained by computing a minimal Gröbner basis of the quotient ideal. The main attraction of this idea is that it generalizes to the parametric case for which a comprehensive Gröbner basis is constructed for the parametric quotient ideal. It is proved that in a minimal comprehensive Gröbner system of a parametric quotient ideal, each branch of specializations corresponds to a principal parametric ideal with a single generator. Using this generator, the parametric GCD of that branch is obtained by division. This algorithm does not need to consider whether parametric polynomials are primitive w.r.t. the main variable. This is in sharp contrast to two algorithms recently proposed by Nagasaka (ISSAC, 2017). The resulting algorithm is not only conceptually simple to understand but is considerably efficient. The proposed algorithm and both of Nagasaka's algorithms have been implemented in Singular (available at http://www.mmrc.iss.ac.cn/~dwang/software.html), and their performance is compared on a number of examples. For more than two polynomials, this process can be repeated by considering pairs of polynomials; the efficiency in that case becomes even more evident. Deepak Kapur, Michael B. Monagan, Yao Sun 0004, Dingkang Wang |
ISSAC | 4 |
| 2018 | The lightest 4 × 4 MDS matrices over GL(4, 𝔽2)
Ting Li 0023, Yao Sun 0004, Dingkang Wang, Dongdai Lin |
Sci. China Inf. Sci. | 3 |
| 2018 | Learning adaptive receptive fields for deep image parsing networksabstractIn this paper, we introduce a novel approach to automatically regulate receptive fields in deep image parsing networks. Unlike previous work which placed much importance on obtaining better receptive fields using manually selected dilated convolutional kernels, our approach uses two affine transformation layers in the network's backbone and operates on feature maps. Feature maps are inflated or shrunk by the new layer, thereby changing the receptive fields in the following layers. By use of end-to-end training, the whole framework is data-driven, without laborious manual intervention. The proposed method is generic across datasets and different tasks. We have conducted extensive experiments on both general image parsing tasks, and face parsing tasks as concrete examples, to demonstrate the method's superior ability to regulate over manual designs. Zhen Wei 0001, Yao Sun 0004, Si Liu 0001 |
Comput. Vis. Media | 2 |
| 2018 | Composing Semantic Collage for Image RetargetingabstractImage retargeting has been applied to display images of any size via devices with various resolutions (e.g., cell phone, TV monitors). To fit an image with the target resolution, certain unimportant regions need to be deleted or distorted and the key problem is to determine the importance of each pixel. Existing methods predict pixel-wise importance in a bottom-up manner via eye fixation estimation or saliency detection. In contrast, the proposed algorithm estimates the pixel-wise importance based on a top-down criterion where the target image maintains the semantic meaning of the original image. To this end, several semantic components corresponding to foreground objects, action contexts, and background regions are extracted. The semantic component maps are integrated by a classification guided fusion network. Specifically, the deep network classifies the original image as object or scene-oriented, and fuses the semantic component maps according to classification results. The network output, referred to as the semantic collage with the same size as the original image, is then fed into any existing optimization method to generate the target image. Extensive experiments are carried out on the RetargetMe dataset and S-Retarget database developed in this work. Experimental results demonstrate the merits of the proposed algorithm over the state-of-the-art image retargeting methods. Si Liu 0001, Zhen Wei 0001, Yao Sun 0004, Xinyu Ou, Bin Liu 0014, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Surveillance Video Parsing with Single Frame SupervisionabstractSurveillance video parsing, which segments the video frames into several labels, e.g., face, pants, left-leg, has wide applications [41, 8]. However, pixel-wisely annotating all frames is tedious and inefficient. In this paper, we develop a Single frame Video Parsing (SVP) method which requires only one labeled frame per video in training stage. To parse one particular frame, the video segment preceding the frame is jointly considered. SVP (i) roughly parses the frames within the video segment, (ii) estimates the optical flow between frames and (iii) fuses the rough parsing results warped by optical flow to produce the refined parsing result. The three components of SVP, namely frame parsing, optical flow estimation and temporal fusion are integrated in an end-to-end manner. Experimental results on two surveillance video datasets show the superiority of SVP over state-of-the-arts. The collected video parsing datasets can be downloaded via http://liusi-group.com/projects/SVP for the further studies. Si Liu 0001, Changhu Wang, Ruihe Qian, Renda Bao, Yao Sun 0004 |
CVPR | 6 |
| 2017 | Learning Adaptive Receptive Fields for Deep Image Parsing NetworkabstractIn this paper, we introduce a novel approach to regulate receptive field in deep image parsing network automatically. Unlike previous works which have stressed much importance on obtaining better receptive fields using manually selected dilated convolutional kernels, our approach uses two affine transformation layers in the networks backbone and operates on feature maps. Feature maps will be inflated/shrinked by the new layer and therefore receptive fields in following layers are changed accordingly. By end-to-end training, the whole framework is data-driven without laborious manual intervention. The proposed method is generic across dataset and different tasks. We conduct extensive experiments on both general parsing task and face parsing task as concrete examples to demonstrate the methods superior regulation ability over manual designs. Zhen Wei 0001, Yao Sun 0004, Jinqiao Wang, Hanjiang Lai, Si Liu 0001 |
CVPR | 2 |
| 2017 | On Checking Linear Dependence of Parametric Vectors
Yao Sun 0004, Dingkang Wang, Yushan Xue |
ICIC (2) | 2 |
| 2017 | Face Aging with Contextual Generative Adversarial NetsabstractFace aging, which renders aging faces for an input face, has attracted extensive attention in the multimedia research. Recently, several conditional Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate images fitting the real face distributions conditioned on each individual age group. However, these methods fail to capture the transition patterns, e.g., the gradual shape and texture changes between adjacent age groups. In this paper, we propose a novel Contextual Generative Adversarial Nets (C-GANs) to specifically take it into consideration. The C-GANs consists of a conditional transformation network and two discriminative networks. The conditional transformation network imitates the aging procedure with several specially designed residual blocks. The age discriminative network guides the synthesized face to fit the real conditional distribution. The transition pattern discriminative network is novel, aiming to distinguish the real transition patterns with the fake ones. It serves as an extra regularization term for the conditional transformation network, ensuring the generated image pairs to fit the corresponding real transition pattern distribution. Experimental results demonstrate the proposed framework produces appealing results by comparing with the state-of-the-art and ground truth. We also observe performance gain for cross-age face verification. Si Liu 0001, Yao Sun 0004, Defa Zhu, Renda Bao, Wei Wang 0108, Xiangbo Shu, Shuicheng Yan |
ACM Multimedia | 2 |
| 2017 | Time Traveler: A Real-time Face Aging SystemabstractFace aging, also known as age progression, is attracting more and more research interests. It has plenty of applications in various domains including cross-age face recognition, finding lost children, and entertainments. In recent years, face aging has witnessed various breakthroughs and a number of face aging models have been proposed. Face aging, however, is still a very challenging task in practice for various reasons. First, faces may have many different expressions and lighting conditions, which pose great challenges to modeling the aging patterns. Besides, the training data are usually very limited and the face images for the same person only cover a narrow range of ages. Lejian Ren, Si Liu 0001, Yao Sun 0004, Jian Dong 0011, Luoqi Liu, Shuicheng Yan |
ACM Multimedia | 3 |
| 2017 | RSVP: A Real-Time Surveillance Video Parsing System with Single Frame SupervisionabstractIn this demo, we present a real-time surveillance video parsing (RSVP) system to parse surveillance videos. Surveillance video parsing, which aims to segment the video frames into several labels, e.g., face, pants, left-legs, has wide applications, especially in security filed. However, it is very tedious and time-consuming to annotate all the frames in a video. We design a RSVP system to parse the surveillance videos in real-time. The RSVP system requires only one labeled frame in training stage. The RSVP system jointly considers the segmentation of preceding frames when parsing one particular frame within the video. The RSVP system is proved to be effective and efficient in real applications. Guanghui Ren, Ruihe Qian, Yao Sun 0004, Changhu Wang, Hanqing Lu, Si Liu 0001 |
ACM Multimedia | 4 |
| 2017 | Automated Reducible Geometric Theorem Proving and Discovery by Gröbner Basis Method
Dingkang Wang, Yao Sun 0004 |
J. Autom. Reason. | 3 |
| 2017 | A weakly supervised method for makeup-invariant face verification
Yao Sun 0004, Lejian Ren, Zhen Wei 0001, Bin Liu 0014, Yanlong Zhai, Si Liu 0001 |
Pattern Recognit. | 1 |
| 2013 | An efficient algorithm for computing a comprehensive Gröbner system of a parametric polynomial system
Deepak Kapur, Yao Sun 0004, Dingkang Wang |
J. Symb. Comput. | 2 |
| 2013 | An efficient method for computing comprehensive Gröbner bases
Deepak Kapur, Yao Sun 0004, Dingkang Wang |
J. Symb. Comput. | 2 |
| 2012 | A signature-based algorithm for computing Gröbner bases in solvable polynomial algebrasabstractSignature-based algorithms, including F5, F5C, G2V and GVW, are efficient algorithms for computing Gröbner bases in commutative polynomial rings. In this paper, we present a signature-based algorithm to compute Gröbner bases in solvable polynomial algebras which include usual commutative polynomial rings and some non-commutative polynomial rings like Weyl algebra. The generalized Rewritten Criterion (discussed in Sun and Wang, ISSAC 2011) is used to reject redundant computations. When this new algorithm uses the partial order implied by GVW, its termination is proved without special assumptions on computing orders of critical pairs. Data structures similar to F5 can be used to speed up this new algorithm, and Gröbner bases of syzygy modules of input polynomials can be obtained from the outputs easily. Experimental data show that most redundant computations can be avoided in this new algorithm. Yao Sun 0004, Dingkang Wang |
ISSAC | 1 |
| 2011 | Computing comprehensive Gröbner systems and comprehensive Gröbner bases simultaneouslyabstractIn Kapur et al (ISSAC, 2010), a new method for computing a comprehensive Grobner system of a parameterized polynomial system was proposed and its efficiency over other known methods was effectively demonstrated. Based on those insights, a new approach is proposed for computing a comprehensive Grobner basis of a parameterized polynomial system. The key new idea is not to simplify a polynomial under various specialization of its parameters, but rather keep track in the polynomial, of the power products whose coefficients vanish; this is achieved by partitioning the polynomial into two parts-nonzero part and zero part for the specialization under consideration. During the computation of a comprehensive Grobner system, for a particular branch corresponding to a specialization of parameter values, nonzero parts of the polynomials dictate the computation, i.e., computing S-polynomials as well as for simplifying a polynomial with respect to other polynomials; but the manipulations on the whole polynomials (including their zero parts) are also performed. Grobner basis computations on such pairs of polynomials can also be viewed as Grobner basis computations on a module. Once a comprehensive Grobner system is generated, both nonzero and zero parts of the polynomials are collected from every branch and the result is a faithful comprehensive Grobner basis, to mean that every polynomial in a comprehensive Grobner basis belongs to the ideal of the original parameterized polynomial system. This technique should be applicable to other algorithms for computing a comprehensive Grobner system as well, thus producing both a comprehensive Grobner system as well as a faithful comprehensive Grobner basis of a parameterized polynomial system simultaneously. The approach is exhibited by adapting the recently proposed method for computing a comprehensive Grobner system in (ISSAC, 2010) for computing a comprehensive Grobner basis. The timings on a collection of examples demonstrate that this new algorithm for computing comprehensive Grobner bases has better performance than other existing algorithms. Deepak Kapur, Yao Sun 0004, Dingkang Wang |
ISSAC | 2 |
| 2011 | A generalized criterion for signature related Gröbner basis algorithmsabstractA generalized criterion for signature related algorithms to compute Gröbner basis is proposed in this paper. Signature related algorithms are a popular kind of algorithms for computing Gröbner basis, including the famous F5 algorithm, the F5C algorithm, the extended F5 algorithm and the GVW algorithm. The main purpose of current paper is to study in theory what kind of criteria is correct in signature related algorithms and provide a generalized method to develop new criteria. For this purpose, a generalized criterion is proposed. The generalized criterion only relies on a general partial order defined on a set of polynomials. When specializing the partial order to appropriate specific orders, the generalized criterion can specialize to almost all existing criteria of signature related algorithms. For admissible partial orders, a proof is presented for the correctness of the algorithm that is based on this generalized criterion. And the partial orders implied by the criteria of F5 and GVW are also shown to be admissible in this paper. More importantly, the generalized criterion provides an effective method to check whether a new criterion is correct as well as to develop new criteria for signature related algorithms. Yao Sun 0004, Dingkang Wang |
ISSAC | 1 |
| 2010 | A new algorithm for computing comprehensive Gröbner systemsabstractA new algorithm for computing a comprehensive Gröbner system of a parametric polynomial ideal over k[U][X] is presented. This algorithm generates fewer branches (segments) compared to Suzuki and Sato's algorithm as well as Nabeshima's algorithm, resulting in considerable efficiency. As a result, the algorithm is able to compute comprehensive Gröbner systems of parametric polynomial ideals arising from applications which have been beyond the reach of other well known algorithms. The starting point of the new algorithm is Weispfenning's algorithm with a key insight by Suzuki and Sato who proposed computing first a Gröbner basis of an ideal over k[U,X] before performing any branches based on parametric constraints. Based on Kalkbrener's results about stability and specialization of Gröbner basis of ideals, the proposed algorithm exploits the result that along any branch in a tree corresponding to a comprehensive Gröbner system, it is only necessary to consider one polynomial for each nondivisible leading power product in k(U)[X] with the condition that the product of their leading coefficients is not 0; other branches correspond to the cases where this product is 0. In addition, for dealing with a disequality parametric constraint, a probabilistic check is employed for radical membership test of an ideal of parametric constraints. This is in contrast to a general expensive check based on Rabinovitch's trick using a new variable as in Nabeshima's algorithm. The proposed algorithm has been implemented in Magma and experimented with a number of examples from different applications. Its performance (vis a vie number of branches and execution timings) has been compared with the Suzuki-Sato's algorithm and Nabeshima's speed-up algorithm. The algorithm has been successfully used to solve the famous P3P problem from computer vision. Deepak Kapur, Yao Sun 0004, Dingkang Wang |
ISSAC | 2 |