VLDB 2026 Research / reviewers in the wild / expert
Ying Wang 0008
dblp:94/3104-8
· DBLP profile ↗
39ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 22 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance FlowabstractRectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive computational overhead, resulting in prolonged inference times. We propose LookFlow, a training-free high-resolution synthesis framework that accelerates inference while preserving visual quality. Building on pretrained text-to-image RFTs, LookFlow employs a dynamic lookahead guidance flow mechanism to refine high-resolution velocity predictions by leveraging multi-timestep lookahead information extracted from a low-resolution flow. Additionally, reusing temporally similar features across consecutive timesteps drastically reduces computation and significantly decreases inference time overhead. Extensive experiments on COCO demonstrate that LookFlow robustly scales resolutions from 4× to 25×, achieving up to a maximum speedup of 2.01× while maintaining competitive visual fidelity. Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang |
AAAI | 5 |
| 2026 | Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language ModelsabstractDespite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluation gap, we introduce EmojiGrid, a novel diagnostic benchmark specifically designed to probe these fine-grained and higher-order skills. Leveraging the universal and semantically rich nature of emojis, we synthesize a grid‑based visual dataset paired with 29,000+ QA pairs. Each pair is explicitly anchored in a three-level cognitive taxonomy comprising (i) Perception and Information Extraction, (ii) Relational and Structural Reasoning, and (iii) Abstraction and Advanced Cognition. These dimensions further decompose into nine categories covering a broad range of cognitive skills, including counting, spatial relations, compositional logic, semantic sentiment, and related higher-order reasoning tasks. Our extensive evaluation of 25 state-of-the-art open-source and proprietary VLMs reveals a significant performance gap between foundational perceptual tasks and higher-level cognitive abilities, particularly in abstraction and advanced emotional reasoning. Notably, all models struggle with compositional logic, spatial consistency, and especially emotional and semantic understanding. EmojiGrid provides a quantifiable, fine-grained benchmark to diagnose VLM limitations and guides future progress toward models that can truly perceive, reason about, and interpret complex, symbol-rich visual scenes. Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang |
AAAI | 5 |
| 2026 | AtmosOceanNet: SST forecast method driven by atmospheric-oceanic multimodal data
Qinxuan Wang, Kun Ding 0001, Ying Wang 0008, Yineng Li, Shiming Xiang, Xiaoqing Chu |
Pattern Recognit. | 4 |
| 2025 | EvoVLMA: Evolutionary Vision-Language Model Adaptation
Kun Ding 0001, Ying Wang 0008, Shiming Xiang |
ACM Multimedia | 2 |
| 2025 | Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and FilteringabstractThe task of Knowlegde-Based Visual Question Answering (KB-VQA) requires the model to understand visual features and retrieve external knowledge. Retrieval-Augmented Generation (RAG) have been employed to address this problem through knowledge base querying. However, existing work demonstrate two limitations: insufficient interactivity during knowledge retrieval and ineffective organization of retrieved information for Visual-Language Model (VLM). To address these challenges, we propose a three-stage visual language model with Process, Retrieve and Filter (VLM-PRF) framework. For interactive retrieval, VLM-PRF uses reinforcement learning (RL) to guide the model to strategically process information via tool-driven operations. For knowledge filtering, our method trains the VLM to transform the raw retrieved information into into task-specific knowledge. With a dual reward as supervisory signals, VLM-PRF successfully enable model to optimize retrieval strategies and answer generation capabilities simultaneously. Experiments on two datasets demonstrate the effectiveness of our framework. Yuyang Hong, Qi Yang 0015, Lubin Fan, Ying Wang 0008, Kun Ding 0001, Shiming Xiang, Jieping Ye |
NeurIPS | 6 |
| 2025 | HAN: An efficient hierarchical self-attention network for skeleton-based gesture recognition
Ying Wang 0008, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 2 |
| 2025 | Transformer with token attention and attribute prediction for image captioning
Lifei Song, Ying Wang 0008, Linsu Shi, Jiazhong Yu, Shiming Xiang |
Pattern Recognit. Lett. | 2 |
| 2024 | Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningabstractWe propose a generalized method for boosting the generalization ability of pre-trained vision-language models (VLMs) while fine-tuning on downstream few-shot tasks. The idea is realized by exploiting out-of-distribution (OOD) detection to predict whether a sample belongs to a base distribution or a novel distribution and then using the score generated by a dedicated competition based scoring function to fuse the zero-shot and few-shot classifier. The fused classifier is dynamic, which will bias towards the zero-shot classifier if a sample is more likely from the distribution pre-trained on, leading to improved base-to-novel generalization ability. Our method is performed only in test stage, which is applicable to boost existing methods without time-consuming re-training. Extensive experiments show that even weak distribution detectors can still improve VLMs' generalization ability. Specifically, with the help of OOD detectors, the harmonic mean of CoOp and ProGrad increase by 2.6 and 1.5 percentage points over 11 recognition datasets in the base-to-novel setting. Kun Ding 0001, Haojian Zhang, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 4 |
| 2024 | Compositional Kronecker Context Optimization for vision-language models
Kun Ding 0001, Ying Wang 0008, Haojian Zhang, Shiming Xiang |
Neurocomputing | 4 |
| 2024 | Multi-task prompt tuning with soft context sharing for vision-language models
Kun Ding 0001, Ying Wang 0008, Pengzhang Liu, Haojian Zhang, Shiming Xiang, Chunhong Pan |
Neurocomputing | 2 |
| 2024 | Image captioning: Semantic selection unit with stacked residual attention
Lifei Song, Ying Wang 0008, Yuanhua Wang, Shiming Xiang |
Image Vis. Comput. | 3 |
| 2024 | Preformer: Simple and Efficient Design for Precipitation Nowcasting With TransformersabstractThe primary objective of precipitation nowcasting is to predict precipitation patterns several hours in advance. Recent studies have emphasized the potential of deep learning methods for this task. To harness the correlations among various meteorological elements, existing frameworks project multiple meteorological elements into a latent space and then utilize convolutional-recurrent networks for future precipitation prediction. Although effective, the escalating model complexity may impede practical applications. This letter develops the Preformer, a streamlined Transformer framework for precipitation nowcasting that efficiently captures global spatiotemporal dependencies among multiple meteorological elements. The Preformer implements an encoder-translator-decoder architecture, where the encoder integrates spatial features of multiple elements, the translator models spatiotemporal dynamics, and the decoder combines spatiotemporal information to forecast future precipitation. Without introducing complex structures or strategies, the Preformer achieves state-of-the-art performance even with the least parameters. Qizhao Jin, Xinbang Zhang, Xinyu Xiao, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Knowledge Mining and Transferring for Domain Adaptive Object DetectionabstractWith the thriving of deep learning, CNN-based object detectors have made great progress in the past decade. However, the domain gap between training and testing data leads to a prominent performance degradation and thus hinders their application in the real world. To alleviate this problem, Knowledge Transfer Network (KTNet) is proposed as a new paradigm for domain adaption. Specifically, KT-Net is constructed on a base detector with intrinsic knowledge mining and relational knowledge constraints. First, we design a foreground/background classifier shared by source domain and target domain to extract the common attribute knowledge of objects in different scenarios. Second, we model the relational knowledge graph and explicitly constrain the consistency of category correlation under source domain, target domain, as well as cross-domain conditions. As a result, the detector is guided to learn object-related and domain-independent representation. Extensive experiments and visualizations confirm that transferring object-specific knowledge can yield notable performance gains. The proposed KTNet achieves state-of-the-art results on three cross-domain detection benchmarks. Chenghao Zhang 0003, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
ICCV | 3 |
| 2020 | Decoupled Representation Learning for Skeleton-Based Gesture RecognitionabstractSkeleton-based gesture recognition is very challenging, as the high-level information in gesture is expressed by a sequence of complexly composite motions. Previous works often learn all the motions with a single model. In this paper, we propose to decouple the gesture into hand posture variations and hand movements, which are then modeled separately. For the former, the skeleton sequence is embedded into a 3D hand posture evolution volume (HPEV) to represent fine-grained posture variations. For the latter, the shifts of hand center and fingertips are arranged as a 2D hand movement map (HMM) to capture holistic movements. To learn from the two inhomogeneous representations for gesture recognition, we propose an end-to-end two-stream network. The HPEV stream integrates both spatial layout and temporal evolution information of hand postures by a dedicated 3D CNN, while the HMM stream develops an efficient 2D CNN to extract hand movement features. Eventually, the predictions of the two streams are aggregated with high efficiency. Extensive experiments on SHREC'17 Track, DHG-14/28 and FPHA datasets demonstrate that our method is competitive with the state-of-the-art. Yongcheng Liu, Ying Wang 0008, Véronique Prinet, Shiming Xiang, Chunhong Pan |
CVPR | 3 |
| 2020 | 3D PostureNet: A unified framework for skeleton-based posture recognition
Ying Wang 0008, Yongcheng Liu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 2 |
| 2019 | Incremental Poisson Surface Reconstruction for Large Scale Three-Dimensional Modeling
Wei Sui, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
PRCV (3) | 3 |
| 2018 | Fast Variational Level Set Based Image Segmentation via Two-Scale Filtering ModelabstractOne major difficulty in medical image segmentation is intensity inhomogeneity, which manifests itself with a slow intensity variation over the whole image domain. Recently, a local binary fitting (LBF) model has been proposed to solve this problem within level set segmentation framework. However, the LBF model has two main problems, i.e., high computational cost and sensitivity to initialization. By analyzing the LBF model, we find that the most computational part is the calculation of two cluster images, which need to be updated in each iteration during the evolution of level set function. With this observation in mind, we propose a novel two-scale filtering (TSF) model, in which the two cluster images can be pre-calculated before evolution. Additionally, we implicitly utilize order constraint to restrict the order of two cluster images. As a result, the proposed TSF model is less sensitive to initialization. Extensive experiments on real medical images illustrate the desirable performances, as compared with the state-of-the-art models. Lingfeng Wang 0002, Ying Wang 0008, Chunhong Pan |
ICASSP | 2 |
| 2017 | Automatic Road Detection and Centerline Extraction via Cascaded End-to-End Convolutional Neural NetworkabstractAccurate road detection and centerline extraction from very high resolution (VHR) remote sensing imagery are of central importance in a wide range of applications. Due to the complex backgrounds and occlusions of trees and cars, most road detection methods bring in the heterogeneous segments; besides for the centerline extraction task, most current approaches fail to extract a wonderful centerline network that appears smooth, complete, as well as single-pixel width. To address the above-mentioned complex issues, we propose a novel deep model, i.e., a cascaded end-to-end convolutional neural network (CasNet), to simultaneously cope with the road detection and centerline extraction tasks. Specifically, CasNet consists of two networks. One aims at the road detection task, whose strong representation ability is well able to tackle the complex backgrounds and occlusions of trees and cars. The other is cascaded to the former one, making full use of the feature maps produced formerly, to obtain the good centerline extraction. Finally, a thinning algorithm is proposed to obtain smooth, complete, and single-pixel width road centerline network. Extensive experiments demonstrate that CasNet outperforms the state-of-the-art methods greatly in learning quality and learning speed. That is, CasNet exceeds the comparing methods by a large margin in quantitative performance, and it is nearly 25 times faster than the comparing methods. Moreover, as another contribution, a large and challenging road centerline data set for the VHR remote sensing image will be publicly available for further studies. Ying Wang 0008, Shibiao Xu, Hongzhen Wang, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Learning-based fully 3D face reconstruction from a single imageabstractThis paper presents an algorithm for fully reconstructing a 3D face from a single image. This task is still highly challenging as most current methods only care about the frontal face, ignoring side face, such as the neck, ears etc. In our algorithm, to get the more detailed texture, we deal with the shape reconstruction and texture recovery respectively. For shape, we estimate the deformation of the 3D model by a set of feature points. For texture, due to the similar facial structure, we divide the full texture into patches and show how sparse learning model can be used to fully recover the texture of the 3D face. Extensive experiment results on the CMU-PIE database and images downloaded from the Internet demonstrate that our method outperforms the state-of-the-art methods. Ying Wang 0008, Feiyun Zhu, Chunhong Pan |
ICASSP | 2 |
| 2016 | Accurate urban road centerline extraction from VHR imagery via multiscale segmentation and tensor voting
Feiyun Zhu, Shiming Xiang, Ying Wang 0008, Chunhong Pan |
Neurocomputing | 4 |
| 2015 | 10, 000+ Times Accelerated Robust Subset SelectionabstractSubset selection from massive data with noised information is increasingly popular for various applications. This problem is still highly challenging as current methods are generally slow in speed and sensitive to outliers. To address the above two issues, we propose an accelerated robust subset selection (ARSS) method. Extensive experiments on ten benchmark datasets verify that our method not only outperforms state of the art methods, but also runs 10,000+ times faster than the most related method. Feiyun Zhu, Bin Fan 0001, Xinliang Zhu, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 4 |
| 2015 | Road extraction via adaptive graph cuts with multiple featuresabstractAccurate road extraction from complex backgrounds plays a fundamental role in a wide range of remote sensing applications. There are two shortcomings for the existing methods: 1) Most of them ignore the spatially contextual information inherent in images; 2) Few existing methods show robustness to the occlusions of cars or trees. To address these two problems, we propose a novel approach via adaptive graph cuts with multiple features. Specifically, for the former problem, we apply multiple features (spectral feature, spatial feature and gradient feature) to obtain not only the spectral characteristic but also the spatially contextual feature. In this way, the structural information of road network can be effectively captured. For the latter, adaptive graph cuts based algorithm is adopted. These two schemes show better performance than state-of-the-art methods under the conditions of occlusions. Experiments on 25 images indicate the validity and effectiveness of our method by comparing with state-of-the-art approaches. Ying Wang 0008, Feiyun Zhu, Chunhong Pan |
ICIP | 2 |
| 2015 | Image Deblurring with Coupled Dictionary Learning
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang |
Int. J. Comput. Vis. | 3 |
| 2015 | Robust Hyperspectral Unmixing With Correntropy-Based MetricabstractHyperspectral unmixing is one of the crucial steps for many hyperspectral applications. The problem of hyperspectral unmixing has proved to be a difficult task in unsupervised work settings where the endmembers and abundances are both unknown. In addition, this task becomes more challenging in the case that the spectral bands are degraded by noise. This paper presents a robust model for unsupervised hyperspectral unmixing. Specifically, our model is developed with the correntropy-based metric where the nonnegative constraints on both endmembers and abundances are imposed to keep physical significance. Besides, a sparsity prior is explicitly formulated to constrain the distribution of the abundances of each endmember. To solve our model, a half-quadratic optimization technique is developed to convert the original complex optimization problem into an iteratively reweighted nonnegative matrix factorization with sparsity constraints. As a result, the optimization of our model can adaptively assign small weights to noisy bands and put more emphasis on noise-free bands. In addition, with sparsity constraints, our model can naturally generate sparse abundances. Experiments on synthetic and real data demonstrate the effectiveness of our model in comparison to the related state-of-the-art unmixing models. Ying Wang 0008, Chunhong Pan, Shiming Xiang, Feiyun Zhu |
IEEE Trans. Image Process. | 1 |
| 2014 | Active Flattening of Curved Document Images via Two Structured BeamsabstractDocument images captured by a digital camera often suffer from serious geometric distortions. In this paper, we propose an active method to correct geometric distortions in a camera-captured document image. Unlike many passive rectification methods that rely on text-lines or features extracted from images, our method uses two structured beams illuminating upon the document page to recover two spatial curves. A developable surface is then interpolated to the curves by finding the correspondence between them. The developable surface is finally flattened onto a plane by solving a system of ordinary differential equations. Our method is a content independent approach and can restore a corrected document image of high accuracy with undistorted contents. Experimental results on a variety of real-captured document images demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Ying Wang 0008, Shenquan Qu, Shiming Xiang, Chunhong Pan |
CVPR | 2 |
| 2014 | Urban road extraction via graph cuts based probability propagationabstractIn this paper, we propose a graph cuts (GC) based probability propagation approach to automatically extract road network from complex remote sensing images. First, the support vector machine (SVM) classifier with a sigmoid model is applied to assign each pixel a posterior probability of being labelled as road class, which avoids the weaknesses of hard labels in general SVM. Then a GC based probability propagation algorithm is employed to keep the extracted road results smooth and coherent, which can reduce the connections between roads and road-like objects. Finally, a road-geometrical prior is considered to refine the extraction result, so that the non-road objects in images can be removed. Experimental results on two remote sensing image datasets indicate the validity and effectiveness of our method by comparing with two other approaches. Ying Wang 0008, Yongchao Gong, Feiyun Zhu, Chunhong Pan |
ICIP | 2 |
| 2014 | Spectral Unmixing via Data-Guided SparsityabstractHyperspectral unmixing, the process of estimating a common set of spectral bases and their corresponding composite percentages at each pixel, is an important task for hyperspectral analysis, visualization, and understanding. From an unsupervised learning perspective, this problem is very challenging-both the spectral bases and their composite percentages are unknown, making the solution space too large. To reduce the solution space, priors. In practice, these priors would easily lead to some unsuitable solution. This is because they are achieved by applying an identical strength of constraints to all the factors, which does not hold in practice. To overcome this limitation, we propose a novel sparsity-based method by learning a data-guided map (DgMap) to describe the individual mixed level of each pixel. Through this DgMap, the l(p) (0 < p < 1) constraint is applied in an adaptive manner. Such implementation not only meets the practical situation, but also guides the spectral bases toward the pixels under highly sparse constraint. What is more, an elegant optimization scheme as well as its convergence proof have been provided in this paper. Extensive experiments on several datasets also demonstrate that the DgMap is feasible, and high quality unmixing results could be obtained by our method. Feiyun Zhu, Ying Wang 0008, Bin Fan 0001, Shiming Xiang, Gaofeng Meng, Chunhong Pan |
IEEE Trans. Image Process. | 2 |
| 2013 | Efficient Image Dehazing with Boundary Constraint and Contextual RegularizationabstractImages captured in foggy weather conditions often suffer from bad visibility. In this paper, we propose an efficient regularization method to remove hazes from a single input image. Our method benefits much from an exploration on the inherent boundary constraint on the transmission function. This constraint, combined with a weighted L_1-norm based contextual regularization, is modeled into an optimization problem to estimate the unknown scene transmission. A quite efficient algorithm based on variable splitting is also presented to solve the problem. The proposed method requires only a few general assumptions and can restore a high-quality haze-free image with faithful colors and fine image details. Experimental results on a variety of haze images demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan |
ICCV | 2 |
| 2013 | Group sparsity based semi-supervised band selection for hyperspectral imagesabstractIn this paper, we propose a novel group sparsity based semi-supervised band selection method. There are three key features in our method. First, it fulfills the band selection task by employing group sparsity on the regression coefficients in a robust linear regression for classification model, so that the selected bands hold lower classification errors. Second, the spatial smoothness prior is incorporated to preserve the similarity of spatial neighbors in band selection. Third, the objective function is efficiently optimized via an alternative iteration algorithm. Comparative results on two hyper-spectral data sets validate the effectiveness of our method, showing higher classification accuracies. Haichang Li, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan |
ICIP | 2 |
| 2013 | Kinect depth restoration via energy minimization with TV21 regularizationabstractDepth maps generated by Kinect cameras often contain a significant amount of missing pixels and strong noise, limiting their usability in many computer vision applications. We present a new energy minimization method to fill the missing regions and remove noise in a depth map, by exploiting the strong correlation between color and depth values in local image neighborhoods. To preserve sharp edges and remove noise from the depth map, we propose to add a TV21regularization term into the energy function. Finally, we show how to effectively minimize the total energy using an alternating optimization approach. Experimental results show that the proposed method outperforms commonly-used depth inpainting approaches. Shaoguo Liu, Ying Wang 0008, Jue Wang 0001, Jixia Zhang, Chunhong Pan |
ICIP | 2 |
| 2013 | Maximum correntropy criterion based 3D head tracking with commodity depth cameraabstract3D head tracking becomes easier with the depth image from Microsoft Kinect. However, the noise from face occlusion and illumination still affects the tracking quality. In this paper, we introduce the robust Maximum Correntropy Criterion (MCC) to the problem of 3D head tracking, to tackle these noises. Fortunately, MCC can handle arbitrarily distributed noises. To solve the MCC based cost function, we develop an effective two-stage optimization scheme with the half-quadric technology. A head tracking system that uses Miscrosoft Kinect is also developed based on the MCC formulation. The system is fully automatic and online, without need of offline training. Experimental results show that the system is very robust against partial occlusion, large motion and sudden illumination variations. Shaoguo Liu, Ying Wang 0008, Jixia Zhang, Chunhong Pan |
ICIP | 3 |
| 2013 | An integrated graph-based face segmentation approach from Kinect videosabstractIn this paper, we present an Integrated Semi-Supervised Graph (IntSSG) approach to automatically segment face from color-depth video. In the first step, IntSSG performs skin color detection and online SIFT matching to initialize some face and non-face pixels. Then, the labels of these pixels are refined by conducting adaptive depth thresholding. Finally, based on a semi-supervised graph framework, IntSSG segments face by propagating the refined labels to other pixels. Experimental results show that IntSSG is able to accurately segment faces in difficult situations such as large pose changes and illumination variations. Jixia Zhang, Shaoguo Liu, Jiangyong Duan, Ying Wang 0008, Chunhong Pan |
ICIP | 5 |
| 2013 | Level set evolution with locally linear classification for image segmentation
Ying Wang 0008, Shiming Xiang, Chunhong Pan, Lingfeng Wang 0002, Gaofeng Meng |
Pattern Recognit. | 1 |
| 2012 | Image Guided Tone Mapping with Locally Nonlinear Model
Huxiang Gu, Ying Wang 0008, Shiming Xiang, Gaofeng Meng, Chunhong Pan |
ECCV (4) | 2 |
| 2012 | Image deblurring with matrix regression and gradient evolution
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang |
Pattern Recognit. | 3 |
| 2011 | Image editing based on Sparse Matrix-Vector multiplicationabstractThis paper presents a unified model for image editing in terms of Sparse Matrix-Vector (SpMV) multiplication. In our framework, we cast image editing as a linear energy minimization problem and address it by solving a sparse linear system, which is able to yield a globally optimal solution. First, three classical image editing operations, including linear filtering, resizing and selecting, are reformulated in the SpMV multiplication form. The SpMV form helps us set up a straightforward mechanism to flexibly and naturally combine various image features (low-level visual features or geometrical features) and constraints together into an integrated energy minimization function under the L2norm. Then, we apply our model to implement the tasks of pan-sharpening, image cloning, image mixed editing and texture transfer, which are now popularly used in the field of digital art. Comparative experiments are reported to validate the effectiveness and efficiency of our model. Ying Wang 0008, Hongping Yan, Chunhong Pan, Shiming Xiang |
ICASSP | 1 |
| 2011 | Level set evolution with locally linear classification for image segmentationabstractThis paper presents a novel local region-based level set model for image segmentation. In each local region, we define a locally weighted least squares energy to fit a linear classification function. The local energy is then integrated over the entire image domain to form an energy functional in terms of level set function. The energy minimization is achieved by level set evolution and estimation of parameters of the locally linear function in an iterative process. By introducing the locally linear functions to separate background and foreground in local regions, our model not only ensures the accuracy of the segmentation results, but also be very robust to initialization. Experiments are reported to demonstrate the effectiveness and efficiency of our model. Ying Wang 0008, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 1 |
| 2008 | Multimodal preserving embedding for face recognitionabstractTraditional dimension reduction approaches always consider the samples in a class are uni-modal. In real world, samples in a class are usually multi-modal, for instance, the manifold of the facial appearance of a person under different illumination, expression, and poses is multi-modal. Recently, dimension reduction approaches based on manifold learning are presented, the main purpose is to preserve the manifold structure on low dimensional space. In this paper, by analyzing the manifold learning methods and the traditional dimension reduction methods, we show that most of these methods can be summarized into a general framework. Based on this framework, we propose a novel dimension reduction approach, called multi-modal preserving embedding (MPE) by utilizing path-based similarity measure. We also describe two useful extensions of our method: KernelMPE and TensorMPE. Comprehensive comparisons and extensive experiments on face recognition are included to demonstrate the effectiveness of our method. Ying Wang 0008, Chunhong Pan |
FG | 1 |
| 2008 | Style preserving Chinese character synthesis based on hierarchical representation of characterabstractwith English, and they are not suitable for Chinese character synthesis. In this paper, we propose an unified approach for modeling and synthesizing Chinese characters. Using a three-level hierarchical representation, each character is decomposed into basic components, which forms the stroke database and radical database. In the synthesis process, we use a wavelet-based approach to select proper strokes and radicals, and some aesthetic constraints are defined based on the relationships between components, then genetic algorithm is employed to search for the optimal results which best match the aesthetic constraints. Experimental results demonstrates the effectiveness of our method. Ying Wang 0008, Chunhong Pan |
ICASSP | 1 |