VLDB 2026 Research / reviewers in the wild / expert
Yinfu Feng
dblp:51/9555
· DBLP profile ↗
26ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-9136-0965ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationabstractThe rapid evolution of the online fashion industry has intensified the demand for interactive fashion retrieval systems capable of precise and flexible searches based on user-specified attribute modifications. However, prevailing fashion retrieval methods often overlook the distinctive distributional properties of fashion images and struggle to preserve semantic consistency during attribute manipulation. To address these limitations, we propose DiSCo, a novel disentangled attribute manipulation retrieval framework via semantic reconstruction and consistency regularization. Our approach comprises three key components: (1) An attribute-aware manipulation network that constructs target fashion embeddings through cross-modal attribute modification deltas, leveraging dedicated fashion attribute encoders; (2) A cross-modal semantic reconstruction network that synthesizes target images directly from modified attribute descriptions, supervised by adversarial and attribute classification losses to ensure interpretable edits; (3) An adaptive fusion mechanism that dynamically integrates attribute-modified embeddings with reconstructed image features. Extensive evaluations on two benchmark datasets (DeepFashion and Shopping100K) demonstrate that DiSCo achieves superior retrieval accuracy over state-of-the-arts while maintaining high-fidelity editing. Quantitative and qualitative analyses further confirm that DiSCo generates more realistic fashion representations, underscoring its effectiveness in attribute-aware retrieval tasks. Min Tan 0005, Guanhao Liu, Huijing Zhan, Yuyu Yin, Zhou Yu 0001, Jiajun Ding, Yinfu Feng |
ACM Multimedia | 7 |
| 2025 | ENCODE: Breaking the Trade-Off Between Performance and Efficiency in Long-Term User Behavior ModelingabstractLong-term user behavior sequences are a goldmine for businesses to explore users’ interests to improve Click-Through Rate (CTR). However, it is very challenging to accurately capture users’ long-term interests from their long-term behavior sequences and give quick responses from the online serving systems. To meet such requirements, existing methods “inadvertently” destroy two basic requirements in long-term sequence modeling:R1) make full use of the entire sequence to keep the information as much as possible;R2) extract information from the most relevant behaviors to keep high relevance between learned interests and current target items. The performance of online serving systems is significantly affected by incomplete and inaccurate user interest information obtained by existing methods. To this end, we propose an efficient two-stage long-term sequence modeling approach, named asEfficieNtClustering based twO-stage interest moDEling (ENCODE), consisting of offline extraction stage and online inference stage. It not only meets the aforementioned two basic requirements but also achieves a desirable balance between online service efficiency and precision. Specifically, in the offline extraction stage, ENCODE clusters the entire behavior sequence and extracts accurate interests. To reduce the overhead of the clustering process, we design a metric learning-based dimension reduction algorithm that preserves the relative pairwise distances of behaviors in the new feature space. While in the online inference stage, ENCODE takes the off-the-shelf user interests to predict the associations with target items. Besides, to further ensure the relevance between user interests and target items, we adopt the same relevance metric throughout the whole pipeline of ENCODE. The extensive experiment and comparison with SOTA on both industrial and public datasets have demonstrated the effectiveness and efficiency of our proposed ENCODE. Yuhang Zheng 0003, Yinfu Feng, Yunan Ye, Rong Xiao 0005, Long Chen 0016, Xiaosong Yang, Jun Xiao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Decomposed Prototype Learning for Few-Shot Scene Graph GenerationabstractToday's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world applications with massive uncommon predicate categories whose annotations are hard to collect. In this article, we focus on Few-Shot SGG (FSSGG) , which encourages SGG models to be able to quickly transfer previous knowledge and recognize unseen predicates well with only a few examples. However, current methods for FSSGG are hindered by the high intra-class variance of predicate categories in SGG: On one hand, each predicate category commonly has multiple semantic meanings under different contexts. On the other hand, the visual appearance of relation triplets with the same predicate differs greatly under different subject–object compositions. Such great variance of inputs makes it hard to learn generalizable representation for each predicate category with current few-shot learning (FSL) methods. However, we found that this intra-class variance of predicates is highly related to the composed subjects and objects. To model the intra-class variance of predicates with subject–object context, we propose a novel Decomposed Prototype Learning (DPL) model for FSSGG. Specifically, we first construct a decomposable prototype space to capture diverse semantics and visual patterns of subjects and objects for predicates by decomposing them into multiple prototypes. Afterwards, we integrate these prototypes with different weights to generate query-adaptive predicate representation with more reliable semantics for each query sample. We conduct extensive experiments and compare with various baseline methods to show the effectiveness of our method. Jun Xiao 0001, Guikun Chen, Yinfu Feng, Yi Yang 0001, Anan Liu, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Multi-Domain Deep Learning from a Multi-View Perspective for Cross-Border E-commerce SearchabstractBuilding click-through rate (CTR) and conversion rate (CVR) prediction models for cross-border e-commerce search requires modeling the correlations among multi-domains. Existing multi-domain methods would suffer severely from poor scalability and low efficiency when number of domains increases. To this end, we propose a Domain-Aware Multi-view mOdel (DAMO), which is domain-number-invariant, to effectively leverage cross-domain relations from a multi-view perspective. Specifically, instead of working in the original feature space defined by different domains, DAMO maps everything to a new low-rank multi-view space. To achieve this, DAMO firstly extracts multi-domain features in an explicit feature-interactive manner. These features are parsed to a multi-view extractor to obtain view-invariant and view-specific features. Then a multi-view predictor inputs these two sets of features and outputs view-based predictions. To enforce view-awareness in the predictor, we further propose a lightweight view-attention estimator to dynamically learn the optimal view-specific weights w.r.t. a view-guided loss. Extensive experiments on public and industrial datasets show that compared with state-of-the-art models, our DAMO achieves better performance with lower storage and computational costs. In addition, deploying DAMO to a large-scale cross-border e-commence platform leads to 1.21%, 1.76%, and 1.66% improvements over the existing CGC-based model in the online AB-testing experiment in terms of CTR, CVR, and Gross Merchandises Value, respectively. Yinfu Feng, Yunan Ye, Min Tan 0005, Rong Xiao 0005, Haihong Tang, Jiajun Ding, Jun Yu 0002 |
AAAI | 2 |
| 2024 | FedSea: Federated Learning via Selective Feature Alignment for Non-IID Multimodal DataabstractThe growing demands for privacy protection challenge the joint training of one model by leveraging multiple datasets. Federated learning (FL) provides a new way to overcome this challenge and has attracted many research interests, which enables multiple parties to collaboratively train a machine learning model without exchanging their local data. Despite some success, the non-independent and identically distributed (non-IID) data distributions in different parties remain challenging and easily damage the performance of FL methods, specifically for the heterogeneous multimodal data. Existing FL studies on non-IID data settings are often dedicated to the label space, neglecting the non-IID issues in feature space, thus limiting their performance when the parties with non-IID multimodal data. This paper proposes a newFederated learning method viaSelective featureAlignment (FedSea) to align representations across multiple parties in the feature space. FedSea uses a domain adversarial learning framework consisting of an affine-transform-based generator and a gradient-reversal-based client discriminator to perform IID transformation and reduce data source distinguishability, respectively. An attention-based mask module and a feature IID confidence quantification method are introduced to effectively address the diverse feature non-IID levels across multimodal data. Comprehensive experiments are conducted on three widely-used public datasets and one large-scale industrial dataset, showing FedSea has: 1) better performance than state-of-the-art FL methods on both multimodal and single-modal datasets; 2) superior feature alignment ability on non-IID datasets, and 3) good model interpretability. Min Tan 0005, Yinfu Feng, Lingqiang Chu, Jingcheng Shi, Rong Xiao 0005, Haihong Tang, Jun Yu 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | UHD Aerial Photograph Categorization by Leveraging Deep Multiattribute Matrix FactorizationabstractThere are thousands of observation satellites orbiting the earth, each of which captures massive-scale photographs covering millions of square kilometers everyday. In practice, these aerial photos are with ultra-high-definitions (UHD) and may contain tens to hundreds of ground objects (e.g., vehicles and rooftops). Understanding the multiple categories of a rich variety of UHD aerial photos is an indispensable technique for many applications, such as intelligent transportation, natural disaster prediction, and smart agriculture. In this work, we propose a novel multi-label UHD aerial photo categorization pipeline, wherein the key is to topologically represent the spatial layouts of the ground objects and further deeply encode them using a deep multi-clue matrix factorization (DMCMF) that robustly handles noisy labels at image-level. More specifically, for each UHD aerial photo, we extract visually/semantically salient object patches inside it. To explicitly encode their spatial layout, we construct a graphlet by linking the spatially adjacent object patches into a small graph. Subsequently, a binary MF is designed to intelligently exploit the semantics of these graphlets, wherein four clues: i) binary hash codes learning, ii) noisy labels refinement, iii) deep image-level semantics, and iv) adaptive data graph updating are incorporated. Such DMCMF can be solved iteratively and each graphlet is then converted into the discrete hash codes. Finally, the hash codes corresponding to graphlets within each UHD aerial photo are quantized into a feature vector by a kernel machine for multi-label categorization. Toward a comprehensive comparative study, we complied a million-scale UHD aerial photo set collected from 100 top-ranking cities worldwide. Experiments have shown that 1) our method is highly competitive in learning categorization model from imperfect labels at image-level, and 2) the four clues are elaborately designed and seamlessly combined to learn hash codes for representing UHD aerial photos. Yinfu Feng, Bing Tu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Cross-Lingual Product Retrieval in E-Commerce Search
Wenya Zhu, Xiaoyu Lv, Baosong Yang, Xu Yong, Linlong Xu, Yinfu Feng, Haibo Zhang 0013, Qing Da, Anxiang Zeng, Ronghua Chen |
PAKDD (2) | 7 |
| 2022 | DHA: Product Title Generation with Discriminative Hierarchical Attention for E-commerce
Wenya Zhu, Yu Zhang 0006, Yu-Hang Zhou, Yinfu Feng, Yuxiang Wu, Qing Da, Anxiang Zeng |
PAKDD (3) | 5 |
| 2021 | A Primal-Dual Online Algorithm for Online Matching Problem in Dynamic EnvironmentsabstractRecently, the online matching problem has attracted much attention due to its wide application on real-world decision-making scenarios. In stationary environments, by adopting the stochastic user arrival model, existing methods are proposed to learn dual optimal prices and are shown to achieve a fast regret bound. However, the stochastic model is no longer a proper assumption when the environment is changing, leading to an optimistic method that may suffer poor performance. In this paper, we study the online matching problem in dynamic environments in which the dual optimal prices are allowed to vary over time. We bound the dynamic regret of online matching problem by the sum of two quantities, including a regret of online max-min problem and a dynamic regret of online convex optimization (OCO) problem. Then we propose a novel online approach named Primal-Dual Online Algorithm (PDOA) to minimize both quantities. In particular, PDOA adopts the primal-dual framework by optimizing dual prices with the online gradient descent (OGD) algorithm to eliminate the online max-min problem's regret. Moreover, it maintains a set of OGD experts and combines them via an expert-tracking algorithm, which gives a sublinear dynamic regret bound for the OCO problem. We show that PDOA achieves an O(K sqrt{T(1+P_T)}) dynamic regret where K is the number of resources, T is the number of iterations and P_T is the path-length of any potential dual price sequence that reflects the dynamic environment. Finally, experiments on real applications exhibit the superiority of our approach. Yu-Hang Zhou, Guangda Huzhang, Yinfu Feng, Qing Da, Xinshang Wang, Anxiang Zeng |
AAAI | 6 |
| 2017 | A human motion feature based on semi-supervised learning of GMM
Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Hanzhi Zhang, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Multim. Syst. | 2 |
| 2016 | A 3D human motion refinement method based on sparse motion bases selectionabstractMotion capture (MOCAP) is an important technique that is widely used in many areas such as computer animation, film industry, physical training and so on. Even with professional MOCAP system, the missing marker problems always occur. Motion refinement is an essential preprocessing step for MOCAP data based applications. Although many existing approaches for motion refinement have been developed, it is still a challenging task due to the complexity and diversity of human motion. A data driven based motion refinement method is proposed in this paper, which modifies the traditional sparse coding process for special task of motion recovery from missing parts. Meanwhile, the objective function is derived by taking both statistical and kinematical property of motion data into account. Poselet model and moving window grouping are applied in the proposed method to achieve a fine-grained feature representation, which preserves the embedded spatial-temporal kinematic information. 5 motion dictionaries are learnt for each kind of poselet from training data in parallel. The motion refine problem is finally solved as an ℓ1-minimization problem. Compared with several state-of-art motion refine methods, the experimental result shows that our approach outperforms the competitors. Yinfu Feng, Shuang Liu 0006, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001 |
CASA | 2 |
| 2016 | Adaptive multi-view feature selection for human motion retrievalabstractHuman motion retrieval plays an important role in many motion data based applications. In the past, many researchers tended to use a single type of visual feature as data representation. Because different visual feature describes different aspects about motion data, and they have dissimilar discriminative power with respect to one particular class of human motion, it led to poor retrieval performance. Thus, it would be beneficial to combine multiple visual features together for motion data representation. In this article, we present an Adaptive Multi-view Feature Selection (AMFS) method for human motion retrieval. Specifically, we first use a local linear regression model to automatically learn multiple view-based Laplacian graphs for preserving the local geometric structure of motion data. Then, these graphs are combined together with a non-negative view-weight vector to exploit the complementary information between different features. Finally, in order to discard the redundant and irrelevant feature components from the original high-dimensional feature representation, we formulate the objective function of AMFS as a general trace ratio optimization problem, and design an effective algorithm to solve the corresponding optimization problem. Extensive experiments on two public human motion database, i.e., HDM05 and MSR Action3D, demonstrate the effectiveness of the proposed AMFS over the state-of-art methods for motion data retrieval. The scalability with large motion dataset, and insensitivity with the algorithm parameters, make our method can be widely used in real-world applications. Yinfu Feng, Tian Qi, Xiaosong Yang, Jian J. Zhang 0001 |
Signal Process. | 2 |
| 2016 | Fast view-based 3D model retrieval via unsupervised multiple feature fusion and online projection learning
Jun Xiao 0001, Yinfu Feng, Mingming Ji, Yueting Zhuang |
Signal Process. | 2 |
| 2015 | A locally weighted sparse graph regularized Non-Negative Matrix Factorization method
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang |
Neurocomputing | 1 |
| 2015 | Efficient semi-supervised multiple feature fusion with out-of-sample extension for 3D model retrieval
Mingming Ji, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Neurocomputing | 2 |
| 2015 | Sparse motion bases selection for human motion denoisingabstractHuman motion denoising is an indispensable step of data preprocessing for many motion data based applications. In this paper, we propose a data-driven based human motion denoising method that sparsely selects the most correlated subset of motion bases for clean motion reconstruction. Meanwhile, it takes the statistic property of two common noises, i.e., Gaussian noise and outliers, into account in deriving the objective functions. In particular, our method firstly divides each human pose into five partitions termed as poselets to gain a much fine-grained pose representation. Then, these poselets are reorganized into multiple overlapped poselet groups using a lagged window moving across the entire motion sequence to preserve the embedded spatial–temporal motion patterns. Afterward, five compacted and representative motion dictionaries are constructed in parallel by means of fast K-SVD in the training phase; they are used to remove the noise and outliers from noisy motion sequences in the testing phase by solving ℓ 1 -minimization problems. Extensive experiments show that our method outperforms its competitors. More importantly, compared with other data-driven based method, our method does not need to specifically choose the training data , it can be more easily applied to real-world applications. Jun Xiao 0001, Yinfu Feng, Mingming Ji, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang |
Signal Process. | 2 |
| 2015 | Sketch-based human motion retrieval via selected 2D geometric posture descriptorabstractSketch-based human motion retrieval is a hot topic in computer animation in recent years. In this paper, we present a novel sketch-based human motion retrieval method via selected 2-dimensional (2D) Geometric Posture Descriptor (2GPD). Specially, we firstly propose a rich 2D pose feature call 2D Geometric Posture Descriptor (2GPD), which is effective in encoding the 2D posture similarity by exploiting the geometric relationships among different human body parts. Since the original 2GPD is of high dimension and redundant, a semi-supervised feature selection algorithm derived from Laplacian Score is then adopted to select the most discriminative feature component of 2GPD as feature representation, and we call it as selected 2GPD. Finally, a posture-by-posture motion retrieval algorithm is used to retrieve a motion sequence by sketching several key postures. Experimental results on CMU human motion database demonstrate the effectiveness of our proposed approach. Jun Xiao 0001, Zhangpeng Tang, Yinfu Feng, Zhidong Xiao |
Signal Process. | 3 |
| 2015 | Mining Spatial-Temporal Patterns and Structural Sparsity for Human Motion Data DenoisingabstractMotion capture is an important technique with a wide range of applications in areas such as computer vision, computer animation, film production, and medical rehabilitation. Even with the professional motion capture systems, the acquired raw data mostly contain inevitable noises and outliers. To denoise the data, numerous methods have been developed, while this problem still remains a challenge due to the high complexity of human motion and the diversity of real-life situations. In this paper, we propose a data-driven-based robust human motion denoising approach by mining the spatial-temporal patterns and the structural sparsity embedded in motion data. We first replace the regularly used entire pose model with a much fine-grained partlet model as feature representation to exploit the abundant local body part posture and movement similarities. Then, a robust dictionary learning algorithm is proposed to learn multiple compact and representative motion dictionaries from the training data in parallel. Finally, we reformulate the human motion denoising problem as a robust structured sparse coding problem in which both the noise distribution information and the temporal smoothness property of human motion have been jointly taken into account. Compared with several state-of-the-art motion denoising methods on both the synthetic and real noisy motion data, our method consistently yields better performance than its counterparts. The outputs of our approach are much more stable than that of the others. In addition, it is much easier to setup the training dataset of our method than that of the other data-driven-based methods. Yinfu Feng, Mingming Ji, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Exploiting temporal stability and low-rank structure for motion capture data refinementabstractInspired by the development of the matrix completion theories and algorithms, a low-rank based motion capture (mocap) data refinement method has been developed, which has achieved encouraging results. However, it does not guarantee a stable outcome if we only consider the low-rank property of the motion data. To solve this problem, we propose to exploit the temporal stability of human motion and convert the mocap data refinement problem into a robust matrix completion problem, where both the low-rank structure and temporal stability properties of the mocap data as well as the noise effect are considered. An efficient optimization method derived from the augmented Lagrange multiplier algorithm is presented to solve the proposed model. Besides, a trust data detection method is also introduced to improve the degree of automation for processing the entire set of the data and boost the performance. Extensive experiments and comparisons with other methods demonstrate the effectiveness of our approaches on both predicting missing data and de-noising. Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001, Rong Song |
Inf. Sci. | 1 |
| 2014 | Human motion retrieval based on freehand sketchabstractABSTRACT In this paper, we present an integrated framework of human motion retrieval based on freehand sketch. With some simple rules, the user can acquire a desired motion by sketching several key postures. To retrieve efficiently and accurately by sketch, the 3D postures are projected onto several 2D planes. The limb direction feature is proposed to represent the input sketch and the projected‐postures. Furthermore, a novel index structure based on k‐d tree is constructed to index the motions in the database, which speeds up the retrieval process. With our posture‐by‐posture retrieval algorithm, a continuous motion can be got directly or generated by using a pre‐computed graph structure. What's more, our system provides an intuitive user interface. The experimental results demonstrate the effectiveness of our method. © 2014 The Authors.Computer Animation and Virtual Worldspublished by John Wiley & Sons, Ltd. Zhangpeng Tang, Jun Xiao 0001, Yinfu Feng, Xiaosong Yang |
Comput. Animat. Virtual Worlds | 3 |
| 2014 | Real-time motion data annotation via action stringabstractABSTRACT Even though there is an explosive growth of motion capture data, there is still a lack of efficient and reliable methods to automatically annotate all the motions in a database. Moreover, because of the popularity of mocap devices in home entertainment systems, real‐time human motion annotation or recognition becomes more and more imperative. This paper presents a new motion annotation method that achieves both the aforementioned two targets at the same time. It uses a probabilistic pose feature based on the Gaussian Mixture Model to represent each pose. After training a clustered pose feature model, a motion clip could be represented as an action string. Then, a dynamic programming‐based string matching method is introduced to compare the differences between action strings. Finally, in order to achieve the real‐time target, we construct a hierarchical action string structure to quickly label each given action string. The experimental results demonstrate the efficacy and efficiency of our method. Copyright © 2014 John Wiley & Sons, Ltd. Qi Tian 0001, Jun Xiao 0001, Yueting Zhuang, Hanzhi Zhang, Xiaosong Yang, Jian J. Zhang 0001, Yinfu Feng |
Comput. Animat. Virtual Worlds | 7 |
| 2014 | Feature Correlation Hypergraph: Exploiting High-order Potentials for Multimodal RecognitionabstractIn computer vision and multimedia analysis, it is common to use multiple features (or multimodal features) to represent an object. For example, to well characterize a natural scene image, we typically extract a set of visual features to represent its color, texture, and shape. However, it is challenging to integrate multimodal features optimally. Since they are usually high-order correlated, e.g., the histogram of gradient (HOG), bag of scale invariant feature transform descriptors, and wavelets are closely related because they collaboratively reflect the image texture. Nevertheless, the existing algorithms fail to capture the high-order correlation among multimodal features. To solve this problem, we present a new multimodal feature integration framework. Particularly, we first define a new measure to capture the high-order correlation among the multimodal features, which can be deemed as a direct extension of the previous binary correlation. Therefore, we construct a feature correlation hypergraph (FCH) to model the high-order relations among multimodal features. Finally, a clustering algorithm is performed on FCH to group the original multimodal features into a set of partitions. Moreover, a multiclass boosting strategy is developed to obtain a strong classifier by combining the weak classifiers learned from each partition. The experimental results on seven popular datasets show the effectiveness of our approach. Yinfu Feng, Jianke Zhu, Deng Cai 0001 |
IEEE Trans. Cybern. | 4 |
| 2013 | A semantic feature for human motion retrievalabstractABSTRACT With the explosive growth of motion capture data, it becomes very imperative in animation production to have an efficient search engine to retrieve motions from large motion repository. However, because of the high dimension of data space and complexity of matching methods, most of the existing approaches cannot return the result in real time. This paper proposes a high level semantic feature in a low dimensional space to represent the essential characteristic of different motion classes. On the basis of the statistic training of Gauss Mixture Model, this feature can effectively achieve motion matching on both global clip level and local frame level. Experiment results show that our approach can retrieve similar motions with rankings from large motion database in real‐time and also can make motion annotation automatically on the fly. Copyright © 2013 John Wiley & Sons, Ltd. Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2012 | Adaptive Unsupervised Multi-view Feature Selection for Visual Concept Recognition
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaoming Liu 0002 |
ACCV (1) | 1 |
| 2012 | Active learning for social image retrieval using Locally Regressive Optimal Design
Yinfu Feng, Jun Xiao 0001, Zhengjun Zha, Yi Yang 0001 |
Neurocomputing | 1 |
| 2011 | Predicting missing markers in human motion capture using l1-sparse representationabstractAbstract Missing marker problem is very common in human motion capture. In contrast to most current methods which handle this problem based on trying to learn a reliable predictor from the observations, we consider it from the perspective of sparse representation and propose a novel method which is namedl1‐sparse representation of missing markers prediction (L1‐SRMMP). We assume that the incomplete pose can be represented by a linear combination of a few poses from the training set and the representation is sparse. Therefore, we cast the predicting missing markers as finding a sparse representation of the observable data of the incomplete pose, and then we use it to predict the missing data. In order to get a sparse representation, we employl1‐norm in our objective function. Moreover, we propose presentation coefficient weighted update (PCWU) algorithm to mitigate the limited capacity problem of the training set. Experimental results demonstrate the effectiveness and efficiency of our method to predict the missing markers in human motion capture. Copyright © 2011 John Wiley & Sons, Ltd. Jun Xiao 0001, Yinfu Feng, Wenyuan Hu |
Comput. Animat. Virtual Worlds | 2 |