Yaqing Wang 0002

dblp:147/1393-2 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0003-1457-1114ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Few-Shot Learning to Data-Efficient Intelligence
abstract
Modern artificial intelligence performs impressively in data-rich settings but still struggles to learn and adapt from only a few examples—a capability central to human intelligence. My research seeks to understand and enable data-efficient generalization, unifying principles across few-shot learning, meta-learning, in-context learning in large language models (LLMs), and adaptive agent behavior. First, I revisit few-shot learning from a foundational perspective, showing why conventional supervised learning breaks down under sparse data and how prior knowledge enables reliable adaptation. I then discuss how these principles extend to real-world scenarios such as scientific discovery and cold-start recommendation, where data are scarce, costly, or dynamically evolving. Finally, I explore how LLMs perform in-context learning and how their adaptive behaviors connect to meta-learning mechanisms. Building on these insights, I develop data-efficient, preference-adaptive agents that quickly align to user needs with minimal interaction.This talk presents a cohesive view of data-efficient intelligence and outlines future directions toward more reliable, human-like learning systems.
Yaqing Wang 0002
AAAI1
2026 Robust heterogeneous network representation learning by multifaceted curriculum training
Zhen Hao Wong, Hansi Yang, Quanming Yao, Yaqing Wang 0002
Neural Networks4
2026 Searching to Modulate for Cold-Start Recommendation
Shiguang Wu 0002, Yaqing Wang 0002, Quanming Yao
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Why In-Context Learning Models are Good Few-Shot Learners?
abstract
We explore in-context learning (ICL) models from a learning-to-learn perspective. Unlike studies that identify specific learning algorithms in ICL models, we compare ICL models with typical meta-learners to understand their superior performance. We theoretically prove the expressiveness of ICL models as learning algorithms and examine their learnability and generalizability. Our findings show that ICL with transformers can effectively construct data-dependent learning algorithms instead of directly follow existing ones (including gradient-based, metric-based, and amortization-based meta-learners). The construction of such learning algorithm is determined by the pre-training process, as a function fitting the training distribution, which raises generalizability as an important issue. With above understanding, we propose strategies to transfer techniques for classical deep networks to meta-level to further improve ICL. As examples, we implement meta-level meta-learning for domain adaptability with limited data and meta-level curriculum learning for accelerated convergence during pre-training, demonstrating their empirical effectiveness.
Shiguang Wu 0002, Yaqing Wang 0002, Quanming Yao
ICLR2
2025 PERSCEN: Learning Personalized Interaction Pattern and Scenario Preference for Multi-Scenario Matching
abstract
With the expansion of business scales and scopes on online platforms, multi-scenario matching has become a mainstream solution to reduce maintenance costs and alleviate data sparsity. The key to effective multi-scenario recommendation lies in capturing both user preferences shared across all scenarios and scenario-aware preferences specific to each scenario. However, existing methods often overlook user-specific modeling, limiting the generation of personalized user representations. To address this, we propose PERSCEN, an innovative approach that incorporates user-specific modeling into multi-scenario matching. PERSCEN constructs a user-specific feature graph based on user characteristics and employs a lightweight graph neural network to capture higher-order interaction patterns, enabling personalized extraction of preferences shared across scenarios. Additionally, we leverage vector quantization techniques to distill scenario-aware preferences from users' behavior sequence within individual scenarios, facilitating user-specific and scenario-aware preference modeling. To enhance efficient and flexible information transfer, we introduce a progressive scenario-aware gated linear unit that allows fine-grained, low-latency fusion. Extensive experiments demonstrate that PERSCEN outperforms existing methods. Further efficiency analysis confirms that PERSCEN effectively balances performance with computational cost, ensuring its practicality for real-world industrial systems.
Haotong Du, Yaqing Wang 0002, Quanming Yao, Zhen Wang 0004
KDD (2)2
2025 Adaptive Preference Arithmetic: A Personalized Agent with Adaptive Preference Arithmetic for Dynamic Preference Modeling
abstract
As large language models (LLMs) are increasingly used as personalized user assistants, effectively adapting to users' evolving preferences is critical for delivering high-quality personalized responses. While user preferences are often stable in content, their relative strengths shift over time due to changing goals and contexts. Therefore, modeling these dynamic preference strengths can enable finer-grained personalization. However, current methods face two major challenges: (i) limited user feedback makes it difficult to estimate preference strengths accurately, and (ii) natural language ambiguity limits the controllability of preference-guided generation. To address these issues, we propose AdaPA-Agent, a LLM-agent personalization framework that models dynamic preference strengths via Adaptive Preference Arithmetic. First, instead of requiring additional user feedback, AdaPA-Agent employs an alignment-based strength estimation module to estimate the strength of user preferences from the existing user-agent interaction. Then, it guides controllable personalized generation by linearly combining next-token distributions, weighted by the estimated strengths of individual preferences. Experiments on two personalization tasks-conversational recommendation and personalized web interaction-demonstrate that AdaPA-Agent better aligning with users' changing intents, and has achieved over 18.9\% and 14.2\% improvements compared to ReAct, the widely-used agent framework.
Hongyi Nie, Yaqing Wang 0002, Feiyang Pan, Quanming Yao, Zhen Wang 0004
NeurIPS2
2025 Learning to Learn with Contrastive Meta-Objective
abstract
Meta-learning enables learning systems to adapt quickly to new tasks, similar to humans. Different meta-learning approaches all work under/with the mini-batch episodic training framework. Such framework naturally gives the information about task identity, which can serve as additional supervision for meta-training to improve generalizability. We propose to exploit task identity as additional supervision in meta-training, inspired by the alignment and discrimination ability which is is intrinsic in human's fast learning. This is achieved by contrasting what meta-learners learn, i.e., model representations. The proposed ConML is evaluating and optimizing the contrastive meta-objective under a problem- and learner-agnostic meta-training framework. We demonstrate that ConML integrates seamlessly with existing meta-learners, as well as in-context learning models, and brings significant boost in performance with small implementation cost.
Shiguang Wu 0002, Yaqing Wang 0002, Yatao Bian, Quanming Yao
NeurIPS2
2024 PACIA: Parameter-Efficient Adapter for Few-Shot Molecular Property Prediction
Shiguang Wu 0002, Yaqing Wang 0002, Quanming Yao
IJCAI2
2024 Warming Up Cold-Start CTR Prediction by Learning Item-Specific Feature Interactions
abstract
In recommendation systems, new items are continuously introduced, initially lacking interaction records but gradually accumulating them over time.Accurately predicting the click-through rate (CTR) for these items is crucial for enhancing both revenue and user experience.While existing methods focus on enhancing item ID embeddings for new items within general CTR models, they tend to adopt a global feature interaction approach, often overshadowing new items with sparse data by those with abundant interactions.Addressing this, our work introduces EmerG, a novel approach that warms up cold-start CTR prediction by learning item-specific feature interaction patterns.EmerG utilizes hypernetworks to generate an item-specific feature graph based on item characteristics, which is then processed by a Graph Neural Network (GNN).This GNN is specially tailored to provably capture feature interactions at any order through a customized message passing mechanism.We further design a meta learning strategy that optimizes parameters of hypernetworks and GNN across various item CTR prediction tasks, while only adjusting a minimal set of item-specific parameters within each task.This strategy effectively reduces the risk of overfitting when dealing with limited data.Extensive experiments on benchmark datasets validate that EmerG consistently performs the best given no, a few and sufficient instances of new items.
Yaqing Wang 0002, Hongming Piao, Daxiang Dong, Quanming Yao, Jingbo Zhou 0003
KDD1
2024 Property-Aware Relation Networks for Few-Shot Molecular Property Prediction
abstract
Molecular property prediction plays a fundamental role in AI-aided drug discovery to identify candidate molecules, which is also essentially a few-shot problem due to lack of labeled data. In this paper, we propose Property-Aware Relation networks (PAR) to handle this problem. We first introduce a property-aware molecular encoder to transform the generic molecular embeddings to property-aware ones. Then, we design a query-dependent relation graph learning module to estimate molecular relation graph and refine molecular embeddings w.r.t. the target property. Thus, the facts that both property-related information and relationships among molecules change across different properties are utilized to better learn and propagate molecular embeddings. Generally, PAR can be regarded as a combination of metric-based and optimization-based few-shot learning method. We further extend PAR to Transferable PAR (T-PAR) to handle the distribution shift, which is common in drug discovery. The keys are joint sampling and relation graph learning schemes, which simultaneously learn molecular embeddings from both source and target domains. Extensive results on benchmark datasets show that PAR and T-PAR consistently outperform existing methods on few-shot and transferable few-shot molecular property prediction tasks, respectively. Besides, ablation and case studies are conducted to validate the rationality of our designs in PAR and T-PAR.
Quanming Yao, Zhenqian Shen, Yaqing Wang 0002, Dejing Dou
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 A Novel Tensor Learning Model for Joint Relational Triplet Extraction
abstract
The relational triplet is a format to represent relational facts in the real world, which consists of two entities and a semantic relation between these two entities. Since the relational triplet is the essential component in a knowledge graph (KG), extracting relational triplets from unstructured texts is vital for KG construction and has attached increasing research interest in recent years. In this work, we find that relation correlation is common in real life and could be beneficial for the relational triplet extraction task. However, existing relational triplet extraction works neglect to explore the relation correlation that bottlenecks the model performance. Therefore, to better explore and take advantage of the correlation among semantic relations, we innovatively utilize a three-dimension word relation tensor to describe relations between words in a sentence. Then, we treat the relation extraction task as a tensor learning problem and propose an end-to-end tensor learning model based on Tucker decomposition. Compared with directly capturing correlation among relations in a sentence, learning the correlation of elements in a three-dimension word relation tensor is more feasible and could be addressed through tensor learning methods. To verify the effectiveness of the proposed model, extensive experiments are also conducted on two widely used benchmark datasets, that is, NYT and WebNLG. Results show that our model outperforms the state-of-the-art by a large margin of F1 scores, such as the developed model has an improvement of 3.2% on the NYT dataset compared to the state-of-the-art. Source codes and data can be found at https://github.com/Sirius11311/TLRel.git.
Zhen Wang 0004, Hongyi Nie, Wei Zheng 0006, Yaqing Wang 0002, Xuelong Li 0001
IEEE Trans. Cybern.4
2024 PreGesNet: Few-Shot Acoustic Gesture Recognition Based on Task-Adaptive Pretrained Networks
abstract
Acoustic-based human gesture recognition (HGR) applications have drawn increasing academic attention in order to overcome the shortcomings of conventional interaction methods on tiny devices. Existing techniques following a learning-based routine requires collecting massive application-specific training data. What is worse, the cross-domain problem induces additional retraining overhead to enable the systems recognize unseen gestures in different environments. This obviously decreases their scalability and prevent them from real-world deployment. Although some recent works propose different few-shot learning solutions to deal with the cross-domain problem in HGR, they possess shortcomings of being application-specific, high training overhead, and/or incapability to recognize unseen gestures. In this paper, we propose PreGesNet, a few-shot acoustic gesture recognition framework based on task-adaptive pretrained networks whose novelty lies in three aspects: i) leveraging pretrained feature extractor which captures generic knowledge of our collected and open-source large-scale gesture datasets; ii) designing task-specific parameter adaptation mechanism to efficiently update the feature extractor to adapt the pretrained feature extractor to each target task; iii) discovering suitable distance metric and task generation strategy which fit HGR application. According to the experiments, when the model is trained with 10 digit gestures, its recognition accuracies of 26 kinds of letter gestures and 8 kinds of other hand gestures can be up to 80.5% and 93.4% with only two shots, respectively. In addition, the average recognition latency of PreGesNet is less than 0.4 second.
Yongpan Zou, Yunshu Wang, Haozhi Dong, Yaqing Wang 0002, Yanbo He, Kaishun Wu
IEEE Trans. Mob. Comput.4
2023 Efficient and Joint Hyperparameter and Architecture Search for Collaborative Filtering
abstract
Automated Machine Learning (AutoML) techniques have recently been introduced to design Collaborative Filtering (CF) models in a data-specific manner. However, existing works either search architectures or hyperparameters while ignoring the fact they are intrinsically related and should be considered together. This motivates us to consider a joint hyperparameter and architecture search method to design CF models. However, this is not easy because of the large search space and high evaluation cost. To solve these challenges, we reduce the space by screening out usefulness hyperparameter choices through a comprehensive understanding of individual hyperparameters. Next, we propose a two-stage search algorithm to find proper configurations from the reduced space. In the first stage, we leverage knowledge from subsampled datasets to reduce evaluation costs; in the second stage, we efficiently fine-tune top candidate models on the whole dataset. Extensive experiments on real-world datasets show better performance can be achieved compared with both hand-designed and previous searched models. Besides, ablation and case studies demonstrate the effectiveness of our search framework.
Chen Gao 0001, Lingling Yi, Liwei Qiu, Yaqing Wang 0002, Yong Li 0008
KDD5
2023 ColdNAS: Search to Modulate for User Cold-Start Recommendation
abstract
Making personalized recommendation for cold-start users, who only have a few interaction histories, is a challenging problem in recommendation systems. Recent works leverage hypernetworks to directly map user interaction histories to user-specific parameters, which are then used to modulate predictor by feature-wise linear modulation function. These works obtain the state-of-the-art performance. However, the physical meaning of scaling and shifting in recommendation data is unclear. Instead of using a fixed modulation function and deciding modulation position by expertise, we propose a modulation framework called ColdNAS for user cold-start problem, where we look for proper modulation structure, including function and position, via neural architecture search. We design a search space which covers broad models and theoretically prove that this search space can be transformed to a much smaller space, enabling an efficient and robust one-shot search algorithm. Extensive experimental results on benchmark datasets show that ColdNAS consistently performs the best. We observe that different modulation functions lead to the best performance on different datasets, which validates the necessity of designing a searching-based method. Codes are available at https://github.com/LARS-research/ColdNAS.
Shiguang Wu 0002, Yaqing Wang 0002, Qinghe Jing, Daxiang Dong, Dejing Dou, Quanming Yao
WWW2
2023 Distilling ensemble of explanations for weakly-supervised pre-training of image segmentation models
Xuhong Li 0002, Haoyi Xiong, Yi Liu 0040, Dingfu Zhou, Yaqing Wang 0002, Dejing Dou
Mach. Learn.6
2023 Feynman: Federated Learning-Based Advertising for Ecosystems-Oriented Mobile Apps Recommendation
abstract
While recommender systems have been ubiquitously used in digital marketing and online business development, the conversions of online advertising for mobile apps installation and activation sometimes are far from satisfactory, due to the lack of feedback from App-related activities, leading to a poor record of Return on Investment (RoI). Though the advertisers, e.g., App operators and App Store, are granted to log users’ app-related activities such as installation, activation, usages, and preferences per the agreement, they usually limit the access to such data from advertisement publishers, due to the privacy concerns. To improve conversions of online advertising under privacy controls, we proposeFeynman—afederated learning-based advertising platform for ecosystems-orientedmobileapps recommendation.Feynmanaims at improving the RoI of mobile app recommendation from an ecosystem's perspective, i.e., per investment in advertising an app (Goal. 1) increasing the number of new installs/users of the app, and then (Goal. 2) increasing the number of new active users (preferably with frequent in-app purchase activities). Incorporating with a federated computing platform,Feynmanleverages users’ records stored in advertisers to refine the pool of targeting users for ads distribution, and jointly builds the predictive models for users’ purchase activities forecasting using features from the Ads publisher and the advertiser. With refined target pools and more accurate models,Feynmanhas successfully helped several mobile apps in China by attracting more than 100 million users to further enlarge their user populations and revenues from in-app purchases. Note that rather than proposing new techniques for federated learning, the design ofFeynmandedicates to show its promising performance in the industrial practices of advertising using federated computing and privacy protected strategies. In three cases that we report in this paper,Feynmanoutperforms the state-of-the-art plans in terms of several key measurements, including Click-Through Rates (CTR), Conversion Rate (CVR), Cost per Action (CPA), and Non-targeting User Hit-Rates (NTHR).
Jiang Bian 0003, Jizhou Huang, Shilei Ji, Yuan Liao 0003, Xuhong Li 0002, Qingzhong Wang, Jingbo Zhou 0003, Dejing Dou, Yaqing Wang 0002, Haoyi Xiong
IEEE Trans. Serv. Comput.9
2022 Simplified Graph Learning for Inductive Short Text Classification
abstract
Short text classification (STC) is hard as short texts lack context information and labeled data is not enough.Graph neural networks obtain the state-of-the-art on STC since they can merge various auxiliary information via the message passing framework.However, existing works conduct transductive learning, which requires retraining to accommodate new samples and takes large memory.In this paper, we present SimpleSTC which handles inductive STC problem but only leverages words.We construct word graph from an external large corpus to compensate for the lack of semantic information, and learn text graph to handle the lack of labeled data.Results show that Sim-pleSTC obtains state-of-the-art performance with lower memory consumption and faster inference speed. 1
Kaixin Zheng, Yaqing Wang 0002, Quanming Yao, Dejing Dou
EMNLP2
2022 Generative Time Series Forecasting with Diffusion, Denoise, and Disentanglement
abstract
Time series forecasting has been a widely explored task of great importance in many applications. However, it is common that real-world time series data are recorded in a short time period, which results in a big gap between the deep model and the limited and noisy time series. In this work, we propose to address the time series forecasting problem with generative modeling and propose a bidirectional variational auto-encoder (BVAE) equipped with diffusion, denoise, and disentanglement, namely D3VAE. Specifically, a coupled diffusion probabilistic model is proposed to augment the time series data without increasing the aleatoric uncertainty and implement a more tractable inference process with BVAE. To ensure the generated series move toward the true target, we further propose to adapt and integrate the multiscale denoising score matching into the diffusion process for time series forecasting. In addition, to enhance the interpretability and stability of the prediction, we treat the latent variable in a multivariate manner and disentangle them on top of minimizing total correlation. Extensive experiments on synthetic and real-world data show that D3VAE outperforms competitive algorithms with remarkable margins. Our implementation is available at https://github.com/PaddlePaddle/PaddleSpatial/tree/main/research/D3VAE.
Xinjiang Lu, Yaqing Wang 0002, Dejing Dou
NeurIPS3
2022 Recognizing Medical Search Query Intent by Few-shot Learning
abstract
Online healthcare services can provide unlimited and in-time medical information to users, which promotes social goods and breaks the barriers of locations. However, understanding the user intents behind the medical related queries is a challenging problem. Medical search queries are usually short and noisy, lack strict syntactic structure, and also require professional background to understand the medical terms. The medical intents are fine-grained, making them hard to recognize. In addition, many intents only have a few labeled data. To handle these problems, we propose a few-shot learning method for medical search query intent recognition called MEDIC. We extract co-click queries from user search logs as weak supervision to compensate for the lack of labeled data. We also design a new query encoder which learns to represent queries as a combination of semantic knowledge recorded in an external medical knowledge graph, syntactic knowledge which marks the grammatical role of each word in the query, and generic knowledge which is captured by language models pretrained from large-scale text corpus. Experimental results on a real medical search query intent recognition dataset validate the effectiveness of MEDIC.
Yaqing Wang 0002, Dejing Dou
SIGIR1
2022 Low-rank Tensor Learning with Nonconvex Overlapped Nuclear Norm Regularization
abstract
Nonconvex regularization has been popularly used in low-rank matrix learning. However, extending it for low-rank tensor learning is still computationally expensive. To address this problem, we develop an efficient solver for use with a nonconvex extension of the overlapped nuclear norm regularizer. Based on the proximal average algorithm, the proposed algorithm can avoid expensive tensor folding/unfolding operations. A special “sparse plus low-rank" structure is maintained throughout the iterations, and allows fast computation of the individual proximal steps. Empirical convergence is further improved with the use of adaptive momentum. We provide convergence guarantees to critical points on smooth losses and also on objectives satisfying the Kurdyka-Lojasiewicz condition. While the optimization problem is nonconvex and nonsmooth, we show that its critical points still have good statistical performance on the tensor completion problem. Experiments on various synthetic and real-world data sets show that the proposed algorithm is efficient in both time and space and more accurate than the existing state-of-the-art.
Quanming Yao, Yaqing Wang 0002, Bo Han 0003, James T. Kwok
J. Mach. Learn. Res.2
2022 Exploring the common principal subspace of deep features in neural networks
Haoyi Xiong, Yaqing Wang 0002, Haozhe An, Dejing Dou, Dongrui Wu
Mach. Learn.3
2021 Hierarchical Heterogeneous Graph Representation Learning for Short Text Classification
abstract
Short text classification is a fundamental task in natural language processing.It is hard due to the lack of context information and labeled data in practice.In this paper, we propose a new method called SHINE, which is based on graph neural network (GNN), for short text classification.First, we model the short text dataset as a hierarchical heterogeneous graph consisting of word-level component graphs which introduce more semantic and syntactic information.Then, we dynamically learn a short document graph that facilitates effective label propagation among similar short texts.Thus, comparing with existing GNN-based methods, SHINE can better exploit interactions between nodes of the same types and capture similarities between short texts.Extensive experiments on various benchmark short text datasets show that SHINE consistently outperforms state-of-the-art methods, especially with fewer labels. 1
Yaqing Wang 0002, Quanming Yao, Dejing Dou
EMNLP (1)1
2021 Property-Aware Relation Networks for Few-Shot Molecular Property Prediction
abstract
Molecular property prediction plays a fundamental role in drug discovery to identify candidate molecules with target properties. However, molecular property prediction is essentially a few-shot problem, which makes it hard to use regular machine learning models. In this paper, we propose Property-Aware Relation networks (PAR) to handle this problem. In comparison to existing works, we leverage the fact that both relevant substructures and relationships among molecules change across different molecular properties. We first introduce a property-aware embedding function to transform the generic molecular embeddings to substructure-aware space relevant to the target property. Further, we design an adaptive relation graph learning module to jointly estimate molecular relation graph and refine molecular embeddings w.r.t. the target property, such that the limited labels can be effectively propagated among similar molecules. We adopt a meta-learning strategy where the parameters are selectively updated within tasks in order to model generic and property-aware knowledge separately. Extensive experiments on benchmark molecular property prediction datasets show that PAR consistently outperforms existing methods and can obtain property-aware molecular embeddings and model molecular relation graph properly.
Yaqing Wang 0002, Abulikemu Abuduweili, Quanming Yao, Dejing Dou
NeurIPS1
2021 A Scalable, Adaptive and Sound Nonconvex Regularizer for Low-rank Matrix Learning
abstract
Matrix learning is at the core of many machine learning problems. A number of real-world applications such as collaborative filtering and text mining can be formulated as a low-rank matrix completion problems, which recovers incomplete matrix using low-rank assumptions. To ensure that the matrix solution has a low rank, a recent trend is to use nonconvex regularizers that adaptively penalize singular values. They offer good recovery performance and have nice theoretical properties, but are computationally expensive due to repeated access to individual singular values. In this paper, based on the key insight that adaptive shrinkage on singular values improve empirical performance, we propose a new nonconvex low-rank regularizer called ”nuclear norm minus Frobenius norm” regularizer, which is scalable, adaptive and sound. We first show it provably holds the adaptive shrinkage property. Further, we discover its factored form which bypasses the computation of singular values and allows fast optimization by general optimization algorithms. Stable recovery and convergence are guaranteed. Extensive low-rank matrix completion experiments on a number of synthetic and real-world data sets show that the proposed method obtains state-of-the-art recovery performance while being the fastest in comparison to existing low-rank matrix learning methods. 1
Yaqing Wang 0002, Quanming Yao, James T. Kwok
WWW1
2020 Generalized Convolutional Sparse Coding With Unknown Noise
abstract
Convolutional sparse coding (CSC) can learn representative shift-invariant patterns from multiple kinds of data. However, existing CSC methods can only model noises from Gaussian distribution, which is restrictive and unrealistic. In this paper, we propose a generalized CSC model capable of dealing with complicated unknown noise. The noise is now modeled by Gaussian mixture model, which can approximate any continuous probability density function. We use the expectation-maximization algorithm to solve the problem and design an efficient method for the weighted CSC problem in maximization step. The crux is to speed up the convolution in the frequency domain while keeping the other computations involving weight matrix in the spatial domain. Besides, we simultaneously update the dictionary and codes by nonconvex accelerated proximal gradient algorithm without bringing in extra alternating loops. The resultant method, called generalized convolutional sparse coding (GCSC), obtains the same space complexity and a smaller running time compared with existing CSC methods. Extensive experiments on synthetic and real-world noisy data sets validate that GCSC can model noise effectively and obtain high-quality filters and representations.
Yaqing Wang 0002, James T. Kwok, Lionel M. Ni
IEEE Trans. Image Process.1
2018 Online Convolutional Sparse Coding with Sample-Dependent Dictionary
abstract
Convolutional sparse coding (CSC) has been popularly used for the learning of shift-invariant dictionaries in image and signal processing. However, existing methods have limited scalability. In this paper, instead of convolving with a dictionary shared by all samples, we propose the use of a sample-dependent dictionary in which each filter is a linear combination of a small set of base filters learned from data. This added flexibility allows a large number of sample-dependent patterns to be captured, which is especially useful in the handling of large or high-dimensional data sets. Computationally, the resultant model can be efficiently learned by online learning. Extensive experimental results on a number of data sets show that the proposed method outperforms existing CSC algorithms with significantly reduced time and space complexities.
Yaqing Wang 0002, Quanming Yao, James T. Kwok, Lionel M. Ni
ICML1
2018 Scalable Online Convolutional Sparse Coding
abstract
Convolutional sparse coding (CSC) improves sparse coding by learning a shift-invariant dictionary from the data. However, most existing CSC algorithms operate in the batch mode and are computationally expensive. In this paper, we alleviate this problem by online learning. The key is a reformulation of the CSC objective so that convolution can be handled easily in the frequency domain, and much smaller history matrices are needed. To solve the resultant optimization problem, we use the alternating direction method of multipliers (ADMMs), and its subproblems have efficient closed-form solutions. Theoretical analysis shows that the learned dictionary converges to a stationary point of the optimization problem. Extensive experiments are performed on both the standard CSC benchmark data sets and much larger data sets such as the ImageNet. Results show that the proposed algorithm outperforms the state-of-the-art batch and online CSC methods. It is more scalable, has faster convergence, and better reconstruction performance.
Yaqing Wang 0002, Quanming Yao, James T. Kwok, Lionel M. Ni
IEEE Trans. Image Process.1
2017 Zero-shot learning with a partial set of observed attributes
abstract
Attributes are human-annotated semantic descriptions of label classes. In zero-shot learning (ZSL), they are often used to construct a semantic embedding for knowledge transfer from known classes to new classes. While collecting all attributes for the new classes is criticized as expensive, a subset of these attributes are often easy to acquire. In this paper, we extend ZSL methods to handle this partial set of observed attributes. We first recover the missing attributes through structured matrix completion. We use the low-rank assumption, and leverage properties of the attributes by extracting their rich semantic information from external sources. The resultant optimization problem can be efficiently solved with alternating minimization, in which each of its subproblems has a simple closed-form solution. The predicted attributes can then be used as semantic embeddings in ZSL. Experimental results show that the proposed method outperform existing methods in recovering the structured missing matrix. Moreover, methods using our predicted attributes in ZSL outperforms methods using either the partial set of observed attributes or other semantic embeddings.
Yaqing Wang 0002, James T. Kwok, Quanming Yao, Lionel M. Ni
IJCNN1