VLDB 2026 Research / reviewers in the wild / expert
Yu Zhang 0006
dblp:50/671-6
· DBLP profile ↗
129ranked-venue papers
27as first author
67since 2021 · last 2026
0000-0003-1100-4835ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 105 · 22 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 36 · 14 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph2Video: Leveraging Video Models to Model Dynamic Graph EvolutionabstractDynamic graphs are common in real‑world systems such as social media, recommender systems, and traffic networks. Existing dynamic graph models for link prediction often fall short in capturing the full complexity of temporal evolution. They tend to overlook fine‑grained variations in interaction order, struggle with dependencies that span long time horizons, and provide limited modeling of pair‑specific relational dynamics. To address those challenges, we propose Graph2Video, a video‑inspired framework that views the temporal neighborhood of a target link as a sequence of “graph frames”. By stacking temporally ordered subgraph frames into a “graph video”, Graph2Video leverages the inductive biases of video foundation models to capture both fine-grained local variations and long-range temporal dynamics. It generates a link-level embedding that serves as a lightweight, plug-and-play, link-centric memory unit. This embedding integrates seamlessly into existing dynamic graph encoders, effectively addressing the limitations of prior approaches. Extensive experiments on benchmark datasets show that Graph2Video outperforms state‑of‑the‑art baselines in the link prediction task on most cases. The results highlight that borrowing spatio‑temporal modeling techniques from computer vision provides a principled and effective avenue for advancing dynamic graph learning. Hua Liu 0008, Yanbin Wei, Tyler Derr, Haoyu Han 0001, Yu Zhang 0006 |
AAAI | 6 |
| 2026 | Dual-balancing for multi-task learning
Baijiong Lin, Weisen Jiang, Feiyang Ye 0001, Yu Zhang 0006, Pengguang Chen, Ying-Cong Chen, Shu Liu 0005, Ivor W. Tsang, James T. Kwok |
Neural Networks | 4 |
| 2026 | Mixture of Cluster-Conditional LoRA Experts for Vision-Language Instruction TuningabstractInstruction tuning of Large Vision-language Models (LVLMs) has revolutionized the development of versatile models with zero-shot generalization across a wide range of downstream vision-language tasks. However, the diversity of different training tasks from various sources and formats would lead to inevitable task conflicts, where different tasks conflict for the same set of model parameters, resulting in sub-optimal instruction-following abilities. To address that, we propose the Mixture of Cluster-conditional LoRA Experts (MoCLE), a novel Mixture of Experts (MoE) architecture designed to activate task-customized model parameters based on instruction clusters. A separate universal expert is further incorporated to improve generalization abilities of MoCLE for novel instructions. Extensive experiments on InstructBLIP and LLaVA demonstrate the effectiveness of MoCLE. Yunhao Gou, Zhili Liu, Kai Chen 0023, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang 0006 |
IEEE Trans. Image Process. | 9 |
| 2026 | KICGPTv2: Large Language Model With Knowledge in Context for Knowledge Graph CompletionabstractKnowledge Graph Completion (KGC) is an essential task aimed at mitigating the issue of incompleteness in knowledge graphs, thereby enhancing their utility for various downstream applications. Existing KGC models predominantly fall into two categories: structure-based and semantic-based approaches. Structure-based methods often encounter challenges with long-tail entities due to the scarcity of structural information and imbalanced entity distributions. Conversely, semantic-based methods, while addressing those limitations, necessitate extensive training of language models and specific finetuning for each knowledge graph, thus constraining their practical efficiency. To alleviate those limitations in both approaches, in this paper, we propose KICGPTv2, an innovative framework that synergizes a large language model (LLM) with traditional KGC methods. This integration effectively mitigates the long-tail entity problem without incurring significant additional training overhead. Central to the KICGPTv2 model is a novel in-context learning strategy, termed Knowledge Prompt, which encodes structural knowledge into demonstrations to effectively guide the LLM. Comprehensive evaluations on various KGC tasks, including link prediction, relation prediction, and triple classification, underscore the efficacy of the KICGPTv2 model, highlighting its ability to achieve competitive performance with reduced training demands and without the need for finetuning Yanbin Wei, Qiushi Huang, James T. Kwok, Yu Zhang 0006 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2026 | MoPD: Mixture-of-Prompts Distillation for Vision-Language ModelsabstractSoft prompt learning methods are effective for adapting vision-language models (VLMs) to downstream tasks. Nevertheless, empirical evidence reveals that existing methods tend to overfit seen classes and exhibit degraded performance on unseen classes. This limitation is due to the inherent bias in the training data towards the seen classes. To address this issue, we propose a novel soft prompt learning method, named Mixture-of-Prompts Distillation (MoPD), which can effectively transfer useful knowledge from hard prompts manually hand-crafted (a.k.a. teacher prompts) to the learnable soft prompt (a.k.a. student prompt), thereby enhancing the generalization ability of soft prompts on unseen classes. Moreover, the proposed MoPD method utilizes a gating network that learns to select hard prompts used for prompt distillation. Extensive experiments demonstrate that the proposed MoPD method outperforms state-of-the-art baselines, especially on unseen classes. Yang Chen 0031, Yu Zhang 0006 |
IEEE Trans. Multim. | 3 |
| 2025 | Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-AlignmentabstractZhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James Kwok. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhili Liu, Yunhao Gou, Kai Chen 0023, Lanqing Hong, Jiahui Gao 0002, Fei Mi, Yu Zhang 0006, Zhenguo Li, Xin Jiang 0002, Qun Liu 0001, James T. Kwok |
ACL (1) | 7 |
| 2025 | Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction TuningabstractYunhao Gou, Hansi Yang, Zhili Liu, Kai Chen, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu, Bo Han, James Kwok, Yu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yunhao Gou, Hansi Yang, Zhili Liu, Kai Chen 0023, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu 0001, Bo Han 0003, James T. Kwok, Yu Zhang 0006 |
EMNLP | 11 |
| 2025 | Sharpness-Aware Black-Box OptimizationabstractBlack-box optimization algorithms have been widely used in various machine learning problems, including reinforcement learning and prompt fine-tuning. However, directly optimizing the training loss value, as commonly done in existing black-box optimization methods, could lead to suboptimal model quality and generalization performance. To address those problems in black-box optimization, we propose a novel Sharpness-Aware Black-box Optimization (SABO) algorithm, which applies a sharpness-aware minimization strategy to improve the model generalization. Specifically, the proposed SABO method first reparameterizes the objective function by its expectation over a Gaussian distribution. Then it iteratively updates the parameterized distribution by approximated stochastic gradients of the maximum objective value within a small neighborhood around the current solution in the Gaussian distribution space. Theoretically, we prove the convergence rate and generalization bound of the proposed SABO algorithm. Empirically, extensive experiments on the black-box prompt fine-tuning tasks demonstrate the effectiveness of the proposed SABO method in improving model generalization performance. Feiyang Ye 0001, Yueming Lyu, Xuehao Wang, Masashi Sugiyama, Yu Zhang 0006, Ivor W. Tsang |
ICLR | 5 |
| 2025 | ComLoRA: A Competitive Learning Approach for Enhancing LoRAabstractWe propose a Competitive Low-Rank Adaptation (ComLoRA) framework to address the limitations of the LoRA method, which either lacks capacity with a single rank-$r$ LoRA or risks inefficiency and overfitting with a larger rank-$Kr$ LoRA, where $K$ is an integer larger than 1. The proposed ComLoRA method initializes $K$ distinct LoRA components, each with rank $r$, and allows them to compete during training. This competition drives each LoRA component to outperform the others, improving overall model performance. The best-performing LoRA is selected based on validation metrics, ensuring that the final model outperforms a single rank-$r$ LoRA and matches the effectiveness of a larger rank-$Kr$ LoRA, all while avoiding extra computational overhead during inference. To the best of our knowledge, this is the first work to introduce and explore competitive learning in the context of LoRA optimization. The ComLoRA's code is available at https://github.com/hqsiswiliam/comlora. Qiushi Huang, Tom Ko, Lilian Tang, Yu Zhang 0006 |
ICLR | 4 |
| 2025 | HiRA: Parameter-Efficient Hadamard High-Rank Adaptation for Large Language ModelsabstractWe propose Hadamard High-Rank Adaptation (HiRA), a parameter-efficient fine-tuning (PEFT) method that enhances the adaptability of Large Language Models (LLMs). While Low-rank Adaptation (LoRA) is widely used to reduce resource demands, its low-rank updates may limit its expressiveness for new tasks. HiRA addresses this by using a Hadamard product to retain high-rank update parameters, improving the model capacity. Empirically, HiRA outperforms LoRA and its variants on several tasks, with extensive ablation studies validating its effectiveness. Our code is available at https://github.com/hqsiswiliam/hira. Qiushi Huang, Tom Ko, Zhan Zhuang, Lilian Tang, Yu Zhang 0006 |
ICLR | 5 |
| 2025 | MTSAM: Multi-Task Fine-Tuning for Segment Anything ModelabstractThe Segment Anything Model (SAM), with its remarkable zero-shot capability, has the potential to be a foundation model for multi-task learning. However, adopting SAM to multi-task learning faces two challenges: (a) SAM has difficulty generating task-specific outputs with different channel numbers, and (b) how to fine-tune SAM to adapt multiple downstream tasks simultaneously remains unexplored. To address these two challenges, in this paper, we propose the Multi-Task SAM (MTSAM) framework, which enables SAM to work as a foundation model for multi-task learning. MTSAM modifies SAM's architecture by removing the prompt encoder and implementing task-specific no-mask embeddings and mask decoders, enabling the generation of task-specific outputs. Furthermore, we introduce Tensorized low-Rank Adaptation (ToRA) to perform multi-task fine-tuning on SAM. Specifically, ToRA injects an update parameter tensor into each layer of the encoder in SAM and leverages a low-rank tensor decomposition method to incorporate both task-shared and task-specific information.
Extensive experiments conducted on benchmark datasets substantiate the efficacy of MTSAM in enhancing the performance of multi-task learning. Our code is available at https://github.com/XuehaoWangFi/MTSAM. Xuehao Wang, Zhan Zhuang, Feiyang Ye 0001, Yu Zhang 0006 |
ICLR | 4 |
| 2025 | Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link PredictionabstractMessage-passing graph neural networks (MPNNs) and structural features (SFs) are cornerstones for the link prediction task. However, as a common and intuitive mode of understanding, the potential of visual perception has been overlooked in the MPNN community. For the first time, we equip MPNNs with vision structural awareness by proposing an effective framework called Graph Vision Network (GVN), along with a more efficient variant (E-GVN). Extensive empirical results demonstrate that with the proposed frameworks, GVN consistently benefits from the vision enhancement across seven link prediction datasets, including challenging large-scale graphs. Such improvements are compatible with existing state-of-the-art (SOTA) methods and GVNs achieve new SOTA results, thereby underscoring a promising novel direction for link prediction. Yanbin Wei, Xuehao Wang, Zhan Zhuang, Yang Chen 0031, Shuhao Chen, Yulong Zhang 0005, James T. Kwok, Yu Zhang 0006 |
ICML | 8 |
| 2025 | Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank AdaptationabstractLow-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters’ activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter’s marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto. Zhan Zhuang, Xiequn Wang, Yulong Zhang 0005, Qiushi Huang, Shuhao Chen, Xuehao Wang, Yanbin Wei, Yuhe Nie, Kede Ma, Yu Zhang 0006, Ying Wei 0001 |
ICML | 11 |
| 2025 | Domain-guided conditional diffusion model for unsupervised domain adaptation
Yulong Zhang 0005, Shuhao Chen, Weisen Jiang, Yu Zhang 0006, Jiangang Lu, James T. Kwok |
Neural Networks | 4 |
| 2025 | Online Test-Time Adaptation of Spatial-Temporal Traffic Flow ForecastingabstractAccurate spatial-temporal traffic flow forecasting is crucial in aiding traffic managers in implementing control measures and assisting drivers in selecting optimal travel routes. Traditional deep-learning based methods for traffic flow forecasting typically rely on historical data to train their models, which are then used to make predictions on future data. However, the performance of the trained model usually degrades due to the temporal drift between the historical and future data. To make the model trained on historical data better adapt to future data in a fully online manner, this paper conducts the first study of the online test-time adaptation techniques for spatial-temporal traffic flow forecasting problems. To this end, we propose anAdaptiveDoubleCorrection bySeriesDecomposition (ADCSD) method, which first decomposes the output of the trained model into seasonal and trend-cyclical parts and then corrects them by two separate modules during the testing phase using the latest observed dataentry by entry. In the proposed ADCSD method, instead of fine-tuning the whole trained model during the testing phase, a lite network is attached after the trained model, and only the lite network is fine-tuned in the testing process each time a data entry is observed. Moreover, to satisfy that different time series variables may have different levels of temporal drift, two adaptive vectors are adopted to provide different weights for different time series variables. Extensive experiments on four real-world traffic flow forecasting datasets demonstrate the effectiveness of the proposed ADCSD method. The code is available athttps://github.com/Pengxin-Guo/ADCSD Pengxin Guo 0001, Pengrong Jin, Ziyue Li 0002, Lei Bai 0001, Yu Zhang 0006 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | A First-Order Multi-Gradient Algorithm for Multi-Objective Bi-Level OptimizationabstractIn this paper, we study the Multi-Objective Bi-Level Optimization (MOBLO) problem, where the upper-level subproblem is a multi-objective optimization problem and the lower-level subproblem is for scalar optimization. Existing gradient-based MOBLO algorithms need to compute the Hessian matrix, causing the computational inefficient problem. To address this, we propose an efficient first-order multi-gradient method for MOBLO, called FORUM. Specifically, we reformulate MOBLO problems as a constrained multi-objective optimization (MOO) problem via the value-function approach. Then we propose a novel multi-gradient aggregation method to solve the challenging constrained MOO problem. Theoretically, we provide the complexity analysis to show the efficiency of the proposed method and a non-asymptotic convergence result. Empirically, extensive experiments demonstrate the effectiveness and efficiency of the proposed FORUM method in different learning problems. In particular, it achieves state-of-the-art performance on three multi-task learning benchmark datasets. The code is available at https://github.com/Baijiong-Lin/FORUM. Feiyang Ye 0001, Baijiong Lin, Xiaofeng Cao 0002, Yu Zhang 0006, Ivor W. Tsang |
ECAI | 4 |
| 2024 | Eyes Closed, Safety on: Protecting Multimodal LLMs via Image-to-Text Transformation
Yunhao Gou, Kai Chen 0023, Zhili Liu, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang 0006 |
ECCV (17) | 9 |
| 2024 | MTMamba: Enhancing Multi-task Dense Scene Understanding by Mamba-Based Decoders
Baijiong Lin, Weisen Jiang, Pengguang Chen, Yu Zhang 0006, Shu Liu 0005, Ying-Cong Chen |
ECCV (70) | 4 |
| 2024 | Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language ModelsabstractWith the prevalence of large-scale pretrained vision-language models (VLMs), such as CLIP, soft-prompt tuning has become a popular method for adapting these models to various downstream tasks. However, few works delve into the inherent properties of learnable soft-prompt vectors, specifically the impact of their norms to the performance of VLMs. This motivates us to pose an unexplored research question: ``Do we need to normalize the soft prompts in VLMs?'' To fill this research gap, we first uncover a phenomenon, called the $\textbf{Low-Norm Effect}$ by performing extensive corruption experiments, suggesting that reducing the norms of certain learned prompts occasionally enhances the performance of VLMs, while increasing them often degrades it. To harness this effect, we propose a novel method named $\textbf{N}$ormalizing th$\textbf{e}$ soft-pro$\textbf{m}$pt v$\textbf{e}$ctors of vi$\textbf{si}$on-language model$\textbf{s}$ ($\textbf{Nemesis}$) to normalize soft-prompt vectors in VLMs. To the best of our knowledge, our work is the first to systematically investigate the role of norms of soft-prompt vector in VLMs, offering valuable insights for future research in soft-prompt tuning. Xiequn Wang, Qiushi Huang, Yu Zhang 0006 |
ICLR | 4 |
| 2024 | Adaptive Stochastic Gradient Algorithm for Black-box Multi-Objective LearningabstractMulti-objective optimization (MOO) has become an influential framework for various machine learning problems, including reinforcement learning and multi-task learning. In this paper, we study the black-box multi-objective optimization problem, where we aim to optimize multiple potentially conflicting objectives with function queries only. To address this challenging problem and find a Pareto optimal solution or the Pareto stationary solution,
we propose a novel adaptive stochastic gradient algorithm for black-box MOO, called ASMG.
Specifically, we use the stochastic gradient approximation method to obtain the gradient for the distribution parameters of the Gaussian smoothed MOO with function queries only. Subsequently, an adaptive weight is employed to aggregate all stochastic gradients to optimize all objective functions effectively.
Theoretically, we explicitly provide the connection between the original MOO problem and the corresponding Gaussian smoothed MOO problem and prove the convergence rate for the proposed ASMG algorithm in both convex and non-convex scenarios.
Empirically, the proposed ASMG method achieves competitive performance on multiple numerical benchmark problems. Additionally, the state-of-the-art performance on the black-box multi-task learning problem demonstrates the effectiveness of the proposed ASMG method. Feiyang Ye 0001, Yueming Lyu, Xuehao Wang, Yu Zhang 0006, Ivor W. Tsang |
ICLR | 4 |
| 2024 | MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsabstractLarge language models (LLMs) have pushed the limits of natural language understanding and exhibited excellent problem-solving ability. Despite the great success, most existing open-source LLMs (\eg, LLaMA-2) are still far away from satisfactory for solving mathematical problems due to the complex reasoning procedures. To bridge this gap, we propose \emph{MetaMath}, a finetuned language model that specializes in mathematical reasoning. Specifically, we start by bootstrapping mathematical questions by rewriting the question from multiple perspectives, which results in a new dataset called MetaMathQA. Then we finetune the LLaMA-2 models on MetaMathQA. Experimental results on two popular benchmarks (\ie, GSM8K and MATH) for mathematical reasoning demonstrate that MetaMath outperforms a suite of open-source LLMs by a significant margin. Our MetaMath-7B model achieves $66.5\%$ on GSM8K and $19.8\%$ on MATH, exceeding the state-of-the-art models of the same size by $11.5\%$ and $8.7\%$. Particularly, MetaMath-70B achieves an accuracy of $82.3\%$ on GSM8K, slightly better than GPT-3.5-Turbo. We release the MetaMathQA dataset, the MetaMath models with different model sizes and the training code for public use. Longhui Yu, Weisen Jiang, Zhengying Liu, Yu Zhang 0006, James T. Kwok, Zhenguo Li, Adrian Weller, Weiyang Liu |
ICLR | 6 |
| 2024 | Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding PerspectiveabstractEmpirical Risk Minimization (ERM) is fragile in scenarios with insufficient labeled samples. A vanilla extension of ERM to unlabeled samples is Entropy Minimization (EntMin), which employs the soft-labels of unlabeled samples to guide their learning. However, EntMin emphasizes prediction discriminability while neglecting prediction diversity. To alleviate this issue, in this paper, we rethink the guidance information to utilize unlabeled samples. By analyzing the learning objective of ERM, we find that the guidance information for labeled samples in a specific category is the corresponding label encoding. Inspired by this finding, we propose a Label-Encoding Risk Minimization (LERM). It first estimates the label encodings through prediction means of unlabeled samples and then aligns them with their corresponding ground-truth label encodings. As a result, the LERM ensures both prediction discriminability and diversity, and it can be integrated into existing methods as a plugin. Theoretically, we analyze the relationships between LERM and ERM as well as EntMin. Empirically, we verify the superiority of the LERM under several label insufficient scenarios. The codes are available at https://github.com/zhangyl660/LERM. Yulong Zhang 0005, Yuan Yao 0016, Shuhao Chen, Pengrong Jin, Yu Zhang 0006, Jiangang Lu |
ICML | 5 |
| 2024 | RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsabstractRecent works show that assembling multiple off-the-shelf large language models (LLMs) can harness their complementary abilities. To achieve this, routing is a promising method, which learns a router to select the most suitable LLM for each query. However, existing routing models are ineffective when multiple LLMs perform well for a query. To address this problem, in this paper, we propose a method called query-based Router by Dual Contrastive learning (RouterDC). The RouterDC model, which consists of an encoder and LLM embeddings, is trained by two proposed contrastive losses (sample-LLM and sample-sample losses). Experimental results show that RouterDC is effective in assembling LLMs and largely outperforms individual top-performing LLMs as well as existing routing methods on both in-distribution (+2.76\%) and out-of-distribution (+1.90\%) tasks. The source code is available at https://github.com/shuhao02/RouterDC. Shuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok, Yu Zhang 0006 |
NeurIPS | 5 |
| 2024 | GITA: Graph to Visual and Textual Integration for Vision-Language Graph ReasoningabstractLarge Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reasoning. The potential benefits and capabilities of representing graph structures as visual images (i.e., $\textit{visual graph}$) are still unexplored. To fill the gap, we innovatively propose an end-to-end framework, called $\textbf{G}$raph to v$\textbf{I}$sual and $\textbf{T}$extual Integr$\textbf{A}$tion (GITA), which firstly incorporates visual graphs into general graph reasoning. Besides, we establish $\textbf{G}$raph-based $\textbf{V}$ision-$\textbf{L}$anguage $\textbf{Q}$uestion $\textbf{A}$nswering (GVLQA) dataset from existing graph data, which is the first vision-language dataset for general graph reasoning purposes. Extensive experiments on the GVLQA dataset and five real-world datasets show that GITA outperforms mainstream LLMs in terms of general graph reasoning capabilities. Moreover, We highlight the effectiveness of the layout augmentation on visual graphs and pretraining on the GVLQA dataset. Yanbin Wei, Weisen Jiang, Zejian Zhang, Zhixiong Zeng, James T. Kwok, Yu Zhang 0006 |
NeurIPS | 8 |
| 2024 | Time-Varying LoRA: Towards Effective Cross-Domain Fine-Tuning of Diffusion ModelsabstractLarge-scale diffusion models are adept at generating high-fidelity images and facilitating image editing and interpolation. However, they have limitations when tasked with generating images in dynamic, evolving domains. In this paper, we introduce Terra, a novel Time-varying low-rank adapter that offers a fine-tuning framework specifically tailored for domain flow generation. The key innovation of Terra lies in its construction of a continuous parameter manifold through a time variable, with its expressive power analyzed theoretically. This framework not only enables interpolation of image content and style but also offers a generation-based approach to address the domain shift problems in unsupervised domain adaptation and domain generalization. Specifically, Terra transforms images from the source domain to the target domain and generates interpolated domains with various styles to bridge the gap between domains and enhance the model generalization, respectively. We conduct extensive experiments on various benchmark datasets, empirically demonstrate the effectiveness of Terra. Our source code is publicly available on https://github.com/zwebzone/terra. Zhan Zhuang, Yulong Zhang 0005, Xuehao Wang, Jiangang Lu, Ying Wei 0001, Yu Zhang 0006 |
NeurIPS | 6 |
| 2024 | Enhancing Sharpness-Aware Minimization by Learning Perturbation Radius
Xuehao Wang, Weisen Jiang, Yu Zhang 0006 |
ECML/PKDD (2) | 4 |
| 2024 | Multi-objective meta-learning
Feiyang Ye 0001, Baijiong Lin, Zhixiong Yue, Yu Zhang 0006, Ivor W. Tsang |
Artif. Intell. | 4 |
| 2024 | Selective Random Walk for Transfer Learning in Heterogeneous Label SpacesabstractTransfer learning has been widely used in different scenarios, especially in those lacking enough labeled data. However, most of the existing transfer learning methods are based on the assumption that the source and target domains should share the label space entirely or partially, which greatly limits their application scopes. In this article, a Selective Random Walk (SRW) method for transfer learning in heterogeneous label spaces is proposed to make full use of unlabeled auxiliary data, which acts as a bridge for knowledge transfer from the source domain to the target domain. The proposed SRW method can explicitly identify transfer sequences between source and target instances via auxiliary instances based on random walk techniques. Since not all of the transfer sequences generated by random walk are credible for the target task, the SRW method can learn to weight transfer sequences adaptively. Based on the weights of the transfer sequences, the SRW method leverages knowledge by forcing adjacent data points in the transfer sequence to be similar and making the target data point in the sequence represented by other data points in the same sequence. Experiments show that the SRW method outperforms state-of-the-art models in plenty of transfer learning tasks with heterogeneous label spaces constructed within and across several benchmark datasets. Qiao Xiao, Yu Zhang 0006, Qiang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | A Versatile Framework for Unsupervised Domain Adaptation Based on Instance WeightingabstractDespite the progress made in domain adaptation, solving Unsupervised Domain Adaptation (UDA) problems with a general method under complex conditions caused by label shifts between domains remains a challenging task. In this work, we comprehensively investigate four distinct UDA settings including closed set domain adaptation, partial domain adaptation, open set domain adaptation, and universal domain adaptation, where shared common classes between source and target domains coexist alongside domain-specific private classes. The prominent challenges inherent in diverse UDA settings center around the discrimination of common/private classes and the precise measurement of domain discrepancy. To surmount these challenges effectively, we propose a novel yet effective method called Learning Instance Weighting for Unsupervised Domain Adaptation (LIWUDA), which caters to various UDA settings. Specifically, the proposed LIWUDA method constructs a weight network to assign weights to each instance based on its probability of belonging to common classes, and designs Weighted Optimal Transport (WOT) for domain alignment by leveraging instance weights. Additionally, the proposed LIWUDA method devises a Separate and Align (SA) loss to separate instances with low similarities and align instances with high similarities. To guide the learning of the weight network, Intra-domain Optimal Transport (IOT) is proposed to enforce the weights of instances in common classes to follow a uniform distribution. Through the integration of those three components, the proposed LIWUDA method demonstrates its capability to address all four UDA settings in a unified manner. Experimental evaluations conducted on four benchmark datasets substantiate the effectiveness of the proposed LIWUDA method. The code is available at https://github.com/JinjingZhu/LIWUDA. Jinjing Zhu, Feiyang Ye 0001, Qiao Xiao, Pengxin Guo 0001, Yu Zhang 0006, Qiang Yang 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Personalized Dialogue Generation with Persona-Adaptive AttentionabstractPersona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs. Qiushi Huang, Yu Zhang 0006, Tom Ko, Xubo Liu 0001, Bo Wu 0018, Wenwu Wang 0001, Lilian Tang |
AAAI | 2 |
| 2023 | Dual-Path Side Information Fusion for Sequential RecommendationabstractSequential recommendations are designed to capture user preferences based on their past actions and predict the items they may interact with in the next moment. Benefiting from the self-attention mechanism, methods that utilize side information (such as item categories or brand) to improve the prediction performance of sequential recommendation have yielded promising results. Previous approaches typically directly fuses side information embeddings into item embeddings as inputs to the model. However, this fusion approach overlooks the distinctions in various types of information in sequential pattern inference, and also failing to fully model the relationship between items and side information. In this work, we propose a Dual-Path Side Information Fusion method (DPIF) to better utilize side information for improved recommendation performance. Our model employs two parallel paths for side information fusion modeling. One path obtains the relationship representation within the items and the side information, and the other path obtains the relationship representation between the items and the side information. Subsequently, an attention-based adaptive fusion module is utilized to combine inter-attribute relationship and intra-attribute relationship representation, generating the final user preferences. Extensive experiments were conducted on four real-world datasets, demonstrating the effectiveness of the introduced model. Our source code is available at https://github.com/ZhangYu-x/DPIF. Yu Zhang 0006, Haiwei Pan, Kejia Zhang 0001, Tianming Zhang, Qingquan Ren |
IEEE Big Data | 1 |
| 2023 | Leveraging per Image-Token Consistency for Vision-Language Pre-trainingabstractMost existing vision-language pre-training (VLP) approaches adopt cross-modal masked language modeling (CMLM) to learn vision-language associations. However, we find that CMLM is insufficient for this purpose according to our observations: (1) Modality bias: a considerable amount of masked tokens in CMLM can be recovered with only the language information, ignoring the visual inputs. (2) Underutilization of the unmasked tokens: CMLM primarily focuses on the masked token but it cannot simultaneously leverage other tokens to learn vision-language associations. To handle those limitations, we propose EPIC (lEveraging Per Image-Token Consistency for vision-language pre-training). In EPIC, for each image-sentence pair, we mask tokens that are salient to the image (i.e., Saliency-based Masking Strategy) and replace them with alternatives sampled from a language model (i.e., Inconsistent Token Generation Procedure), and then the model is required to determine for each token in the sentence whether it is consistent with the image (i.e., Image-Token Consistency Task). The proposed EPIC method is easily combined with pre-training methods. Extensive experiments show that the combination of the EPIC method and state-of-the-art pre-training approaches, including ViLT, ALBEF, METER, and X-VLM, leads to significant improvements on downstream tasks. Our coude is released at https://github.com/gyhdog99/epic Yunhao Gou, Tom Ko, Hansi Yang, James T. Kwok, Yu Zhang 0006, Mingxuan Wang |
CVPR | 5 |
| 2023 | Learning Retrieval Augmentation for Personalized Dialogue GenerationabstractPersonalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications.However, persona profiles, a prevalent setting in current personalized dialogue datasets, typically composed of merely four to five sentences, may not offer comprehensive descriptions of the persona about the agent, posing a challenge to generate truly personalized dialogues.To handle this problem, we propose Learning Retrieval Augmentation for Personalized DialOgue Generation (LAPDOG), which studies the potential of leveraging external knowledge for persona dialogue generation.Specifically, the proposed LAPDOG model consists of a story retriever and a dialogue generator.The story retriever uses a given persona profile as queries to retrieve relevant information from the story document, which serves as a supplementary context to augment the persona profile.The dialogue generator utilizes both the dialogue history and the augmented persona profile to generate personalized responses.For optimization, we adopt a joint training framework that collaboratively learns the story retriever and dialogue generator, where the story retriever is optimized towards desired ultimate metrics (e.g., BLEU) to retrieve content for the dialogue generator to generate personalized responses.Experiments conducted on the CONVAI2 dataset with ROCStory as a supplementary data source show that the proposed LAPDOG method substantially outperforms the baselines, indicating the effectiveness of the proposed method.The LAPDOG model code is publicly available for further exploration. Qiushi Huang, Xubo Liu 0001, Wenwu Wang 0001, Tom Ko, Yu Zhang 0006, Lilian Tang |
EMNLP | 6 |
| 2023 | An Adaptive Policy to Employ Sharpness-Aware Minimization
Weisen Jiang, Hansi Yang, Yu Zhang 0006, James T. Kwok |
ICLR | 3 |
| 2023 | Effective Structured Prompting by Meta-Learning and Representative VerbalizerabstractPrompt tuning for pre-trained masked language models (MLM) has shown promising performance in natural language processing tasks with few labeled examples. It tunes a prompt for the downstream task, and a verbalizer is used to bridge the predicted token and label prediction. Due to the limited training data, prompt initialization is crucial for prompt tuning. Recently, MetaPrompting (Hou et al., 2022) uses meta-learning to learn a shared initialization for all task-specific prompts. However, a single initialization is insufficient to obtain good prompts for all tasks and samples when the tasks are complex. Moreover, MetaPrompting requires tuning the whole MLM, causing a heavy burden on computation and memory as the MLM is usually large. To address these issues, we use a prompt pool to extract more task knowledge and construct instance-dependent prompts via attention. We further propose a novel soft verbalizer (RepVerb) which constructs label embedding from feature embeddings directly. Combining meta-learning the prompt pool and RepVerb, we propose MetaPrompter for effective structured prompting. MetaPrompter is parameter-efficient as only the pool is required to be tuned. Experimental results demonstrate that MetaPrompter performs better than the recent state-of-the-arts and RepVerb outperforms existing soft verbalizers. Weisen Jiang, Yu Zhang 0006, James T. Kwok |
ICML | 2 |
| 2023 | Multi-Task Learning via Time-Aware Neural ODEabstractMulti-Task Learning (MTL) is a well-established paradigm for learning shared models for a diverse set of tasks. Moreover, MTL improves data efficiency by jointly training all tasks simultaneously. However, directly optimizing the losses of all the tasks may lead to imbalanced performance on all the tasks due to the competition among tasks for the shared parameters in MTL models. Many MTL methods try to mitigate this problem by dynamically weighting task losses or manipulating task gradients. Different from existing studies, in this paper, we propose a Neural Ordinal diffeRential equation based Multi-tAsk Learning (NORMAL) method to alleviate this issue by modeling task-specific feature transformations from the perspective of dynamic flows built on the Neural Ordinary Differential Equation (NODE). Specifically, the proposed NORMAL model designs a time-aware neural ODE block to learn task-specific time information, which determines task positions of feature transformations in the dynamic flow, in NODE automatically via gradient descent methods. In this way, the proposed NORMAL model handles the problem of competing shared parameters by learning task positions. Moreover, the learned task positions can be used to measure the relevance among different tasks. Extensive experiments show that the proposed NORMAL model outperforms state-of-the-art MTL models. Feiyang Ye 0001, Xuehao Wang, Yu Zhang 0006, Ivor W. Tsang |
IJCAI | 3 |
| 2023 | Partially-Labeled Domain Generalization via Multi-Dimensional Domain AdaptationabstractDomain generalization deals with a challenging setting where several labeled source domains are given, and the goal is to train machine learning models that can generalize to an unseen test domain. However, in practice, labeled samples are often difficult and expensive to obtain. Thus the source domains would not always be labeled. When only some source domains are labeled and others are unlabeled, we formally introduce this domain generalization problem as Partially-Labeled Domain Generalization (PLDG). In this paper, we study the most chal- lenging setting in PLDG problems, where only one source domain is labeled and a few unlabeled source domains are available. To enable generalization, we assume that all source domains follow certain domain index information that can reflect their domain relationships. With this domain index information, we propose a Multi-Dimensional Domain Adaptation (MDDA) method to address this PLDG problem. Specifically, the MDDA method first trains multiple domain adaptation models to adapt from the labeled source domain to all the unlabeled source domains via adversarial learning. Then those domain adaptation models and the source-only model trained on the labeled source domain only are distilled into the target model used for the unseen target domain. Theoretically, we provide a generalization bound of the MDDA method. The experiments on four real-world datasets demonstrate the effectiveness of the proposed MDDA method. Feiyang Ye 0001, Jianghan Bao, Yu Zhang 0006 |
IJCNN | 3 |
| 2023 | Visually-Aware Audio Captioning With Adaptive Audio-Visual AttentionabstractAudio captioning aims to generate text descriptions of audio clips.In the real world, many objects produce similar sounds.How to accurately recognize ambiguous sounds is a major challenge for audio captioning.In this work, inspired by inherent human multimodal perception, we propose visuallyaware audio captioning, which makes use of visual information to help the description of ambiguous sounding objects.Specifically, we introduce an off-the-shelf visual encoder to extract video features and incorporate the visual features into an audio captioning system.Furthermore, to better exploit complementary audio-visual contexts, we propose an audio-visual attention mechanism that adaptively integrates audio and visual context and removes the redundant information in the latent space.Experimental results on AudioCaps, the largest audio captioning dataset, show that our proposed method achieves state-of-theart results on machine translation metrics. Xubo Liu 0001, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang 0006, Lilian Tang, Mark D. Plumbley, Volkan Kilic, Wenwu Wang 0001 |
INTERSPEECH | 9 |
| 2023 | Unsupervised Domain Adaptation via Bidirectional Cross-Attention Transformer
Pengxin Guo 0001, Yu Zhang 0006 |
ECML/PKDD (5) | 3 |
| 2023 | LibMTL: A Python Library for Deep Multi-Task LearningabstractThis paper presents LibMTL, an open-source Python library built on PyTorch, which provides a unified, comprehensive, reproducible, and extensible implementation framework for Multi-Task Learning (MTL). LibMTL considers different settings and approaches in MTL, and it supports a large number of state-of-the-art MTL methods, including 13 optimization strategies and 8 architectures. Moreover, the modular design in LibMTL makes it easy to use and well-extensible, thus users can easily and fast develop new MTL methods, compare with existing MTL methods fairly, or apply MTL algorithms to real-world applications with the support of LibMTL. The source code and detailed documentations of LibMTL are available at https://github.com/median-research-group/LibMTL and https://libmtl.readthedocs.io, respectively. Baijiong Lin, Yu Zhang 0006 |
J. Mach. Learn. Res. | 2 |
| 2023 | Superpixelwise Low-Rank Approximation-Based Partial Label Learning for Hyperspectral Image ClassificationabstractInsufficient prior knowledge of a captured hyperspectral image (HSI) scene may lead the experts or the automatic labeling systems to offer incorrect labels or ambiguous labels (i.e., assigning each training sample to a group of candidate labels, among which only one of them is valid; this is also known as partial label learning) during the labeling process. Accordingly, how to learn from such data with ambiguous labels is a problem of great practical importance. In this letter, we propose a novel superpixelwise low-rank approximation (LRA)-based partial label learning method, namely SLAP, which is the first to take into account partial label learning in HSI classification. SLAP is mainly composed of two phases: disambiguating the training labels and acquiring the predictive model. Specifically, in the first phase, we propose a superpixelwise LRA-based model, preparing the affinity graph for the subsequent label propagation process while extracting the discriminative representation to enhance the following classification task of the second phase. Then to disambiguate the training labels, label propagation propagates the labeling information via the affinity graph of training pixels. In the second phase, we take advantage of the resulting disambiguated training labels and the discriminative representations to enhance the classification performance. The extensive experiments validate the advantage of the proposed SLAP method over state-of-the-art methods. Shujun Yang, Yu Zhang 0006, Yao Ding 0010, Danfeng Hong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Learning Linear and Nonlinear Low-Rank Structure in Multi-Task LearningabstractAs the trace norm can discover low-rank structures in a matrix, it has been widely used in multi-task learning to recover the low-rank structure contained in the parameter matrix. Recently, with the emerging of big complex datasets and the popularity of deep learning techniques, tensor trace norms have been used for deep multi-task models. However, existing tensor trace norms exhibit some limitations. For example, they cannot discover all the low-rank structures in a tensor, they require users to manually specify the importance of each component in the corresponding tensor trace norm, and they only capture the linear low-rank structure. To solve the first issue, in this paper, we propose a Generalized Tensor Trace Norm (GTTN). The GTTN is defined as a convex combination of matrix trace norms of all possible tensor flattenings and hence it can discover all the possible low-rank structures. For the second issue, in the induced objective function with the GTTN, we propose four strategies to learn combination coefficients in the GTTN. Furthermore, we propose the Nonlinear GTTN (NGTTN) to capture nonlinear low-rank structure among all the tasks. Experiments on benchmark datasets demonstrate the effectiveness of the proposed GTTN and NGTTN. Yu Zhang 0006, Wei Wang 0028 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Multisource Heterogeneous Domain Adaptation With Conditional Weighting Adversarial NetworkabstractHeterogeneous domain adaptation (HDA) tackles the learning of cross-domain samples with both different probability distributions and feature representations. Most of the existing HDA studies focus on the single-source scenario. In reality, however, it is not uncommon to obtain samples from multiple heterogeneous domains. In this article, we study the multisource HDA problem and propose a conditional weighting adversarial network (CWAN) to address it. The proposed CWAN adversarially learns a feature transformer, a label classifier, and a domain discriminator. To quantify the importance of different source domains, CWAN introduces a sophisticated conditional weighting scheme to calculate the weights of the source domains according to the conditional distribution divergence between the source and target domains. Different from existing weighting schemes, the proposed conditional weighting scheme not only weights the source domains but also implicitly aligns the conditional distributions during the optimization process. Experimental results clearly demonstrate that the proposed CWAN performs much better than several state-of-the-art methods on four real-world datasets. Yuan Yao 0016, Xutao Li 0003, Yu Zhang 0006, Yunming Ye |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingabstractJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei |
ACL (1) | 10 |
| 2022 | Selective Partial Domain Adaptation
Pengxin Guo 0001, Jinjing Zhu, Yu Zhang 0006 |
BMVC | 3 |
| 2022 | Multi-View Self-Attention Based Transformer for Speaker RecognitionabstractInitially developed for natural language processing (NLP), Transformer model is now widely used for speech processing tasks such as speaker recognition, due to its powerful sequence modeling capabilities. However, conventional self-attention mechanisms are originally designed for modeling textual sequence without considering the characteristics of speech and speaker modeling. Besides, different Transformer variants for speaker recognition have not been well studied. In this work, we propose a novel multi-view self-attention mechanism and present an empirical study of different Transformer variants with or without the proposed attention mechanism for speaker recognition. Specifically, to balance the capabilities of capturing global dependencies and modeling the locality, we propose a multi-view self-attention mechanism for speaker Transformer, in which different attention heads can attend to different ranges of the receptive field. Furthermore, we introduce and compare five Transformer variants with different network architectures, embedding locations, and pooling methods to learn speaker embeddings. Experimental results on the VoxCeleb1 and VoxCeleb2 datasets show that the proposed multi-view self-attention mechanism achieves improvement in the performance of speaker recognition, and the proposed speaker Transformer network attains excellent results compared with state-of-the-art models. Rui Wang 0073, Junyi Ao, Shujie Liu 0001, Zhihua Wei 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006 |
ICASSP | 8 |
| 2022 | Exploring Machine Speech Chain For Domain AdaptationabstractMachine Speech Chain integrates both end-to-end (E2E) automatic speech recognition (ASR) and neural text-to-speech (TTS) into one circle for joint training. It has been proven that it can effectively leverage a large amount of unpaired data in the spirit of data augmentation. In this paper, we explore the TTS→ASR pipeline in machine speech chain to perform domain adaptation for both E2E ASR and neural TTS models with only text data from the target domain. We conduct experiments by adapting from audiobook domain (i.e., LibriSpeech) to presentation domain (i.e., TED-LIUM). There is a relative word error rate (WER) reduction of 19.7% for the E2E ASR model on the TED-LIUM test set, and a relative WER reduction of 29.4% in synthetic speech generated by neural TTS in the presentation domain. Moreover, we observe that the gains from the proposed method and conventional adaptation methods of language models are additive. Fengpeng Yue, Lei He 0005, Tom Ko, Yu Zhang 0006 |
ICASSP | 5 |
| 2022 | Subspace Learning for Effective Meta-LearningabstractMeta-learning aims to extract meta-knowledge from historical tasks to accelerate learning on new tasks. Typical meta-learning algorithms like MAML learn a globally-shared meta-model for all tasks. However, when the task environments are complex, task model parameters are diverse and a common meta-model is insufficient to capture all the meta-knowledge. To address this challenge, in this paper, task model parameters are structured into multiple subspaces, and each subspace represents one type of meta-knowledge. We propose an algorithm to learn the meta-parameters (\ie, subspace bases). We theoretically study the generalization properties of the learned subspaces. Experiments on regression and classification meta-learning datasets verify the effectiveness of the proposed algorithm. Weisen Jiang, James T. Kwok, Yu Zhang 0006 |
ICML | 3 |
| 2022 | Learning Feature Alignment Architecture for Domain AdaptationabstractIn domain adaptation, where the feature distributions of the source and target domains are different, various distance-based methods have been proposed to handle the domain shift by minimizing the discrepancy between the source and target domains. These methods use hand-crafted bottleneck networks, which might hinder the alignment of hidden feature representations extracted from both domains. In this paper, we propose a new method called Alignment Architecture Search with Population Correlation (AASPC) to automatically learn the architecture of the bottleneck network that can align the source and target domains. The proposed AASPC method introduces a new similarity function called Population Correlation (PC) to measure the domain discrepancy. The proposed AASPC method leverages PC to learn the alignment architecture and domaininvariant feature representation. Experiments on several benchmark datasets, including Office-31, Office-Home, and VisDA-2017, show the effectiveness of the proposed AASPC method. Zhixiong Yue, Pengxin Guo 0001, Yu Zhang 0006, Christy Jie Liang |
IJCNN | 3 |
| 2022 | Effective, Efficient and Robust Neural Architecture SearchabstractDesigning neural network architecture for embedded devices is practical but challenging because the models are expected to be not only accurate but also enough lightweight and robust. However, it is challenging to balance those trade-offs manually because of the large search space. To solve this problem, we propose an Effective, Efficient, and Robust Neural Architecture Search (E2RNAS) method to automatically search a neural network architecture that balances the performance, robustness, and resource consumption. Unlike previous studies, the objective function of the proposed E2RNAS method is formulated as a multi-objective bi-level optimization problem with the upper-level subproblem as a multi-objective optimization problem that considers the performance, robustness, and resource consumption. To solve the proposed objective function, we integrate the multiple-gradient descent algorithm, a widely studied gradient-based multi-objective optimization algorithm, with the bi-level optimization. Experiments on benchmark datasets show that the proposed E2RNAS method can find robust architecture with low resource consumption and comparable classification accuracy. Zhixiong Yue, Baijiong Lin, Yu Zhang 0006, Christy Jie Liang |
IJCNN | 3 |
| 2022 | A Study of Modeling Rising Intonation in Cantonese Neural Speech SynthesisabstractIn human speech, the attitude of a speaker cannot be fully expressed only by the textual content.It has to come along with the intonation.Declarative questions are commonly used in daily Cantonese conversations, and they are usually uttered with rising intonation.Vanilla neural text-to-speech (TTS) systems are not capable of synthesizing rising intonation for these sentences due to the loss of semantic information.Though it has become more common to complement the systems with extra language models, their performance in modeling rising intonation is not well studied.In this paper, we propose to complement the Cantonese TTS model with a BERT-based statement/question classifier.We design different training strategies and compare their performance.We conduct our experiments on a Cantonese corpus named CanTTS.Empirical results show that the separate training approach obtains the best generalization performance and feasibility. Qibing Bai, Tom Ko, Yu Zhang 0006 |
INTERSPEECH | 3 |
| 2022 | Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech TranslationabstractDirect Speech-to-speech translation (S2ST) has drawn more and more attention recently. The task is very challenging due to data scarcity and complex speech-to-speech mapping. In this paper, we report our recent achievements in S2ST. Firstly, we build a S2ST Transformer baseline which outperforms the original Translatotron. Secondly, we utilize the external data by pseudo-labeling and obtain a new state-of-the-art result on the Fisher English-to-Spanish test set. Indeed, we exploit the pseudo data with a combination of popular techniques which are not trivial when applied to S2ST. Moreover, we evaluate our approach on both syntactically similar (Spanish-English) and distant (English-Chinese) language pairs. Our implementation is available at https://github.com/fengpeng-yue/speech-to-speech-translation. Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, Yu Zhang 0006 |
INTERSPEECH | 6 |
| 2022 | LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERTabstractSelf-supervised speech representation learning has shown promising results in various speech processing tasks.However, the pre-trained models, e.g., HuBERT, are storage-intensive Transformers, limiting their scope of applications under lowresource settings.To this end, we propose LightHuBERT, a once-for-all Transformer compression framework, to find the desired architectures automatically by pruning structured parameters.More precisely, we create a Transformer-based supernet that is nested with thousands of weight-sharing subnets and design a two-stage distillation strategy to leverage the contextualized latent representations from HuBERT.Experiments on automatic speech recognition (ASR) and the SU-PERB benchmark show the proposed LightHuBERT enables over 10 9 architectures concerning the embedding dimension, attention dimension, head number, feed-forward network ratio, and network depth.LightHuBERT outperforms the original HuBERT on ASR and five SUPERB tasks with the Hu-BERT size, achieves comparable performance to the teacher model in most tasks with a reduction of 29% parameters, and obtains a 3.5× compression ratio in three SUPERB tasks, e.g., automatic speaker verification, keyword spotting, and intent classification, with a slight accuracy loss.The code and pre-trained models are available at https://github.com/mechanicalsea/lighthubert. Rui Wang 0073, Qibing Bai, Junyi Ao, Zhixiang Xiong, Zhihua Wei 0001, Yu Zhang 0006, Tom Ko, Haizhou Li 0001 |
INTERSPEECH | 7 |
| 2022 | Dynamic Sparse Network for Time Series Classification: Learning What to "See"abstractThe receptive field (RF), which determines the region of time series to be “seen” and used, is critical to improve the performance for time series classification (TSC). However, the variation of signal scales across and within time series data, makes it challenging to decide on proper RF sizes for TSC. In this paper, we propose a dynamic sparse network (DSN) with sparse connections for TSC, which can learn to cover various RF without cumbersome hyper-parameters tuning. The kernels in each sparse layer are sparse and can be explored under the constraint regions by dynamic sparse training, which makes it possible to reduce the resource cost. The experimental results show that the proposed DSN model can achieve state-of-art performance on both univariate and multivariate TSC datasets with less than 50% computational cost compared with recent baseline methods, opening the path towards more accurate resource-aware methods for time series analyses. Our code is publicly available at: https://github.com/QiaoXiao7282/DSN. Qiao Xiao, Boqian Wu, Yu Zhang 0006, Shiwei Liu 0003, Mykola Pechenizkiy, Elena Mocanu, Decebal Constantin Mocanu |
NeurIPS | 3 |
| 2022 | DHA: Product Title Generation with Discriminative Hierarchical Attention for E-commerce
Wenya Zhu, Yu Zhang 0006, Yu-Hang Zhou, Yinfu Feng, Yuxiang Wu, Qing Da, Anxiang Zeng |
PAKDD (3) | 3 |
| 2022 | Adversarial VAE with Normalizing Flows for Multi-Dimensional Classification
Wenbo Zhang 0010, Yunhao Gou, Yuepeng Jiang, Yu Zhang 0006 |
PRCV (1) | 4 |
| 2022 | A context-enhanced sentence representation learning method for close domains with topic modelingabstractSentence representation approaches have been widely used and proven to be effective in many text modeling tasks and downstream applications. Many recent proposals are available on learning sentence representations based on deep neural frameworks. However, these methods are pre-trained in open domains and depend on the availability of large-scale data for model fitting. As a result, they may fail in some special scenarios, where data are sparse and embedding interpretations are required, such as legal, medical, or technical fields. In this paper, we present an unsupervised learning method to exploit representations of sentences for some closed domains via topic modeling. We reformulate the inference process of the sentences with the corresponding contextual sentences and the associated words, and propose an effective context-enhanced process called the bi-Directional Context-enhanced Sentence Representation Learning (bi-DCSR). This method takes advantage of the semantic distributions of the nearby contextual sentences and the associated words to form a context-enhanced sentence representation. To support the bi-DCSR, we develop a novel Bayesian topic model to embed sentences and words into the same latent interpretable topic space called the Hybrid Priors Topic Model (HPTM). Based on the defined topic space by the HPTM, the bi-DCSR method learns the embedding of a sentence by the two-directional contextual sentences and the words in it, which allows us to efficiently learn high-quality sentence representations in such closed domains. In addition to an open-domain dataset from Wikipedia, our method is validated using three closed-domain datasets from legal cases, electronic medical records, and technical reports. Our experiments indicate that the HPTM significantly outperforms on language modeling and topic coherence, compared with the existing topic models. Meanwhile, the bi-DCSR method does not only outperform the state-of-the-art unsupervised learning methods on closed domain sentence classification tasks, but also yields competitive performance compared to these established approaches on the open domain. Additionally, the visualizations of the semantics of sentences and words demonstrate the interpretable capacity of our model. Shuangyin Li, Yu Zhang 0006, Gansen Zhao, Zhenhua Huang 0001, Yong Tang 0001 |
Inf. Sci. | 3 |
| 2022 | A Survey on Multi-Task LearningabstractMulti-Task Learning (MTL) is a learning paradigm in machine learning and its aim is to leverage useful information contained in multiple related tasks to help improve the generalization performance of all the tasks. In this paper, we give a survey for MTL from the perspective of algorithmic modeling, applications and theoretical analyses. For algorithmic modeling, we give a definition of MTL and then classify different MTL algorithms into five categories, including feature learning approach, low-rank approach, task clustering approach, task relation learning approach and decomposition approach as well as discussing the characteristics of each approach. In order to improve the performance of learning tasks further, MTL can be combined with other learning paradigms including semi-supervised learning, active learning, unsupervised learning, reinforcement learning, multi-view learning and graphical models. When the number of tasks is large or the data dimensionality is high, we review online, parallel and distributed MTL models as well as dimensionality reduction and feature hashing to reveal their computational and storage advantages. Many real-world applications use MTL to boost their performance and we review representative works in this paper. Finally, we present theoretical analyses and discuss several future directions for MTL. Yu Zhang 0006, Qiang Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Distant Transfer Learning via Deep Random WalkabstractTransfer learning, which is to improve the learning performance in the target domain by leveraging useful knowledge from the source domain, often requires that those two domains are very close, which limits its application scope. Recently, distant transfer learning has been studied to transfer knowledge between two distant or even totally unrelated domains via unlabeled auxiliary domains that act as a bridge in the spirit of human transitive inference that two completely unrelated concepts can be connected through gradual knowledge transfer. In this paper, we study distant transfer learning by proposing a DeEp Random Walk basEd distaNt Transfer (DERWENT) method. Different from existing distant transfer learning models that implicitly identify the path of knowledge transfer between the source and target instances through auxiliary instances, the proposed DERWENT model can explicitly learn such paths via the deep random walk technique. Specifically, based on sequences identified by the random walk technique on a data graph where source and target data have no direct connection, the proposed DERWENT model enforces adjacent data points in a sequence to be similar, makes the ending data point be represented by other data points in the same sequence, and considers weighted classification losses of source data. Empirical studies on several benchmark datasets demonstrate that the proposed DERWENT algorithm yields the state-of-the-art performance. Qiao Xiao, Yu Zhang 0006 |
AAAI | 2 |
| 2021 | A Simple Approach to Balance Task Loss in Multi-Task LearningabstractIn multi-task learning, the training losses of different tasks are varying. There are many works to handle this situation and we classify them into five categories. In this paper, we propose a Balanced Multi-Task Learning (BMTL) framework. Different from existing studies which rely on task weighting, the BMTL framework proposes to transform the training loss of each task to balance different tasks based on an intuitive idea that tasks with larger training losses will receive more attention during the optimization procedure. We analyze the transformation function and derive necessary conditions as well as some properties. The proposed BMTL framework is very simple and it can be combined with most multi-task learning models. Empirical studies show the state-of-the-art performance of the proposed BMTL framework. Sicong Liang, Chang Deng, Yu Zhang 0006 |
IEEE BigData | 3 |
| 2021 | Time-Aware Recommender System via Continuous-Time ModelingabstractThe overload of information on the Internet becomes ubiquitous nowadays, which makes the role of recommender systems more important. In recommender systems, the interest of users and popularity of items are not static, but can change drastically. Thus modeling the temporal dynamic of user-item interactions is crucial in recommender systems. The newly proposed Neural Ordinary Differential Equation (NODE) method is able to modeling the temporal mechanism of a system with neural networks. By using the ODE-LSTM method, which unites the ability of NODE to handle continuous time and that of LSTM to address sequential data, in this paper we achieve significant improvements for the recommendation task on several real-world datasets with the time irregularity. To handle sessions with different timestamps in ODE-LSTM, we propose a collective timeline technique that contributes a lot to the performance improvement. Moreover, we find that reducing the scale of time intervals in sessions significantly improves the recommendation performance. Jianghan Bao, Yu Zhang 0006 |
CIKM | 2 |
| 2021 | Region Semantically Aligned Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize unseen classes based on the knowledge of seen classes. Previous methods focused on learning direct embeddings from global features to the semantic space in hope of knowledge transfer from seen classes to unseen classes. However, an unseen class shares local visual features with a set of seen classes and leveraging global visual features makes the knowledge transfer ineffective. To tackle this problem, we propose a Region Semantically Aligned Network (RSAN), which maps local features of unseen classes to their semantic attributes. Instead of using global features which are obtained by an average pooling layer after an image encoder, we directly utilize the output of the image encoder which maintains local information of the image. Concretely, we obtain each attribute from a specific region of the output and exploit these attributes for recognition. As a result, the knowledge of seen classes can be successfully transferred to unseen classes in a region-bases manner. In addition, we regularize the image encoder through attribute regression with a semantic knowledge to extract robust and attribute-related visual features. Experiments on several standard ZSL datasets reveal the benefit of the proposed RSAN method, outperforming state-of-the-art methods. Yunhao Gou, Jingjing Li 0001, Yu Zhang 0006, Yang Yang 0002 |
CIKM | 4 |
| 2021 | SEEN: Few-Shot Classification with SElf-ENsembleabstractFew-shot classification aims at learning new concepts with only a few labeled examples. In this paper, we focus on metric-based methods that have achieved state-of-the-art performance. However, they classify query examples based on embeddings extracted from only the last layer. These embeddings tend to be class-specific and may not generalize well to novel classes or domains. To alleviate this problem, we propose the SElf-ENsemble (SEEN) that leverages embeddings from multiple layers. Specifically, a base classifier is built for each of the last few layers, and the resultant base classifiers are then combined together. Experiments on various benchmark datasets demonstrate that the proposed SEEN method outperforms existing methods in both standard few-shot classification and cross-domain few-shot classification scenarios. Weisen Jiang, Yu Zhang 0006, James T. Kwok |
IJCNN | 2 |
| 2021 | Multi-Task Learning via Generalized Tensor Trace NormabstractThe trace norm is widely used in multi-task learning as it can discover low-rank structures among tasks in terms of model parameters. Nowadays, with the emerging of big complex datasets and the popularity of deep learning techniques, tensor trace norms have been used for deep multi-task models. However, existing tensor trace norms cannot discover all the low-rank structures and they require users to determine the importance of their components manually. To solve those two issues, in this paper, we propose a Generalized Tensor Trace Norm (GTTN). The GTTN is defined as a convex combination of matrix trace norms of all possible tensor flattenings and hence it can discover all the possible low-rank structures. Based on the induced objective function with the GTTN, we can learn combination coefficients in the GTTN with several strategies. Experiments on real-world datasets demonstrate the effectiveness of the proposed GTTN. Yu Zhang 0006, Wei Wang 0028 |
KDD | 2 |
| 2021 | Effective Meta-Regularization by Kernelized Proximal RegularizationabstractWe study the problem of meta-learning, which has proved to be advantageous to accelerate learning new tasks with a few samples. The recent approaches based on deep kernels achieve the state-of-the-art performance. However, the regularizers in their base learners are not learnable. In this paper, we propose an algorithm called MetaProx to learn a proximal regularizer for the base learner. We theoretically establish the convergence of MetaProx. Experimental results confirm the advantage of the proposed algorithm. Weisen Jiang, James T. Kwok, Yu Zhang 0006 |
NeurIPS | 3 |
| 2021 | Multi-Objective Meta LearningabstractMeta learning with multiple objectives has been attracted much attention recently since many applications need to consider multiple factors when designing learning models. Existing gradient-based works on meta learning with multiple objectives mainly combine multiple objectives into a single objective in a weighted sum manner. This simple strategy usually works but it requires to tune the weights associated with all the objectives, which could be time consuming. Different from those works, in this paper, we propose a gradient-based Multi-Objective Meta Learning (MOML) framework without manually tuning weights. Specifically, MOML formulates the objective function of meta learning with multiple objectives as a Multi-Objective Bi-Level optimization Problem (MOBLP) where the upper-level subproblem is to solve several possibly conflicting objectives for the meta learner. To solve the MOBLP, we devise the first gradient-based optimization algorithm by alternatively solving the lower-level and upper-level subproblems via the gradient descent method and the gradient-based multi-objective optimization method, respectively. Theoretically, we prove the convergence properties of the proposed gradient-based optimization algorithm. Empirically, we show the effectiveness of the proposed MOML framework in several meta learning problems, including few-shot learning, domain adaptation, multi-task learning, and neural architecture search. The source code of MOML is available at https://github.com/Baijiong-Lin/MOML. Feiyang Ye 0001, Baijiong Lin, Zhixiong Yue, Pengxin Guo 0001, Qiao Xiao, Yu Zhang 0006 |
NeurIPS | 6 |
| 2021 | Deep Multi-task Augmented Feature Learning via Hierarchical Graph Neural Network
Pengxin Guo 0001, Chang Deng, Linjie Xu, Xiaonan Huang, Yu Zhang 0006 |
ECML/PKDD (1) | 5 |
| 2020 | Learn to Cross-lingual Transfer with Meta Graph Learning Across Heterogeneous LanguagesabstractThe recent emergence of multilingual pretraining language model (mPLM) has enabled breakthroughs on various downstream cross-lingual transfer (CLT) tasks. However, mPLM-based methods usually involve two problems: (1) simply fine-tuning may not adapt general-purpose multilingual representations to be task-aware on low-resource languages; (2) ignore how cross-lingual adaptation happens for downstream tasks. To address the issues, we propose a meta graph learning (MGL) method. Unlike prior works that transfer from scratch, MGL can learn to cross-lingual transfer by extracting meta-knowledge from historical CLT experiences (tasks), making mPLM insensitive to low-resource languages. Besides, for each CLT task, MGL formulates its transfer process as information propagation over a dynamic graph, where the geometric structure can automatically capture intrinsic language relationships to guide cross-lingual transfer explicitly. Empirically, extensive experiments on both public and real-world datasets demonstrate the effectiveness of the MGL method. © 2020 Association for Computational Linguistics Zheng Li 0018, Mukul Kumar, William Headden, Ying Wei 0001, Yu Zhang 0006, Qiang Yang 0001 |
EMNLP (1) | 6 |
| 2020 | WISE: Word-Level Interaction-Based Multimodal Fusion for Speech Emotion Recognition
Guang Shen, Riwei Lai, Rui Chen 0012, Yu Zhang 0006, Kejia Zhang 0001, Qilong Han |
INTERSPEECH | 4 |
| 2020 | Fisher Deep Domain AdaptationabstractDeep domain adaptation models learn a neural network in an unlabeled target domain by leveraging the knowledge from a labeled source domain. This can be achieved by learning a domain-invariant feature space. Though the learned representations are separable in the source domain, they usually have a large variance and samples with different class labels tend to overlap in the target domain, which yields suboptimal adaptation performance. To fill the gap, a Fisher loss is proposed to learn discriminative representations which are within-class compact and between-class separable. Experimental results on two benchmark datasets show that the Fisher loss is a general and effective loss for deep domain adaptation. Noticeable improvements are brought when it is used together with widely adopted transfer criteria, including MMD, CORAL and domain adversarial loss. For example, an absolute improvement of 6.67% in terms of the mean accuracy is attained when the Fisher loss is used together with the domain adversarial loss on the Office-Home dataset. Yu Zhang 0006, Ying Wei 0001, Yangqiu Song, Qiang Yang 0001 |
SDM | 2 |
| 2020 | Adaptive Probabilistic Word EmbeddingabstractWord embeddings have been widely used and proven to be effective in many natural language processing and text modeling tasks. It is obvious that one ambiguous word could have very different semantics in various contexts, which is called polysemy. Most existing works aim at generating only one single embedding for each word while a few works build a limited number of embeddings to present different meanings for each word. However, it is hard to determine the exact number of senses for each word as the word meaning is dependent on contexts. To address this problem, we propose a novel Adaptive Probabilistic Word Embedding (APWE) model, where the word polysemy is defined over a latent interpretable semantic space. Specifically, at first each word is represented by an embedding in the latent semantic space and then based on the proposed APWE model, the word embedding can be adaptively adjusted and updated based on different contexts to obtain the tailored word embedding. Empirical comparisons with state-of-the-art models demonstrate the superiority of the proposed APWE model. Shuangyin Li, Yu Zhang 0006, Kaixiang Mo |
WWW | 2 |
| 2020 | Discriminative distribution alignment: A unified framework for heterogeneous domain adaptation
Yuan Yao 0016, Yu Zhang 0006, Xutao Li 0003, Yunming Ye |
Pattern Recognit. | 2 |
| 2020 | Bi-Directional Recurrent Attentional Topic ModelabstractIn a document, the topic distribution of a sentence depends on both the topics of its neighbored sentences and its own content, and it is usually affected by the topics of the neighbored sentences with different weights. The neighbored sentences of a sentence include the preceding sentences and the subsequent sentences. Meanwhile, it is natural that a document can be treated as a sequence of sentences. Most existing works for Bayesian document modeling do not take these points into consideration. To fill this gap, we propose a bi-Directional Recurrent Attentional Topic Model (bi-RATM) for document embedding. The bi-RATM not only takes advantage of the sequential orders among sentences but also uses the attention mechanism to model the relations among successive sentences. To support to the bi-RATM, we propose a bi-Directional Recurrent Attentional Bayesian Process (bi-RABP) to handle the sequences. Based on the bi-RABP, bi-RATM fully utilizes the bi-directional sequential information of the sentences in a document. Online bi-RATM is proposed to handle large-scale corpus. Experiments on two corpora show that the proposed model outperforms state-of-the-art methods on document modeling and classification. Shuangyin Li, Yu Zhang 0006 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Exploiting Coarse-to-Fine Task Transfer for Aspect-Level Sentiment ClassificationabstractAspect-level sentiment classification (ASC) aims at identifying sentiment polarities towards aspects in a sentence, where the aspect can behave as a general Aspect Category (AC) or a specific Aspect Term (AT). However, due to the especially expensive and labor-intensive labeling, existing public corpora in AT-level are all relatively small. Meanwhile, most of the previous methods rely on complicated structures with given scarce data, which largely limits the efficacy of the neural models. In this paper, we exploit a new direction named coarse-to-fine task transfer, which aims to leverage knowledge learned from a rich-resource source domain of the coarse-grained AC task, which is more easily accessible, to improve the learning in a low-resource target domain of the fine-grained AT task. To resolve both the aspect granularity inconsistency and feature mismatch between domains, we propose a Multi-Granularity Alignment Network (MGAN). In MGAN, a novel Coarse2Fine attention guided by an auxiliary task can help the AC task modeling at the same finegrained level with the AT task. To alleviate the feature false alignment, a contrastive feature alignment method is adopted to align aspect-specific feature representations semantically. In addition, a large-scale multi-domain dataset for the AC task is provided. Empirically, extensive experiments demonstrate the effectiveness of the MGAN. Zheng Li 0018, Ying Wei 0001, Yu Zhang 0006, Xiang Zhang 0001, Xin Li 0056 |
AAAI | 3 |
| 2019 | Transferable End-to-End Aspect-based Sentiment Analysis with Selective Adversarial LearningabstractZheng Li, Xin Li, Ying Wei, Lidong Bing, Yu Zhang, Qiang Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zheng Li 0018, Xin Li 0056, Ying Wei 0001, Lidong Bing, Yu Zhang 0006, Qiang Yang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Heterogeneous Domain Adaptation via Soft Transfer NetworkabstractHeterogeneous domain adaptation (HDA) aims to facilitate the learning task in a target domain by borrowing knowledge from a heterogeneous source domain. In this paper, we propose a Soft Transfer Network (STN), which jointly learns a domain-shared classifier and a domain-invariant subspace in an end-to-end manner, for addressing the HDA problem. The proposed STN not only aligns the discriminative directions of domains but also matches both the marginal and conditional distributions across domains. To circumvent negative transfer, STN aligns the conditional distributions by using the soft-label strategy of unlabeled target data, which prevents the hard assignment of each unlabeled target data to only one category that may be incorrect. Further, STN introduces an adaptive coefficient to gradually increase the importance of the soft-labels since they will become more and more accurate as the number of iterations increases. We perform experiments on the transfer tasks of image-to-image, text-to-image, and text-to-text. Experimental results testify that the STN significantly outperforms several state-of-the-art approaches. Yuan Yao 0016, Yu Zhang 0006, Xutao Li 0003, Yunming Ye |
ACM Multimedia | 2 |
| 2019 | Parameter Transfer Unit for Deep Neural Networks
Yu Zhang 0006, Qiang Yang 0001 |
PAKDD (2) | 2 |
| 2019 | Transfer Meets Hybrid: A Synthetic Approach for Cross-Domain Collaborative Filtering with TextabstractCollaborative Filtering (CF) is the key technique for recommender systems. CF exploits user-item behavior interactions (e.g., clicks) only and hence suffers from the data sparsity issue. One research thread is to integrate auxiliary information such as product reviews and news titles, leading to hybrid filtering methods. Another thread is to transfer knowledge from source domains such as improving the movie recommendation with the knowledge from the book domain, leading to transfer learning methods. In real-world applications, a user registers for multiple services across websites. Thus it motivates us to exploit both auxiliary and source information for recommendation in this paper. To achieve this, we propose a Transfer Meeting Hybrid (TMH) model for cross-domain recommendation with unstructured text. The proposed TMH model attentively extracts useful content from unstructured text via a memory network and selectively transfers knowledge from a source domain via a transfer network. On two real-world datasets, TMH shows better performance in terms of three ranking metrics by comparing with various baselines. We conduct thorough analyses to understand how the text content and transferred knowledge help the proposed model. Guang-Neng Hu, Yu Zhang 0006, Qiang Yang 0001 |
WWW | 2 |
| 2019 | Low-resolution image categorization via heterogeneous domain adaptation
Yuan Yao 0016, Xutao Li 0003, Yunming Ye, Feng Liu 0034, Michael Kwok-Po Ng, Zhichao Huang 0001, Yu Zhang 0006 |
Knowl. Based Syst. | 7 |
| 2018 | Hierarchical Attention Transfer Network for Cross-Domain Sentiment ClassificationabstractCross-domain sentiment classification aims to leverage useful information in a source domain to help do sentiment classification in a target domain that has no or little supervised information. Existing cross-domain sentiment classification methods cannot automatically capture non-pivots, i.e., the domain-specific sentiment words, and pivots, i.e., the domain-shared sentiment words, simultaneously. In order to solve this problem, we propose a Hierarchical Attention Transfer Network (HATN) for cross-domain sentiment classification. The proposed HATN provides a hierarchical attention transfer mechanism which can transfer attentions for emotions across domains by automatically capturing pivots and non-pivots. Besides, the hierarchy of the attention mechanism mirrors the hierarchical structure of documents, which can help locate the pivots and non-pivots better. The proposed HATN consists of two hierarchical attention networks, with one named P-net aiming to find the pivots and the other named NP-net aligning the non-pivots by using the pivots as a bridge. Specifically, P-net firstly conducts individual attention learning to provide positive and negative pivots for NP-net. Then, P-net and NP-net conduct joint attention learning such that the HATN can simultaneously capture pivots and non-pivots and realize transferring attentions for emotions across domains. Experiments on the Amazon review dataset demonstrate the effectiveness of HATN. Zheng Li 0018, Ying Wei 0001, Yu Zhang 0006, Qiang Yang 0001 |
AAAI | 3 |
| 2018 | Transferable Contextual Bandit for Cross-Domain RecommendationabstractTraditional recommendation systems (RecSys) suffer from two problems: the exploitation-exploration dilemma and the cold-start problem. One solution to solving the exploitation-exploration dilemma is the contextual bandit policy, which adaptively exploits and explores user interests. As a result, the contextual bandit policy achieves increased rewards in the long run. The contextual bandit policy, however, may cause the system to explore more than needed in the cold-start situations, which can lead to worse short-term rewards. Cross-domain RecSys methods adopt transfer learning to leverage prior knowledge in a source RecSys domain to jump start the cold-start target RecSys. To solve the two problems together, in this paper, we propose the first applicable transferable contextual bandit (TCB) policy for the cross-domain recommendation. TCB not only benefits the exploitation but also accelerates the exploration in the target RecSys. TCB's exploration, in turn, helps to learn how to transfer between different domains. TCB is a general algorithm for both homogeneous and heterogeneous domains. We perform both theoretical regret analysis and empirical experiments. The empirical results show that TCB outperforms the state-of-the-art algorithms over time. Bo Liu 0015, Ying Wei 0001, Yu Zhang 0006, Zhixian Yan, Qiang Yang 0001 |
AAAI | 3 |
| 2018 | Personalizing a Dialogue System With Transfer Reinforcement LearningabstractIt is difficult to train a personalized task-oriented dialogue system because the data collected from each individual is often insufficient. Personalized dialogue systems trained on a small dataset is likely to overfit and make it difficult to adapt to different user needs. One way to solve this problem is to consider a collection of multiple users as a source domain and an individual user as a target domain, and to perform transfer learning from the source domain to the target domain. By following this idea, we propose a PErsonalized Task-oriented diALogue (PETAL) system, a transfer reinforcement learning framework based on POMDP, to construct a personalized dialogue system. The PETAL system first learns common dialogue knowledge from the source domain and then adapts this knowledge to the target domain. The proposed PETAL system can avoid the negative transfer problem by considering differences between the source and target users in a personalized Q-function. Experimental results on a real-world coffee-shopping data and simulation data show that the proposed PETAL system can learn optimal policies for different users, and thus effectively improve the dialogue quality under the personalized setting. Kaixiang Mo, Yu Zhang 0006, Shuangyin Li, Qiang Yang 0001 |
AAAI | 2 |
| 2018 | CoNet: Collaborative Cross Networks for Cross-Domain RecommendationabstractThe cross-domain recommendation technique is an effective way of alleviating the data sparse issue in recommender systems by leveraging the knowledge from relevant domains. Transfer learning is a class of algorithms underlying these techniques. In this paper, we propose a novel transfer learning approach for cross-domain recommendation by using neural networks as the base model. In contrast to the matrix factorization based cross-domain techniques, our method is deep transfer learning, which can learn complex user-item interaction relationships. We assume that hidden layers in two base networks are connected by cross mappings, leading to the collaborative cross networks (CoNet). CoNet enables dual knowledge transfer across domains by introducing cross connections from one base network to another and vice versa. CoNet is achieved in multi-layer feedforward networks by adding dual connections and joint loss functions, which can be trained efficiently by back-propagation. The proposed model is thoroughly evaluated on two large real-world datasets. It outperforms baselines by relative improvements of 7.84% in NDCG. We demonstrate the necessity of adaptively selecting representations to transfer. Our model can reduce tens of thousands training examples comparing with non-transfer methods and still has the competitive performance with them. Guang-Neng Hu, Yu Zhang 0006, Qiang Yang 0001 |
CIKM | 2 |
| 2018 | Transfer Learning via Learning to TransferabstractIn transfer learning, what and how to transfer are two primary issues to be addressed, as different transfer learning algorithms applied between a source and a target domain result in different knowledge transferred and thereby the performance improvement in the target domain. Determining the optimal one that maximizes the performance improvement requires either exhaustive exploration or considerable expertise. Meanwhile, it is widely accepted in educational psychology that human beings improve transfer learning skills of deciding what to transfer through meta-cognitive reflection on inductive transfer learning practices. Motivated by this, we propose a novel transfer learning framework known as Learning to Transfer (L2T) to automatically determine what and how to transfer are the best by leveraging previous transfer learning experiences. We establish the L2T framework in two stages: 1) we learn a reflection function encrypting transfer learning skills from experiences; and 2) we infer what and how to transfer are the best for a future pair of domains by optimizing the reflection function. We also theoretically analyse the algorithmic stability and generalization bound of L2T, and empirically demonstrate its superiority over several state-of-the-art transfer learning algorithms. Ying Wei 0001, Yu Zhang 0006, Junzhou Huang, Qiang Yang 0001 |
ICML | 2 |
| 2018 | Learning to MultitaskabstractMultitask learning has shown promising performance in many applications and many multitask models have been proposed. In order to identify an effective multitask model for a given multitask problem, we propose a learning framework called Learning to MultiTask (L2MT). To achieve the goal, L2MT exploits historical multitask experience which is organized as a training set consisting of several tuples, each of which contains a multitask problem with multiple tasks, a multitask model, and the relative test error. Based on such training set, L2MT first uses a proposed layerwise graph neural network to learn task embeddings for all the tasks in a multitask problem and then learns an estimation function to estimate the relative test error based on task embeddings and the representation of the multitask model based on a unified formulation. Given a new multitask problem, the estimation function is used to identify a suitable multitask model. Experiments on benchmark datasets show the effectiveness of the proposed L2MT framework. Yu Zhang 0006, Ying Wei 0001, Qiang Yang 0001 |
NeurIPS | 1 |
| 2017 | Recurrent Attentional Topic ModelabstractIn a document, the topic distribution of a sentence depends on both the topics of preceding sentences and its own content, and it is usually affected by the topics of the preceding sentences with different weights. It is natural that a document can be treated as a sequence of sentences. Most existing works for Bayesian document modeling do not take these points into consideration. To fill this gap, we propose a Recurrent Attentional Topic Model (RATM) for document embedding. The RATM not only takes advantage of the sequential orders among sentence but also use the attention mechanism to model the relations among successive sentences. In RATM, we propose a Recurrent Attentional Bayesian Process (RABP) to handle the sequences. Based on the RABP, RATM fully utilizes the sequential information of the sentences in a document. Experiments on two copora show that our model outperforms state-of-the-art methods on document modeling and classification. Shuangyin Li, Yu Zhang 0006, Mingzhi Mao, Yang Yang 0002 |
AAAI | 2 |
| 2017 | Distant Domain Transfer LearningabstractIn this paper, we study a novel transfer learning problem termed Distant Domain Transfer Learning (DDTL). Different from existing transfer learning problems which assume that there is a close relation between the source domain and the target domain, in the DDTL problem, the target domain can be totally different from the source domain. For example, the source domain classifies face images but the target domain distinguishes plane images. Inspired by the cognitive processof human where two seemingly unrelated concepts can be connected by learning intermediate concepts gradually, we propose a Selective Learning Algorithm (SLA) to solve the DDTL problem with supervised autoencoder or supervised convolutional autoencoder as a base model for handling different types of inputs. Intuitively, the SLA algorithm selects usefully unlabeled data gradually from intermediate domains as a bridge to break the large distribution gap for transferring knowledge between two distant domains. Empirical studies on image classification problems demonstrate the effectiveness of the proposed algorithm, and on some tasks the improvement in terms of the classification accuracy is up to 17% over “non-transfer” methods. Ben Tan, Yu Zhang 0006, Sinno Jialin Pan, Qiang Yang 0001 |
AAAI | 2 |
| 2017 | Learning Sparse Task Relations in Multi-Task LearningabstractIn multi-task learning, when the number of tasks is large, pairwise task relations exhibit sparse patterns since usually a task cannot be helpful to all of the other tasks and moreover, sparse task relations can reduce the risk of overfitting compared with the dense ones. In this paper, we focus on learning sparse task relations. Based on a regularization framework which can learn task relations among multiple tasks, we propose a SParse covAriance based mulTi-taSk (SPATS) model to learn a sparse covariance by using the ℓl regularization. The resulting objective function of the SPATS method is convex, which allows us to devise an alternating method to solve it. Moreover, some theoretical properties of the proposed model are studied. Experiments on synthetic and real-world datasets demonstrate the effectiveness of the proposed method. Yu Zhang 0006, Qiang Yang 0001 |
AAAI | 1 |
| 2017 | End-to-End Adversarial Memory Network for Cross-domain Sentiment ClassificationabstractDomain adaptation tasks such as cross-domain sentiment classification have raised much attention in recent years. Due to the domain discrepancy, a sentiment classifier trained in a source domain may not work well when directly applied to a target domain. Traditional methods need to manually select pivots, which behave in the same way for discriminative learning in both domains. Recently, deep learning methods have been proposed to learn a representation shared by domains. However, they lack the interpretability to directly identify the pivots. To address the problem, we introduce an end-to-end Adversarial Memory Network (AMN) for cross-domain sentiment classification. Unlike existing methods, our approach can automatically capture the pivots using an attention mechanism. Our framework consists of two parameter-shared memory networks: one is for sentiment classification and the other is for domain classification. The two networks are jointly trained so that the selected features minimize the sentiment classification error and at the same time make the domain classifier indiscriminative between the representations from the source or target domains. Moreover, unlike deep learning methods that cannot tell us which words are the pivots, our approach can offer a direct visualization of them. Experiments on the Amazon review dataset demonstrate that our approach can significantly outperform state-of-the-art methods. Zheng Li 0018, Yu Zhang 0006, Ying Wei 0001, Yuxiang Wu, Qiang Yang 0001 |
IJCAI | 2 |
| 2017 | Deep Neural Networks for High Dimension, Low Sample Size DataabstractDeep neural networks (DNN) have achieved breakthroughs in applications with large sample size. However, when facing high dimension, low sample size (HDLSS) data, such as the phenotype prediction problem using genetic data in bioinformatics, DNN suffers from overfitting and high-variance gradients. In this paper, we propose a DNN model tailored for the HDLSS data, named Deep Neural Pursuit (DNP). DNP selects a subset of high dimensional features for the alleviation of overfitting and takes the average over multiple dropouts to calculate gradients with low variance. As the first DNN method applied on the HDLSS data, DNP enjoys the advantages of the high nonlinearity, the robustness to high dimensionality, the capability of learning from a small number of samples, the stability in feature selection, and the end-to-end training. We demonstrate these advantages of DNP via empirical results on both synthetic and real-world biological datasets. Bo Liu 0015, Ying Wei 0001, Yu Zhang 0006, Qiang Yang 0001 |
IJCAI | 3 |
| 2017 | Multimodal Linear Discriminant Analysis via Structural SparsityabstractLinear discriminant analysis (LDA) is a widely used supervised dimensionality reduction technique. Even though the LDA method has many real-world applications, it has some limitations such as the single-modal problem that each class follows a normal distribution. To solve this problem, we propose a method called multimodal linear discriminant analysis (MLDA). By generalizing the between-class and within-class scatter matrices, the MLDA model can allow each data point to have its own class mean which is called the instance-specific class mean. Then in each class, data points which share the same or similar instance-specific class means are considered to form one cluster or modal. In order to learn the instance-specific class means, we use the ratio of the proposed generalized between-class scatter measure over the proposed generalized within-class scatter measure, which encourages the class separability, as a criterion. The observation that each class will have a limited number of clusters inspires us to use a structural sparse regularizor to control the number of unique instance-specific class means in each class. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed MLDA method. Yu Zhang 0006, Yuan Jiang 0001 |
IJCAI | 1 |
| 2017 | A Component-Based Diffusion Model With Structural Diversity for Social NetworksabstractDiffusion on social networks refers to the process where opinions are spread via the connected nodes. Given a set of observed information cascades, one can infer the underlying diffusion process for social network analysis. The independent cascade model (IC model) is a widely adopted diffusion model where a node is assumed to be activated independently by any one of its neighbors. In reality, how a node will be activated also depends on how its neighbors are connected and activated. For instance, the opinions from the neighbors of the same social group are often similar and thus redundant. In this paper, we extend the IC model by considering that: 1) the information coming from the connected neighbors are similar and 2) the underlying redundancy can be modeled using a dynamic structural diversity measure of the neighbors. Our proposed model assumes each node to be activated independently by different communities (or components) of its parent nodes, each weighted by its effective size. An expectation maximization algorithm is derived to infer the model parameters. We compare the performance of the proposed model with the basic IC model and its variants using both synthetic data sets and a real-world data set containing news stories and Web blogs. Our empirical results show that incorporating the community structure of neighbors and the structural diversity measure into the diffusion model significantly improves the accuracy of the model, at the expense of only a reasonable increase in run-time. Qing Bao, William Kwok-Wai Cheung, Yu Zhang 0006, Jiming Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | Multi-Stage Multi-Task Learning with Reduced RankabstractMulti-task learning (MTL) seeks to improve the generalization performance by sharing information among multiple tasks. Many existing MTL approaches aim to learn the low-rank structure on the weight matrix, which stores the model parameters of all tasks, to achieve task sharing, and as a consequence the trace norm regularization is widely used in the MTL literature. A major limitation of these approaches based on trace norm regularization is that all the singular values of the weight matrix are penalized simultaneously, leading to impaired estimation on recovering the larger singular values in the weight matrix. To address the issue, we propose a Reduced rAnk MUlti-Stage multi-tAsk learning (RAMUSA) method based on the recently proposed capped norms. Different from existing trace-norm-based MTL approaches which minimize the sum of all the singular values, the RAMUSA method uses a capped trace norm regularizer to minimize only the singular values smaller than some threshold. Due to the non-convexity of the capped trace norm, we develop a simple but well guaranteed multi-stage algorithm to learn the weight matrix iteratively. We theoretically prove that the estimation error at each stage in the proposed algorithm shrinks and finally achieves a lower upper-bound as the number of stages becomes large enough. Empirical studies on synthetic and real-world datasets demonstrate the effectiveness of the RAMUSA method in comparison with the state-of-the-art methods. Lei Han 0001, Yu Zhang 0006 |
AAAI | 2 |
| 2016 | Reduction Techniques for Graph-Based Convex ClusteringabstractThe Graph-based Convex Clustering (GCC) method has gained increasing attention recently. The GCC method adopts a fused regularizer to learn the cluster centers and obtains a geometric clusterpath by varying the regularization parameter. One major limitation is that solving the GCC model is computationally expensive. In this paper, we develop efficient graph reduction techniques for the GCC model to eliminate edges, each of which corresponds to two data points from the same cluster, without solving the optimization problem in the GCC method, leading to improved computational efficiency. Specifically, two reduction techniques are proposed according to tree-based and cyclic-graph-based convex clustering methods separately. The proposed reduction techniques are appealing since they only need to scan the data once with negligibly additional cost and they are independent of solvers for the GCC method, making them capable of improving the efficiency of any existing solver. Experiments on both synthetic and real-world datasets show that our methods can largely improve the efficiency of the GCC model. Lei Han 0001, Yu Zhang 0006 |
AAAI | 2 |
| 2016 | Generalized Hierarchical Sparse Model for Arbitrary-Order Interactive Antigenic Sites Identification in Flu Virus DataabstractRecent statistical evidence has shown that a regression model by incorporating the interactions among the original covariates (features) can significantly improve the interpretability for biological data. One major challenge is the exponentially expanded feature space when adding high-order feature interactions to the model. To tackle the huge dimensionality, Hierarchical Sparse Models (HSM) are developed by enforcing sparsity under heredity structures in the interactions among the covariates. However, existing methods only consider pairwise interactions, making the discovery of important high-order interactions a non-trivial open problem. In this paper, we propose a Generalized Hierarchical Sparse Model (GHSM) as a generalization of the HSM models to learn arbitrary-order interactions. The GHSM applies the l1 penalty to all the model coefficients under a constraint that given any covariate, if none of its associated kth-order interactions contribute to the regression model, then neither do its associated higher-order interactions. The resulting objective function is non-convex with a challenge lying in the coupled variables appearing in the arbitrary-order hierarchical constraints and we devise an efficient optimization algorithm to directly solve it. Specifically, we decouple the variables in the constraints via both the GIST and ADMM methods into three subproblems, each of which is proved to admit an efficiently analytical solution. We evaluate the GHSM method in both synthetic problem and the antigenic sites identification problem for the flu virus data, where we expand the feature space up to the 5th-order interactions. Empirical results demonstrate the effectiveness and efficiency of the proposed method and the learned high-order interactions have meaningful synergistic covariate patterns in the virus antigenicity. Lei Han 0001, Yu Zhang 0006, Xiu-Feng Wan, Tong Zhang 0001 |
KDD | 2 |
| 2016 | Fast Component Pursuit for Large-Scale Inverse Covariance EstimationabstractThe maximum likelihood estimation (MLE) for the Gaussian graphical model, which is also known as the inverse covariance estimation problem, has gained increasing interest recently. Most existing works assume that inverse covariance estimators contain sparse structure and then construct models with the l 1 regularization. In this paper, different from existing works, we study the inverse covariance estimation problem from another perspective by efficiently modeling the low-rank structure in the inverse covariance, which is assumed to be a combination of a low-rank part and a diagonal matrix. One motivation for this assumption is that the low-rank structure is common in many applications including the climate and financial analysis, and another one is that such assumption can reduce the computational complexity when computing its inverse. Specifically, we propose an efficient COmponent Pursuit (COP) method to obtain the low-rank part, where each component can be sparse. For optimization, the COP method greedily learns a rank-one component in each iteration by maximizing the log-likelihood. Moreover, the COP algorithm enjoys several appealing properties including the existence of an efficient solution in each iteration and the theoretical guarantee on the convergence of this greedy approach. Experiments on large-scale synthetic and real-world datasets including thousands of millions variables show that the COP method is faster than the state-of-the-art techniques for the inverse covariance estimation problem when achieving comparable log-likelihood on test data. Lei Han 0001, Yu Zhang 0006, Tong Zhang 0001 |
KDD | 2 |
| 2016 | Correlated Tag Learning in Topic Model
Shuangyin Li, Yu Zhang 0006, Qiang Yang 0001 |
UAI | 3 |
| 2015 | Discriminative Feature GroupingabstractFeature grouping has been demonstrated to be promising in learning with high-dimensional data. It helps reduce the variances in the estimation and improves the stability of feature selection. One major limitation of existing feature grouping approaches is that some similar but different feature groups are often mis-fused, leading to impaired performance. In this paper, we propose a Discriminative Feature Grouping (DFG) method to discover the feature groups with enhanced discrimination. Different from existing methods, DFG adopts a novel regularizer for the feature coefficients to trade-off between fusing and discriminating feature groups. The proposed regularizer consists of a ell_1 norm to enforce feature sparsity and a pairwise ell_infty norm to encourage the absolute differences among any three feature coefficients to be similar. To achieve better asymptotic property, we generalize the proposed regularizer to an adaptive one where the feature coefficients are weighted based on the solution of some estimator with root-n consistency. For optimization, we employ the alternating direction method of multipliers to solve the proposed methods efficiently. Experimental results on synthetic and real-world datasets demonstrate that the proposed methods have good performance compared with the state-of-the-art feature grouping methods. Lei Han 0001, Yu Zhang 0006 |
AAAI | 2 |
| 2015 | Learning Multi-Level Task Groups in Multi-Task LearningabstractIn multi-task learning (MTL), multiple related tasks are learned jointly by sharing information across them. Many MTL algorithms have been proposed to learn the underlying task groups. However, those methods are limited to learn the task groups at only a single level, which may be not sufficient to model the complex structure among tasks in many real-world applications. In this paper, we propose a Multi-Level Task Grouping (MeTaG) method to learn the multi-level grouping structure instead of only one level among tasks. Specifically, by assuming the number of levels to be H, we decompose the parameter matrix into a sum of H component matrices, each of which is regularized with a l2 norm on the pairwise difference among parameters of all the tasks to construct level-specific task groups. For optimization, we employ the smoothing proximal gradient method to efficiently solve the objective function of the MeTaG model. Moreover, we provide theoretical analysis to show that under certain conditions the MeTaG model can recover the true parameter matrix and the true task groups in each level with high probability. We experiment our approach on both synthetic and real-world datasets, showing competitive performance over state-of-the-art MTL methods. Lei Han 0001, Yu Zhang 0006 |
AAAI | 2 |
| 2015 | Parallel Multi-task LearningabstractIn this paper, we develop parallel algorithms for a family of regularized multi-task methods which can model task relations under the regularization framework. Since those multi-task methods cannot be parallelized directly, we use the FISTA algorithm, which in each iteration constructs a surrogate function of the original problem by utilizing the Lipschitz structure of the objective function based on the solution in the last iteration, to solve it. Specifically, we investigate the dual form of the objective function in those methods by adopting the hinge, e-insensitive, and square losses to deal with multi-task classification and regression problems, and then utilize the Lipschitz structure to construct the surrogate function for the dual forms. The surrogate functions constructed in the FISTA algorithm are founded to be decomposable, leading to parallel designs for those multi-task methods. Experiments on several benchmark datasets show that the convergence of the proposed algorithms is as fast as that of SMO-style algorithms and the parallel design can speedup the computation. Yu Zhang 0006 |
ICDM | 1 |
| 2015 | Differentially Private High-Dimensional Data Publication via Sampling-Based InferenceabstractReleasing high-dimensional data enables a wide spectrum of data mining tasks. Yet, individual privacy has been a major obstacle to data sharing. In this paper, we consider the problem of releasing high-dimensional data with differential privacy guarantees. We propose a novel solution to preserve the joint distribution of a high-dimensional dataset. We first develop a robust sampling-based framework to systematically explore the dependencies among all attributes and subsequently build a dependency graph. This framework is coupled with a generic threshold mechanism to significantly improve accuracy. We then identify a set of marginal tables from the dependency graph to approximate the joint distribution based on the solid inference foundation of the junction tree algorithm while minimizing the resultant error. We prove that selecting the optimal marginals with the goal of minimizing error is NP-hard and, thus, design an approximation algorithm using an integer programming relaxation and the constrained concave-convex procedure. Extensive experiments on real datasets demonstrate that our solution substantially outperforms the state-of-the-art competitors. Rui Chen 0012, Qian Xiao 0002, Yu Zhang 0006, Jianliang Xu |
KDD | 3 |
| 2015 | Learning Tree Structure in Multi-Task LearningabstractIn multi-task learning (MTL), multiple related tasks are learned jointly by sharing information according to task relations. One promising approach is to utilize the given tree structure, which describes the hierarchical relations among tasks, to learn model parameters under the regularization framework. However, such a priori information is rarely available in most applications. To the best of our knowledge, there is no work to learn the tree structure among tasks and model parameters simultaneously under the regularization framework and in this paper, we develop a TAsk Tree (TAT) model for MTL to achieve this. By specifying the number of layers in the tree as H, the TAT method decomposes the parameter matrix into H component matrices, each of which corresponds to the model parameters in each layer of the tree. In order to learn the tree structure, we devise sequential constraints to make the distance between the parameters in the component matrices corresponding to each pair of tasks decrease over layers, and hence the component parameters will keep fused until the topmost layer, once they become fused in a layer. Moreover, to make the component parameters have chance to fuse in different layers, we develop a structural sparsity regularizer, which is the sum of the l2 norm on the pairwise difference among the component parameters, to learn layer-specific task structure. In order to solve the resulting non-convex objective function, we use the general iterative shrinkage and thresholding (GIST) method. By using the alternating direction method of multipliers (ADMM) method, we decompose the proximal problem in the GIST method into three independent subproblems, where a key subproblem with the sequential constraints has an efficient solution as the other two subproblems do. We also provide some theoretical analysis for the TAT model. Experiments on both synthetic and real-world datasets show the effectiveness of the TAT model. Lei Han 0001, Yu Zhang 0006 |
KDD | 2 |
| 2015 | Instance-specific canonical correlation analysis
Deming Zhai, Yu Zhang 0006, Dit-Yan Yeung, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001 |
Neurocomputing | 2 |
| 2015 | A Unified Framework for Epidemic Prediction based on Poisson RegressionabstractEpidemic prediction is an important problem in epidemic control. Poisson regression methods are often adopted in existing works, mostly with only the (intra-)regional environmental factors considered. As the diffusion of epidemics is affected by not only the intra-regional factors but also inter-regional and external ones, a unified framework based on Poisson regression with the three types of factors incorporated is proposed for the prediction. Specifically, we propose a Poisson-regression-based model first with the intra-regional and inter-regional factors included. The intra-regional factor in a particular time interval is represented by one feature vector with the regionally environmental and social factors considered. The inter-regional factor is modeled by a diffusion matrix which describes the possibilities that the epidemics can spread from one region to another, which in turn accounts for the propagating effects of the infected cases. To learn the structure of the diffusion matrix, we propose two approaches-utilizing some a priori knowledge (e.g., transportation network) and estimating it from scratch via a sparse structure assumption. The resulting optimization problem of the maximum a posterior solution is a convex one and can be efficiently solved by the alternating direction method of multipliers (ADMM). In addition, we incorporate also the external factor, i.e., the imported cases. With one fact that the distribution of the number of infected cases over a year is (approximately) unimodal for most epidemics and one assumption that the importing rate has a small variance over the year, we can approximate the effect of the external factor with a parametric function (e.g., a quadratic function) over time. The resulting optimization problem is still convex and can be also solved by the ADMM algorithm. Empirical evaluations are conducted based on a real data set which records the 16-days-reported cases in the Yunnan province of China for seven years, from 2005 to 2011. The experimental results demonstrate the effectiveness of our proposed models. Yu Zhang 0006, William Kwok-Wai Cheung, Jiming Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Encoding Tree Sparsity in Multi-Task Learning: A Probabilistic FrameworkabstractMulti-task learning seeks to improve the generalization performance by sharing common information among multiple related tasks. A key assumption in most MTL algorithms is that all tasks are related, which, however, may not hold in many real-world applications. Existing techniques, which attempt to address this issue, aim to identify groups of related tasks using group sparsity. In this paper, we propose a probabilistic tree sparsity (PTS) model to utilize the tree structure to obtain the sparse solution instead of the group structure. Specifically, each model coefficient in the learning model is decomposed into a product of multiple component coefficients each of which corresponds to a node in the tree. Based on the decomposition, Gaussian and Cauchy distributions are placed on the component coefficients as priors to restrict the model complexity. We devise an efficient expectation maximization algorithm to learn the model parameters. Experiments conducted on both synthetic and real-world problems show the effectiveness of our model compared with state-of-the-art baselines. Lei Han 0001, Yu Zhang 0006, Guojie Song, Kunqing Xie |
AAAI | 2 |
| 2013 | Learning High-Order Task Relationships in Multi-Task Learning
Yu Zhang 0006, Dit-Yan Yeung |
IJCAI | 1 |
| 2013 | Incorporating Structural Diversity of Neighbors in a Diffusion Model for Social NetworksabstractDiffusion is known to be an important process governing the behaviours observed in network environments like social networks, contact networks, etc. For modeling the diffusion process, the Independent Cascade Model (IC Model) is commonly adopted and algorithms have been proposed for recovering the hidden diffusion network based on observed cascades. However, the IC Model assumes the effects of multiple neighbors on a node to be independent and does not consider the structural diversity of nodes' neighbourhood. In this paper, we propose an extension of the IC Model with the community structure of node neighbours incorporated. We derive an expectation maximization (EM) algorithm to infer the model parameters. To evaluate the effectiveness and efficiency of the proposed method, we compared it with the IC model and its variants that do not consider the structural properties. Our empirical results based on the MemeTracker dataset, shows that after incorporating the structural diversity, there is a significant improvement in the modelling accuracy, with reasonable increase in run-time. Qing Bao, William Kwok-Wai Cheung, Yu Zhang 0006 |
Web Intelligence | 3 |
| 2013 | Multilabel relationship learningabstractMultilabel learning problems are commonly found in many applications. A characteristic shared by many multilabel learning problems is that some labels have significant correlations between them. In this article, we propose a novel multilabel learning method, called MultiLabel Relationship Learning (MLRL), which extends the conventional support vector machine by explicitly learning and utilizing the relationships between labels. Specifically, we model the label relationships using a label covariance matrix and use it to define a new regularization term for the optimization problem. MLRL learns the model parameters and the label covariance matrix simultaneously based on a unified convex formulation. To solve the convex optimization problem, we use an alternating method in which each subproblem can be solved efficiently. The relationship between MLRL and two widely used maximum margin methods for multilabel learning is investigated. Moreover, we also propose a semisupervised extension of MLRL, called SSMLRL, to demonstrate how to make use of unlabeled data to help learn the label covariance matrix. Through experiments conducted on some multilabel applications, we find that MLRL not only gives higher classification accuracy but also has better interpretability as revealed by the label covariance matrix. Yu Zhang 0006, Dit-Yan Yeung |
ACM Trans. Knowl. Discov. Data | 1 |
| 2013 | A Regularization Approach to Learning Task Relationships in Multitask LearningabstractMultitask learning is a learning paradigm that seeks to improve the generalization performance of a learning task with the help of some other related tasks. In this article, we propose a regularization approach to learning the relationships between tasks in multitask learning. This approach can be viewed as a novel generalization of the regularized formulation for single-task learning. Besides modeling positive task correlation, our approach—multitask relationship learning (MTRL)—can also describe negative task correlation and identify outlier tasks based on the same underlying principle. By utilizing a matrix-variate normal distribution as a prior on the model parameters of all tasks, our MTRL method has a jointly convex objective function. For efficiency, we use an alternating method to learn the optimal model parameters for each task as well as the relationships between tasks. We study MTRL in the symmetric multitask learning setting and then generalize it to the asymmetric setting as well. We also discuss some variants of the regularization approach to demonstrate the use of other matrix-variate priors for learning task relationships. Moreover, to gain more insight into our model, we also study the relationships between MTRL and some existing multitask learning methods. Experiments conducted on a toy problem as well as several benchmark datasets demonstrate the effectiveness of MTRL as well as its high interpretability revealed by the task covariance matrix. Yu Zhang 0006, Dit-Yan Yeung |
ACM Trans. Knowl. Discov. Data | 1 |
| 2012 | Supervised Probabilistic Robust Embedding with Sparse NoiseabstractMany noise models do not faithfully reflect the noise processes introduced during data collection in many real-world applications. In particular, we argue that a type of noise referred to as sparse noise is quite commonly found in many applications and many existing works have been proposed to model such sparse noise. However, all the existing works only focus on unsupervised learning without considering the supervised information, i.e., label information. In this paper, we consider how to model and handle sparse noise in the context of embedding high-dimensional data under a probabilistic formulation for supervised learning. We propose a supervised probabilistic robust embedding (SPRE) model in which data are corrupted either by sparse noise or by a combination of Gaussian and sparse noises. By using the Laplace distribution as a prior to model sparse noise, we devise a two-fold variational EM learning algorithm in which the update of model parameters has analytical solution. We report some classification experiments to compare SPRE with several related models. Yu Zhang 0006, Dit-Yan Yeung, Eric P. Xing |
AAAI | 1 |
| 2012 | Overlapping community detection via bounded nonnegative matrix tri-factorizationabstractComplex networks are ubiquitous in our daily life, with the World Wide Web, social networks, and academic citation networks being some of the common examples. It is well understood that modeling and understanding the network structure is of crucial importance to revealing the network functions. One important problem, known as community detection, is to detect and extract the community structure of networks. More recently, the focus in this research topic has been switched to the detection of overlapping communities. In this paper, based on the matrix factorization approach, we propose a method called bounded nonnegative matrix tri-factorization (BNMTF). Using three factors in the factorization, we can explicitly model and learn the community membership of each node as well as the interaction among communities. Based on a unified formulation for both directed and undirected networks, the optimization problem underlying BNMTF can use either the squared loss or the generalized KL-divergence as its loss function. In addition, to address the sparsity problem as a result of missing edges, we also propose another setting in which the loss function is defined only on the observed edges. We report some experiments on real-world datasets to demonstrate the superiority of BNMTF over other related matrix factorization methods. Yu Zhang 0006, Dit-Yan Yeung |
KDD | 1 |
| 2012 | Multi-Task Boosting by Exploiting Task Relationships
Yu Zhang 0006, Dit-Yan Yeung |
ECML/PKDD (1) | 1 |
| 2012 | Transfer Metric Learning with Semi-Supervised ExtensionabstractDistance metric learning plays a very crucial role in many data mining algorithms because the performance of an algorithm relies heavily on choosing a good metric. However, the labeled data available in many applications is scarce, and hence the metrics learned are often unsatisfactory. In this article, we consider a transfer-learning setting in which some related source tasks with labeled data are available to help the learning of the target task. We first propose a convex formulation for multitask metric learning by modeling the task relationships in the form of a task covariance matrix. Then we regard transfer learning as a special case of multitask learning and adapt the formulation of multitask metric learning to the transfer-learning setting for our method, called transfer metric learning (TML). In TML, we learn the metric and the task covariances between the source tasks and the target task under a unified convex formulation. To solve the convex optimization problem, we use an alternating method in which each subproblem has an efficient solution. Moreover, in many applications, some unlabeled data is also available in the target task, and so we propose a semi-supervised extension of TML called STML to further improve the generalization performance by exploiting the unlabeled data based on the manifold assumption. Experimental results on some commonly used transfer-learning applications demonstrate the effectiveness of our method. Yu Zhang 0006, Dit-Yan Yeung |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Multi-Task Learning in Heterogeneous Feature SpacesabstractMulti-task learning aims at improving the generalization performance of a learning task with the help of some other related tasks. Although many multi-task learning methods have been proposed, they are all based on the assumption that all tasks share the same data representation. This assumption is too restrictive for general applications. In this paper, we propose a multi-task extension of linear discriminant analysis (LDA), called multi-task discriminant analysis (MTDA), which can deal with learning tasks with different data representations. For each task, MTDA learns a separate transformation which consists of two parts, one specific to the task and one common to all tasks. A by-product of MTDA is that it can alleviate the labeled data deficiency problem of LDA. Moreover, unlike many existing multi-task learning methods, MTDA can handle binary and multi-class problems for each task in a generic way. Experimental results on face recognition show that MTDA consistently outperforms related methods. Yu Zhang 0006, Dit-Yan Yeung |
AAAI | 1 |
| 2011 | Discriminative Experimental Design
Yu Zhang 0006, Dit-Yan Yeung |
ECML/PKDD (3) | 1 |
| 2011 | Semisupervised Generalized Discriminant AnalysisabstractGeneralized discriminant analysis (GDA) is a commonly used method for dimensionality reduction. In its general form, it seeks a nonlinear projection that simultaneously maximizes the between-class dissimilarity and minimizes the within-class dissimilarity to increase class separability. In real-world applications where labeled data are scarce, GDA may not work very well. However, unlabeled data are often available in large quantities at very low cost. In this paper, we propose a novel GDA algorithm which is abbreviated as semisupervised generalized discriminant analysis (SSGDA). We utilize unlabeled data to maximize an optimality criterion of GDA and formulate the problem as an optimization problem that is solved using the constrained concave-convex procedure. The optimization procedure leads to estimation of the class labels for the unlabeled data. We propose a novel confidence measure and a method for selecting those unlabeled data points whose labels are estimated with high confidence. The selected unlabeled data can then be used to augment the original labeled dataset for performing GDA. We also propose a variant of SSGDA, called M-SSGDA, which adopts the manifold assumption to utilize the unlabeled data. Extensive experiments on many benchmark datasets demonstrate the effectiveness of our proposed methods. Yu Zhang 0006, Dit-Yan Yeung |
IEEE Trans. Neural Networks | 1 |
| 2010 | Adaptive Transfer LearningabstractTransfer learning aims at reusing the knowledge in some source tasks to improve the learning of a target task. Many transfer learning methods assume that the source tasks and the target task be related, even though many tasks are not related in reality. However, when two tasks are unrelated, the knowledge extracted from a source task may not help, and even hurt, the performance of a target task. Thus, how to avoid negative transfer and then ensure a "safe transfer" of knowledge is crucial in transfer learning. In this paper, we propose an Adaptive Transfer learning algorithm based on Gaussian Processes (AT-GP), which can be used to adapt the transfer learning schemes by automatically estimating the similarity between a source and a target task. The main contribution of our work is that we propose a new semi-parametric transfer kernel for transfer learning from a Bayesian perspective, and propose to learn the model with respect to the target task, rather than all tasks as in multi-task learning. We can formulate the transfer learning problem as a unified Gaussian Process (GP) model. The adaptive transfer ability of our approach is verified on both synthetic and real-world datasets. Bin Cao 0001, Sinno Jialin Pan, Yu Zhang 0006, Dit-Yan Yeung, Qiang Yang 0001 |
AAAI | 3 |
| 2010 | Transductive Learning on Adaptive GraphsabstractGraph-based semi-supervised learning methods are based on some smoothness assumption about the data. As a discrete approximation of the data manifold, the graph plays a crucial role in the success of such graph-based methods. In most existing methods, graph construction makes use of a predefined weighting function without utilizing label information even when it is available. In this work, by incorporating label information, we seek to enhance the performance of graph-based semi-supervised learning by learning the graph and label inference simultaneously. In particular, we consider a particular setting of semi-supervised learning called transductive learning. Using the LogDet divergence to define the objective function, we propose an iterative algorithm to solve the optimization problem which has closed-form solution in each step. We perform experiments on both synthetic and real data to demonstrate improvement in the graph and in terms of classification accuracy. Yan-Ming Zhang 0001, Yu Zhang 0006, Dit-Yan Yeung, Cheng-Lin Liu 0001, Xinwen Hou |
AAAI | 2 |
| 2010 | Multi-task warped Gaussian process for personalized age estimationabstractAutomatic age estimation from facial images has aroused research interests in recent years due to its promising potential for some computer vision applications. Among the methods proposed to date, personalized age estimation methods generally outperform global age estimation methods by learning a separate age estimator for each person in the training data set. However, since typical age databases only contain very limited training data for each person, training a separate age estimator using only training data for that person runs a high risk of overfitting the data and hence the prediction performance is limited. In this paper, we propose a novel approach to age estimation by formulating the problem as a multi-task learning problem. Based on a variant of the Gaussian process (GP) called warped Gaussian process (WGP), we propose a multi-task extension called multi-task warped Gaussian process (MTWGP). Age estimation is formulated as a multi-task regression problem in which each learning task refers to estimation of the age function for each person. While MTWGP models common features shared by different tasks (persons), it also allows task-specific (person-specific) features to be learned automatically. Moreover, unlike previous age estimation methods which need to specify the form of the regression functions or determine many parameters in the functions using inefficient methods such as cross validation, the form of the regression functions in MTWGP is implicitly defined by the kernel function and all its model parameters can be learned from data automatically. We have conducted experiments on two publicly available age databases, FG-NET and MORPH. The experimental results are very promising in showing that MTWGP compares favorably with state-of-the-art age estimation methods. Yu Zhang 0006, Dit-Yan Yeung |
CVPR | 1 |
| 2010 | Transfer metric learning by learning task relationshipsabstractDistance metric learning plays a very crucial role in many data mining algorithms because the performance of an algorithm relies heavily on choosing a good metric. However, the labeled data available in many applications is scarce and hence the metrics learned are often unsatisfactory. In this paper, we consider a transfer learning setting in which some related source tasks with labeled data are available to help the learning of the target task. We first propose a convex formulation for multi-task metric learning by modeling the task relationships in the form of a task covariance matrix. Then we regard transfer learning as a special case of multi-task learning and adapt the formulation of multi-task metric learning to the transfer learning setting for our method, called transfer metric learning (TML). In TML, we learn the metric and the task covariances between the source tasks and the target task under a unified convex formulation. To solve the convex optimization problem, we use an alternating method in which each subproblem has an efficient solution. Experimental results on some commonly used transfer learning applications demonstrate the effectiveness of our method. Yu Zhang 0006, Dit-Yan Yeung |
KDD | 1 |
| 2010 | Worst-Case Linear Discriminant AnalysisabstractDimensionality reduction is often needed in many applications due to the high dimensionality of the data involved. In this paper, we first analyze the scatter measures used in the conventional linear discriminant analysis~(LDA) model and note that the formulation is based on the average-case view. Based on this analysis, we then propose a new dimensionality reduction method called worst-case linear discriminant analysis~(WLDA) by defining new between-class and within-class scatter measures. This new model adopts the worst-case view which arguably is more suitable for applications such as classification. When the number of training data points or the number of features is not very large, we relax the optimization problem involved and formulate it as a metric learning problem. Otherwise, we take a greedy approach by finding one direction of the transformation at a time. Moreover, we also analyze a special case of WLDA to show its relationship with conventional LDA. Experiments conducted on several benchmark datasets demonstrate the effectiveness of WLDA when compared with some related dimensionality reduction methods. Yu Zhang 0006, Dit-Yan Yeung |
NIPS | 1 |
| 2010 | Probabilistic Multi-Task Feature SelectionabstractRecently, some variants of the $l_1$ norm, particularly matrix norms such as the $l_{1,2}$ and $l_{1,\infty}$ norms, have been widely used in multi-task learning, compressed sensing and other related areas to enforce sparsity via joint regularization. In this paper, we unify the $l_{1,2}$ and $l_{1,\infty}$ norms by considering a family of $l_{1,q}$ norms for $1 < q\le\infty$ and study the problem of determining the most appropriate sparsity enforcing norm to use in the context of multi-task feature selection. Using the generalized normal distribution, we provide a probabilistic interpretation of the general multi-task feature selection problem using the $l_{1,q}$ norm. Based on this probabilistic interpretation, we develop a probabilistic model using the noninformative Jeffreys prior. We also extend the model to learn and exploit more general types of pairwise relationships between tasks. For both versions of the model, we devise expectation-maximization~(EM) algorithms to learn all model parameters, including $q$, automatically. Experiments have been conducted on two cancer classification applications using microarray gene expression data. Yu Zhang 0006, Dit-Yan Yeung |
NIPS | 1 |
| 2010 | Multi-Domain Collaborative Filtering
Yu Zhang 0006, Bin Cao 0001, Dit-Yan Yeung |
UAI | 1 |
| 2010 | A Convex Formulation for Learning Task Relationships in Multi-Task Learning
Yu Zhang 0006, Dit-Yan Yeung |
UAI | 1 |
| 2009 | Heteroscedastic Probabilistic Linear Discriminant Analysis with Semi-supervised Extension
Yu Zhang 0006, Dit-Yan Yeung |
ECML/PKDD (2) | 1 |
| 2009 | Semi-Supervised Multi-Task Regression
Yu Zhang 0006, Dit-Yan Yeung |
ECML/PKDD (2) | 1 |
| 2008 | Semi-Supervised Discriminant Analysis using robust path-based similarityabstractLinear discriminant analysis (LDA), which works by maximizing the within-class similarity and minimizing the between-class similarity simultaneously, is a popular dimensionality reduction technique in pattern recognition and machine learning. In real-world applications when labeled data are limited, LDA does not work well. Under many situations, however, it is easy to obtain unlabeled data in large quantities. In this paper, we propose a novel dimensionality reduction method, called semi-supervised discriminant analysis (SSDA), which can utilize both labeled and unlabeled data to perform dimensionality reduction in the semi-supervised setting. Our method uses a robust path-based similarity measure to capture the manifold structure of the data and then uses the obtained similarity to maximize the separability between different classes. A kernel extension of the proposed method for nonlinear dimensionality reduction in the semi-supervised setting is also presented. Experiments on face recognition demonstrate the effectiveness of the proposed method. Yu Zhang 0006, Dit-Yan Yeung |
CVPR | 1 |
| 2008 | Semi-supervised Discriminant Analysis Via CCCP
Yu Zhang 0006, Dit-Yan Yeung |
ECML/PKDD (2) | 1 |
| 2006 | Learning from facial aging patterns for automatic age estimationabstractAge Specific Human-Computer Interaction (ASHCI) has vast potential applications in daily life. However, automatic age estimation technique is still underdeveloped. One of the main reasons is that the aging effects on human faces present several unique characteristics which make age estimation a challenging task that requires non-standard classification approaches. According to the speciality of the facial aging effects, this paper proposes the AGES (AGing pattErn Subspace) method for automatic age estimation. The basic idea is to model the aging pattern, which is defined as a sequence of personal aging face images, by learning a representative subspace. The proper aging pattern for an unseen face image is then determined by the projection in the subspace that can best reconstruct the face image, while the position of the face image in that aging pattern will indicate its age. The AGES method has shown encouraging performance in the comparative experiments either as an age estimator or as an age range estimator. Xin Geng 0001, Zhi-Hua Zhou, Yu Zhang 0006, Gang Li 0009, Honghua Dai 0001 |
ACM Multimedia | 3 |