Changshui Zhang

dblp:z/ChangshuiZhang · also Chang-shui Zhang · DBLP profile ↗
← Back
364ranked-venue papers
7as first author
68since 2021 · last 2026
0000-0002-8088-367XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 261 · 6 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 116 · 18 since 2021Databases, data management, data science and information retrieval · 47 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual Refinement
abstract
Chain-of-Thought prompting has remarkably advanced LLM reasoning by generating explicit step-by-step tokens, yet its discrete nature inherently limits expressiveness and efficiency, struggling with abstract, ambiguous, or semantically divergent cognition beyond linguistic tokens. Latent reasoning offers a promising alternative by operating in the model’s internal continuous space for richer cognitive representations. However, existing methods typically rely on finetuning or token interpolation to bridge latent and input spaces, introducing training difficulty or semantic degradation. To this end, we propose Dynamic Latent Reasoning (DyLaR), a training-free framework that preserves semantic fidelity to latent space. DyLaR introduces a Semantic Residual Refinement module that progressively refines latent inputs by integrating semantic residuals from prior hidden states, thus capturing expressive semantic hierarchies that closely approximate continuous latent representations. To enhance flexibility, DyLaR further incorporates a dynamic switching policy that allows LLMs to alternate between discrete and latent reasoning based on model uncertainty, favoring explicit reasoning when confident and latent exploration under ambiguity. Empirical experiments across knowledge- and reasoning-intensive tasks demonstrate that DyLaR consistently outperforms strong baselines in both effectiveness and token efficiency. Qualitative analyses further illustrate its interpretability and flexibility in navigating complex reasoning scenarios.
Fangrui Lv, Ruixin Hong, Tingting Gao, Guorui Zhou, Changshui Zhang
AAAI8
2026 Kronos: A Foundation Model for the Language of Financial Markets
abstract
The success of large-scale pre-training paradigm, exemplified by Large Language Models (LLMs), has inspired the development of Time Series Foundation Models (TSFMs). However, their application to financial candlestick (K-line) data remains limited, often underperforming non-pre-trained architectures. Moreover, existing TSFMs often overlook crucial downstream tasks such as volatility prediction and synthetic data generation. To address these limitations, we propose Kronos, a unified, scalable pre-training framework tailored to financial K-line modeling. Kronos introduces a specialized tokenizer that discretizes continuous market information into token sequences, preserving both price dynamics and trade activity patterns. We pre-train Kronos using an autoregressive objective on a massive, multi-market corpus of over 12 billion K-line records from 45 global exchanges, enabling it to learn nuanced temporal and cross-asset representations. Kronos excels in a zero-shot setting across a diverse set of financial tasks. On benchmark datasets, Kronos boosts price series forecasting RankIC by 93% over the leading TSFM and 87% over the best non-pre-trained baseline. It also achieves a 9% lower MAE in volatility forecasting and a 22% improvement in generative fidelity for synthetic K-line sequences. These results establish Kronos as a robust, versatile foundation model for end-to-end financial time series analysis.
Zongliang Fu, Shuo Chen 0008, Bohan Zhao, Changshui Zhang
AAAI6
2025 Physics Reasoner: Knowledge-Augmented Reasoning for Solving Physics Problems with Large Language Models
abstract
Physics problems constitute a significant aspect of reasoning, necessitating complicated reasoning ability and abundant physics knowledge. However, existing large language models (LLMs) frequently fail due to a lack of knowledge or incorrect knowledge application. To mitigate these issues, we propose Physics Reasoner, a knowledge-augmented framework to solve physics problems with LLMs. Specifically, the proposed framework constructs a comprehensive formula set to provide explicit physics knowledge and utilizes checklists containing detailed instructions to guide effective knowledge application. Namely, given a physics problem, Physics Reasoner solves it through three stages: problem analysis, formula retrieval, and guided reasoning. During the process, checklists are employed to enhance LLMs’ self-improvement in the analysis and reasoning stages. Empirically, Physics Reasoner mitigates the issues of insufficient knowledge and incorrect application, achieving state-of-the-art performance on SciBench with an average accuracy improvement of 5.8%.
Xinyu Pang, Ruixin Hong, Zhanke Zhou, Fangrui Lv, Zhilong Liang, Bo Han 0003, Changshui Zhang
COLING8
2025 Learning without Isolation: Pathway Protection for Continual Learning
abstract
Deep networks are prone to catastrophic forgetting during sequential task learning, i.e., losing the knowledge about old tasks upon learning new tasks. To this end, continual learning (CL) has emerged, whose existing methods focus mostly on regulating or protecting the parameters associated with the previous tasks. However, parameter protection is often impractical, since the size of parameters for storing the old-task knowledge increases linearly with the number of tasks, otherwise it is hard to preserve the parameters related to the old-task knowledge. In this work, we bring a dual opinion from neuroscience and physics to CL: in the whole networks, the pathways matter more than the parameters when concerning the knowledge acquired from the old tasks. Following this opinion, we propose a novel CL framework, learning without isolation (LwI), where model fusion is formulated as graph matching and the pathways occupied by the old tasks are protected without being isolated. Thanks to the sparsity of activation channels in a deep network, LwI can adaptively allocate available pathways for a new task, realizing pathway protection and addressing catastrophic forgetting in a parameter-effcient manner. Experiments on popular benchmark datasets demonstrate the superiority of the proposed LwI.
Zhikang Chen, Abudukelimu Wuerkaixi, Sen Cui, Haoxuan Li 0001, Jingfeng Zhang, Bo Han 0003, Gang Niu 0001, Houfang Liu, Yi Yang 0039, Sifan Yang, Changshui Zhang
ICML12
2025 Adaptive Localization of Knowledge Negation for Continual LLM Unlearning
abstract
With the growing deployment of large language models (LLMs) across diverse domains, concerns regarding their safety have grown substantially. LLM unlearning has emerged as a pivotal approach to removing harmful or unlawful contents while maintaining utility. Despite increasing interest, the challenges of continual unlearning, which is common in real-world scenarios, remain underexplored. Successive unlearning tasks often lead to intensified utility degradation. To effectively unlearn targeted knowledge while preserving LLM utility, it is essential to minimize changes in model parameters by selectively updating those linked to the target knowledge, thereby ensuring other knowledge remains unaffected. Building on the task vector framework, we propose a new method named ALKN (Adaptive Localization of Knowledge Negation), which uses dynamic masking to sparsify training gradients and adaptively adjusts unlearning intensity based on inter-task relationships. Comprehensive experiments across three well-established LLM unlearning datasets demonstrate that our approach consistently outperforms baseline methods in both unlearning effectiveness and utility retention under continual unlearning settings.
Abudukelimu Wuerkaixi, Sen Cui, Wutong Xu, Bo Han 0003, Gang Niu 0001, Masashi Sugiyama, Changshui Zhang
ICML8
2025 CALM: Consensus-Aware Localized Merging for Multi-Task Learning
abstract
Model merging aims to integrate the strengths of multiple fine-tuned models into a unified model while preserving task-specific capabilities. Existing methods, represented by task arithmetic, are typically classified into global- and local-aware methods. However, global-aware methods inevitably cause parameter interference, while local-aware methods struggle to maintain the effectiveness of task-specific details in the merged model. To address these limitations, we propose a Consensus Aware Localized Merging (CALM) method which incorporates localized information aligned with global task consensus, ensuring its effectiveness post-merging. CALM consists of three key components: (1) class-balanced entropy minimization sampling, providing a more flexible and reliable way to leverage unsupervised data; (2) an efficient-aware framework, selecting a small set of tasks for sequential merging with high scalability; (3) a consensus-aware mask optimization, aligning localized binary masks with global task consensus and merging them conflict-free. Experiments demonstrate the superiority and robustness of our CALM, significantly outperforming existing methods and achieving performance close to traditional MTL.
Kunda Yan, Min Zhang 0068, Sen Cui, Zikun Qu, Bo Jiang 0016, Changshui Zhang
ICML7
2025 Meta-Modulation: A General Learning Framework for Cross-Task Adaptation
abstract
Building learning systems possessing adaptive flexibility to different tasks is critical and challenging. In this article, we propose a novel and general meta-learning framework, called meta-modulation (MeMo), to foster the adaptation capability of a base learner across different tasks where only a few training data are available per task. For one independent task, MeMo proceeds like a "feedback regulation system," which achieves an adaptive modulation on the so-called definitive embeddings of query data to maximize the corresponding task objective. Specifically, we devise a type of efficient feedback information, definitive embedding feedback (DEF), to mathematize and quantify the unsuitability between the few training data and the base learner as well as the promising adjustment direction to reduce this unsuitability. The DEFs are encoded into high-level representation and temporarily stored as task-specific modulator templates by a modulation encoder. For coming query data, we develop an attention mechanism acting upon these modulator templates and combine both task/data-level modulation to generate the final data-specific meta-modulator. This meta-modulator is then used to modulate the query's embedding for correct decision-making. Our framework is scalable for various base learner models like multi-layer perceptron (MLP), long short-term memory (LSTM), convolutional neural network (CNN), and transformer, and applicable to different learning problems like language modeling and image recognition. Experimental results on a 2-D point synthetic dataset and various benchmarks in language and vision domains demonstrate the effectiveness and competitiveness of our framework.
Jiang Lu, Changming Xiao, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2025 DreamArrangement: Learning Language-Conditioned Robotic Rearrangement of Objects via Denoising Diffusion and VLM Planner
abstract
The capability for robotic systems to rearrange objects based on human instructions represents a critical step toward realizing embodied intelligence. Recently, diffusion-based learning has shown significant advancements in the field of data generation while prompt-based learning has proven effective in formulating robot manipulation strategies. However, prior solutions for robotic rearrangement have overlooked the significance of integrating human preferences and optimizing for rearrangement efficiency. Additionally, traditional prompt-based approaches struggle with complex, semantically meaningful rearrangement tasks without predefined target states for objects. To address these challenges, our work first introduces a comprehensive two dimensional (2-D) tabletop rearrangement dataset, utilizing a physical simulator to capture interobject relationships and semantic configurations. Then, we present DreamArrangement, a novel language-conditioned object rearrangement scheme, consisting of two primary processes: employing a transformer-based multimodal denoising diffusion model to envisage the desired arrangement of objects, and leveraging a vision–language foundational model to derive actionable policies from text, alongside initial and target visual information. In particular, we introduce an efficiency-oriented learning strategy to minimize the average motion distance of objects. Given few-shot instruction examples, the learned policy from our synthetic dataset can be transferred to the real world without extra human intervention. Extensive simulations validate DreamArrangement’s superior rearrangement quality and efficiency. Moreover, real-world robotic experiments confirm that our method can adeptly execute a range of challenging, language-conditioned, and long-horizon tasks with a singular model. The demonstration video can be found at https://youtu.be/fq25-DjrbQE
Changming Xiao, Fuchun Sun 0001, Changshui Zhang, Jianwei Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2024 CLOMO: Counterfactual Logical Modification with Large Language Models
abstract
Yinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao, Zhicheng Yang, Dong Yu, Changshui Zhang, Xiaodan Liang, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yinya Huang, Ruixin Hong, Hongming Zhang 0009, Wei Shao 0009, Dong Yu 0001, Changshui Zhang, Xiaodan Liang, Linqi Song
ACL (1)7
2024 Subjective Topic meets LLMs: Unleashing Comprehensive, Reflective and Creative Thinking through the Negation of Negation
abstract
Large language models (LLMs) exhibit powerful reasoning capacity, as evidenced by prior studies focusing on objective topics that with unique standard answer such as arithmetic and commonsense reasoning.However, the reasoning to definite answers emphasizes more on logical thinking, and falls short in effectively reflecting the comprehensive, reflective, and creative thinking that is also critical for the overall reasoning prowess of LLMs.In light of this, we build a dataset SJTP comprising diverse SubJective ToPics with free responses, as well as three evaluation indicators to fully explore LLM's reasoning ability.We observe that a sole emphasis on logical thinking falls short in effectively tackling subjective challenges.Therefore, we introduce a framework grounded in the principle of the Negation of Negation (NeoN) to unleash the potential comprehensive, reflective, and creative thinking abilities of LLMs.Comprehensive experiments on SJTP demonstrate the efficacy of NeoN, and the enhanced performance on various objective reasoning tasks unequivocally underscores the benefits of stimulating LLM's subjective thinking in augmenting overall reasoning capabilities.
Fangrui Lv, Kaixiong Gong, Jian Liang 0002, Xinyu Pang, Changshui Zhang
EMNLP5
2024 Learned ISTA with Error-Based Thresholding for Adaptive Sparse Coding
abstract
Drawing on theoretical insights, we advocate an error-based thresholding (EBT) mechanism for learned ISTA (LISTA), which utilizes a function of the layer-wise reconstruction error to suggest a specific threshold for each observation in the shrinkage function of each layer. We show that the proposed EBT mechanism well disentangles the learnable parameters in the shrinkage functions from the reconstruction errors, endowing the obtained models with improved adaptivity to possible data variations. With rigorous analyses, we further show that the proposed EBT also leads to a faster convergence on the basis of LISTA or its variants, in addition to its higher adaptivity. Extensive experimental results confirm our theoretical analyses and verify the effectiveness of our methods.
Kailun Wu, Yiwen Guo, Changshui Zhang
ICASSP4
2024 CLAP: Collaborative Adaptation for Patchwork Learning
abstract
In this paper, we investigate a new practical learning scenario, where the data distributed in different sources/clients are typically generated with various modalities. Existing research on learning from multi-source data mostly assume that each client owns the data of all modalities, which may largely limit its practicability. In light of the expensiveness and sparsity of multimodal data, we propose patchwork learning to jointly learn from fragmented multimodal data in distributed clients. Considering the concerns on data privacy, patchwork learning aims to impute incomplete multimodal data for diverse downstream tasks without accessing the raw data directly. Local clients could miss different modality combinations. Due to the statistical heterogeneity induced by non-i.i.d. data, the imputation is more challenging since the learned dependencies fail to adapt to the imputation of other clients. In this paper, we provide a novel imputation framework to tackle modality combination heterogeneity and statistical heterogeneity simultaneously, called ``collaborative adaptation''. In particular, for two observed modality combinations from two clients, we learn the transformations between their maximal intersection and other modalities by proposing a novel ELBO. We improve the worst-performing required transformations through a Pareto min-max optimization framework. In extensive experiments, we demonstrate the superiority of the proposed method compared to existing related methods on benchmark data sets and a real-world clinical data set.
Sen Cui, Abudukelimu Wuerkaixi, Weishen Pan, Jian Liang 0002, Changshui Zhang, Fei Wang 0001
ICLR6
2024 Accurate Forgetting for Heterogeneous Federated Continual Learning
abstract
Recent years have witnessed a burgeoning interest in federated learning (FL). However, the contexts in which clients engage in sequential learning remain under- explored. Bridging FL and continual learning (CL) gives rise to a challenging practical problem: federated continual learning (FCL). Existing research in FCL primarily focuses on mitigating the catastrophic forgetting issue of continual learning while collaborating with other clients. We argue that forgetting phenomena are not invariably detrimental. In this paper, we consider a more practical and challenging FCL setting characterized by potentially unrelated or even antagonistic data/tasks across different clients. In the FL scenario, statistical heterogeneity and data noise among clients may exhibit spurious correlations which result in biased feature learning. While existing CL strategies focus on the complete utilization of previous knowledge, we found that forgetting biased information was beneficial in our study. Therefore, we propose a new concept accurate forgetting (AF) and develop a novel generative-replay method AF-FCL that selectively utilizes previous knowledge in federated networks. We employ a probabilistic framework based on a normalizing flow model to quantify the credibility of previous knowledge. Comprehensive experiments affirm the superiority of our method over baselines.
Abudukelimu Wuerkaixi, Sen Cui, Jingfeng Zhang, Kunda Yan, Bo Han 0003, Gang Niu 0001, Changshui Zhang, Masashi Sugiyama
ICLR8
2024 Balancing Similarity and Complementarity for Federated Learning
abstract
In mobile and IoT systems, Federated Learning (FL) is increasingly important for effectively using data while maintaining user privacy. One key challenge in FL is managing statistical heterogeneity, such as non-i.i.d. data, arising from numerous clients and diverse data sources. This requires strategic cooperation, often with clients having similar characteristics. However, we are interested in a fundamental question: does achieving optimal cooperation necessarily entail cooperating with the most similar clients? Typically, significant model performance improvements are often realized not by partnering with the most similar models, but through leveraging complementary data. Our theoretical and empirical analyses suggest that optimal cooperation is achieved by enhancing complementarity in feature distribution while restricting the disparity in the correlation between features and targets. Accordingly, we introduce a novel framework, FedSaC, which balances similarity and complementarity in FL cooperation. Our framework aims to approximate an optimal cooperation network for each client by optimizing a weighted sum of model similarity and feature complementarity. The strength of FedSaC lies in its adaptability to various levels of data heterogeneity and multimodal scenarios. Our comprehensive unimodal and multimodal experiments demonstrate that FedSaC markedly surpasses other state-of-the-art FL methods.
Kunda Yan, Sen Cui, Abudukelimu Wuerkaixi, Jingfeng Zhang, Bo Han 0003, Gang Niu 0001, Masashi Sugiyama, Changshui Zhang
ICML8
2024 A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
abstract
Ruixin Hong, Hongming Zhang, Xinyu Pang, Dong Yu, Changshui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Ruixin Hong, Hongming Zhang 0009, Xinyu Pang, Dong Yu 0001, Changshui Zhang
NAACL-HLT5
2024 Neural Collapse Inspired Feature Alignment for Out-of-Distribution Generalization
abstract
The spurious correlation between the background features of the image and its label arises due to that the samples labeled with the same class in the training set often co-occurs with a specific background, which will cause the encoder to extract non-semantic features for classification, resulting in poor out-of-distribution generalization performance. Although many studies have been proposed to address this challenge, the semantic and spurious features are still difficult to accurately decouple from the original image and fail to achieve high performance with deep learning models. This paper proposes a novel perspective inspired by neural collapse to solve the spurious correlation problem through the alternate execution of environment partitioning and learning semantic masks. Specifically, we propose to assign an environment to each sample by learning a local model for each environment and using maximum likelihood probability. At the same time, we require that the learned semantic mask neurally collapses to the same simplex equiangular tight frame (ETF) in each environment after being applied to the original input. We conduct extensive experiments on four datasets, and the results demonstrate that our method significantly improves out-of-distribution performance.
Zhikang Chen, Min Zhang 0068, Sen Cui, Haoxuan Li 0001, Gang Niu 0001, Mingming Gong, Changshui Zhang, Kun Zhang 0001
NeurIPS7
2024 Source-Free Domain Adaptation via Target Prediction Distribution Searching
abstract
Abstract Existing Source-Free Domain Adaptation (SFDA) methods typically adopt the feature distribution alignment paradigm via mining auxiliary information (eg., pseudo-labelling, source domain data generation). However, they are largely limited due to that the auxiliary information is usually error-prone whilst lacking effective error-mitigation mechanisms. To overcome this fundamental limitation, in this paper we propose a novel Target Prediction Distribution Searching (TPDS) paradigm. Theoretically, we prove that in case of sufficient small distribution shift, the domain transfer error could be well bounded. To satisfy this condition, we introduce a flow of proxy distributions that facilitates the bridging of typically large distribution shift from the source domain to the target domain. This results in a progressive searching on the geodesic path where adjacent proxy distributions are regularized to have small shift so that the overall errors can be minimized. To account for the sequential correlation between proxy distributions, we develop a new pairwise alignment with category consistency algorithm for minimizing the adaptation errors. Specifically, a manifold geometry guided cross-distribution neighbour search is designed to detect the data pairs supporting the Wasserstein distance based shift measurement. Mutual information maximization is then adopted over these pairs for shift regularization. Extensive experiments on five challenging SFDA benchmarks show that our TPDS achieves new state-of-the-art performance. The code and datasets are available at https://github.com/tntek/TPDS .
Song Tang 0001, An Chang, Fabian Zhang, Xiatian Zhu, Mao Ye 0001, Changshui Zhang
Int. J. Comput. Vis.6
2024 From text to mask: Localizing entities using the attention of text-to-image diffusion models
Changming Xiao, Changshui Zhang
Neurocomputing4
2024 A Theoretical View of Linear Backpropagation and its Convergence
abstract
Backpropagation (BP) is widely used for calculating gradients in deep neural networks (DNNs). Applied often along with stochastic gradient descent (SGD) or its variants, BP is considered as a de-facto choice in a variety of machine learning tasks including DNN training and adversarial attack/defense. Recently, a linear variant of BP named LinBP was introduced for generating more transferable adversarial examples for performing black-box attacks, by (Guo et al. 2020). Although it has been shown empirically effective in black-box attacks, theoretical studies and convergence analyses of such a method is lacking. This paper serves as a complement and somewhat an extension to Guo et al. (2020) paper, by providing theoretical analyses on LinBP in neural-network-involved learning tasks, including adversarial attack and model training. We demonstrate that, somewhat surprisingly, LinBP can lead to faster convergence in these tasks in the same hyper-parameter settings, compared to BP. We confirm our theoretical results with extensive experiments.
Yiwen Guo, Haodi Liu, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Weak Augmentation Guided Relational Self-Supervised Learning
abstract
Self-supervised Learning (SSL) including the mainstream contrastive learning has achieved great success in learning visual representations without data annotations. However, most methods mainly focus on the instance level information (i.e., the different augmented images of the same instance should have the same feature or cluster into the same class), but there is a lack of attention on the relationships between different instances. In this paper, we introduce a novel SSL paradigm, which we term as relational self-supervised learning (ReSSL) framework that learns representations by modeling the relationship between different instances. Specifically, our proposed method employs sharpened distribution of pairwise similarities among different instances as relation metric, which is thus utilized to match the feature embeddings of different augmentations. To boost the performance, we argue that weak augmentations matter to represent a more reliable relation, and leverage momentum strategy for practical efficiency. The designed asymmetric predictor head and an InfoNCE warm-up strategy enhance the robustness to hyper-parameters and benefit the resulting performance. Experimental results show that our proposed ReSSL substantially outperforms the state-of-the-art methods across different network architectures, including various lightweight networks (e.g., EfficientNet and MobileNet).
Mingkai Zheng, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0005, Chang Xu 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 A Dexterous Hand-Arm Teleoperation System Based on Hand Pose Estimation and Active Vision
abstract
Markerless vision-based teleoperation that leverages innovations in computer vision offers the advantages of allowing natural and noninvasive finger motions for multifingered robot hands. However, current pose estimation methods still face inaccuracy issues due to the self-occlusion of the fingers. Herein, we develop a novel vision-based hand-arm teleoperation system that captures the human hands from the best viewpoint and at a suitable distance. This teleoperation system consists of an end-to-end hand pose regression network and a controlled active vision system. The end-to-end pose regression network (Transteleop), combined with an auxiliary reconstruction loss function, captures the human hand through a low-cost depth camera and predicts joint commands of the robot based on the image-to-image translation method. To obtain the optimal observation of the human hand, an active vision system is implemented by a robot arm at the local site that ensures the high accuracy of the proposed neural network. Human arm motions are simultaneously mapped to the slave robot arm under relative control. Quantitative network evaluation and a variety of complex manipulation tasks, for example, tower building, pouring, and multitable cup stacking, demonstrate the practicality and stability of the proposed teleoperation system.
Shuang Li 0014, Norman Hendrich, Hongzhuo Liang, Philipp Ruppel, Changshui Zhang, Jianwei Zhang 0001
IEEE Trans. Cybern.5
2024 Progressive Source-Aware Transformer for Generalized Source-Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) tends to forget the source domain, suffering from limitations in real-world scenarios. Recently, generalized source-free domain adaptation (GSFDA) problem naturally emerges, aiming for good performance on both target and source domains. The existing methods attempt to retain model parameters associated with the source domain to prevent such forgetting. However, this strategy is not conducive to improving cross-domain performance on the target domain, prioritizing mitigating forgetting on the source domain. This article introduces a Progressive Source-Aware Transformer approach for GSFDA, dubbed PSAT-GDA. Our core idea is to enforce the domain adaptation process to remember the source domain by imposing source guidance, offering a target domain-centric anti-forgetting mechanism. Specifically, for each epoch, a Transformer-based deep network is adapted to do domain alignment like the traditional SFDA method, because the transformer working on the image patch sequence helps to reduce image noise caused by domain shift. Meanwhile, another Transformer is designed to generate source guidance supervising domain alignment. By augmenting target sample and mining the source information from the historical models before current epoch, source injected feature group is constructed. Based on the Transformer mechanism, the attention block can select useful source information for each target sample. From it, we devise neighbour-based and augmentation-based regularizations to shape the source guidance. Experiments on three challenging datasets show that our method can achieve evident cross-domain improvement on the target domains. Also, it can mitigate forgetting on all domains after adapting to single or multiple target domains.
Song Tang 0001, Yuji Shi, Mao Ye 0001, Changshui Zhang, Jianwei Zhang 0001
IEEE Trans. Multim.5
2024 CATRO: Channel Pruning via Class-Aware Trace Ratio Optimization
abstract
Deep convolutional neural networks are shown to be overkill with high parametric and computational redundancy in many application scenarios, and an increasing number of works have explored model pruning to obtain lightweight and efficient networks. However, most existing pruning approaches are driven by empirical heuristics and rarely consider the joint impact of channels, leading to unguaranteed and suboptimal performance. In this article, we propose a novel channel pruning method via c lass-aware t race r atio o ptimization (CATRO) to reduce the computational burden and accelerate the model inference. Utilizing class information from a few samples, CATRO measures the joint impact of multiple channels by feature space discriminations and consolidates the layerwise impact of preserved channels. By formulating channel pruning as a submodular set function maximization problem, CATRO solves it efficiently via a two-stage greedy iterative optimization procedure. More importantly, we present theoretical justifications on convergence of CATRO and performance of pruned networks. Experimental results demonstrate that CATRO achieves higher accuracy with similar computation cost or lower computation cost with similar accuracy than other state-of-the-art channel pruning algorithms. In addition, because of its class-aware property, CATRO is suitable to prune efficient networks adaptively for various classification subtasks, enhancing handy deployment and usage of deep networks in real-world applications.
Wenzheng Hu, Zhengping Che, Ning Liu 0007, Jian Tang 0008, Changshui Zhang, Jianqiang Wang 0003
IEEE Trans. Neural Networks Learn. Syst.6
2024 On the Equivalence of Linear Discriminant Analysis and Least Squares Regression
abstract
Studying the relationship between linear discriminant analysis (LDA) and least squares regression (LSR) is of great theoretical and practical significance. It is well-known that the two-class LDA is equivalent to an LSR problem, and directly casting multiclass LDA as an LSR problem, however, becomes more challenging. Recent study reveals that the equivalence between multiclass LDA and LSR can be established based on a special class indicator matrix, but under a mild condition which may not hold under the scenarios with low-dimensional or oversampled data. In this article, we show that the equivalence between multiclass LDA and LSR can be established based on arbitrary linearly independent class indicator vectors and without any condition. In addition, we show that LDA is also equivalent to a constrained LSR based on the data-dependent indicator vectors. It can be concluded that under exactly the same mild condition, such two regressions are both equivalent to the null space LDA method. Illuminated by the equivalence of LDA and LSR, we propose a direct LDA classifier to replace the conventional framework of LDA plus extra classifier. Extensive experiments well validate the above theoretic analysis.
Feiping Nie 0001, Hong Chen 0015, Shiming Xiang, Changshui Zhang, Shuicheng Yan, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 IEEE Transactions on Neural Networks and Learning Systems Special Issue on Causal Discovery and Causality-Inspired Machine Learning
abstract
Causality is a fundamental notion in science and engineering. It has attracted much interest across research communities in statistics, machine learning (ML), healthcare, and artificial intelligence (AI), and is becoming increasingly recognized as a vital research area. One of the fundamental problems in causality is how to find the causal structure or the underlying causal model. Accordingly, one focus of this Special Issue is oncausal discovery, i.e., how can we discover causal structure over a set of variables from observational data with automated procedures? Besides learning causality, another focus is on using causality to help understand and advance ML, that is, causality-inspired ML.
Kun Zhang 0001, Ilya Shpitser, Sara Magliacane, Davide Bacciu, Fei Wu 0001, Changshui Zhang, Peter Spirtes
IEEE Trans. Neural Networks Learn. Syst.6
2023 Structured BFGS Method for Optimal Doubly Stochastic Matrix Approximation
abstract
Doubly stochastic matrix plays an essential role in several areas such as statistics and machine learning. In this paper we consider the optimal approximation of a square matrix in the set of doubly stochastic matrices. A structured BFGS method is proposed to solve the dual of the primal problem. The resulting algorithm builds curvature information into the diagonal components of the true Hessian, so that it takes only additional linear cost to obtain the descent direction based on the gradient information without having to explicitly store the inverse Hessian approximation. The cost is substantially fewer than quadratic complexity of the classical BFGS algorithm. Meanwhile, a Newton-based line search method is presented for finding a suitable step size, which in practice uses the existing knowledge and takes only one iteration. The global convergence of our algorithm is established. We verify the advantages of our approach on both synthetic data and real data sets. The experimental results demonstrate that our algorithm outperforms the state-of-the-art solvers and enjoys outstanding scalability.
Dejun Chu, Changshui Zhang, Shiliang Sun, Qing Tao 0001
AAAI2
2023 Faithful Question Answering with Monte-Carlo Planning
abstract
Although large language models demonstrate remarkable question-answering performances, revealing the intermediate reasoning steps that the models faithfully follow remains challenging.In this paper, we propose FAME (FAithful question answering with MontE-carlo planning) to answer questions based on faithful reasoning steps.The reasoning steps are organized as a structured entailment tree, which shows how premises are used to produce intermediate conclusions that can prove the correctness of the answer.We formulate the task as a discrete decision-making problem and solve it through the interaction of a reasoning environment and a controller.The environment is modular and contains several basic task-oriented modules, while the controller proposes actions to assemble the modules.Since the search space could be large, we introduce a Monte-Carlo planning algorithm to do a look-ahead search and select actions that will eventually lead to highquality steps.FAME achieves advanced performance on the standard benchmark.It can produce valid and faithful reasoning steps compared with large language models with a much smaller model size.
Ruixin Hong, Hongming Zhang 0009, Dong Yu 0001, Changshui Zhang
ACL (1)5
2023 Bipartite Ranking Fairness Through a Model Agnostic Ordering Adjustment
abstract
Recently, with the applications of algorithms in various risky scenarios, algorithmic fairness has been a serious concern and received lots of interest in machine learning community. In this article, we focus on the bipartite ranking scenario, where the instances come from either the positive or negative class and the goal is to learn a ranking function that ranks positive instances higher than negative ones. We are interested in whether the learned ranking function can cause systematic disparity across different protected groups defined by sensitive attributes. While there could be a trade-off between fairness and performance, we propose a model agnostic post-processing framework xOrder for achieving fairness in bipartite ranking and maintaining the algorithm classification performance. In particular, we optimize a weighted sum of the utility as identifying an optimal warping path across different protected groups and solve it through a dynamic programming process. xOrder is compatible with various classification models and ranking fairness metrics, including supervised and unsupervised fairness metrics. In addition to binary groups, xOrder can be applied to multiple protected groups. We evaluate our proposed algorithm on four benchmark data sets and two real-world patient electronic health record repositories. xOrder consistently achieves a better balance between the algorithm utility and ranking fairness on a variety of datasets with different metrics. From the visualization of the calibrated ranking scores, xOrder mitigates the score distribution shifts of different groups compared with baselines. Moreover, additional analytical results verify that xOrder achieves a robust performance when faced with fewer samples and a bigger difference between training and testing ranking score distributions.
Sen Cui, Weishen Pan, Changshui Zhang, Fei Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Searching for Network Width With Bilaterally Coupled Network
abstract
Searching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfil the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this article, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we propose to reduce the redundant search space and present the BCNetV2 as the enhanced supernet to ensure rigorous training fairness over channels. Furthermore, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. We also propose a new open-source width search benchmark on macro structures named Channel-Bench-Macro for the better comparisons of the width search algorithms with MobileNet- and ResNet-like architectures. Extensive experiments on the benchmark datasets demonstrate that our method can achieve state-of-the-art performance.
Xiu Su, Shan You, Jiyang Xie 0001, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 A survey on machine learning from few samples
Jiang Lu, Pinghua Gong, Jieping Ye, Jianwei Zhang 0001, Changshui Zhang
Pattern Recognit.5
2023 Where you edit is what you get: Text-guided image editing with region-based attention
abstract
Leveraging the abundant knowledge learned from pre-trained multi-modal models like CLIP has recently proved to be effective for text-guided image editing. Though convincing results have been made when combining the image generator StyleGAN with CLIP, most methods need to train separate models for different prompts, and irrelevant regions are often changed after editing due to the lack of spatial disentanglement. We propose a novel framework that can edit different images according to different prompts in one model. Besides, an innovative region-based spatial attention mechanism is adopted to explicitly guarantee the locality of editing. Experiments mainly in the face domain verify the feasibility of our framework and show that when multi-text editing and local editing are accomplishable, our method can complete practical applications like sequential editing and regional style transfer.
Changming Xiao, Xiaoqiang Xu, Jianwei Zhang 0001, Changshui Zhang
Pattern Recognit.6
2022 GreedyNASv2: Greedier Search with a Greedy Path Filter
abstract
Training a good supernet in one-shot NAS methods is difficult since the search space is usually considerably huge$(\mathrm{e}.\mathrm{g}.,\ 13^{21})$. In order to enhance the supernet's evaluation ability, one greedy strategy is to sample good paths, and let the supernet lean towards the good ones and ease its evaluation burden as a result. However, in practice the search can be still quite inefficient since the identification of good paths is not accurate enough and sampled paths still scatter around the whole search space. In this paper, we leverage an explicit path filter to capture the characteristics of paths and directly filter those weak ones, so that the search can be thus implemented on the shrunk space more greedily and efficiently. Concretely, based on the fact that good paths are much less than the weak ones in the space, we argue that the label of “weak paths” will be more confident and reliable than that of “good paths” in multi-path sampling. In this way, we thus cast the training of path filter in the positive and unlabeled (PU) learning paradigm, and also encourage a path embedding as better path/operation representation to enhance the identification capacity of the learned filter. By dint of this embedding, we can further shrink the search space by aggregating similar operations with similar embeddings, and the search can be more efficient and accurate. Extensive experiments validate the effectiveness of the proposed method GreedyNASv2. For example, our obtained GreedyNASv2-L achieves 81.1% Top-1 accuracy on ImageNet dataset, significantly outperforming the ResNet-50 strong baselines.
Tao Huang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002
CVPR5
2022 ViTAS: Vision Transformer Architecture Search
Xiu Su, Shan You, Jiyang Xie 0001, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002
ECCV (21)7
2022 MetaLogic: Logical Reasoning Explanations with Fine-Grained Structure
abstract
In this paper, we propose a comprehensive benchmark to investigate models' logical reasoning capabilities in complex real-life scenarios.Current explanation datasets often employ synthetic data with simple reasoning structures.Therefore, it cannot express more complex reasoning processes, such as the rebuttal to a reasoning step and the degree of certainty of the evidence.To this end, we propose a comprehensive logical reasoning explanation form.Based on the multi-hop chain of reasoning, the explanation form includes three main components:(1) The condition of rebuttal that the reasoning node can be challenged; (2) Logical formulae that uncover the internal texture of reasoning nodes; (3) Reasoning strength indicated by degrees of certainty.The fine-grained structure conforms to the real logical reasoning scenario, better fitting the human cognitive process but, simultaneously, is more challenging for the current models.We evaluate the current best models' performance on this new explanation form.The experimental results show that generating reasoning graphs remains a challenging task for current models, even with the help of giant pre-trained language models.
Yinya Huang, Hongming Zhang 0009, Ruixin Hong, Xiaodan Liang, Changshui Zhang, Dong Yu 0001
EMNLP5
2022 Leveraging Sparse Coding for EEG Based Emotion Recognition in Shooting
abstract
Emotion recognition in shooting is of great importance for improving athletes’ training methods. However, there is no open and high confident electroencephalography (EEG) dataset about shooting due to the difficulty of data acquisition, which made it a challenge for related studies. In this paper, we collected EEG of novice shooters and high-level shooters in different emotion states, and established two shooting datasets. Furthermore, instead of adopting the common convolutional neural network, we are the first to leverage sparse coding for EEG based emotion recognition in shooting process. Our proposed method can effectively solve the problem of low accuracy caused by data with low signal-noise ratio and small training set. The experimental results demonstrate that our method outperforms other representative deep learning based methods.
Yulu Wang, Changshui Zhang
ICASSP4
2022 Data Agnostic Filter Gating For Efficient Deep Networks
abstract
Filter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of the CNN’s loss function, thereby ignoring the variance of input batches to differences in sparse structure over the filters. In this paper, we propose a data-agnostic filter pruning method that uses an auxiliary network named Dagger module to induce pruning with the pre-trained weights as input. Besides, to help prune filters with a preset FLOPs constraint, we utilize an explicit FLOPs-aware regularisation mechanism to directly promote pruning filters toward the target FLOPs. Experimental results on CIFAR-10 and ImageNet datasets show that the proposed filter pruning method surpasses the state-of-the-art.
Hongyan Xu 0002, Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002, Dadong Wang, Arcot Sowmya
ICASSP7
2022 Collaboration Equilibrium in Federated Learning
abstract
Federated learning (FL) refers to the paradigm of learning models over a collaborative research network involving multiple clients without sacrificing privacy. Recently, there have been rising concerns on the distributional discrepancies across different clients, which could even cause counterproductive consequences when collaborating with others. While it is not necessarily that collaborating with all clients will achieve the best performance, in this paper, we study a rational collaboration called "collaboration equilibrium'' (CE), where smaller collaboration coalitions are formed. Each client collaborates with certain members who maximally improve the model learning and isolates the others who make little contribution. We propose the concept of benefit graph which describes how each client can benefit from collaborating with other clients and advance a Pareto optimization approach to identify the optimal collaborators. Then we theoretically prove that we can reach a CE from the benefit graph through an iterative graph operation. Our framework provides a new way of setting up collaborations in a research network. Experiments on both synthetic and real world data sets are provided to demonstrate the effectiveness of our method.
Sen Cui, Jian Liang 0002, Weishen Pan, Kun Chen 0002, Changshui Zhang, Fei Wang 0001
KDD5
2022 DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization
abstract
Speaker diarization aims to determine “who spoke when” in multi-speaker scenarios. Audio-visual speaker diarization leverages visual information in addition to audio signals and has shown improved performance. Existing audio-visual methods extract speaker embeddings for each video clip using audio and facial features, and then perform clustering according to their similarity. However, this approach would not work well for noisy or overlapped speech where audio features are corrupted, nor for off-screen speakers where visual features are missing. In this work, we propose dynamic vision-guided speaker embedding (DyViSE), a novel method for leveraging visual information to extract speaker embeddings in a multi-stage system. DyViSE uses dynamic lip movement information to denoise audio in a latent space and integrates facial features to obtain an identity-discriminative embedding for each speaking segment. DyViSE is trained with a deep clustering loss along with an exemplary loss. DyViSE demonstrates remarkable performance on both real-world videos and artificially assembled videos. Our code is available at https://github.com/urkax/DyViSE.
Abudukelimu Wuerkaixi, Kunda Yan, You Zhang 0001, Zhiyao Duan, Changshui Zhang
MMSP5
2022 Synergy-of-Experts: Collaborate to Improve Adversarial Robustness
abstract
Learning adversarially robust models require invariant predictions to a small neighborhood of its natural inputs, often encountering insufficient model capacity. There is research showing that learning multiple sub-models in an ensemble could mitigate this insufficiency, further improving the generalization and the robustness. However, the ensemble's voting-based strategy excludes the possibility that the true predictions remain with the minority. Therefore, this paper further improves the ensemble through a collaboration scheme---Synergy-of-Experts (SoE). Compared with the voting-based strategy, the SoE enables the possibility of correct predictions even if there exists a single correct sub-model. In SoE, every sub-model fits its specific vulnerability area and reserves the rest of the sub-models to fit other vulnerability areas, which effectively optimizes the utilization of the model capacity. Empirical experiments verify that SoE outperforms various ensemble methods against white-box and transfer-based adversarial attacks.
Sen Cui, Jingfeng Zhang, Jian Liang 0002, Bo Han 0003, Masashi Sugiyama, Changshui Zhang
NeurIPS6
2022 Recent Advances in Large Margin Learning
abstract
This paper serves as a survey of recent advances in large margin training and its theoretical foundations, mostly for (nonlinear) deep neural networks (DNNs) that are probably the most prominent machine learning models for large-scale data in the community over the past decade. We generalize the formulation of classification margins from classical research to latest DNNs, summarize theoretical connections between the margin, network generalization, and robustness, and introduce recent efforts in enlarging the margins for DNNs comprehensively. Since the viewpoint of different methods is discrepant, we categorize them into groups for ease of comparison and discussion in the paper. Hopefully, our discussions and overview inspire new research work in the community that aim to improve the performance of DNNs, and we also point to directions where the large margin principle can be verified to provide theoretical evidence why certain regularizations for DNNs function well in practice. We managed to shorten the paper such that the crucial spirit of large margin learning and related methods are better emphasized.
Yiwen Guo, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Model-Protected Multi-Task Learning
abstract
Multi-task learning (MTL) refers to the paradigm of learning multiple related tasks together. In contrast, in single-task learning (STL) each individual task is learned independently. MTL often leads to better trained models because they can leverage the commonalities among related tasks. However, because MTL algorithms can "leak" information from different models across different tasks, MTL poses a potential security risk. Specifically, an adversary may participate in the MTL process through one task and thereby acquire the model information for another task. The previously proposed privacy-preserving MTL methods protect data instances rather than models, and some of them may underperform in comparison with STL methods. In this paper, we propose a privacy-preserving MTL framework to prevent information from each model leaking to other models based on a perturbation of the covariance matrix of the model matrix. We study two popular MTL approaches for instantiation, namely, learning the low-rank and group-sparse patterns of the model matrix. Our algorithms can be guaranteed not to underperform compared with STL methods. We build our methods based upon tools for differential privacy, and privacy guarantees, utility bounds are provided, and heterogeneous privacy budgets are considered. The experiments demonstrate that our algorithms outperform the baseline methods constructed by existing privacy-preserving MTL methods on the proposed model-protection problem.
Jian Liang 0002, Xiaoqian Jiang, Changshui Zhang, Fei Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Self-Reinforcing Unsupervised Matching
abstract
Remarkable gains in deep learning usually benefit from large-scale supervised data. Ensuring the intra-class modality diversity in training set is critical for generalization capability of cutting-edge deep models, but it burdens human with heavy manual labor on data collection and annotation. In addition, some rare or unexpected modalities are new for the current model, causing reduced performance under such emerging modalities. Inspired by the achievements in speech recognition, psychology and behavioristics, we present a practical solution, self-reinforcing unsupervised matching (SUM), to annotate the images with 2D structure-preserving property in an emerging modality by cross-modality matching. Specifically, we propose a dynamic programming algorithm, dynamic position warping (DPW), to reveal the underlying element correspondence relationship between two matrix-form data in an order-preserving fashion, and devise a local feature adapter (LoFA) to allow for cross-modality similarity measurement. On these bases, we develop a two-tier self-reinforcing learning mechanism on both feature level and image level to optimize the LoFA. The proposed SUM framework requires no any supervision in emerging modality and only one template in seen modality, providing a promising route towards incremental learning and continual learning. Extensive experimental evaluation on two proposed challenging one-template visual matching tasks demonstrate its efficiency and superiority.
Jiang Lu, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 VD-PCR: Improving visual dialog with pronoun coreference resolution
Xintong Yu 0002, Hongming Zhang 0009, Ruixin Hong, Yangqiu Song, Changshui Zhang
Pattern Recognit.5
2022 Learning Event Extraction From a Few Guideline Examples
abstract
Existing fully supervised event extraction models achieve advanced performance with large-scale labeled data. However, when new event types emerge and annotations are scarce, it is hard for the supervised models to master the new types with limited annotations. In contrast, humans can learn to understand new event types with only a few examples in the event extraction guideline. In this paper, we work on a challenging yet more realistic setting, the few-example event extraction. It requires models to learn event extraction with only a few sentences in guidelines as training data, so that we do not need to collect large-scale annotations each time when new event types emerge. As models tend to overfit when trained with only a few examples, we propose knowledge-guided data augmentation to generate valid and diverse sentences from the guideline examples. To help models better leverage the augmented data, we add a consistency regularization to guarantee consistent representations between the augmented sentences and the original ones. Experiments on the standard benchmark ACE-2005 indicate that our method can extract event triggers and arguments effectively with only a few guideline examples.
Ruixin Hong, Hongming Zhang 0009, Xintong Yu 0002, Changshui Zhang
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 CoDriver ETA: Combine Driver Information in Estimated Time of Arrival by Driving Style Learning Auxiliary Task
abstract
Estimated time of arrival (ETA) is one of the most important services in intelligent transportation systems (ITS). Precise ETA ensures proper travel scheduling of passengers as well as guarantees efficient decision-making on ride-hailing platforms, which are used by an explosively growing number of people in the past few years. Recently, machine learning-based methods have been widely adopted to solve this time estimation problem and become state-of-the-art. However, they do not well explore the personalization information, as many drivers are short of personalized data and do not have sufficient trajectory data in real applications. This data sparsity problem prevents existing methods from obtaining higher prediction accuracy. In this article, we propose a novel deep learning method to solve this problem. We introduce an auxiliary task to learn an embedding of the personalized driving information under multi-task learning framework. In this task, we discriminatively learn the embedding of driving preference that preserves the historical statistics of driving speed. For this purpose, we adapt the triplet network from face recognition to learn the embedding by constructing triplets in the feature space. This simultaneously learned embedding can effectively boost the prediction accuracy of the travel time. We evaluate our method on two large-scale real-world datasets from Didi Chuxing platform. The extensive experimental results on billions of historical vehicle travel data demonstrate that the proposed method outperforms state-of-the-art algorithms.
Kun Fu 0002, Zheng Wang 0010, Donghua Zhou, Kailun Wu, Jieping Ye, Changshui Zhang
IEEE Trans. Intell. Transp. Syst.7
2022 Alleviating Data Sparsity Problems in Estimated Time of Arrival via Auxiliary Metric Learning
abstract
With millions of people using ride-hailing platforms for daily travel, estimated time of arrival (ETA) has become a significant problem in intelligent transportation systems and attracted considerable attention recently. Deep learning-based ETA methods have achieved promising results using massive spatial-temporal data. However, we find that the prediction accuracy is not satisfactory in practical applications due to the prevalent data sparsity problems. Instead of focusing on the average prediction performance as many other methods, this study aims to alleviate the data sparsity problems in ETA to enhance user experience. In general, the data sparsity problems arise from two aspects. The first is the road network, where many links are only traversed by few floating cars. The second aspect is drivers, where many drivers’ trajectories are too scarce (e.g., with only 3 trip records). To alleviate the sparsity in road network, we propose a Road Network Metric Learning framework for ETA (RNML-ETA), where an auxiliary metric learning task is used to improve the link-embedding, especially for links with insufficient data. A novel triangle loss is proposed to improve metric learning effectiveness for links. Experiments on massive real-world data show that RNML-ETA outperforms competing methods by promoting the cold links with limited data. Furthermore, we propose a novel unified framework to Alleviate Data Sparsity problems in ETA (ADS-ETA) by extending RNML-ETA with an additional auxiliary task for driver ID embedding. Results with extensive experiments demonstrate that ADS-ETA can effectively alleviate the data sparsity problems caused by road network and driver sparsity.
Wenzheng Hu, Donghua Zhou, Baichuan Mo, Kun Fu 0002, Zhengping Che, Zheng Wang 0010, Shenhao Wang, Jinhua Zhao 0001, Jieping Ye, Jian Tang 0008, Changshui Zhang
IEEE Trans. Intell. Transp. Syst.12
2021 Learning a Proposal Classifier for Multiple Object Tracking
abstract
The recent trend in multiple object tracking (MOT) is heading towards leveraging deep learning to boost the tracking performance. However, it is not trivial to solve the data-association problem in an end-to-end fashion. In this paper, we propose a novel proposal-based learnable framework, which models MOT as a proposal generation, proposal scoring and trajectory inference paradigm on an affinity graph. This framework is similar to the two-stage object detector Faster RCNN, and can solve the MOT problem in a data-driven way. For proposal generation, we propose an iterative graph clustering method to reduce the computational cost while maintaining the quality of the generated proposals. For proposal scoring, we deploy a trainable graph-convolutional-network (GCN) to learn the structural patterns of the generated proposals and rank them according to the estimated quality scores. For trajectory inference, a simple deoverlapping strategy is adopted to generate tracking output while complying with the constraints that no detection can be assigned to more than one track. We experimentally demonstrate that the proposed method achieves a clear performance improvement in both MOTA and IDF1 with respect to previous state-of-the-art on two public benchmarks. Our code is available at https://github.com/daip13/LPC_MOT.git.
Renliang Weng, Wongun Choi, Changshui Zhang, Zhangping He
CVPR4
2021 Prioritized Architecture Sampling With Monto-Carlo Tree Search
abstract
One-shot neural architecture search (NAS) methods significantly reduce the search cost by considering the whole search space as one network, which only needs to be trained once. However, current methods select each operation independently without considering previous layers. Besides, the historical information obtained with huge computation costs is usually used only once and then discarded. In this paper, we introduce a sampling strategy based on Monte Carlo tree search (MCTS) with the search space modeled as a Monte Carlo tree (MCT), which captures the dependency among layers. Furthermore, intermediate results are stored in the MCT for future decisions and a better exploration-exploitation balance. Concretely, MCT is updated using the training loss as a reward to the architecture performance; for accurately evaluating the numerous nodes, we propose node communication and hierarchical node selection methods in the training and search stages, respectively, making better uses of the operation rewards and hierarchical information. Moreover, for a fair comparison of different NAS methods, we construct an open-source NAS benchmark of a macro search space evaluated on CIFAR-10, namely NAS-Bench-Macro. Extensive experiments on NAS-Bench-Macro and ImageNet demonstrate that our method significantly improves search efficiency and performance. For example, by only searching 20 architectures, our obtained architecture achieves 78.0% top-1 accuracy with 442M FLOPs on ImageNet. Code (Benchmark) is available at: https://github.com/xiusu/NAS-Bench-Macro.
Xiu Su, Tao Huang 0020, Yanxi Li 0001, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
CVPR7
2021 BCNet: Searching for Network Width With Bilaterally Coupled Network
abstract
Searching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfill the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this paper, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. Extensive experiments on benchmark CIFAR-10 and ImageNet datasets indicate that our method can achieve state-of-the-art or competing performance over other baseline methods. Moreover, our method turns out to further boost the performance of NAS models by refining their network widths. For example, with the same FLOPs budget, our obtained EfficientNet-B0 achieves 77.36% Top-1 accuracy on ImageNet dataset, surpassing the performance of original setting by 0.48%.
Xiu Su, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
CVPR5
2021 Exophoric Pronoun Resolution in Dialogues with Topic Regularization
abstract
Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem.Previous works on pronoun coreference resolution (PCR) mostly focus on resolving pronouns to mentions in text while ignoring the exophoric scenario.Exophoric pronouns are common in daily communications, where speakers may directly use pronouns to refer to some objects present in the environment without introducing the objects first.Although such objects are not mentioned in the dialogue text, they can often be disambiguated by the general topics of the dialogue.Motivated by this, we propose to jointly leverage the local context and global topics of dialogues to solve the out-of-text PCR problem.Extensive experiments demonstrate the effectiveness of adding topic regularization for resolving exophoric pronouns.
Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Changshui Zhang, Kun Xu 0005, Dong Yu 0001
EMNLP (1)4
2021 A Trans-regional Online and Offline Fusion Lab Teaching Practice Through Cross-university Cooperation
abstract
Work in Progress: This Innovative Practice Work in Progress Paper presents how a cross-regional online and offline mixed teaching practice has been carried out by coordinating multiple local universities' laboratory resources. Owing to the COVID-19 epidemic, students could not go back to the campus but stay home all over the country. To work with an electronic system design and implementation project in the Electronic Technology Projects course, students in each team need a public physical workplace equipped with the necessary tools and instruments for circuit debugging and implementation. By utilizing local universities' laboratory resources near their homes, students of the same group could have face-to-face discussions and get offline support from local university laboratory teachers. Each team could also communicate online with course teachers on the technical scheme, detailed design, and fault debugging. While online education can share virtual teaching resources, cross-regional online and offline fusion education can further realize the sharing of entity teaching resources. Twenty-three students have fulfilled their projects in eight local universities under online and offline guidance. Such a teaching attempt has also promoted in-depth cooperation between teachers and students across universities.
Yanpin Ren, Changshui Zhang
FIE2
2021 FMA-ETA: Estimating Travel Time Entirely Based on FFN with Attention
abstract
Estimated time of arrival (ETA) is one of the most important services in intelligent transportation systems (ITS) and becomes a challenging spatial-temporal (ST) data mining task in recent years. Nowadays, deep learning based methods, specifically recurrent neural networks (RNN) based ones are adapted to model the ST patterns from massive data for ETA and become the state-of-the-art. However, RNN is suffering from slow training and inference speed, as its structure is unfriendly to parallel computing. To solve this problem, we propose a novel, brief and effective framework mainly based on feed-forward network (FFN) for ETA, FFN with Multifactor Attention (FMA-ETA). The novel Multi-factor Attention mechanism is proposed to deal with different category features and aggregate the information purposefully. Extensive experimental results on the real-world vehicle travel dataset show FMA-ETA is competitive with state-of-the-art methods in terms of the prediction accuracy with significantly better inference speed.
Yulu Wang, Kun Fu 0002, Zheng Wang 0010, Ziang Yan, Changshui Zhang, Jieping Ye
ICASSP6
2021 Learning with Privileged Tasks
abstract
Multi-objective multi-task learning aims to boost the performance of all tasks by leveraging their correlation and conflict appropriately. Nevertheless, in real practice, users may have preference for certain tasks, and other tasks simply serve as privileged or auxiliary tasks to assist the training of target tasks. The privileged tasks thus possess less or even no priority in the final task assessment by users. Motivated by this, we propose a privileged multiple descent algorithm to arbitrate the learning of target tasks and privileged tasks. Concretely, we introduce a privileged parameter so that the optimization direction does not necessarily follow the gradient from the privileged tasks, but concentrates more on the target tasks. Besides, we also encourage a priority parameter for the target tasks to control the potential distraction of optimization direction from the privileged tasks. In this way, the optimization direction can be more aggressively determined by weighting the gradients among target and privileged tasks, and thus highlight more the performance of target tasks under the unified multi-task learning context. Extensive experiments on synthetic and real-world datasets indicate that our method can achieve versatile Pareto solutions under varying preference for the target tasks.
Yuru Song, Zan Lou, Shan You, Erkun Yang, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001
ICCV7
2021 Weakly Supervised Contrastive Learning
abstract
Unsupervised visual representation learning has gained much attention from the computer vision community because of the recent achievement of contrastive learning. Most of the existing contrastive learning frameworks adopt the instance discrimination as the pretext task, which treating every single instance as a different class. However, such method will inevitably cause class collision problems, which hurts the quality of the learned representation. Motivated by this observation, we introduced a weakly supervised contrastive learning framework (WCL) to tackle this issue. Specifically, our proposed framework is based on two projection heads, one of which will perform the regular instance discrimination task. The other head will use a graph-based method to explore similar samples and generate a weak label, then perform a supervised contrastive learning task based on the weak label to pull the similar images closer. We further introduced a K-Nearest Neighbor based multi-crop strategy to expand the number of positive samples. Extensive experimental results demonstrate WCL improves the quality of self-supervised representations across different datasets. Notably, we get a new state-of-the-art result for semi-supervised learning. With only 1% and 10% labeled examples, WCL achieves 65% and 72% ImageNet Top-1 Accuracy using ResNet50, which is even higher than SimCLRv2 with ResNet101.
Mingkai Zheng, Fei Wang 0032, Shan You, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002
ICCV5
2021 Locally Free Weight Sharing for Network Width Search
Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
ICLR6
2021 Policy-Driven Attack: Learning to Query for Hard-label Black-box Adversarial Examples
Ziang Yan, Yiwen Guo, Jian Liang 0002, Changshui Zhang
ICLR4
2021 K-shot NAS: Learnable Weight-Sharing for NAS with K-shot Supernets
abstract
In one-shot weight sharing for NAS, the weights of each operation (at each layer) are supposed to be identical for all architectures (paths) in the supernet. However, this rules out the possibility of adjusting operation weights to cater for different paths, which limits the reliability of the evaluation results. In this paper, instead of counting on a single supernet, we introduce $K$-shot supernets and take their weights for each operation as a dictionary. The operation weight for each path is represented as a convex combination of items in a dictionary with a simplex code. This enables a matrix approximation of the stand-alone weight matrix with a higher rank ($K>1$). A \textit{simplex-net} is introduced to produce architecture-customized code for each path. As a result, all paths can adaptively learn how to share weights in the $K$-shot supernets and acquire corresponding weights for better evaluation. $K$-shot supernets and simplex-net can be iteratively trained, and we further extend the search to the channel dimension. Extensive experiments on benchmark datasets validate that K-shot NAS significantly improves the evaluation accuracy of paths and thus brings in impressive performance improvements.
Xiu Su, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
ICML6
2021 Towards Model-Agnostic Post-Hoc Adjustment for Balancing Ranking Fairness and Algorithm Utility
abstract
Bipartite ranking, which aims to learn a scoring function that ranks positive individuals higher than negative ones from labeled data, is widely adopted in various applications where sample prioritization is needed. Recently, there have been rising concerns on whether the learned scoring function can cause systematic disparity across different protected groups defined by sensitive attributes. While there could be trade-off between fairness and performance, in this paper we propose a model agnostic post-processing framework for balancing them in the bipartite ranking scenario. Specifically, we maximize a weighted sum of the utility and fairness by directly adjusting the relative ordering of samples across groups. By formulating this problem as the identification of an optimal warping path across different protected groups, we propose a non-parametric method to search for such an optimal path through a dynamic programming process. Our method is compatible with various classification models and applicable to a variety of ranking fairness metrics. Comprehensive experiments on a suite of benchmark data sets and two real-world patient electronic health record repositories show that our method can achieve a great balance between the algorithm utility and ranking fairness. Furthermore, we experimentally verify the robustness of our method when faced with the fewer training samples and the difference between training and testing ranking score distributions.
Sen Cui, Weishen Pan, Changshui Zhang, Fei Wang 0001
KDD3
2021 Explaining Algorithmic Fairness Through Fairness-Aware Causal Path Decomposition
abstract
Algorithmic fairness has aroused considerable interests in data mining and machine learning communities recently. So far the existing research has been mostly focusing on the development of quantitative metrics to measure algorithm disparities across different protected groups, and approaches for adjusting the algorithm output to reduce such disparities. In this paper, we propose to study the problem of identification of the source of model disparities. Unlike existing interpretation methods which typically learn feature importance, we consider the causal relationships among feature variables and propose a novel framework to decompose the disparity into the sum of contributions from fairness-aware causal paths, which are paths linking the sensitive attribute and the final predictions, on the graph. We also consider the scenario when the directions on certain edges within those paths cannot be determined. Our framework is also model agnostic and applicable to a variety of quantitative disparity measures. Empirical evaluations on both synthetic and real-world data sets are provided to show that our method can provide precise and comprehensive explanations to the model disparities.
Weishen Pan, Sen Cui, Jiang Bian 0001, Changshui Zhang, Fei Wang 0001
KDD4
2021 Workshop on Model Mining
abstract
How to mine the knowledge in the pretrained models is of significance in achieving more promising performance, since practitioners have access to many pretrained models easily. This Workshop on Model Mining aims to investigate more diverse and advanced manners in mining knowledge within models, which tends to leverage the pretrained models more wisely, elegantly and systematically. There are many topics related to this workshop, such as distilling a lightweight model from a well-trained heavy model via teacher-student paradigm, and boosting the performance of the model by carefully designing the predecessor tasks, e.g., pre-training, self-supervised and contrastive learning. Model mining as a special way of data mining is relevant to SIGKDD, and its audience including researchers and engineers will benefit a lot for designing more advanced algorithms for their tasks.
Shan You, Chang Xu 0002, Fei Wang 0032, Changshui Zhang
KDD4
2021 Addressing Algorithmic Disparity and Performance Inconsistency in Federated Learning
abstract
Federated learning (FL) has gain growing interests for its capability of learning from distributed data sources collectively without the need of accessing the raw data samples across different sources. So far FL research has mostly focused on improving the performance, how the algorithmic disparity will be impacted for the model learned from FL and the impact of algorithmic disparity on the utility inconsistency are largely unexplored. In this paper, we propose an FL framework to jointly consider performance consistency and algorithmic fairness across different local clients (data sources). We derive our framework from a constrained multi-objective optimization perspective, in which we learn a model satisfying fairness constraints on all clients with consistent performance. Specifically, we treat the algorithm prediction loss at each local client as an objective and maximize the worst-performing client with fairness constraints through optimizing a surrogate maximum function with all objectives involved. A gradient-based procedure is employed to achieve the Pareto optimality of this optimization problem. Theoretical analysis is provided to prove that our method can converge to a Pareto solution that achieves the min-max performance with fairness constraints on all clients. Comprehensive experiments on synthetic and real-world datasets demonstrate the superiority that our approach over baselines and its effectiveness in achieving both fairness and consistency across all local clients.
Sen Cui, Weishen Pan, Jian Liang 0002, Changshui Zhang, Fei Wang 0001
NeurIPS4
2021 ReSSL: Relational Self-Supervised Learning with Weak Augmentation
abstract
Self-supervised Learning (SSL) including the mainstream contrastive learning has achieved great success in learning visual representations without data annotations. However, most of methods mainly focus on the instance level information (\ie, the different augmented images of the same instance should have the same feature or cluster into the same class), but there is a lack of attention on the relationships between different instances. In this paper, we introduced a novel SSL paradigm, which we term as relational self-supervised learning (ReSSL) framework that learns representations by modeling the relationship between different instances. Specifically, our proposed method employs sharpened distribution of pairwise similarities among different instances as \textit{relation} metric, which is thus utilized to match the feature embeddings of different augmentations. Moreover, to boost the performance, we argue that weak augmentations matter to represent a more reliable relation, and leverage momentum strategy for practical efficiency. Experimental results show that our proposed ReSSL significantly outperforms the previous state-of-the-art algorithms in terms of both performance and training efficiency.
Mingkai Zheng, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002
NeurIPS5
2021 On Connections Between Regularizations for Improving DNN Robustness
abstract
This paper analyzes regularization terms proposed recently for improving the adversarial robustness of deep neural networks (DNNs), from a theoretical point of view. Specifically, we study possible connections between several effective methods, including input-gradient regularization, Jacobian regularization, curvature regularization, and a cross-Lipschitz functional. We investigate them on DNNs with general rectified linear activations, which constitute one of the most prevalent families of models for image classification and a host of other machine learning applications. We shed light on essential ingredients of these regularizations and re-interpret their functionality. Through the lens of our study, more principled and efficient regularizations can possibly be invented in the near future.
Yiwen Guo, Yurong Chen 0001, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Adversarial Margin Maximization Networks
abstract
The tremendous recent success of deep neural networks (DNNs) has sparked a surge of interest in understanding their predictive ability. Unlike the human visual system which is able to generalize robustly and learn with little supervision, DNNs normally require a massive amount of data to learn new concepts. In addition, research works also show that DNNs are vulnerable to adversarial examples-maliciously generated images which seem perceptually similar to the natural ones but are actually formed to fool learning models, which means the models have problem generalizing to unseen data with certain type of distortions. In this paper, we analyze the generalization ability of DNNs comprehensively and attempt to improve it from a geometric point of view. We propose adversarial margin maximization (AMM), a learning-based regularization which exploits an adversarial perturbation as a proxy. It encourages a large margin in the input space, just like the support vector machines. With a differentiable formulation of the perturbation, we train the regularized DNNs simply through back-propagation in an end-to-end manner. Experimental results on various datasets (including MNIST, CIFAR-10/100, SVHN and ImageNet) and different DNN architectures demonstrate the superiority of our method over previous state-of-the-arts. Code and models for reproducing our results will be made publicly available.
Ziang Yan, Yiwen Guo, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 DCR: Disentangled component representation for sketch generation
Zhong Cao 0001, Sen Cui, Changshui Zhang
Pattern Recognit. Lett.3
2021 DiFNet: Densely High-Frequency Convolutional Neural Networks
abstract
Deep convolutional neural networks have achieved great success in many computer vision tasks. However, they can be attacked by adversarial examples which are input-data with small intentional feature perturbations to fool machine learning models. This vulnerability to adversarial examples poses a potential threat to their widespread application, especially in security-sensitive scenarios. In this paper, we revisit adversarial examples in the frequency domain referring to the computational theory of edge detection, and propose a noveldensely high-frequency convolution neural network(DiFNet) to effectively defend against adversarial attacks. DiFNet introduces classical edge detection operations into the network structure to enhance the detection ability for high-frequency components in the image. It works well to defend against the imperceptible perturbations attacks even without adversarial examples. Experiments demonstrate that DiFNet outperforms handcraft-designed CNNs in terms of prediction accuracy with improved robustness against state-of-the-art adversarial attacks (FGSM, PGD, etc.).
Wenzheng Hu, Zheng Wang 0010, Jianqiang Wang 0003, Changshui Zhang
IEEE Signal Process. Lett.5
2021 Deep Likelihood Network for Image Restoration With Multiple Degradation Levels
abstract
Convolutional neural networks have been proven effective in a variety of image restoration tasks. Most state-of-the-art solutions, however, are trained using images with a single particular degradation level, and their performance deteriorates drastically when applied to other degradation settings. In this paper, we propose deep likelihood network (DL-Net), aiming at generalizing off-the-shelf image restoration networks to succeed over a spectrum of degradation levels. We slightly modify an off-the-shelf network by appending a simple recursive module, which is derived from a fidelity term, for disentangling the computation for multiple degradation levels. Extensive experimental results on image inpainting, interpolation, and super-resolution show the effectiveness of our DL-Net.
Yiwen Guo, Ming Lu 0002, Wangmeng Zuo, Changshui Zhang, Yurong Chen 0001
IEEE Trans. Image Process.4
2021 Robust Few-Shot Learning for User-Provided Data
abstract
Few-shot learning (FSL) focuses on distilling transferrable knowledge from existing experience to cope with novel concepts for which the labeled data are scarce. A typical assumption in FSL is that the training examples of novel classes are all clean with no outlier interference. In many realistic applications where examples are provided by users, however, data are potentially noisy or unreadable. In this context, we introduce a novel research topic, robust FSL (RFSL), where we aim to address two types of outliers within user-provided data: the representation outlier (RO) and the label outlier (LO). Moreover, we introduce a metric for estimating robustness and use it to investigate the performance of several advanced methods to FSL when faced with user-provided outliers. In addition, we propose robust attentive profile networks (RapNets) to achieve outlier suppression. The results of a comprehensive evaluation of benchmark data sets demonstrate the shortcomings of current FSL methods and the superiority of the proposed RapNets when dealing with RFSL problems, establishing a benchmark for follow-up studies.
Jiang Lu, Sheng Jin 0007, Jian Liang 0002, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2020 RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement Learning
abstract
This paper presents a deep reinforcement learning algorithm for online accompaniment generation, with potential for real-time interactive human-machine duet improvisation. Different from offline music generation and harmonization, online music accompaniment requires the algorithm to respond to human input and generate the machine counterpart in a sequential order. We cast this as a reinforcement learning problem, where the generation agent learns a policy to generate a musical note (action) based on previously generated context (state). The key of this algorithm is the well-functioning reward model. Instead of defining it using music composition rules, we learn this model from monophonic and polyphonic training data. This model considers the compatibility of the machine-generated note with both the machine-generated context and the human-generated context. Experiments show that this algorithm is able to respond to the human part and generate a melodic, harmonic and diverse machine part. Subjective evaluations on preferences show that the proposed algorithm generates music pieces of higher quality than the baseline method.
Nan Jiang 0023, Sheng Jin 0007, Zhiyao Duan, Changshui Zhang
AAAI4
2020 Reborn Filters: Pruning Convolutional Neural Networks with Limited Data
abstract
Channel pruning is effective in compressing the pretrained CNNs for their deployment on low-end edge devices. Most existing methods independently prune some of the original channels and need the complete original dataset to fix the performance drop after pruning. However, due to commercial protection or data privacy, users may only have access to a tiny portion of training examples, which could be insufficient for the performance recovery. In this paper, for pruning with limited data, we propose to use all original filters to directly develop new compact filters, named reborn filters, so that all useful structure priors in the original filters can be well preserved into the pruned networks, alleviating the performance drop accordingly. During training, reborn filters can be easily implemented via 1×1 convolutional layers and then be fused in the inference stage for acceleration. Based on reborn filters, the proposed channel pruning algorithm shows its effectiveness and superiority on extensive experiments.
Yehui Tang 0001, Shan You, Chang Xu 0002, Jin Han 0001, Chen Qian 0006, Boxin Shi, Chao Xu 0006, Changshui Zhang
AAAI8
2020 Few Sample Knowledge Distillation for Efficient Network Compression
abstract
Deep neural network compression techniques such as pruning and weight tensor decomposition usually require fine-tuning to recover the prediction accuracy when the compression ratio is high. However, conventional fine-tuning suffers from the requirement of a large training set and the time-consuming training procedure. This paper proposes a novel solution for knowledge distillation from label-free few samples to realize both data efficiency and training/processing efficiency. We treat the original network as "teacher-net" and the compressed network as "student-net". A 1x1 convolution layer is added at the end of each layer block of the student-net, and we fit the block-level outputs of the student-net to the teacher-net by estimating the parameters of the added layers. We prove that the added layer can be merged without adding extra parameters and computation cost during inference. Experiments on multiple datasets and network architectures verify the method's effectiveness on student-nets obtained by various network pruning and weight decomposition methods. Our method can recover student-net's accuracy to the same level as conventional fine-tuning methods in minutes while using only 1% label-free data of the full training data.
Tianhong Li, Zhuang Liu 0003, Changshui Zhang
CVPR4
2020 Boosting Semantic Human Matting With Coarse Annotations
abstract
Semantic human matting aims to estimate the per-pixel opacity of the foreground human regions. It is quite challenging that usually requires user interactive trimaps and plenty of high quality annotated data. Annotating such kind of data is labor intensive and requires great skills beyond normal users, especially considering the very detailed hair part of humans. In contrast, coarse annotated human dataset is much easier to acquire and collect from the public dataset. In this paper, we propose to leverage coarse annotated data coupled with fine annotated data to boost end-to-end semantic human matting without trimaps as extra input. Specifically, We train a mask prediction network to estimate the coarse semantic mask using the hybrid data, and then propose a quality unification network to unify the quality of the previous coarse mask outputs. A matting refinement network takes the unified mask and the input image to predict the final alpha matte. The collected coarse annotated dataset enriches our dataset significantly, allows generating high quality alpha matte for real images. Experimental results show that the proposed method performs comparably against state-of-the-art methods. Moreover, the proposed method can be used for refining coarse annotated public dataset, as well as semantic segmentation methods, which reduces the cost of annotating high quality human data to a great extent.
Jinlin Liu, Yuan Yao 0013, Wendi Hou, Miaomiao Cui, Xuansong Xie, Changshui Zhang, Xian-Sheng Hua 0001
CVPR6
2020 GreedyNAS: Towards Fast One-Shot NAS With Greedy Supernet
abstract
Training a supernet matters for one-shot neural architecture search (NAS) methods since it serves as a basic performance estimator for different architectures (paths). Current methods mainly hold the assumption that a supernet should give a reasonable ranking over all paths. They thus treat all paths equally, and spare much effort to train paths. However, it is harsh for a single supernet to evaluate accurately on such a huge-scale search space (e.g., 7^21). In this paper, instead of covering all paths, we ease the burden of supernet by encouraging it to focus more on evaluation of those potentially-good ones, which are identified using a surrogate portion of validation data. Concretely, during training, we propose a multi-path sampling strategy with rejection, and greedily filter the weak paths. The training efficiency is thus boosted since the training space has been greedily shrunk from all paths to those potentially-good ones. Moreover, we further adopt an exploration and exploitation policy by introducing an empirical candidate path pool. Our proposed method GreedyNAS is easy-to-follow, and experimental results on ImageNet dataset indicate that it can achieve better Top-1 accuracy under same search space and FLOPs or latency level, but with only ~60% of supernet training cost. By searching on a larger space, our GreedyNAS can also obtain new state-of-the-art architectures.
Shan You, Tao Huang 0020, Mingmin Yang, Fei Wang 0032, Chen Qian 0006, Changshui Zhang
CVPR6
2020 An Overview of Signal Processing Engineering Education in China
abstract
This Innovative Practice Full Paper presents a survey on how signal processing engineering education has been carried out in China. Although courses on signal processing vary from university to university in China, they mainly fall into four categories: courses on principle, method and technology, application, and implementation. Along with the theoretical teaching of these courses, increasing emphasis has been put on the practical part. Engineering education on signal processing has been put into practice over decades through the design and implementation of multi-level in-curriculum experiments, project-driven training, cross-disciplinary contest, and a whole-semester graduation project. Manifold tools and platforms have been adopted or established to enhance the diverse laboratory teaching of signal processing. These tools and platforms include hardware circuits, software tools, and hardware platforms based on embedded processors.
Yanpin Ren, Qing Zhuo, Yi Yang 0039, Changshui Zhang
FIE4
2020 Adversarial Attacks on Deep Unfolded Networks for Sparse Coding
abstract
Previous works have discovered that general DNNs are vulnerable under some subtle and specific perturbations on classification tasks. In recent years, deep neural networks (DNNs) unfolded through sparse coding algorithms have achieved great success in sparse coding problem. Some applications of sparse coding have important strategic significance in many cases, and the security of learned models is vital. However, it has not achieved enough attentions. Our paper is the first work to study the adversarial performance on unfolded DNNs for sparse coding. We first verify the effectiveness of the existing attack or defense strategies, and surprisingly discover the defense strategies are useless. In addition, we propose a special attack strategy to eliminate the element on certain dimension of the sparse output code of DNNs. Furthermore, we propose a succinct black-box attack strategy, which could generate adversarial perturbations without knowing the parameters of DNNs and data.
Yulu Wang, Kailun Wu, Changshui Zhang
ICASSP3
2020 Sparse Coding with Gated Learned ISTA
Kailun Wu, Yiwen Guo, Changshui Zhang
ICLR4
2020 Semismooth Newton Algorithm for Efficient Projections onto ℓ1, ∞-norm Ball
Dejun Chu, Changshui Zhang, Shiliang Sun, Qing Tao 0001
ICML2
2020 Road Network Metric Learning for Estimated Time of Arrival
abstract
Recently, deep learning have achieved promising results in Estimated Time of Arrival (ETA), which is considered as predicting the travel time from the origin to the destination along a given path. One of the key techniques is to use embedding vectors to represent the elements of road network, such as the links (road segments). However, the embedding suffers from the data sparsity problem that many links in the road network are traversed by too few floating cars even in large ride-hailing platforms like Uber and DiDi. Insufficient data makes the embedding vectors in an under-fitting status, which undermines the accuracy of ETA prediction. To address the data sparsity problem, we propose the Road Network Metric Learning framework for ETA (RNML-ETA). It consists of two components: (1) a main regression task to predict the travel time, and (2) an auxiliary metric learning task to improve the quality of link embedding vectors. We further propose the triangle loss, a novel loss function to improve the efficiency of metric learning. We validated the effectiveness of RNML-ETA on large scale realworld datasets, by showing that our method outperforms the state-of-the-art model and the promotion concentrates on the cold links with few data.
Kun Fu 0002, Zheng Wang 0010, Changshui Zhang, Jieping Ye
ICPR4
2020 Constructing Geographic and Long-term Temporal Graph for Traffic Forecasting
abstract
Traffic forecasting influences various intelligent transportation system (ITS) services and is of great significance for user experience as well as urban traffic control. It is challenging due to the fact that the road network contains complex and time-varying spatial-temporal dependencies. Recently, deep learning based methods have achieved promising results by adopting graph convolutional network (GCN) to extract the spatial correlations and recurrent neural network (RNN) to capture the temporal dependencies. However, the existing methods often construct the graph only based on road network connectivity, which limits the interaction between roads. In this work, we propose Geographic and Long-term Temporal Graph Convolutional Recurrent Neural Network (GLT-GCRNN), a novel framework for traffic forecasting that learns the rich interactions between roads sharing similar geographic or longterm temporal patterns. Extensive experiments on a real-world traffic state dataset validate the effectiveness of our method by showing that GLT-GCRNN outperforms the state-of-the-art methods in terms of different metrics.
Yulu Wang, Kun Fu 0002, Zheng Wang 0010, Changshui Zhang, Jieping Ye
ICPR5
2020 Diversity in Neural Architecture Search
abstract
Neural architecture search (NAS) is usually divided into two phases: model search, where candidate architectures go through an early training for a small number of epochs (e.g., 20) and a search strategy is used to find one or multiple top candidates, and model tuning, where the top candidates are trained fully (e.g., for 600 epochs) and one final best architecture is chosen. The top M-best strategy (M-Best) is typically used to help find better candidates during model search. However, the top M best solutions may concentrate in narrow similar areas and do not have enough diversity. Furthermore, empirical evidence suggests that performance distribution of the models which only go through the early training does not have a strong correlation with that of the models trained fully. Therefore, many of the M best solutions may turn out to be sub-optimal simultaneously because of their similarity, which limits the ability to find true top architectures. To alleviate the problems, we define diverse M-best architectures that are both of high quality and sufficiently different from each other based on a novel graph-based architecture distance. The concept is very general and is applicable to existing architecture search methods using top M-Best. To the best of our knowledge, this is the first time that diversity is introduced into architecture search. We applied the method in the progressive neural architecture search (PNAS) algorithm (Liu et al. 2018a). Experimental results show that our diverse M-Best is indeed beneficial for finding better architectures.
Wenzheng Hu, Changhe Yuan, Changshui Zhang, Jianqiang Wang 0003
IJCNN4
2020 Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient Space
abstract
Distilling knowledge from an ensemble of teacher models is expected to have a more promising performance than that from a single one. Current methods mainly adopt a vanilla average rule, i.e., to simply take the average of all teacher losses for training the student network. However, this approach treats teachers equally and ignores the diversity among them. When conflicts or competitions exist among teachers, which is common, the inner compromise might hurt the distillation performance. In this paper, we examine the diversity of teacher models in the gradient space and regard the ensemble knowledge distillation as a multi-objective optimization problem so that we can determine a better optimization direction for the training of student network. Besides, we also introduce a tolerance parameter to accommodate disagreement among teachers. In this way, our method can be seen as a dynamic weighting method for each teacher in the ensemble. Extensive experiments validate the effectiveness of our method for both logits-based and feature-based cases.
Shangchen Du, Shan You, Jianlong Wu, Fei Wang 0032, Chen Qian 0006, Changshui Zhang
NeurIPS7
2020 When Counterpoint Meets Chinese Folk Melodies
abstract
Counterpoint is an important concept in Western music theory. In the past century, there have been significant interests in incorporating counterpoint into Chinese folk music composition. In this paper, we propose a reinforcement learning-based system, named FolkDuet, towards the online countermelody generation for Chinese folk melodies. With no existing data of Chinese folk duets, FolkDuet employs two reward models based on out-of-domain data, i.e. Bach chorales, and monophonic Chinese folk melodies. An interaction reward model is trained on the duets formed from outer parts of Bach chorales to model counterpoint interaction, while a style reward model is trained on monophonic melodies of Chinese folk songs to model melodic patterns. With both rewards, the generator of FolkDuet is trained to generate countermelodies while maintaining the Chinese folk style. The entire generation process is performed in an online fashion, allowing real-time interactive human-machine duet improvisation. Experiments show that the proposed algorithm achieves better subjective and objective results than the baselines.
Nan Jiang 0023, Sheng Jin 0007, Zhiyao Duan, Changshui Zhang
NeurIPS4
2020 Zero-shot Handwritten Chinese Character Recognition with hierarchical decomposition embedding
Zhong Cao 0001, Jiang Lu, Sen Cui, Changshui Zhang
Pattern Recognit.4
2020 Deep Gesture Video Generation With Learning on Regions of Interest
abstract
Generating videos with semantic meaning, such as gestures in sign language, is a challenging problem. The model should not only learn to generate videos with realistic appearance, but also take notice of crucial details in frames to convey precise information. In this paper, we focus on the problem of generating long-term gesture videos containing precise and complete semantic meanings. We develop a novel architecture to learn the temporal and spatial transforms in regions of interest, i.e., gesticulating hands or face in our case. We adopt a hierarchical approach for generating gesture videos, by first making predictions on future pose configurations, and then using the encoder-decoder architecture to synthesize future frames based on the predicted pose structures. We develop the scheme of action progress in our architecture to represent how far the action has been performed during its expected execution, and to instruct our model to synthesize actions with various paces. Our approach is evaluated on two challenging datasets for the task of gesture video generation. Experimental results show that our method can produce gesture videos with more realistic appearance and precise meaning than the state-of-the-art video generation approaches.
Runpeng Cui, Zhong Cao 0001, Weishen Pan, Changshui Zhang, Jianqiang Wang 0003
IEEE Trans. Multim.4
2020 Automatically Design Convolutional Neural Networks by Optimization With Submodularity and Supermodularity
abstract
The architecture of convolutional neural networks (CNNs) is a key factor of influencing their performance. Although deep CNNs perform well in many difficult problems, how to intelligently design the architecture is still a challenging problem. Focusing on two practical architectural design problems: to maximize the accuracy with a given forward running time and to minimize the forward running time with a given accuracy requirement, we innovatively utilize prior knowledge to convert architecture optimization problems into submodular optimization problems. We propose efficient Greedy algorithms to solve them and give theoretical bounds of our algorithms. Specifically, we employ the techniques on some public data sets and compare our algorithms with some other hyperparameter optimization methods. Experiments show our algorithms' efficiency.
Wenzheng Hu, Junqi Jin, Tie-Yan Liu, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2020 Compressing Deep Neural Networks With Sparse Matrix Factorization
abstract
Modern deep neural networks (DNNs) are usually overparameterized and composed of a large number of learnable parameters. One of a few effective solutions attempts to compress DNN models via learning sparse weights and connections. In this article, we follow this line of research and present an alternative framework of learning sparse DNNs, with the assistance of matrix factorization. We provide an underlying principle for substituting the original parameter matrices with the multiplications of highly sparse ones, which constitutes the theoretical basis of our method. Experimental results demonstrate that our method substantially outperforms previous states of the arts for compressing various DNNs, giving rich empirical evidence in support of its effectiveness. It is also worth mentioning that, unlike many other works that focus on feedforward networks like multi-layer perceptrons and convolutional neural networks only, we also evaluate our method on a series of recurrent networks in practice.
Kailun Wu, Yiwen Guo, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2019 What You See is What You Get: Visual Pronoun Coreference Resolution in Dialogues
abstract
Xintong Yu, Hongming Zhang, Yangqiu Song, Yan Song, Changshui Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Yan Song 0003, Changshui Zhang
EMNLP/IJCNLP (1)5
2019 Image Captioning with Partially Rewarded Imitation Learning
abstract
Current state-of-the-art image captioning algorithms have achieved great progress via reinforcement learning or generative adversarial nets, with hand-craft metrics such as CIDEr as the reward for the former and signals from adversarial discriminative networks for the latter. Despite the high scores on metrics or improvement in diversity gained from the application of these methods, they suffer from distinction with human-written sentences and drop of ratings on metrics respectively.In this paper, we propose a novel training objective for image captioning that consists of two parts representing explicit and implicit knowledge respectively. Optimizing the new reward partially with imitation learning, we devise an algorithm in which the caption generator is trained to maximize the combination of CIDEr and predictions from adversarial discriminator. Experiments on MSCOCO dataset demonstrate that the proposed method can integrate the strengths of state-of-the-arts, producing more human-like captions while maintaining comparable performance on traditional metrics.
Xintong Yu 0002, Tszhang Guo, Kun Fu 0002, Changshui Zhang, Jianwei Zhang 0001
IJCNN5
2019 Subspace Attack: Exploiting Promising Subspaces for Query-Efficient Black-box Attacks
abstract
Unlike the white-box counterparts that are widely studied and readily accessible, adversarial examples in black-box settings are generally more Herculean on account of the difficulty of estimating gradients. Many methods achieve the task by issuing numerous queries to target classification systems, which makes the whole procedure costly and suspicious to the systems. In this paper, we aim at reducing the query complexity of black-box attacks in this category. We propose to exploit gradients of a few reference models which arguably span some promising search subspaces. Experimental results show that, in comparison with the state-of-the-arts, our method can gain up to 2x and 4x reductions in the requisite mean and medium numbers of queries with much lower failure rates even if the reference models are trained on a small and inadequate dataset disjoint to the one for training the victim model. Code and models for reproducing our results will be made publicly available.
Yiwen Guo, Ziang Yan, Changshui Zhang
NeurIPS3
2019 Hierarchical automatic curriculum learning: Converting a sparse reward navigation task into dense reward
Nan Jiang 0023, Sheng Jin 0007, Changshui Zhang
Neurocomputing3
2019 Audio-Visual Deep Clustering for Speech Separation
abstract
Speech separation aims to separate individual voices from an audio mixture of multiple simultaneous talkers. Audio-only approaches show unsatisfactory performance when the speakers are of the same gender or share similar voice characteristics. This is due to challenges on learning appropriate feature representations for separating voices in single frames and streaming voices across time. Visual signals of speech (e.g., lip movements), if available, can be leveraged to learn better feature representations for separation. In this paper, we propose a novel audio-visual deep clustering model (AVDC) to integrate visual information into the process of learning better feature representations (embeddings) for Time-Frequency (T-F) bin clustering. It employs a two-stage audio-visual fusion strategy where speaker-wise audio-visual T-F embeddings are first computed after the first-stage fusion to model the audio-visual correspondence for each speaker. In the second-stage fusion, audio-visual embeddings of all speakers and audio embeddings calculated by deep clustering from the audio mixture are concatenated to form the final T-F embedding for clustering. Through a series of experiments, the proposed AVDC model is shown to outperform the audio-only deep clustering and utterance-level permutation invariant training baselines and three other state-of-the-art audio-visual approaches. Further analyses show that the AVDC model learns a better T-F embedding for alleviating the source permutation problem across frames. Other experiments show that the AVDC model is able to generalize across different numbers of speakers between training and testing and shows some robustness when visual information is partially missing.
Rui Lu 0003, Zhiyao Duan, Changshui Zhang
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A Deep Neural Framework for Continuous Sign Language Recognition by Iterative Training
abstract
This work develops a continuous sign language (SL) recognition framework with deep neural networks, which directly transcribes videos of SL sentences to sequences of ordered gloss labels. Previous methods dealing with continuous SL recognition usually employ hidden Markov models with limited capacity to capture the temporal information. In contrast, our proposed architecture adopts deep convolutional neural networks with stacked temporal fusion layers as the feature extraction module, and bidirectional recurrent neural networks as the sequence learning module. We propose an iterative optimization process for our architecture to fully exploit the representation capability of deep neural networks with limited data. We first train the end-to-end recognition model for alignment proposal, and then use the alignment proposal as strong supervisory information to directly tune the feature extraction module. This training process can run iteratively to achieve improvements on the recognition performance. We further contribute by exploring the multimodal fusion of RGB images and optical flow in sign language. Our method is evaluated on two challenging SL recognition benchmarks, and outperforms the state of the art by a relative improvement of more than 15% on both databases.
Runpeng Cui, Hu Liu 0001, Changshui Zhang
IEEE Trans. Multim.3
2019 An Accelerated Linearly Convergent Stochastic L-BFGS Algorithm
abstract
The limited memory version of the Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm is the most popular quasi-Newton algorithm in machine learning and optimization. Recently, it was shown that the stochastic L-BFGS (sL-BFGS) algorithm with the variance-reduced stochastic gradient converges linearly. In this paper, we propose a new sL-BFGS algorithm by importing a proper momentum. We prove an accelerated linear convergence rate under mild conditions. The experimental results on different data sets also verify this acceleration advantage.
Daqing Chang, Shiliang Sun, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2018 Multi-Scale Recurrent Neural Network for Sound Event Detection
abstract
Sound event detection (SED) in real life is an interesting but challenging task due to the polyphonic and long-term dependent nature of sound events. Recently, multi-label recurrent neural networks (RNNs) have shown promises. However, even equipped with long short-term memory (LSTM) or gated recurrent unit (GRU) cells, RNNs are still limited to model the long-term dependency. In this paper, we propose a multiscale RNN to address this issue. By integrating information from different time resolutions, we can better capture both the fine-grained and long-term dependencies of sound events. We experiment on the development sets of Task3 of DCASE2016 and DCASE2017. Compared to our previously proposed single-scale RNN that won the third place among the 13 teams in Task3 of DCASE2017, the proposed multiscale model achieves statistically significantly better performance on the development datasets of both DECASE2016 and DCASE2017.
Rui Lu 0003, Zhiyao Duan, Changshui Zhang
ICASSP3
2018 Deep Generative Adversarial Networks for the Sparse Signal Denoising
abstract
In many practical denoising problems, noisy signal contains lots of sparse information, which is helpful for denoising. However, common denoising methods such as low-pass filtering methods and Wavelet denoising methods have a number of limitations like the rigid projection space or loss of high frequency component. In this paper, we propose a deep learning framework based on Generative Adversarial Networks (GANs) to deal with the sparse denoising tasks. We design the Generative Network (G-net) as denoising model with three parts, which are encoding part, denoising part and linear recovery part. To maintain the original features of the data, we utilize the Discriminator Network (D-net) to help the denoising model G-net learn. The experimental results show that our framework is more effective than some traditional methods and state-of-art deep learning methods. In particular, sparse denoising GANs can recover details of picture better in the MNIST image tasks.
Kailun Wu, Changshui Zhang
ICPR2
2018 Boosting Few-Shot Image Recognition Via Domain Alignment Prototypical Networks
abstract
Human has the ability of drawing inferences about other things from only one instance. Few-shot learning is aimed at imitating this generalized learning behavior of human beings, where the learning machine is expected to recognize novel categories not seen in the training set, given only a few training data for each novel category. In this paper, we enhance the Prototypical Network for few-shot learning tasks by introducing a domain alignment module, which takes into account the domain shifts existing between different categories. Compared to original Prototypical Network (PN), the most excellent model for few-shot learning at present, our proposed Domain Alignment Prototypical Network (DA-PN) is able to abate the distribution differences among the data of training and test classes, further optimizing the embedding space of prototype feature for each category and then boosting few-shot recognition. Comprehensive empirical evidence demonstrates that the proposed DA-PN can yield state-of-the-art few-shot recognition performance on the public benchmark dataset mini-ImageNet as well as a novel proposed few-shot dataset MNIST&CIFAR10.
Jiang Lu, Zhong Cao 0001, Kailun Wu, Changshui Zhang
ICTAI5
2018 Sparse DNNs with Improved Adversarial Robustness
abstract
Deep neural networks (DNNs) are computationally/memory-intensive and vulnerable to adversarial attacks, making them prohibitive in some real-world applications. By converting dense models into sparse ones, pruning appears to be a promising solution to reducing the computation/memory cost. This paper studies classification models, especially DNN-based ones, to demonstrate that there exists intrinsic relationships between their sparsity and adversarial robustness. Our analyses reveal, both theoretically and empirically, that nonlinear DNN-based classifiers behave differently under $l_2$ attacks from some linear ones. We further demonstrate that an appropriately higher model sparsity implies better robustness of nonlinear DNNs, whereas over-sparsified models can be more difficult to resist adversarial examples.
Yiwen Guo, Changshui Zhang, Yurong Chen 0001
NeurIPS3
2018 Connectionist Temporal Classification with Maximum Entropy Regularization
abstract
Connectionist Temporal Classification (CTC) is an objective function for end-to-end sequence learning, which adopts dynamic programming algorithms to directly learn the mapping between sequences. CTC has shown promising results in many sequence learning applications including speech recognition and scene text recognition. However, CTC tends to produce highly peaky and overconfident distributions, which is a symptom of overfitting. To remedy this, we propose a regularization method based on maximum conditional entropy which penalizes peaky distributions and encourages exploration. We also introduce an entropy-based pruning method to dramatically reduce the number of CTC feasible paths by ruling out unreasonable alignments. Experiments on scene text recognition show that our proposed methods consistently improve over the CTC baseline without the need to adjust training settings. Code has been made publicly available at: https://github.com/liuhu-bigeye/enctc.crnn.
Hu Liu 0001, Sheng Jin 0007, Changshui Zhang
NeurIPS3
2018 Deep Defense: Training DNNs with Improved Adversarial Robustness
abstract
Despite the efficacy on a variety of computer vision tasks, deep neural networks (DNNs) are vulnerable to adversarial attacks, limiting their applications in security-critical systems. Recent works have shown the possibility of generating imperceptibly perturbed image inputs (a.k.a., adversarial examples) to fool well-trained DNN classifiers into making arbitrary predictions. To address this problem, we propose a training recipe named "deep defense". Our core idea is to integrate an adversarial perturbation-based regularizer into the classification objective, such that the obtained models learn to resist potential attacks, directly and precisely. The whole optimization problem is solved just like training a recursive network. Experimental results demonstrate that our method outperforms training with adversarial/Parseval regularizations by large margins on various datasets (including MNIST, CIFAR-10 and ImageNet) and different DNN architectures. Code and models for reproducing our results are available at https://github.com/ZiangYan/deepdefense.pytorch.
Ziang Yan, Yiwen Guo, Changshui Zhang
NeurIPS3
2018 Robust finite mixture regression for heterogeneous targets
Jian Liang 0002, Kun Chen 0002, Ming Lin 0002, Changshui Zhang, Fei Wang 0001
Data Min. Knowl. Discov.4
2018 Deep transformer: A framework for 2D text image rectification from planar transformations
Chengzhe Yan, Jie Hu 0003, Changshui Zhang
Neurocomputing3
2018 Attribute-Based Synthetic Network (ABS-Net): Learning more from pseudo feature representations
Jiang Lu, Jin Li 0020, Ziang Yan, Fenghua Mei, Changshui Zhang
Pattern Recognit.5
2018 Listen and Look: Audio-Visual Matching Assisted Speech Source Separation
abstract
Source permutation, i.e., assigning separated signal snippets to wrong sources over time, is a major issue in the state-of-the-art speaker-independent speech source separation methods. In addition to auditory cues, humans also leverage visual cues to solve this problem at cocktail parties: matching lip movements with voice fluctuations helps humans to better pay attention to the speaker of interest. In this letter, we propose an audio-visual matching network to learn the correspondence between voice fluctuations and lip movements. We then propose a framework to apply this network to address the source permutation problem and improve over audio-only speech separation methods. The modular design of this framework makes it easy to apply the matching network to any audio-only speech separation method. Experiments on two-talker mixtures show that the proposed approach significantly improves the separation quality over the state-of-the-art audio-only method. This improvement is especially pronounced on mixtures that the audio-only method fails, in which the speakers often have similar voice characteristics.
Rui Lu 0003, Zhiyao Duan, Changshui Zhang
IEEE Signal Process. Lett.3
2018 A New Anchor-Labeling Method For Oriented Text Detection Using Dense Detection Framework
abstract
This letter proposes a new method for dense scene text detection anchor box labeling using single-shot multibox detection (SSD) as the base framework and VGG16 as the backbone, enhanced for scene text detection. This method can be further generalized to other detection tasks with various aspect ratios. We argue that the IoU criterion used by the dense object detection framework may have low recall ratios in extreme aspect ratio cases and oriented objects, and we propose a new criterion of the anchor-labeling method for these kinds of objects. The result shows that this method has better performance on public datasets compared with the previous labeling methods.
Chengzhe Yan, Kailun Wu, Changshui Zhang
IEEE Signal Process. Lett.3
2018 On the Generalization Ability of Online Gradient Descent Algorithm Under the Quadratic Growth Condition
abstract
Online learning has been successfully applied in various machine learning problems. Conventional analysis of online learning achieves a sharp generalization bound with a strongly convex assumption. In this paper, we study the generalization ability of the classic online gradient descent algorithm under the quadratic growth condition (QGC), a strictly weaker condition than strong convexity. Under some mild assumptions, we prove that the excess risk converges no worse than $O(\log T/T)$ when the data are independently and identically distributed (i.i.d.). When the data are generated from a $\phi $ -mixing process, we achieve the excess risk bound $O(\log T /T+\phi (\tau))$ , where $\phi (\tau)$ is the mixing coefficient capturing the non-i.i.d. attribute. Our key technique is based on the combination of the QGC and the martingale concentrations. Our results indicate that the strong convexity is not necessary to achieve the sharp $O(\log {T}/T)$ convergence rate in online learning. We verify our theories on both synthetic and real-world data.
Daqing Chang, Ming Lin 0002, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2018 Optimizing Top-k Multiclass SVM via Semismooth Newton Algorithm
abstract
Top- performance has recently received increasing attention in large data categories. Advances, like a top- multiclass support vector machine (SVM), have consistently improved the top- accuracy. However, the key ingredient in the state-of-the-art optimization scheme based upon stochastic dual coordinate ascent relies on the sorting method, which yields complexity. In this paper, we leverage the semismoothness of the problem and propose an optimized top- multiclass SVM algorithm, which employs semismooth Newton algorithm for the key building block to improve the training speed. Our method enjoys a local superlinear convergence rate in theory. In practice, experimental results confirm the validity. Our algorithm is four times faster than the existing method in large synthetic problems; Moreover, on real-world data sets it also shows significant improvement in training time.
Dejun Chu, Rui Lu 0003, Jin Li 0020, Xintong Yu 0002, Changshui Zhang, Qing Tao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2018 Image-Text Surgery: Efficient Concept Learning in Image Captioning by Generating Pseudopairs
abstract
Image captioning aims to generate natural language sentences to describe the salient parts of a given image. Although neural networks have recently achieved promising results, a key problem is that they can only describe concepts seen in the training image-sentence pairs. Efficient learning of novel concepts has thus been a topic of recent interest to alleviate the expensive manpower of labeling data. In this paper, we propose a novel method, Image-Text Surgery, to synthesize pseudoimage-sentence pairs. The pseudopairs are generated under the guidance of a knowledge base, with syntax from a seed data set (i.e., MSCOCO) and visual information from an existing large-scale image base (i.e., ImageNet). Via pseudodata, the captioning model learns novel concepts without any corresponding human-labeled pairs. We further introduce adaptive visual replacement, which adaptively filters unnecessary visual features in pseudodata with an attention mechanism. We evaluate our approach on a held-out subset of the MSCOCO data set. The experimental results demonstrate that the proposed approach provides significant performance improvements over state-of-the-art methods in terms of F1 score and sentence quality. An ablation study and the qualitative results further validate the effectiveness of our approach.
Kun Fu 0002, Jin Li 0020, Junqi Jin, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2017 Recurrent Convolutional Neural Networks for Continuous Sign Language Recognition by Staged Optimization
abstract
This work presents a weakly supervised framework with deep neural networks for vision-based continuous sign language recognition, where the ordered gloss labels but no exact temporal locations are available with the video of sign sentence, and the amount of labeled sentences for training is limited. Our approach addresses the mapping of video segments to glosses by introducing recurrent convolutional neural network for spatio-temporal feature extraction and sequence learning. We design a three-stage optimization process for our architecture. First, we develop an end-to-end sequence learning scheme and employ connectionist temporal classification (CTC) as the objective function for alignment proposal. Second, we take the alignment proposal as stronger supervision to tune our feature extractor. Finally, we optimize the sequence learning model with the improved feature representations, and design a weakly supervised detection network for regularization. We apply the proposed approach to a real-world continuous sign language recognition benchmark, and our method, with no extra supervision, achieves results comparable to the state-of-the-art.
Runpeng Cui, Hu Liu 0001, Changshui Zhang
CVPR3
2017 Deep ranking: Triplet MatchNet for music metric learning
abstract
Metric learning for music is an important problem for many music information retrieval (MIR) applications such as music generation, analysis, retrieval, classification and recommendation. Traditional music metrics are mostly defined on linear transformations of handcrafted audio features, and may be improper in many situations given the large variety of music styles and instrumentations. In this paper, we propose a deep neural network named Triplet MatchNet to learn metrics directly from raw audio signals of triplets of music excerpts with human-annotated relative similarity in a supervised fashion. It has the advantage of learning highly nonlinear feature representations and metrics in this end-to-end architecture. Experiments on a widely used music similarity measure dataset show that our method significantly outperforms three state-of-the-art music metric learning methods. Experiments also show that the learned features better preserve the partial orders of the relative similarity than handcrafted features.
Rui Lu 0003, Kailun Wu, Zhiyao Duan, Changshui Zhang
ICASSP4
2017 Learning Efficient Convolutional Networks through Network Slimming
abstract
The deployment of deep convolutional neural networks (CNNs) in many real world applications is largely hindered by their high computational cost. In this paper, we propose a novel learning scheme for CNNs to simultaneously 1) reduce the model size; 2) decrease the run-time memory footprint; and 3) lower the number of computing operations, without compromising accuracy. This is achieved by enforcing channel-level sparsity in the network in a simple but effective way. Different from many existing approaches, the proposed method directly applies to modern CNN architectures, introduces minimum overhead to the training process, and requires no special software/hardware accelerators for the resulting models. We call our approach network slimming, which takes wide and large networks as input models, but during training insignificant channels are automatically identified and pruned afterwards, yielding thin and compact models with comparable accuracy. We empirically demonstrate the effectiveness of our approach with several state-of-the-art CNN models, including VGGNet, ResNet and DenseNet, on various image classification datasets. For VGGNet, a multi-pass version of network slimming gives a 20× reduction in model size and a 5× reduction in computing operations.
Zhuang Liu 0003, Gao Huang 0001, Shoumeng Yan, Changshui Zhang
ICCV6
2017 Robust text image alignment with template for information retrieval
abstract
Alignment is an important preprocessing step for image information retrieval. With template information, images should be aligned precisely to retrieve information inside. These images may contain repeated characters or radicals, and their textures may be not appropriate for keypoint extractions. In this paper, a novel robust image alignment algorithm is proposed for images with text template using clustering and robust subspace learning. Experiment shows it can perform well even in extreme cases.
Chengzhe Yan, Jie Hu 0003, Runpeng Cui, Changshui Zhang
SMC4
2017 Rare Chinese character recognition by Radical Extraction Network
abstract
Building a modern Optical Character Recognition (OCR) system for Chinese is hard due to the large Chinese vocabulary list. Training images for rare Chinese characters are extremely expensive to obtain. Radical-based OCR systems tackle this problem by first extracting and recognizing basic graphical components (i.e., radicals) of a Chinese character. However, how to reliably recognize radicals still remains an open challenge. In this paper, we propose a novel Radical Extraction Network (REN) to extract and recognize radicals using deep Convolutional Neural Network (CNN). REN is end-to-end trainable, and it needs less hand-tunning compared with previous segmentation-based approaches. Deep appearance models for radicals are learned from data in a weakly supervised fashion, and no radical-level annotations are required. We learn to recognize different radicals on commonly used Chinese characters, and transfer the learned deep appearance models to rarely used Chinese characters. Experimental results show that the proposed method helps the classifier to recognize rare Chinese characters.
Ziang Yan, Chengzhe Yan, Changshui Zhang
SMC3
2017 Preface
Hang Li 0001, Xiang Bai, Xuanjing Huang 0001, Changshui Zhang
J. Comput. Sci. Technol.4
2017 Current trends in the development of intelligent unmanned autonomous systems
abstract
Intelligent unmanned autonomous systems are some of the most important applications of artificial intelligence (AI). The development of such systems can significantly promote innovation in AI technologies. This paper introduces the trends in the development of intelligent unmanned autonomous systems by summarizing the main achievements in each technological platform. Furthermore, we classify the relevant technologies into seven areas, including AI technologies, unmanned vehicles, unmanned aerial vehicles, service robots, space robots, marine robots, and unmanned workshops/intelligent plants. Current trends and developments in each area are introduced.
Tao Zhang 0006, Qing Li 0010, Changshui Zhang, Hua-wei Liang, Ping Li 0057, Tianmiao Wang, Shuo Li 0001, Yun-long Zhu
Frontiers Inf. Technol. Electron. Eng.3
2017 Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific Contexts
abstract
Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an image captioning system that exploits the parallel structures between images and sentences. In our model, the process of generating the next word, given the previously generated ones, is aligned with the visual perception experience where the attention shifts among the visual regions-such transitions impose a thread of ordering in visual perception. This alignment characterizes the flow of latent meaning, which encodes what is semantically shared by both the visual scene and the text description. Our system also makes another novel modeling contribution by introducing scene-specific contexts that capture higher-level semantic information encoded in an image. The contexts adapt language models for word generation to specific scene types. We benchmark our system and contrast to published results on several popular datasets, using both automatic evaluation metrics and human evaluation. We show that either region-based attention or scene-specific contexts improves systems without those components. Furthermore, combining these two modeling ingredients attains the state-of-the-art performance.
Kun Fu 0002, Junqi Jin, Runpeng Cui, Fei Sha, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 A faster cutting plane algorithm with accelerated line search for linear SVM
Dejun Chu, Changshui Zhang, Qing Tao 0001
Pattern Recognit.2
2016 Learning with only multiple instance positive bags
abstract
In traditional multiple instance learning (MIL), both positive and negative bags are required to learn a prediction function. However, a high human cost is needed to know the label of each bag—positive or negative. Only positive bags contain our focus (positive instances) while negative bags consist of noise or background (negative instances). So we do not expect to spend too much to label the negative bags. Contrary to our expectation, nearly all existing MIL methods require enough negative bags besides positive ones. In this paper we propose an algorithm called “Positive Multiple Instance” (PMI), which learns a classifier given only a set of positive bags. So the annotation of negative bags becomes unnecessary in our method. PMI is constructed based on the assumption that the unknown positive instances in positive bags be similar each other and constitute one compact cluster in feature space and the negative instances locate outside this cluster. The experimental results demonstrate that PMI achieves the performances close to or a little worse than those of the traditional MIL algorithms on benchmark and real data sets. However, the number of training bags in PMI is reduced significantly compared with traditional MIL algorithms.
Changshui Zhang
IJCNN3
2016 Online kernel learning with nearly constant support vectors
Ming Lin 0002, Lijun Zhang 0005, Rong Jin 0001, Shifeng Weng, Changshui Zhang
Neurocomputing5
2016 DeepFish: Accurate underwater live fish recognition with a deep architecture
Hongwei Qin, Xiu Li 0001, Jian Liang 0002, YiGang Peng, Changshui Zhang
Neurocomputing5
2016 Real-Time Traffic Light Detection With Adaptive Background Suppression Filter
abstract
Traffic light detection plays an important role in intelligent transportation system, and many detection methods have been proposed in recent years. However, illumination variation effect is still of its major technical problem in real urban driving environments. In this paper, we propose a novel vision-based traffic light detection method for driving vehicles, which is fast and robust under different illumination conditions. The proposed method contains two stages: the candidate extraction stage and the recognition stage. On the candidate extraction stage, we propose an adaptive background suppression algorithm to highlight the traffic light candidate regions while suppressing the undesired backgrounds. On the recognition stage, each candidate region is verified and is further classified into different traffic light semantic classes. We evaluate our method on video sequences (more than 5000 frames and labels) captured from urban streets and suburb roads in varying illumination and compared with other vision-based traffic detection approaches. The experiment shows that the proposed method can achieve a desired detection result with high quality and robustness; simultaneously, the whole detection system can meet the real-time processing requirement of about 15 fps on video sequences.
Zhenwei Shi 0001, Zhengxia Zou, Changshui Zhang
IEEE Trans. Intell. Transp. Syst.3
2015 Visual Tracking via Part-Based Model with Auxiliary Objects
abstract
In most model-free tracking algorithms, context of the target is usually taken as source of negative examples for training the appearance model and thus not fully utilized. In fact, the target is embedded in its context, and there are spatial constraints between them with the potential motion correlations and relative locations. This paper presents a part-based model to describe the spatial constraints between the target object and its context. Auxiliary objects are selected as parts of the part based model to represent the context information. These auxiliary objects should have some properties: (i) co-occurrence with the target, (ii) having distinct appearance, and (iii) having homogeneous motion as an invisible visual region. A template-based model concerning about the object appearance is implemented to help estimate the reliability of tracking and make the part-based model updated properly. We test the proposed method on a recent benchmark with 50 video sequences evaluated, and a superior performance is achieved compared with state-of-the-arts.
Runpeng Cui, Guangdong Hou, Changshui Zhang
SMC3
2015 A Real-Time Hand Pose Estimation System with Retrieval
abstract
In this paper, we propose a real-time system of hand pose estimation that infers the hand pose and shows the finger positions in each frame. The system is designed through data driven methodology. For the system to perform in real-time, we employ a retrieval method based on an inverted-file index with edge-based descriptors. To strengthen the discriminability, we combine a robust orientation assignment method with the descriptors. A novel alignment method is designed to transfer pose information from the retrieval results to the recognized image. To refine the retrieval results in the measurement of pose similarity, a mixed criterion that considers both outline deviation and descriptor matching is also proposed. The proposed system is quite scalable with respect to poses, since unobserved poses can be added freely. To evaluate the retrieval performance, we apply criteria that assess about the deviation of the corresponding positions of the query pose from the retrieval results. The evaluation is carried out on a retrieval dataset with manually marked images. Our proposed method effectively estimates pose at 20 frames per second, as demonstrated in our experiments on both real and publicly available datasets.
Guangdong Hou, Runpeng Cui, Changshui Zhang
SMC3
2015 Image Deblurring with Coupled Dictionary Learning
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang
Int. J. Comput. Vis.5
2015 Tree-based compact hashing for approximate nearest neighbor search
Guangdong Hou, Runpeng Cui, Changshui Zhang
Neurocomputing4
2015 Damping proximal coordinate descent algorithm for non-convex regularization
Ming Lin 0002, Guangdong Hou, Changshui Zhang
Neurocomputing4
2015 Solving one-class problem with outlier examples by SVM
Shifeng Weng, Changshui Zhang
Neurocomputing4
2015 Incremental multiple instance outlier detection
Shifeng Weng, Changshui Zhang
Neural Comput. Appl.4
2015 A novel unsupervised approach to discovering regions of interest in traffic images
Zhenyu An, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang
Pattern Recognit.4
2015 Large-scale eigenvector approximation via Hilbert Space Embedding Nyström
Ming Lin 0002, Fei Wang 0001, Changshui Zhang
Pattern Recognit.3
2015 Relaxed sparse eigenvalue conditions for sparse estimation via non-convex regularized regression
Changshui Zhang
Pattern Recognit.2
2015 Sparse Unmixing of Hyperspectral Data Using Spectral A Priori Information
abstract
Given a spectral library, sparse unmixing aims at finding the optimal subset of endmembers from it to model each pixel in the hyperspectral scene. However, sparse unmixing still remains a challenging task due to the usually high mutual coherence of the spectral library. In this paper, we exploit the spectral a priori information in the hyperspectral image to alleviate this difficulty. It assumes that some materials in the spectral library are known to exist in the scene. Such information can be obtained via field investigation or hyperspectral data analysis. Then, we propose a novel model to incorporate the spectral a priori information into sparse unmixing. Based on the alternating direction method of multipliers, we present a new algorithm, which is termed sparse unmixing using spectral a priori information (SUnSPI), to solve the model. Experimental results on both synthetic and real data demonstrate that the spectral a priori information is beneficial to sparse unmixing and that SUnSPI can exploit this information effectively to improve the abundance estimation.
Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang
IEEE Trans. Geosci. Remote. Sens.4
2015 Dependent Online Kernel Learning With Constant Number of Random Fourier Features
abstract
Traditional online kernel learning analysis assumes independently identically distributed (i.i.d.) about the training sequence. Recent studies reveal that when the loss function is smooth and strongly convex, given T i.i.d. training instances, a constant sampling complexity of random Fourier features is sufficient to ensure O(logT/T) convergence rate of excess risk, which is optimal in online kernel learning up to a logT factor. However, the i.i.d. hypothesis is too strong in practice, which greatly impairs their value. In this paper, we study the sampling complexity of random Fourier features in online kernel learning under non-i.i.d. assumptions. We prove that the sampling complexity under non-i.i.d. settings is also constant, but the convergence rate of excess risk is O(logT/T+ ϕ) , where ϕ is the mixing coefficient measuring the extent of non-i.i.d. of training sequence. We conduct experiments both on artificial and real large-scale data sets to verify our theories.
Ming Lin 0002, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2014 Non-Asymptotic Analysis of Relational Learning with One Network
abstract
This theoretical paper is concerned with a rigorous non-asymptotic analysis of relational learning applied to a single network. Under suitable and intuitive conditions on features and clique dependencies over the network, we present the first probably approximately correct (PAC) bound for maximum likelihood estimation (MLE). To our best knowledge, this is the first sample complexity result of this problem. We propose a novel combinational approach to analyze complex dependencies of relational data, which is crucial to our non-asymptotic analysis. The consistency of MLE under our conditions is also proved as the consequence of our sample complexity bound. Finally, our combinational method for analyzing dependent data can be easily generalized to treat other generalized maximum likelihood estimators for relational learning.
Changshui Zhang
AISTATS2
2014 Exploiting the Limits of Structure Learning via Inherent Symmetry
abstract
This theoretical paper is concerned with the structure learning limit for Gaussian Markov random fields from i.i.d. samples. The common strategy is applying the Fano method to a family of restricted ensembles. The efficiency of this method, however, depends crucially on selected restricted ensembles. To break through this limitation, we analyze the whole graph ensemble from high-dimensional geometric and group-theoretical perspectives. The key ingredients of our approach are the geometric property of concentration matrices and the invariance of orthogonal group actions on the symmetric Kullback-Leibler divergence. We then establish the connection of the learning limit and eigenvalues of concentration matrices, which leads to a sharper structure learning limit. To our best knowledge, this is the first paper to consider the structure learning problem via inherent symmetries of the whole ensemble. Finally, our approach can be applicable to other graphical structure learning problems.
Changshui Zhang
AISTATS2
2014 Seeing the Arrow of Time
abstract
We explore whether we can observe Time's Arrow in a temporal sequence - is it possible to tell whether a video is running forwards or backwards? We investigate this somewhat philosophical question using computer vision and machine learning techniques. We explore three methods by which we might detect Time's Arrow in video sequences, based on distinct ways in which motion in video sequences might be asymmetric in time. We demonstrate good video forwards/backwards classification results on a selection of YouTube video clips, and on natively-captured sequences (with no temporally-dependent video compression), and examine what motions the models have learned that help discriminate forwards from backwards time.
Lyndsey C. Pickup, Donglai Wei 0001, Changshui Zhang, Andrew Zisserman, Bernhard Schölkopf, William T. Freeman
CVPR5
2014 Domain Transfer via Multiple Sources Regularization
Shaofeng Hu, Jiangtao Ren, Changshui Zhang, Chaogui Zhang
PAKDD (2)3
2014 Efficient Sparse Recovery via Adaptive Non-Convex Regularizers with Oracle Property
Ming Lin 0002, Rong Jin 0001, Changshui Zhang
UAI3
2014 Special issue on the Sino-foreign-interchange workshop on intelligence science and intelligent data engineering
Yanning Zhang 0001, Zhi-Hua Zhou, Changshui Zhang
Neurocomputing3
2014 Single Remote Sensing Image Dehazing
abstract
Remote sensing images are widely used in various fields. However, they usually suffer from the poor contrast caused by haze. In this letter, we propose a simple, but effective, way to eliminate the haze effect on remote sensing images. Our work is based on the dark channel prior and a common haze imaging model. In order to eliminate halo artifacts, we use a low-pass Gaussian filter to refine the coarse estimated atmospheric veil. We then redefine the transmission, with the aim of preventing the color distortion of the recovered images. The main advantage of the proposed algorithm is its fast speed, while it can also achieve good results. The experimental results demonstrate that our algorithm produces visually appealing dehazing images and retains the very fine details. Moreover, for images containing partly clear and partly hazy areas, our algorithm can also achieve good results.
Jiao Long, Zhenwei Shi 0001, Wei Tang 0016, Changshui Zhang
IEEE Geosci. Remote. Sens. Lett.4
2014 Learning high-dimensional correspondence via manifold learning and local approximation
Chenping Hou, Feiping Nie 0001, Hua Wang 0007, Dongyun Yi, Changshui Zhang
Neural Comput. Appl.5
2014 Multiple rank multi-linear SVM for matrix data classification
Chenping Hou, Feiping Nie 0001, Changshui Zhang, Dongyun Yi, Yi Wu 0003
Pattern Recognit.3
2014 Regularized Tree Partitioning and Its Application to Unsupervised Image Segmentation
abstract
In this paper, we propose regularized tree partitioning approaches. We study normalized cut (NCut) and average cut (ACut) criteria over a tree, forming two approaches: 1) normalized tree partitioning (NTP) and 2) average tree partitioning (ATP). We give the properties that result in an efficient algorithm for NTP and ATP. In addition, we present the relations between the solutions of NTP and ATP over the maximum weight spanning tree of a graph and NCut and ACut over this graph. To demonstrate the effectiveness of the proposed approaches, we show its application to image segmentation over the Berkeley image segmentation data set and present qualitative and quantitative comparisons with state-of-the-art methods.
Jingdong Wang 0001, Huaizu Jiang, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan
IEEE Trans. Image Process.5
2014 Traffic Sign Recognition With Hinge Loss Trained Convolutional Neural Networks
abstract
Traffic sign recognition (TSR) is an important and challenging task for intelligent transportation systems. We describe the details of our model's architecture for TSR and suggest a hinge loss stochastic gradient descent (HLSGD) method to train convolutional neural networks (CNNs). Our CNN consists of three stages (70–110–180) with 1 162 284 trainable parameters. The HLSGD is evaluated on the German Traffic Sign Recognition Benchmark, which offers a faster and more stable convergence and a state-of-the-art recognition rate of 99.65%. We write a graphics processing unit package to train several CNNs and establish the final classifier in an ensemble way.
Junqi Jin, Kun Fu 0002, Changshui Zhang
IEEE Trans. Intell. Transp. Syst.3
2013 High-dimensional Inference via Lipschitz Sparsity-Yielding Regularizers
abstract
Non-convex regularizers are more and more applied to high-dimensional inference with sparsity prior knowledge. In general, the non-convex regularizer is superior to the convex ones in inference but it suffers the difficulties brought by local optimums and massive computation. A "good" regularizer should perform well in both inference and optimization. In this paper, we prove that some non-convex regularizers can be such "good" regularizers. They are a family of sparsity-yielding penalties with proper Lipschitz subgradients. These regularizers keep the superiority of non-convex regularizers in inference. Their estimation conditions based on sparse eigenvalues are weaker than the convex regularizers. Meanwhile, if properly tuned, they behave like convex regularizers since standard proximal methods guarantee to give stationary solutions. These stationary solutions, if sparse enough, are identical to the global solutions. If the solution sequence provided by proximal methods is along a sparse path, the convergence rate to the global optimum is on the order of 1/k where k is the number of iterations.
Changshui Zhang
AISTATS2
2013 Efficient blind separation of reflection layers with nonparametric transformations
abstract
Superimposed images are very common when taking photos behind glass. We address the reflection separation problem using multiple superimposed images photographed in different viewpoints. With viewpoints changing, the reflected scenes could contain arbitrarily complicated variations between mixtures, like human's motions or other nonrigid motions. In this article, we propose a moderate hypothesis to tackle the reflected scenes' arbitrary variations as well as the parametric transformations of transmitted scenes. To rapidly recover high-quality image layers, we propose an Efficient Superimposition Recovering Algorithm (ESRA) by extending the framework of accelerated gradient method. Our recovering method has good converging performance and is more than 30 times faster than state-of-the-art methods. Experimental results on synthetic and real world images demonstrate that our method is promising.
Han Li 0005, Kun Gai, Pinghua Gong, Changshui Zhang
ICASSP4
2013 A General Iterative Shrinkage and Thresholding Algorithm for Non-convex Regularized Optimization Problems
abstract
Non-convex sparsity-inducing penalties have recently received considerable attentions in sparse learning. Recent theoretical investigations have demonstrated their superiority over the convex counterparts in several sparse learning settings. However, solving the non-convex optimization problems associated with non-convex penalties remains a big challenge. A commonly used approach is the Multi-Stage (MS) convex relaxation (or DC programming), which relaxes the original non-convex problem to a sequence of convex problems. This approach is usually not very practical for large-scale problems because its computational cost is a multiple of solving a single convex problem. In this paper, we propose a General Iterative Shrinkage and Thresholding (GIST) algorithm to solve the nonconvex optimization problem for a large class of non-convex penalties. The GIST algorithm iteratively solves a proximal operator problem, which in turn has a closed-form solution for many commonly used penalties. At each outer iteration of the algorithm, we use a line search initialized by the Barzilai-Borwein (BB) rule that allows finding an appropriate step size quickly. The paper also presents a detailed convergence analysis of the GIST algorithm. The efficiency of the proposed algorithm is demonstrated by extensive experiments on large-scale data sets.
Pinghua Gong, Changshui Zhang, Zhaosong Lu, Jieping Ye
ICML (2)2
2013 Robust non-negative matrix factorization via joint sparse and graph regularization
abstract
In real world applications, we often have to deal with some high-dimensional, sparse and noisy data. In this paper, we aim to handle this kind of complex data by a Robust Non-negative Matrix Factorization via joint Sparse and Graph regularization model (RSGNMF). We provide a novel efficient and elegant iterative updating algorithm with rigorous convergence analysis for RSGNMF model. Experimental results on image data sets demonstrate that our RSGNMF model outperforms existing start-of-art methods.
Shizhun Yang, Chenping Hou, Changshui Zhang, Yi Wu 0003, Shifeng Weng
IJCNN3
2013 SVM-SVDD: A New Method to Solve Data Description Problem with Negative Examples
Changshui Zhang
ISNN (1)3
2013 Learning a subspace for clustering via pattern shrinking
Chenping Hou, Feiping Nie 0001, Yuanyuan Jiao, Changshui Zhang, Yi Wu 0003
Inf. Process. Manag.4
2013 Multi-stage multi-task feature learning
Pinghua Gong, Jieping Ye, Changshui Zhang
J. Mach. Learn. Res.3
2013 Robust non-negative matrix factorization via joint sparse and graph regularization for transfer learning
Shizhun Yang, Chenping Hou, Changshui Zhang, Yi Wu 0003
Neural Comput. Appl.3
2013 On the Sample Complexity of Random Fourier Features for Online Learning: How Many Random Fourier Features Do We Need?
abstract
We study the sample complexity of random Fourier features for online kernel learning—that is, the number of random Fourier features required to achieve good generalization performance. We show that when the loss function is strongly convex and smooth, online kernel learning with random Fourier features can achieve an O (log T / T ) bound for the excess risk with only O (1/λ 2 ) random Fourier features, where T is the number of training examples and λ is the modulus of strong convexity. This is a significant improvement compared to the existing result for batch kernel learning that requires O ( T ) random Fourier features to achieve a generalization bound O (1/√T). Our empirical study verifies that online kernel learning with a limited number of random Fourier features can achieve similar generalization performance as online learning using full kernel matrix. We also present an enhanced online learning algorithm with random Fourier features that improves the classification performance by multiple passes of training examples and a partial average.
Ming Lin 0002, Shifeng Weng, Changshui Zhang
ACM Trans. Knowl. Discov. Data3
2012 Efficient Multi-Stage Conjugate Gradient for Trust Region Step
abstract
The trust region step problem, by solving a sphere constrained quadratic programming, plays a critical role in the trust region Newton method. In this paper, we propose an efficient Multi-Stage Conjugate Gradient (MSCG) algorithm to compute the trust region step in a multi-stage manner. Specifically, when the iterative solution is in the interior of the sphere, we perform the conjugate gradient procedure. Otherwise, we perform a gradient descent procedure which points to the inner of the sphere and can make the next iterative solution be a interior point. Subsequently, we proceed with the conjugate gradient procedure again. We repeat the above procedures until convergence. We also present a theoretical analysis which shows that the MSCG algorithm converges. Moreover, the proposed MSCG algorithm can generate a solution in any prescribed precision controlled by a tolerance parameter which is the only parameter we need. Experimental results on large-scale text data sets demonstrate our proposed MSCG algorithm has a faster convergence speed compared with the state-of-the-art algorithms.
Pinghua Gong, Changshui Zhang
AAAI2
2012 Learning similarity metric with SVM
abstract
In this paper, we show how to learn a good similarity metric for SVM classification. We present a novel approach to simultaneously learn a Mahalanobis similarity metric and an SVM classifier. Different from previous approaches, we optimize the Mahalanobis metric directly for minimizing the SVM classification error. Our formulation generalizes the traditional large margin principle used in standard SVM, that is, we maximize the margin-radius-ratio. The learned similarity metric significantly improves the classification performance of standard SVM. Empirical studies on real datasets show the proposed approach achieves higher or comparable classification accuracies compared with state-of-the-art similarity learning methods.
Xiaoqiang Zhu, Pinghua Gong, Changshui Zhang
IJCNN4
2012 Robust multi-task feature learning
abstract
Multi-task learning (MTL) aims to improve the performance of multiple related tasks by exploiting the intrinsic relationships among them. Recently, multi-task feature learning algorithms have received increasing attention and they have been successfully applied to many applications involving high-dimensional data. However, they assume that all tasks share a common set of features, which is too restrictive and may not hold in real-world applications, since outlier tasks often exist. In this paper, we propose a Robust MultiTask Feature Learning algorithm (rMTFL) which simultaneously captures a common set of features among relevant tasks and identifies outlier tasks. Specifically, we decompose the weight (model) matrix for all tasks into two components. We impose the well-known group Lasso penalty on row groups of the first component for capturing the shared features among relevant tasks. To simultaneously identify the outlier tasks, we impose the same group Lasso penalty but on column groups of the second component. We propose to employ the accelerated gradient descent to efficiently solve the optimization problem in rMTFL, and show that the proposed algorithm is scalable to large-size problems. In addition, we provide a detailed theoretical analysis on the proposed rMTFL formulation. Specifically, we present a theoretical bound to measure how well our proposed rMTFL approximates the true evaluation, and provide bounds to measure the error between the estimated weights of rMTFL and the underlying true weights. Moreover, by assuming that the underlying true weights are above the noise level, we present a sound theoretical result to show how to obtain the underlying true shared features and outlier tasks (sparsity patterns). Empirical studies on both synthetic and real-world data demonstrate that our proposed rMTFL is capable of simultaneously capturing shared features among tasks and identifying outlier tasks.
Pinghua Gong, Jieping Ye, Changshui Zhang
KDD3
2012 Multi-Stage Multi-Task Feature Learning
abstract
Multi-task sparse feature learning aims to improve the generalization performance by exploiting the shared features among tasks. It has been successfully applied to many applications including computer vision and biomedical informatics. Most of the existing multi-task sparse feature learning algorithms are formulated as a convex sparse regularization problem, which is usually suboptimal, due to its looseness for approximating an $\ell_0$-type regularizer. In this paper, we propose a non-convex formulation for multi-task sparse feature learning based on a novel regularizer. To solve the non-convex optimization problem, we propose a Multi-Stage Multi-Task Feature Learning (MSMTFL) algorithm. Moreover, we present a detailed theoretical analysis showing that MSMTFL achieves a better parameter estimation error bound than the convex formulation. Empirical studies on both synthetic and real-world data sets demonstrate the effectiveness of MSMTFL in comparison with the state of the art multi-task sparse feature learning algorithms.
Pinghua Gong, Jieping Ye, Changshui Zhang
NIPS3
2012 A brief introduction to the special issue for ISNN2010
Liqing Zhang 0001, James T. Kwok, Changshui Zhang
Neurocomputing3
2012 Gabor face recognition by multi-channel classifier fusion of supervised kernel manifold learning
Zeng-Guang Hou, Changshui Zhang
Neurocomputing5
2012 A general framework for transfer sparse subspace learning
Shizhun Yang, Ming Lin 0002, Chenping Hou, Changshui Zhang, Yi Wu 0003
Neural Comput. Appl.4
2012 Blind Separation of Superimposed Moving Images Using Image Statistics
abstract
We address the problem of blind separation of multiple source layers from their linear mixtures with unknown mixing coefficients and unknown layer motions. Such mixtures can occur when one takes photos through a transparent medium, like a window glass, and the camera or the medium moves between snapshots. To understand how to achieve correct separation, we study the statistics of natural images in the Labelme data set. We not only confirm the well-known sparsity of image gradients, but also discover new joint behavior patterns of image gradients. Based on these statistical properties, we develop a sparse blind separation algorithm to estimate both layer motions and linear mixing coefficients and then recover all layers. This method can handle general parameterized motions, including translations, scalings, rotations, and other transformations. In addition, the number of layers is automatically identified, and all layers can be recovered, even in the underdetermined case where mixtures are fewer than layers. The effectiveness of this technology is shown in experiments on both simulated and real superimposed images.
Kun Gai, Zhenwei Shi 0001, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Efficient Nonnegative Matrix Factorization via projected Newton method
Pinghua Gong, Changshui Zhang
Pattern Recognit.2
2012 Image deblurring with matrix regression and gradient evolution
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang
Pattern Recognit.5
2012 Orthogonal vs. uncorrelated least squares discriminant analysis for feature extraction
Feiping Nie 0001, Shiming Xiang, Yun Liu 0021, Chenping Hou, Changshui Zhang
Pattern Recognit. Lett.5
2012 Discriminative Least Squares Regression for Multiclass Classification and Feature Selection
abstract
This paper presents a framework of discriminative least squares regression (LSR) for multiclass classification and feature selection. The core idea is to enlarge the distance between different classes under the conceptual framework of LSR. First, a technique called ε-dragging is introduced to force the regression targets of different classes moving along opposite directions such that the distances between classes can be enlarged. Then, the ε-draggings are integrated into the LSR model for multiclass classification. Our learning framework, referred to as discriminative LSR, has a compact model form, where there is no need to train two-class machines that are independent of each other. With its compact form, this model can be naturally extended for feature selection. This goal is achieved in terms of L2,1 norm of matrix, generating a sparse learning model for feature selection. The model for multiclass classification and its extension for feature selection are finally solved elegantly and efficiently. Experimental evaluation over a range of benchmark datasets indicates the validity of our method.
Shiming Xiang, Feiping Nie 0001, Gaofeng Meng, Chunhong Pan, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.5
2011 A Fast Dual Projected Newton Method for l1-Regularized Least Squares
abstract
L1-regularized least squares, with the ability of discovering sparse representations, is quite prevalent in the field of machine learning, statistics and signal processing. In this paper, we propose a novel algorithm called Dual Projected Newton Method (DPNM) to solve the l1-regularized least squares problem. In DPNM, we first derive a new dual problem as a box constrained quadratic programming. Then, a projected Newton method is utilized to solve the dual problem, achieving a quadratic convergence rate. Moreover, we propose to utilize some practical techniques, thus it greatly reduces the computational cost and makes DPNM more efficient. Experimental results on six real-world data sets indicate that DPNM is very efficient for solving the l1-regularized least squares problem, by comparing it with state of the art methods.
Pinghua Gong, Changshui Zhang
IJCAI2
2011 Topic model with constrainted word burstiness intensities
abstract
Word burstiness phenomenon, which means that if a word occurs once in a document it is likely to occur repeatedly, has interested the text analysis field recently. Dirichlet Compound Multinomial Latent Dirichlet Allocation (DCMLDA) introduces this word burstiness mechanism into Latent Dirichlet Allocation (LDA). However, in DCMLDA, there is no restriction on the word burstiness intensity of each topic. Consequently, as shown in this paper, the burstiness intensities of words in major topics will become extremely low and the topics' ability to represent different semantic meanings will be impaired. In order to get topics that represent semantic meanings of documents well, we introduce constraints on topics' word burstiness intensities. Experiments demonstrate that DCMLDA with constrained word burstiness intensities achieves better performance than the original one without constraints. Besides, these additional constraints help to reveal the relationship between two key properties inherited from DCM and LDA respectively. These two properties have a great influence on the combined model's performance and their relationship revealed by this paper is an important guidance for further study of topic models.
Shaoze Lei, Shifeng Weng, Changshui Zhang
IJCNN4
2011 Efficient Euclidean projections via Piecewise Root Finding and its application in gradient projection
Pinghua Gong, Kun Gai, Changshui Zhang
Neurocomputing3
2011 Multitask Bregman clustering
Changshui Zhang
Neurocomputing2
2011 Semi-supervised ranking aggregation
Shouchun Chen, Fei Wang 0001, Yangqiu Song, Changshui Zhang
Inf. Process. Manag.4
2011 Active learning with adaptive regularization
Zheng Wang 0010, Shuicheng Yan, Changshui Zhang
Pattern Recognit.3
2011 Interactive Image Segmentation With Multiple Linear Reconstructions in Windows
abstract
This paper proposes an algorithm for interactive image segmentation. The task is formulated as a problem of graph-based transductive classification. Specifically, given an image window, the color of each pixel in it will be reconstructed linearly with those of the remaining pixels in this window. The optimal reconstruction weights will be kept unchanged to linearly reconstruct their class labels. The label reconstruction errors are estimated in each window. These errors are further collected together to develop a learning model. Then, the class information about the user specified foreground and background pixels are integrated into a regularization framework. Under this framework, a globally optimal labeling is finally obtained. The computational complexity is analyzed, and an approach for speeding up the algorithm is presented. Comparative experimental results illustrate the validity of our algorithm.
Shiming Xiang, Chunhong Pan, Feiping Nie 0001, Changshui Zhang
IEEE Trans. Multim.4
2011 Semisupervised Learning Using Negative Labels
abstract
The problem of semisupervised learning has aroused considerable research interests in the past few years. Most of these methods aim to learn from a partially labeled dataset, i.e., they assume that the exact labels of some data are already known. In this paper, we propose to use a novel type of supervision information to guide the process of semisupervised learning, which indicates whether a point does not belong to a specific category. We call this kind of information negative label (NL) and propose a novel approach called NL propagation (NLP) to efficiently make use of this type of information to assist the process of semisupervised learning. Specifically, NLP assumes that nearby points should have similar class indicators. The data labels are propagated under the guidance of NL information and the geometric structure revealed by both labeled and unlabeled points, by employing some specified initialization and parameter matrices. The convergence analysis, out-of-sample extension, parameter determination, computational complexity, and relations to other approaches are presented. We also interpret the proposed approach within the framework of regularization. Promising experimental results on image, digit, spoken letter, and text classification tasks are provided to show the effectiveness of our method.
Chenping Hou, Feiping Nie 0001, Fei Wang 0001, Changshui Zhang, Yi Wu 0003
IEEE Trans. Neural Networks4
2011 Spectral Embedded Clustering: A Framework for In-Sample and Out-of-Sample Spectral Clustering
abstract
Spectral clustering (SC) methods have been successfully applied to many real-world applications. The success of these SC methods is largely based on the manifold assumption, namely, that two nearby data points in the high-density region of a low-dimensional data manifold have the same cluster label. However, such an assumption might not always hold on high-dimensional data. When the data do not exhibit a clear low-dimensional manifold structure (e.g., high-dimensional and sparse data), the clustering performance of SC will be degraded and become even worse than K -means clustering. In this paper, motivated by the observation that the true cluster assignment matrix for high-dimensional data can be always embedded in a linear space spanned by the data, we propose the spectral embedded clustering (SEC) framework, in which a linearity regularization is explicitly added into the objective function of SC methods. More importantly, the proposed SEC framework can naturally deal with out-of-sample data. We also present a new Laplacian matrix constructed from a local regression of each pattern and incorporate it into our SEC framework to capture both local and global discriminative information for clustering. Comprehensive experiments on eight real-world high-dimensional datasets demonstrate the effectiveness and advantages of our SEC framework over existing SC methods and K-means-based clustering methods. Our SEC framework significantly outperforms SC using the Nyström algorithm on unseen data.
Feiping Nie 0001, Ivor W. Tsang, Dong Xu 0001, Changshui Zhang
IEEE Trans. Neural Networks5
2011 Unsupervised Large Margin Discriminative Projection
abstract
We propose a new dimensionality reduction method called maximum margin projection (MMP), which aims to project data samples into the most discriminative subspace, where clusters are most well-separated. Specifically, MMP projects input patterns onto the normal of the maximum margin separating hyperplanes. As a result, MMP only depends on the geometry of the optimal decision boundary and not on the distribution of those data points lying further away from this boundary. Technically, MMP is formulated as an integer programming problem and we propose a column generation algorithm to solve it. Moreover, through a combination of theoretical results and empirical observations we show that the computation time needed for MMP can be treated as linear in the dataset size. Experimental results on both toy and real-world datasets demonstrate the effectiveness of MMP.
Fei Wang 0001, Bin Zhao 0004, Changshui Zhang
IEEE Trans. Neural Networks3
2011 Guest Editorial Introduction to the Special Issue on Pattern Recognition Technologies for Anti-Terrorism Applications
abstract
The five papers in this special issue focus on pattern recognition technologies for anti-terrorism applications.
Sos S. Agaian, Jinshan Tang, Sabah Jassim, C. L. Philip Chen, Changshui Zhang, Yongyan Cao
IEEE Trans. Syst. Man Cybern. Part C5
2011 Regression Reformulations of LLE and LTSA With Locally Linear Transformation
abstract
Locally linear embedding (LLE) and local tangent space alignment (LTSA) are two fundamental algorithms in manifold learning. Both LLE and LTSA employ linear methods to achieve their goals but with different motivations and formulations. LLE is developed by locally linear reconstructions in both high- and low-dimensional spaces, while LTSA is developed with the combinations of tangent space projections and locally linear alignments. This paper gives the regression reformulations of the LLE and LTSA algorithms in terms of locally linear transformations. The reformulations can help us to bridge them together, with which both of them can be addressed into a unified framework. Under this framework, the connections and differences between LLE and LTSA are explained. Illuminated by the connections and differences, an improved LLE algorithm is presented in this paper. Our algorithm learns the manifold in way of LLE but can significantly improve the performance. Experiments are conducted to illustrate this fact.
Shiming Xiang, Feiping Nie 0001, Chunhong Pan, Changshui Zhang
IEEE Trans. Syst. Man Cybern. Part B4
2010 What if the Irresponsible Teachers Are Dominating?
abstract
As the Internet-based crowdsourcing services become more and more popular, learning from multiple teachers or sources has received more attention of the researchers in the machine learning area. In this setting, the learning system is dealing with samples and labels provided by multiple teachers, who in common cases, are non-expert. Their labeling styles and behaviors are usually diverse, some of which are even detrimental to the learning system. Thus, simply putting them together and utilizing the algorithms designed for single-teacher scenario would be not only improper, but also damaging. The problem calls for more specific methods. Our work focuses on a case where the teachers are composed of good ones and irresponsible ones. By irresponsible, we mean the teacher who takes the labeling task not seriously and label the sample at random without inspecting the sample itself. This behavior is quite common when the task is not attractive enough and the teacher just wants to finish it as soon as possible. Sometimes, the irresponsible teachers could take a considerable part among all the teachers. If we do not take out their effects, our learning system would be ruined with no doubt. In this paper, we propose a method for picking out the good teachers with promising experimental results. It works even when the irresponsible teachers are dominating in numbers.
Shuo Chen 0008, Guangyun Chen, Changshui Zhang
AAAI4
2010 Learning Discriminative Piecewise Linear Models with Boundary Points
abstract
We introduce a new discriminative piecewise linear model for classification. A two-step method is developed to construct the model. In the first step, we sample some boundary points that lie between positive and negative data, as well as corresponding directions from negative data to positive data. The sampling result gives a discriminative nonparametric decision surface, which preserves enough information to correctly classify all training data. To simplify this surface, in the second step we propose a nonparametric approach for linear surface segmentation using Dirichlet process mixtures. The final result is a piecewise linear model, in which the number of linear surface pieces is automatically determined by the Bayesian inference according to data. Experiments on both synthetic and real data verify the effectiveness of the proposed model.
Kun Gai, Changshui Zhang
AAAI2
2010 Multitask Bregman Clustering
abstract
Traditional clustering methods deal with a single clustering task on a single data set. However, in some newly emerging applications, multiple similar clustering tasks are involved simultaneously. In this case, we not only desire a partition for each task, but also want to discover the relationship among clusters of different tasks. It's also expected that the learnt relationship among tasks can improve performance of each single task. In this paper, we propose a general framework for this problem and further suggest a specific approach. In our approach, we alternatively update clusters and learn relationship between clusters of different tasks, and the two phases boost each other. Our approach is based on the general Bregman divergence, hence it's suitable for a large family of assumptions on data distributions and divergences. Empirical results on several benchmark data sets validate the approach.
Changshui Zhang
AAAI2
2010 Compressed Learning with Regular Concept
Jiawei Lv, Fei Wang 0001, Zheng Wang 0010, Changshui Zhang
ALT5
2010 Homotopy Regularization for Boosting
abstract
In this paper, we present a homotopy regularization algorithm for boosting. We introduce a regularization term with adaptive weight into the boosting framework and compose a homotopy objective function. Optimization of this objective approximately composes a solution path for the regularized boosting. Following this path, we can find suitable solution efficiently using early stopping. Experiments show that this adaptive regularization method gives a more efficient parameter selection strategy than regularized boosting and semi supervised boosting algorithms, and significantly improves the performances of traditional AdaBoost and related methods.
Zheng Wang 0010, Yangqiu Song, Changshui Zhang
ICDM3
2010 Evolutionary hierarchical dirichlet processes for multiple correlated time-varying corpora
abstract
Mining cluster evolution from multiple correlated time-varying text corpora is important in exploratory text analytics. In this paper, we propose an approach called evolutionary hierarchical Dirichlet processes (EvoHDP) to discover interesting cluster evolution patterns from such text data. We formulate the EvoHDP as a series of hierarchical Dirichlet processes~(HDP) by adding time dependencies to the adjacent epochs, and propose a cascaded Gibbs sampling scheme to infer the model. This approach can discover different evolving patterns of clusters, including emergence, disappearance, evolution within a corpus and across different corpora. Experiments over synthetic and real-world multiple correlated time-varying data sets illustrate the effectiveness of EvoHDP on discovering cluster evolution patterns.
Yangqiu Song, Changshui Zhang, Shixia Liu
KDD3
2010 Learning Kernels with Radiuses of Minimum Enclosing Balls
abstract
In this paper, we point out that there exist scaling and initialization problems in most existing multiple kernel learning (MKL) approaches, which employ the large margin principle to jointly learn both a kernel and an SVM classifier. The reason is that the margin itself can not well describe how good a kernel is due to the negligence of the scaling. We use the ratio between the margin and the radius of the minimum enclosing ball to measure the goodness of a kernel, and present a new minimization formulation for kernel learning. This formulation is invariant to scalings of learned kernels, and when learning linear combination of basis kernels it is also invariant to scalings of basis kernels and to the types (e.g., L1 or L2) of norm constraints on combination coefficients. We establish the differentiability of our formulation, and propose a gradient projection algorithm for kernel learning. Experiments show that our method significantly outperforms both SVM with the uniform combination of basis kernels and other state-of-art MKL approaches.
Kun Gai, Guangyun Chen, Changshui Zhang
NIPS3
2010 Boosting with pairwise constraints
Changshui Zhang, Qutang Cai 0001, Yangqiu Song
Neurocomputing1
2010 A general kernelization framework for learning algorithms based on kernel PCA
Changshui Zhang, Feiping Nie 0001, Shiming Xiang
Neurocomputing1
2010 Collaborative Learning by Boosting in Distributed Environments
abstract
In human society, people learn from each other and knowledge is accumulated from generation to generation. This provides some hints to distributed learning. For distributed applications, each site has its own data. If we can build a local model for each site and improve the model based on models learned by its neighbor sites with low communication cost, then it would be very helpful to the distributed applications. In this paper, we propose a new distributed learning method called distributed network boosting (DNB) algorithm for distributed applications. The learned hypotheses are exchanged between neighboring sites during learning process. Theoretical analysis shows that the DNB algorithm minimizes the cost function through collaborative functional gradient descent in hypotheses space. We also give upper bounds of training error and generalization error of the DNB algorithm. Comparison results of the DNB algorithm with other algorithms on real data sets with different sizes show the effectiveness of the proposed algorithm for distributed applications. In order to show the influence of network topology on the performance of the DNB algorithm, we tested it on random graphs and scale-free networks. Bias-variance decomposition shows that the network topology plays an important role in controlling the diversity of the learned classifier ensemble.
Changshui Zhang
Int. J. Pattern Recognit. Artif. Intell.2
2010 A general graph-based semi-supervised learning with novel class discovery
Feiping Nie 0001, Shiming Xiang, Yun Liu 0021, Changshui Zhang
Neural Comput. Appl.4
2010 Semi-Supervised Classification via Local Spline Regression
abstract
This paper presents local spline regression for semi-supervised classification. The core idea in our approach is to introduce splines developed in Sobolev space to map the data points directly to be class labels. The spline is composed of polynomials and Green's functions. It is smooth, nonlinear, and able to interpolate the scattered data points with high accuracy. Specifically, in each neighborhood, an optimal spline is estimated via regularized least squares regression. With this spline, each of the neighboring data points is mapped to be a class label. Then, the regularized loss is evaluated and further formulated in terms of class label vector. Finally, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on the labeled and unlabeled data. To achieve the goal of semi-supervised classification, an objective function is constructed by combining together the global loss of the local spline regressions and the squared errors of the class labels of the labeled data. In this way, a transductive classification algorithm is developed in which a globally optimal classification can be finally obtained. In the semi-supervised learning setting, the proposed algorithm is analyzed and addressed into the Laplacian regularization framework. Comparative classification experiments on many public data sets and applications to interactive image segmentation and image matting illustrate the validity of our method.
Shiming Xiang, Feiping Nie 0001, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Multiple view semi-supervised dimensionality reduction
Chenping Hou, Changshui Zhang, Yi Wu 0003, Feiping Nie 0001
Pattern Recognit.2
2010 A multilevel approach for learning from labeled and unlabeled data on graphs
Changshui Zhang, Fei Wang 0001
Pattern Recognit.1
2010 Interactive localized content based image retrieval with multiple-instance active learning
Dan Zhang 0007, Fei Wang 0001, Zhenwei Shi 0001, Changshui Zhang
Pattern Recognit.4
2010 Multiple Fundamental Frequency Estimation by Modeling Spectral Peaks and Non-Peak Regions
abstract
This paper presents a maximum-likelihood approach to multiple fundamental frequency (F0) estimation for a mixture of harmonic sound sources, where the power spectrum of a time frame is the observation and the F0s are the parameters to be estimated. When defining the likelihood model, the proposed method models both spectral peaks and non-peak regions (frequencies further than a musical quarter tone from all observed peaks). It is shown that the peak likelihood and the non-peak region likelihood act as a complementary pair. The former helps find F0s that have harmonics that explain peaks, while the latter helps avoid F0s that have harmonics in non-peak regions. Parameters of these models are learned from monophonic and polyphonic training data. This paper proposes an iterative greedy search strategy to estimate F0s one by one, to avoid the combinatorial problem of concurrent F0 estimation. It also proposes a polyphony estimation method to terminate the iterative process. Finally, this paper proposes a postprocessing method to refine polyphony and F0 estimates using neighboring frames. This paper also analyzes the relative contributions of different components of the proposed method. It is shown that the refinement component eliminates many inconsistent estimation errors. Evaluations are done on ten recorded four-part J. S. Bach chorales. Results show that the proposed method shows superior F0 estimation and polyphony estimation compared to two state-of-the-art algorithms.
Zhiyao Duan, Bryan Pardo, Changshui Zhang
IEEE Trans. Speech Audio Process.3
2010 Flexible Manifold Embedding: A Framework for Semi-Supervised and Unsupervised Dimension Reduction
abstract
We propose a unified manifold learning framework for semi-supervised and unsupervised dimension reduction by employing a simple but effective linear regression function to map the new data points. For semi-supervised dimension reduction, we aim to find the optimal prediction labels F for all the training samples X, the linear regression function h(X) and the regression residue F(0) = F - h(X) simultaneously. Our new objective function integrates two terms related to label fitness and manifold smoothness as well as a flexible penalty term defined on the residue F(0). Our Semi-Supervised learning framework, referred to as flexible manifold embedding (FME), can effectively utilize label information from labeled data as well as a manifold structure from both labeled and unlabeled data. By modeling the mismatch between h(X) and F, we show that FME relaxes the hard linear constraint F = h(X) in manifold regularization (MR), making it better cope with the data sampled from a nonlinear manifold. In addition, we propose a simplified version (referred to as FME/U) for unsupervised dimension reduction. We also show that our proposed framework provides a unified view to explain and understand many semi-supervised, supervised and unsupervised dimension reduction techniques. Comprehensive experiments on several benchmark databases demonstrate the significant improvement over existing dimension reduction algorithms.
Feiping Nie 0001, Dong Xu 0001, Ivor W. Tsang, Changshui Zhang
IEEE Trans. Image Process.4
2010 TurboPixel Segmentation Using Eigen-Images
abstract
TurboPixel (TP) is a powerful tool for image over-segmentation. It is fast and can yield a lattice-like structure of superpixel regions with uniform size. This paper presents a method to learn eigen-images from the image to be segmented. Such eigen-images are used to generate the evolution speed in the TP framework. The task is formulated as a problem of pixel clustering. Specifically, for the pixels in each local window, a linear transformation is introduced to map their color vectors to be the cluster indicator vectors. The errors under all such linear transformations are estimated and summed together to obtain an objective function, from which a global optimum is finally obtained. In this process, the eigen-images are constructed. Based on these eigen-images, multidimensional image gradient operator is defined to evaluate the gradient, which is supplied to the TP algorithm to obtain the final superpixel segmentations. The computational issues are discussed, and an image pyramid is introduced to speed up the computation. Comparative experiments illustrate the effectiveness of our method.
Shiming Xiang, Chunhong Pan, Feiping Nie 0001, Changshui Zhang
IEEE Trans. Image Process.4
2010 Linear time maximum margin clustering
abstract
Maximum margin clustering (MMC) is a newly proposed clustering method which has shown promising performance in recent studies. It extends the computational techniques of support vector machine (SVM) to the unsupervised scenario. Traditionally, MMC is formulated as a nonconvex integer programming problem which makes it difficult to solve. Several methods have been proposed in the literature to solve the MMC problem based on either semidefinite programming (SDP) or alternating optimization. However, these methods are still time demanding when handling large scale data sets, which limits its application in real-world problems. In this paper, we propose a cutting plane maximum margin clustering (CPMMC) algorithm. It first decomposes the nonconvex MMC problem into a series of convex subproblems by making use of the constrained concave-convex procedure (CCCP), then for each subproblem, our algorithm adopts the cutting plane algorithm to solve it. Moreover, we show that the CPMMC algorithm takes O(sn) time to converge with guaranteed accuracy, where n is the number of samples in the data set and s is the sparsity of the data set, i.e., the average number of nonzero features of the data samples. We also derive the multiclass version of our CPMMC algorithm. Experimental evaluations on several real-world data sets show that CPMMC performs better than existing MMC methods, both in efficiency and accuracy.
Fei Wang 0001, Bin Zhao 0004, Changshui Zhang
IEEE Trans. Neural Networks3
2009 Instance- and bag-level manifold regularization for aggregate outputs classification
abstract
Aggregate outputs learning differs from the classical supervised learning setting in that, training samples are packed into bags with only the aggregate outputs (labels for classification or real values for regression) known. This setting of the problem is associated with several kinds of application background. We focus on the aggregate outputs classification problem in this paper, and set up a manifold regularization framework to deal with it. The framework can be of both instance level and bag level for different testing goals. We propose four concrete algorithms based on our framework, each of which can cope with both binary and multi-class scenarios. The experimental results on several datasets suggest that our algorithms outperform the stateof-art technique.
Shuo Chen 0008, Bin Liu 0015, Mingjie Qian, Changshui Zhang
CIKM4
2009 Efficient multi-class unlabeled constrained semi-supervised SVM
abstract
Semi-supervised learning has been successfully applied to many fields such as knowledge management, information retrieval and data mining as it can utilize both labeled and unlabeled data. In this paper, we propose a general semi-supervised framework for multi-class categorization. Many classical supervised and semi-supervised method dealing with binary classification or multi-class classification including the standard regularization and the manifold regularization can be viewed as special cases of this framework. Based on this framework, we propose a novel method called multi-class unlabeled constrained SVM(MCUCSVM) and its special case: multi-class Laplacian SVM(MCLapSVM). We then put forward a general kernel version semi-supervised dual coordinate descent algorithm to efficiently solve MCUCSVM and makes it more applicable to problems with large number of classes and large scale labeled data. Both rigorous theory and promising experimental results on four real datasets show the great performance and remarkable efficiency of MCUCSVM and MCLapSVM.
Mingjie Qian, Feiping Nie 0001, Changshui Zhang
CIKM3
2009 Efficient multi-label classification with hypergraph regularization
abstract
Many computer vision applications, such as image classification and video indexing, are usually multi-label classification problems in which an instance can be assigned to more than one category. In this paper, we present a novel multi-label classification approach with hypergraph regularization that addresses the correlations among different categories. First, a hypergraph is constructed to capture the correlations among different categories, in which each vertex represents one training instance and each hyperedge for one category contains all the instances belonging to the same category. Then, an improved SVM like learning system incorporating the hypergraph regularization, called Rank-HLapSVM, is proposed to handle the multi-label classification problems. We find that the corresponding optimization problem can be efficiently solved by the dual coordinate descent method. Many promising experimental results on the real datasets including ImageCLEF and MediaMill demonstrate the effectiveness and efficiency of the proposed algorithm.
Gang Chen 0003, Fei Wang 0001, Changshui Zhang, Yuli Gao
CVPR4
2009 Blind separation of superimposed images with unknown motions
abstract
We consider the blind separation of source layers from superimposed mixtures thereof, involving unknown motions and unknown mixing coefficients of layers in each mixture. Previous blind separation approaches for such problems assume motions to be uniform translations, and hence are limited for real world applications. In this paper, we develop a sparse blind separation algorithm to estimate both parameterized motions and mixing coefficients. Then, a novel reconstruction approach is presented to recover all layers, by utilizing not only the mixing model but also the statistical properties of natural images. The whole method can handle more general motions than translations, including scalings, rotations and other transformations. In addition, the number of layers is automatically identified, and all layers can be recovered even in the under-determined case where mixtures are fewer than layers. The effectiveness of this technology is shown in the experiments on two simulated mixtures of four layers, real photos containing transparency and reflections, and real crossfade images from videos.
Kun Gai, Zhenwei Shi 0001, Changshui Zhang
CVPR3
2009 Unsupervised Maximum Margin Feature Selection with manifold regularization
abstract
Feature selection plays a fundamental role in many pattern recognition problems. However, most efforts have been focused on the supervised scenario, while unsupervised feature selection remains as a rarely touched research topic. In this paper, we propose manifold-based maximum margin feature selection (M3FS) to select the most discriminative features for clustering. M3FS targets to find those features that would result in the maximal separation of different clusters and incorporates manifold information by enforcing smoothness constraint on the clustering function. Specifically, we define scale factor for each feature to measure its relevance to clustering, and irrelevant features are identified by assigning zero weights. Feature selection is then achieved by the sparsity constraints on scale factors. Computationally, M3FS is formulated as an integer programming problem and we propose a cutting plane algorithm to efficiently solve it. Experimental results on both toy and real-world data sets demonstrate its effectiveness.
Bin Zhao 0004, James T. Kwok, Fei Wang 0001, Changshui Zhang
CVPR4
2009 Beyond Banditron: A Conservative and Efficient Reduction for Online Multiclass Prediction with Bandit Setting Model
abstract
In this paper, we consider a recently proposed supervised learning problem, called online multiclass prediction with bandit setting model. Aiming at learning from partial feedback of online classification results, i.e. ¿true¿ when the predicting label is right or ¿false¿ when the predicting label is wrong, this new kind of problems arouses much of researchers' interest due to its close relations to real world internet applications and human cognitive procedure. While some algorithms have been brought forward, we propose a novel algorithm to deal with such problems. First, we reduce the multiclass prediction problem to binary based on Conservative one-versus-all others Reduction scheme; Then Online Passive-Aggressive Algorithm is embedded as binary learning algorithm to solve the reduced problem. Also we derive a pleasing cumulative mistake bound for our algorithm and a time complexity bound linear to the sample size. Further experimental evaluation on several real world multiclass datasets including RCV1, MNIST, 20 Newsgroups and USPS shows that our method outperforms the existing algorithms with a great improvement.
Guangyun Chen, Gang Chen 0003, Shuo Chen 0008, Changshui Zhang
ICDM5
2009 Sparse Norm-Regularized Reconstructive Coefficients Learning
abstract
Inspired by the fact that the final decision rule is mainly affected by a small subset of the training samples, i.e., Support Vector Machine (SVM) shows that the decision function relies on the few samples that are on or over the margin. We propose a new framework that explicitly strengthen this intuitive fact by adding an l1-norm regularizer. We give different formulations for our framework in different scenarios, and the experiments show that our framework can not only lead to high sparse solutions but also better performance than traditional methods.
Bin Liu 0015, Shuo Chen 0008, Mingjie Qian, Changshui Zhang
ICDM4
2009 Maximum Margin Clustering with Multivariate Loss Function
abstract
This paper presents a simple but powerful extension of the maximum margin clustering (MMC) algorithm that optimizes multivariate performance measure specifically defined for clustering, including normalized mutual information, rand index and F-measure. Different from previous MMC algorithms that always employ the error rate as the loss function, our formulation involves a multivariate loss function that is a non-linear combination of the individual clustering results. Computationally, we propose a cutting plane algorithm to approximately solve the resulting optimization problem with a guaranteed accuracy. Experimental evaluations show clear improvements in clustering performance of our method over previous maximum margin clustering algorithms.
Bin Zhao 0004, James T. Kwok, Changshui Zhang
ICDM3
2009 Selecting Informative Universum Sample for Semi-Supervised Learning
Shuo Chen 0008, Changshui Zhang
IJCAI2
2009 Semi-Supervised Classification on Evolutionary Data
Yangqing Jia, Shuicheng Yan, Changshui Zhang
IJCAI3
2009 Spectral Embedded Clustering
Feiping Nie 0001, Dong Xu 0001, Ivor W. Tsang, Changshui Zhang
IJCAI4
2009 Knowledge Transfer on Hybrid Graph
Zheng Wang 0010, Yangqiu Song, Changshui Zhang
IJCAI3
2009 On-line Evolutionary Exponential Family Mixture
Yangqiu Song, Gang Chen 0003, Changshui Zhang
IJCAI4
2009 A fast mean-field method for large-scale high-dimensional data and its application in colonic polyp detection at CT colonography
abstract
In this paper, we propose a fast mean-field method called LHMF to handle probabilistic models of large-scale data in high dimensional space. By using diffusion map locally linear embedding method which is a non-linear dimensionality reduction method, we first embed the high dimensional data into a low dimensional space. Then we construct a coarse-grained graph which preserves the spectral properties of original weighted graph in the high dimensional space by clustering. A new spin model is defined in the diffusion space and the geometric centroids of clusters represent variables in the new spin model. The calculation demand of mean-field methods can be reduced greatly on the coarse-grained spin model. The final marginal moments of original variables are derived from the states of geometric centroids by using geometric harmonics. We first tested the proposed method on the MNIST hand-written digits dataset. Experimental results show that the LHMF method is competent with consistency approach, a state-of-the-art semi-supervised learning method. Then we applied the proposed method to a large-scale colonic polyp dataset from computed tomography (CT) scans. Free-response operator characteristic analysis shows that our method achieves higher sensitivity with lower false positive rate compared with support vector machines.
Ronald M. Summers, Changshui Zhang
IJCNN3
2009 Efficient Active Learning with Boosting
abstract
This paper presents an active learning strategy for boosting. In this strategy, we construct a novel objective function to unify semi-supervised learning and active learning boosting. Minimization of this objective is achieved through alternating optimization with respect to the classifier ensemble and the queried data set iteratively. Previous semi-supervised learning or active learning methods based on boosting can be viewed as special cases under this framework. More important, we derive an efficient active learning algorithm under this framework, based on a novel query mechanism called query by incremental committee. It does not only save considerable computational cost, but also outperforms conventional active learning methods based on boosting. We report the experimental results on both boosting benchmarks and real-world database, which show the efficiency of our algorithm and verify our theoretical analysis.
Zheng Wang 0010, Yangqiu Song, Changshui Zhang
SDM3
2009 Multiple Kernel Clustering
abstract
Maximum margin clustering (MMC) has recently attracted considerable interests in both the data mining and machine learning communities. It first projects data samples to a kernel-induced feature space and then performs clustering by finding the maximum margin hyperplane over all possible cluster labelings. As in other kernel methods, choosing a suitable kernel function is imperative to the success of maximum margin clustering. In this paper, we propose a multiple kernel clustering (MKC) algorithm that simultaneously finds the maximum margin hyperplane, the best cluster labeling, and the optimal kernel. Moreover, we provide detailed analysis on the time complexity of the MKC algorithm and also extend multiple kernel clustering to the multi-class scenario. Experimental results on both toy and real-world data sets demonstrate the effectiveness and efficiency of the MKC algorithm.
Bin Zhao 0004, James T. Kwok, Changshui Zhang
SDM3
2009 Analysis of classification margin for classification accuracy with applications
Qutang Cai 0001, Changshui Zhang, Chunyi Peng 0001
Neurocomputing2
2009 Collaborative filtering using orthogonal nonnegative matrix tri-factorization
Gang Chen 0003, Fei Wang 0001, Changshui Zhang
Inf. Process. Manag.3
2009 A new approach to discover interlacing data structures in high-dimensional space
Tao Ban, Changshui Zhang, Shigeo Abe
J. Intell. Inf. Syst.2
2009 Embedding new data points for manifold learning via coordinate propagation
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang, Chunxia Zhang 0001
Knowl. Inf. Syst.4
2009 Soft Constraint Harmonic Energy Minimization for Transductive Learning and its Two Interpretations
Changshui Zhang, Feiping Nie 0001, Shiming Xiang, Chenping Hou
Neural Process. Lett.1
2009 Semi-supervised discriminative classification with application to tumorous tissues segmentation of MR brain images
Yangqiu Song, Changshui Zhang, Jianguo Lee, Fei Wang 0001, Shiming Xiang, Dan Zhang 0007
Pattern Anal. Appl.2
2009 Linear Neighborhood Propagation and Its Applications
abstract
In this paper, a novel graph-based transductive classification approach, called Linear Neighborhood Propagation, is proposed. The basic idea is to predict the label of a data point according to its neighbors in a linear way. This method can be cast into a second-order intrinsic Gaussian Markov random field framework. Its result corresponds to a solution to an approximate inhomogeneous biharmonic equation with Dirichlet boundary conditions. Different from existing approaches, our approach provides a novel graph structure construction method by introducing multiple-wise edges instead of pairwise edges, and presents an effective scheme to estimate the weights for such multiple-wise edges. To the best of our knowledge, these two contributions are novel for semi-supervised classification. The experimental results on image segmentation and transductive classification demonstrate the effectiveness and efficiency of the proposed approach.
Jingdong Wang 0001, Fei Wang 0001, Changshui Zhang, Helen C. Shen, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Stable local dimensionality reduction approaches
Chenping Hou, Changshui Zhang, Yi Wu 0003, Yuanyuan Jiao
Pattern Recognit.2
2009 Front-view vehicle detection by Markov chain Monte Carlo method
Yangqing Jia, Changshui Zhang
Pattern Recognit.2
2009 Semi-supervised orthogonal discriminant analysis via label propagation
Feiping Nie 0001, Shiming Xiang, Yangqing Jia, Changshui Zhang
Pattern Recognit.4
2009 Extracting the optimal dimensionality for local tensor discriminant analysis
Feiping Nie 0001, Shiming Xiang, Yangqiu Song, Changshui Zhang
Pattern Recognit.4
2009 Fast nonlinear autocorrelation algorithm for source separation
Zhenwei Shi 0001, Changshui Zhang
Pattern Recognit.2
2009 marginFace: A novel face recognition method by average neighborhood margin maximization
Fei Wang 0001, Xin Wang 0013, Daoqiang Zhang, Changshui Zhang, Tao Li 0001
Pattern Recognit.4
2009 Learning an Orthogonal and Smooth Subspace for Image Classification
abstract
The recent years have witnessed a surge of interests of learning a subspace for image classification, which has aroused considerable researches from the pattern recognition and signal processing fields. However, for image classification, the accuracies of previous methods are not so high since they neglect some particular characters of the image data. In this paper, we propose a new subspace learning method. It constrains that the transformation basis is orthonormal and the derived coefficients are spatially smooth. Classification is then performed in the image subspace. The proposed method can not only represent the intrinsic structure of the image data, but also avoid over-fitting. More importantly, it can be considered as a general framework, within which the performances of other subspace learning methods can be improved in the same way. Some related analyses of the proposed approach are presented. Promising experimental results on different kinds of real images demonstrate the effectiveness of our algorithm for image classification.
Chenping Hou, Feiping Nie 0001, Changshui Zhang, Yi Wu 0003
IEEE Signal Process. Lett.3
2009 Interactive Natural Image Segmentation via Spline Regression
abstract
This paper presents an interactive algorithm for segmentation of natural images. The task is formulated as a problem of spline regression, in which the spline is derived in Sobolev space and has a form of a combination of linear and Green's functions. Besides its nonlinear representation capability, one advantage of this spline in usage is that, once it has been constructed, no parameters need to be tuned to data. We define this spline on the user specified foreground and background pixels, and solve its parameters (the combination coefficients of functions) from a group of linear equations. To speed up spline construction, K-means clustering algorithm is employed to cluster the user specified pixels. By taking the cluster centers as representatives, this spline can be easily constructed. The foreground object is finally cut out from its background via spline interpolation. The computational complexity of the proposed algorithm is linear in the number of the pixels to be segmented. Experiments on diverse natural images, with comparison to existing algorithms, illustrate the validity of our method.
Shiming Xiang, Feiping Nie 0001, Chunxia Zhang 0001, Changshui Zhang
IEEE Trans. Image Process.4
2009 Clustering with Local and Global Regularization
abstract
Clustering is an old research topic in data mining and machine learning. Most of the traditional clustering methods can be categorized as local or global ones. In this paper, a novel clustering method that can explore both the local and global information in the data set is proposed. The method, Clustering with Local and Global Regularization (CLGR), aims to minimize a cost function that properly trades off the local and global costs. We show that such an optimization problem can be solved by the eigenvalue decomposition of a sparse symmetric matrix, which can be done efficiently using iterative methods. Finally, the experimental results on several data sets are presented to show the effectiveness of our method.
Fei Wang 0001, Changshui Zhang, Tao Li 0001
IEEE Trans. Knowl. Data Eng.2
2009 Nonlinear Dimensionality Reduction with Local Spline Embedding
abstract
This paper presents a new algorithm for Nonlinear Dimensionality Reduction (NLDR). Our algorithm is developed under the conceptual framework of compatible mapping. Each such mapping is a compound of a tangent space projection and a group of splines. Tangent space projection is estimated at each data point on the manifold, through which the data point itself and its neighbors are represented in tangent space with local coordinates. Splines are then constructed to guarantee that each of the local coordinates can be mapped to its own single global coordinate with respect to the underlying manifold. Thus, the compatibility between local alignments is ensured. In such a work setting, we develop an optimization framework based on reconstruction error analysis, which can yield a global optimum. The proposed algorithm is also extended to embed out of samples via spline interpolation. Experiments on toy data sets and real-world data sets illustrate the validity of our method.
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2009 Trace Ratio Problem Revisited
abstract
Dimensionality reduction is an important issue in many machine learning and pattern recognition applications, and the trace ratio (TR) problem is an optimization problem involved in many dimensionality reduction algorithms. Conventionally, the solution is approximated via generalized eigenvalue decomposition due to the difficulty of the original problem. However, prior works have indicated that it is more reasonable to solve it directly than via the conventional way. In this brief, we propose a theoretical overview of the global optimum solution to the TR problem via the equivalent trace difference problem. Eigenvalue perturbation theory is introduced to derive an efficient algorithm based on the Newton-Raphson method. Theoretical issues on the convergence and efficiency of our algorithm compared with prior literature are proposed, and are further supported by extensive empirical results.
Yangqing Jia, Feiping Nie 0001, Changshui Zhang
IEEE Trans. Neural Networks3
2009 Block-Quantized Support Vector Ordinal Regression
abstract
Support vector ordinal regression (SVOR) is a recently proposed ordinal regression (OR) algorithm. Despite its theoretical and empirical success, the method has one major bottleneck, which is the high computational complexity. In this brief, we propose a both practical and theoretical guaranteed algorithm, block-quantized support vector ordinal regression (BQSVOR), where we approximate the kernel matrix K with K that is composed of k2 constant blocks. We provide detailed theoretical justification on the approximation accuracy of BQSVOR. Moreover, we prove theoretically that the OR problem with the block-quantized kernel matrix K could be solved by first separating the data samples in the training set into k clusters with kernel k-means and then performing SVOR on the k cluster representatives. Hence, the algorithm leads to an optimization problem that scales only with the number of clusters, instead of the data set size. Finally, experiments on several real-world data sets support the previous analysis and demonstrate that BQSVOR improves the speed of SVOR significantly with guaranteed accuracy.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang
IEEE Trans. Neural Networks3
2008 Instance-level Semisupervised Multiple Instance Learning
Yangqing Jia, Changshui Zhang
AAAI2
2008 Trace Ratio Criterion for Feature Selection
Feiping Nie 0001, Shiming Xiang, Yangqing Jia, Changshui Zhang, Shuicheng Yan
AAAI4
2008 Semi-supervised Classification Using Local and Global Regularization
Fei Wang 0001, Tao Li 0001, Gang Wang 0004, Changshui Zhang
AAAI4
2008 On Discriminative Semi-Supervised Classification
Fei Wang 0001, Changshui Zhang
AAAI2
2008 Multi-View Local Learning
Dan Zhang 0007, Fei Wang 0001, Changshui Zhang, Tao Li 0001
AAAI3
2008 Semi-supervised ranking aggregation
abstract
Ranking aggregation is important in data mining and information retrieval. In this paper, we proposed a semi-supervised ranking aggregation method, in which the order of several item pairs are labeled as side information. The core idea is to learn a ranking function based on the ordering agreement of different rankers. The ranking scores assigned by this ranking function on the labeled data are consistent with the given pairwise order constraints while the ranking scores on the unlabeled data obey the intrinsic manifold structure of the rank items. The experiment results show our method work well.
Shouchun Chen, Fei Wang 0001, Yangqiu Song, Changshui Zhang
CIKM4
2008 Semi-supervised metric learning by maximizing constraint margin
abstract
Distance metric learning is an old problem that has been researched in the supervised learning field for a very long time. In this paper, we consider the problem of learning a proper distance metric under the guidance of some weak supervisory information. Specifically, those information are in the form of pairwise constraints which specify whether a pair of data points are in the same class (must link constraints) or in the different classes (cannot link constraints). Given those constraints, our algorithm aims to learn a distance metric under which the points with must link constraints are pushed as close as possible, while simultaneously the points with cannot link constraints are pulled away as far as possible. Finally the experimental results are presented to show the effectiveness of our method.
Fei Wang 0001, Shouchun Chen, Changshui Zhang, Tao Li 0001
CIKM3
2008 Transductive object cutout
abstract
In this paper, we address the issue of transducing the object cutout model from an example image to novel image instances. We observe that although object and background are very likely to contain similar colors in natural images, it is much less probable that they share similar color configurations. Motivated by this observation, we propose a local color pattern model to characterize the color configuration in a robust way. Additionally, we propose an edge profile model to modulate the contrast of the image, which enhances edges along object boundaries and attenuates edges inside object or background. The local color pattern model and edge model are integrated in a graph-cut framework. Higher accuracy and improved robustness of the proposed method are demonstrated through experimental comparison with state-of-the-art algorithms.
Jingyu Cui, Qiong Yang, Fang Wen 0001, Qiying Wu, Changshui Zhang, Luc Van Gool, Xiaoou Tang
CVPR5
2008 Blindly separating mixtures of multiple layers with spatial shifts
abstract
We address the problem of blindly separating mixtures of multiple layer images with unknown spatial shifts and mixing coefficients. Our proposed method can handle the over-determined, determined and under-determined cases where mixtures are more than, as many as and fewer than layers, respectively. The method is fast in over-determined and determined cases, with the same complexity as the fast Fourier transform (FFT), and can separate more layers from fewer mixtures in the under-determined case. It consists of two main steps. First, a novel sparse blind separation algorithm is applied, to estimate the spatial shifts, the mixing coefficients and the edge image of each layer. Second, all layers are reconstructed, by large scale linear programming in the under-determined case, or by least-squares solutions in other cases. The effectiveness of this technology is shown in the experiments on two simulated mixtures of four layers with spatial shifts, real mixture photos containing transparency and reflections, and real mixture images in a dissolve from a video.
Kun Gai, Zhenwei Shi 0001, Changshui Zhang
CVPR3
2008 Normalized tree partitioning for image segmentation
abstract
In this paper, we propose a novel graph based clustering approach with satisfactory clustering performance and low computational cost. It consists of two main steps: tree fitting and partitioning. We first introduce a probabilistic method to fit a tree to a data graph under the sense of minimum entropy. Then, we propose a novel tree partitioning method under a normalized cut criterion, called Normalized Tree Partitioning (NTP), in which a fast combinatorial algorithm is designed for exact bipartitioning. Moreover, we extend it to k-way tree partitioning by proposing an efficient best-first recursive bipartitioning scheme. Compared with spectral clustering, NTP produces the exact global optimal bipartition, introduces fewer approximations for k-way partitioning and can intrinsically produce superior performance. Compared with bottom-up aggregation methods, NTP adopts a global criterion and hence performs better. Last, experimental results on image segmentation demonstrate that our approach is more powerful compared with existing graph-based approaches.
Jingdong Wang 0001, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan
CVPR4
2008 A weighted subspace approach for improving bagging performance
abstract
Bagging is an ensemble method that uses random resampling of a dataset to construct models. In classification scenarios, the random resampling procedure in bagging induces some classification margin over the dataset. In addition, when perform bagging in different feature subspaces, the resulting classification margins are likely to be diverse. We take into account the diversity of classification margins in feature sub- spaces for improving the performance of bagging. We first study the average error rate of bagging, convert our task into an optimization problem for determining some weights for feature subspaces, and then assign the weights to the sub- spaces via a randomized technique in classifier construction. Experimental results demonstrate that our method is able to further improve the classification accuracy of bagging, and also outperforms several other ensemble methods including AdaBoost, random forests and random subspace method.
Qutang Cai 0001, Chunyi Peng 0001, Changshui Zhang
ICASSP3
2008 Cost-sensitive boosting algorithms as gradient descent
abstract
AdaBoost is a well known boosting method for generating strong ensemble of weak base learners. The procedure of AdaBoost can be fitted in a gradient descent optimization framework, which is important for analyzing and devising its procedure. Cost sensitive boosting (CSB) is an emerging subject extending the boosting methods for cost sensitive classification applications. Most CSB methods are performed by directly modifying the original AdaBoost procedure. Unfortunately, the effectiveness of most cost sensitive boosting methods are checked only by experiments. It remains unclear whether these methods can be viewed as gradient descent procedures like AdaBoost. In this paper, we show that several typical CSB methods can also be view as gradient descent for minimizing a unified objective function. We then deduce a general greedy boosting procedure. Experimental results also validate the effectiveness of the proposed procedure.
Qutang Cai 0001, Yangqiu Song, Changshui Zhang
ICASSP3
2008 Towards optimal query design for relevance feedback in image retrieval
abstract
We analyze the sub-optimality of traditional greedy active learning based relevance feedback methods in image retrieval, and propose a novel active learning approach to query labels of multiple images together, which minimize the needed round of feedbacks and achieve satisfactory result in a near optimal manner. Our experiments on real image retrieval demonstrate that our solution can yield comparable precession/recall rate by significantly less relevance feedbacks.
Jingyu Cui, Changshui Zhang
ICASSP2
2008 Active model selection for Graph-Based Semi-Supervised Learning
abstract
The recent years have witnessed a surge of interest in graph-based semi-supervised learning (GBSSL). However, despite its extensive research, there has been little work on graph construction, which is at the heart of GBSSL. In this study, we propose a novel active learning method, active model selection (AMS), which aims at learning both data labels and the optimal graph by allowing the learner the flexibility to choose samples for labeling. AMS minimizes the regularization function in GBSSL by iterating between the active sample selection step and the graph reconstruction step, where the samples querying which leads to the optimal graph are selected. Experimental results on four real-world datasets are provided to demonstrate the effectiveness of AMS.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang, Yangqiu Song
ICASSP3
2008 Maximum Margin Embedding
abstract
We propose a new dimensionality reduction method called Maximum Margin Embedding (MME), which targets to projecting data samples into the most discriminative subspace, where clusters are most well-separated. Specifically, MME projects input patterns onto the normal of the maximum margin separating hyperplanes. As a result, MME only depends on the geometry of the optimal decision boundary and not on the distribution of those data points lying further away from this boundary. Technically, MME is formulated as an integer programming problem and we propose a cutting plane algorithm to solve it. Moreover, we prove theoretically that the computational time of MME scales linearly with the dataset size. Experimental results on both toy and real world datasets demonstrate the effectiveness of MME.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang
ICDM3
2008 Equivalence Knowledge Mass and Approximate Reasoning in -Logic (I)
Yalin Zheng, Guang Yang 0002, Changshui Zhang, Yunpeng Xu
ICIC (2)3
2008 Augmented tree partitioning for interactive image segmentation
abstract
In this paper, we propose a new fast semi-supervised image segmentation method based on augmented tree partitioning. Unlike many existing methods that use a graph structure to model the image, we use a tree-based structure called the augmented tree, which is built up by augmenting several abstract label nodes to the minimum spanning tree of the original graph. We then model image segmentation as the partitioning problem on the augmented tree. Dynamic programming is used to efficiently solve the optimization problem. Experimental results show that our method gives competitive segmentation results, and the speed is much faster than graph-based methods.
Yangqing Jia, Jingdong Wang 0001, Changshui Zhang, Xian-Sheng Hua 0001
ICIP3
2008 Learning distance metric for semi-supervised image segmentation
abstract
Semi-supervised image segmentation is an important issue in many image processing applications, and has been a popular research area recently, the most popular are graph-based methods. However, parameter selection in these methods is still largely heuristic. In this paper, we introduce distance metric learning into graph-based semi-supervised segmentation to automatically obtain good results for images with different appearances. We first derive the optimization problem with respect to the distance metric as well as the segmentation labels, and use gradient descent method to find a local optimum solution. Experiments on general images and the fungal disease analysis application have shown that our method provides a steady performance under casual user annotations and different image appearances.
Yangqing Jia, Changshui Zhang
ICIP2
2008 Localized content based image retrieval by multiple instance active learning
abstract
In this paper, we propose two general multiple instance active learning (MIAL) algorithms, multiple-instance active learning with a simple margin strategy (S-MIAL) and multiple- instance active learning with fisher information (F-MIAL), and apply them to the relevance feedback in localized content based image retrieval (LCBIR). S-MIAL considers the most ambiguous picture as the most valuable one, while F-MIAL can utilize the fisher information and analyze the value of the unlabeled pictures by assigning different labels to them. We show that F-MIAL can be integrated more naturally into the multiple instance learning scenario. In experiments, we will show their superior performances on some real-world image datasets.
Dan Zhang 0007, Fei Wang 0001, Zhenwei Shi 0001, Changshui Zhang
ICIP4
2008 Audio tonality mode classification without tonic annotations
abstract
Traditional tonality mode (major or minor) classification or audio key finding algorithms often rely on tonic annotations (key names) of the training songs. However, unlike classical music whose keys are usually explicitly labeled in their titles, the keys of numerous popular music are hard to obtain. In contrast, it is much easier to only label the mode for each song. With only modes labeled, traditional approaches to key or mode classification cannot be directly applied, due to the lack of the reference point to transpose and align the chroma features with different keys. In this paper, we present an alignment approach to transpose chroma features within each mode to a reference (but unknown) tonic. Then several methods, including Single Profile Correlation, Multiple Profile Correlation and Support Vector Machine, are exploited to address mode learning and classification. Experimental results show the feasibility of the proposed approach.
Zhiyao Duan, Lie Lu, Changshui Zhang
ICME3
2008 Efficient multiclass maximum margin clustering
abstract
This paper presents a cutting plane algorithm for multiclass maximum margin clustering (MMC). The proposed algorithm constructs a nested sequence of successively tighter relaxations of the original MMC problem, and each optimization problem in this sequence could be efficiently solved using the constrained concave-convex procedure (CCCP). Experimental evaluations on several real world datasets show that our algorithm converges much faster than existing MMC methods with guaranteed accuracy, and can thus handle much larger datasets efficiently.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang
ICML3
2008 Local Regularized Least-Square Dimensionality Reduction
abstract
In this paper, we propose a new nonlinear dimensionality reduction algorithm by adopting regularized least-square criterion on local areas of the data distribution. We first propose a local linear model to describe the characteristic of the low-dimensional coordinates of the neighborhood centered in each data point, and use regularized least-square criterion to evaluate the fitness of the low-dimensional embedding. Next, we form an optimization task similar to the graph Laplacian and efficiently retrieve the solution via eigenvalue decomposition. The relationship between our method and the Laplacian Eigenmaps are discussed, and experimental results are presented.
Yangqing Jia, Changshui Zhang
ICPR2
2008 Collaborative learning by boosting in distributed environments
abstract
In this paper we propose a new distributed learning method called distributed network boosting (DNB) algorithm for distributed applications. The learned hypotheses are exchanged between neighboring sites during learning process. Theoretical analysis shows that the DNB algorithm minimizes the cost function through the collaborative functional gradient descent in hypotheses space. Comparison results of the DNB algorithm with other distributed learning methods on real data sets with different sizes show its effectiveness.
Changshui Zhang
ICPR2
2008 Cuts3vm: a fast semi-supervised svm algorithm
abstract
Semi-supervised support vector machine (S3VM) attempts to learn a decision boundary that traverses through low data density regions by maximizing the margin over labeled and unlabeled examples. Traditionally, S3VM is formulated as a non-convex integer programming problem and is thus difficult to solve. In this paper, we propose the cutting plane semi-supervised support vector machine (CutS3VM) algorithm, to solve the S3VM problem. Specifically, we construct a nested sequence of successively tighter relaxations of the original S3VM problem, and each optimization problem in this sequence could be efficiently solved using the constrained concave-convex procedure (CCCP). Moreover, we prove theoretically that the CutS3VM algorithm takes time O(sn) to converge with guaranteed accuracy, where n is the total number of samples in the dataset and s is the average number of non-zero features, i.e. the sparsity. Experimental evaluations on several real world datasets show that CutS3VM performs better than existing S3VM methods, both in efficiency and accuracy.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang
KDD3
2008 Finding image exemplars using fast sparse affinity propagation
abstract
In this paper, we propose a novel approach to organize image search results obtained from state-of-the-art image search engines in order to improve user experience. We aim to discover exemplars from search results and simultaneously group the images. The exemplars are delivered to the user as a summary of search results instead of the large amount of unorganized images. This gives the user a brief overview of search results with a small amount of images, and helps the user to further find the images of interest. We adopt the idea of affinity propagation and design a fast sparse affinity propagation algorithm to find exemplars that best represent the image search results. Experiments on real-world data demonstrate the effectiveness of our method both visually and quantitatively.
Yangqing Jia, Jingdong Wang 0001, Changshui Zhang, Xian-Sheng Hua 0001
ACM Multimedia3
2008 Distortion-Free Nonlinear Dimensionality Reduction
Yangqing Jia, Zheng Wang 0010, Changshui Zhang
ECML/PKDD (1)3
2008 Transferred Dimensionality Reduction
Zheng Wang 0010, Yangqiu Song, Changshui Zhang
ECML/PKDD (2)3
2008 Semi-supervised Multi-label Learning by Solving a Sylvester Equation
abstract
Multi-label learning refers to the problems where an instance can be assigned to more than one category. In this paper, we present a novel Semi-supervised algorithm for Multi-label learning by solving a Sylvester Equation (SMSE). Two graphs are first constructed on instance level and category level respectively. For instance level, a graph is defined based on both labeled and unlabeled instances, where each node represents one instance and each edge weight reflects the similarity between corresponding pairwise instances. Similarly, for category level, a graph is also built based on all the categories, where each node represents one category and each edge weight reflects the similarity between corresponding pairwise categories. A regularization framework combining two regularization terms for the two graphs is suggested. The regularization term for instance graph measures the smoothness of the labels of instances, and the regularization term for category graph measures the smoothness of the labels of categories. We show that the labels of unlabeled data finally can be obtained by solving a Sylvester Equation. Experiments on RCV1 data set show that SMSE can make full use of the unlabeled data information as well as the correlations among categories and achieve good performance. In addition, we give a SMSE's extended application on collaborative filtering.
Gang Chen 0003, Yangqiu Song, Fei Wang 0001, Changshui Zhang
SDM4
2008 Semi-Supervised Clustering via Matrix Factorization
abstract
The recent years have witnessed a surge of interests of semi-supervised clustering methods, which aim to cluster the data set under the guidance of some supervisory information. Usually those supervisory information takes the form of pairwise constraints that indicate the similarity/dissimilarity between the two points. In this paper, we propose a novel matrix factorization based approach for semi-supervised clustering. In addition, we extend our algorithm to co-cluster the data sets of different types with constraints. Finally the experiments on UCI data sets and real world Bulletin Board Systems (BBS) data sets show the superiority of our proposed method.
Fei Wang 0001, Tao Li 0001, Changshui Zhang
SDM3
2008 Semi-Supervised Classification with Universum
abstract
The Universum data, defined as a collection of “non-examples” that do not belong to any class of interest, have been shown to encode some prior knowledge by representing meaningful concepts in the same domain as the problem at hand. In this paper, we address a novel semi-supervised classification problem, called semi-supervised Universum, that can simultaneously utilize the labeled data, unlabeled data and the Universum data to improve the classification performance. We propose a graph based method to make use of the Universum data to help depict the prior information for possible classifiers. Like conventional graph based semi-supervised methods, the graph regularization is also utilized to favor the consistency between the labels. Furthermore, since the proposed method is a graph based one, it can be easily extended to the multiclass case. The empirical experiments on the USPS and MNIST datasets are presented to show that the proposed method can obtain superior performances over conventional supervised and semi-supervised methods.
Dan Zhang 0007, Jingdong Wang 0001, Fei Wang 0001, Changshui Zhang
SDM4
2008 Efficient Maximum Margin Clustering via Cutting Plane Algorithm
abstract
Maximum margin clustering (MMC) is a recently proposed clustering method, which extends the theory of support vector machine to the unsupervised scenario and aims at finding the maximum margin hyperplane which separates the data from different classes. Traditionally, MMC is formulated as a non-convex integer programming problem and is thus difficult to solve. Several methods have been proposed in the literature to solve the MMC problem based on either semidefinite programming or alternative optimization. However, these methods are time demanding while handling large scale datasets and therefore unsuitable for real world applications. In this paper, we propose the cutting plane maximum margin clustering (CPMMC) algorithm, to solve the MMC problem. Specifically, we construct a nested sequence of successively tighter relaxations of the original MMC problem, and each optimization problem in this sequence could be efficiently solved using the constrained concave-convex procedure (CCCP). Moreover, we prove theoretically that the CPMMC algorithm takes time O(sn) to converge with guaranteed accuracy, where n is the total number of samples in the dataset and s is the average number of non-zero features, i.e. the sparsity. Experimental evaluations on several real world datasets show that CPMMC performs better than existing MMC methods, both in efficiency and accuracy.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang
SDM3
2008 MACBSE: Extracting signals with linear autocorrelations
Zhenwei Shi 0001, Dan Zhang 0007, Changshui Zhang
Neurocomputing3
2008 A unified framework for semi-supervised dimensionality reduction
Yangqiu Song, Feiping Nie 0001, Changshui Zhang, Shiming Xiang
Pattern Recognit.3
2008 The random electrode selection ensemble for EEG signal classification
Shiliang Sun, Changshui Zhang, Yue Lu 0001
Pattern Recognit.2
2008 Contour graph based human tracking and action sequence recognition
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang
Pattern Recognit.4
2008 Learning a Mahalanobis distance metric for data clustering and classification
Shiming Xiang, Feiping Nie 0001, Changshui Zhang
Pattern Recognit.3
2008 Semi-supervised sub-manifold discriminant analysis
Yangqiu Song, Feiping Nie 0001, Changshui Zhang
Pattern Recognit. Lett.3
2008 Unsupervised Single-Channel Music Source Separation by Average Harmonic Structure Modeling
abstract
Source separation of musical signals is an appealing but difficult problem, especially in the single-channel case. In this paper, an unsupervised single-channel music source separation algorithm based on average harmonic structure modeling is proposed. Under the assumption of playing in narrow pitch ranges, different harmonic instrumental sources in a piece of music often have different but stable harmonic structures; thus, sources can be characterized uniquely by harmonic structure models. Given the number of instrumental sources, the proposed algorithm learns these models directly from the mixed signal by clustering the harmonic structures extracted from different frames. The corresponding sources are then extracted from the mixed signal using the models. Experiments on several mixed signals, including synthesized instrumental sources, real instrumental sources, and singing voices, show that this algorithm outperforms the general nonnegative matrix factorization (NMF)-based source separation algorithm, and yields good subjective listening quality. As a side effect, this algorithm estimates the pitches of the harmonic instrumental sources. The number of concurrent sounds in each frame is also computed, which is a difficult task for general multipitch estimation (MPE) algorithms.
Zhiyao Duan, Yungang Zhang, Changshui Zhang, Zhenwei Shi 0001
IEEE Trans. Speech Audio Process.3
2008 Label Propagation through Linear Neighborhoods
abstract
In many practical data mining applications such as text classification, unlabeled training examples are readily available, but labeled ones are fairly expensive to obtain. Therefore, semi supervised learning algorithms have aroused considerable interests from the data mining and machine learning fields. In recent years, graph-based semi supervised learning has been becoming one of the most active research areas in the semi supervised learning community. In this paper, a novel graph-based semi supervised learning approach is proposed based on a linear neighborhood model, which assumes that each data point can be linearly reconstructed from its neighborhood. Our algorithm, named linear neighborhood propagation (LNP), can propagate the labels from the labeled points to the whole data set using these linear neighborhoods with sufficient smoothness. A theoretical analysis of the properties of LNP is presented in this paper. Furthermore, we also derive an easy way to extend LNP to out-of-sample data. Promising experimental results are presented for synthetic data, digit, and text classification tasks.
Fei Wang 0001, Changshui Zhang
IEEE Trans. Knowl. Data Eng.2
2008 Content-Based Information Fusion for Semi-Supervised Music Genre Classification
abstract
In this paper, we propose an information fusion framework for the semi-supervised distance-based music genre classification problem. We make use of the regularized least-square framework as the basic classifier, which only involves the similarity scores among different music tracks. We present a similarity score that multiplies different scores based on different distance measures. Particularly the distance measures are not restricted to the Euclidean distance. By adding a weight to each single distance based score, we propose an expectation-maximization (EM) algorithm to adaptively learn the fusion scores. Experiments on real music data set show that our approach can give promising results.
Yangqiu Song, Changshui Zhang
IEEE Trans. Multim.2
2008 Semisupervised Learning Based on Generalized Point Charge Models
abstract
The recent years have witnessed a surge of interest in semisupervised learning. Numerous methods have been proposed for learning from partially labeled data. In this brief, a novel semisupervised learning approach based on an electrostatic field model is proposed. We treat the labeled data points as point charges, therefore the remaining unlabeled data points are placed in the electrostatic fields generated by these charges. The labels of these unlabeled data points can be regarded as the electric potentials of the electrostatic field at their corresponding places. Moreover, we also develop an efficient way to extend our method for out-of-sample data and analyze theoretically the relationship between our method and the traditional graph-based methods. Finally, the experimental results on both toy and real-world data sets are provided to show the effectiveness of our method.
Fei Wang 0001, Changshui Zhang
IEEE Trans. Neural Networks2
2007 Clustering with Local and Global Regularization
Fei Wang 0001, Changshui Zhang, Tao Li 0001
AAAI2
2007 Localized Content-Based Image Retrieval Using Semi-Supervised Multiple Instance Learning
Dan Zhang 0007, Zhenwei Shi 0001, Yangqiu Song, Changshui Zhang
ACCV (1)4
2007 Ranking with semi-supervised distance metric learning and its application to housing potential estimation
abstract
This paper proposes a semi-supervised distance metric learning algorithm for the ranking problem. Instead of giving the computer what are the important factors that affect the final rank value, we only give several most certainly ranked points which implicitly contain the knowledge of the ranking factors. Then the computer can automatically use the most certain points and plenty of unlabeded data to learn an informative metric for ranking. This metric not only can help to regress an order in the observed data, but also can be used to retrieve the data by querying new test points. Moreover, the lower-rank distance metric can be used to visualize high-dimensional data. We also present an application to the housing potential estimation problem. It is shown that the algorithm is efficient to help consultants to refine their consulting work.
Yangqiu Song, Bin Zhang 0013, Wenjun Yin, Changshui Zhang
CIKM4
2007 Optimal Dimensionality Discriminant Analysis and Its Application to Image Recognition
abstract
Dimensionality reduction is an important issue when facing high-dimensional data. For supervised dimensionality reduction, linear discriminant analysis (LDA) is one of the most popular methods and has been successfully applied in many classification problems. However, there are several drawbacks in LDA. First, it suffers from the singularity problem, which makes it hard to preform. Second, LDA has the distribution assumption which may make it fail in applications where the distribution is more complex than Gaussian. Third, LDA can not determine the optimal dimensionality for discriminant analysis, which is an important issue but has often been neglected previously. In this paper, we propose a new algorithm and endeavor to solve all these three problems. Furthermore, we present that our method can be extended to the two-dimensional case, in which the optimal dimensionalities of the two projection matrices can be determined simultaneously. Experimental results show that our methods are effective and demonstrate much higher performance in comparison to LDA.
Feiping Nie 0001, Shiming Xiang, Yangqiu Song, Changshui Zhang
CVPR4
2007 Feature Extraction by Maximizing the Average Neighborhood Margin
abstract
A novel algorithm called Average Neighborhood Margin Maximization (ANMM) is proposed for supervised linear feature extraction. For each data point, ANMM aims at pulling the neighboring points with the same class label towards it as near as possible, while simultaneously pushing the neighboring points with different labels away from it as far as possible. We will show that features extracted from ANMM can separate the data from different classes well, and it avoids the small sample size problem existed in traditional Linear Discriminant Analysis (LDA). The kernelized (nonlinear) counterpart of ANMM is also established in this paper. Moreover, as in many computer vision applications the data are more naturally represented by higher order tensors (e.g. images and videos), we develop a tensorized (multilinear) form of ANMM, which can directly extract features from tensors. The experimental results of applying ANMM to face recognition are presented to show the effectiveness of our method.
Fei Wang 0001, Changshui Zhang
CVPR2
2007 Discriminant Additive Tangent Spaces for Object Recognition
abstract
Pattern variation is a major factor that affects the performance of recognition systems. In this paper, a novel manifold tangent modeling method called discriminant additive tangent spaces (DATS) is proposed for invariant pattern recognition. In DATS, intra-class variations for traditional tangent learning are called positive tangent samples. In addition, extra-class variations are introduced as negative tangent samples. We use log-odds to measure the significance of samples being positive or negative, and then directly characterizes this log-odds using generalized additive models (GAM). This model is estimated to maximally discriminate positive and negative samples. Besides, since traditional GAM fitting algorithm can not handle the high dimensional data in visual recognition tasks, we also present an efficient, sparse solution for GAM estimation. The resulting DATS is a nonparametric discriminant model based on quite weak prior hypotheses, hence it can depict various pattern variations effectively. Experiments demonstrate the effectiveness of our method in several recognition tasks.
Liang Xiong, Changshui Zhang
CVPR3
2007 Semi-definite Manifold Alignment
Liang Xiong, Fei Wang 0001, Changshui Zhang
ECML3
2007 Analysis of Max-Min Eigenvalue of Constrained Linear Combinations of Symmetric Matrices
abstract
This paper studies the problem whether the smallest eigenvalue of constrained linear combinations of symmetric matrices can reach a desirable value, which actually extends the mathematical problem of finding a positive definite linear combination of symmetric matrices(PDLC), and provides a universal framework to maximize the minimal eigenvalue of linear combined symmetric matrices. For solving this problem, we cast an equivalent optimization task, and propose one general algorithm framework that is proved to be globally optimal and convergent. Both theoretical analysis and experiments under a typical constraint verify our algorithm's validity and efficiency.
Qutang Cai 0001, Chunyi Peng 0001, Changshui Zhang
ICASSP (3)3
2007 Extracting the Optimal Dimensionality for Discriminant Analysis
abstract
For classification task, supervised dimensionality reduction is a very important method when facing with high-dimensional data. Linear discriminant analysis (LDA) is one of the most popular method for supervised dimensionality reduction. However, LDA suffers from the singularity problem, which makes it hard to work. Another problem is the determination of optimal dimensionality for discriminant analysis, which is an important issue but often been neglected previously. In this paper, we propose a new algorithm to address these two problems. Experiments show the effectiveness of our method and demonstrate much higher performance in comparison to LDA.
Feiping Nie 0001, Shiming Xiang, Yangqiu Song, Changshui Zhang
ICASSP (2)4
2007 Semi-Supervised Music Genre Classification
abstract
Music genre classification is a hot topic in pattern recognition and signal processing. Classical supervised methods need lost of labeled music data to train a classifier. In this paper, we propose a semi-supervised genre classification algorithm which is developed on several labeled music tracks and lots of unlabelled tracks. Three features are extracted from the each music track and manifold regularization method is used to design the classifier. Experiments on a large number of test music data show that semi-supervised method can improve the classification accuracy.
Yangqiu Song, Changshui Zhang, Shiming Xiang
ICASSP (2)2
2007 Smoothness Maximization via Gradient Descents
abstract
The recent years have witnessed a surge of interest in graph based semi-supervised learning. However, despite its extensive research, there has been little work on graph construction. In this study, employing the idea of gradient descent, we propose a novel method called iterative smoothness maximization (ISM), to learn an optimal graph automatically for a semi-supervised learning task. The main procedure of ISM is to minimize the upper bound of semi-supervised classification error through an iterative gradient descent approach. We also prove the convergence of ISM theoretically, and finally experimental results on two real-world data sets are provided to demonstrate the effectiveness of ISM.
Bin Zhao 0004, Fei Wang 0001, Changshui Zhang
ICASSP (2)3
2007 Multilevel Belief Propagation for Fast Inference on Markov Random Fields
abstract
Graph-based inference plays an important role in many mining and learning tasks. Among all the solvers for this problem, belief propagation (BP) provides a general and efficient way to derive approximate solutions. However, for large scale graphs the computational cost of BP is still demanding. In this paper, we propose a multilevel algorithm to accelerate belief propagation on Markov Random Fields (MRF). First, we coarsen the original graph to get a smaller one. Then, BP is applied on the new graph to get a coarse result. Finally the coarse solution is efficiently refined back to derive the original solution. Unlike traditional multi- resolution approaches, our method features adaptive coarsening and efficient refinement. The above process can be recursively applied to reduce the computational cost remarkably. We theoretically justify the feasibility of our method on Gaussian MRFs, and empirically show that it is also effectual on discrete MRFs. The effectiveness of our method is verified in experiments on various inference tasks.
Liang Xiong, Fei Wang 0001, Changshui Zhang
ICDM3
2007 Multi-Pitch Estimation Based on Partial Event and Support Transfer
abstract
This paper proposes a method for the multi-pitch estimation of polyphonic music signals. Instead of on the frame level, the estimation is based on the partial event, which is defined like the note event in MIDI. All partial events in a piece of music are extracted dynamically in the process of the frame by frame short time Fourier transform (STFT). For each event, net support degree received from other events is calculated and the events with the highest support degrees are selected to be the fundamental frequency (FO) events. From another point of view, the support is transferred from higher frequency partial events to lower ones and finally concentrated on the FO events. This method can estimate the number of concurrent sounds, the onset and offset times of the notes. Experiments on both randomly mixed chord signals and synthesized ensemble music signals in "wav" format are conducted and the results are promising.
Zhiyao Duan, Dan Zhang 0007, Changshui Zhang, Zhenwei Shi 0001
ICME3
2007 A Fixed-Point Algorithm for Blind Separation of Temporally Correlated Sources
abstract
In this paper we develop a new method for blind separation of temporally correlated sources, possibly dependent signals from linear mixtures of them. The proposed algorithm is based on the mutual independency of the innovations of source signals instead of original signals. This algorithm takes into account both the temporal structure and the high-order statistics of source signals and in contrast to the most known blind separation algorithms only exploiting the second order statistics or the non-Gaussianity. In this framework, a fixed-point algorithm is introduced. The fixed-point algorithm is computationally very simple, converge fast, and does not need choose any learning step sizes. Extensive computer simulations with speech signals and images confirm the validity and high performance of the proposed algorithm.
Zhenwei Shi 0001, Dan Zhang 0007, Changshui Zhang
ICME3
2007 Generalized Additive Bayesian Network Classifiers
Changshui Zhang, Tao Wang 0003, Yimin Zhang 0002
IJCAI2
2007 Neighborhood MinMax Projections
Feiping Nie 0001, Shiming Xiang, Changshui Zhang
IJCAI3
2007 Combining stroke-based and selection-based relevance feedback for content-based image retrieval
abstract
We propose a flexible interaction mechanism for CBIR by enabling relevance feedback inside images through drawling strokes. User's interest is obtained from an easy-to-use user interface, and fused seamlessly with traditional feedback information in a semi-supervised learning framework. Retrieval performance is boosted due to more precise description of the query concept. Region segmentation is also improved based on the collected strokes, and further enhances the retrieval precision. We implement our system Flexible Image Search Tool (FIST) based on the ideas above. Experiments on two real world data sets demonstrate the effectiveness of our approach.
Jingyu Cui, Changshui Zhang
ACM Multimedia2
2007 Interactive Visual Object Extraction Based on Belief Propagation
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001
MMM (1)3
2007 Embedding New Data Points for Manifold Learning Via Coordinate Propagation
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang, Chunxia Zhang 0001
PAKDD4
2007 Fast Multilevel Transduction on Graphs
abstract
The recent years have witnessed a surge of interest in graph-based semi-supervised learning methods. The common denominator of these methods is that the data are represented by the nodes of a graph, the edges of which encode the pairwise similarities of the data. Despite the theoretical and empirical success, these methods have one major bottleneck which is the high computational complexity (since they usually require the computation of matrix inverse). In this paper, we propose a multilevel scheme for speeding up the traditional graph based semi-supervised learning methods. Unlike other accelerating approaches based on pure mathematical derivations, our method has explicit physical meanings with some graph intuitions. We also analyze the relationship of our method with multigrid methods, and provide a theoretical guarantee of the performance of our method. Finally the experimental results are presented to show the effectiveness of our method.
Fei Wang 0001, Changshui Zhang
SDM2
2007 Regularized clustering for documents
abstract
In recent years, document clustering has been receiving more and more attentions as an important and fundamental technique for unsupervised document organization, automatictopic extraction, and fast information retrieval or filtering. In this paper, we propose a novel method for clustering documents using regularization. Unlike traditional globally regularized clustering methods, our method first construct a local regularized linear label predictor for each document vector, and then combine all those local regularizers with a global smoothness regularizer. So we call our algorithm Clustering with Local and Global Regularization (CLGR). We will show that the cluster memberships of the documents can be achieved by eigenvalue decomposition of a sparse symmetric matrix, which can be efficiently solved by iterative methods. Finally our experimental evaluations on several datasets are presented to show the superiorities of CLGR over traditional document clustering methods.
Fei Wang 0001, Changshui Zhang, Tao Li 0001
SIGIR2
2007 Semi-blind source extraction for fetal electrocardiogram extraction by combining non-Gaussianity and time-correlation
Zhenwei Shi 0001, Changshui Zhang
Neurocomputing2
2007 Nonlinear innovation to blind source separation
Zhenwei Shi 0001, Changshui Zhang
Neurocomputing2
2007 Robust self-tuning semi-supervised learning
Fei Wang 0001, Changshui Zhang
Neurocomputing2
2007 Face recognition using spectral features
Fei Wang 0001, Jingdong Wang 0001, Changshui Zhang, James T. Kwok
Pattern Recognit.3
2007 An evolutionary system for near-regular texture synthesis
Fei Wu 0003, Changshui Zhang, Jingrui He
Pattern Recognit.2
2007 An experimental evaluation of ensemble methods for EEG signal classification
Shiliang Sun, Changshui Zhang, Dan Zhang 0007
Pattern Recognit. Lett.2
2007 The Selective Random Subspace Predictor for Traffic Flow Forecasting
abstract
Traffic flow forecasting is an important issue for the application of Intelligent Transportation Systems. Due to practical limitations, traffic flow data may be incomplete (partially missing or substantially contaminated by noises), which will aggravate the difficulties for traffic flow forecasting. In this paper, a new approach, termed the selective random subspace predictor (SRSP), is developed, which is capable of implementing traffic flow forecasting effectively whether incomplete data exist or not. It integrates the entire spatial and temporal traffic flow information in a transportation network to carry out traffic flow forecasting. To forecast the traffic flow at an object road link, the Pearson correlation coefficient is adopted to select some candidate input variables that compose the selective input space. Then, a number of subsets of the input variables in the selective input space are randomly selected to, respectively, serve as specific inputs for prediction. The multiple outputs are combined through a fusion methodology to make final decisions. Both theoretical analysis and experimental results demonstrate the effectiveness and robustness of the SRSP for traffic flow forecasting, whether for complete data or for incomplete data
Shiliang Sun, Changshui Zhang
IEEE Trans. Intell. Transp. Syst.2
2007 Blind Source Extraction Using Generalized Autocorrelations
abstract
This letter addresses blind (semiblind) source extraction (BSE) problem when a desired source signal has temporal structures, such as linear or nonlinear autocorrelations. Using the temporal characteristics of sources, we develop objective functions based on the generalized autocorrelations of primary sources. Maximizing the objective functions, we propose simple fixed-point source extraction algorithms. We give the stability analysis and prove convergence properties of the algorithms as the generalized autocorrelation function is linear or nonlinear. Especially, as the generalized autocorrelation function is linear, the algorithm has interesting character of "one-iteration" convergence under some conditions. Computer simulations and real-data application experiments show that the algorithms are appealing BSE methods for temporal signals of interest by capturing the linear or nonlinear autocorrelations of the desired sources.
Zhenwei Shi 0001, Changshui Zhang
IEEE Trans. Neural Networks2
2006 Texture Image Segmentation: An Interactive Framework Based on Adaptive Features and Transductive Learning
Shiming Xiang, Feiping Nie 0001, Changshui Zhang
ACCV (1)3
2006 Exemplar-Based Human Contour Tracking
Shiming Xiang, Feiping Nie 0001, Changshui Zhang
ACCV (1)3
2006 Contour Matching Based on Belief Propagation
Shiming Xiang, Feiping Nie 0001, Changshui Zhang
ACCV (2)3
2006 Semi-Supervised Classification Using Linear Neighborhood Propagation
abstract
In this paper, we address the general problem of learning from both labeled and unlabeled data. Based on the reasonable assumption that the label of each data can be linearly reconstructed from its neighbors’ labels, we develop a novel approach, called Linear Neighborhood Propagation (LNP), to learn the linear construction weights and predict the labels of the unlabeled data from the labeled data. The main procedure is composed of two steps: (1) it computes the local neighborhood geometry according to the image information by solving a constrained least squares problem, (2) the obtained geometry will be used to predict the labels of unlabeled data by solving a linearly constrained quadratic optimization problem. LNP has at least three characteristics. Firstly, it is much practical in that only one free parameter, the size of neighborhood, requires user-tuning. Secondly it can be easily extended to out-of-sample examples. Finally the labels achieved from LNP can be sufficiently smooth with respect to the intrinsic structure collectively revealed by the labeled and unlabeled points. Many promising experimental results, such as object recognition, data ranking and interactive image segmentation, demonstrate the effectiveness and efficiency of the proposed approach.
Fei Wang 0001, Changshui Zhang, Helen C. Shen, Jingdong Wang 0001
CVPR (1)2
2006 Spline Embedding for Nonlinear Dimensionality Reduction
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001
ECML3
2006 Label propagation through linear neighborhoods
abstract
A novel semi-supervised learning approach is proposed based on a linear neighborhood model, which assumes that each data point can be linearly reconstructed from its neighborhood. Our algorithm, named Linear Neighborhood Propagation (LNP), can propagate the labels from the labeled points to the whole dataset using these linear neighborhoods with sufficient smoothness. We also derive an easy way to extend LNP to out-of-sample data. Promising experimental results are presented for synthetic data, digit and text classification tasks.
Fei Wang 0001, Changshui Zhang
ICML2
2006 Machine Learning Way for Boosting Accuracy in Canonical Correlation Analysis based Frequency Recognition
abstract
Canonical Correlation Analysis (CCA) is used to frequency recognition of multichannel signals. The unknown signals are compared against known templates and their frequencies are recognized by simply comparing the biggest coefficients of their CCA coefficient vectors. This strategy is straightforward but may not give optimal results. To boost the accuracy of recognition we reformulate the approach in views of machine learning. In this paper, we propose a new strategy based on supervised learning. We also employ feature selection within this framework to adopt efficient coefficients which may not be the largest coefficients for the features vectors. The recognition method is validated by results with real world data.
Zhonglin Lin, Changshui Zhang, Xiaorong Gao
IJCNN2
2006 Estimating Fractal Intrinsic Dimension from the Neighborhood
Qutang Cai 0001, Changshui Zhang
ISNN (1)2
2006 A Random Walks Method for Text Classification
abstract
Practical text classification system should be able to utilize information from both expensive labelled documents and large volumes of cheap unlabelled documents. It should also easily deal with newly input samples. In this paper, we propose a random walks method for text classification, in which the classification problem is formulated as solving the absorption probabilities of Markov random walks on a weighted graph. Then the Laplacian operator for asymmetric graphs is derived and utilized for asymmetric transition matrix. We also develop an induction algorithm for the newly input documents based on the random walks method. Meanwhile, to make full use of text information, a difference measure for text data based on language model and KL-divergence is proposed, as well as a new smoothing technique for it. Finally an algorithm for elimination of ambiguous states is proposed to address the problem of noisy data. Experiments on two well-known data sets: WebKB and 20Newsgroup demonstrate the effectivity of the proposed random walks method.
Yunpeng Xu, Xing Yi, Changshui Zhang
SDM3
2006 Gaussian moments for noisy complexity pursuit
Zhenwei Shi 0001, Changshui Zhang
Neurocomputing2
2006 An optimal kernel feature extractor and its application to EEG signal classification
Shiliang Sun, Changshui Zhang
Neurocomputing2
2006 Classification of gene-expression data: The manifold-based metric learning way
Jianguo Lee, Changshui Zhang
Pattern Recognit.2
2006 Generalized Manifold-Ranking-Based Image Retrieval
abstract
In this paper, we propose a general transductive learning framework named generalized manifold-ranking-based image retrieval (gMRBIR) for image retrieval. Comparing with an existing transductive learning method named MRBIR [12], our method could work well whether or not the query image is in the database; thus, it is more applicable for real applications. Given a query image, gMRBIR first initializes a pseudo seed vector based on neighborhood relationship and then spread its scores via manifold ranking to all the unlabeled images in the database. Furthermore, in gMRBIR, we also make use of relevance feedback and active learning to refine the retrieval result so that it converges to the query concept as fast as possible. Systematic experiments on a general-purpose image database consisting of 5,000 Corel images demonstrate the superiority of gMRBIR over state-of-the-art techniques.
Jingrui He, Mingjing Li, HongJiang Zhang, Hanghang Tong, Changshui Zhang
IEEE Trans. Image Process.5
2006 A bayesian network approach to traffic flow forecasting
abstract
A new approach based on Bayesian networks for traffic flow forecasting is proposed. In this paper, traffic flows among adjacent road links in a transportation network are modeled as a Bayesian network. The joint probability distribution between the cause nodes (data utilized for forecasting) and the effect node (data to be forecasted) in a constructed Bayesian network is described as a Gaussian mixture model (GMM) whose parameters are estimated via the competitive expectation maximization (CEM) algorithm. Finally, traffic flow forecasting is performed under the criterion of minimum mean square error (mmse). The approach departs from many existing traffic flow forecasting models in that it explicitly includes information from adjacent road links to analyze the trends of the current link statistically. Furthermore, it also encompasses the issue of traffic flow forecasting when incomplete data exist. Comprehensive experiments on urban vehicular traffic flow data of Beijing and comparisons with several other methods show that the Bayesian network is a very promising and effective approach for traffic flow modeling and forecasting, both for complete data and incomplete data.
Shiliang Sun, Changshui Zhang, Guoqiang Yu
IEEE Trans. Intell. Transp. Syst.2
2005 A Unified Optimization Based Learning Method for Image Retrieval
abstract
In this paper, an optimization based learning method is proposed for image retrieval from graph model point of view. Firstly, image retrieval is formulated as a regularized optimization problem, which simultaneously considers the constraints from low-level feature, online relevance feedback and offline semantic information. Then, the global optimal solution is developed in both closed form and iterative form, providing that the latter converges to the former. The proposed method is unified in the senses that 1) it makes use of the information from various aspects in a global optimization manner so that the retrieval performance might be maximally improved; 2) it provides a natural way to support two typical query scenarios in image retrieval. The proposed method has a solid mathematical ground. Systematic experimental results on a general-purpose image database demonstrate that it achieves significant improvements over existing methods.
Hanghang Tong, Jingrui He, Mingjing Li, Wei-Ying Ma, Changshui Zhang, HongJiang Zhang
CVPR (2)5
2005 Learning On-Line Classification via Decorrelated LMS Algorithm: Application to Brain-Computer Interfaces
Shiliang Sun, Changshui Zhang
Discovery Science2
2005 Network Boosting for BCI Applications
Zhonglin Lin, Changshui Zhang
Discovery Science3
2005 Network Game and Boosting
Changshui Zhang
ECML2
2005 Traffic Flow Forecasting Using a Spatio-temporal Bayesian Network Predictor
Shiliang Sun, Changshui Zhang, Yi Zhang 0029
ICANN (2)2
2005 Boosting Web Image Search by Co-Ranking
abstract
To maximally improve the precision among top-ranked images returned by a Web image search engine without putting extra burden on the user, we propose in this paper a novel co-ranking framework which re-ranks the retrieved images to move the irrelevant ones to the tail of the list. The characteristics of the proposed framework can be summarized as follows: (1) making use of the decisions from multi-view of images to boost retrieval performance; (2) generalizing present multi-view algorithms which need labeled data for initialization to the unsupervised case so that no extra interaction is required. To implement the framework, we use one-class support vector machines to train the basic learner, and propose different schemes for combination. Experimental results demonstrate the effectiveness of the proposed framework.
Jingrui He, Changshui Zhang, Nanyuan Zhao, Hanghang Tong
ICASSP (2)2
2005 Assessing features for electroencephalographic signal categorization
abstract
The classification of electroencephalographic (EEG) signals is an important issue in the ongoing research of brain-computer interface (BCI) technology. One such BCI uses slow cortical potential measures to infer user intent from the original brain activity. Seven features based on the standard low-level signal properties are evaluated for their ability to classify brain activities, and thus make up for the scarcity of signal features for the current EEG signal categorization. In addition, a paradigm is proposed to select effective low-level features for EEG signal classification. Combining the features selected by the paradigm with the DC value of slow cortical potentials for categorization based on a Bayesian classifier, we obtained significant improvement on classification accuracy for data set Ia of the BCI competition 2003, which is a typical representative of one kind of BCI data.
Shiliang Sun, Changshui Zhang
ICASSP (5)2
2005 On the Scale-Free Intersection Graphs
Xin Yao 0003, Changshui Zhang, Yanda Li
ICCSA (2)2
2005 Separation of Voice and Music by Harmonic Structure Stability Analysis
abstract
Separation of voice and music is an interesting but difficult problem. It is useful for many other researches such as audio content analysis. In this paper, the difference between voice and music signals is carefully studied. It is proposed that the harmonic structure stability is the key difference between them. A separation algorithm based on this theory is proposed. The main idea is to learn the average harmonic structure of the music, and then separate signals by using it to distinguish voice and music harmonic structures. Experimental results show that the algorithm can separate mixed signals and obtains not only a very high signal-to-noise ratio (SNR) but also a rather good subjective audio quality
Yungang Zhang, Changshui Zhang
ICME2
2005 Learning probability density functions from marginal distributions with applications to Gaussian mixtures
abstract
Probability density function (PDF) estimation is a constantly important topic in the fields related to artificial intelligence and machine learning. This paper is dedicated to considering problems on the estimation of a density function simply from its marginal distributions. The possibility of the learning problem is first investigated and a uniqueness proposition involving a large family of distribution functions is proposed. The learning problem is then reformulated into an optimization task which is studied and applied to Gaussian mixture models (GMM) via the generalized expectation maximization procedure (GEM) and Monte Carlo method. Experimental results show that our approach for GMM, only using partial information of the coordinates of the samples, can obtain satisfactory performance, which in turn verifies the proposed reformulation and proposition.
Qutang Cai 0001, Changshui Zhang, Chunyi Peng 0001
IJCNN2
2005 New boosting methods of Gaussian processes for regression
abstract
Feedforward neural networks are popular tools for nonlinear regression and classification problems. Gaussian Process (GP) can be viewed as an RBF neural network which have infinite number of hidden neurons. On regression problems, they can predict both the mean value and the variance of the given sample. Boosting is one of the most important recent developments in machine learning. Classification problems have dominated research on boosting to date. On the other hand, the application of boosting of regression has received less investigation. In this paper, we develop two boosting methods of GPs for regression according to the characteristic of them. We compare the performance of our ensembles with other boosting algorithms and find that our methods are more stable and essentially have less over-fitting problems than the other methods.
Yangqiu Song, Changshui Zhang
IJCNN2
2005 Spectral feature analysis
abstract
We have seen a surge of interest in spectral-based methods and kernel-based methods for machine learning and data mining. Despite the significant research, these methods remain only loosely related. In this paper, we give theoretically an explicit relation between spectral clustering and weighted kernel principal component analysis (WKPCA). We show that spectral clustering is not only a method for data clustering, but also for feature extraction. We are then able to reinterpret the spectral clustering algorithm in terms of WKPCA and propose our spectral feature analysis (SFA) method. The spectral features extracted by SFA can capture the distinguishing information of data from different classes effectively. Finally some experimental results are presented to show the effectiveness of our method.
Fei Wang 0001, Jingdong Wang 0001, Changshui Zhang
IJCNN3
2005 Graph based multi-modality learning
abstract
To better understand the content of multimedia, a lot of research efforts have been made on how to learn from multi-modal feature. In this paper, it is studied from a graph point of view: each kind of feature from one modality is represented as one independent graph; and the learning task is formulated as inferring from the constraints in every graph as well as supervision information (if available). For semi-supervised learning, two different fusion schemes, namely linear form and sequential form, are proposed. For each scheme, it is derived from optimization point of view; and further justified from two sides: similarity propagation and Bayesian interpretation. By doing so, we reveal the regular optimization nature, transductive learning nature as well as prior fusion nature of the proposed schemes, respectively. Moreover, the proposed method can be easily extended to unsupervised learning, including clustering and embedding. Systematic experimental results validate the effectiveness of the proposed method.
Hanghang Tong, Jingrui He, Mingjing Li, Changshui Zhang, Wei-Ying Ma
ACM Multimedia4
2005 Learning No-Reference Quality Metric by Examples
abstract
In this paper, a novel learning based method is proposed for No-Reference image quality assessment. Instead of examining the exact prior knowledge for the given type of distortion and finding a suitable way to represent it, our method aims to directly get the quality metric by means of learning. At first, some training examples are prepared for both high-quality and low-quality classes; then a binary classifier is built on the training set; finally the quality metric of an un-labeled example is denoted by the extent to which it belongs to these two classes. Different schemes to acquire examples from a given image, to build the binary classifier and to model the quality metric are proposed and investigated. While most existing methods are tailored for some specific distortion type, the proposed method might provide a general solution for No-Reference image quality assessment. Experimental results on JPEG and JPEG2000 compressed images validate the effectiveness of the proposed method.
Hanghang Tong, Mingjing Li, HongJiang Zhang, Changshui Zhang, Jingrui He, Wei-Ying Ma
MMM4
2005 Separation of Music Signals by Harmonic Structure Modeling
abstract
Separation of music signals is an interesting but difficult problem. It is helpful for many other music researches such as audio content analysis. In this paper, a new music signal separation method is proposed, which is based on harmonic structure modeling. The main idea of harmonic structure modeling is that the harmonic structure of a music signal is stable, so a music signal can be represented by a harmonic structure model. Accordingly, a corresponding separation algorithm is proposed. The main idea is to learn a harmonic structure model for each music signal in the mixture, and then separate signals by using these models to distinguish harmonic structures of different signals. Experimental results show that the algorithm can separate signals and obtain not only a very high Signalto-Noise Ratio (SNR) but also a rather good subjective audio quality.
Yungang Zhang, Changshui Zhang
NIPS2
2005 Visual object recognition using probabilistic kernel subspace similarity
Jianguo Lee, Jingdong Wang 0001, Changshui Zhang, Zhaoqi Bian
Pattern Recognit.3
2005 Exploring the structure of supervised data by Discriminant Isometric Mapping
Shifeng Weng, Changshui Zhang, Zhonglin Lin
Pattern Recognit.2
2005 Active curve axis Gaussian mixture models
Baibo Zhang, Changshui Zhang, Xing Yi
Pattern Recognit.2
2004 Bayesian Network Methods for Traffic Flow Forecasting with Incomplete Data
Shiliang Sun, Changshui Zhang, Guoqiang Yu, Naijiang Lu
ECML2
2004 Mamdanian logic
abstract
In the framework of stratified fuzzy propositional logic F, the regular harmonious neighborhood structure determined by the Mamdanian approximate function Hgt is the typical and best paradigm of approximate reasoning neighborhood algorithms. The advantage of Mamdanian algorithm comes from the regular harmoniousness property of Hgt presented by us in this paper, which guarantees that when the input of the obtained new approximate knowledge is close enough to that of the standard knowledge, not only the approximate reasoning consequence of the new knowledge is adequately close to that of the standard knowledge, but also the new knowledge itself is close enough to the standard knowledge.
Yalin Zheng, Changshui Zhang, Xing Yi
FUZZ-IEEE2
2004 Switching ARIMA model based forecasting for traffic flow
abstract
Switching dynamic linear models are commonly used methods to describe change in an evolving time series, where the switching ARIMA (autoregressive integrated moving average) model is a special case. Short-term forecasting of traffic flows is an essential part of intelligent traffic systems (ITS). We apply the switching ARIMA model to a traffic flow series. We have observed that the conventional switching model is inappropriate to describe the pattern changing. Thus, the variable of duration is introduced and we use the sigmoid function to describe the influence of duration to the transition probability of the patterns. Based on the switching ARIMA model, a forecasting algorithm is presented. We apply the proposed model to real data obtained from UTC/SCOOT systems in Beijing's traffic management bureau. The experiments show that our proposed model is applicable and effective.
Guoqiang Yu, Changshui Zhang
ICASSP (2)2
2004 Symmetry feature in content-based image retrieval
abstract
In this paper, we first apply the theory of wallpaper groups to natural images and extract a novel feature to depict the symmetry property of natural images. The original proposed algorithm takes autocorrelation and correlation as a preprocessing step, which is very time-consuming. Through further analysis, we develop a set of schemes to accelerate this algorithm. Experimental results demonstrate that in performing content-based image retrieval, the proposed symmetry feature outperforms the wavelet feature, which is a widely accepted descriptor of texture, and water-filling feature. The accelerated version of the algorithm improves the processing speed by a large margin while it produces little degradation in retrieval performance.
Jingrui He, Mingjing Li, HongJiang Zhang, Changshui Zhang
ICIP4
2004 No-reference quality assessment for JPEG2000 compressed images
abstract
No-reference quality assessment is a relatively new topic and has been attracting more and more attention in recent years. Due to the limited understanding of the human vision system, most of the existing methods focus on measuring to what extent the image has been distorted. In this paper, by viewing all edge points in JPEG2000 compressed images as 'distorted' or 'un-distorted', we propose using principal component analysis (PCA) to extract the local feature of a given edge point, which indicates both blurring and ringing. We also propose using the probabilities of the given edge point being 'distorted' and 'un-distorted' to model the local distortion metric, which is straightforward and can be easily applied to any type of local feature. Experimental results demonstrate the effectiveness of our scheme.
Hanghang Tong, Mingjing Li, HongJiang Zhang, Changshui Zhang
ICIP4
2004 Blur detection for digital images using wavelet transform
abstract
With the prevalence of digital cameras, the number of digital images increases quickly, which raises the demand for image quality assessment in terms of blur. Based on the edge type and sharpness analysis, using the Harr wavelet transform, a new blur detection scheme is proposed in this paper, which can determine whether an image is blurred or not and to what extent an image is blurred. Experimental results demonstrate the effectiveness of the proposed scheme.
Hanghang Tong, Mingjing Li, HongJiang Zhang, Changshui Zhang
ICME4
2004 Multi-view EM algorithm and its application to color image segmentation
abstract
We propose a new algorithm, the multi-view expectation and maximization algorithm (multi-view EM), to deal with real-world learning problems where there are some natural split of features. Multi-view EM does feature split in the same manner as co-training and co-EM, two successful semi-supervised learning algorithms in text learning tasks, but it considers the multi-view learning problem in the framework of the EM algorithm. The multi-view EM algorithm has impressive advantages compared with co-training and co-EM: its convergence is theoretically guaranteed; and it can deal with multiple views instead of only two views. We utilize it for color image segmentation and discuss the phenomenon that different weights for the color view and coordinate view lead to different segmentation results.
Xing Yi, Changshui Zhang, Jingdong Wang 0001
ICME2
2004 Probabilistic tangent subspace: a unified view
abstract
Tangent Distance (TD) is one classical method for invariant pattern classification. However, conventional TD need pre-obtain tangent vectors, which is difficult except for image objects. This paper extends TD to more general pattern classification tasks. The basic assumption is that tangent vectors can be approximately represented by the pattern variations. We propose three probabilistic subspace models to encode the variations: the linear subspace, nonlinear subspace, and manifold subspace models. These three models are addressed in a unified view, namely Probabilistic Tangent Subspace (PTS). Experiments show that PTS can achieve promising classification performance in non-image data sets.
Jianguo Lee, Jingdong Wang 0001, Changshui Zhang, Zhaoqi Bian
ICML3
2004 Learning probabilistic kernel feature subspace with side-information for classification
abstract
Kernel PCA is an efficient method for nonlinear feature extraction. We address two issues in kernel PCA based feature extraction and classification. First, it extracts features without utilizing sample label information. Second, it does not provide a practical means to choose the dimensionality for principal subspace. In this paper, one kind of side-information is incorporated into kernel PCA to solve the first problem. And a complete probabilistic density function is estimated in kernel space so that the choice of dimensionality for principal subspace becomes less important. The proposed model is named probabilistic kernel feature subspace (PKFS). Experiments show that it achieves promising performance and outperforms many other algorithms in classification.
Jianguo Lee, Changshui Zhang, Zhaoqi Bian
IJCNN2
2004 Self-enhanced relevant component analysis with side-information and unlabeled data
abstract
Relevant component analysis (RCA) is a powerful tool for relevant linear feature extraction with side-information, a new focus in machine learning fields. But its only utilizing positive constraints weakens this algorithm's performance and robustness, especially when there are few positive constraints - a common case in practice. To overcome this drawback, in this paper we propose an extended algorithm named self-enhanced relevant component analysis (SERCA). Through a boosting procedure in the product space, it efficiently uses both the given side-information and unlabeled data. The experimental results on several data sets show that SERCA achieves an obvious improvement compared with RCA.
Fei Wu 0003, Yonglei Zhou, Changshui Zhang
IJCNN3
2004 Learning the Supervised NLDR Mapping for Classification
Zhonglin Lin, Shifeng Weng, Changshui Zhang, Naijiang Lu, Zhimin Xia
ISNN (1)3
2004 A New Adaptive Self-Organizing Map
Shifeng Weng, Fai Wong, Changshui Zhang
ISNN (1)3
2004 Short-Term Traffic Flow Forecasting Using Expanded Bayesian Network for Incomplete Data
Changshui Zhang, Shiliang Sun, Guoqiang Yu
ISNN (2)1
2004 Manifold-ranking based image retrieval
abstract
In this paper, we propose a novel transductive learning framework named manifold-ranking based image retrieval (MRBIR). Given a query image, MRBIR first makes use of a manifold ranking algorithm to explore the relationship among all the data points in the feature space, and then measures relevance between the query and all the images in the database accordingly, which is different from traditional similarity metrics based on pair-wise distance. In relevance feedback, if only positive examples are available, they are added to the query set to improve the retrieval result; if examples of both labels can be obtained, MRBIR discriminately spreads the ranking scores of positive and negative examples, considering the asymmetry between these two types of images. Furthermore, three active learning methods are incorporated into MRBIR, which select images in each round of relevance feedback according to different principles, aiming to maximally improve the ranking result. Experimental results on a general-purpose image database show that MRBIR attains a significant improvement over existing systems from all aspects.
Jingrui He, Mingjing Li, HongJiang Zhang, Hanghang Tong, Changshui Zhang
ACM Multimedia5
2004 Extended Nearest Feature Line Classifier
Yonglei Zhou, Changshui Zhang, Jingchun Wang
PRICAI2
2004 Reconstruction and analysis of multi-pose face images based on nonlinear dimensionality reduction
Changshui Zhang, Nanyuan Zhao, David Zhang 0001
Pattern Recognit.1
2004 Competitive EM algorithm for finite mixture models
Baibo Zhang, Changshui Zhang, Xing Yi
Pattern Recognit.2
2004 Distance metric learning by knowledge embedding
Yungang Zhang, Changshui Zhang, David Zhang 0001
Pattern Recognit.2
2004 Erratum to "Distance metric learning by knowledge embedding" [Pattern Recognition 37(1)161-163(2004)]
Yungang Zhang, Changshui Zhang, David Zhang 0001
Pattern Recognit.2
2003 Kernel Trick Embedded Gaussian Mixture Model
Jingdong Wang 0001, Jianguo Lee, Changshui Zhang
ALT3
2003 An Analytical Mapping for LLE and Its Application in Multi-Pose Face Synthesis
abstract
Locally Linear Embedding (LLE) is a nonlinear dimensionality reduction method proposed recently. It can reveal the intrinsic manifold of data, which can’t be provided by classical linear dimensionality reduc tion methods. The application of LLE, however, is limited because of its lack of a mapping between the observation and the low-dimensional output. In this paper, we propose a method to establish an analytical mapping for LLE and validate its efficiency with the application in multi-pose face synth esis. Furthermore, through learning the similarity for the same kind of pose change mode of different persons, we generalize our method to small set cases with methods of statistical learning theory. The experiments of multi-p ose face synthesis on small sets prove that our idea and method are correct.
Changshui Zhang, Zhongbao Kou
BMVC2
2003 Discovery of Relationships between Interests from Bulletin Board System by Dissimilarity Reconstruction
Zhongbao Kou, Ban Tao, Changshui Zhang
Discovery Science3
2003 Color Image Segmentation: Kernel Do the Feature Space
Jianguo Lee, Jingdong Wang 0001, Changshui Zhang
ECML3
2003 Clustering in Knowledge Embedded Space
Yungang Zhang, Changshui Zhang
ECML2
2003 Kernel GMM and its application to image binarization
abstract
Gaussian mixture model (GMM) is an efficient method for parametric clustering. However, traditional GMM can't perform clustering well on data set with complex structure such as images. In this paper, kernel trick, successfully used by SVM and kernel PCA, is introduced into EM algorithm for solving parameter estimation of GMM, which is so called kernel GMM (kGMM). The basic idea of kernel GMM is to apply kernel based GMM in feature space instead of in input data space. In order to avoid the curse of dimension, the proposed kGMM also embeds a step to automatically select discriminative features in feature space. kGMM is employed for the task of image binarization. Result shows that the proposed approach is feasible.
Jingdong Wang 0001, Jianguo Lee, Changshui Zhang
ICME3
2002 Hierarchical Shape Modeling for Automatic Face Localization
Ce Liu 0001, Harry Shum, Changshui Zhang
ECCV (2)3
2001 A Two-Step Approach to Hallucinating Faces: Global Parametric Model and Local Nonparametric Model
abstract
In this paper, we study face hallucination, or synthesizing a high-resolution face image from low-resolution input, with the help of a large collection of high-resolution face images. We develop a two-step statistical modeling approach that integrates both a global parametric model and a local nonparametric model. First, we derive a global linear model to learn the relationship between the high-resolution face images and their smoothed and down-sampled lower resolution ones. Second, the residual between an original high-resolution image and the reconstructed high-resolution image by a learned linear model is modeled by a patch-based nonparametric Markov network, to capture the high-frequency content of faces. By integrating both global and local models, we can generate photorealistic face images. Our approach is demonstrated by extensive experiments with high-quality hallucinated faces.
Ce Liu 0001, Harry Shum, Changshui Zhang
CVPR (1)3
2001 Automated Cartridge Identification for Firearm Authentication
abstract
In this paper, an algorithm of automated cartridge identification for firearm authentication is proposed. The ejector impression is used to calibrate the cartridge image. Features of the firing pin impression and the breach face impression are extracted using an active snake model and local orientation analysis, respectively. Different features are then integrated to make a final decision using a support vector machine. Experimental results illustrate the effectiveness of our algorithm.
Jie Zhou 0001, Le-ping Xin, Dashan Gao 0001, Changshui Zhang, David Zhang 0001
CVPR (1)4
2001 Palmprint recognition using crease
abstract
The palmprint has one special salient feature which is not salient in fingerprint. That is the crease. In a palmprint the creases are large in number and comparatively easy to extract. Creases are also approximately stable in a person's whole life, which qualifies them as features in palmprint recognition. We give creases an accurate definition which is fit for algorithm implementation. We devised a rather exquisite algorithm to extract all the creases in a palmprint whose success is mainly from a new different direction computing method and a thorough local analysis and a robust search algorithm. Based on the extracted creases, we devised a robust palmprint matching algorithm which is rotation and translation invariant. The crease extraction results and palmprint matching results show that the crease can be extracted successfully and crease-based palmprint matching is robust and accurate.
Changshui Zhang, Gang Rong
ICIP (3)2
2000 A Novel Algorithm for Rotated Human Face Detection
abstract
We present a framework to detect human faces rotated within the image plane. A novel orientation histogram of an image is constructed using local orientation analysis. From the view of an oriented pattern, a histogram of the human face shows symmetry with respect to the orientation, /spl beta/, of the principal axis, and a local peak also occurs at the orientation orthogonal to /spl beta/. We analyze the histogram to find the orientation of the face principal axis, with which conventional upright face detectors can be used to verify the supposed rotated face. Experimental results show that the algorithm is effective.
Xiao-guang Lv, Jie Zhou 0001, Changshui Zhang
CVPR3
2000 AR model for keystroker verification
abstract
The fast development of the Internet means that the security problems of information systems become more important. The traditional security mechanisms are not enough. Keystroke dynamics identification is a good means to strengthen the security of the computer and the network. The paper provides a novel method, the AR model to identify keystrokers.
Changshui Zhang, Yanhua Sun
SMC1
1998 A fully automated face recognition system under different conditions
abstract
We present a fully automated face recognition system (FAFRS) which contains two components: eyes detection and face recognition. Firstly, a hybrid neural method is proposed to localize human eyes. Then, a dual eigenspaces method (DEM) is developed to extract algebraic features from normalized facial images and perform the recognition task. Experimental results show that the proposed system is effective and robust under different conditions.
Changshui Zhang, Zhaoqi Bian
ICPR2