VLDB 2026 Research / reviewers in the wild / expert
Dijun Luo
dblp:47/2883
· DBLP profile ↗
38ranked-venue papers
24as first author
6since 2021 · last 2022
0009-0004-0566-1948ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 18 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14 · 10 first-authorGraphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 1 since 2021Computer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Reinforcement learning · 61% Representation and self-supervised learning · 22% Graph learning · 6% | |
| Databases, data mining, and information retrieval
13 papers |
Data mining · 79% Information retrieval · 12% Recommender systems · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 74% Environmental and earth informatics · 26% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 61% Routing and switching · 39% |
Topics — the 30 heaviest of 64, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
clustering |
0.9 | 8 | 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network Modules · ICDM 2016 Cluster Indicator Decomposition for Efficient Matrix Factorization · IJCAI 2011 Consensus spectral clustering in near-linear time · ICDE 2011 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative reinforcement learning |
0.6 | 1 | 2022 | Structured Cooperative Reinforcement Learning With Time-Varying Composite Action Space · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Reinforcement learning › markov decision process
factored action spaces |
0.6 | 1 | 2022 | Structured Cooperative Reinforcement Learning With Time-Varying Composite Action Space · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Reinforcement learning
markov decision process |
0.6 | 1 | 2022 | iGrow: A Smart Agriculture Solution to Autonomous Greenhouse Control · AAAI 2022 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.6 | 1 | 2022 | Structured Cooperative Reinforcement Learning With Time-Varying Composite Action Space · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
0.5 | 1 | 2021 | FOCAL: Efficient Fully-Offline Meta-Reinforcement Learning via Distance Metric Learning and Behavior Regularization · ICLR 2021 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.5 | 1 | 2021 | FOCAL: Efficient Fully-Offline Meta-Reinforcement Learning via Distance Metric Learning and Behavior Regularization · ICLR 2021 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.5 | 1 | 2021 | Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation · ICRA 2021 |
Machine learning › Reinforcement learning › meta-reinforcement learning
offline meta-reinforcement learning |
0.5 | 1 | 2021 | FOCAL: Efficient Fully-Offline Meta-Reinforcement Learning via Distance Metric Learning and Behavior Regularization · ICLR 2021 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.5 | 1 | 2021 | Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation · ICRA 2021 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.3 | 3 | 2011 | Discriminative high order SVD: Adaptive tensor subspace selection for image classification, clustering, and retrieval · ICCV 2011 Linear Discriminant Analysis: New Formulations and Overfit Analysis · AAAI 2011 Non-negative Laplacian Embedding · ICDM 2009 |
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization |
0.3 | 3 | 2011 | Low-order tensor decompositions for social tagging recommendation · WSDM 2011 Multi-Level Cluster Indicator Decompositions of Matrices and Tensors · AAAI 2011 Simultaneous tensor subspace selection and clustering: the equivalence of high order svd and k-means clustering · KDD 2008 |
Bioinformatics and computational biology › neuroscience › neuroinformatics
brain network analysis |
0.2 | 1 | 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network Modules · ICDM 2016 |
Bioinformatics and computational biology › neuroscience
neuroinformatics |
0.2 | 1 | 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network Modules · ICDM 2016 |
Data mining › clustering
graph clustering |
0.2 | 1 | 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network Modules · ICDM 2016 |
Data mining › structured data mining › graph mining › community detection
stochastic block model |
0.2 | 1 | 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network Modules · ICDM 2016 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › discriminant analysis
linear discriminant analysis |
0.2 | 2 | 2011 | Linear Discriminant Analysis: New Formulations and Overfit Analysis · AAAI 2011 Symmetric two dimensional linear discriminant analysis (2DLDA) · CVPR 2009 |
Data mining › clustering
spectral clustering |
0.2 | 2 | 2011 | Consensus spectral clustering in near-linear time · ICDE 2011 Non-negative Laplacian Embedding · ICDM 2009 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
manifold denoising |
0.2 | 1 | 2014 | Video Motion Segmentation Using New Adaptive Manifold Denoising Model · CVPR 2014 |
Computer vision › Video understanding and tracking
motion segmentation |
0.2 | 1 | 2014 | Video Motion Segmentation Using New Adaptive Manifold Denoising Model · CVPR 2014 |
Environmental and earth informatics › agriculture
smart agriculture |
0.2 | 1 | 2022 | iGrow: A Smart Agriculture Solution to Autonomous Greenhouse Control · AAAI 2022 |
Machine learning › Graph learning
graph construction |
0.1 | 1 | 2012 | Forging The Graphs: A Low Rank and Positive Semidefinite Graph Learning Approach · NIPS 2012 |
Parallel and multicore computing
parallel data mining |
0.1 | 1 | 2012 | Parallelization with Multiplicative Algorithms for Big Data Mining · ICDM 2012 |
Parallel and multicore computing › parallel computing › parallel machine learning
parallel learning algorithms |
0.1 | 1 | 2012 | Parallelization with Multiplicative Algorithms for Big Data Mining · ICDM 2012 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2011 | Discriminative high order SVD: Adaptive tensor subspace selection for image classification, clustering, and retrieval · ICCV 2011 |
Machine learning › Graph learning
network embedding |
0.1 | 1 | 2011 | Cauchy Graph Embedding · ICML 2011 |
Data mining
dimensionality reduction |
0.1 | 1 | 2011 | Robust Principal Component Analysis with Non-Greedy l1-Norm Maximization · IJCAI 2011 |
Data mining › clustering
ensemble clustering |
0.1 | 1 | 2011 | Consensus spectral clustering in near-linear time · ICDE 2011 |
Data mining › clustering
large-scale clustering |
0.1 | 1 | 2011 | Consensus spectral clustering in near-linear time · ICDE 2011 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2011 | Ball Ranking Machine for Content-Based Multimedia Retrieval · IJCAI 2011 |
Methods — techniques the papers use, named apart from their topics
neural network simulator · 1.1bi-level optimization · 1.1variational autoencoder · 0.6graph attention network · 0.6centralized critic decentralized actor · 0.6weighted exploitation · 0.5probabilistic multi-graph decomposition · 0.5optimization · 0.5optimistic exploration · 0.5model ensemble · 0.5distance metric learning · 0.5behavior regularization · 0.5low-rank approximation · 0.3singular value decomposition · 0.2non-negative matrix factorization · 0.2higher-order SVD · 0.2eckart-young bound · 0.2d-1 tensor factorization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | iGrow: A Smart Agriculture Solution to Autonomous Greenhouse ControlabstractAgriculture is the foundation of human civilization. However, the rapid increase of the global population poses a challenge on this cornerstone by demanding more food. Modern autonomous greenhouses, equipped with sensors and actuators, provide a promising solution to the problem by empowering precise control for high-efficient food production. However, the optimal control of autonomous greenhouses is challenging, requiring decision-making based on high-dimensional sensory data, and the scaling of production is limited by the scarcity of labor capable of handling this task. With the advances of artificial intelligence (AI), the internet of things (IoT), and cloud computing technologies, we are hopeful to provide a solution to automate and smarten greenhouse control to address the above challenges. In this paper, we propose a smart agriculture solution named iGrow, for autonomous greenhouse control (AGC): (1) for the first time, we formulate the AGC problem as a Markov decision process (MDP) optimization problem; (2) we design a neural network-based simulator incorporated with the incremental mechanism to simulate the complete planting process of an autonomous greenhouse, which provides a testbed for the optimization of control strategies; (3) we propose a closed-loop bi-level optimization algorithm, which can dynamically re-optimize the greenhouse control strategy with newly observed data during real-world production. We not only conduct simulation experiments but also deploy iGrow in real scenarios, and experimental results demonstrate the effectiveness and superiority of iGrow in autonomous greenhouse simulation and optimal control. Particularly, compelling results from the tomato pilot project in real autonomous greenhouses show that our solution significantly increases crop yield (+10.15%) and net profit (+92.70%) with statistical significance compared to planting experts. Our solution opens up a new avenue for greenhouse production. The code is available at https://github.com/holmescao/iGrow.git. Xiaoyan Cao, Yao Yao 0006, Lanqing Li, Wanpeng Zhang 0002, Zhicheng An, Zhong Zhang 0014, Li Xiao 0009, Shihui Guo, Meihong Wu, Dijun Luo |
AAAI | 11 |
| 2022 | Structured Cooperative Reinforcement Learning With Time-Varying Composite Action SpaceabstractIn recent years, reinforcement learning has achieved excellent results in low-dimensional static action spaces such as games and simple robotics. However, the action space is usually composite, composed of multiple sub-action with different functions, and time-varying for practical tasks. The existing sub-actions might be temporarily invalid due to the external environment, while unseen sub-actions can be added to the current system. To solve the robustness and transferability problems in time-varying composite action spaces, we propose a structured cooperative reinforcement learning algorithm based on the centralized critic and decentralized actor framework, called SCORE. We model the single-agent problem with composite action space as a fully cooperative partially observable stochastic game and further employ a graph attention network to capture the dependencies between heterogeneous sub-actions. To promote tighter cooperation between the decomposed heterogeneous agents, SCORE introduces a hierarchical variational autoencoder, which maps the heterogeneous sub-action space into a common latent action space. We also incorporate an implicit credit assignment structure into the SCORE to overcome the multi-agent credit assignment problem in the fully cooperative partially observable stochastic game. Performance experiments on the proof-of-concept task and precision agriculture task show that SCORE has significant advantages in robustness and transferability for time-varying composite action space. Wenhao Li 0001, Xiangfeng Wang 0001, Bo Jin 0003, Dijun Luo, Hongyuan Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Hierarchical Multiagent Reinforcement Learning for Allocating Guaranteed Display AdsabstractIn this article, we study the problem of guaranteed display ads (GDAs) allocation, which requires proactively allocate display ads to different impressions to fulfill their impression demands indicated in the contracts. Existing methods for this problem either assume the impressions that are static or solely consider a specific ad's benefits. Thus, it is hard to generalize to the industrial production scenario where the impressions are dynamical and large-scale, and the overall allocation optimality of all the considered GDAs is required. To bridge this gap, we formulate this problem as a sequential decision-making problem in the scope of multiagent reinforcement learning (MARL), by assigning an allocation agent to each ad and coordinating all the agents for allocating GDAs. The inputs are the states (e.g., the demands of the ad and the remaining time steps for displaying the ads) of each ad and the impressions at different time steps, and the outputs are the display ratios of each ad for each impression. Specifically, we propose a novel hierarchical MARL (HMARL) method that creates hierarchies over the agent policies to handle a large number of ads and the dynamics of impressions. HMARL contains: 1) a manager policy to navigate the agent to choose an appropriate subpolicy and 2) a set of subpolicies that let the agents perform diverse conditioning on their states. Extensive experiments on three real-world data sets from the Tencent advertising platform with tens of millions of records demonstrate significant improvements of HMARL over state-of-the-art approaches. Lu Wang 0029, Xinru Chen, Chengchang Li, Junzhou Huang, Weinan Zhang 0001, Wei Zhang 0056, Dijun Luo |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2021 | Robust Model-based Reinforcement Learning for Autonomous Greenhouse ControlabstractDue to the high efficiency and less weather dependency, autonomous greenhouses provide an ideal solution to meet the increasing demand for fresh food. However, managers are faced with some challenges in finding appropriate control strategies for crop growth, since the decision space of the greenhouse control problem is an astronomical number. Therefore, an intelligent closed-loop control framework is highly desired to generate an automatic control policy. As a powerful tool for optimal control, reinforcement learning (RL) algorithms can surpass human beings’ decision-making and can also be seamlessly integrated into the closed-loop control framework. However, in complex real-world scenarios such as agricultural automation control, where the interaction with the environment is time-consuming and expensive, the application of RL algorithms encounters two main challenges, i.e., sample efficiency and safety. Although model-based RL methods can greatly mitigate the efficiency problem of greenhouse control, the safety problem has not got too much attention. In this paper, we present a model-based robust RL framework for autonomous greenhouse control to meet the sample efficiency and safety challenges. Specifically, our framework introduces an ensemble of environment models to work as a simulator and assist in policy optimization, thereby addressing the low sample efficiency problem. As for the safety concern, we propose a sample dropout module to focus more on worst-case samples, which can help improve the adaptability of the greenhouse planting policy in extreme cases. Experimental results demonstrate that our approach can learn a more effective greenhouse planting policy with better robustness than existing methods. Wanpeng Zhang 0002, Xiaoyan Cao, Yao Yao 0006, Zhicheng An, Xi Xiao 0001, Dijun Luo |
ACML | 6 |
| 2021 | FOCAL: Efficient Fully-Offline Meta-Reinforcement Learning via Distance Metric Learning and Behavior Regularization
Lanqing Li, Dijun Luo |
ICLR | 3 |
| 2021 | Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and ExploitationabstractModel-based deep reinforcement learning has achieved success in various domains that require high sample efficiencies, such as Go and robotics. However, there are some remaining issues, such as planning efficient explorations to learn more accurate dynamic models, evaluating the uncertainty of the learned models, and more rational utilization of models. To mitigate these issues, we present MEEE, a model-ensemble method that consists of optimistic exploration and weighted exploitation. During exploration, unlike prior methods directly selecting the optimal action that maximizes the expected accumulative return, our agent first generates a set of action candidates and then seeks out the optimal action that takes both expected return and future observation novelty into account. During exploitation, different discounted weights are assigned to imagined transition tuples according to their model uncertainty respectively, which will prevent model predictive error propagation in agent training. Experiments on several challenging continuous control benchmark tasks demonstrated that our approach outperforms other model-free and model-based state-of-the-art methods, especially in sample complexity. Yao Yao 0006, Li Xiao 0006, Zhicheng An, Wanpeng Zhang 0002, Dijun Luo |
ICRA | 5 |
| 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network ModulesabstractMany recent scientific efforts have been devoted to constructing the human connectome using Diffusion Tensor Imaging (DTI) data for understanding large-scale brain networks that underlie higher-level cognition in human. However, suitable network analysis computational tools are still lacking in human brain connectivity research. To address this problem, we propose a novel probabilistic multi-graph decomposition model to identify consistent network modules from the brain connectivity networks of the studied subjects. At first, we propose a new probabilistic graph decomposition model to address the high computational complexity issue in existing stochastic block models. After that, we further extend our new probabilistic graph decomposition model for multiple networks/graphs to identify the shared modules cross multiple brain networks by simultaneously incorporating multiple networks and predicting the hidden block state variables. We also derive an efficient optimization algorithm to solve the proposed objective and estimate the model parameters. We validate our method by analyzing both the weighted fiber connectivity networks constructed from DTI images and the standard human face image clustering benchmark data sets. The promising empirical results demonstrate the superior performance of our proposed method. Dijun Luo, Zhouyuan Huo, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
ICDM | 1 |
| 2014 | Video Motion Segmentation Using New Adaptive Manifold Denoising ModelabstractVideo motion segmentation techniques automatically segment and track objects and regions from videos or image sequences as a primary processing step for many computer vision applications. We propose a novel motion segmentation approach for both rigid and non-rigid objects using adaptive manifold denoising. We first introduce an adaptive kernel space in which two feature trajectories are mapped into the same point if they belong to the same rigid object. After that, we employ an embedded manifold denoising approach with the adaptive kernel to segment the motion of rigid and non-rigid objects. The major observation is that the non-rigid objects often lie on a smooth manifold with deviations which can be removed by manifold denoising. We also show that performing manifold denoising on the kernel space is equivalent to doing so on its range space, which theoretically justifies the embedded manifold denoising on the adaptive kernel space. Experimental results indicate that our algorithm, named Adaptive Manifold Denoising (AMD), is suitable for both rigid and non-rigid motion segmentation. Our algorithm works well in many cases where several state-of-the-art algorithms fail. Dijun Luo, Heng Huang 0001 |
CVPR | 1 |
| 2014 | Building Maximum Lifetime Shortest Path Data Aggregation Trees in Wireless Sensor NetworksabstractIn wireless sensor networks, the spanning tree is usually used as a routing structure to collect data. In some situations, nodes do in-network aggregation to reduce transmissions, save energy, and maximize network lifetime. Because of the restricted energy of sensor nodes, how to build an aggregation tree of maximum lifetime is an important issue. It has been proved to be NP-complete in previous works. As shortest path spanning trees intuitively have short delay, it is imperative to find an energy-efficient shortest path tree for time-critical applications. In this article, we first study the problem of building maximum lifetime shortest path aggregation trees in wireless sensor networks. We show that when restricted to shortest path trees, building maximum lifetime aggregation trees can be solved in polynomial time. We present a centralized algorithm and design a distributed protocol for building such trees. Simulation results show that our approaches greatly improve the lifetime of the network and are very effective compared to other solutions. We extend our discussion to networks without aggregation and present interesting results. Mengfan Shan, Guihai Chen, Dijun Luo, Xiaojun Zhu 0001, Xiaobing Wu |
ACM Trans. Sens. Networks | 3 |
| 2013 | Toward structural sparsity: an explicit ℓ2/ℓ0 approach
Dijun Luo, Chris Ding, Heng Huang 0001 |
Knowl. Inf. Syst. | 1 |
| 2012 | Combining Knowledge and Data Driven Insights for Identifying Risk Factors using Electronic Health Records
Jimeng Sun 0001, Jianying Hu, Dijun Luo, Marianthi Markatou, Fei Wang 0001, Shahram Ebadollahi, Zahra Daar, Walter F. Stewart |
AMIA | 3 |
| 2012 | Parallelization with Multiplicative Algorithms for Big Data MiningabstractWe propose a nontrivial strategy to parallelize a series of data mining and machine learning problems, including 1-class and 2-class support vector machines, nonnegative least square problems, and $\ell_1$ regularized regression (LASSO) problems. Our strategy fortunately leads to extremely simple multiplicative algorithms which can be straightforwardly implemented in parallel computational environments, such as Map Reduce, or CUDA. We provide rigorous analysis of the correctness and convergence of the algorithm. We demonstrate the scalability and accuracy of our algorithms in comparison with other current leading algorithms. Dijun Luo, Chris Ding, Heng Huang 0001 |
ICDM | 1 |
| 2012 | Forging The Graphs: A Low Rank and Positive Semidefinite Graph Learning ApproachabstractIn many graph-based machine learning and data mining approaches, the quality of the graph is critical. However, in real-world applications, especially in semi-supervised learning and unsupervised learning, the evaluation of the quality of a graph is often expensive and sometimes even impossible, due the cost or the unavailability of ground truth. In this paper, we proposed a robust approach with convex optimization to ``forge'' a graph: with an input of a graph, to learn a graph with higher quality. Our major concern is that an ideal graph shall satisfy all the following constraints: non-negative, symmetric, low rank, and positive semidefinite. We develop a graph learning algorithm by solving a convex optimization problem and further develop an efficient optimization to obtain global optimal solutions with theoretical guarantees. With only one non-sensitive parameter, our method is shown by experimental results to be robust and achieve higher accuracy in semi-supervised learning and clustering under various settings. As a preprocessing of graphs, our method has a wide range of potential applications machine learning and data mining. Dijun Luo, Chris Ding, Heng Huang 0001, Feiping Nie 0001 |
NIPS | 1 |
| 2012 | SOR: Scalable Orthogonal Regression for Low-Redundancy Feature Selection and its Healthcare ApplicationsabstractAs more clinical information with increasing diversity become available for analysis, a large number of features can be constructed and leveraged for predictive modeling. Feature selection is a classic analytic component that faces new challenges due to the new applications: How to handle a diverse set of high dimensional features? How to select features with high predictive power, but low redundant information? How to design methods that can select globally optimal features with theoretical guarantee? How to incorporate and extend existing knowledge driven approach? In this paper, we present Scalable Orthogonal Regression (SOR), an optimization-based feature selection method with the following novelties: 1) Scalability: SOR achieves nearly linear scale-up with respect to the number of input features and the number of samples; 2) Optimality: SOR is formulated as an alternative convex optimization problem with theoretical convergence and global optimality guarantee; 3) Low-redundancy: thanks to the orthogonality objective, SOR is designed specifically to select less redundant features without sacrificing quality; 4) Extendability: SOR can enhance an existing set of preselected features by adding additional features that complement the existing feature set but still with strong predictive power. We present evaluation results showing that SOR consistently outperforms state of the art feature selection methods in a range of quality metrics on several real world data sets. We demonstrate a case study of a large-scale clinical application for predicting early onset of Heart Failure (HF) using real Electronic Health Records (EHRs) data of over 10K patients for over 7 years. Leveraging SOR, we are able to construct accurate and robust predictive models and derive potential clinical insights. Dijun Luo, Fei Wang 0001, Jimeng Sun 0001, Marianthi Markatou, Jianying Hu, Shahram Ebadollahi |
SDM | 1 |
| 2011 | Linear Discriminant Analysis: New Formulations and Overfit AnalysisabstractIn this paper, we will present a unified view for LDA. We will (1) emphasize that standard LDA solutions are not unique, (2) propose several new LDA formulations: St-orthonormal LDA, Sw-orthonormal LDA and orthogonal LDA which have unique solutions, and (3) show that with St-orthonormal LDA and Sw-orthonormal LDA formulations, solutions to all four major LDA objective functions are identical. Furthermore, we perform an indepth analysis to show that the LDA sometimes performs poorly due to over-fitting, i.e., it picks up PCA dimensions with small eigenvalues. From this analysis, we propose a stable LDA which uses PCA first to reduce to a small PCA subspace and do LDA in the subspace. Dijun Luo, Chris Ding, Heng Huang 0001 |
AAAI | 1 |
| 2011 | Multi-Level Cluster Indicator Decompositions of Matrices and TensorsabstractA main challenging problem for many machine learning and data mining applications is that the amount of data and features are very large, so that low-rank approximations of original data are often required for efficient computation. We propose new multi-level clustering based low-rank matrix approximations which are comparable and even more compact than Singular Value Decomposition (SVD). We utilize the cluster indicators of data clustering results to form the subspaces, hence our decomposition results are more interpretable. We further generalize our clustering based matrix decompositions to tensor decompositions that are useful in high-order data analysis. We also provide an upper bound for the approximation error of our tensor decomposition algorithm. In all experimental results, our methods significantly outperform traditional decomposition methods such as SVD and high-order SVD. Dijun Luo, Chris Ding, Heng Huang 0001 |
AAAI | 1 |
| 2011 | Discriminative high order SVD: Adaptive tensor subspace selection for image classification, clustering, and retrievalabstractTensor based dimensionality reduction has recently attracted attention from computer vision and pattern recognition communities for both feature extraction and data compression. As an unsupervised method, High-Order Singular Value Decomposition (HOSVD) searches for low-rank subspaces such that the low-rank approximation error is minimized. In this paper, we propose a new unsupervised high-order tensor decomposition approach which employs the strength of discriminative analysis and K-means clustering to adaptively select subspaces that improve the clustering, classification, and retrieval capabilities of HOSVD. We provide both theoretical analysis to guarantee that our new method generates more discriminative subspaces and empirical studies on several public computer vision data sets to show the consistent improvement of our method over existing methods. Dijun Luo, Heng Huang 0001, Chris Ding |
ICCV | 1 |
| 2011 | Consensus spectral clustering in near-linear timeabstractThis paper addresses the scalability issue in spectral analysis which has been widely used in data management applications. Spectral analysis techniques enjoy powerful clustering capability while suffer from high computational complexity. In most of previous research, the bottleneck of computational complexity of spectral analysis stems from the construction of pairwise similarity matrix among objects, which costs at least O(n2) where n is the number of the data points. In this paper, we propose a novel estimator of the similarity matrix using K-means accumulative consensus matrix which is intrinsically sparse. The computational cost of the accumulative consensus matrix is O(nlogn). We further develop a Non-negative Matrix Factorization approach to derive clustering assignment. The overall complexity of our approach remains O(nlogn). In order to validate our method, we (1) theoretically show the local preserving and convergent property of the similarity estimator, (2) validate it by a large number of real world datasets and compare the results to other state-of-the-art spectral analysis, and (3) apply it to large-scale data clustering problems. Results show that our approach uses much less computational time than other state-of-the-art clustering methods, meanwhile provides comparable clustering qualities. We also successfully apply our approach to a 5-million dataset on a single machine using reasonable time. Our techniques open a new direction for high-quality large-scale data analysis. Dijun Luo, Chris Ding, Heng Huang 0001, Feiping Nie 0001 |
ICDE | 1 |
| 2011 | Cauchy Graph Embedding
Dijun Luo, Chris Ding, Feiping Nie 0001, Heng Huang 0001 |
ICML | 1 |
| 2011 | Cluster Indicator Decomposition for Efficient Matrix Factorization
Dijun Luo, Chris Ding, Heng Huang 0001 |
IJCAI | 1 |
| 2011 | Ball Ranking Machine for Content-Based Multimedia RetrievalabstractIn this paper, we propose the new Ball Ranking Machines (BRMs) to address the supervised ranking problems. In previous work, supervised ranking methods have been successfully applied in various information retrieval tasks. Among these methodologies, the Ranking Support Vector Machines (Rank SVMs) are well investigated. However, one major fact limiting their applications is that Ranking SVMs need optimize a margin-based objective function over all possible document pairs within all queries on the training set. In consequence, Ranking SVMs need select a large number of support vectors among a huge number of support vector candidates. This paper introduces a new model of of Ranking SVMs and develops an efficient approximation algorithm, which decreases the training time and generates much fewer support vectors. Empirical studies on synthetic data and content-based image/video retrieval data show that our method is comparable to Ranking SVMs in accuracy, but use much fewer ranking support vectors and significantly less training time. Dijun Luo, Heng Huang 0001 |
IJCAI | 1 |
| 2011 | Robust Principal Component Analysis with Non-Greedy l1-Norm Maximization
Feiping Nie 0001, Heng Huang 0001, Chris Ding, Dijun Luo, Hua Wang 0007 |
IJCAI | 4 |
| 2011 | Maximizing lifetime for the shortest path aggregation tree in wireless sensor networksabstractIn many applications of wireless sensor networks, a sensor node senses the environment to get data and delivers them to the sink via a single hop or multi-hop path. Many systems use a tree rooted at the sink as the underlying routing structure. Since the sensor node is energy constrained, how to construct a good tree to prolong the lifetime of the network is an important problem. We consider this problem under the scenario where nodes have different initial energy, and they can do in-network aggregation. In previous works, it has been proved that finding a maximum lifetime tree from all feasible spanning trees is NP-complete. Since delay is also an important element in time-critical applications, and shortest path trees intuitively have short delay, it is imperative to find a shortest path tree with long lifetime. This paper studies the problem of maximizing the lifetime of data aggregation trees, which are limited to shortest path trees. We find that when it is restricted to shortest path trees, the original problem is in P. We transform the problem into a general version of semi-matching problem, and show that the problem can be solved by min-cost max-flow approach in polynomial time. Also we design a distributed solution. Simulation results show that our approach greatly improves the lifetime of the network and is more competitive when it is applied in a dense network. Dijun Luo, Xiaojun Zhu 0001, Xiaobing Wu, Guihai Chen |
INFOCOM | 1 |
| 2011 | Are Tensor Decomposition Solutions Unique? On the Global Convergence HOSVD and ParaFac Algorithms
Dijun Luo, Chris Ding, Heng Huang 0001 |
PAKDD (1) | 1 |
| 2011 | Graph Evolution via Social Diffusion Processes
Dijun Luo, Chris Ding, Heng Huang 0001 |
ECML/PKDD (2) | 1 |
| 2011 | Multi-Subspace Representation and Discovery
Dijun Luo, Feiping Nie 0001, Chris Ding, Heng Huang 0001 |
ECML/PKDD (2) | 1 |
| 2011 | Low-order tensor decompositions for social tagging recommendationabstractSocial tagging recommendation is an urgent and useful enabling technology for Web 2.0. In this paper, we present a systematic study of low-order tensor decomposition approach that are specifically targeted at the very sparse data problem in tagging recommendation problem. Low-order polynomials have low functional complexity, are uniquely capable of enhancing statistics and also avoids over-fitting than traditional tensor decompositions such as Tucker and Parafac decompositions. We perform extensive experiments on several datasets and compared with 6 existing methods. Experimental results demonstrate that our approach outperforms existing approaches. Yuanzhe Cai, Dijun Luo, Chris Ding, Sharma Chakravarthy |
WSDM | 3 |
| 2010 | Towards Structural Sparsity: An Explicit l2/l0 ApproachabstractIn many cases of machine learning or data mining applications, we are not only aimed to establish accurate black box predictors, we are also interested in discovering predictive patterns in data which enhance our interpretation and understanding of underlying physical, biological and other natural processes. Sparse representation is one of the focuses in this direction. More recently, structural sparsity has attracted increasing attentions. The structural sparsity is often achieved by imposing ℓ2/ℓ1norms. In this paper, we present the explicit ℓ2/ℓ0norm to directly achieve structural sparsity. To tackle the problem of intractable ℓ2/ℓ0optimization, we develop a general Lipschitz auxiliary function which leads to simple iterative algorithms. In each iteration, optimal solution is achieved for the induced sub-problem and a guarantee of convergence is provided. Further more, the local convergent rate is also theoretically bounded. We test our optimization techniques in the multi-task feature learning problem. Experimental results suggest that our approaches outperform other approaches in both synthetic and real world data sets. Dijun Luo, Chris Ding, Heng Huang 0001 |
ICDM | 1 |
| 2010 | Improved MinMax Cut Graph Clustering with Nonnegative Relaxation
Feiping Nie 0001, Chris Ding, Dijun Luo, Heng Huang 0001 |
ECML/PKDD (2) | 3 |
| 2010 | On the eigenvectors of p-Laplacian
Dijun Luo, Heng Huang 0001, Chris Ding, Feiping Nie 0001 |
Mach. Learn. | 1 |
| 2009 | Symmetric two dimensional linear discriminant analysis (2DLDA)abstractLinear discriminant analysis (LDA) has been successfully applied into computer vision and pattern recognition for effective feature extraction. High-dimensional objects such as images are usually transform as 1D vectors before the LDA transformation. Recently, two-dimension LDA (2DLDA) methods have been proposed which reduced the dimensionality of images without transforming the matrices into vectors. However, the objective function for 2DLDA remains an unresolved problem. In this paper, we (1) propose a symmetric LDA formulation which resolves the ambiguity problem, and (2) propose an effective algorithm to solve the symmetric 2DLDA objective. Experiments on UMIST, CMU PIE, and YaleB images databases show that our approach outperforms the other 2DLDA methods in terms of both classification accuracy and objective function results. Dijun Luo, Chris Ding, Heng Huang 0001 |
CVPR | 1 |
| 2009 | Non-negative Laplacian EmbeddingabstractLaplacian embedding provides a low dimensional representation for a matrix of pairwise similarity data using the eigenvectors of the Laplacian matrix. The true power of Laplacian embedding is that it provides an approximation of the ratio cut clustering. However, ratio cut clustering requires the solution to be nonnegative. In this paper, we propose a new approach, nonnegative Laplacian embedding, which approximates ratio cut clustering in a more direct way than traditional approaches. From the solution of our approach, clustering structures can be read off directly. We also propose an efficient algorithm to optimize the objective function utilized in our approach. Empirical studies on many real world datasets show that our approach leads to more accurate ratio cut solution and improves clustering accuracy at the same time. Dijun Luo, Chris Ding, Heng Huang 0001, Tao Li 0001 |
ICDM | 1 |
| 2009 | Link prediction of multimedia social network via unsupervised face recognitionabstractWe propose a new challenge for predicting links of social networks by unsupervised face recognition on photo albums. We solve the task by formulating it into Kernel Set Discovery problem. We enhance Affinity Propagation algorithm to tackle the problem with more constraints. More specifically, the face cannot appear more than once in the same photo and we impose constraints such that detected face images in the same photograph are never clustered into the same person. We construct a synthetic dataset based on AT\&T image benchmark for empirical validation. Moreover, we validate our algorithms by a real world application which contains a real friend relation on the Web 2.0 social network system. Results indicate our Constraint Affinity Propagation method is suitable to unsupervisedly predict links of social network. Dijun Luo, Heng Huang 0001 |
ACM Multimedia | 1 |
| 2008 | Tensor reduction error analysis - Applications to video compression and classificationabstractTensor based dimensionality reduction has recently been extensively studied for computer vision applications. To our knowledge, however, there exist no rigorous error analysis on these methods. Here we provide the first error analysis of these methods and provide error bound results similar to Eckart-Young Theorem which plays critical role in the development and application of singular value decomposition (SVD). Beside performance guarantee, these error bounds are useful for subspace size determination according to the required video/image reconstruction error. Furthermore, video surveillance/retrieval, 3D/4D medical image analysis, and other computer vision applications require particular reduction in spatio-temporal space, but not along data index dimension. This motivates a D-1 tensor reduction. Standard method such as high order SVD (HOSVD) compress data in all index dimensions and thus can not perform the classification and pattern recognition tasks. We provide algorithm and error bound analysis of the D-1 factorization for spatio-temporal data dimensionality. Experiments on video sequences demonstrate our approach outperforms the previous dimensionality deduction methods for spatio temporal data. Chris Ding, Heng Huang 0001, Dijun Luo |
CVPR | 3 |
| 2008 | Simultaneous tensor subspace selection and clustering: the equivalence of high order svd and k-means clusteringabstractSingular Value Decomposition (SVD)/Principal Component Analysis (PCA) have played a vital role in finding patterns from many datasets. Recently tensor factorization has been used for data mining and pattern recognition in high index/order data. High Order SVD (HOSVD) is a commonly used tensor factorization method and has recently been used in numerous applications like graphs, videos, social networks, etc. Heng Huang 0001, Chris Ding, Dijun Luo, Tao Li 0001 |
KDD | 3 |
| 2008 | Posterior probabilistic clustering using NMFabstractWe introduce the posterior probabilistic clustering (PPC), which provides a rigorous posterior probability interpretation for Nonnegative Matrix Factorization (NMF) and removes the uncertainty in clustering assignment. Furthermore, PPC is closely related to probabilistic latent semantic indexing (PLSI). Chris Ding, Tao Li 0001, Dijun Luo, Wei Peng 0001 |
SIGIR | 3 |
| 2008 | An improved error-correcting output coding framework with kernel-based decoding
Dijun Luo, Rong Xiong |
Neurocomputing | 1 |
| 2006 | Distance Function Learning in Error-Correcting Output Coding Framework
Dijun Luo, Rong Xiong |
ICONIP (2) | 1 |