Li Wang 0033

dblp:58/6810-33 · DBLP profile ↗
← Back
52ranked-venue papers
14as first author
17since 2021 · last 2026
0000-0003-2658-4262ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Transductive t-SNE (tt-SNE) for classification and data visualization
Joseph Balderas, Li Wang 0033, Andrzej Korzeniowski, Ren-Cang Li
Pattern Recognit. Lett.2
2023 Human Identification at a Distance: Challenges, Methods and Results on HID 2023
abstract
Human Identification at a Distance (HID) is an important research area due to its importance (especially in biometrics) and inherent challenges within this domain. To mitigate some of the constraints, we have introduced the HID challenge. This paper presents an overview of the 4th International Competition on Human Identification at a Distance (HID 2023), which serves as a benchmark for evaluating various methods in the field of human identification at a distance. We have introduced a new dataset, SUSTech-Competition, engulfing a cross-domain challenge. This dataset has 859 subjects, having various variations of clothing, carrying conditions, occlusions, and view angles. With a substantial participation of 254 registered teams, HID 2023 has attracted considerable attention and yielded highly encouraging results. Notably, the top-performing teams achieved significantly good accuracies. In this paper, we provide an introduction to the competition, encompassing the dataset, experimental settings, and competition organization, as well as an analysis of the results obtained by the top teams. Additionally, we delve into the methodologies employed by these leading teams. The progress demonstrated in this competition offers an optimistic outlook on the advancements in gait recognition, highlighting its potential for robust real applications.
Shiqi Yu 0001, Chenye Wang, Li Wang 0033, Qing Li 0015, Runsheng Wang, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Md. Atiqur Rahman Ahad
IJCB4
2023 ImpDet: Exploring Implicit Fields for 3D Object Detection
abstract
Conventional 3D object detection approaches concentrate on bounding boxes representation learning with several parameters, i.e., localization, dimension, and orientation. Despite its popularity and universality, such a straightforward paradigm is sensitive to slight numerical deviations, especially in localization. By exploiting the property that point clouds are naturally captured on the surface of objects along with accurate location and intensity information, we introduce a new perspective that views bounding box regression as an implicit function. This leads to our proposed framework, termed Implicit Detection or ImpDet, which leverages implicit field learning for 3D object detection. Our ImpDet assigns specific values to points in different local 3D spaces, thereby high-quality boundaries can be generated by classifying points inside or outside the boundary. To solve the problem of sparsity on the object surface, we further present a simple yet efficient virtual sampling strategy to not only fill the empty region, but also learn rich semantic features to help refine the boundaries. Extensive experimental results on KITTI and Waymo benchmarks demonstrate the effectiveness and robustness of unifying implicit fields into object detection.
Xuelin Qian, Li Wang 0033, Yi Zhu 0001, Li Zhang 0040, Yanwei Fu 0001, Xiangyang Xue 0001
WACV2
2023 Multiview Orthonormalized Partial Least Squares: Regularizations and Deep Extensions
abstract
In this article, we establish a family of subspace-based learning methods for multiview learning using least squares as the fundamental basis. Specifically, we propose a novel unified multiview learning framework called multiview orthonormalized partial least squares (MvOPLSs) to learn a classifier over a common latent space shared by all views. The regularization technique is further leveraged to unleash the power of the proposed framework by providing three types of regularizers on its basic ingredients, including model parameters, decision values, and latent projected points. With a set of regularizers derived from various priors, we not only recast most existing multiview learning methods into the proposed framework with properly chosen regularizers but also propose two novel models. To further improve the performance of the proposed framework, we propose to learn nonlinear transformations parameterized by deep networks. Extensive experiments are conducted on multiview datasets in terms of both feature extraction and cross-modal retrieval. Results show that the subspace-based learning for a common latent space is effective and its nonlinear extension can further boost performance, and more importantly, one of two proposed methods with nonlinear extension can achieve better results than all compared methods.
Li Wang 0033, Ren-Cang Li, Wen-Wei Lin
IEEE Trans. Neural Networks Learn. Syst.1
2023 Density-Based Distance Preserving Graph: Theoretical and Practical Analyses
abstract
This brief aims to provide theoretical guarantee and practical guidance on constructing a type of graphs from input data via distance preserving criterion. Unlike the graphs constructed by other methods, the targeted graphs are hidden through estimating a density function of latent variables such that the pairwise distances in both the input space and the latent space are retained, and they have been successfully applied to various learning scenarios. However, previous work heuristically treated the multipliers in the dual as the graph weights, so the interpretation of this graph from a theoretical perspective is still missing. In this brief, we fill up this gap by presenting a detailed interpretation based on optimality conditions and their connections to neighborhood graphs. We further provide a systematic way to set up proper hyperparameters to prevent trivial graphs and achieve varied levels of sparsity. Three extensions are explored to leverage different measure functions, refine/reweigh an initial graph, and reduce computation cost for medium-sized graph. Extensive experiments on both synthetic and real datasets were conducted and experimental results verify our theoretical findings and the showcase of the studied graph in semisupervised learning provides competitive results to those of compared methods with their best graph.
Li Wang 0033, Haian Yin, Jin Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2022 Exploring Latent Sparse Graph for Large-Scale Semi-supervised Learning
Li Wang 0033, Raymond Chan 0001, Tieyong Zeng
ECML/PKDD (4)2
2022 Orthogonal multi-view analysis by successive approximations via eigenvectors
Li Wang 0033, Lei-Hong Zhang, Chungen Shen, Ren-Cang Li
Neurocomputing1
2022 Predicting brain structural network using functional connectivity
Lu Zhang 0050, Li Wang 0033, Dajiang Zhu
Medical Image Anal.2
2022 A Self-Consistent-Field Iteration for Orthogonal Canonical Correlation Analysis
abstract
We propose an efficient algorithm for solving orthogonal canonical correlation analysis (OCCA) in the form of trace-fractional structure and orthogonal linear projections. Even though orthogonality has been widely used and proved to be a useful criterion for visualization, pattern recognition and feature extraction, existing methods for solving OCCA problem are either numerically unstable by relying on a deflation scheme, or less efficient by directly using generic optimization methods. In this paper, we propose an alternating numerical scheme whose core is the sub-maximization problem in the trace-fractional form with an orthogonality constraint. A customized self-consistent-field (SCF) iteration for this sub-maximization problem is devised. It is proved that the SCF iteration is globally convergent to a KKT point and that the alternating numerical scheme always converges. We further formulate a new trace-fractional maximization problem for orthogonal multiset CCA and propose an efficient algorithm with an either Jacobi-style or Gauss-Seidel-style updating scheme based on the SCF iteration. Extensive experiments are conducted to evaluate the proposed algorithms against existing methods, including real-world applications of multi-label classification and multi-view feature extraction. Experimental results show that our methods not only perform competitively to or better than the existing methods but also are more efficient.
Lei-Hong Zhang, Li Wang 0033, Zhaojun Bai, Ren-Cang Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 A Scalable Algorithm for Large-Scale Unsupervised Multi-View Partial Least Squares
abstract
We present an unsupervised multi-view partial least squares (PLS) by learning a common latent space from given multi-view data. Although PLS is a frequently used technique for analyzing relationships between two datasets, its extension to more than two views in unsupervised setting is seldom studied. In this article, we fill up the gap, and our model bears similarity to the extension of canonical correlation analysis (CCA) to more than two sets of variables and is built on the findings from analyzing PLS, CCA, and its variants. The resulting problem involves a set of orthogonality constraints on view-specific projection matrices, and is numerically challenging to existing methods that may have numerical instabilities and offer no orthogonality guarantee on view-specific projection matrices. To solve this problem, we propose a stable deflation algorithm that relies on proven numerical linear algebra techniques, can guarantee the orthogonality constraints, and simultaneously maximizes the covariance in the common space. We further adapt our algorithm to efficiently handle large-scale high-dimensional data. Extensive experiments have been conducted to evaluate the algorithm through performing two learning tasks, cross-modal retrieval, and multi-view feature extraction. The results demonstrate that the proposed algorithm outperforms the baselines and is scalable for large-scale high-dimensional datasets.
Li Wang 0033, Ren-Cang Li
IEEE Trans. Big Data1
2022 Deep Tensor CCA for Multi-View Learning
abstract
We present Deep Tensor Canonical Correlation Analysis (DTCCA), a method to learn complex nonlinear transformations of multiple views (more than two) of data such that the resulting representations are linearly correlated in high order. The high-order correlation of given multiple views is modeled by covariance tensor, which is different from most CCA formulations relying solely on the pairwise correlations. Parameters of transformations of each view are jointly learned by maximizing the high-order canonical correlation. To solve the resulting problem, we reformulate it as the best sum of rank-1 approximation, which can be efficiently solved by existing tensor decomposition method. DTCCA is a nonlinear extension of tensor CCA (TCCA) via deep networks. Comparing with kernel TCCA, DTCCA not only can deal with arbitrary dimensions of the input data, but also does not need to maintain the training data for computing representations of any given data point. Hence, DTCCA as a unified model can efficiently overcome the scalable issue of TCCA for either high-dimensional multi-view data or a large amount of views, and it also naturally extends TCCA for learning nonlinear representation. Extensive experiments on four multi-view data sets demonstrate the effectiveness of the proposed method.
Hok Shing Wong, Li Wang 0033, Raymond Chan 0001, Tieyong Zeng
IEEE Trans. Big Data2
2021 Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object Detection
abstract
The objective of this paper is to learn context- and depth- aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message propagation (DDMP) network to effectively integrate the multi-scale depth information with the image context; (ii) this is achieved by first adaptively sampling context-aware nodes in the image context and then dynamically predicting hybrid depth-dependent filter weights and affinity matrices for propagating information; (Hi) by augmenting a center-aware depth encoding (CDE) task, our method successfully alleviates the inaccurate depth prior; (iv) we thoroughly demonstrate the effectiveness of our proposed approach and show state-of-the-art results among the monocular-based approaches on the KITTI benchmark dataset. Particularly, we rank 1stin the highly competitive KITTI monocular 3D object detection track on the submission day (November 16th, 2020). Code and models are released at https: //github.com/fudan-zvg/DDMP
Li Wang 0033, Liang Du 0004, Xiaoqing Ye, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001, Jianfeng Feng, Li Zhang 0040
CVPR1
2021 Progressive Coordinate Transforms for Monocular 3D Object Detection
abstract
Recognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for 3D object detection given only a monocular image. While there exist different alternatives for tackling this problem, it is found that they are either equipped with heavy networks to fuse RGB and depth information or empirically ineffective to process millions of pseudo-LiDAR points. With in-depth examination, we realize that these limitations are rooted in inaccurate object localization. In this paper, we propose a novel and lightweight approach, dubbed {\em Progressive Coordinate Transforms} (PCT) to facilitate learning coordinate representations. Specifically, a localization boosting mechanism with confidence-aware loss is introduced to progressively refine the localization prediction. In addition, semantic image representation is also exploited to compensate for the usage of patch proposals. Despite being lightweight and simple, our strategy allows us to establish a new state-of-the-art among the monocular 3D detectors on the competitive KITTI benchmark. At the same time, our proposed PCT shows great generalization to most coordinate-based 3D detection frameworks.
Li Wang 0033, Li Zhang 0040, Yi Zhu 0001, Zhi Zhang 0005, Tong He 0002, Mu Li 0003, Xiangyang Xue 0001
NeurIPS1
2021 Deep Fusion of Brain Structure-Function in Mild Cognitive Impairment
Lu Zhang 0050, Li Wang 0033, Jean Gao, Shannon L. Risacher, Gang Li 0001, Tianming Liu 0001, Dajiang Zhu
Medical Image Anal.2
2021 Probabilistic Structure Learning for EEG/MEG Source Imaging With Hierarchical Graph Priors
abstract
Brain source imaging is an important method for noninvasively characterizing brain activity using Electroencephalogram (EEG) or Magnetoencephalography (MEG) recordings. Traditional EEG/MEG Source Imaging (ESI) methods usually assume the source activities at different time points are unrelated, and do not utilize the temporal structure in the source activation, making the ESI analysis sensitive to noise. Some methods may encourage very similar activation patterns across the entire time course and may be incapable of accounting the variation along the time course. To effectively deal with noise while maintaining flexibility and continuity among brain activation patterns, we propose a novel probabilistic ESI model based on a hierarchical graph prior. Under our method, a spanning tree constraint ensures that activity patterns have spatiotemporal continuity. An efficient algorithm based on an alternating convex search is presented to solve the resulting problem of the proposed model with guaranteed convergence. Comprehensive numerical studies using synthetic data on a realistic brain model are conducted under different levels of signal-to-noise ratio (SNR) from both sensor and source spaces. We also examine the EEG/MEG datasets in two real applications, in which our ESI reconstructions are neurologically plausible. All the results demonstrate significant improvements of the proposed method over benchmark methods in terms of source localization performance, especially at high noise levels.
Feng Liu 0011, Li Wang 0033, Yifei Lou, Ren-Cang Li, Patrick L. Purdon
IEEE Trans. Medical Imaging2
2021 Multi-Material Decomposition for Single Energy CT Using Material Sparsity Constraint
abstract
Multi-material decomposition (MMD) decomposes CT images into basis material images, and is a promising technique in clinical diagnostic CT to identify material compositions within the human body. MMD could be implemented on measurements obtained from spectral CT protocol, although spectral CT data acquisition is not readily available in most clinical environments. MMD methods using single energy CT (SECT), broadly applied in radiological departments of most hospitals, have been proposed in the literature while challenged by the inferior decomposition accuracy and the limited number of material bases due to the constrained material information in the SECT measurement. In this paper, we propose an image-domain SECT MMD method using material sparsity as an assistance under the condition that each voxel of the CT image contains at most two different elemental materials. L0norm represents the material sparsity constraint (MSC) and is integrated into the decomposition objective function with a least-square data fidelity term, total variation term, and a sum-to-one constraint of material volume fractions. An accelerated primal-dual (APD) algorithm with line-search scheme is applied to solve the problem. The pixelwise direct inversion method with the two-material assumption (TMA) is applied to estimate the initials. We validate the proposed method on phantom and patient data. Compared with the TMA method, the proposed MSC method increases the volume fraction accuracy (VFA) from 92.0% to 98.5% in the phantom study. In the patient study, the calcification area can be clearly visualized in the virtual non-contrast image generated by the proposed method, and has a similar shape to that in the ground-truth contrast-free CT image. The high decomposition image quality from the proposed method substantially facilitates the SECT-based MMD clinical applications.
Yi Xue 0002, Wenjian Qin, Yangkang Jiang, Tiffany Tsui, Hongjian He, Li Wang 0033, Jiale Qin, Yaoqin Xie, Tianye Niu
IEEE Trans. Medical Imaging8
2021 Probabilistic Semi-Supervised Learning via Sparse Graph Structure Learning
abstract
We present a probabilistic semi-supervised learning (SSL) framework based on sparse graph structure learning. Different from existing SSL methods with either a predefined weighted graph heuristically constructed from the input data or a learned graph based on the locally linear embedding assumption, the proposed SSL model is capable of learning a sparse weighted graph from the unlabeled high-dimensional data and a small amount of labeled data, as well as dealing with the noise of the input data. Our representation of the weighted graph is indirectly derived from a unified model of density estimation and pairwise distance preservation in terms of various distance measurements, where latent embeddings are assumed to be random variables following an unknown density function to be learned, and pairwise distances are then calculated as the expectations over the density for the model robustness to the data noise. Moreover, the labeled data based on the same distance representations are leveraged to guide the estimated density for better class separation and sparse graph structure learning. A simple inference approach for the embeddings of unlabeled data based on point estimation and kernel representation is presented. Extensive experiments on various data sets show promising results in the setting of SSL compared with many existing methods and significant improvements on small amounts of labeled data.
Li Wang 0033, Raymond Chan 0001, Tieyong Zeng
IEEE Trans. Neural Networks Learn. Syst.1
2020 Recovering Brain Structural Connectivity from Functional Connectivity via Multi-GCN Based Generative Adversarial Network
Lu Zhang 0050, Li Wang 0033, Dajiang Zhu
MICCAI (7)2
2020 Learning Low-Dimensional Latent Graph Structures: A Density Estimation Approach
abstract
We aim to automatically learn a latent graph structure in a low-dimensional space from high-dimensional, unsupervised data based on a unified density estimation framework for both feature extraction and feature selection, where the latent structure is considered as a compact and informative representation of the high-dimensional data. Based on this framework, two novel methods are proposed with very different but intuitive learning criteria from existing methods. The proposed feature extraction method can learn a set of embedded points in a low-dimensional space by naturally integrating the discriminative information of the input data with structure learning so that multiple disconnected embedding structures of data can be uncovered. The proposed feature selection method preserves the pairwise distances only on the optimal set of features and selects these features simultaneously. It not only obtains the optimal set of features but also learns both the structure and embeddings for visualization. Extensive experiments demonstrate that our proposed methods can achieve competitive quantitative (often better) results in terms of discriminant evaluation performance and are able to obtain the embeddings of smooth skeleton structures and select optimal features to unveil the correct graph structures of high-dimensional data sets.
Li Wang 0033, Ren-Cang Li
IEEE Trans. Neural Networks Learn. Syst.1
2019 CODA: Counting Objects via Scale-Aware Adversarial Density Adaption
abstract
Recent advances in crowd counting have achieved promising results with increasingly complex convolutional neural network designs. However, due to the unpredictable domain shift, generalizing trained model to unseen scenarios is often suboptimal. Inspired by the observation that density maps of different scenarios share similar local structures, we propose a novel adversarial learning approach in this paper, i.e., CODA (Counting Objects via scale-aware adversarial Density Adaption). To deal with different object scales and density distributions, we perform adversarial training with pyramid patches of multi-scales from both source-and target-domain. Along with a ranking constraint across levels of the pyramid input, consistent object counts can be produced for different scales. Extensive experiments demonstrate that our network produces much better results on unseen datasets compared with existing counting adaption models. Notably, the performance of our CODA is comparable with the state-of-the-art fully-supervised models that are trained on the target dataset. Further analysis indicates that our density adaption framework can effortlessly extend to scenarios with different objects. The code is available at https://github.com/Willy0919/CODA.
Li Wang 0033, Xiangyang Xue 0001
ICME1
2019 Probabilistic Dimensionality Reduction via Structure Learning
abstract
We propose an alternative probabilistic dimensionality reduction framework that can naturally integrate the generative model and the locality information of data. Based on this framework, we present a new model, which is able to learn a set of embedding points in a low-dimensional space by retaining the inherent structure from high-dimensional data. The objective function of this new model can be equivalently interpreted as two coupled learning problems, i.e., structure learning and the learning of projection matrix. Inspired by this interesting interpretation, we propose another model, which finds a set of embedding points that can directly form an explicit graph structure. We proved that the model by learning explicit graphs generalizes the reversed graph embedding method, but leads to a natural interpretation from Bayesian perspective. This can greatly facilitate data visualization and scientific discovery in downstream analysis. Extensive experiments are performed that demonstrate that the proposed framework is able to retain the inherent structure of datasets and achieve competitive quantitative results in terms of various performance evaluation criteria.
Li Wang 0033, Qi Mao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Efficient Test-Time Predictor Learning With Group-Based Budget
Li Wang 0033, Dajiang Zhu, Yujie Chi
AAAI1
2018 Learning Multiviewpoint Context-Aware Representation for RGB-D Scene Classification
abstract
Effective visual representation plays an important role in the scene classification systems. While many existing methods are focused on the generic descriptors extracted from the RGB color channels, we argue the importance of depth context, since scenes are composed with spatial variability and depth is an essential component in understanding the geometry. In this letter, we present a novel depth representation for RGB-D scene classification based on a specific designed convolutional neural network (CNN). Contrast to previous deep models that transfer from pretrained RGB CNN models, we harness model by using the multiviewpoint depth image augmentation to overcome the data scarcity problem. The proposed CNN framework contains the dilated convolutions to expand the receptive field and a subsequent spatial pooling to aggregate multiscale contextual information. The combination of contextual design and multiviewpoint depth images are important toward a more compact representation, compared to directly using original depth images or off-the-shelf networks. Through extensive experiments on SUN RGB-D dataset, we demonstrate that the representation outperforms recent state of the arts, and combining it with standard CNN-based RGB features can lead to further improvements.
Yingbin Zheng, Hao Ye 0005, Li Wang 0033, Jian Pu
IEEE Signal Process. Lett.3
2018 Arbitrary-Oriented Scene Text Detection via Rotation Proposals
abstract
This paper introduces a novel rotation-based framework for arbitrary-oriented text detection in natural scene images. We present theRotation Region Proposal Networks, which are designed to generate inclined proposals with text orientation angle information. The angle information is then adapted for bounding box regression to make the proposals more accurately fit into the text region in terms of the orientation. TheRotation Region-of-Interestpooling layer is proposed to project arbitrary-oriented proposals to a feature map for a text region classifier. The whole framework is built upon a region-proposal-based architecture, which ensures the computational efficiency of the arbitrary-oriented text detection compared with previous text detection systems. We conduct experiments using the rotation-based framework on three real-world scene text detection datasets and demonstrate its superiority in terms of effectiveness and efficiency over previous approaches.
Jianqi Ma, Weiyuan Shao, Hao Ye 0005, Li Wang 0033, Hong Wang 0014, Yingbin Zheng, Xiangyang Xue 0001
IEEE Trans. Multim.4
2017 Latent Smooth Skeleton Embedding
Li Wang 0033, Qi Mao 0001, Ivor W. Tsang
AAAI1
2017 On the Flatness of Loss Surface for Two-layered ReLU Networks
abstract
Deep learning has achieved unprecedented practical success in many applications. Despite its empirical success, however, the theoretical understanding of deep neural networks still remains a major open problem. In this paper, we explore properties of two-layered ReLU networks. For simplicity, we assume that the optimal model parameters (also called ground-truth parameters) are known. We then assume that a network receives Gaussian input and is trained by minimizing the expected squared loss between the prediction function of the network and a target function. To conduct the analysis, we propose a normal equation for critical points, and study the invariances under three kinds of transformations, namely, scale transformation, rotation transformation and perturbation transformation. We prove that these transformations can keep the loss of a critical point invariant, thus can incur flat regions. Consequently, how to escape from flat regions is vital in training neural networks.
Jiezhang Cao, Qingyao Wu, Yuguang Yan, Li Wang 0033, Mingkui Tan
ACML4
2017 UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
The rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001
AVSS21
2017 Evolving boxes for fast vehicle detection
abstract
We perform fast vehicle detection from traffic surveillance cameras. A novel deep learning framework, namely Evolving Boxes, is developed that proposes and refines the object boxes under different feature representations. Specifically, our framework is embedded with a light-weight proposal network to generate initial anchor boxes as well as to early discard unlikely regions; a fine-turning network produces detailed features for these candidate boxes. We show intriguingly that by applying different feature fusion techniques, the initial boxes can be refined for both localization and recognition. We evaluate our network on the recent DETRAC benchmark and obtain a significant improvement over the state-of-the-art Faster RCNN by 9.5% mAP. Further, our network achieves 9–13 FPS detection speed on a moderate commercial GPU.
Li Wang 0033, Yao Lu 0028, Hong Wang 0014, Yingbin Zheng, Hao Ye 0005, Xiangyang Xue 0001
ICME1
2017 A unified probabilistic framework for robust manifold learning and embedding
Qi Mao 0001, Li Wang 0033, Ivor W. Tsang
Mach. Learn.2
2017 Principal Graph and Structure Learning Based on Reversed Graph Embedding
abstract
Many scientific datasets are of high dimension, and the analysis usually requires retaining the most important structures of data. Principal curve is a widely used approach for this purpose. However, many existing methods work only for data with structures that are mathematically formulated by curves, which is quite restrictive for real applications. A few methods can overcome the above problem, but they either require complicated human-made rules for a specific task with lack of adaption flexibility to different tasks, or cannot obtain explicit structures of data. To address these issues, we develop a novel principal graph and structure learning framework that captures the local information of the underlying graph structure based on reversed graph embedding. As showcases, models that can learn a spanning tree or a weighted undirected `1 graph are proposed, and a new learning algorithm is developed that learns a set of principal points and a graph structure from data, simultaneously. The new algorithm is simple with guaranteed convergence. We then extend the proposed framework to deal with large-scale data. Experimental results on various synthetic and six real world datasets show that the proposed method compares favorably with baselines and can uncover the underlying structure correctly.
Qi Mao 0001, Li Wang 0033, Ivor W. Tsang, Yijun Sun
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Learning Sparse Confidence-Weighted Classifier on Very High Dimensional Data
abstract
Confidence-weighted (CW) learning is a successful online learning paradigm which maintains a Gaussian distribution over classifier weights and adopts a covariancematrix to represent the uncertainties of the weight vectors. However, there are two deficiencies in existing full CW learning paradigms, these being the sensitivity to irrelevant features, and the poor scalability to high dimensional data due to the maintenance of the covariance structure. In this paper, we begin by presenting an online-batch CW learning scheme, and then present a novel paradigm to learn sparse CW classifiers. The proposed paradigm essentially identifies feature groups and naturally builds a block diagonal covariance structure, making it very suitable for CW learning over very high-dimensional data.Extensive experimental results demonstrate the superior performance of the proposed methods over state-of-the-art counterparts on classification and feature selection tasks.
Mingkui Tan, Yan Yan 0006, Li Wang 0033, Anton van den Hengel, Ivor W. Tsang, Qinfeng Shi
AAAI3
2016 Face Recognition via Active Annotation and Learning
abstract
In this paper, we introduce an active annotation and learning framework for the face recognition task. Starting with an initial label deficient face image training set, we iteratively train a deep neural network and use this model to choose the examples for further manual annotation. We follow the active learning strategy and derive the Value of Information criterion to actively select candidate annotation images. During these iterations, the deep neural network is incrementally updated. Experimental results conducted on LFW benchmark and MS-Celeb-1M challenge demonstrate the effectiveness of our proposed framework.
Hao Ye 0005, Weiyuan Shao, Hong Wang 0014, Jianqi Ma, Li Wang 0033, Yingbin Zheng, Xiangyang Xue 0001
ACM Multimedia5
2015 Parallel Hierarchical Clustering in Linearithmic Time for Large-Scale Sequence Analysis
abstract
The rapid development of sequencing technology has led to an explosive accumulation of genomics data. Clustering is often the first step to perform in sequence analysis, and hierarchical clustering is one of the most commonly used approaches for this purpose. However, the standard hierarchical clustering method scales poorly due to its quadratic time and space complexities stemming mainly from the need of computing and storing a pairwise distance matrix. It is thus necessary to minimize the number of pairwise distances computed without degrading clustering performance. On the other hand, as high-performance computing systems are becoming widely accessible, it is highly desirable that a clustering method can be easily adapted to parallel computing environments for further speedup, which is not a trivial task for hierarchical clustering. We proposed a new hierarchical clustering method that achieves good clustering performance and high scalability on large sequence datasets. It consists of two stages. In the first stage, a new landmark-based active hierarchical divisive clustering method was proposed that partitions a large-scale sequence dataset into groups, and in the second stage, a fast hierarchical agglomerative clustering method is applied to each group. By assembling hierarchies from both stages, the hierarchy of the data can be easily recovered. Theoretical results showed that our method can recover the true hierarchy with a high probability under some mild conditions and has a linearithmic time complexity with respect to the number of input sequences. The proposed method also facilitates an efficient parallel implementation. Empirical results on various datasets showed that our method achieved clustering accuracy comparable to ESPRIT-Tree and ran faster than greedy heuristic methods.
Qi Mao 0001, Wei Zheng 0010, Li Wang 0033, Yunpeng Cai, Volker Mai, Yijun Sun
ICDM3
2015 Dimensionality Reduction Via Graph Structure Learning
abstract
We present a new dimensionality reduction setting for a large family of real-world problems. Unlike traditional methods, the new setting aims to explicitly represent and learn an intrinsic structure from data in a high-dimensional space, which can greatly facilitate data visualization and scientific discovery in downstream analysis. We propose a new dimensionality-reduction framework that involves the learning of a mapping function that projects data points in the original high-dimensional space to latent points in a low-dimensional space that are then used directly to construct a graph. Local geometric information of the projected data is naturally captured by the constructed graph. As a showcase, we develop a new method to obtain a discriminative and compact feature representation for clustering problems. In contrast to assumptions used in traditional clustering methods, we assume that centers of clusters should be close to each other if they are connected in a learned graph, and other cluster centers should be distant. Extensive experiments are performed that demonstrate that the proposed method is able to obtain discriminative feature representations yielding superior clustering performance, and correctly recover the intrinsic structures of various real-world datasets including curves, hierarchies and a cancer progression path.
Qi Mao 0001, Li Wang 0033, Steve Goodison, Yijun Sun
KDD2
2015 SimplePPT: A Simple Principal Tree Algorithm
abstract
Many scientific datasets are of high dimension, and the analysis usually requires visual manipulation by retaining the most important structures of data. Principal curve is a widely used approach for this purpose. However, many existing methods work only for data with structures that are not self-intersected, which is quite restrictive for real applications. To address this issue, we develop a new model, which captures the local information of the underlying graph structure based on reversed graph embedding. A generalization bound is derived that show that the model is consistent if the number of data points is sufficiently large. As a special case, a principal tree model is proposed and a new algorithm is developed that learns a tree structure automatically from data. The new algorithm is simple and parameter-free with guaranteed convergence. Experimental results on synthetic and breast cancer datasets show that the proposed method compares favorably with baselines and can discover a breast cancer progression path with multiple branches.
Qi Mao 0001, Li Wang 0033, Steve Goodison, Yijun Sun
SDM3
2015 Generalized Multiple Kernel Learning With Data-Dependent Priors
abstract
Multiple kernel learning (MKL) and classifier ensemble are two mainstream methods for solving learning problems in which some sets of features/views are more informative than others, or the features/views within a given set are inconsistent. In this paper, we first present a novel probabilistic interpretation of MKL such that maximum entropy discrimination with a noninformative prior over multiple views is equivalent to the formulation of MKL. Instead of using the noninformative prior, we introduce a novel data-dependent prior based on an ensemble of kernel predictors, which enhances the prediction performance of MKL by leveraging the merits of the classifier ensemble. With the proposed probabilistic framework of MKL, we propose a hierarchical Bayesian model to learn the proposed data-dependent prior and classification model simultaneously. The resultant problem is convex and other information (e.g., instances with either missing views or missing labels) can be seamlessly incorporated into the data-dependent priors. Furthermore, a variety of existing MKL models can be recovered under the proposed MKL framework and can be readily extended to incorporate these priors. Extensive experiments demonstrate the benefits of our proposed framework in supervised and semisupervised settings, as well as in tasks with partial correspondence among multiple views.
Qi Mao 0001, Ivor W. Tsang, Shenghua Gao, Li Wang 0033
IEEE Trans. Neural Networks Learn. Syst.4
2014 Riemannian Pursuit for Big Matrix Recovery
abstract
Low rank matrix recovery is a fundamental task in many real-world applications. The performance of existing methods, however, deteriorates significantly when applied to ill-conditioned or large-scale matrices. In this paper, we therefore propose an efficient method, called Riemannian Pursuit (RP), that aims to address these two problems simultaneously. Our method consists of a sequence of fixed-rank optimization problems. Each subproblem, solved by a nonlinear Riemannian conjugate gradient method, aims to correct the solution in the most important subspace of increasing size. Theoretically, RP converges linearly under mild conditions and experimental results show that it substantially outperforms existing methods when applied to large-scale and ill-conditioned matrices.
Mingkui Tan, Ivor W. Tsang, Li Wang 0033, Bart Vandereycken, Sinno Jialin Pan
ICML3
2014 Recognizing flu-like symptoms from videos
abstract
BACKGROUND: Vision-based surveillance and monitoring is a potential alternative for early detection of respiratory disease outbreaks in urban areas complementing molecular diagnostics and hospital and doctor visit-based alert systems. Visible actions representing typical flu-like symptoms include sneeze and cough that are associated with changing patterns of hand to head distances, among others. The technical difficulties lie in the high complexity and large variation of those actions as well as numerous similar background actions such as scratching head, cell phone use, eating, drinking and so on. RESULTS: In this paper, we make a first attempt at the challenging problem of recognizing flu-like symptoms from videos. Since there was no related dataset available, we created a new public health dataset for action recognition that includes two major flu-like symptom related actions (sneeze and cough) and a number of background actions. We also developed a suitable novel algorithm by introducing two types of Action Matching Kernels, where both types aim to integrate two aspects of local features, namely the space-time layout and the Bag-of-Words representations. In particular, we show that the Pyramid Match Kernel and Spatial Pyramid Matching are both special cases of our proposed kernels. Besides experimenting on standard testbed, the proposed algorithm is evaluated also on the new sneeze and cough set. Empirically, we observe that our approach achieves competitive performance compared to the state-of-the-arts, while recognition on the new public health dataset is shown to be a non-trivial task even with simple single person unobstructed view. CONCLUSIONS: Our sneeze and cough video dataset and newly developed action recognition algorithm is the first of its kind and aims to kick-start the field of action recognition of flu-like symptoms from videos. It will be challenging but necessary in future developments to consider more complex real-life scenario of detecting these actions simultaneously from multiple persons in possibly crowded environments.
Tuan Hue Thi, Li Wang 0033, Jian Zhang 0002, Sebastian Maurer-Stroh, Li Cheng 0001
BMC Bioinform.2
2014 Minimizing rational functions by exact Jacobian SDP relaxation applicable to finite singularities
Li Wang 0033, Guangming Zhou
J. Glob. Optim.2
2014 Towards ultrahigh dimensional feature selection for big data
Mingkui Tan, Ivor W. Tsang, Li Wang 0033
J. Mach. Learn. Res.3
2013 Minimax Sparse Logistic Regression for Very High-Dimensional Feature Selection
abstract
Because of the strong convexity and probabilistic underpinnings, logistic regression (LR) is widely used in many real-world applications. However, in many problems, such as bioinformatics, choosing a small subset of features with the most discriminative power are desirable for interpreting the prediction model, robust predictions or deeper analysis. To achieve a sparse solution with respect to input features, many sparse LR models are proposed. However, it is still challenging for them to efficiently obtain unbiased sparse solutions to very high-dimensional problems (e.g., identifying the most discriminative subset from millions of features). In this paper, we propose a new minimax sparse LR model for very high-dimensional feature selections, which can be efficiently solved by a cutting plane algorithm. To solve the resultant nonsmooth minimax subproblems, a smoothing coordinate descent method is presented. Numerical issues and convergence rate of this method are carefully studied. Experimental results on several synthetic and real-world datasets show that the proposed method can obtain better prediction accuracy with the same number of selected features and has better or competitive scalability on very high-dimensional problems compared with the baseline methods, including the l1-regularized LR.
Mingkui Tan, Ivor W. Tsang, Li Wang 0033
IEEE Trans. Neural Networks Learn. Syst.3
2012 Convex Matching Pursuit for Large-Scale Sparse Coding and Subset Selection
abstract
In this paper, a new convex matching pursuit scheme is proposed for tackling large-scale sparse coding and subset selection problems. In contrast with current matching pursuit algorithms such as subspace pursuit (SP), the proposed algorithm has a convex formulation and guarantees that the objective value can be monotonically decreased. Moreover, theoretical analysis and experimental results show that the proposed method achieves better scalability while maintaining similar or better decoding ability compared with state-of-the-art methods on large-scale problems.
Mingkui Tan, Ivor W. Tsang, Li Wang 0033
AAAI3
2012 Integrating local action elements for action analysis
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
Comput. Vis. Image Underst.4
2012 Structured learning of local features for human action classification and localization
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
Image Vis. Comput.4
2011 Human Action Segmentation and Recognition Using Discriminative Semi-Markov Models
Qinfeng Shi, Li Cheng 0001, Li Wang 0033, Alexander J. Smola
Int. J. Comput. Vis.3
2011 Elastic Sequence Correlation for Human Action Analysis
abstract
This paper addresses the problem of automatically analyzing and understanding human actions from video footage. An "action correlation" framework, elastic sequence correlation (ESC), is proposed to identify action subsequences from a database of (possibly long) video sequences that are similar to a given query video action clip. In particular, we show that two well-known algorithms, namely approximate pattern matching in computer and information sciences and dynamic time warping (DTW) method in signal processing, are special cases of our ESC framework. The proposed framework is applied to two important real-world applications: action pattern retrieval, as well as action segmentation and recognition, where, on average, its run time speed (in matlab) is about 3.3 frames per second. In addition, comparing with the state-of-the-art algorithms on a number of challenging data sets, our approach is demonstrated to perform competitively.
Li Wang 0033, Li Cheng 0001, Liang Wang 0001
IEEE Trans. Image Process.1
2010 Human Action Recognition and Localization in Video Using Structured Learning of Local Space-Time Features
abstract
This paper presents a unified framework for human action classification and localization in video using structured learning of local space-time features. Each human action class is represented by a set of its own compact set of local patches. In our approach, we first use a discriminative hierarchical Bayesian classifier to select those space-time interest points that are constructive for each particular action. Those concise local features are then passed to a Support Vector Machine with Principal Component Analysis projection for the classification task. Meanwhile, the action localization is done using Dynamic Conditional Random Fields developed to incorporate the spatial and temporal structure constraints of superpixels extracted around those features. Each superpixel in the video is defined by the shape and motion information of its corresponding feature region. Compelling results obtained from experiments on KTH [22], Weizmann [1], HOHA [13] and TRECVid [23] datasets have proven the efficiency and robustness of our framework for the task of human action recognition and localization in video.
Tuan Hue Thi, Jian Zhang 0002, Li Cheng 0001, Li Wang 0033, Shin'ichi Satoh 0001
AVSS4
2010 Implicit Motion-Shape Model: A generic approach for action matching
abstract
We develop a robust technique to find similar matches of human actions in video. Given a query video, Motion History Images (MHI) are constructed for consecutive keyframes. This is followed by dividing the MHI into local Motion-Shape regions, which allows us to analyze the action as a set of sparse space-time patches in 3D. Inspired by the idea of Generalized Hough Transform, we develop the Implicit Motion-Shape Model that allows the integration of these local patches to describe the dynamic characteristics of the query action. In the same way we retrieve motion segments from video candidates, then project them onto the Hough Space built by the query model. This produces the matching score by running Parzen window density estimation under different scales. Empirical experiments on popular datasets demonstrate the efficiency of this approach, where highly accurate matches are returned within acceptable processing time.
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033
ICIP4
2010 Learning Sparse SVM for Feature Selection on Very High Dimensional Datasets
Mingkui Tan, Li Wang 0033, Ivor W. Tsang
ICML2
2010 Weakly Supervised Action Recognition Using Implicit Shape Models
abstract
In this paper, we present a robust framework for action recognition in video, that is able to perform competitively against the state-of-the-art methods, yet does not rely on sophisticated background subtraction preprocess to remove background features. In particular, we extend the Implicit Shape Modeling (ISM) of [10] for object recognition to 3D to integrate local spatiotemporal features, which are produced by a weakly supervised Bayesian kernel filter. Experiments on benchmark datasets (including KTH and Weizmann) verifies the effectiveness of our approach.
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
ICPR4
2009 Human Body Articulation for Action Recognition in Video Sequences
abstract
This paper presents a new technique for action recognition in video using human body part-based approach, combining both local feature description of each body part, and global graphical model structure of the human action. The human body is divided into elementary points from which a Decomposable Triangulated Graph will be built. The temporal variation of human activity is encoded in the velocity distribution of each node in the graph, while the graph structure shows the spatial configuration of all the nodes in the action. Tracking trajectories of unlabeled good feature points are correctly labeled using Maximum a Posterior probability. Dynamic Programming is then implemented to boost up the exhaustive search for the optimal labeling of unknown body parts and the best possible action. A simple and efficient technique for building the optimal structure of the human action graph is also implemented. Experimental results on the KTH dataset proves the success and potential applications of this proposed technique.
Tuan Hue Thi, Sijun Lu, Jian Zhang 0002, Li Cheng 0001, Li Wang 0033
AVSS5
2008 Discriminative human action segmentation and recognition using semi-Markov model
abstract
Given an input video sequence of one person conducting a sequence of continuous actions, we consider the problem of jointly segmenting and recognizing actions. We propose a discriminative approach to this problem under a semi-Markov model framework, where we are able to define a set of features over input-output space that captures the characteristics on boundary frames, action segments and neighboring action segments, respectively. In addition, we show that this method can also be used to recognize the person who performs in this video sequence. A Viterbi-like algorithm is devised to help efficiently solve the induced optimization problem. Experiments on a variety of datasets demonstrate the effectiveness of the proposed method.
Qinfeng Shi, Li Wang 0033, Li Cheng 0001, Alexander J. Smola
CVPR2