Juhua Hu

dblp:147/2228 · also Ju-Hua Hu · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
15since 2021 · last 2024
0000-0001-5869-3549ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 13 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2024 Text-Guided Mixup Towards Long-Tailed Image Categorization
Richard Franklin, Jiawei Yao, Deyang Zhong, Qi Qian 0001, Juhua Hu
BMVC5
2024 Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
abstract
Multiple clustering has gained significant attention in recent years due to its potential to reveal multiple hidden structures of data from different perspectives. The advent of deep multiple clustering techniques has notably advanced the performance by uncovering complex patterns and relationships within large datasets. However, a major challenge arises as users often do not need all the clusterings that algorithms generate, and figuring out the one needed requires a substantial understanding of each clustering result. Traditionally, aligning a user's brief keyword of interest with the corresponding vision components was challenging, but the emergence of multi-modal and large language models (LLMs) has begun to bridge this gap. In response, given unlabeled target visual data, we propose Multi-MaP, a novel method employing a multi-modal proxy learning process. It leverages CLIP encoders to extract coherent text and image embeddings, with GPT-4 integrating users' interests to formulate effective textual contexts. Moreover, reference word constraint and concept-level constraint are designed to learn the optimal text proxy according to the user's interest. Multi-MaP not only adeptly captures a user's interest via a keyword but also facilitates identifying relevant clusterings. Our extensive experiments show that Multi-MaP consistently outperforms state-of-the-art methods in all benchmark multi-clustering vision tasks. Our code is available at https://github.com/Alexander-Yao/Multi-MaP.
Jiawei Yao, Qi Qian 0001, Juhua Hu
CVPR3
2024 Online Zero-Shot Classification with CLIP
Qi Qian 0001, Juhua Hu
ECCV (77)2
2024 SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning
Qi Qian 0001, Yuanhong Xu, Juhua Hu
ECCV (78)3
2024 Curbside Parking Occupancy Detection - Dashcam-Based Solutions
abstract
Accurate assessment of curbside parking occupancy is essential for policymakers to optimize public resource allocation. It is also beneficial for driver/autonomous vehicles’ parking planning. Despite its importance, there is still no comprehensive solution of curbside parked vehicle detection for different road and traffic scenarios. We therefore propose two computer vision based solutions for efficiently quantify parked vehicles in simple and complex scenarios, respectively, from the street videos taken by off-the-shelf dash cameras. The proposed AI pipelines encompass multiple tasks, including vehicle detection and tracking, road surface detection, and lane line detection. The interplay between detected vehicles, road surface, and lane lines enhances the robustness of feature engineering. Through evaluations, our solutions demonstrate their capability to handle diverse road and traffic scenarios, including busy main roads, quiet side roads, and residential areas.
Hanming Zhang, Juhua Hu
MDM3
2024 Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning
abstract
Multiple clustering aims to discover various latent structures of data from different aspects. Deep multiple clustering methods have achieved remarkable performance by exploiting complex patterns and relationships in data. However, existing works struggle to flexibly adapt to diverse user-specific needs in data grouping, which may require manual understanding of each clustering. To address these limitations, we introduce Multi-Sub, a novel end-to-end multiple clustering approach that incorporates a multi-modal subspace proxy learning framework in this work. Utilizing the synergistic capabilities of CLIP and GPT-4, Multi-Sub aligns textual prompts expressing user preferences with their corresponding visual representations. This is achieved by automatically generating proxy words from large language models that act as subspace bases, thus allowing for the customized representation of data in terms specific to the user’s interests. Our method consistently outperforms existing baselines across a broad set of datasets in visual multiple clustering tasks. Our code is available at https://github.com/Alexander-Yao/Multi-Sub.
Jiawei Yao, Qi Qian 0001, Juhua Hu
NeurIPS3
2024 Dual-disentangled Deep Multiple Clustering
abstract
Multiple clustering has gathered significant attention in recent years due to its potential to reveal multiple hidden structures of the data from different perspectives. Most of multiple clustering methods first derive feature representations by controlling the dissimilarity among them, subsequently employing traditional clustering methods (e.g., k-means) to achieve the final multiple clustering outcomes. However, the learned feature representations can exhibit a weak relevance to the ultimate goal of distinct clustering. Moreover, these features are often not explicitly learned for the purpose of clustering. Therefore, in this paper, we propose a novel Dual-Disentangled deep Multiple Clustering method named DDMC by learning disentangled representations. Specifically, DDMC is achieved by a variational Expectation-Maximization (EM) framework. In the E-step, the disentanglement learning module employs coarse-grained and fine-grained disentangled representations to obtain a more diverse set of latent factors from the data. In the M-step, the cluster assignment module utilizes a cluster objective function to augment the effectiveness of the cluster output. Our extensive experiments demonstrate that DDMC consistently outperforms state-of-the-art methods across seven commonly used tasks. Our code is available at https://github.com/Alexander-Yao/DDMC.
Jiawei Yao, Juhua Hu
SDM2
2023 NPRL: Nightly Profile Representation Learning for Early Sepsis Onset Prediction in ICU Trauma Patients
abstract
Sepsis is a syndrome that develops in the body in response to the presence of an infection. Characterized by severe organ dysfunction, sepsis is one of the leading causes of mortality in Intensive Care Units (ICUs) worldwide. These complications can be reduced through early application of antibiotics. Hence, the ability to anticipate the onset of sepsis early is crucial to the survival and well-being of patients. Current machine learning algorithms deployed inside medical infrastructures have demonstrated poor performance and are insufficient for anticipating sepsis onset early. Recently, deep learning methodologies have been proposed to predict sepsis, but some fail to capture the time of onset (e.g., classifying patients’ entire visits as developing sepsis or not) and others are unrealistic for deployment in clinical settings (e.g., creating training instances using a fixed time to onset, where the time of onset needs to be known apriori). In this paper, we first propose a novel but realistic prediction framework that predicts each morning whether sepsis onset will occur within the next 24 hours using the most recent data collected the previous night, when patient-provider ratios are higher due to cross-coverage resulting in limited observation to each patient. However, as we increase the prediction rate into daily, the number of negative instances will increase, while that of positive instances remain the same. This causes a severe class imbalance problem making it hard to capture these rare sepsis cases. To address this, we propose a nightly profile representation learning (NPRL) approach. We prove that NPRL can theoretically alleviate the rare event problem and our empirical study using data from a level-1 trauma center demonstrates the effectiveness of our proposal.
Tucker Stewart, Katherine Stern, Grant E. O'Keefe, Ankur Teredesai, Juhua Hu
IEEE Big Data5
2023 Dictionary-Guided Text Recognition for Smart Street Parking
Deyang Zhong, Juhua Hu
BMVC4
2023 Improved Visual Fine-tuning with Natural Language Supervision
abstract
Fine-tuning a visual pre-trained model can leverage the semantic information from large-scale pre-training data and mitigate the over-fitting problem on downstream vision tasks with limited training examples. While the problem of catastrophic forgetting in pre-trained backbone has been extensively studied for fine-tuning, its potential bias from the corresponding pre-training task and data, attracts less attention. In this work, we investigate this problem by demonstrating that the obtained classifier after fine-tuning will be close to that induced by the pre-trained model. To reduce the bias in the classifier effectively, we introduce a reference distribution obtained from a fixed text classifier, which can help regularize the learned vision classifier. The proposed method, Text Supervised fine-tuning (TeS), is evaluated with diverse pre-trained vision models including ResNet and ViT, and text encoders including BERT and CLIP, on 11 downstream tasks. The consistent improvement with a clear margin over distinct scenarios confirms the effectiveness of our proposal. Code is available at https://github.com/idstcv/TeS.
Junyang Wang 0001, Yuanhong Xu, Juhua Hu, Ming Yan 0008, Jitao Sang 0001, Qi Qian 0001
ICCV3
2023 Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIP
abstract
Vision-language pre-training methods, e.g., CLIP, demonstrate an impressive zero-shot performance on visual categorizations with the class proxy from the text embedding of the class name. However, the modality gap between the text and vision space can result in a sub-optimal performance. We theoretically show that the gap cannot be reduced sufficiently by minimizing the contrastive loss in CLIP and the optimal proxy for vision tasks may reside only in the vision space. Therefore, given unlabeled target vision data, we propose to learn the vision proxy directly with the help from the text proxy for zero-shot transfer. Moreover, according to our theoretical analysis, strategies are developed to further refine the pseudo label obtained by the text proxy to facilitate the intra-modal proxy learning (InMaP) for vision. Experiments on extensive downstream tasks confirm the effectiveness and efficiency of our proposal. Concretely, InMaP can obtain the vision proxy within one minute on a single GPU while improving the zero-shot accuracy from $77.02\%$ to $80.21\%$ on ImageNet with ViT-L/14@336 pre-trained by CLIP.
Qi Qian 0001, Yuanhong Xu, Juhua Hu
NeurIPS3
2022 Sub-Sequence Graph Representation Learning on High Variability Data for Dynamic Risk Prediction in Critical Care
abstract
Sepsis is an extreme inflammatory response of the body to an infection. It is one of the leading causes of death in critical care and ICUs worldwide, resulting in approximately 25% mortality in critically ill populations. Early identification and intervention are crucial to reducing sepsis-associated mortality and improving patient prognosis because severe sepsis cases can lead to organ failure and other life-threatening complications. Diagnosis of Sepsis is challenging in terms of diagnostic accuracy and timeliness due to ambiguous symptoms and individual differences which are captured in heterogeneous varied data sources with irregular time sequences. In recent years, numerous efforts have been made using machine learning methods for sepsis prediction. However, there are still very limited successful industry scale implementations due to limitations of consistency of input data, which often does not fit the characteristics of uncertain time intervals and large number of missing values in real-world critical care settings. In this study, we demonstrate an innovative approach to predict sepsis occurrence in real-time, on heterogeneous sets of variables with multiple time-granularity, using a flexible graph structure to model patient health records. Our modeling task is to predict future sepsis risk at any time after the first 48 hours of admission using observations data from any time window within the past 12 hours. To the best of our knowledge, our proposed approach is the first ever implementation that is dynamic, uses a graph representation to overcome the problem of irregular input features, and offers continuous risk prediction, thereby improving the compatibility of the model in clinical settings. Unlike prior efforts that report results on de-identified curated public datasets with synthetic data filters, our experiments and results are validated by clinical experts, and are based on a large level-1 trauma center’s multi-year real-world longitudinal data.
Ankur Teredesai, Sijin Huang, Tucker Stewart, Juhua Hu, Armaan Thakker, Katherine Stern, Grant E. O'Keefe
IEEE Big Data4
2022 Unsupervised Visual Representation Learning by Online Constrained K-Means
abstract
Cluster discrimination is an effective pretext task for unsupervised representation learning, which often consists of two phases: clustering and discrimination. Clustering is to assign each instance a pseudo label that will be used to learn representations in discrimination. The main challenge resides in clustering since prevalent clustering methods (e.g., k-means) have to run in a batch mode. Besides, there can be a trivial solution consisting of a dominating cluster. To address these challenges, we first investigate the objective of clustering-based representation learning. Based on this, we propose a novel clustering-based pretext task with online Constrained K-means (CoKe). Compared with the balanced clustering that each cluster has exactly the same size, we only constrain the minimal size of each cluster to flexibly capture the inherent data structure. More importantly, our online assignment method has a theoretical guarantee to approach the global optimum. By decoupling clustering and discrimination, CoKe can achieve competitive performance when optimizing with only a single view from each instance. Extensive experiments on ImageNet and other benchmark data sets verify both the efficacy and efficiency of our proposal.
Qi Qian 0001, Yuanhong Xu, Juhua Hu, Hao Li 0030, Rong Jin 0001
CVPR3
2022 Improved Knowledge Distillation via Full Kernel Matrix Transfer
abstract
Knowledge distillation is an effective way for model compression in deep learning. Given a large model (i.e., teacher model), it aims to improve the performance of a compact model (i.e., student model) by transferring the information from the teacher. Various information for distillation has been studied. Recently, a number of works propose to transfer the pairwise similarity between examples to distill relative information. However, most of efforts are devoted to developing different similarity measurements, while only a small matrix consisting of examples within a mini-batch is transferred at each iteration that can be inefficient for optimizing the pairwise similarity over the whole data set. In this work, we aim to transfer the full similarity matrix effectively. The main challenge is from the size of the full matrix that is quadratic to the number of examples. To address the challenge, we decompose the original full matrix with Nyström method. By selecting appropriate landmark points, our theoretical analysis indicates that the loss for transfer can be further simplified. Concretely, we find that the difference between the original full kernel matrices between teacher and student can be well bounded by that of the corresponding partial matrices, which only consists of similarities between original examples and landmark points. Compared with the full matrix, the size of the partial matrix is linear in the number of examples, which improves the efficiency of optimization significantly. The empirical study on benchmark data sets demonstrates the effectiveness of the proposed algorithm.
Qi Qian 0001, Hao Li 0030, Juhua Hu
SDM3
2021 Weakly Supervised Representation Learning with Coarse Labels
abstract
With the development of computational power and techniques for data collection, deep learning demonstrates a superior performance over most existing algorithms on visual benchmark data sets. Many efforts have been devoted to studying the mechanism of deep learning. One important observation is that deep learning can learn the discriminative patterns from raw materials directly in a task-dependent manner. Therefore, the representations obtained by deep learning outperform hand-crafted features significantly. However, for some real-world applications, it is too expensive to collect the task-specific labels, such as visual search in online shopping. Compared to the limited availability of these task-specific labels, their coarse-class labels are much more affordable, but representations learned from them can be suboptimal for the target task. To mitigate this challenge, we propose an algorithm to learn the fine-grained patterns for the target task, when only its coarse-class labels are available. More importantly, we provide a theoretical guarantee for this. Extensive experiments on real-world data sets demonstrate that the proposed method can significantly improve the performance of learned representations on the target task, when only coarse-class information is available for training.
Yuanhong Xu, Qi Qian 0001, Hao Li 0030, Rong Jin 0001, Juhua Hu
ICCV5
2020 Hierarchically Robust Representation Learning
abstract
With the tremendous success of deep learning in visual tasks, the representations extracted from intermediate layers of learned models, that is, deep features, attract much attention of researchers. Previous empirical analysis shows that those features can contain appropriate semantic information. Therefore, with a model trained on a large-scale benchmark data set (e.g., ImageNet), the extracted features can work well on other tasks. In this work, we investigate this phenomenon and demonstrate that deep features can be suboptimal due to the fact that they are learned by minimizing the empirical risk. When the data distribution of the target task is different from that of the benchmark data set, the performance of deep features can degrade. Hence, we propose a hierarchically robust optimization method to learn more generic features. Considering the example-level and concept-level robustness simultaneously, we formulate the problem as a distributionally robust optimization problem with Wasserstein ambiguity set constraints, and an efficient algorithm with the conventional training pipeline is proposed. Experiments on benchmark data sets demonstrate the effectiveness of the robust deep representations.
Qi Qian 0001, Juhua Hu, Hao Li 0030
CVPR2
2019 Cost-adaptive Neural Networks for Peak Volume Prediction with EMM Filtering
abstract
As the emergence of the Internet of Things (IoT) and the growing number of IoT devices, a stable connection service has become one of the key factors concerning the Quality of Service (QoS) provision. How to anticipate the peak traffic volume is essential. If the resource allocation is under provisioned, the service becomes susceptible to failure or security breach. Unfortunately, peak volumes are not captured in the systematic components of data and as a result conventional trend prediction methods have proven insufficient. We propose a framework that implements neural networks with filtering and a cost-adaptive loss function to improve the ability to predict peak volumes. Implementing this method on a real Domain Name Server (DNS) traffic data, we observe not only the improvement in the prediction performance but also a shorter lag time to predict peak values, which demonstrates our proposed method.
Giovanna Graciani, Anderson C. A. Nascimento, Juhua Hu
IEEE BigData4
2019 SoftTriple Loss: Deep Metric Learning Without Triplet Sampling
abstract
Distance metric learning (DML) is to learn the embeddings where examples from the same class are closer than examples from different classes. It can be cast as an optimization problem with triplet constraints. Due to the vast number of triplet constraints, a sampling strategy is essential for DML. With the tremendous success of deep learning in classifications, it has been applied for DML. When learning embeddings with deep neural networks (DNNs), only a mini-batch of data is available at each iteration. The set of triplet constraints has to be sampled within the mini-batch. Since a mini-batch cannot capture the neighbors in the original set well, it makes the learned embeddings sub-optimal. On the contrary, optimizing SoftMax loss, which is a classification loss, with DNN shows a superior performance in certain DML tasks. It inspires us to investigate the formulation of SoftMax. Our analysis shows that SoftMax loss is equivalent to a smoothed triplet loss where each class has a single center. In real-world data, one class can contain several local clusters rather than a single one, e.g., birds of different poses. Therefore, we propose the SoftTriple loss to extend the SoftMax loss with multiple centers for each class. Compared with conventional deep metric learning algorithms, optimizing SoftTriple loss can learn the embeddings without the sampling phase by mildly increasing the size of the last fully connected layer. Experiments on the benchmark fine-grained data sets demonstrate the effectiveness of the proposed loss function.
Qi Qian 0001, Baigui Sun, Juhua Hu, Tacoma Tacoma, Hao Li 0030, Rong Jin 0001
ICCV4
2018 Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
abstract
Strong intelligent machines powered by deep neural networks are increasingly deployed as black boxes to make decisions in risk-sensitive domains, such as finance and medical. To reduce potential risk and build trust with users, it is critical to interpret how such machines make their decisions. Existing works interpret a pre-trained neural network by analyzing hidden neurons, mimicking pre-trained models or approximating local predictions. However, these methods do not provide a guarantee on the exactness and consistency of their interpretations. In this paper, we propose an elegant closed form solution named $OpenBox$ to compute exact and consistent interpretations for the family of Piecewise Linear Neural Networks (PLNN). The major idea is to first transform a PLNN into a mathematically equivalent set of linear classifiers, then interpret each linear classifier by the features that dominate its prediction. We further apply $OpenBox$ to demonstrate the effectiveness of non-negative and sparse constraints on improving the interpretability of PLNNs. The extensive experiments on both synthetic and real world data sets clearly demonstrate the exactness and consistency of our interpretation.
Lingyang Chu, Juhua Hu, Lanjun Wang, Jian Pei 0001
KDD3
2018 Subspace multi-clustering: a review
Juhua Hu, Jian Pei 0001
Knowl. Inf. Syst.1
2017 Finding multiple stable clusterings
Juhua Hu, Qi Qian 0001, Jian Pei 0001, Rong Jin 0001, Shenghuo Zhu
Knowl. Inf. Syst.1
2015 Finding Multiple Stable Clusterings
abstract
Multi-clustering, which tries to find multiple independent ways to partition a data set into groups, has enjoyed many applications, such as customer relationship management, bioinformatics and healthcare informatics. This paper addresses two fundamental questions in multi-clustering: how to model the quality of clusterings and how to find multiple stable clusterings. We introduce to multi-clustering the notion of clustering stability based on Laplacian eigengap, which was originally used in the regularized spectral learning method for similarity matrix learning. We mathematically prove that the larger the eigengap, the more stable the clustering. Consequently, we propose a novel multi-clustering method MSC (for Multiple Stable Clustering). An advantage of our method comparing to the existing multi-clustering methods is that our method does not need any parameter about the number of alternative clusterings in the data set. Our method can heuristically estimate the number of meaningful clusterings in a data set, which is infeasible in the existing multi-clustering methods. We report an empirical study that clearly demonstrates the effectiveness of our method.
Juhua Hu, Qi Qian 0001, Jian Pei 0001, Rong Jin 0001, Shenghuo Zhu
ICDM1
2015 Pairwised Specific Distance Learning from Physical Linkages
abstract
In real tasks, usually a good classification performance can only be obtained when a good distance metric is obtained; therefore, distance metric learning has attracted significant attention in the past few years. Typical studies of distance metric learning evaluate how to construct an appropriate distance metric that is able to separate training data points from different classes or satisfy a set of constraints (e.g., must-links and/or cannot-links). It is noteworthy that this task becomes challenging when there are only limited labeled training data points and no constraints are given explicitly. Moreover, most existing approaches aim to construct a global distance metric that is applicable to all data points. However, different data points may have different properties and may require different distance metrics. We notice that data points in real tasks are often connected by physical links (e.g., people are linked with each other in social networks; personal webpages are often connected to other webpages, including nonpersonal webpages), but the linkage information has not been exploited in distance metric learning. In this article, we develop a pairwised specific distance (PSD) approach that exploits the structures of physical linkages and in particular captures the key observations that nonmetric and clique linkages imply the appearance of different or unique semantics, respectively. It is noteworthy that, rather than generating a global distance, PSD generates different distances for different pairs of data points; this property is desired in applications involving complicated data semantics. We mainly present PSD for multi-class learning and further extend it to multi-label learning. Experimental results validate the effectiveness of PSD, especially in the scenarios in which there are very limited labeled training data points and no explicit constraints are given.
Juhua Hu, De-Chuan Zhan, Xintao Wu, Yuan Jiang 0001, Zhi-Hua Zhou
ACM Trans. Knowl. Discov. Data1
2014 Distance metric learning using dropout: a structured regularization approach
abstract
Distance metric learning (DML) aims to learn a distance metric better than Euclidean distance. It has been successfully applied to various tasks, e.g., classification, clustering and information retrieval. Many DML algorithms suffer from the over-fitting problem because of a large number of parameters to be determined in DML. In this paper, we exploit the dropout technique, which has been successfully applied in deep learning to alleviate the over-fitting problem, for DML. Different from the previous studies that only apply dropout to training data, we apply dropout to both the learned metrics and the training data. We illustrate that application of dropout to DML is essentially equivalent to matrix norm based regularization. Compared with the standard regularization scheme in DML, dropout is advantageous in simulating the structured regularizers which have shown consistently better performance than non structured regularizers. We verify, both empirically and theoretically, that dropout is effective in regulating the learned metric to avoid the over-fitting problem. Last, we examine the idea of wrapping the dropout technique in the state-of-art DML methods and observe that the dropout technique can significantly improve the performance of the original DML methods.
Qi Qian 0001, Juhua Hu, Rong Jin 0001, Jian Pei 0001, Shenghuo Zhu
KDD2
2014 How Can I Index My Thousands of Photos Effectively and Automatically? An Unsupervised Feature Selection Approach
abstract
Given a large photo collection without domain knowledge (e.g., tourism photos, conference photos, event photos, images wrapped from webpages), it is not easy for human beings to organize or only view them within a reasonable time. In this paper, we propose to automatically extract meaningful semantics from a photo collection named “dimensions” to help people view, search and organize photos conveniently and efficiently. However, due to the lack of additional domain knowledge or content information, existing image retrieval techniques are not applicable. To tackle the problem, we first propose a simple strategy to extract all meaningful semantics from original photos/images as candidate dimensions, and then propose an efficient unsupervised feature/dimension selection method to select a sufficient dimension subset to uniquely index each photo within this collection. Our experiments on several real-world photo/image collections validate both the efficiency and effectiveness of our proposed method.
Juhua Hu, Jian Pei 0001, Jie Tang 0001
SDM1
2012 Towards Discovering What Patterns Trigger What Labels
abstract
In many real applications, especially those involving data objects with complicated semantics, it is generally desirable to discover the relation between patterns in the input space and labels corresponding to different semantics in the output space. This task becomes feasible with MIML (Multi-Instance Multi-Label learning), a recently developed learning framework, where each data object is represented by multiple instances and is allowed to be associated with multiple labels simultaneously. In this paper, we propose KISAR, an MIML algorithm that is able to discover what instances trigger what labels. By considering the fact that highly relevant labels usually share some patterns, we develop a convex optimization formulation and provide an alternating optimization solution. Experiments show that KISAR is able to discover reasonable relations between input patterns and output labels, and achieves performances that are highly competitive with many state-of-the-art MIML algorithms.
Yufeng Li 0008, Juhua Hu, Yuan Jiang 0001, Zhi-Hua Zhou
AAAI2