Harry Zhang

dblp:12/2414 · DBLP profile ↗
← Back
64ranked-venue papers
14as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 11 first-author · 9 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 6 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 CRISP: Object Pose and Shape Estimation with Test-Time Adaptation
abstract
We consider the problem of estimating object pose and shape from an RGB-D image. Our first contribution is to introduce CRISP, a category-agnostic object pose and shape estimation pipeline. The pipeline implements an encoder-decoder model for shape estimation. It uses FiLM-conditioning for implicit shape reconstruction and a DPT-based network for estimating pose-normalized points for pose estimation. As a second contribution, we propose an optimization-based pose and shape corrector that can correct estimation errors caused by a domain gap. Observing that the shape decoder is well behaved in the convex hull of known shapes, we approximate the shape decoder with an active shape model, and show that this reduces the shape correction problem to a constrained linear least squares problem, which can be solved efficiently by an interior point algorithm. Third, we introduce a self-training pipeline to perform self-supervised domain adaptation of CRISP. The self-training is based on a correct-and-certify approach, which leverages the corrector to generate pseudo-labels at test time, and uses them to self-train CRISP. We demonstrate CRISP (and the self-training) on YCBV, SPE3R, and NOCS datasets. CRISP shows high performance on all the datasets. Moreover, our self-training is capable of bridging a large domain gap. Finally, CRISP also shows an ability to generalize to unseen objects. Code, pre-trained models and videos of sample results are available on the project webpage.1
Jingnan Shi, Rajat Talak, Harry Zhang, David Jin, Luca Carlone
CVPR3
2025 An Efficient Private Set Frequency Query Scheme Under Local Differential Privacy
abstract
Crowdsourcing has become a widely utilized method for data collection and analysis; however, privacy concerns remain a significant challenge. In this paper, we introduce a novel and efficient private set frequency (PSF) query scheme designed for crowdsourcing scenarios. Our proposed scheme is based on edge computing and leverages local differential privacy (LDP) and Bloom filter techniques to ensure both query privacy and high communication efficiency. Specifically, we employ two non-colluding edge devices to assist the server in achieving highaccuracy query result estimation while preserving the privacy of both the server's query set and users' sensitive data. A comprehensive security analysis confirms that the query value remains confidential, and users' privacy is guaranteed under$\varepsilon$-LDP. Additionally, performance evaluations demonstrate the efficiency and improved accuracy of our proposed scheme.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
ICC4
2025 CHAMP: Conformalized 3D Human Multi-Hypothesis Pose Estimators
abstract
We introduce CHAMP, a novel method for learning sequence-to-sequence, multi-hypothesis 3D human poses from 2D keypoints by leveraging a conditional distribution with a diffusion model. To predict a single output 3D pose sequence, we generate and aggregate multiple 3D pose hypotheses. For better aggregation results, we develop a method to score these hypotheses during training, effectively integrating conformal prediction into the learning process. This process results in a differentiable conformal predictor that is trained end-to-end with the 3D pose estimator. Post-training, the learned scoring model is used as the conformity score, and the 3D pose estimator is combined with a conformal predictor to select the most accurate hypotheses for downstream aggregation. Our results indicate that using a simple mean aggregation on the conformal prediction-filtered hypotheses set yields competitive results. When integrated with more sophisticated aggregation techniques, our method achieves state-of-the-art performance across various metrics and datasets while inheriting the probabilistic guarantees of conformal prediction.
Harry Zhang, Luca Carlone
ICLR1
2025 CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep Uncertainty
abstract
We introduce CUPS, a novel method for learning sequence-to-sequence 3D human shapes and poses from RGB videos with uncertainty quantification. To improve on top of prior work, we develop a method to generate and score multiple hypotheses during training, effectively integrating uncertainty quantification into the learning process. This process results in a deep uncertainty function that is trained end-to-end with the 3D pose estimator. Post-training, the learned deep uncertainty model is used as the conformity score, which can be used to calibrate a conformal predictor in order to assess the quality of the output prediction. Since the data in human pose-shape learning is not fully exchangeable, we also present two practical bounds for the coverage gap in conformal prediction, developing theoretical backing for the uncertainty bound of our model. Our results indicate that by taking advantage of deep uncertainty with conformal prediction, our method achieves state-of-the-art performance across various metrics and datasets while inheriting the probabilistic guarantees of conformal prediction. Interactive 3D visualization, code, and data will be available at https://sites.google.com/view/champpp.
Harry Zhang, Luca Carlone
ICML1
2025 Enhancing Autonomous Navigation by Imaging Hidden Objects Using Single-Photon LiDAR
abstract
Robust autonomous navigation in environments with limited visibility remains a critical challenge in robotics. We present a novel approach that leverages Non-Line-of-Sight (NLOS) sensing using single-photon LiDAR to improve visibility and enhance autonomous navigation. Our method enables mobile robots to “see around corners” by utilizing multi-bounce light information, effectively expanding their perceptual range without additional infrastructure. We propose a three-module pipeline: (1) Sensing, which captures multi-bounce histograms using SPAD-based LiDAR; (2) Perception, which estimates occupancy maps of hidden regions from these histograms using a convolutional neural network; and (3) Control, which allows a robot to follow safe paths based on the estimated occupancy. We evaluate our approach through simulations and real-world experiments on a mobile robot navigating an L-shaped corridor with hidden obstacles. Our work represents the first experimental demonstration of NLOS imaging for autonomous navigation, paving the way for safer and more efficient robotic systems operating in complex environments. We also contribute a novel dynamics-integrated transient rendering framework for simulating NLOS scenarios, facilitating future research in this domain.
Aaron Young, Nevindu Batagoda, Harry Zhang, Akshat Dave, Adithya Kumar Pediredla, Dan Negrut, Ramesh Raskar
ICRA3
2025 GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
Harry Zhang, Kurt Partridge, Pai Zhu, Neng Chen, Hyun Jin Park, Dhruuv Agarwal
INTERSPEECH1
2025 Max Entropy Moment Kalman Filter for Polynomial Systems with Arbitrary Noise
abstract
Designing optimal Bayes filters for nonlinear non-Gaussian systems is a challenging task. The main difficulties are: 1) representing complex beliefs, 2) handling non-Gaussian noise, and 3) marginalizing past states. To address these challenges, we focus on polynomial systems and propose the Max Entropy Moment Kalman Filter (MEM-KF). To address 1), we represent arbitrary beliefs by a Moment-Constrained Max-Entropy Distribution (MED). The MED can asymptotically approximate almost any distribution given an increasing number of moment constraints. To address 2), we model the noise in the process and observation model as MED. To address 3), we propagate the moments through the process model and recover the distribution as MED, thus avoiding symbolic integration, which is generally intractable. All the steps in MEM-KF, including the extraction of a point estimate, can be solved via convex optimization. We showcase the MEM-KF in challenging robotics tasks, such as localization with unknown data association.
Sangli Teng, Harry Zhang, David Jin, Ashkan Jasour, Ramanarayan Vasudevan, Maani Ghaffari Jadidi, Luca Carlone
NeurIPS2
2025 Optimized Sparse Vector Aggregation Under Local Differential Privacy
abstract
In crowdsourcing applications, gathering and analyzing users’ strong positive (1) or negative (-1) reactions to a large number of items is crucial for improving service quality, particularly in recommendation systems. However, protecting users’ privacy while handling diverse sparse patterns in contexts with a large dimension sizedposes significant challenges for efficient and privacy-preserving data aggregation. To address these challenges, in this paper, we propose an optimizedk-sparse vector mean estimation scheme under Local Differential Privacy (LDP), ensuring that each user’s entire set of up tokprivate values from {−1, 1} satisfies ε-LDP. Specifically, our proposed scheme employs a seed mining technique in conjunction with PRNG Randomizer, which allows users to send their data only once while enabling the server to accurately estimate any value’s mean in the domain. Our scheme achieves an asymptotically optimal error ofO( 1/ε√n), equivalent to that of a 1-sparse case, while also ensuring efficient communication costs. The communication cost remains at a minimal level ofO(1) (only 2 bytes per user’s report) for smallerkvalues and scales toO(k) for largerk, due to efficient binning strategies. Extensive experimental results confirm that our results align with theoretical expectations, demonstrating that our scheme not only preserves user privacy but also ensures higher accuracy compared to other schemes.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
IEEE Trans. Inf. Forensics Secur.4
2024 A Communication-efficient Conjunctive Query Scheme under Local Differential Privacy
abstract
Crowdsourcing has become a widely used method for data collection and analysis, yet its privacy remains a challenge. In this paper, we present a new efficient and privacy-preserving conjunctive query scheme for crowdsourcing scenarios. The scheme employs the Local Differential Privacy (LDP) technique to ensure both query privacy and high communication efficiency. Specifically, when an aggregator launches a conjunctive query to a set of crowdsourcing users, the query condition will not be leaked. To respond the query, each user just needs to return one bit back to the aggregator. By integrating prefix encoding technique, our proposed scheme can also efficiently support conjunctive queries with one range query condition. Detailed security analysis shows our proposed scheme can achieve desirable security requirements. In addition, performance evaluations also indicate its efficiency. Furthermore, extensive experiments demonstrate our proposed scheme can achieve high accuracy while ensuring ε-LDP.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
GLOBECOM4
2024 An Efficient Range Sum Query Scheme Under Local Differential Privacy
abstract
Crowdsourcing has received considerable attention in recent years; however, privacy in crowdsourcing remains a challenge. In this paper, we present a privacy-preserving range sum query scheme under Local Differential Privacy (LDP) that not only enhances accuracy but also guarantees privacy in crowdsourcing applications. Specifically, our proposed scheme employs keyed hash, prefix encoding, and garbled bloom filter techniques to convert a large query range into a small domain, independent of the range length, thus improving accuracy. For the query response, the Optimal Unary Encoding (OUE) technique is applied to achieve ε-LDP. Security analysis shows that our proposed scheme can achieve the desirable privacy requirement for users' private items and the server's query range. In addition, performance evaluations also confirm the efficiency of our scheme in terms of computational costs and communication overhead. Furthermore, extensive experiments validate that our proposed scheme outperforms a potential strawman solution in terms of accuracy.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
ICC5
2024 Multi-Model 3D Registration: Finding Multiple Moving Objects in Cluttered Point Clouds
abstract
We investigate a variation of the 3D registration problem, named multi-model 3D registration. In the multi-model registration problem, we are given two point clouds picturing a set of objects at different poses (and possibly including points belonging to the background) and we want to simultaneously reconstruct how all objects moved between the two point clouds. This setup generalizes standard 3D registration where one wants to reconstruct a single pose, e.g., the motion of the sensor picturing a static scene. Moreover, it provides a mathematically grounded formulation for relevant robotics applications, e.g., where a depth sensor onboard a robot perceives a dynamic scene and has the goal of estimating its own motion (from the static portion of the scene) while simultaneously recovering the motion of all dynamic objects. We assume a correspondence-based setup where we have putative matches between the two point clouds and consider the practical case where these correspondences are plagued with outliers. We then propose a simple approach based on Expectation-Maximization (EM) and establish theoretical conditions under which the EM approach converges to the ground truth. We evaluate the approach in simulated and real datasets ranging from table-top scenes to self-driving scenarios and demonstrate its effectiveness when combined with state-of-the-art scene flow methods to establish dense correspondences.
David Jin, Sushrut Karmalkar, Harry Zhang, Luca Carlone
ICRA3
2024 DiffCLIP: Leveraging Stable Diffusion for Language Grounded 3D Classification
abstract
Large pre-trained models have revolutionized the field of computer vision by facilitating multi-modal learning. Notably, the CLIP model has exhibited remarkable proficiency in tasks such as image classification, object detection, and semantic segmentation. Nevertheless, its efficacy in processing 3D point clouds is restricted by the domain gap between the depth maps derived from 3D projection and the training images of CLIP.This paper introduces DiffCLIP, a novel pre-training framework that seamlessly integrates stable diffusion with ControlNet. The primary objective of DiffCLIP is to bridge the domain gap inherent in the visual branch. Furthermore, to address few-shot tasks in the textual branch, we incorporate a style-prompt generation module.Extensive experiments on the ModelNet10, ModelNet40, and ScanObjectNN datasets show that DiffCLIP has strong abilities for 3D understanding. By using stable diffusion and style-prompt generation, DiffCLIP achieves an accuracy of 43.2% for zero-shot classification on OBJ_BG of ScanObjectNN, which is state-of-the-art performance, and an accuracy of 82.4% for zero-shot classification on Model-Net10, which is also state-of-the-art performance.
Sitian Shen, Zilin Zhu, Linqian Fan, Harry Zhang, Xinxiao Wu
WACV4
2024 An Efficient Heap Tree-Based Range Query Scheme Under Local Differential Privacy
abstract
Crowdsourcing, which is regarded as one of the most important data collection techniques in Internet of Things (IoT) and Big Data era, has received significant attention in recent years. However, privacy concerns persist across various crowdsourcing scenarios. In this paper, aiming to address users’ privacy issues in crowdsourcing scenarios, we propose an efficient and privacy-preserving range query scheme under Local Differential Privacy (LDP) setting. Specifically, given a domain V = {0, 1, 2, ..., d – 1} where d = 2w, our proposed scheme integrates binary heap tree, prefix encoding, randomized response, and pseudo-random number generator techniques to enable each user to report only w bits as a query response, which is sufficient for a server to efficiently compute the range query result for any range [a, b] in the domain V. Security analysis demonstrates that our proposed scheme can achieve ε-LDP, effectively preserving the privacy of users’ private items. In addition to its low communication overhead, performance evaluation also indicates our proposed scheme is computationally efficient when the pre-computation is implemented at the server. Furthermore, our proposed scheme exhibits higher accuracy compared to previously reported flat and tree-based methods, especially for a large domain size d and a large range length m = b – a + 1.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
IEEE Internet Things J.4
2023 The Gift of Feedback: Improving ASR Model Quality by Learning from User Corrections Through Federated Learning
abstract
Automatic speech recognition (ASR) models are typically trained on large datasets of transcribed speech. As language evolves and new terms come into use, these models can become outdated and stale. In the context of models trained on the server but deployed on edge devices, errors may result from the mismatch between server training data and actual on-device usage. In this work, we seek to continually learn from on-device user corrections through Federated Learning (FL) to address this issue. We explore techniques to target fresh terms that the model has not previously encountered, learn long-tail words, and mitigate catastrophic forgetting. In experimental evaluations, we find that the proposed techniques improve model recognition of fresh terms, while preserving quality on the overall language distribution.
Lillian Zhou, Mingqing Chen, Harry Zhang, Rohit Prabhavalkar, Dhruv Guliani, Giovanni Motta, Rajiv Mathews
ASRU4
2023 ERQ: An Efficient Range Query Scheme Under Local Differential Privacy
abstract
Crowdsourcing has recently become a popular method of outsourcing tasks to many individuals in data-oriented applications. However, privacy is still a significant concern as crowdsourcing relies on individual responses. This paper focuses on range queries in crowd sourcing scenarios and proposes an efficient range query scheme, called ERQ, under Local Differential Privacy (LDP). ERQ is characterized by employing i) the accumulated encoding and perturbation techniques to protect the privacy of user data, and ii) the k-anonymity technique to conceal the real endpoints of a range query. Detailed security analysis shows that ERQ can achieve the desirable privacy requirements. In addition, performance evaluation also indicates ERQ is efficient in terms of low computational costs and communication overhead, and can achieve better accuracy while preserving privacy.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
GLOBECOM4
2023 An Efficient Bloom Filter-based Range Query Scheme Under Local Differential Privacy
abstract
While crowdsourcing for data collection has become increasingly popular in data-driven applications, privacy remains a significant challenge. This paper presents an effective scheme for conducting range queries under local differential privacy (LDP) in crowdsourcing applications, which addresses the privacy challenges that arise in such scenarios. In particular, our proposed scheme utilizes Prefix Encoding (PE) and Bloom Filter (BF) techniques to convert a large domain into a binary domain for improved query accuracy. When responding to the query, individual users can check a Bloom filter to determine whether their private item is within the query range and use the Basic Randomized Response (BRR) technique to perturb their result for achieving ε-LDP. Detailed security analysis shows that our proposed scheme can preserve user’s item privacy and also keep an external passive attacker from learning the query range. In addition, performance evaluation shows that the proposed scheme is efficient in terms of computational cost and communication overhead, while effectively balancing range query accuracy and privacy.
Ellen Z. Zhang, Yunguo Guan, Rongxing Lu, Harry Zhang
PIMRC4
2022 Enabling On-Device Training of Speech Recognition Models With Federated Dropout
abstract
Federated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients’ devices. These costs are strongly correlated with the size of the model being trained, and are significant for state-of-the-art automatic speech recognition models.We propose using federated dropout to reduce the size of client models while training a full-size model server-side. We provide empirical evidence of the effectiveness of federated dropout, and propose a novel approach to vary the dropout rate applied at each layer. Furthermore, we find that federated dropout enables a set of smaller sub-models within the larger model to independently have low word error rates, making it easier to dynamically adjust the size of the model deployed for inference.
Dhruv Guliani, Lillian Zhou, Changwan Ryu, Tien-Ju Yang, Harry Zhang, Yonghui Xiao, Françoise Beaufays, Giovanni Motta
ICASSP5
2022 The Unmet Data Visualization Needs of Decision Makers Within Organizations
abstract
When an organization chooses one course of action over alternatives, this task typically falls on a decision maker with relevant knowledge, experience, and understanding of context. Decision makers rely on data analysis, which is either delegated to analysts, or done on their own. Often the decision maker combines data, likely uncertain or incomplete, with non-formalized knowledge within a multi-objective problem space, weighing the recommendations of analysts within broader contexts and goals. As most past research in visual analytics has focused on understanding the needs and challenges of data analysts, less is known about the tasks and challenges of organizational decision makers, and how visualization support tools might help. Here we characterize the decision maker as a domain expert, review relevant literature in management theories, and report the results of an empirical survey and interviews with people who make organizational decisions. We identify challenges and opportunities for novel visualization tools, including trade-off overviews, scenario-based analysis, interrogation tools, flexible data input and collaboration support. Our findings stress the need to expand visualization design beyond data analysis into tools for information management.
Evanthia Dimara, Harry Zhang, Melanie Tory, Steven Franconeri
IEEE Trans. Vis. Comput. Graph.2
2021 Robots of the Lost Arc: Self-Supervised Learning to Dynamically Manipulate Fixed-Endpoint Cables
abstract
We explore how high-speed robot arm motions can dynamically manipulate ropes and cables to vault over obstacles, knock objects from pedestals, and weave between obstacles. In this paper, we propose a self-supervised learning framework that enables a UR5 robot to perform these three tasks. The framework finds a 3D apex point for the robot arm, which, together with a task-specific trajectory function, defines an arcing motion that dynamically manipulates the cable to perform a task with varying obstacle and target locations. The trajectory function computes minimum-jerk motions that are constrained to remain within joint limits and to travel through the 3D apex point by repeatedly solving quadratic programs to find the shortest and fastest feasible motion. We experiment with 5 physical cables with different thickness and mass and compare performance against two baselines in which a human chooses the apex point. Results suggest that a baseline with a fixed apex across the three tasks achieves respective success rates of 51.7 %, 36.7 %, and 15.0 %, and a baseline with human-specified, task-specific apex points achieves 66.7 %, 56.7 %, and 15.0 % success rate respectively, while the robot using the learned apex point can achieve success rates of 81.7 % in vaulting, 65.0 % in knocking, and 60.0 % in weaving. Code, data, and supplementary materials are available at https://sites.google.com/berkeley.edu/dynrope/home.
Harry Zhang, Jeffrey Ichnowski, Daniel Seita, Kenneth Y. Goldberg
ICRA1
2020 Dex-Net AR: Distributed Deep Grasp Planning Using a Commodity Cellphone and Augmented Reality App
abstract
Consumer demand for augmented reality (AR) in mobile phone applications, such as the Apple ARKit. Such applications have potential to expand access to robot grasp planning systems such as Dex-Net. AR apps use structure from motion methods to compute a point cloud from a sequence of RGB images taken by the camera as it is moved around an object. However, the resulting point clouds are often noisy due to estimation errors. We present a distributed pipeline, Dex-Net AR, that allows point clouds to be uploaded to a server in our lab, cleaned, and evaluated by Dex-Net grasp planner to generate a grasp axis that is returned and displayed as an overlay on the object. We implement Dex-Net AR using the iPhone and ARKit and compare results with those generated with high-performance depth sensors. The success rates with AR on harder adversarial objects are higher than traditional depth images. The server URL is https://sites.google.com/berkeley.edu/dex-net-ar/home.
Harry Zhang, Jeffrey Ichnowski, Yahav Avigal, Joseph Gonzalez 0001, Ion Stoica, Kenneth Y. Goldberg
ICRA1
2019 Personalization of End-to-End Speech Recognition on Mobile Devices for Named Entities
abstract
We study the effectiveness of several techniques to personalize end-to-end speech models and improve the recognition of proper names relevant to the user. These techniques differ in the amounts of user effort required to provide supervision, and are evaluated on how they impact speech recognition performance. We propose using keyword-dependent precision and recall metrics to measure vocabulary acquisition performance. We evaluate the algorithms on a dataset that we designed to contain names of persons that are difficult to recognize. Therefore, the baseline recall rate for proper names in this dataset is very low: 2.4%. A data synthesis approach we developed brings it to 48.6%, with no need for speech input from the user. With speech input, if the user corrects only the names, the name recall rate improves to 64.4%. If the user corrects all the recognition errors, we achieve the best recall of 73.5%. To eliminate the need to upload user data and store personalized models on a server, we focus on performing the entire personalization workflow on a mobile device.
Khe Chai Sim, Leif Johnson, Giovanni Motta, Lillian Zhou, Françoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang
ASRU12
2017 A weighted-resampling based transfer learning algorithm
abstract
Transfer learning has attracted more and more attention, and many scholars proposed some useful strategies. Boosting is the main strategy for transfer learning. In boosting, resampling is preferred over reweighting, and it can be applied to any base learner. In this paper, we propose a weighted-resampling method for transfer learning, called TrResampling. Firstly, resampling is applied to the data with heaven weight in the source domain, and the resampled data is used with the target data as the training data to build a classifier. Then the TrAdaBoost algorithm is used to adjust the weights of source data and target data. We discuss Decision Tree, Naive Bayes, and SVM as the base learner in TrResampling, and choose the suitable for TrResampling. In order to illustrate the performance of the proposed algorithm, we compare TrResampling with the state-of-the-art algorithm TrAdaBoost and the base learner Decision Tree, experimental results on UCI data sets indicate that TrResampling is superior to TrAdaBoost and Decision Tree on many data sets.
Xiaobo Liu 0001, Zhentao Liu 0001, Guangjun Wang, Zhihua Cai, Harry Zhang
IJCNN5
2016 Improving Deep Belief Networks via Delta Rule for Sentiment Classification
abstract
Sentiment classification has received much attention in both engineering and academic fields. Deep belief networks (DBN) has proved powerful in many domains including natural language processing. In this paper, DBN is applied in sentiment classification, while we propose a new way to improve the DBN based on the unsupervised training phase of restricted Boltzmann machines (RBM). That is, the RBM generates the hidden layer in an unsupervised fashion, and then we use this hidden layer as the output of a single-layer neural network, which is trained using the delta rule. The new weights trained from delta rule are then transmitted into the whole back propagation. This way keeps much more correction signal information for each layer in back propagation compared to that in the same network structure. Consequently, our experimental results demonstrate that the new learning method performs relatively better on ten sentiment datasets, which further proves the delta rule improves DBN performance for natural language processing tasks.
Harry Zhang, Donglei Du
ICTAI2
2014 A Novel Distance Function: frequency difference Metric
abstract
A high quality distance function that measures the difference between instances is essential in many real-world applications and research fields. For example, in instance-based learning, the distance function plays the most important role. A large number of distance functions have been proposed. For nominal attributes, Value Difference Metric (VDM) is one of the state-of-the-art and widely used distance functions. However, it needs to estimate the conditional probabilities, which drops its efficiency in computing the distance between instances. Besides, a practical issue that arises in estimating the conditional probabilities is that the denominators can be zero or very small. This makes them either undefined or very large. Therefore, an efficient distance function that can measure the difference between two instances but without the practical issue confronting VDM is desirable. In this paper, we propose a novel distance function: Frequency Difference Metric (FDM). FDM is just based on the joint frequencies of class labels and attribute values, instead of the conditional probabilities. Extensive empirical studies show that FDM performs almost as well as VDM in terms of accuracy, but significantly outperforms VDM in terms of efficiency. This work provides a very simple, efficient, and effective distance function that can be widely used in many real-world applications and research fields.
Liangxiao Jiang, Chaoqun Li 0001, Harry Zhang, Zhihua Cai
Int. J. Pattern Recognit. Artif. Intell.3
2013 Sampled Bayesian Network Classifiers for Class-Imbalance and Cost-Sensitive Learning
abstract
In many real-world applications, it is often the case that the class distribution of instances is imbalanced and the costs of misclassification are different. Thus, class-imbalance and cost-sensitive learning have attracted much attention from researchers. Sampling is one of the widely used approaches in dealing with the class imbalance problem, which alters the class distribution of instances so that the minority class is well represented in the training data. In this paper, we study the effect of sampling the natural training data on state-of-the-art Bayesian network classifiers, such as Naive Bayes (NB), Tree Augmented Naïve Bayes (TAN), Averaged One-Dependence Estimators (AODE), Weighted Average of One-Dependence Estimators (WAODE), and Hidden naive Bayes (HNB) and propose sampled Bayesian network classifiers. Our experimental results on a large number of UCI datasets show that our sampled Bayesian network classifiers perform much better than the ones trained from the natural training data especially when the natural training data is highly imbalanced and the cost ratio is high enough.
Liangxiao Jiang, Chaoqun Li 0001, Zhihua Cai, Harry Zhang
ICTAI4
2013 Naive Bayes text classifiers: a locally weighted learning approach
abstract
Due to being fast, easy to implement and relatively effective, some state-of-the-art naive Bayes text classifiers with the strong assumption of conditional independence among attributes, such as multinomial naive Bayes, complement naive Bayes and the one-versus-all-but-one model, have received a great deal of attention from researchers in the domain of text classification. In this article, we revisit these naive Bayes text classifiers and empirically compare their classification performance on a large number of widely used text classification benchmark datasets. Then, we propose a locally weighted learning approach to these naive Bayes text classifiers. We call our new approach locally weighted naive Bayes text classifiers (LWNBTC). LWNBTC weakens the attribute conditional independence assumption made by these naive Bayes text classifiers by applying the locally weighted learning approach. The experimental results show that our locally weighted versions significantly outperform these state-of-the-art naive Bayes text classifiers in terms of classification accuracy.
Liangxiao Jiang, Zhihua Cai, Harry Zhang, Dianhong Wang
J. Exp. Theor. Artif. Intell.3
2012 A Tri-training Based Transfer Learning Algorithm
abstract
The lack of labeled training data is a common issue in many machine learning applications. Semi-supervised learning addresses this issue by self-labeling unlabelled examples. Transfer learning tackles it from a different way: borrow labeled examples from a different but related domain (source domain) by assigning weights to those examples based on their suitability on the new domain (target domain). However, it is quite challenging to figure out the suitability. In this paper, we propose a different way for utilizing the labeled examples from source domain. That is, we use them only for labelling the unlabelled examples in the target domain. In this self-labelling, we use the idea of Tri-training. We call our new algorithm: TriTransfer. In TriTransfer, three initial classifiers are generated from the source data and the originally labeled data in the target domain, and an unlabeled example is labeled and added to the labeled data for a classifier if other two classifiers agree on its label. After an expanded labeled data set is obtained, we re-train the classifier. We repeat this process until no more change can be made. At the end, the final classifier, which is a weighted combination of the three classifiers, is output. We conduct an extensive empirical study on 34 UCI datasets, which shows that TriTransfer performs better than the state-of-art algorithms Transfer Boost, Tritraining, and NaiveBayes.
Xiaobo Liu 0001, Harry Zhang, Zhihua Cai, Guangjun Wang
ICTAI2
2012 Not so greedy: Randomly Selected Naive Bayes
Liangxiao Jiang, Zhihua Cai, Harry Zhang, Dianhong Wang
Expert Syst. Appl.3
2012 Weighted average of one-dependence estimators†
abstract
Naive Bayes (NB) is a probability-based classification model which is based on the attribute independence assumption. However, in many real-world data mining applications, its attribute independence assumption is often violated. Responding to this fact, researchers have made a substantial amount of effort to improve the classification accuracy of NB by weakening its attribute independence assumption. For a recent example, averaged one-dependence estimators (AODE) is proposed, which weakens its attribute independence assumption by averaging all models from a restricted class of one-dependence classifiers. However, all one-dependence classifiers in AODE have same weights and are treated equally. According to our observation, different one-dependence classifiers should have different weights. Therefore, in this article, we proposed an improved model called weighted average of one-dependence estimators (WAODE) by assigning different weights to these one-dependence classifiers. In our WAODE, four different weighting approaches are designed and thus four different versions are created. For simplicity, we respectively denote them by WAODE-MI, WAODE-ACC, WAODE-CLL and WAODE-AUC. The experimental results on a large number of UCI datasets published on the main website of Weka platform show that our WAODE significantly outperform AODE.
Liangxiao Jiang, Harry Zhang, Zhihua Cai, Dianhong Wang
J. Exp. Theor. Artif. Intell.2
2012 Improving Tree augmented Naive Bayes for class probability estimation
Liangxiao Jiang, Zhihua Cai, Dianhong Wang, Harry Zhang
Knowl. Based Syst.4
2011 The Unsymmetrical-Style Co-training
Harry Zhang, Bruce Spencer
PAKDD (1)2
2010 Experience, adjustment, and engagement: the role of video in law enforcement
abstract
Questions about the effectiveness of increasingly ubiquitous video technology in law enforcement have prompted an examination of the practices surrounding this technology. We present the results of a multi-site study aimed at understanding the use of video in several phases of law enforcement, from crime prevention and response to investigation and prosecution. Our findings show that while video has provided numerous benefits to law enforcement agencies, in many cases the technology either fails to support key facets of work or introduces new tasks that present an additional burden to workers. We discuss the need to incorporate human experience and tacit knowledge, operator engagement, and the greater ecosystem of work into video deployments.
Joe Tullio, Elaine M. Huang, David Wheatley, Harry Zhang, Claudia V. S. Guerrero, Amruta Tamdoo
CHI4
2010 An Extensive Empirical Study on Semi-supervised Learning
abstract
Semi-supervised classification methods utilize unlabeled data to help learn better classifiers, when only a small amount of labeled data is available. Many semi-supervised learning methods have been proposed in the past decade. However, some questions have not been well answered, e.g., whether semi-supervised learning methods outperform base classifiers learned only from the labeled data, when different base classifiers are used, whether selecting unlabeled data with efforts is superior to random selection, and how the quality of the learned classifier changes at each iteration of learning process. This paper conducts an extensive empirical study on the performance of several commonly used semi-supervised learning methods when different Bayesian classifiers (NB, NBTree, TAN, HGC, HNB, and DNB) are used as the base classifier, respectively. Results on Transductive SVM and a graph-based semi-supervised learning method LLGC are also studied for comparison. The experimental results on 26 UCI datasets and 6 widely used benchmark datasets show that these semi-supervised learning methods generally do not obtain better performance than classifiers learned only from the labeled data. Moreover, for standard self-training and co-training, when selecting the most confident unlabeled instances during learning process, the performance is not necessarily better than that of random selection of unlabeled instances. We especially discovered interesting outcomes when drawing learning curves for using NB in self-training on some UCI datasets. The accuracy of the learned classifier on the testing set may fluctuate or decrease as more unlabeled instances are used. Also on the mushroom dataset, even when all the selected unlabeled instances are correctly labeled in each iteration, the accuracy on the testing set still goes down.
Xiaoda Niu, Harry Zhang
ICDM3
2010 Usability evaluation of beep-to-the-box
abstract
Radio Frequency Identification (RFID) provides various opportunities to increase the productivity of retail business. In this paper, we describe a usability evaluation study for an RFID-based location tracking application, called Beep-To-The-Box (BTTB). The experiment was conducted in a simulated retail store to gain in-depth understanding of the usefulness and usability of the prototype in determining visual and audio user interface features. We describe the features of the BTTB, report the experimental results, and discuss insights gained to provide design recommendations for the final product design.
Young Seok Lee, Santosh Basapur, Harry Zhang, Claudia V. S. Guerrero, Noel Massey
Mobile HCI3
2010 Learning Decision Trees with log Conditional Likelihood
abstract
In machine learning and data mining, traditional learning models aim for high classification accuracy. However, accurate class probability prediction is more desirable than classification accuracy in many practical applications, such as medical diagnosis. Although it is known that decision trees can be adapted to be class probability estimators in a variety of approaches, and the resulting models are uniformly called Probability Estimation Trees (PETs), the performances of these PETs in class probability estimation, have not yet been investigated. We begin our research by empirically studying PETs in terms of class probability estimation, measured by Log Conditional Likelihood (LCL). We also compare a PET called C4.4 with other representative models, including Naïve Bayes, Naïve Bayes Tree, Bayesian Network, KNN and SVM, in LCL. From our experiments, we draw several valuable conclusions. First, among various tree-based models, C4.4 is the best in yielding precise class probability prediction measured by LCL. We provide an explanation for this and reveal the nature of LCL. Second, compared with non tree-based models, C4.4 also performs best. Finally, LCL does not dominate another well-established relevant metric — AUC, which suggests that different decision-tree learning models should be used for different objectives. Our experiments are conducted on the basis of 36 UCI sample sets. We run all the models within a machine learning platform — Weka. We also explore an approach to improve the class probability estimation of Naïve Bayes Tree. We propose a greedy and recursive learning algorithm, where at each step, LCL is used as the scoring function to expand the decision tree. The algorithm uses Naïve Bayes created at leaves to estimate class probabilities of test samples. The whole tree encodes the posterior class probability in its structure. One benefit of improving class probability estimation is that both classification accuracy and AUC can be possibly scaled up. We call the new model LCL Tree (LCLT). Our experiments on 33 UCI sample sets show that LCLT outperforms all state-of-the-art learning models, such as Naïve Bayes Tree, significantly in accurate class probability prediction measured by LCL, as well as in classification accuracy and AUC.
Yuhong Yan, Harry Zhang
Int. J. Pattern Recognit. Artif. Intell.3
2009 A Novel Bayes Model: Hidden Naive Bayes
abstract
Because learning an optimal Bayesian network classifier is an NP-hard problem, learning-improved naive Bayes has attracted much attention from researchers. In this paper, we summarize the existing improved algorithms and propose a novel Bayes model: hidden naive Bayes (HNB). In HNB, a hidden parent is created for each attribute which combines the influences from all other attributes. We experimentally test HNB in terms of classification accuracy, using the 36 UCI data sets selected by Weka, and compare it to naive Bayes (NB), selective Bayesian classifiers (SBC), naive Bayes tree (NBTree), tree-augmented naive Bayes (TAN), and averaged one-dependence estimators (AODE). The experimental results show that HNB significantly outperforms NB, SBC, NBTree, TAN, and AODE. In many data mining applications, an accurate class probability estimation and ranking are also desirable. We study the class probability estimation and ranking performance, measured by conditional log likelihood (CLL) and the area under the ROC curve (AUC), respectively, of naive Bayes and its improved models, such as SBC, NBTree, TAN, and AODE, and then compare HNB to them in terms of CLL and AUC. Our experiments show that HNB also significantly outperforms all of them.
Liangxiao Jiang, Harry Zhang, Zhihua Cai
IEEE Trans. Knowl. Data Eng.2
2008 Switching among Non-Weighting, Clause Weighting, and Variable Weighting in Local Search for SAT
Wanxia Wei, Chu Min Li 0001, Harry Zhang
CP3
2008 Discriminative parameter learning for Bayesian networks
abstract
Bayesian network classifiers have been widely used for classification problems. Given a fixed Bayesian network structure, parameters learning can take two different approaches: generative and discriminative learning. While generative parameter learning is more efficient, discriminative parameter learning is more effective. In this paper, we propose a simple, efficient, and effective discriminative parameter learning method, called Discriminative Frequency Estimate (DFE), which learns parameters by discriminatively computing frequencies from data. Empirical studies show that the DFE algorithm integrates the advantages of both generative and discriminative learning: it performs as well as the state-of-the-art discriminative parameter learning method ELR in accuracy, but is significantly more efficient.
Jiang Su, Harry Zhang, Charles Ling 0001, Stan Matwin
ICML2
2008 Proper Model Selection with Significance Test
Charles Ling 0001, Harry Zhang, Stan Matwin
ECML/PKDD (1)3
2008 Using Instance cloning to Improve Naive Bayes for Ranking
abstract
Improving naive Bayes (simply NB)15,28 for classification has received significant attention. Related work can be broadly divided into two approaches: eager learning and lazy learning.1 Different from eager learning, the key idea for extending naive Bayes using lazy learning is to learn an improved naive Bayes for each test instance. In recent years, several lazy extensions of naive Bayes have been proposed. For example, LBR,30 SNNB,27 and LWNB.8 All these algorithms aim to improve naive Bayes' classification performance. Indeed, they achieve significant improvement in terms of classification, measured by accuracy. In many real-world data mining applications, however, an accurate ranking is more desirable than an accurate classification. Thus a natural question is whether they also achieve significant improvement in terms of ranking, measured by AUC (the area under the ROC curve).2,11,17 Responding to this question, we conduct experiments on the 36 UCI data sets18 selected by Weka12 to investigate their ranking performance and find that they do not significantly improve the ranking performance of naive Bayes. Aiming at scaling up naive Bayes' ranking performance, we present a novel lazy method ICNB (instance cloned naive Bayes) and develop three ICNB algorithms using different instance cloning strategies. We empirically compare them with naive Bayes. The experimental results show that our algorithms achieve significant improvement in terms of AUC. Our research provides a simple but effective method for the applications where an accurate ranking is desirable.
Liangxiao Jiang, Dianhong Wang, Harry Zhang, Zhihua Cai, Bo Huang 0005
Int. J. Pattern Recognit. Artif. Intell.3
2008 Naive Bayes for optimal ranking
abstract
It is well known that naive Bayes performs surprisingly well in classification, but its probability estimation is poor. AUC (the area under the receiver operating characteristics curve) is a measure different from classification accuracy and probability estimation, which is often used to measure the quality of rankings. Indeed, an accurate ranking of examples is often more desirable than a mere classification. What is the general performance of naive Bayes in yielding optimal ranking, measured by AUC? In this paper, we study it systematically by both empirical experiments and theoretical analysis. In our experiments, we compare naive Bayes with a state-of-the-art decision-tree learning algorithm C4.4 for ranking, and some popular extensions of naive Bayes which achieve a significant improvement over naive Bayes in classification, such as the selective Bayesian classifier (SBC) and tree-augmented naive Bayes (TAN). Our experimental results show that naive Bayes performs significantly better than C4.4 and comparably with TAN. This provides empirical evidence that naive Bayes performs well in ranking. Then we analyse theoretically the optimality of naive Bayes in ranking. We study two example problems: conjunctive concepts and m-of-n concepts, which have been used in analysing the performance of naive Bayes in classification. Surprisingly, naive Bayes performs optimally on them in ranking, even though it does not in classification. We present and prove a sufficient condition for the optimality of naive Bayes in ranking. From both empirical and theoretical studies, we believe that naive Bayes is a competitive model for ranking.
Harry Zhang, Jiang Su
J. Exp. Theor. Artif. Intell.1
2007 Learning Locally Weighted C4.4 for Class Probability Estimation
Liangxiao Jiang, Harry Zhang, Dianhong Wang, Zhihua Cai
Discovery Science2
2007 Naturalistic use of cell phones in driving and context-based user assistance
abstract
A field study has been conducted to investigate the naturalistic use of cell phone applications in driving, home, work, and school and during daytime and nighttime. GPS coordinates are used to determine whether cell phone users are driving. The frequency and duration of use of various cell phone applications such as incoming or outgoing voice calls, music player, calendar, SMS, camera, and the Internet are analyzed separately for driving and non-driving. The present results provide fundamental data for adequately assessing the distraction potential of mobile devices and guiding the design of context-based assistance systems.
Harry Zhang, Christopher Schreiner, Keshu Zhang, Kari Torkkola
Mobile HCI1
2007 Combining Adaptive Noise and Look-Ahead in Local Search for SAT
Chu Min Li 0001, Wanxia Wei, Harry Zhang
SAT3
2006 A Fast Decision Tree Learning Algorithm
Jiang Su, Harry Zhang
AAAI2
2006 Improving the Ranking Performance of Decision Trees
Harry Zhang
ECML2
2006 Full Bayesian network classifiers
abstract
The structure of a Bayesian network (BN) encodes variable independence. Learning the structure of a BN, however, is typically of high computational complexity. In this paper, we explore and represent variable independence in learning conditional probability tables (CPTs), instead of in learning structure. A full Bayesian network is used as the structure and a decision tree is learned for each CPT. The resulting model is called full Bayesian network classifiers (FBCs). In learning an FBC, learning the decision trees for CPTs captures essentially both variable independence and context-specific independence. We present a novel, efficient decision tree learning, which is also effective in the context of FBC learning. In our experiments, the FBC learning algorithm demonstrates better performance in both classification and ranking compared with other state-of-the-art learning algorithms. In addition, its reduced effort on structure learning makes its time complexity quite low as well.
Jiang Su, Harry Zhang
ICML2
2006 Decision Trees for Probability Estimation: An Empirical Study
abstract
Accurate probability estimation generated by learning models is desirable in some practical applications, such as medical diagnosis. In this paper, we empirically study traditional decision-tree learning models and their variants in terms of probability estimation, measured by conditional log likelihood (CLL). Furthermore, we also compare decision tree learning with other kinds of representative learning: Naive Bayes, Naive Bayes tree, Bayesian network, K-nearest neighbors and support vector machine with respect to probability estimation. From our experiments, we have several interesting observations. First, among various decision-tree learning models, C4.4 is the best in yielding precise probability estimation measured by CLL, although its performance is not good in terms of other evaluation criteria, such as accuracy and ranking. We provide an explanation for this and reveal the nature of CLL. Second, compared with other popular models, C4.4 achieves the best CLL. Finally, CLL does not dominate another well-established relevant measurement AUC (the area under the curve of receiver operating characteristics), which suggests that different decision-tree learning models should be used for different objectives. Our experiments are conducted on the basis of 36 UCI sample sets that cover a wide range of domains and data characteristics. We run all the models within a machine learning platform $Weka
Harry Zhang, Yuhong Yan
ICTAI2
2006 Weightily Averaged One-Dependence Estimators
Liangxiao Jiang, Harry Zhang
PRICAI2
2006 Learning probabilistic decision trees for AUC
Harry Zhang, Jiang Su
Pattern Recognit. Lett.1
2005 Representing Conditional Independence Using Decision Trees
Jiang Su, Harry Zhang
AAAI2
2005 Hidden Naive Bayes
Harry Zhang, Liangxiao Jiang, Jiang Su
AAAI1
2005 One Dependence Augmented Naive Bayes
Liangxiao Jiang, Harry Zhang, Zhihua Cai, Jiang Su
ADMA2
2005 Learning k-Nearest Neighbor Naive Bayes for Ranking
Liangxiao Jiang, Harry Zhang, Jiang Su
ADMA2
2005 Driver State Monitor from DELPHI
abstract
We present an automotive-grade, real-time, vision-based driver state monitor. Upon detecting and tracking the driver's facial features, the system analyzes eye-closures and head pose to infer his/her fatigue or distraction. This information is used to warn the driver and to modulate the actions of other safety systems. The purpose of this monitor is to increase road safety by preventing drivers from falling asleep or from being overly distracted, and to improve the effectiveness of other safety systems.
N. Edenborough, Riad I. Hammoud, A. Harbach, A. Ingold, Branislav Kisacanin, Phillip Malawey, T. Newman, G. Scharenbroch, S. Skiver, Matthew R. H. Smith, Andrew Wilhelm, Gerald J. Witt, E. Yoder, Harry Zhang
CVPR (2)14
2005 Learning Tree Augmented Naive Bayes for Ranking
Liangxiao Jiang, Harry Zhang, Zhihua Cai, Jiang Su
DASFAA2
2005 Learning Instance Greedily Cloning Naive Bayes for Ranking
abstract
Naive Bayes (simply NB) (Langley et al., 1992) has been widely used in machine learning and data mining as a simple and effective classification algorithm. Since its conditional independence assumption is rarely true, researchers have made a substantial amount of effort to improve naive Bayes. The related research work can be broadly divided into two approaches: eager learning and lazy learning, depending on when the major computation occurs. Different from eager approach, the key idea for extending naive Bayes from the lazy approach is to learn a naive Bayes for each testing example. In recent years, some lazy extensions of naive Bayes have been proposed. For example, SNNB, LWNB, and LBR. All are aiming at improving the classification accuracy of naive Bayes. In many real-world machine learning and data mining applications, however, an accurate ranking is more desirable than an accurate classification. Responding to this fact, we present a lazy learning algorithm called instance greedily cloning naive Bayes (simply IGCNB) in this paper. Our motivation is to improve naive Bayes' ranking performance measured by AUC (Bradley, 1997; Provost and Fawcett, 1997). We experimentally tested our algorithm, using the whole 36 UCI datasets recommended by Weka, and compared it to C4.4 (Provost and Domingos, 2003), NB (Langley et al., 1992), SNNB (Xie, 2002) and LWNB (Frank, 2003). The experimental results show that our algorithm outperforms all the other algorithms used to compare significantly in yielding accurate ranking.
Liangxiao Jiang, Harry Zhang
ICDM2
2005 Augmenting naive Bayes for ranking
abstract
Naive Bayes is an effective and efficient learning algorithm in classification. In many applications, however, an accurate ranking of instances based on the class probability is more desirable. Unfortunately, naive Bayes has been found to produce poor probability estimates. Numerous techniques have been proposed to extend naive Bayes for better classification accuracy, of which selective Bayesian classifiers (SBC) (Langley & Sage, 1994), tree-augmented naive Bayes (TAN) (Friedman et al., 1997), NBTree (Kohavi, 1996), boosted naive Bayes (Elkan, 1997), and AODE (Webb et al., 2005) achieve remarkable improvement over naive Bayes in terms of classification accuracy. An interesting question is: Do these techniques also produce accurate ranking? In this paper, we first conduct a systematic experimental study on their efficacy for ranking. Then, we propose a new approach to augmenting naive Bayes for generating accurate ranking, called hidden naive Bayes (HNB). In an HNB, a hidden parent is created for each attribute to represent the influences from all other attributes, and thus a more accurate ranking is expected. HNB inherits the structural simplicity of naive Bayes and can be easily learned without structure learning. Our experiments show that HNB outperforms naive Bayes, SBC, boosted naive Bayes, NBTree, and TAN significantly, and performs slightly better than AODE in ranking.
Harry Zhang, Liangxiao Jiang, Jiang Su
ICML1
2005 Exploring Conditions For The Optimality Of Naïve Bayes
abstract
Naïve Bayes is one of the most efficient and effective inductive learning algorithms for machine learning and data mining. Its competitive performance in classification is surprising, because the conditional independence assumption on which it is based is rarely true in real-world applications. An open question is: what is the true reason for the surprisingly good performance of Naïve Bayes in classification? In this paper, we propose a novel explanation for the good classification performance of Naïve Bayes. We show that, essentially, dependence distribution plays a crucial role. Here dependence distribution means how the local dependence of an attribute distributes in each class, evenly or unevenly, and how the local dependences of all attributes work together, consistently (supporting a certain classification) or inconsistently (canceling each other out). Specifically, we show that no matter how strong the dependences among attributes are, Naïve Bayes can still be optimal if the dependences distribute evenly in classes, or if the dependences cancel each other out. We propose and prove a sufficient and necessary condition for the optimality of Naïve Bayes. Further, we investigate the optimality of Naïve Bayes under the Gaussian distribution. We present and prove a sufficient condition for the optimality of Naïve Bayes, in which the dependences among attributes exist. This provides evidence that dependences may cancel each other out. Our theoretic analysis can be used in designing learning algorithms. In fact, a major class of learning algorithms for Bayesian networks are conditional independence-based (or CI-based), which are essentially based on dependence. We design a dependence distribution-based algorithm by extending the ChowLiu algorithm, a widely used CI based algorithm. Our experiments show that the new algorithm outperforms the ChowLiu algorithm, which also provides empirical evidence to support our new explanation.
Harry Zhang
Int. J. Pattern Recognit. Artif. Intell.1
2004 Naive Bayesian Classifiers for Ranking
Harry Zhang, Jiang Su
ECML1
2004 Conditional Independence Trees
Harry Zhang, Jiang Su
ECML1
2004 Learning Conditional Independence Tree for Ranking
abstract
Accurate ranking is desired in many real-world data mining applications. Traditional learning algorithms, however, aim only at high classification accuracy. It has been observed that both traditional decision trees and naive Bayes produce good classification accuracy but poor probability estimates. In this paper, we use a new model, conditional independence tree (CITree), which is a combination of decision tree and naive Bayes and more suitable for ranking and more learnable in practice. We propose a novel algorithm for learning CITree for ranking, and the experiments show that the CITree algorithm outperforms the state-of-the-art decision tree learning algorithm C4.4 and naive Bayes significantly in yielding accurate rankings. Our work provides an effective data mining algorithm for applications in which an accurate ranking is required.
Jiang Su, Harry Zhang
ICDM2
2004 Learning Weighted Naive Bayes with Accurate Ranking
abstract
Naive Bayes is one of most effective classification algorithms. In many applications, however, a ranking of examples are more desirable than just classification. How to extend naive Bayes to improve its ranking performance is an interesting and useful question in practice. Weighted naive Bayes is an extension of naive Bayes, in which attributes have different weights. This paper investigates how to learn a weighted naive Bayes with accurate ranking from data, or more precisely, how to learn the weights of a weighted naive Bayes to produce accurate ranking. We explore various methods: the gain ratio method, the hill climbing method, and the Markov chain Monte Carlo method, the hill climbing method combined with the gain ratio method, and the Markov chain Monte Carlo method combined with the gain ratio method. Our experiments show that a weighted naive Bayes trained to produce accurate ranking outperforms naive Bayes.
Harry Zhang, Shengli Sheng
ICDM1
2003 AUC: a Statistically Consistent and more Discriminating Measure than Accuracy
Charles Ling 0001, Harry Zhang
IJCAI3