EDBT 2026 Demo / reviewers in the wild / expert
Hantian Zhang
dblp:155/9999
· DBLP profile ↗
11ranked-venue papers
4as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Falcon: Fair Active Learning using Multi-armed BanditsabstractBiased data can lead to unfair machine learning models, highlighting the importance of embedding fairness at the beginning of data analysis, particularly during dataset curation and labeling. In response, we propose Falcon, a scalable fair active learning framework. Falcon adopts a data-centric approach that improves machine learning model fairness via strategic sample selection. Given a user-specified group fairness measure, Falcon identifies samples from "target groups" (e.g., (attribute=female, label=positive)) that are the most informative for improving fairness. However, a challenge arises since these target groups are defined using ground truth labels that are not available during sample selection. To handle this, we propose a novel trial-and-error method, where we postpone using a sample if the predicted label is different from the expected one and falls outside the target group. We also observe the trade-off that selecting more informative samples results in higher likelihood of postponing due to undesired label prediction, and the optimal balance varies per dataset. We capture the trade-off between informativeness and postpone rate as policies and propose to automatically select the best policy using adversarial multi-armed bandit methods, given their computational efficiency and theoretical guarantees. Experiments show that Falcon significantly outperforms existing fair active learning approaches in terms of fairness and accuracy and is more efficient. In particular, only Falcon supports a proper trade-off between accuracy and fairness where its maximum fairness score is 1.8--4.5x higher than the second-best results. Ki Hyun Tae, Hantian Zhang, Jaeyoung Park 0003, Kexin Rong 0001, Steven Euijong Whang |
Proc. VLDB Endow. | 2 |
| 2024 | BAI-Net: Individualized Anatomical Cerebral Cartography Using Graph Neural NetworkabstractBrain atlas is an important tool in the diagnosis and treatment of neurological disorders. However, due to large variations in the organizational principles of individual brains, many challenges remain in clinical applications. Brain atlas individualization network (BAI-Net) is an algorithm that subdivides individual cerebral cortex into segregated areas using brain morphology and connectomes. The presented method integrates group priors derived from a population atlas, adjusts areal probabilities using the context of connectivity fingerprints derived from the fiber-tract embedding of tractography, and provides reliable and explainable individualized brain areas across multiple sessions and scanners. We demonstrate that BAI-Net outperforms the conventional iterative clustering approach by capturing significantly heritable topographic variations in individualized cartographies. The topographic variability of BAI-Net cartographies has shown strong associations with individual variability in brain morphology, connectivity as well as higher relationship on individual cognitive behaviors and genetics. This study provides an explainable framework for individualized brain cartography that may be useful in the precise localization of neuromodulation and treatments on individual brains. Yu Zhang 0115, Hantian Zhang, Luqi Cheng, Zhengyi Yang 0002, Yuheng Lu, Weiyang Shi, Wen Li 0021, Junjie Zhuo, Jiaojian Wang, Lingzhong Fan, Tianzi Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | iFlipper: Label Flipping for Individual FairnessabstractAs machine learning becomes prevalent, mitigating any unfairness present in the training data becomes critical. Among the various notions of fairness, this paper focuses on the well-known individual fairness, which states that similar individuals should be treated similarly. While individual fairness can be improved when training a model (in-processing), we contend that fixing the data before model training (pre-processing) is a more fundamental solution. In particular, we show that label flipping is an effective pre-processing technique for improving individual fairness. Our system iFlipper solves the optimization problem of minimally flipping labels given a limit to the individual fairness violations, where a violation occurs when two similar examples in the training data have different labels. We first prove that the problem is NP-hard. We then propose an approximate linear programming algorithm and provide theoretical guarantees on how close its result is to the optimal solution in terms of the number of label flips. We also propose techniques for making the linear programming solution more optimal without exceeding the violations limit. Experiments on real datasets show that iFlipper significantly outperforms other pre-processing baselines in terms of individual fairness and accuracy on unseen test sets. In addition, iFlipper can be combined with in-processing techniques for even better results. Hantian Zhang, Ki Hyun Tae, Jaeyoung Park 0003, Xu Chu 0002, Steven Euijong Whang |
Proc. ACM Manag. Data | 1 |
| 2022 | LifeRec: A Mobile App for Lifelog Recording and Ubiquitous RecommendationabstractIn recent years, context information has played an increasingly significant role in recommendation systems. With the rapid growth of portable sensor devices, lifelog data, such as mood, location, and daily activity, has been recorded and used for ubiquitous recommendation tasks. However, since the multi-modal lifelog data contains objective context information and subjective user labeling, it is challenging to record the lifelog thoroughly and perform personalized recommendations in real-time. In this work, we design a mobile application (App), LifeRec, to record multi-modal lifelog data and perform personalized recommendations by communicating with the remote server. The App helps users collect various lifelog information (e.g., location, diet, activity, and mood) and receive real-time recommendation with privacy protection and little effort. It is useful for lifelog data collection, user status monitoring, and various ubiquitous recommendation tasks. We examine LifeRec in a one-week field study with seven subjects. The users’ experience feedback and recording results show great usability and task completeness with our App. Jiayu Li 0001, Hantian Zhang, Zhiyu He 0001, Rongwu Xu, Pingfei Wu, Min Zhang 0006, Yiqun Liu 0001, Shaoping Ma |
CHIIR | 2 |
| 2021 | OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine LearningabstractMachine learning (ML) is increasingly being used to make decisions in our society. ML models, however, can be unfair to certain demographic groups (e.g., African Americans or females) according to various fairness metrics. Existing techniques for producing fair ML models either are limited to the type of fairness constraints they can handle (e.g., preprocessing) or require nontrivial modifications to downstream ML training algorithms (e.g., in-processing). Hantian Zhang, Xu Chu 0002, Abolfazl Asudeh, Shamkant B. Navathe |
SIGMOD Conference | 1 |
| 2020 | ALEX: An Updatable Adaptive Learned IndexabstractRecent work on "learned indexes" has changed the way we look at the decades-old field of DBMS indexing. The key idea is that indexes can be thought of as "models" that predict the position of a key in a dataset. Indexes can, thus, be learned. The original work by Kraska et al. shows that a learned index beats a B+ tree by a factor of up to three in search time and by an order of magnitude in memory footprint. However, it is limited to static, read-only workloads. In this paper, we present a new learned index called ALEX which addresses practical issues that arise when implementing learned indexes for workloads that contain a mix of point lookups, short range queries, inserts, updates, and deletes. ALEX effectively combines the core insights from learned indexes with proven storage and indexing techniques to achieve high performance and low memory footprint. On read-only workloads, ALEX beats the learned index from Kraska et al. by up to 2.2X on performance with up to 15X smaller index size. Across the spectrum of read-write workloads, ALEX beats B+ trees by up to 4.1X while never performing worse, with up to 2000X smaller index size. We believe ALEX presents a key step towards making learned indexes practical for a broader class of database workloads with dynamic updates. Jialin Ding 0001, Umar Farooq Minhas, Jia Yu 0001, Chi Wang 0001, Jaeyoung Do, Yinan Li 0009, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David B. Lomet, Tim Kraska |
SIGMOD Conference | 7 |
| 2019 | Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-precision LearningabstractLearning from the data stored in a database is an important function increasingly available in relational engines. Methods using lower precision input data are of special interest given their overall higher efficiency. However, in databases, these methods have a hidden cost: the quantization of the real value into a smaller number is an expensive step. To address this issue, we present ML-Weaving, a data structure and hardware acceleration technique intended to speed up learning of generalized linear models over low precision data. MLWeaving provides a compact in-memory representation that enables the retrieval of data at any level of precision. MLWeaving also provides a highly efficient implementation of stochastic gradient descent on FPGAs and enables the dynamic tuning of precision, instead of using a fixed precision level during learning. Experimental results show that MLWeaving converges up to 16 x faster than low-precision implementations of first-order methods on CPUs. Zeke Wang, Kaan Kara, Hantian Zhang, Gustavo Alonso, Ce Zhang 0001, Onur Mutlu |
Proc. VLDB Endow. | 3 |
| 2018 | MLBench: Benchmarking Machine Learning Services Against Human ExpertsabstractModern machine learning services and systems are complicated data systems --- the process of designing such systems is an art of compromising between functionality , performance , and quality . Providing different levels of system supports for different functionalities, such as automatic feature engineering, model selection and ensemble, and hyperparameter tuning, could improve the quality, but also introduce additional cost and system complexity. In this paper, we try to facilitate the process of asking the following type of questions: How much will the users lose if we remove the support of functionality x from a machine learning service? Answering this type of questions using existing datasets, such as the UCI datasets, is challenging. The main contribution of this work is a novel dataset, MLBench, harvested from Kaggle competitions. Unlike existing datasets, MLBench contains not only the raw features for a machine learning task, but also those used by the winning teams of Kaggle competitions. The winning features serve as a baseline of best human effort that enables multiple ways to measure the quality of machine learning services that cannot be supported by existing datasets, such as relative ranking on Kaggle and relative accuracy compared with best-effort systems. We then conduct an empirical study using MLBench to understand example machine learning services from Amazon and Microsoft Azure, and showcase how MLBench enables a comparative study revealing the strength and weakness of these existing machine learning services quantitatively and systematically. The full version of this paper can be found at arxiv.org/abs/1707.09562 Yu Liu 0075, Hantian Zhang, Luyuan Zeng, Wentao Wu 0001, Ce Zhang 0001 |
Proc. VLDB Endow. | 2 |
| 2017 | How good are machine learning clouds for binary classification with good features?: extended abstractabstractIn spite of the recent advancement of machine learning research, modern machine learning systems are still far from easy to use, at least from the perspective of business users or even scientists without a computer science background. Recently, there is a trend toward pushing machine learning onto the cloud as a "service," a.k.a. machine learning clouds. By putting a set of machine learning primitives on the cloud, these services significantly raise the level of abstraction for machine learning. For example, with Amazon Machine Learning, users only need to upload the dataset and specify the type of task (classification or regression). The cloud will then train machine learning models without any user intervention. Hantian Zhang, Luyuan Zeng, Wentao Wu 0001, Ce Zhang 0001 |
SoCC | 1 |
| 2017 | Scalable inference of decision tree ensembles: Flexible design for CPU-FPGA platformsabstractDecision tree ensembles are commonly used in a wide range of applications and becoming the de facto algorithm for decision tree based classifiers. Different trees in an ensemble can be processed in parallel during tree inference, making them a suitable use case for FPGAs. Large tree ensembles, however, require careful mapping of trees to on-chip memory and management of memory accesses. As a result, existing FPGA solutions suffer from the inability to scale beyond tens of trees and lack the flexibility to support different tree ensembles. In this paper we present an FPGA tree ensemble classifier together with a software driver to efficiently manage the FPGA's memory resources. The classifier architecture efficiently utilizes the FPGA's resources to fit half a million tree nodes in on-chip memory, delivering up to 20× speedup over a 10-threaded CPU implementation when fully processing the tree ensemble on the FPGA. It can also combine the CPU and FPGA to scale to tree ensembles that do not fit in on-chip memory, achieving up to an order of magnitude speedup compared to a pure CPU implementation. In addition, the classifier architecture can be programmed at runtime to process varying tree ensemble sizes. Muhsen Owaida, Hantian Zhang, Ce Zhang 0001, Gustavo Alonso |
FPL | 2 |
| 2017 | ZipML: Training Linear Models with End-to-End Low Precision, and a Little Bit of Deep LearningabstractRecently there has been significant interest in training machine-learning models at low precision: by reducing precision, one can reduce computation and communication by one order of magnitude. We examine training at reduced precision, both from a theoretical and practical perspective, and ask: is it possible to train models at end-to-end low precision with provable guarantees? Can this lead to consistent order-of-magnitude speedups? We mainly focus on linear models, and the answer is yes for linear models. We develop a simple framework called ZipML based on one simple but novel strategy called double sampling. Our ZipML framework is able to execute training at low precision with no bias, guaranteeing convergence, whereas naive quantization would introduce significant bias. We validate our framework across a range of applications, and show that it enables an FPGA prototype that is up to $6.5\times$ faster than an implementation using full 32-bit precision. We further develop a variance-optimal stochastic quantization strategy and show that it can make a significant difference in a variety of settings. When applied to linear models together with double sampling, we save up to another $1.7\times$ in data movement compared with uniform quantization. When training deep networks with quantized models, we achieve higher accuracy than the state-of-the-art XNOR-Net. Hantian Zhang, Jerry Li 0001, Kaan Kara, Dan Alistarh, Ji Liu 0002, Ce Zhang 0001 |
ICML | 1 |