Haozhe An

dblp:263/7358 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 On the Mutual Influence of Gender and Occupation in LLM Representations
abstract
We examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names in LLMs influence each other mutually.We find that LLMs' first-name gender representations correlate with real-world gender statistics associated with the name, and are influenced by the co-occurrence of stereotypically feminine or masculine occupations.Additionally, we study the influence of firstname gender representations on LLMs in a downstream occupation prediction task and their potential as an internal metric to identify extrinsic model biases.While feminine firstname embeddings often raise the probabilities for female-dominated jobs (and vice versa for male-dominated jobs), reliably using these internal gender representations for bias detection remains challenging.Is Jody male or female? Female MaleJody is a nurse.Is Jody male or female? Female MaleJody is a comedian.Is Jody male or female? Female Male
Haozhe An, Connor Baumler, Abhilasha Sancheti, Rachel Rudinger
ACL (1)1
2024 Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the US
abstract
Recent work has highlighted the culturallycontingent nature of commonsense knowledge (Shen et al., 2024).We introduce AMAMMERE (/A:.mA:.mu:.reI/),from the Akan word meaning 'culture.'This test set of 525 multiple-choice questions is designed to evaluate the commonsense knowledge of English LLMs, relative to the cultural contexts of Ghana and the United States.To create AMAMMERE, we select a set of multiplechoice questions (MCQs) from existing commonsense datasets and rewrite them in a multistage process involving surveys of Ghanaian and U.S. participants.In three rounds of surveys, participants from both pools are solicited to ( 1) write correct and incorrect answer choices, ( 2) rate individual answer choices on a 5-point Likert scale, and (3) select the best answer choice from the newly-constructed MCQ items, in a final validation step.By engaging participants at multiple stages, our procedure ensures that participant perspectives are incorporated both in the creation and validation of test items, resulting in high levels of agreement within each pool.We evaluate several off-the-shelf English LLMs on AMAMMERE. 1 Uniformly, models prefer answers choices that align with the preferences of U.S. annotators over Ghanaian annotators.Additionally, when test items specify a cultural context (Ghana or the U.S.), models exhibit some ability to adapt, but performance is consistently better in U.S. contexts than Ghanaian.As large resources are devoted to the advancement of English LLMs, our findings underscore the need for culturally adaptable models and evaluations to meet the needs of diverse English-speaking populations around the world.
Christabel Acquaye, Haozhe An, Rachel Rudinger
EMNLP2
2024 On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
abstract
We study the presence of heteronormative biases and prejudice against interracial romantic relationships in large language models by performing controlled name-replacement experiments for the task of relationship prediction.We show that models are less likely to predict romantic relationships for (a) same-gender character pairs than different-gender pairs; and (b) intra/inter-racial character pairs involving Asian names as compared to Black, Hispanic, or White names.We examine the contextualized embeddings of first names and find that gender for Asian names is less discernible than non-Asian names.We discuss the social implications of our findings, underlining the need to prioritize the development of inclusive and equitable technology.2 Except for Hispanic wherein we did not get any names in 5 -10% bin and only 1 name in 25 -50% bin.480 0 -2 2 -5 5 -1 0 1 0 -2 5 2 5 -5 0 5 0 -7 5 7 5 -9 0 9 0 -9 5 9 5 -9 8 9 8 -1 0 0 % Female 0-2 2-5 5-10 10-25 25-50 50-75 75-90 90-95 95-98 98-100 % Female 0.56 0.55 0.59 0.58 0.61 0.59 0.69 0.68 0.64 0.72 0.56 0.52 0.57 0.58 0.59 0.58 0.65 0.66 0.63 0.69 0.61 0.58 0.63 0.60 0.64 0.62 0.70 0.69 0.67 0.72 0.58 0.57 0.59 0.60 0.62 0.59 0.67 0.66 0.63 0.70 0.61 0.60 0.64 0.62 0.65 0.62 0.69 0.69 0.66 0.69 0.59 0.56 0.61 0.60 0.61 0.60 0.66 0.64 0.62 0.67 0.66 0.64 0.67 0.66 0.67 0.64 0.69 0.66 0.65 0.67 0.67 0.65 0.68 0.66 0.68 0.64 0.67 0.67 0.64 0.67 0.64 0.61 0.65 0.63 0.65 0.61 0.65 0.63 0.62 0.65 0.70 0.66 0.70 0.68 0.68 0.66 0.68 0.67 0.65 0.68 Male Neutral Female Male Neutral Female Asian (Recall) 0 -2 2 -5 5 -1 0 1 0 -2 5 2 5 -5 0 5 0 -7 5 7 5 -9 0 9 0 -9 5 9 5 -9 8 9 8 -1 0 0
Abhilasha Sancheti, Haozhe An, Rachel Rudinger
EMNLP2
2023 SODAPOP: Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models
abstract
A common limitation of diagnostic tests for detecting social biases in NLP models is that they may only detect stereotypic associations that are pre-specified by the designer of the test.Since enumerating all possible problematic associations is infeasible, it is likely these tests fail to detect biases that are present in a model but not pre-specified by the designer.To address this limitation, we propose SODAPOP 1 (SOcial bias Discovery from Answers about PeOPle), an approach for automatic social bias discovery in social commonsense question-answering.The SODAPOP pipeline generates modified instances from the Social IQa dataset (Sap et al., 2019b) by ( 1) substituting names associated with different demographic groups, and (2) generating many distractor answers from a masked language model.By using a social commonsense model to score the generated distractors, we are able to uncover the model's stereotypic associations between demographic groups and an open set of words.We also test SODAPOP on debiased models and show the limitations of multiple state-of-the-art debiasing algorithms.
Haozhe An, Zongxia Li, Jieyu Zhao 0001, Rachel Rudinger
EACL1
2022 The Fine-Grained Complexity of Multi-Dimensional Ordering Properties
Haozhe An, Mohit Gurumukhani, Russell Impagliazzo, Michael Jaber, Marvin Künnemann, Maria Paula Parga Nina
Algorithmica1
2022 Exploring the common principal subspace of deep features in neural networks
Haoyi Xiong, Yaqing Wang 0002, Haozhe An, Dejing Dou, Dongrui Wu
Mach. Learn.4
2022 COLAM: Co-Learning of Deep Neural Networks and Soft Labels via Alternating Minimization
Xingjian Li 0002, Haoyi Xiong, Haozhe An, Cheng-Zhong Xu 0001, Dejing Dou
Neural Process. Lett.3
2021 The Fine-Grained Complexity of Multi-Dimensional Ordering Properties
abstract
We define a class of problems whose input is an n-sized set of d-dimensional vectors, and where the problem is first-order definable using comparisons between coordinates. This class captures a wide variety of tasks, such as complex types of orthogonal range search, model-checking first-order properties on geometric intersection graphs, and elementary questions on multidimensional data like verifying Pareto optimality of a choice of data points. Focusing on constant dimension d, we show that any k-quantifier, d-dimensional such problem is solvable in O(n^{k-1} log^{d-1} n) time. Furthermore, this algorithm is conditionally tight up to subpolynomial factors: we show that assuming the 3-uniform hyperclique hypothesis, there is a k-quantifier, (3k-3)-dimensional problem in this class that requires time Ω(n^{k-1-o(1)}). Towards identifying a single representative problem for this class, we study the existence of complete problems for the 3-quantifier setting (since 2-quantifier problems can already be solved in near-linear time O(nlog^{d-1} n), and k-quantifier problems with k > 3 reduce to the 3-quantifier case). We define a problem Vector Concatenated Non-Domination VCND_d (Given three sets of vectors X,Y and Z of dimension d,d and 2d, respectively, is there an x ∈ X and a y ∈ Y so that their concatenation x∘y is not dominated by any z ∈ Z, where vector u is dominated by vector v if u_i ≤ v_i for each coordinate 1 ≤ i ≤ d), and determine it as the "unique" candidate to be complete for this class (under fine-grained assumptions).
Haozhe An, Mohit Gurumukhani, Russell Impagliazzo, Michael Jaber, Marvin Künnemann, Maria Paula Parga Nina
IPEC1
2020 Quasi-optimal Data Placement for Secure Multi-tenant Data Federation on the Cloud
abstract
As it is difficult to directly share data among different organizations, data federation brings new opportunities to the data-related cooperation among different organizations by providing abstract data interfaces. With the development of Cloud computing, organizations store data on the Cloud to achieve elasticity and scalability for data processing. The existing data placement approaches generally only consider one aspect, which is either communication cost or time cost, and do not consider the features of jobs that process the data. In this paper, we propose an approach to enable secure data processing on the Cloud with the data from different organizations. The approach consists of a data federation platform for secure data processing on the Cloud named FedCube and a greedy data placement algorithm that creates a plan to store data on the Cloud in order to achieve multiple objectives based on a cost model. The cost model is composed of two objectives, i.e., reducing both monetary cost and execution time. We present an experimental evaluation by comparing our data placement algorithm with the existing methods based on the data federation platform. The experiments show that our proposed algorithm significantly reduce the total cost (up to 69.8%).
Ji Liu 0003, Haoyi Xiong, Haozhe An, Xingjian Li 0002, Zhi Feng, Licheng Wang 0004, Dejing Dou
IEEE BigData5
2020 An Investigation of Containment Measures Against the COVID-19 Pandemic in Mainland China
abstract
As the recent COVID-19 outbreak rapidly expands all over the world, various containment measures have been carried out to fight against the COVID-19 pandemic. In Mainland China, the containment measures consist of three types, i.e., Wuhan travel ban, intra-city quarantine and isolation, and intercity travel restriction. In order to carry out the measures, local economy and information acquisition play an important role. In this paper, we investigate the correlation of local economy and the information acquisition on the execution of containment measures to fight against the COVID-19 pandemic in Mainland China. First, we use a parsimonious model, i.e., SIR-X model to estimate the parameters, which represent the execution of intra-city quarantine and isolation in major cities of Mainland China. In order to understand the execution of intra-city quarantine and isolation, we analyze the correlation between the representative parameters including local economy, mobility, and information acquisition. To this end, we collect the data of Gross Domestic Product (GDP), the inflows from Wuhan and outflows, and the COVID-19 related search frequency from a widely-used Web mapping service, i.e., Baidu Maps, and Web search engine, i.e., Baidu Search Engine, in Mainland China. Based on the analysis, we confirm the strong correlation between the local economy and the execution of information acquisition in major cities of Mainland China. We further evidence that, although the cities with high GDP per capita attract more inflows from Wuhan, people are more likely to conduct the quarantine measure and to reduce travelling to other cities. Finally, the correlation analysis using search data shows that well-informed individuals are likely to carry out containment measures.
Ji Liu 0003, Xiakai Wang, Haoyi Xiong, Jizhou Huang, Siyu Huang, Haozhe An, Dejing Dou, Haifeng Wang 0001
IEEE BigData6
2020 RIFLE: Backpropagation in Depth for Deep Transfer Learning through Re-Initializing the Fully-connected LayEr
abstract
Fine-tuning the deep convolution neural network (CNN) using a pre-trained model helps transfer knowledge learned from larger datasets to the target task. While the accuracy could be largely improved even when the training dataset is small, the transfer learning outcome is similar with the pre-trained one with closed CNN weights[17], as the backpropagation here brings less updates to deeper CNN layers. In this work, we propose RIFLE - a simple yet effective strategy that deepens backpropagation in transfer learning settings, through periodically ReInitializing the Fully-connected LayEr with random scratch during the fine-tuning procedure. RIFLE brings significant perturbation to the backpropagation process and leads to deep CNN weights update, while the affects of perturbation can be easily converged throughout the overall learning procedure. The experiments show that the use of RIFLE significantly improves deep transfer learning accuracy on a wide range of datasets, outperforming known tricks for the similar purpose, such as dropout, dropconnect, stochastic depth, and cyclic learning rate, under the same settings with 0.5%-2% higher testing accuracy. Empirical cases and ablation studies further indicate RIFLE brings meaningful updates to deep CNN layers with accuracy improved.
Xingjian Li 0002, Haoyi Xiong, Haozhe An, Cheng-Zhong Xu 0001, Dejing Dou
ICML3