Yiqun Chen 0001

dblp:59/1143-1 · also Yiqun T. Chen · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-4100-1507ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 67% Web and social media mining · 33%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 50% Design research and methods · 50%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%
Artificial intelligence
2 papers
Generative modeling · 50% Trustworthy machine learning · 25% Probabilistic and Bayesian machine learning · 25%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.712023
Selective inference for k-means clustering · J. Mach. Learn. Res. 2023
Data mining › clustering
k-means clustering
0.712023
Selective inference for k-means clustering · J. Mach. Learn. Res. 2023
Web and social media mining › social media analysis
social image analysis
0.712023
TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter · NeurIPS 2023
Software testing
fault detection
0.412020
Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set Size · ASE 2020
Software testing › test adequacy
test adequacy criteria
0.412020
Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set Size · ASE 2020
Machine learning › Probabilistic and Bayesian machine learning › clustering
cluster validity
0.212023
Selective inference for k-means clustering · J. Mach. Learn. Res. 2023
Machine learning › Generative modeling
diffusion model
0.212023
TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.212023
Selective inference for k-means clustering · J. Mach. Learn. Res. 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.212023
TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter · NeurIPS 2023
Software testing › test coverage
coverage-based testing
0.112020
Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set Size · ASE 2020

Methods — techniques the papers use, named apart from their topics

metadata analysis · 2.0comparative image analysis · 2.0systematic review · 1.3survey · 1.3selective inference · 1.3hypothesis testing · 1.3bootstrap · 1.3probabilistic coupling · 0.4empirical evaluation · 0.4
YearPublicationVenuePosition
2023 Why, when, and from whom: considerations for collecting and reporting race and ethnicity data in HCI
abstract
Engaging diverse participants in HCI research is critical for creating safe, inclusive, and equitable technology. However, there is a lack of guidelines on when, why, and how HCI researchers collect study participants’ race and ethnicity. Our paper aims to take the first step toward such guidelines by providing a systematic review and discussion of the status quo of race and ethnicity data collection in HCI. Through an analysis of 2016–2021 CHI proceedings and a survey with 15 authors who published in these proceedings, we found that reporting race and ethnicity of participants is very rare (<3%) and that researchers are far from consensus. Drawing from multidisciplinary literature and our findings, we devise considerations for HCI researchers to decide why, when, and from whom to collect race and ethnicity data. For truly inclusive, equitable technologies, we encourage deliberate decisions rather than default omissions.
Yiqun Chen 0001, Angela D. R. Smith, Katharina Reinecke, Alexandra To
CHI1
2023 TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter
abstract
Recent progress in generative artificial intelligence (gen-AI) has enabled the generation of photo-realistic and artistically-inspiring photos at a single click, catering to millions of users online. To explore how people use gen-AI models such as DALLE and StableDiffusion, it is critical to understand the themes, contents, and variations present in the AI-generated photos. In this work, we introduce TWIGMA (TWItter Generative-ai images with MetadatA), a comprehensive dataset encompassing over 800,000 gen-AI images collected from Jan 2021 to March 2023 on Twitter, with associated metadata (e.g., tweet text, creation date, number of likes). Through a comparative analysis of TWIGMA with natural images and human artwork, we find that gen-AI images possess distinctive characteristics and exhibit, on average, lower variability when compared to their non-gen-AI counterparts. Additionally, we find that the similarity between a gen-AI image and natural images is inversely correlated with the number of likes. Finally, we observe a longitudinal shift in the themes of AI-generated images on Twitter, with users increasingly sharing artistically sophisticated content such as intricate human portraits, whereas their interest in simple subjects such as natural scenes and animals has decreased. Our analyses and findings underscore the significance of TWIGMA as a unique data resource for studying AI-generated images.
Yiqun Chen 0001, James Zou 0001
NeurIPS1
2023 Selective inference for k-means clustering
abstract
We consider the problem of testing for a difference in means between clusters of observations identified via k-means clustering. In this setting, classical hypothesis tests lead to an inflated Type I error rate. In recent work, Gao et al. (2022) considered a related problem in the context of hierarchical clustering. Unfortunately, their solution is highly-tailored to the context of hierarchical clustering, and thus cannot be applied in the setting of k-means clustering. In this paper, we propose a p-value that conditions on all of the intermediate clustering assignments in the k-means algorithm. We show that the p-value controls the selective Type I error for a test of the difference in means between a pair of clusters obtained using k-means clustering in finite samples, and can be efficiently computed. We apply our proposal on hand-written digits data and on single-cell RNA-sequencing data.
Yiqun Chen 0001, Daniela M. Witten
J. Mach. Learn. Res.1
2020 Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set Size
abstract
The research community has long recognized a complex interrelationship between fault detection, test adequacy criteria, and test set size. However, there is substantial confusion about whether and how to experimentally control for test set size when assessing how well an adequacy criterion is correlated with fault detection and when comparing test adequacy criteria. Resolving the confusion, this paper makes the following contributions: (1) A review of contradictory analyses of the relationships between fault detection, test adequacy criteria, and test set size. Specifically, this paper addresses the supposed contradiction of prior work and explains why test set size is neither a confounding variable, as previously suggested, nor an independent variable that should be experimentally manipulated. (2) An explication and discussion of the experimental designs of prior work, together with a discussion of conceptual and statistical problems, as well as specific guidelines for future work. (3) A methodology for comparing test adequacy criteria on an equal basis, which accounts for test set size without directly manipulating it through unrealistic stratification. (4) An empirical evaluation that compares the effectiveness of coverage-based testing, mutation-based testing, and random testing. Additionally, this paper proposes probabilistic coupling, a methodology for assessing the representativeness of a set of test goals for a given fault and for approximating the fault-detection probability of adequate test sets.
Yiqun Chen 0001, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst, Reid Holmes, Gordon Fraser 0001, Paul Ammann, René Just
ASE1