EDBT 2026 Demo / reviewers in the wild / expert
Gary Weiss 0001
dblp:37/1010 · also Gary M. Weiss
· DBLP profile ↗
25ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0001-5009-7101ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-authorArtificial intelligence and machine learning · 6 · 5 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unveiling Bias: Analyzing Race and Gender Disparities in AI-Generated ImageryabstractThis study explores gender and racial bias in AI-generated images. DALL-E 3 was used to generate 2800 images based on prompts related to occupations, activities, and positive/negative personal characteristics, and human reviewers classified the generated images by gender and race. Our analysis reveals that certain prompts are disproportionately associated with specific races and/or genders, suggesting that the AI model may be biased. Race and gender statistics are compared with real-world statistics to determine whether the generated images mirror existing societal biases or introduce new biases. Our findings raise ethical concerns about fairness and representation in AI technologies and discuss the consequences of biased image generation. This research is motivated by the growing integration of AI in media generation and the associated risks of perpetuating and amplifying existing biases. The dataset used in this study is provided via a GitHub repository to support reproducibility, transparency, and broader studies in the research community. David Cordero, Priscilla Diaz, Gary Weiss 0001 |
COMPSAC | 4 |
| 2024 | The Construction and Analysis of Course Grades Across Public Universities
Hyun Jeong, Gary Weiss 0001, Audrey Leung, Daniel D. Leeds |
EDM | 2 |
| 2024 | GPT vs. Llama2: Which Comes Closer to Human Writing?
Fernando Martínez-López, Gary Weiss 0001, Miguel Palma, Haoran Xue, Alexander Borelli |
EDM | 2 |
| 2024 | Predicting GRE Scores from Application Materials in Test-Optional Admissions
Zhengxin Qi, Son Tung Do, John Grossi, Jee Hun Kang, Gary Weiss 0001 |
EDM | 6 |
| 2023 | An Analysis of Grading Patterns in Undergraduate University CoursesabstractUniversity undergraduate course grades have several purposes: they provide feedback to the student and motivation to perform well; serve as admission criteria for entering a major; and are used as selection criteria for future employers and graduate programs. Accurate assignment of grades is therefore important and critical to ensure fairness. However, grades may also impact the student’s assessment of the instructor, which leads to a conflict of interest when such assessments are a component of employment, salary, or tenure decisions. This paper performs a detailed descriptive analysis of undergraduate grades collected over an eight year period from a major metropolitan university. Interesting grading patterns are identified and discussed, and the analysis suggests that grading policies vary substantially at the department, course, and instructor level. A connection is observed between course/department enrollment and average grades assigned. A particular focus of this study involves describing the grading behavior of instructors, with the goal of identifying instructors that assign grades that are statistically far above or below the norm. The analysis performed in this study can be applied to grade data from other universities using our publicly available Python-based analytics tool. The results of these analyses can be used to better understand existing grading policies, identify potential sources of grading inequities, and, when appropriate, take corrective action. Gary Weiss 0001, Luisa A. L. Rosa, Hyun Jeong, Daniel D. Leeds |
COMPSAC | 1 |
| 2022 | Assessing Instructor Effectiveness Based on Future Student Performance
Gary Weiss 0001, Erik Brown, Michael Riad-Zaky, Ruby Iannone, Daniel D. Leeds |
EDM | 1 |
| 2022 | The Impact of Semester Gaps on Student Grades
Gary Weiss 0001, Joseph Denham, Daniel D. Leeds |
EDM | 1 |
| 2022 | Generalized Sequential Pattern Mining of Undergraduate Courses
Daniel D. Leeds, Cody Chen, Fiza Metla, James Guest, Gary Weiss 0001 |
EDM | 6 |
| 2021 | Identifying Hubs in Undergraduate Course Networks Based on Scaled Co-Enrollments
Gary Weiss 0001, Karla Dominguez, Daniel D. Leeds |
EDM | 1 |
| 2021 | Measuring the Academic Impact of Course Sequencing using Student Grade Data
Tess Gutenbrunner, Daniel D. Leeds, Spencer Ross, Michael Riad-Zaky, Gary Weiss 0001 |
EDM | 5 |
| 2021 | Mining Course Groupings based on Academic Performance
Daniel D. Leeds, Gary Weiss 0001 |
EDM | 3 |
| 2020 | Predicting Student Performance in a Master's Program in Data Science using Admissions Data
Qiangwen Xu, Gary Weiss 0001 |
EDM | 4 |
| 2020 | A College Major Recommendation SystemabstractCollege students are required to select a major but are often provided with only a modest amount of support in making this important decision. A poor decision is detrimental to the student, since it may result in the student later switching to a different major with a delay in graduation—or even result in the student leaving the university. This also impacts the university since time to graduation and retention rate are used to evaluate the quality of a university. There is a general lack of research on recommender systems for college majors, with the most relevant systems focusing on course-level recommendations. This study describes and evaluates a recommender system for selecting an undergraduate major, utilizing nine years of historical student data from a large university. The system bases its recommendations on the courses that the student takes in the first few years of college, and how well they performed in these courses. The system is designed to recommend majors that the student is likely to be interested in and will perform well in. Recommendations are evaluated based on the likelihood that the student's actual major was in the top five recommended majors, and whether the student performed above average in that major. The recommendation system dramatically outperforms the baseline strategy of randomly selecting a major, and when the recommendation is followed the student is 12% more likely to perform above average in the major. Samuel A. Stein, Gary Weiss 0001, Daniel D. Leeds |
RecSys | 2 |
| 2020 | Event Detection Through Differential Pattern Mining in Cyber-Physical SystemsabstractExtracting knowledge from sensor data for various purposes has received a great deal of attention by the data mining community. For the purpose of event detection in cyber-physical systems (CPS), e.g., damage in building or aerospace vehicles from the continuous arriving data is challenging due to the detection quality. Traditional data mining schemes are used to reduce data that often use metrics, association rules, and binary values for frequent patterns as indicators for finding interesting knowledge about an event. However, these may not be directly applicable to the network due to certain constraints (communication, computation, bandwidth). We discover that, the indicators may not reveal meaningful information for event detection in practice. In this paper, we propose a comprehensive data mining framework for event detection in the CPS named DPminer, which functions in a distributed and parallel manner (data in a partitioned database processed by one or more sensor processors) and is able to extract a pattern of sensors that may have event information with a low communication cost. To achieve this, we introduce a new sensor behavioral pattern mining technique called differential sensor pattern (DSP) which considers different frequencies and values (non-binary) with a set of sensors, instead of traditional binary patterns. We present an algorithm for data preparation and then use a highly-compact data tree structure (called DP-Tree) for generating the DSP. An important tradeoff between the communication and computation costs for the event detection via data mining is made. Evaluation results show that DPminer can be very useful for networked sensing with a superior performance in terms of communication cost and event detection quality compared to existing data mining schemes. Md. Zakirul Alam Bhuiyan, Jie Wu 0001, Gary Weiss 0001, Thaier Hayajneh, Tian Wang 0001, Guojun Wang 0001 |
IEEE Trans. Big Data | 3 |
| 2016 | Actitracker: A Smartphone-Based Activity Recognition System for Improving Health and Well-BeingabstractActitracker is a smartphone-based activity-monitoring service to help people ensure they receive sufficient activity to maintain proper health. This free service allowed people to set personal activity goals and monitor their progress toward these goals. Actitracker uses machine learning methods to recognize a user's activities. It initially employs a "universal" model generated from labeled activity data from a panel of users, but will automatically shift to a much more accurate personalized model once a user completes a simple training phase. Detailed activity reports and statistics are maintained and provided to the user. Actitracker is a research-based system that began in 2011, before fitness trackers like Fitbit were popular, and was deployed for public use from 2012 until 2015, during which period it had 1,000 registered users. This paper describes the Actitracker system, its use of machine learning, and user experiences. While activity recognition has now entered the mainstream, this paper provides insights into applied activity recognition, something that commercial companies rarely share. Gary Weiss 0001, Jeffrey W. Lockhart, Tony T. Pulickal, Paul T. McHugh, Isaac H. Ronan, Jessica L. Timko |
DSAA | 1 |
| 2014 | The Benefits of Personalized Smartphone-Based Activity Recognition ModelsabstractActivity recognition allows ubiquitous mobile devices like smartphones to be context-aware and also enables new applications, such as mobile health applications that track a user's activities over time. However, it is difficult for smartphone-based activity recognition models to perform well, since only a single body location is instrumented. Most research focuses on universal/impersonal activity recognition models, where the model is trained using data from a panel of representative users. In this paper we compare the performance of these impersonal models with those of personal models, which are trained using labeled data from the intended user, and hybrid models, which combine aspects of both types of models. Our analysis indicates that personal training data is required for high accuracy but that only a very small amount of training data is necessary. This conclusion led us to implement a self-training capability into our Actitracker smartphone-based activity recognition system[1], and we believe personal models can also benefit other activity recognition systems as well. Jeffrey W. Lockhart, Gary Weiss 0001 |
SDM | 2 |
| 2012 | Applications of mobile activity recognitionabstractActivity Recognition (AR), which identifies the activity that a user performs, is attracting a tremendous amount of attention, especially with the recent explosion of smart mobile devices. These ubiquitous mobile devices, most notably but not exclusively smartphones, provide the sensors, processing, and communication capabilities that enable the development of diverse and innovative activity recognition-based applications. However, although there has been a great deal of research into activity recognition, surprisingly little practical work has been done in the area of applications in mobile devices. In this paper we describe and categorize a variety of activity recognition-based applications. Our hope is that this work will encourage the development of such applications and also influence the direction of activity recognition research. Jeffrey W. Lockhart, Tony T. Pulickal, Gary Weiss 0001 |
UbiComp | 3 |
| 2012 | A comparison of alternative client/server architectures for ubiquitous mobile sensor-based applicationsabstractMobile devices such as smart phones, tablet computers, and music players are ubiquitous. These devices typically contain many sensors, such as vision sensors (cameras), audio sensors (microphones), acceleration sensors (accelerometers) and location sensors (e.g., GPS), and also have some capability to send and receive data wirelessly. Sensor arrays on these mobile devices make innovative applications possible, especially when data mining is applied to the sensor data. But a key design decision is how best to distribute the responsibilities between the client (e.g., smartphone) and any servers. In this paper we investigate alternative architectures, ranging from a "dumb" client, where virtually all processing takes place on the server, to a "smart" client, where no server is needed. We describe the advantages and disadvantages of these alternative architectures and describe under what circumstances each is most appropriate. We use our own WISDM (WIreless Sensor Data Mining) architecture to provide concrete examples of the various alternatives. Gary Weiss 0001, Jeffrey W. Lockhart |
UbiComp | 1 |
| 2009 | Quantification and semi-supervised classification methods for handling changes in class distributionabstractIn realistic settings the prevalence of a class may change after a classifier is induced and this will degrade the performance of the classifier. Further complicating this scenario is the fact that labeled data is often scarce and expensive. In this paper we address the problem where the class distribution changes and only unlabeled examples are available from the new distribution. We design and evaluate a number of methods for coping with this problem and compare the performance of these methods. Our quantification-based methods estimate the class distribution of the unlabeled data from the changed distribution and adjust the original classifier accordingly, while our semi-supervised methods build a new classifier using the examples from the new (unlabeled) distribution which are supplemented with predicted class values. We also introduce a hybrid method that utilizes both quantification and semi-supervised learning. All methods are evaluated using accuracy and F-measure on a set of benchmark data sets. Our results demonstrate that our methods yield substantial improvements in accuracy and F-measure. Jack Chongjie Xue, Gary Weiss 0001 |
KDD | 2 |
| 2008 | Maximizing classifier utility when there are data acquisition and modeling costs
Gary Weiss 0001 |
Data Min. Knowl. Discov. | 1 |
| 2008 | Guest editorial: special issue on utility-based data mining
Gary Weiss 0001, Bianca Zadrozny, Maytal Saar-Tsechansky |
Data Min. Knowl. Discov. | 1 |
| 2003 | Learning When Training Data are Costly: The Effect of Class Distribution on Tree InductionabstractFor large, real-world inductive learning problems, the number of training examples often must be limited due to the costs associated with procuring, preparing, and storing the training examples and/or the computational costs associated with learning from them. In such circumstances, one question of practical importance is: if only n training examples can be selected, in what proportion should the classes be represented? In this article we help to answer this question by analyzing, for a fixed training-set size, the relationship between the class distribution of the training data and the performance of classification trees induced from these data. We study twenty-six data sets and, for each, determine the best class distribution for learning. The naturally occurring class distribution is shown to generally perform well when classifier performance is evaluated using undifferentiated error rate (0/1 loss). However, when the area under the ROC curve is used to evaluate classifier performance, a balanced distribution is shown to perform well. Since neither of these choices for class distribution always generates the best-performing classifier, we introduce a budget-sensitive progressive sampling algorithm for selecting training examples based on the class associated with each example. An empirical analysis of this algorithm shows that the class distribution of the resulting training set yields classifiers with good (nearly-optimal) classification performance. Gary Weiss 0001, Foster J. Provost |
J. Artif. Intell. Res. | 1 |
| 1998 | The Problem with Noise and Small Disjuncts
Gary Weiss 0001, Haym Hirsh |
ICML | 1 |
| 1998 | Learning to Predict Rare Events in Event Sequences
Gary Weiss 0001, Haym Hirsh |
KDD | 1 |
| 1995 | Learning with Rare Cases and Small Disjuncts
Gary Weiss 0001 |
ICML | 1 |