Helen Chen

dblp:71/6283 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Operating systems · 44% Compilers and program optimization · 44% Runtime systems and virtual machines · 13%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 56% Rendering · 44%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
memoization
0.612022
Floo: automatic, lightweight memoization for faster mobile apps · MobiSys 2022
Operating systems
mobile systems
0.612022
Floo: automatic, lightweight memoization for faster mobile apps · MobiSys 2022
Rendering › rendering optimization
overdraw reduction
0.312018
Using Animation to Alleviate Overdraw in Multiclass Scatterplot Matrices · CHI 2018
Visualization and visual analytics › scatterplot
scatterplot matrix
0.312018
Using Animation to Alleviate Overdraw in Multiclass Scatterplot Matrices · CHI 2018
Visualization and visual analytics › visualization evaluation
user study
0.112018
Using Animation to Alleviate Overdraw in Multiclass Scatterplot Matrices · CHI 2018

Methods — techniques the papers use, named apart from their topics

memoization · 0.6cache lookup optimization · 0.6animation · 0.3
YearPublicationVenuePosition
2025 SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
abstract
The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks, due to difficulties in acquiring sensitive data. To address this, we introduce SynQP, an open framework for benchmarking privacy in synthetic data generation (SDG) using simulated sensitive data, ensuring that original data remains confidential. We also highlight the need for privacy metrics that fairly account for the probabilistic nature of machine learning models. As a demonstration, we use SynQPto benchmark CTGAN and propose a new identity disclosure risk metric that offers a more accurate estimation of privacy risks compared to existing approaches. Our work provides a critical tool for improving the transparency and reliability of privacy evaluations, enabling safer use of synthetic data in health-related applications. Our privacy assessments (Table II) reveal that DP consistently lowers both identity disclosure risk (SD-IDR) and membershipinference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold.Code available at https://github.com/CAN-SYNH/SynQP
Asma Bahamyirou, Helen Chen
PST4
2024 Synthetic Data from Diffusion Models Improves Drug Discovery Prediction
abstract
Artificial intelligence (AI) is increasingly used in every stage of drug development. Continuing breakthroughs in AI-based methods for drug discovery require the creation, improvement, and refinement of drug discovery data. We posit a new data challenge that slows the advancement of drug discovery AI: datasets are often collected independently from each other, often with little overlap, creating data sparsity. Data sparsity makes data curation difficult for researchers looking to answer key research questions requiring values posed across multiple datasets. We propose a novel diffusion GNN model Syngand capable of generating ligand and pharmacokinetic data end-to-end. We show and provide a methodology for sampling pharmacokinetic data for existing ligands using our Syngand model. We show the initial promising results on the efficacy of the Syngand-generated synthetic target property data on downstream regression tasks with AqSolDB, LD50, and hERG central. Using our proposed model and methodology, researchers can easily generate synthetic ligand data to help them explore research questions that require data spanning multiple datasets.
Ashish Saragadam, Anita Layton, Helen Chen
BIBM4
2022 Floo: automatic, lightweight memoization for faster mobile apps
abstract
Owing to growing feature sets and sluggish improvements to smartphone CPUs (relative to mobile networks), mobile app response times have increasingly become bottlenecked on client-side computations. In designing a solution to this emerging issue, our primary insight is that app computations exhibit substantial stability over time in that they are entirely performed in rarely-updated codebases within app binaries and the OS. Building on this, we present Floo, a system that aims to automatically reuse (or memoize) computation results during app operation in an effort to reduce the amount of compute needed to handle user interactions. To ensure practicality - the struggle with any memoization effort - in the face of limited mobile device resources and the short-lived nature of each app computation, Floo embeds several new techniques that collectively enable it to mask cache lookup overheads and ensure high cache hit rates, all the while guaranteeing correctness for any reused computations. Across a wide range of apps, live networks, phones, and interaction traces, Floo reduces median and 95th percentile interaction response times by 32.7% and 72.3%.
Murali Ramanujam, Helen Chen, Shaghayegh Mardani, Ravi Netravali
MobiSys2
2018 Using Animation to Alleviate Overdraw in Multiclass Scatterplot Matrices
abstract
The scatterplot matrix (SPLOM) is a commonly used technique for visualizing multiclass multivariate data. However, multiclass SPLOMs have issues with overdraw (overlapping points), and most existing techniques for alleviating overdraw focus on individual scatterplots with a single class. This paper explores whether animation using flickering points is an effective way to alleviate overdraw in these multiclass SPLOMs. In a user study with 69 participants, we found that users not only performed better at identifying dense regions using animated SPLOMs, but also found them easier to interpret and preferred them to static SPLOMs. These results open up new directions for future work on alleviating overdraw for multiclass SPLOMs, and provide insights for applying animation to alleviate overdraw in other settings.
Helen Chen, Sophie Engle, Alark Joshi, Eric D. Ragan, Beste F. Yuksel, Lane Harrison
CHI1
2018 Reflective Design Practice: A Novel Assessment of the Impact of Design-based Courses on Students
abstract
This innovative practice work in progress paper responds to the question, how might we prepare design students to transfer their practice from academic contexts to applied contexts? This mixed methods study describes a new reflective assessment called Reflective Design Practice. This assessment prompts student to deliberately reflect on concrete artifacts created during the course. Twenty university students enrolled in design-based and project-based courses used Reflective Design Practice to more deeply understand both their design practice and the way environment impacts design. This assessment differs from existing assessments in that it can be applied to virtually any student work created during a course. Furthermore, it grounds abstract metacognition in tangible output.
Adam Royalty, Helen Chen, Sheri D. Sheppard
FIE2
2009 On optimizing I/O through InfiniBand RDMA for commodity clusters
abstract
Our goal was to enable pNFS as a highperformance parallel file system by using network file system (NFS) storage objects and InfiniBand remote direct memory access (RDMA) transport in the Linux mainline. One obstacle to this approach was the performance bottleneck in NFS/RDMA streaming-writes from the compute nodes. We benchmarked, tuned, and improved the streaming-write efficiency of the Linux NFS client. However, deeper analyses of the benchmarks and the various I/O short-circuit schemes established upper bounds on the performance of the NFS client — even with an infinitely fast network — so that the performance was substantially less than the theoretical streaming bandwidth of the fast interconnections. The complex interactions between the Linux virtual file system, Linux virtual memory management, and the IB network subsystems apparently impose a limit on further improvement.
Benjamin A. Allan, Helen Chen, Scott Cranford, Ron Minnich, Don W. Rudish, Lee Ward
CLUSTER2
2007 Advancing translational research with the Semantic Web
abstract
BACKGROUND: A fundamental goal of the U.S. National Institute of Health (NIH) "Roadmap" is to strengthen Translational Research, defined as the movement of discoveries in basic research to application at the clinical level. A significant barrier to translational research is the lack of uniformly structured data across related biomedical domains. The Semantic Web is an extension of the current Web that enables navigation and meaningful use of digital resources by automatic processes. It is based on common formats that support aggregation and integration of data drawn from diverse sources. A variety of technologies have been built on this foundation that, together, support identifying, representing, and reasoning across a wide range of biomedical data. The Semantic Web Health Care and Life Sciences Interest Group (HCLSIG), set up within the framework of the World Wide Web Consortium, was launched to explore the application of these technologies in a variety of areas. Subgroups focus on making biomedical data available in RDF, working with biomedical ontologies, prototyping clinical decision support systems, working on drug safety and efficacy communication, and supporting disease researchers navigating and annotating the large amount of potentially relevant literature. RESULTS: We present a scenario that shows the value of the information environment the Semantic Web can support for aiding neuroscience researchers. We then report on several projects by members of the HCLSIG, in the process illustrating the range of Semantic Web technologies that have applications in areas of biomedicine. CONCLUSION: Semantic Web technologies present both promise and challenges. Current tools and standards are already adequate to implement components of the bench-to-bedside vision. On the other hand, these technologies are young. Gaps in standards and implementations still exist and adoption is limited by typical problems with early technology, such as the need for a critical mass of practitioners and installed base, and growing pains as the technology is scaled up. Still, the potential of interoperable knowledge sources for biomedicine, at the scale of the World Wide Web, merits continued work.
Alan Ruttenberg, Tim Clark, William J. Bug, Matthias Samwald, Olivier Bodenreider, Helen Chen, Donald Doherty, Kerstin Forsberg, Vipul Kashyap, June Kinoshita, Joanne S. Luciano, M. Scott Marshall, Chimezie Ogbuji, Jonathan Rees, Susie Stephens, Gwendolyn T. Wong, Elizabeth Wu, Davide Zaccagnini, Tonya Hongsermeier, Eric Neumann, Ivan Herman, Kei-Hoi Cheung
BMC Bioinform.6
2004 Comparative Performance Evaluation of iSCSI Protocol over Metro, Local, and Wide Area Networks
Ismail Dalgic, Kadir Ozdemir, Rajkumar Velpuri, Jason Weber, Helen Chen, Umesh Kukreja
MSST5