Fangyi Chen

dblp:153/0049 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
abstract
Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologies like trees or graphs. To address this, we introduce STELAR-Vision, a training framework for topology-aware reasoning. At its core is TopoAug, a synthetic data pipeline that enriches training with diverse topological structures. Using supervised fine-tuning and reinforcement learning, we post-train Qwen2VL models with both accuracy and efficiency in mind. Additionally, we propose Frugal Learning, which reduces output length with minimal accuracy loss. On MATH-V and VLM_S2H, STELAR-Vision improves accuracy by 9.7% over its base model and surpasses the larger Qwen2VL-72B-Instruct by 7.3%. On five out-of-distribution benchmarks, it outperforms Phi-4-Multimodal-Instruct by up to 28.4% and LLaMA-3.2-11B-Vision-Instruct by up to 13.2%, demonstrating strong generalization. Compared to Chain-Only training, our approach achieves 4.3% higher overall accuracy on in-distribution datasets and consistently outperforms across all OOD benchmarks.
Han Zhang 0048, Zhantao Yang, Fangyi Chen, Anudeepsekhar Bolimera, Marios Savvides
AAAI4
2026 A critical evaluation of generative query expansion on biomedical literature retrieval
abstract
OBJECTIVE: To evaluate the effectiveness of generative query expansion for biomedical literature retrieval. MATERIALS AND METHODS: We thoroughly examined eight generative query expansion methods using three large language models across five datasets for biomedical literature retrieval. We further performed a quantitative analysis, including performance comparisons, rank transition analysis, and article-type effect analysis. We also conducted a qualitative examination of representative cases, from which we derived an error taxonomy. RESULTS: On BioASQ-Y/N, GPT-4o-based query expansion shifts Recall@10 to 0.417-0.512 and nDCG@10 to 0.358-0.479, relative to a baseline of 0.491 and 0.456. For PubMedQA, Precision@1 ranges from 0.764 to 0.876 and nDCG@10 from 0.847 to 0.931, compared with baseline values of 0.893 and 0.935. For 2019-Trec-PM, query expansion yields Recall@100 of 0.217-0.256 and nDCG@100 of 0.272-0.312, versus a baseline of 0.227 and 0.274. Similarly, for 2018-TREC-PM, Recall@100 spans 0.169-0.227 and nDCG@100 spans 0.195-0.250, relative to baseline scores of 0.164 and 0.191. For 2017-TREC-PM, Recall@100 and nDCG@100 fall within 0.111-0.139 and 0.154-0.191 under query expansion, compared with baseline metrics of 0.102 and 0.147. Both general-purpose and domain-specific Llama-based models demonstrate similar performance to GPT-4o. DISCUSSION AND CONCLUSION: The impact of query expansion varies significantly by the expansion methods and type of evidence, but is relatively agnostic to backbone model choice. Notably, query expansion primarily affects article ranking but has a limited impact on the screening stage. Our findings underscore the unique challenges of biomedical literature retrieval and highlight the need to develop domain-specific information retrieval techniques.
Yilu Fang, Fangyi Chen, Yifan Peng 0002, Chunhua Weng
J. Am. Medical Informatics Assoc.3
2026 A data-driven method for research trend analysis in a scientific discipline: Application to the journal of biomedical informatics
abstract
OBJECTIVE: Accurately characterizing research trends is critical for identifying cutting-edge scientific breakthroughs in their infancy and informing strategic priorities. This research contributes a pipeline that utilizes generative AI technologies to develop research topic taxonomies from publication keywords and analyze keyword evolution within topics, methodological and domain trends, and topic co-occurrences. We demonstrated the pipeline by conducting a retrospective analysis of biomedical informatics research trends in the Journal of Biomedical Informatics (JBI). METHODS: We identified the JBI publications with keywords available on PubMed, spanning 2011-2025. We downloaded all the keywords and categorized them into methodological innovations and health domains, identified topics, assigned topic names, and constructed their hierarchies, all using large-language models (LLMs). We introduced an automated method for evaluating topics, leveraging MeSH terminology as the underlying knowledge base. RESULTS: Using 6,930 unique keywords from 2,427 publications, we derived 1,028 distinct topics related to methodological innovations, with each topic associated with medians of four keywords (Q1: 2, Q3: 13) and six publications (Q1: 2, Q3: 19). We identified 904 topics related to health domains, with each topic associated with three keywords (Q1: 1, Q3: 11) and four publications (Q1: 1, Q3: 15). Based on the topics, we analyzed the prominent research areas, trends in publication volume, evolution of keyword distributions within each topic, and patterns of co-occurring topics. Among the 2,379 eligible publications, 2,009 (84.4%) exhibited overlap between the keyword-derived MeSH terms and the MeSH terms assigned to the publication by the National Library of Medicine. CONCLUSION: This study presents a method that leverages modern generative AI technologies for retrospective analysis of a scientific field to identify emerging topics and to detect shifts in scholarly focus. Illustrated by data for JBI and correlated with historical background events and policy changes, our findings demonstrate the effectiveness and utility of the methods while providing a powerful lens to understand the evolution of biomedical informatics research priorities in JBI.
Yilu Fang, Samir Sanchez Tejada, Fangyi Chen, Edward H. Shortliffe, Vimla L. Patel, Mor Peleg, Chunhua Weng
J. Biomed. Informatics4
2025 SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
abstract
Efficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the representation capacity of the latent space. When applied to Transformer-based architectures, our approach compresses 256×256 and 512×512 images using as few as 32 or 64 1-dimensional tokens. Not only does SoftVQ-VAE show consistent and high-quality reconstruction, more importantly, it also achieves state-of-the-art and significantly faster image generation results across different denoising-based generative models. Remarkably, SoftVQ-VAE improves inference throughput by up to 18x for generating 256×256 images and 55x for 512×512 images while achieving competitive FID scores of 1.78 and 2.21 for SiT-XL. It also improves the training efficiency of the generative models by reducing the number of training iterations by 2.3x while maintaining comparable performance. With its fully-differentiable design and semantic-rich latent space, our experiment demonstrates that SoftVQ-VAE achieves efficient tokenization without compromising generation quality, paving the way for more efficient generative models. Code and model are released1.
Hao Chen 0102, Ze Wang 0008, Xiang Li 0106, Ximeng Sun, Fangyi Chen, Jiang Liu 0014, Jindong Wang 0001, Bhiksha Raj, Zicheng Liu 0001, Emad Barsoum
CVPR5
2025 Masked Autoencoders Are Effective Tokenizers for Diffusion Models
abstract
Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76× faster training and 31× higher inference throughput for 512×512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models will be released.
Hao Chen 0102, Yujin Han, Fangyi Chen, Xiang Li 0106, Yidong Wang 0003, Jindong Wang 0001, Ze Wang 0008, Zicheng Liu 0001, Difan Zou, Bhiksha Raj
ICML3
2025 Semi-supervised learning from small annotated data and large unlabeled data for fine-grained Participants, Intervention, Comparison, and Outcomes entity recognition
abstract
OBJECTIVE: Extracting PICO elements-Participants, Intervention, Comparison, and Outcomes-from clinical trial literature is essential for clinical evidence retrieval, appraisal, and synthesis. Existing approaches do not distinguish the attributes of PICO entities. This study aims to develop a named entity recognition (NER) model to extract PICO entities with fine granularities. MATERIALS AND METHODS: Using a corpus of 2511 abstracts with PICO mentions from 4 public datasets, we developed a semi-supervised method to facilitate the training of a NER model, FinePICO, by combining limited annotated data of PICO entities and abundant unlabeled data. For evaluation, we divided the entire dataset into 2 subsets: a smaller group with annotations and a larger group without annotations. We then established the theoretical lower and upper performance bounds based on the performance of supervised learning models trained solely on the small, annotated subset and on the entire set with complete annotations, respectively. Finally, we evaluated FinePICO on both the smaller annotated subset and the larger, initially unannotated subset. We measured the performance of FinePICO using precision, recall, and F1. RESULTS: Our method achieved precision/recall/F1 of 0.567/0.636/0.60, respectively, using a small set of annotated samples, outperforming the baseline model (F1: 0.437) by more than 16%. The model demonstrates generalizability to a different PICO framework and to another corpus, which consistently outperforms the benchmark in diverse experimental settings (P-value < .001). DISCUSSION: We developed FinePICO to recognize fine-grained PICO entities from text and validated its performance across diverse experimental settings, highlighting the feasibility of using semi-supervised learning (SSL) techniques to enhance PICO entities extraction. Future work can focus on optimizing SSL algorithms to improve efficiency and reduce computational costs. CONCLUSION: This study contributes a generalizable and effective semi-supervised approach leveraging large unlabeled data together with small, annotated data for fine-grained PICO extraction.
Fangyi Chen, Yilu Fang, Yifan Peng 0002, Chunhua Weng
J. Am. Medical Informatics Assoc.1
2025 Mini-mental status examination phenotyping for Alzheimer's disease patients using both structured and narrative electronic health record features
abstract
OBJECTIVE: This study aims to automate the prediction of Mini-Mental State Examination (MMSE) scores, a widely adopted standard for cognitive assessment in patients with Alzheimer's disease, using natural language processing (NLP) and machine learning (ML) on structured and unstructured EHR data. MATERIALS AND METHODS: We extracted demographic data, diagnoses, medications, and unstructured clinical visit notes from the EHRs. We used Latent Dirichlet Allocation (LDA) for topic modeling and Term-Frequency Inverse Document Frequency (TF-IDF) for n-grams. In addition, we extracted meta-features such as age, ethnicity, and race. Model training and evaluation employed eXtreme Gradient Boosting (XGBoost), Stochastic Gradient Descent Regressor (SGDRegressor), and Multi-Layer Perceptron (MLP). RESULTS: We analyzed 1654 clinical visit notes collected between September 2019 and June 2023 for 1000 Alzheimer's disease patients. The average MMSE score was 20, with patients averaging 76.4 years old, 54.7% female, and 54.7% identifying as White. The best-performing model (ie, lowest root mean squared error (RMSE)) is MLP, which achieved an RMSE of 5.53 on the validation set using n-grams, indicating superior prediction performance over other models and feature sets. The RMSE on the test set was 5.85. DISCUSSION: This study developed a ML method to predict MMSE scores from unstructured clinical notes, demonstrating the feasibility of utilizing NLP to support cognitive assessment. Future work should focus on refining the model and evaluating its clinical relevance across diverse settings. CONCLUSION: We contributed a model for automating MMSE estimation using EHR features, potentially transforming cognitive assessment for Alzheimer's patients and paving the way for more informed clinical decisions and cohort identification.
Betina Ross S. Idnay, Fangyi Chen, Casey N. Ta, Matthew W. Schelke, Karen Marder, Chunhua Weng
J. Am. Medical Informatics Assoc.3
2025 CLEAR: A vision to support clinical evidence lifecycle with continuous learning
abstract
Human knowledge of diseases, treatments, and prevention techniques is constantly evolving. The generation of clinical evidence using randomized controlled trials on human subjects occurs notably slowly and inefficiently. The Learning Health System (LHS) has been proposed to facilitate the continuous improvement of individual and population health through a cycle of knowledge, practice, and data. However, the gap between the demand for high-quality evidence to support clinical decisions and the available evidence continues to enlarge. While the current LHS vision articulates the integration of Real-World Data (RWD), the rapid generation of RWD often outpaces the rate of effective evidence synthesis and implementation. Considering this, we propose a new framework that more effectively leverages RWD to support the entire clinical evidence lifecycle through a continuous learning mechanism. This framework, powered by modern data science and informatics, offers enhanced scalability and efficiency. In this vision, specifically, RWD is integrated into the clinical evidence lifecycle via four closed feedback loops: 1) guiding research prioritization and study design, 2) facilitating clinical guideline development, 3) assisting guideline evaluation, and 4) supporting shared decision-making. Our framework enables rapid responsiveness to emerging health data and evolving healthcare needs, timely development of clinical guidelines to optimize clinical recommendations, and sustained improvements in clinical practice and patient outcomes. This vision calls for informatics support for an efficient, scalable, and stakeholder-aware clinical evidence lifecycle.
Yilu Fang, Fangyi Chen, George Hripcsak, Yifan Peng 0002, Patrick B. Ryan, Chunhua Weng
J. Biomed. Informatics3
2025 Scalable scientific interest profiling using large language models
Yilun Liang, Edward Sun, Betina Ross S. Idnay, Yilu Fang, Fangyi Chen, Casey N. Ta, Yifan Peng 0002, Chunhua Weng
J. Biomed. Informatics6
2024 A Reference-Based 3D Semantic-Aware Framework for Accurate Local Facial Attribute Editing
abstract
Facial attribute editing plays a crucial role in synthesizing realistic faces with specific characteristics while maintaining realistic appearances. Despite advancements, challenges persist in achieving precise, 3D-aware attribute modifications, which are crucial for consistent and accurate representations of faces from different angles. Current methods struggle with semantic entanglement and lack effective guidance for incorporating attributes while maintaining image integrity. To address these issues, we introduce a novel framework that merges the strengths of latent-based and reference-based editing methods. Our approach employs a 3D GAN inversion technique to embed attributes from the reference image into a tri-plane space, ensuring 3D consistency and realistic viewing from multiple perspectives. We utilize blending techniques and predicted semantic masks to locate precise edit regions, merging them with the contextual guidance from the reference image. A coarse-to-fine inpainting strategy is then applied to preserve the integrity of untargeted areas, significantly enhancing realism. Our evaluations demonstrate superior performance across diverse editing tasks, validating our framework’s effectiveness in realistic and applicable facial attribute editing.
Yutong Zheng, Yen-Shuo Su, Anudeepsekhar Bolimera, Han Zhang 0048, Fangyi Chen, Marios Savvides
IJCB6
2024 Criteria2Query 3.0: Leveraging generative large language models for clinical trial eligibility query generation
Jimyung Park, Yilu Fang, Casey N. Ta, Betina Ross S. Idnay, Fangyi Chen, Rebecca Shyu, Emily R. Gordon, Matthew E. Spotnitz, Chunhua Weng
J. Biomed. Informatics6
2023 Enhanced Training of Query-Based Object Detection via Selective Query Recollection
abstract
This paper investigates a phenomenon where query-based object detectors mispredict at the last decoding stage while predicting correctly at an intermediate stage. We review the training process and attribute the overlooked phenomenon to two limitations: lack of training emphasis and cascading errors from decoding sequence. We design and present Selective Query Recollection (SQR), a simple and effective training strategy for query-based object detectors. It cumulatively collects intermediate queries as decoding stages go deeper and selectively forwards the queries to the downstream stages aside from the sequential structure. Such-wise, SQR places training emphasis on later stages and allows later stages to work with intermediate queries from earlier stages directly. SQR can be easily plugged into various query-based object detectors and significantly enhances their performance while leaving the inference pipeline unchanged. As a result, we apply SQR on Adamixer, DAB-DETR, and Deformable-DETR across various settings (backbone, number of queries, schedule) and consistently brings 1.4 ~2.8 AP improvement. Code is available at https://github.com/Fangyi-Chen/SQR
Fangyi Chen, Han Zhang 0048, Kai Hu 0010, Chenchen Zhu, Marios Savvides
CVPR1
2023 The Ethical Evaluation Method of Algorithmic Behavior Based on Computational Experiments
Fangyi Chen, Xiao Xue 0001, Xiao Wang 0002
PRICAI (2)1
2023 Highly transparent material classification using the refractive index, reflectivity, and transmissivity features from an imaging model of a time-of-flight camera
Shinan Lang, Fangyi Chen, Yiheng Cai
Mach. Vis. Appl.2
2023 Material classification of polishing and convex surface objects based on photon accumulation point spread function (PAPSF) from imaging model of binocular pulsed time-of-flight camera
Shinan Lang, Jizhong Zhang, Fangyi Chen, Yiheng Cai, Qiang Wu 0020
Mach. Vis. Appl.3
2023 From SOA to VOA: A Shift in Understanding the Operation and Evolution of Service Ecosystem
abstract
With the development of ICT (information and communications technology) and service economy, service ecosystem is emerging in a lot of fields, including E-commerce, O2O(Online To Offline) life service, healthcare service, cloud manufacturing, and so on. As a complex socio-technical system, the evolution of service ecosystem is the joint result of the interaction of the three heterogeneous networks, including social network, service network and value network. Under such circumstances, the traditional SOA (Service Oriented Architecture)-based analysis model is powerless. As a result, how to analyze the laws behind the evolution of service ecosystem is still a serious challenge in the field. This paper proposes a value oriented analysis framework (VOA) of service ecosystem, which can use value as a clue to describe the interaction of the three heterogeneous networks. In addition, a computational experiment system is established to verify the effectiveness of the VOA framework, which stimulates the effect of different intervention strategies on service ecosystem. The result shows that our analysis framework can provide new means and ideas for the analysis of service ecosystem.
Xiao Xue 0001, Deyu Zhou 0001, Fangyi Chen, Xiangning Yu 0001, Zhiyong Feng 0002, Yucong Duan, Lin Meng 0001, Mu Zhang 0013
IEEE Trans. Serv. Comput.3
2022 Unitail: Detecting, Reading, and Matching in Retail Scene
Fangyi Chen, Han Zhang 0048, Zaiwang Li, Jiachen Dou, Shentong Mo, Hao Chen 0102, Uzair Ahmed, Chenchen Zhu, Marios Savvides
ECCV (7)1
2022 Computational Experiments for Complex Social Systems - Part I: The Customization of Computational Model
abstract
Computational experiments have emerged as a new method for quantitative analysis of complex social systems. It has been applied to many interdisciplinary research fields, such as economics, finance, and epidemiology. Though the representation form of computational experiments is relatively flexible, the real system is more complex. Therefore, it is important to seek a balance between the flexibility of computational modeling and the credibility of conclusion. This article proposes a customized design framework for computational experiment models, so as to meet the diverse application demands of computational experiments in different fields. Finally, this article outlines some typical applications of computational experiments to provide a roadmap for its rapid development and widespread application.
Xiao Xue 0001, Fangyi Chen, Deyu Zhou 0001, Xiao Wang 0002, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.2
2021 Semantic Relation Reasoning for Shot-Stable Few-Shot Object Detection
abstract
Few-shot object detection is an imperative and long-lasting problem due to the inherent long-tail distribution of real-world data. Its performance is largely affected by the data scarcity of novel classes. But the semantic relation between the novel classes and the base classes is constant regardless of the data availability. In this work, we investigate utilizing this semantic relation together with the visual information and introduce explicit relation reasoning into the learning of novel object detection. Specifically, we represent each class concept by a semantic embedding learned from a large corpus of text. The detector is trained to project the image representations of objects into this embedding space. We also identify the problems of trivially using the raw embeddings with a heuristic knowledge graph and propose to augment the embeddings with a dynamic relation graph. As a result, our few-shot detector, termed SRR-FSD, is robust and stable to the variation of shots of novel objects. Experiments show that SRR-FSD can achieve competitive results at higher shots, and more importantly, a significantly better performance given both lower explicit and implicit shots. The benchmark protocol with implicit shots removed from the pretrained classification dataset can serve as a more realistic setting for future research.
Chenchen Zhu, Fangyi Chen, Uzair Ahmed, Marios Savvides
CVPR2
2020 Soft Anchor-Point Object Detection
Chenchen Zhu, Fangyi Chen, Marios Savvides
ECCV (9)2
2020 Solving Missing-Annotation Object Detection with Background Recalibration Loss
abstract
This paper focuses on a novel and challenging detection scenario: A majority of true objects/instances is unlabeled in the datasets, so these missing-labeled areas will be regarded as the background during training. Previous art [1] on this problem has proposed to use soft sampling to re-weight the gradients of RoIs based on the overlaps with positive instances, while their method is mainly based on the two-stage detector (i.e. Faster RCNN) which is more robust and friendly for the missing label scenario. In this paper, we introduce a superior solution called Background Recalibration Loss (BRL) that can automatically re-calibrate the loss signals according to the pre-defined IoU threshold and input image. Our design is built on the one-stage detector which is faster and lighter. Inspired by the Focal Loss [2] formulation, we make several significant modifications to fit on the missing-annotation circumstance. We conduct extensive experiments on the curated PASCAL VOC [3] and MS COCO [4] datasets. The results demonstrate that our proposed method outperforms the baseline and other state-of-the-arts by a large margin.
Han Zhang 0048, Fangyi Chen, Qiqi Hao, Chenchen Zhu, Marios Savvides
ICASSP2
2020 NCMS: Towards accurate anchor free object detection through ℓ2 norm calibration and multi-feature selection
abstract
We present simple and flexible drop-in modules in feature pyramids for general object detection, which can be easily generalized to other anchor-free detectors without introducing extra parameters, and only involves negligible computational cost on training and testing. The proposed detector, called NCMS, inserts a simple norm calibration (NC) operation between the feature pyramids and detection head to alleviate and balance the norm bias caused by feature pyramid network (FPN). Furthermore, the NCMS leverages an enhanced multi-feature selective strategy (MS) during training to assign the ground-truth to particular feature pyramid levels as supervisions, in order to obtain more discriminative representation for objects. By generalizing to the state-of-the-art FSAF module (Zhu et al., 2019), our NCMS improves it by 1.6% on COCO val set without bells and whistles. The resulting best model achieves 44.0% mAP with single-model and single-scale testing, which is a fairly competitive result.
Fangyi Chen, Chenchen Zhu, Han Zhang 0048, Marios Savvides
Comput. Vis. Image Underst.1
2019 A Novel Collaborative Control Strategy for Enhanced Training of Vehicle Recognition
abstract
Deep learning methods support vehicular technology in various aspects. How to efficiently and effectively optimize deep learning models remains a challenge. It is known that the learning rate is an important hyper-parameter to optimize models and the batch size is one of the keys to speed up training. This paper empirically studies the principles of scheduling batch size and learning rate during training, and proposes the Collaborative Control Strategy (CCS), which practically improves both classification accuracy and training speed. Instead of stepwise decreasing learning rate and keeping batch size unchanged, we asynchronously adjust batch size and learning rate based on time-division and restart policy. We study and analyze the proposed CCS on general image recognition benchmarks. Compared to traditional training strategy, without bells and whistles CCS decreases error rates by absolute 0.78% and 0.75% on CIFAR-10 and CIFAR-100 respectively, and reduces 40% of training time. For vehicular applications, we demonstrate its advantage on Stanford car-196 dataset with different architectures, showing consistent training speedup and accuracy improvement.
Fangyi Chen, Chenchen Zhu, Marios Savvides
VTC Fall1