VLDB 2026 Research / reviewers in the wild / expert
Qi Ye 0004
dblp:19/6124-4
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-2512-1266ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale RecognitionabstractVisual scale recognition is a fundamental aspect for humans to perceive physical quantities in the real world, and it is crucial for enabling human-like intelligence in multimodal large language models (MLLMs). However, existing benchmarks typically focus on a single type of quantity (e.g., time) or a specific format (e.g., dials), lacking a comprehensive evaluation of scale recognition capabilities. To address these problems, we propose ScaleBench, a visual scale recognition benchmark built using images from COCO, Open Images, and Flickr, designed to comprehensively evaluate the scale recognition capabilities of MLLMs. To ensure high data quality, we develop detailed annotation guidelines and procedures, resulting in a total of 6,574 annotated samples. Based on this benchmark, we evaluate multiple closed-source and open-source MLLMs. Experimental results reveal that the best-performing model achieves only 42.60% accuracy, far lower than the 97.40% of humans. Furthermore, we conduct in-depth experimental analyses and provide future research directions. Our benchmark and implementation codes are available at https://github.com/Sonder-hang/ScaleBench. Jihang Jin, Ronghao Chen, Huacan Wang, Qi Ye 0004 |
ACL (1) | 6 |
| 2025 | IMQC: A Large Language Model Platform for Medical Quality ControlabstractMedical quality control (MQC) indicators are essential for evaluating the performance of healthcare institutions to ensure high-quality patient care. In this paper, we report the design, implementation, and deployment of the Intelligent EMR-LLM platform for Medical Quality Control (IMQC), a large language model (LLM)-empowered system for automatically computing MQC indicators for enhancing the quality of medical services in Shanghai. It consists of an LLM (i.e., EMR-LLM) for processing electronic medical records (EMRs). With EMR-LLM, IMQC translates existing MQC indicators into a standardized representation language and automatically computes them based on EMRs. Since its deployment in February 2024, IMQC has been adopted by the Shanghai Medical Quality Management Center and associated hospitals. So far, it has processed 1,245 medical quality indicators for secondary- and tertiary-level hospitals, achieving an MQC evaluation accuracy of 93.31%, which is comparable to human experts. It has significantly improved efficiency, increasing from 10 EMRs per hour per human expert to over 1,000 EMRs per hour on average using one single H800 GPU. Over the first round of deployment in Shanghai, it is estimated that IMQC saves around 3.42 million RMB per month in manpower costs compared to traditional reporting methods. The successful deployment of IMQC sets a precedence for other regions to adopt similar AI-driven solutions to enhance medical quality control. Qi Ye 0004, Guangya Yu, Erzhen Chen, Chenjie Dong, Xiaosheng Lin, Zelei Liu, Han Yu 0001, Tong Ruan |
AAAI | 1 |
| 2024 | Alignment of Chinese-English Medical Terminology in Small-Sample Scenarios: A Two-Stage ApproachabstractCross-lingual terminology alignment is an important task in the field of medical terminology. Through cross-lingual alignment, medical terms from different languages can be accurately mapped to their corresponding concepts, thus establishing a unified and multilingual medical terminology fusion system. However, the scarcity of annotated parallel corpora for medical terminology in Chinese and English languages poses a challenge during the training process of Chinese-English terminology alignment models, resulting in decreased alignment accuracy. To address this issue, this paper proposes a hybrid approach that combines Large Language Models(LLMs) and pretrained Language Models(PLMs), leveraging the rich multilingual knowledge embedded in LLMs to obtain more comprehensive Chinese-English terminology information for assisting cross-lingual terminology alignment. However, LLMs require significant computational resources and memory, making them less suitable for large-scale alignment tasks. To tackle this problem, a confidence sampling strategy is introduced, delegating challenging samples to the LLMs for re-ranking, thereby reducing resource costs. Additionally, a prompt strategy tailored for terminology alignment tasks is proposed to enhance the accuracy of predictions made by the LLMs.We evaluate our method on mapping files of two medical open terminologies, and the experimental results demonstrate that our method outperforms baseline methods by 2% in terms of Hits@1 and Hits@10 metrics.Our code and data are available at https://github.com/Bruce-Y12/two-stage-ranking. Qi Ye 0004, Zicheng Yao, Peihong Hu, Tong Ruan, Ruihui Hou |
BIBM | 1 |
| 2024 | A multi-view representation learning framework for commonsense knowledge bases
Weiyan Zhang, Qi Ye 0004, Tong Ruan |
Inf. Sci. | 5 |
| 2024 | Emotion-wise feature interaction analysis-based visual emotion distribution learning
Jing Zhang 0041, Qiuge Qin, Qi Ye 0004, Wen Du |
Vis. Comput. | 4 |
| 2023 | AF Adapter: Continual Pretraining for Building Chinese Biomedical Language ModelabstractContinual pretraining is a popular way of building a domain-specific pretrained language model from a general-domain language model. In spite of its high efficiency, continual pretraining suffers from catastrophic forgetting, which may harm the model’s performance in downstream tasks. To alleviate the issue, in this paper, we propose a continual pretraining method for the BERT-based model, named Attention-FFN Adapter. Its main idea is to introduce a small number of attention heads and hidden units inside each self-attention layer and feed-forward network. Furthermore, we train a domain-specific language model named AF Adapter based RoBERTa for the Chinese biomedical domain. In experiments, models are applied to downstream tasks for evaluation. The results demonstrate that with only about 17% of model parameters trained, AF Adapter achieves 0.6%, 2% gain in performance on average, compared to strong baselines. Further experimental results show that our method alleviates the catastrophic forgetting problem by 11% compared to the fine-tuning method. Code is available at https://github.com/yanyongyu/AF-Adapter. Yongyu Yan, Kui Xue, Qi Ye 0004, Tong Ruan |
BIBM | 4 |
| 2023 | Integration of multiple terminology bases: a multi-view alignment method using the hierarchical structureabstractMOTIVATION: In the medical field, multiple terminology bases coexist across different institutions and contexts, often resulting in the presence of redundant terms. The identification of overlapping terms among these bases holds significant potential for harmonizing multiple standards and establishing unified framework, which enhances user access to comprehensive and well-structured medical information. However, the majority of terminology bases exhibit differences not only in semantic aspects but also in the hierarchy of their classification systems. The conventional approaches that rely on neighborhood-based methods such as GCN may introduce errors due to the presence of different superordinate and subordinate terms. Therefore, it is imperative to explore novel methods to tackle this structural challenge. RESULTS: To address this heterogeneity issue, this paper proposes a multi-view alignment approach that incorporates the hierarchical structure of terminologies. We utilize BERT-based model to capture the recursive relationships among different levels of hierarchy and consider the interaction information of name, neighbors, and hierarchy between different terminologies. We test our method on mapping files of three medical open terminologies, and the experimental results demonstrate that our method outperforms baseline methods in terms of Hits@1 and Hits@10 metrics by 2%. AVAILABILITY AND IMPLEMENTATION: The source code will be available at https://github.com/Ulricab/Bert-Path upon publication. Peihong Hu, Qi Ye 0004, Weiyan Zhang, Tong Ruan |
Bioinform. | 2 |
| 2022 | Image sentiment classification via multi-level sentiment region correlation analysis
Jing Zhang 0041, Qi Ye 0004, Zhe Wang 0002 |
Neurocomputing | 4 |
| 2021 | An Integrated Resampling Methods for Imbalanced Sporadic Temporal Data in EHRsabstractMost real-world applications in EHRs involve temporal data with skewed distributions. The imbalanced classification problem becomes more difficult in sporadic temporal data that variables exist on correlation and have some missing values. A common solution to classification tasks with imbalanced data is the oversampling methods, which generate new samples to re-balancing the classes. However, traditional oversampling methods usually change the distribution, thereby leading to bias. This paper proposed a self-adaptive integrated oversampling method for imbalanced sporadic temporal data in EHRs. The masking vectors and density vectors have been introduced to measure missing value distribution of samples, and the minority samples are divided into high density samples and sparse density samples. We extend the resampling strategies combining a subsample alignment method and structure preserving oversampling method. The weight of sample difference is used to improve classification performance. Furthermore, the filter mechanism is proposed to remove the noise samples with good efficiency. The experimental results show that the proposed method increases performance compared to traditional resampling methods in terms of AUC, F1, and G-mean evaluation metrics. Qi Ye 0004, Tomohiro Kuroda, Tong Ruan, Xiaoling Ge |
BIBM | 1 |
| 2021 | A combined recall and rank framework with online negative sampling for Chinese procedure terminology normalizationabstractMOTIVATION: Medical terminology normalization aims to map the clinical mention to terminologies coming from a knowledge base, which plays an important role in analyzing electronic health record and many downstream tasks. In this article, we focus on Chinese procedure terminology normalization. The expressions of terminology are various and one medical mention may be linked to multiple terminologies. Existing studies based on learning to rank does not fully consider the quality of negative samples during model training and the importance of keywords in this domain-specific task. RESULTS: We propose a combined recall and rank framework to solve these problems. A pair-wise Bert model with deep metric learning is used to recall candidates. Previous methods either train Bert in a point-wise way or based on a multi-class classification problem, which may lead serious efficiency problems or not be effective enough. During model training, we design a novel online negative sampling algorithm to activate the pair-wise method. To deal with multi-implication scenarios, we train the task of implication number prediction together with the recall task in a multi-task learning setting, since these two tasks are highly complementary. In rank step, we propose a keywords attentive mechanism to focus on domain-specific information such as procedure sites and procedure types. Finally, a fusion block merges the results of the recall and the rank model. Detailed experimental analysis shows our proposed framework has a remarkable improvement on both performance and efficiency. AVAILABILITY AND IMPLEMENTATION: The source code will be available at https://github.com/sxthunder/CMTN upon publication. Kui Xue, Qi Ye 0004, Tong Ruan |
Bioinform. | 3 |
| 2018 | An Effective Standardization Method for the Lab Indicators in Regional Medical Health Platform Using N-grams and Stacking
Qi Wang 0020, Yangming Zhou, Qi Ye 0004, Jiahui Qiu |
BIBM | 5 |