Jingcheng Wu

dblp:229/2981 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 30% Knowledge representation and reasoning · 24% Trustworthy machine learning · 23%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
1.922026
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning · AAAI 2026
Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guarantees · EMNLP 2025
Computer vision › Vision and language
visual entity recognition
1.012026
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning · AAAI 2026
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
0.912025
Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guarantees · EMNLP 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guarantees · EMNLP 2025
Machine learning › Efficient and distributed learning › federated learning
asynchronous federated learning
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Machine learning › Optimization for machine learning
convergence analysis
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Machine learning › Efficient and distributed learning
federated learning
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Audio and music processing
music generation
0.812024
SongCreator: Lyrics-based Universal Song Generation · NeurIPS 2024
Audio and music processing › music generation
song generation
0.812024
SongCreator: Lyrics-based Universal Song Generation · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

zero-shot recognition · 1.0contrastive learning · 1.0non-conformity measure · 0.9conformal prediction · 0.9dual-sequence language model · 0.8convergence analysis · 0.8cached update calibration · 0.8attention masking · 0.8
YearPublicationVenuePosition
2026 Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
abstract
Open-domain visual entity recognition aims to identify and link entities depicted in images to a vast and evolving set of real-world concepts, such as those found in Wikidata. Unlike conventional classification tasks with fixed label sets, it operates under open-set conditions, where most target entities are unseen during training and exhibit long-tail distributions. This makes the task inherently challenging due to limited supervision, high visual ambiguity, and the need for semantic disambiguation. We propose a Knowledge-guided Contrastive Learning (KnowCoL) framework that combines both images and text descriptions into a shared semantic space grounded by structured information from Wikidata. By abstracting visual and textual inputs to a conceptual level, the model leverages entity descriptions, type hierarchies, and relational context to support zero-shot entity recognition. We evaluate our approach on the OVEN benchmark, a large-scale open-domain visual recognition dataset with Wikidata IDs as the label space. Our experiments show that using visual, textual, and structured knowledge greatly improves accuracy, especially for rare and unseen entities. Our smallest model improves the accuracy on unseen entities by 10.5% compared to the state-of-the-art, despite being 35 times smaller.
Lavdim Halilaj, Sebastian Monka, Stefan Schmid 0002, Yuqicheng Zhu, Jingcheng Wu, Nadeem Nazer, Steffen Staab
AAAI6
2025 Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guarantees
abstract
Uncertain knowledge graph embedding (Un-KGE) methods learn vector representations that capture both structural and uncertainty information to predict scores of unseen triples.However, existing methods produce only point estimates, without quantifying predictive uncertainty-limiting their reliability in high-stakes applications where understanding confidence in predictions is crucial.To address this limitation, we propose UNKGCP, a framework that generates prediction intervals guaranteed to contain the true score with a user-specified level of confidence.The length of the intervals reflects the model's predictive uncertainty.UNKGCP builds on the conformal prediction framework but introduces a novel nonconformity measure tailored to UnKGE methods and an efficient procedure for interval construction.We provide theoretical guarantees for the intervals and empirically verify these guarantees.Extensive experiments on standard benchmarks across diverse UnKGE methods further demonstrate that the intervals are sharp and effectively capture predictive uncertainty.To support future research on this topic, we release our code 1 .
Yuqicheng Zhu, Jingcheng Wu, Jiaoyan Chen 0001, Evgeny Kharlamov, Steffen Staab
EMNLP2
2024 Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration
abstract
Asynchronous federated learning, which enables local clients to send their model update asynchronously to the server without waiting for others, has recently emerged for its improved efficiency and scalability over traditional synchronized federated learning. In this paper, we study how the asynchronous delay affects the convergence of asynchronous federated learning under non-i.i.d. distributed data across clients. Through the theoretical convergence analysis of one representative asynchronous federated learning algorithm under standard nonconvex stochastic settings, we show that the asynchronous delay can largely slow down the convergence, especially with high data heterogeneity. To further improve the convergence of asynchronous federated learning under heterogeneous data distributions, we propose a novel asynchronous federated learning method with a cached update calibration. Specifically, we let the server cache the latest update for each client and reuse these variables for calibrating the global update at each round. We theoretically prove the convergence acceleration for our proposed method under nonconvex stochastic settings. Extensive experiments on several vision and language tasks demonstrate our superior performances compared to other asynchronous federated learning baselines.
Yuanpu Cao, Jingcheng Wu
ICLR3
2024 SongCreator: Lyrics-based Universal Song Generation
abstract
Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and a series of attention mask strategies for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various songrelated generation tasks by utilizing specific attention masks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different audio prompts, exhibiting its potential applicability. Our samples are available at https://thuhcsi.github.io/SongCreator/.
Shun Lei, Yixuan Zhou 0002, Boshi Tang, Max W. Y. Lam, Jingcheng Wu, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng
NeurIPS7
2024 CovEpiAb: a comprehensive database and analysis resource for immune epitopes and antibodies of human coronaviruses
abstract
Coronaviruses have threatened humans repeatedly, especially COVID-19 caused by SARS-CoV-2, which has posed a substantial threat to global public health. SARS-CoV-2 continuously evolves through random mutation, resulting in a significant decrease in the efficacy of existing vaccines and neutralizing antibody drugs. It is critical to assess immune escape caused by viral mutations and develop broad-spectrum vaccines and neutralizing antibodies targeting conserved epitopes. Thus, we constructed CovEpiAb, a comprehensive database and analysis resource of human coronavirus (HCoVs) immune epitopes and antibodies. CovEpiAb contains information on over 60 000 experimentally validated epitopes and over 12 000 antibodies for HCoVs and SARS-CoV-2 variants. The database is unique in (1) classifying and annotating cross-reactive epitopes from different viruses and variants; (2) providing molecular and experimental interaction profiles of antibodies, including structure-based binding sites and around 70 000 data on binding affinity and neutralizing activity; (3) providing virological characteristics of current and past circulating SARS-CoV-2 variants and in vitro activity of various therapeutics; and (4) offering site-level annotations of key functional features, including antibody binding, immunological epitopes, SARS-CoV-2 mutations and conservation across HCoVs. In addition, we developed an integrated pipeline for epitope prediction named COVEP, which is available from the webpage of CovEpiAb. CovEpiAb is freely accessible at https://pgx.zju.edu.cn/covepiab/.
Jingcheng Wu, Yuanyuan Luo, Yilin Wang 0032, Ruiying Kong, Ying Chi, Yisheng Sun, Qiaojun He, Zhan Zhou
Briefings Bioinform.2
2021 CanDriS: posterior profiling of cancer-driving sites based on two-component evolutionary model
abstract
Current cancer genomics databases have accumulated millions of somatic mutations that remain to be further explored. Due to the over-excess mutations unrelated to cancer, the great challenge is to identify somatic mutations that are cancer-driven. Under the notion that carcinogenesis is a form of somatic-cell evolution, we developed a two-component mixture model: while the ground component corresponds to passenger mutations, the rapidly evolving component corresponds to driver mutations. Then, we implemented an empirical Bayesian procedure to calculate the posterior probability of a site being cancer-driven. Based on these, we developed a software CanDriS (Cancer Driver Sites) to profile the potential cancer-driving sites for thousands of tumor samples from the Cancer Genome Atlas and International Cancer Genome Consortium across tumor types and pan-cancer level. As a result, we identified that approximately 1% of the sites have posterior probabilities larger than 0.90 and listed potential cancer-wide and cancer-specific driver mutations. By comprehensively profiling all potential cancer-driving sites, CanDriS greatly enhances our ability to refine our knowledge of the genetic basis of cancer and might guide clinical medication in the upcoming era of precision medicine. The results were displayed in a database CandrisDB (http://biopharm.zju.edu.cn/candrisdb/).
Wenyi Zhao, Jingcheng Wu, Guoxing Cai, Jeffrey Haltom, Weijia Su, Michael J. Dong, Jian Wu 0001, Zhan Zhou, Xun Gu 0002
Briefings Bioinform.3