VLDB 2026 Research / reviewers in the wild / expert
Linjiang Chen
dblp:409/8092
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 33% Language models and text generation · 29% Learning paradigms · 29% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
conditional generation |
1.0 | 1 | 2026 | MSAnchor: De Novo Molecular Generation from Mass Spectrometry Data with Anchor-Extended Molecular Scaffolds · AAAI 2026 |
Bioinformatics and computational biology › structural bioinformatics
molecular structure prediction |
1.0 | 1 | 2026 | MSAnchor: De Novo Molecular Generation from Mass Spectrometry Data with Anchor-Extended Molecular Scaffolds · AAAI 2026 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation |
1.0 | 1 | 2026 | MSAnchor: De Novo Molecular Generation from Mass Spectrometry Data with Anchor-Extended Molecular Scaffolds · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models · NeurIPS 2025 |
Machine learning › Learning paradigms
long-tailed recognition |
0.9 | 1 | 2025 | Deciphering the Extremes: A Novel Approach for Pathological Long-tailed Recognition in Scientific Discovery · NeurIPS 2025 |
Bioinformatics and computational biology › machine learning for biology
molecular relational learning |
0.9 | 1 | 2025 | ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models · NeurIPS 2025 |
Machine learning › Graph learning › graph representation learning
molecular graph representation |
0.3 | 1 | 2025 | ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
conditional information bottleneck · 2.0anchor-extended molecular scaffold · 2.0smooth objective regularization · 1.7balanced supervised contrastive learning · 1.73d conformation encoders · 1.72d graph encoders · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MSAnchor: De Novo Molecular Generation from Mass Spectrometry Data with Anchor-Extended Molecular ScaffoldsabstractTandem mass spectrometry (MS/MS) is a critical tool for identifying molecular structures. By efficiently separating molecular fragments based on their mass-to-charge (m/z) ratios, it facilitates molecular generation and subsequent scientific discoveries. However, de novo molecular generation from MS/MS spectra remains fundamentally constrained by two paramount challenges: the vast chemical space requires effective structural constraints, and the absence of fine-grained substructural generation weakens the correspondences between spectral features and molecular structures. In this work, we propose MSAnchor, a novel two-stage framework for MS/MS-based molecular structure generation. We mitigate the search space challenge through the introduction of Anchor-Extended Molecular Scaffold (AEMS) representation that explicitly encodes side-chain anchoring points, thereby dramatically reducing combinatorial complexity. Leveraging the explicit attachment sites provided by AEMS, we develop anchor-specific priors that establish effective alignments between spectral features and molecular substructures. This fine-grained substructural correspondence is further enhanced by a modified Conditional Information Bottleneck (CIB) module that extracts the most informative spectral components in a structure-aware manner. These innovations enable MSAnchor to generate molecular structures that closely reflect spectral characteristics while constraining combinatorial complexity. Extensive experiments on the CANOPUS and MassSpecGym datasets demonstrate that MSAnchor achieves state-of-the-art performance in molecular structure prediction from MS/MS spectra, with performance improvements that are particularly more pronounced for molecules with higher complexity. Xiaohan Qin, Zhengyang Zhou, Linjiang Chen, Wenjie Du 0003, Yang Wang 0015 |
AAAI | 4 |
| 2025 | ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language ModelsabstractMolecular Relational Learning (MRL) aims to understand interactions between molecular pairs, playing a critical role in advancing biochemical research. With the recent development of large language models (LLMs), a growing number of studies have explored the integration of MRL with LLMs and achieved promising results. However, the increasing availability of diverse LLMs and molecular structure encoders has significantly expanded the model space, presenting major challenges for benchmarking. Currently, there is no LLM framework that supports both flexible molecular input formats and dynamic architectural switching. To address these challenges, reduce redundant coding, and ensure fair model comparison, we propose ModuLM, a framework designed to support flexible LLM-based model construction and diverse molecular representations. ModuLM provides a rich suite of modular components, including 8 types of 2D molecular graph encoders, 11 types of 3D molecular conformation encoders, 7 types of interaction layers, and 7 mainstream LLM backbones. Owing to its highly flexible model assembly mechanism, ModuLM enables the dynamic construction of over 50,000 distinct model configurations. In addition, we provide comprehensive benchmark results to demonstrate the effectiveness of ModuLM in supporting LLM-based MRL tasks. Yizhen Zheng, Huan Yee Koh, Hongxin Xiang, Linjiang Chen, Wenjie Du 0003, Yang Wang 0015 |
NeurIPS | 5 |
| 2025 | Deciphering the Extremes: A Novel Approach for Pathological Long-tailed Recognition in Scientific DiscoveryabstractScientific discovery across diverse fields increasingly grapples with datasets exhibiting pathological long-tailed distributions: a few common phenomena overshadow a multitude of rare yet scientifically critical instances. Unlike standard benchmarks, these scientific datasets often feature extreme imbalance coupled with a modest number of classes and limited overall sample volume, rendering existing long-tailed recognition (LTR) techniques ineffective. Such methods, biased by majority classes or prone to overfitting on scarce tail data, frequently fail to identify the very instances—novel materials, rare disease biomarkers, faint astronomical signals—that drive scientific breakthroughs. This paper introduces a novel, end-to-end framework explicitly designed to address pathological long-tailed recognition in scientific contexts. Our approach synergizes a Balanced Supervised Contrastive Learning (B-SCL) mechanism, which enhances the representation of tail classes by dynamically re-weighting their contributions, with a Smooth Objective Regularization (SOR) strategy that manages the inherent tension between tail-class focus and overall classification performance. We introduce and analyze the real-world ZincFluor chemical dataset ($\mathcal{T}=137.54$) and synthetic benchmarks with controllable extreme imbalances (CIFAR-LT variants). Extensive evaluations demonstrate our method's superior ability to decipher these extremes. Notably, on ZincFluor, our approach achieves a Tail Top-2 accuracy of $66.84\%$, significantly outperforming existing techniques. On CIFAR-10-LT with an imbalance ratio of $1000$ ($\mathcal{T}=100$), our method achieves a tail-class accuracy of $38.99\%$, substantially leading the next best. These results underscore our framework's potential to unlock novel insights from complex, imbalanced scientific datasets, thereby accelerating discovery. Zhe Zhao 0008, Haibin Wen, Xianfu Liu, Pengkun Wang 0001, Liheng Yu, Linjiang Chen, Bo An 0001, Qingfu Zhang 0001, Yang Wang 0015 |
NeurIPS | 7 |