EDBT 2026 Demo / reviewers in the wild / expert
Weian Mao
dblp:289/1631
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Segmentation and scene understanding · 30% Video understanding and tracking · 17% Generative modeling · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 24 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein design |
1.5 | 2 | 2024 | Floating Anchor Diffusion Model for Multi-motif Scaffolding · ICML 2024 De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Computer vision › Segmentation and scene understanding
instance segmentation |
1.4 | 2 | 2024 | Generative Active Learning for Long-tailed Instance Segmentation · ICML 2024 SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023 |
Computer vision › Face, body and person analysis
human pose estimation |
1.1 | 2 | 2022 | Poseur: Direct Human Pose Regression with Transformers · ECCV (6) 2022 FCPose: Fully Convolutional Multi-Person Pose Estimation With Dynamic Instance-Aware Convolutions · CVPR 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.9 | 1 | 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation Models · ICLR 2025 |
Machine learning › Generative modeling › molecular generation
inverse folding models |
0.9 | 1 | 2025 | Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein Interactions · ICLR 2025 |
Bioinformatics and computational biology › sequence analysis › sequence modeling
DNA sequence modeling |
0.9 | 1 | 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation Models · ICLR 2025 |
Bioinformatics and computational biology
genomics |
0.9 | 1 | 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation Models · ICLR 2025 |
Bioinformatics and computational biology
protein engineering |
0.9 | 1 | 2025 | Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein Interactions · ICLR 2025 |
Machine learning › Efficient and distributed learning
active learning |
0.8 | 1 | 2024 | Generative Active Learning for Long-tailed Instance Segmentation · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Floating Anchor Diffusion Model for Multi-motif Scaffolding · ICML 2024 |
Computer vision › Segmentation and scene understanding › instance segmentation
long-tailed instance segmentation |
0.8 | 1 | 2024 | Generative Active Learning for Long-tailed Instance Segmentation · ICML 2024 |
Bioinformatics and computational biology › protein design
de novo protein design |
0.8 | 1 | 2024 | De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Bioinformatics and computational biology › protein design
motif scaffolding |
0.8 | 1 | 2024 | Floating Anchor Diffusion Model for Multi-motif Scaffolding · ICML 2024 |
Bioinformatics and computational biology › structural bioinformatics
protein structure |
0.8 | 1 | 2024 | De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Computer vision › Segmentation and scene understanding › object segmentation
class-agnostic segmentation |
0.7 | 1 | 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023 |
Computer vision › Video understanding and tracking
instance embedding |
0.7 | 1 | 2023 | CTVIS: Consistent Training for Online Video Instance Segmentation · ICCV 2023 |
Computer vision › Video understanding and tracking › video instance segmentation
online video instance segmentation |
0.7 | 1 | 2023 | CTVIS: Consistent Training for Online Video Instance Segmentation · ICCV 2023 |
Computer vision › Segmentation and scene understanding › instance segmentation
open-world instance segmentation |
0.7 | 1 | 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023 |
Computer vision › Video understanding and tracking
video instance segmentation |
0.7 | 1 | 2023 | CTVIS: Consistent Training for Online Video Instance Segmentation · ICCV 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.6 | 1 | 2022 | Poseur: Direct Human Pose Regression with Transformers · ECCV (6) 2022 |
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation |
0.5 | 1 | 2021 | FCPose: Fully Convolutional Multi-Person Pose Estimation With Dynamic Instance-Aware Convolutions · CVPR 2021 |
Bioinformatics and computational biology › protein structure prediction
protein-protein docking |
0.3 | 1 | 2025 | Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein Interactions · ICLR 2025 |
Computer vision › 3D vision
geometric deep learning |
0.2 | 1 | 2024 | De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-dataset transfer |
0.2 | 1 | 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.3thermodynamic cycle · 1.7state space model · 1.7inverse folding · 1.7gated convolution · 1.7dilated convolution · 1.7boltzmann alignment · 1.7protein diffusion · 1.5attention aggregation · 1.5ESM · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolutionary profiles for protein fitness predictionabstractMOTIVATION: Predicting the fitness impact of mutations is central to protein engineering but constrained by limited assays relative to the size of sequence space. Protein language models (pLMs) trained with masked language modeling (MLM) exhibit strong zero-shot fitness prediction; we provide an interpretive lens by regarding natural evolution as implicit reward maximization and MLM as inverse reinforcement learning (IRL), in which extant sequences act as expert demonstrations and pLM log-odds serve as fitness estimates. RESULTS: Building on this perspective, we introduce EvoIF, a lightweight model that integrates two complementary sources of evolutionary signal: (i) evolutionary profiles from retrieved homologs and (ii) inverse folding (IF) profiles distilled from IF logits. EvoIF fuses sequence-structure representations with these profiles via a compact transition block, yielding calibrated probabilities for log-odds scoring. On ProteinGym (217 mutational assays; >2.5M mutants), EvoIF and its MSA-enabled variant achieve competitive performance while using only 0.15% of the training data and fewer parameters than recent large models. Ablations confirm that evolutionary and IF profiles are complementary, improving robustness across function types, MSA depths, taxa, and mutation depths. AVAILABILITY AND IMPLEMENTATION: Code is archived on Zenodo at https://doi.org/10.5281/zenodo.20139484. Xiaoran Jiao, Shengdong Lin, Jigang Fan, Zhanming Liang, Weian Mao, Hao Cheng 0012, Chunhua Shen |
Bioinform. | 5 |
| 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation ModelsabstractIn recent years, A variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models.
However, there is a lack of comparison between these recent approaches and the classical architecture—convolutional networks (CNNs)—on foundation model benchmarks.
This raises the question: are CNNs truly being surpassed by these recent approaches based on transformer and SSM architectures? In this paper, we develop a simple but well-designed CNN-based method, termed ConvNova. ConvNova identifies and proposes three effective designs: 1) dilated convolutions, 2) gated convolutions, and 3) a dual-branch framework for gating mechanisms.
Through extensive empirical experiments, we demonstrate that ConvNova significantly outperforms recent methods on more than half of the tasks across several foundation model benchmarks. For example, in histone-related tasks, ConvNova exceeds the second-best method by an average of 5.8\%, while generally utilizing fewer parameters and enabling faster computation. In addition, the experiments observed findings that may be related to biological characteristics. This indicates that CNNs are still a strong competitor compared to Transformers and SSMs. We anticipate that this work will spark renewed interest in CNN-based methods for DNA foundation models. Yu Bo, Weian Mao, Yanjun Shao, Weiqiang Bai, Peng Ye 0006, Xinzhu Ma, Hao Chen 0041, Chunhua Shen |
ICLR | 2 |
| 2025 | Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein InteractionsabstractPredicting the change in binding free energy ($\Delta \Delta G$) is crucial for understanding and modulating protein-protein interactions, which are critical in drug design.
Due to the scarcity of experimental $\Delta\Delta G$ data,
existing methods focus on pre-training,
while neglecting the importance of alignment.
In this work, we propose Boltzmann Alignment technique to transfer knowledge from pre-trained inverse folding models to prediction of $\Delta\Delta G$.
We begin by analyzing the thermodynamic definition of $\Delta\Delta G$ and introducing the Boltzmann distribution to connect energy to the protein conformational distribution.
However, the protein conformational distribution is intractable. Therefore, we employ Bayes’ theorem to circumvent direct estimation and instead utilize the log-likelihood provided by protein inverse folding models for the estimation of $\Delta\Delta G$.
Compared to previous methods based on inverse folding, our method explicitly accounts for the unbound state of the protein complex in the $\Delta \Delta G$ thermodynamic cycle, introducing a physical inductive bias and achieving supervised and unsupervised state-of-the-art (SoTA) performance.
Experimental results on SKEMPI v2 indicate that our method achieves Spearman coefficients of 0.3201 (unsupervised) and 0.5134 (supervised) on SKEMPI v2, significantly surpassing the previously reported
%SoTA values
SoTA results
of 0.2632 and 0.4324, respectively.
Furthermore, we demonstrate the capability of our method in binding
energy prediction, protein-protein docking, and antibody optimization tasks.
Code is available at [https://github.com/aim-uofa/BA-DDG](https://github.com/aim-uofa/BA-DDG) Xiaoran Jiao, Weian Mao, Wengong Jin, Peiyuan Yang, Hao Chen 0041, Chunhua Shen |
ICLR | 2 |
| 2024 | De novo Protein Design Using Geometric Vector Field NetworksabstractAdvances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Only a few basic encoders, like IPA, have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we introduce the Vector Field Network (VFN), that enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential $\textit{universal encoder}$. In protein diffusion (frame modeling), VFN exhibits a impressive performance advantage over IPA, excelling in terms of both designability ($\textbf{67.04}$\% vs. 53.58\%) and diversity ($\textbf{66.54}$\% vs. 51.98\%). In inverse folding(frame and atom modeling), VFN outperforms the previous SoTA model, PiFold ($\textbf{54.7}$\% vs. 51.66\%), on sequence recovery rate; we also propose a method of equipping VFN with the ESM model, which significantly surpasses the previous ESM-based SoTA ($\textbf{62.67}$\% vs. 55.65\%), LM-Design, by a substantial margin. Code is available at https://github.com/aim-uofa/VFN Weian Mao, Muzhi Zhu, Shuaike Shen, Lin Wu 0001, Hao Chen 0041, Chunhua Shen |
ICLR | 1 |
| 2024 | Floating Anchor Diffusion Model for Multi-motif ScaffoldingabstractMotif scaffolding seeks to design scaffold structures for constructing proteins with functions derived from the desired motif, which is crucial for the design of vaccines and enzymes. Previous works approach the problem by inpainting or conditional generation. Both of them can only scaffold motifs with fixed positions, and the conditional generation cannot guarantee the presence of motifs. However, prior knowledge of the relative motif positions in a protein is not readily available, and constructing a protein with multiple functions in one protein is more general and significant because of the synergies between functions. We propose a Floating Anchor Diffusion (FADiff) model. FADiff allows motifs to float rigidly and independently in the process of diffusion, which guarantees the presence of motifs and automates the motif position design. Our experiments demonstrate the efficacy of FADiff with high success rates and designable novel scaffolds. To the best of our knowledge, FADiff is the first work to tackle the challenge of scaffolding multiple motifs without relying on the expertise of relative motif positions in the protein. Code is available at https://github.com/aim-uofa/FADiff. Ke Liu 0012, Weian Mao, Shuaike Shen, Xiaoran Jiao, Hao Cheng 0012, Chunhua Shen |
ICML | 2 |
| 2024 | Generative Active Learning for Long-tailed Instance SegmentationabstractRecently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do not thoroughly explore how to better select and utilize generated data. On the other hand, there is still a lack of research oriented towards active learning on generated data. In this paper, we explore how to perform active learning specifically for generated data in the long-tailed instance segmentation task. Subsequently, we propose BSGAL, a new algorithm that estimates the contribution of the current batch-generated data based on gradient cache. BSGAL is meticulously designed to cater for unlimited generated data and complex downstream segmentation tasks. BSGAL outperforms the baseline approach and effectually improves the performance of long-tailed segmentation. Muzhi Zhu, Chengxiang Fan, Hao Chen 0041, Yang Liu 0357, Weian Mao, Xiaogang Xu 0002, Chunhua Shen |
ICML | 5 |
| 2023 | CTVIS: Consistent Training for Online Video Instance SegmentationabstractThe discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones. Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen |
ICCV | 3 |
| 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt LearningabstractCurrent closed-set instance segmentation models rely on pre-defined class labels for each mask during training and evaluation, largely limiting their ability to detect novel objects. Open-world instance segmentation (OWIS) models address this challenge by detecting unknown objects in a class-agnostic manner. However, previous OWIS approaches completely erase category information during training to keep the model’s ability to generalize to unknown objects. In this work, we propose a novel training mechanism termed SegPrompt that uses category information to improve the model’s class-agnostic segmentation ability for both known and unknown categories. In addition, the previous OWIS training setting exposes the unknown classes to the training set and brings information leakage, which is unreasonable in the real world. Therefore, we provide a new open-world benchmark closer to a real-world scenario by dividing the dataset classes into known-seen-unseen parts. For the first time, we focus on the model’s ability to discover objects that never appear in the training set images.Experiments show that SegPrompt can improve the overall and unseen detection performance by 5.6% and 6.1% in AR on our new benchmark without affecting the inference efficiency. We further demonstrate the effectiveness of our method on existing cross-dataset transfer and strongly supervised settings, leading to 5.5% and 12.3% relative improvement. Code and data are released at: https://github.com/aim-uofa/SegPrompt Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chengxiang Fan, Weian Mao, Chenchen Jing, Yifan Liu 0001, Chunhua Shen |
ICCV | 5 |
| 2022 | Poseur: Direct Human Pose Regression with Transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Zhibin Wang 0004, Anton van den Hengel |
ECCV (6) | 1 |
| 2021 | FCPose: Fully Convolutional Multi-Person Pose Estimation With Dynamic Instance-Aware ConvolutionsabstractWe propose a fully convolutional multi-person pose estimation framework using dynamic instance-aware convolutions, termed FCPose. Different from existing methods, which often require ROI (Region of Interest) operations and/or grouping post-processing, FCPose eliminates the ROIs and grouping post-processing with dynamic instance-aware keypoint estimation heads. The dynamic keypoint heads are conditioned on each instance (person), and can encode the instance concept in the dynamically-generated weights of their filters. Moreover, with the strong representation capacity of dynamic convolutions, the keypoint heads in FCPose are designed to be very compact, resulting in fast inference and making FCPose have almost constant inference time regardless of the number of persons in the image. For example, on the COCO dataset, a real-time version of FCPose using the DLA-34 backbone infers about 4.5×faster than Mask R-CNN (ResNet-101) (41.67 FPS vs. 9.26 FPS) while achieving improved performance (64.8% APkpvs. 64.3% APkp). FCPose also offers better speed/accuracy trade-off than other state-of-the-art methods. Our experiment results show that FCPose is a simple yet effective multi-person pose estimation framework. Code is available at: https://git.io/AdelaiDet Weian Mao, Zhi Tian, Chunhua Shen |
CVPR | 1 |