VLDB 2026 Research / reviewers in the wild / expert
Yehui Yang
dblp:161/7207
· DBLP profile ↗
19ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Image to Pixels: Towards Fine-Grained Medical Vision-Language ModelsabstractMultimodal large language models (MLLMs) offer immense potential for biomedical AI, yet current applications remain limited to coarse-grained image understanding and basic textual queries-falling short of the fine-grained reasoning required in clinical contexts. In this work, we present a comprehensive solution spanning data, model, and training innovations to advance pixel-level multimodal intelligence in biomedicine. First, we construct MeCoVQA, a new visual-language benchmark that spans eight medical imaging modalities and four core tasks, supporting both spatially-grounded reasoning and fine-grained diagnostic comprehension. Building on this, we introduce MedPLIB, an end-to-end biomedical MLLM equipped with pixel-level visual understanding. MedPLIB supports diverse multimodal tasks-including VQA, point- and region-based querying, grounding, and segmentation-through unified modeling. To further accommodate the heterogeneous nature of biomedical tasks, we design a task-specialized Mixture-of-Experts (MoE) architecture, where each expert is tailored to a specific task and jointly optimized via unified fine-tuning. This modular design accommodates diverse biomedical tasks while maintaining a unified and efficient architecture. By integrating retrieval-augmented generation (RAG) and in-context learning (ICL), MedPLIB also demonstrates strong generalization on out-of-distribution (OOD) medical image segmentation. Experiments across multiple benchmarks show that MedPLIB sets a new state-of-the-art on biomedical vision-language tasks; notably, it outperforms the best existing small and large models by 19.7 and 15.6 mDice in zero-shot pixel-level grounding, highlighting its clinical utility and generalization strength. Lingdong Shen, Xiaoshuang Huang, Fangxin Shang, Yehui Yang, Bin Fan 0001, Shiming Xiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Towards a Multimodal Large Language Model with Pixel-Level Insight for BiomedicineabstractIn recent years, Multimodal Large Language Models (MLLM) have achieved notable advancements, demonstrating the feasibility of developing an intelligent biomedical assistant. However, current biomedical MLLMs predominantly focus on image-level understanding and restrict interactions to textual commands, thus limiting their capability boundaries and the flexibility of usage. In this paper, we introduce a novel end-to-end multimodal large language model for the biomedical domain, named MedPLIB, which possesses pixel-level understanding. Excitingly, it supports visual question answering (VQA), arbitrary pixel-level prompts (points, bounding boxes, and free-form shapes), and pixel-level grounding. We propose a novel Mixture-of-Experts (MoE) multi-stage training strategy, which divides MoE into separate training phases for a visual-language expert model and a pixel-grounding expert model, followed by fine-tuning using MoE. This strategy effectively coordinates multitask learning while maintaining the computational cost at inference equivalent to that of a single expert model. To advance the research of biomedical MLLMs, we introduce the Medical Complex Vision Question Answering Dataset (MeCoVQA), which comprises an array of 8 modalities for complex medical imaging question answering and image region understanding. Experimental results indicate that MedPLIB has achieved state-of-the-art outcomes across multiple medical visual language tasks. More importantly, in zero-shot evaluations for the pixel grounding task, MedPLIB leads the best small and large models by margins of 19.7 and 15.6 respectively on the mDice metric. Xiaoshuang Huang, Lingdong Shen, Fangxin Shang, Yehui Yang |
AAAI | 7 |
| 2024 | A Refer-and-Ground Multimodal Large Language Model for Biomedicine
Xiaoshuang Huang, Lingdong Shen, Yehui Yang, Fangxin Shang, Jia Liu 0010 |
MICCAI (12) | 4 |
| 2024 | An Embeddable Implicit IUVD Representation for Part-Based 3D Human Surface ReconstructionabstractTo reconstruct a 3D human surface from a single image, it is crucial to simultaneously consider human pose, shape, and clothing details. Recent approaches have combined parametric body models (such as SMPL), which capture body pose and shape priors, with neural implicit functions that flexibly learn clothing details. However, this combined representation introduces additional computation, e.g. signed distance calculation in 3D body feature extraction, leading to redundancy in the implicit query-and-infer process and failing to preserve the underlying body shape prior. To address these issues, we propose a novel IUVD-Feedback representation, consisting of an IUVD occupancy function and a feedback query algorithm. This representation replaces the time-consuming signed distance calculation with a simple linear transformation in the IUVD space, leveraging the SMPL UV maps. Additionally, it reduces redundant query points through a feedback mechanism, leading to more reasonable 3D body features and more effective query points, thereby preserving the parametric body prior. Moreover, the IUVD-Feedback representation can be embedded into any existing implicit human reconstruction pipeline without requiring modifications to the trained neural networks. Experiments on the THuman2.0 dataset demonstrate that the proposed IUVD-Feedback representation improves the robustness of results and achieves three times faster acceleration in the query-and-infer process. Furthermore, this representation holds potential for generative applications by leveraging its inherent semantic information from the parametric body model. Baoxing Li, Yehui Yang, Xu Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Temporally consistent reconstruction of 3D clothed human surface with warp field
Baoxing Li, Yehui Yang, Xu Zhao 0001 |
Image Vis. Comput. | 3 |
| 2022 | A Speaker-Aware Co-Attention Framework for Medical Dialogue Information ExtractionabstractYuan Xia, Zhenhui Shi, Jingbo Zhou, Jiayu Xu, Chao Lu, Yehui Yang, Lei Wang, Haifeng Huang, Xia Zhang, Junwei Liu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Yuan Xia, Zhenhui Shi, Yehui Yang |
EMNLP | 6 |
| 2022 | SeATrans: Learning Segmentation-Assisted Diagnosis Model via Transformer
Huihui Fang, Fangxin Shang, Dalu Yang, Zhaowei Wang 0004, Yehui Yang, Yanwu Xu 0001 |
MICCAI (2) | 7 |
| 2022 | Learning Self-calibrated Optic Disc and Cup Segmentation from Multi-rater Annotations
Huihui Fang, Zhaowei Wang 0004, Dalu Yang, Yehui Yang, Fangxin Shang, Wenshuo Zhou, Yanwu Xu 0001 |
MICCAI (2) | 5 |
| 2022 | Opinions Vary? Diagnosis First!
Huihui Fang, Dalu Yang, Zhaowei Wang 0004, Wenshuo Zhou, Fangxin Shang, Yehui Yang, Yanwu Xu 0001 |
MICCAI (2) | 7 |
| 2022 | Spherical Interpolated Convolutional Network With Distance-Feature Density for 3-D Semantic Segmentation of Point CloudsabstractThe semantic segmentation of point clouds is an important part of the environment perception for robots. However, it is difficult to directly adopt the traditional 3-D convolution kernel to extract features from raw 3-D point clouds because of the unstructured property of point clouds. In this article, a spherical interpolated convolution operator is proposed to replace the traditional grid-shaped 3-D convolution operator. In addition, this article analyzes the defect of point cloud interpolation methods based on the distance as the interpolation weight and proposes the self-learned distance-feature density by combining the distance and the feature correlation. The proposed method makes the feature extraction of the spherical interpolated convolution network more rational and effective. The effectiveness of the proposed network is demonstrated on the 3-D semantic segmentation task of point clouds. Experiments show that the proposed method achieves good performance on the ScanNet dataset and Paris-Lille-3D dataset. The comparison experiments with the traditional grid-shaped 3-D convolution operator demonstrated that the newly proposed feature extraction operator improves the accuracy of the network and reduces the parameters of the network. The source codes will be released on https://github.com/IRMVLab/SIConv. Guangming Wang 0001, Yehui Yang, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Robust Collaborative Learning of Patch-Level and Image-Level Annotations for Diabetic Retinopathy Grading From Fundus ImageabstractDiabetic retinopathy (DR) grading from fundus images has attracted increasing interest in both academic and industrial communities. Most convolutional neural network-based algorithms treat DR grading as a classification task via image-level annotations. However, these algorithms have not fully explored the valuable information in the DR-related lesions. In this article, we present a robust framework, which collaboratively utilizes patch-level and image-level annotations, for DR severity grading. By an end-to-end optimization, this framework can bidirectionally exchange the fine-grained lesion and image-level grade information. As a result, it exploits more discriminative features for DR grading. The proposed framework shows better performance than the recent state-of-the-art algorithms and three clinical ophthalmologists with over nine years of experience. By testing on datasets of different distributions (such as label and camera), we prove that our algorithm is robust when facing image quality and distribution variations that commonly exist in real-world practice. We inspect the proposed framework through extensive ablation studies to indicate the effectiveness and necessity of each motivation. The code and some valuable annotations are now publicly available. Yehui Yang, Fangxin Shang, Binghong Wu, Dalu Yang, Yanwu Xu 0001, Wensheng Zhang 0002, Tianzhu Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Residual-CycleGAN Based Camera Adaptation for Robust Diabetic Retinopathy Screening
Dalu Yang, Yehui Yang, Tiantian Huang, Binghong Wu, Yanwu Xu 0001 |
MICCAI (2) | 2 |
| 2019 | Multi-task Character-Level Attentional Networks for Medical Concept Normalization
Jinghao Niu, Yehui Yang, Siheng Zhang, Zhengya Sun, Wensheng Zhang 0002 |
Neural Process. Lett. | 2 |
| 2017 | Lesion Detection and Grading of Diabetic Retinopathy via Two-Stages Deep Convolutional Neural Networks
Yehui Yang, Wensi Li, Haishan Wu, Wensheng Zhang 0002 |
MICCAI (3) | 1 |
| 2017 | Discriminative Reverse Sparse Tracking via Weighted Multitask LearningabstractMultitask learning has shown great potentiality for visual tracking under a particle filter framework. However, the recent multitask trackers, which exploit the similarity between all candidates by imposing group sparsity on the candidate representations, have a limitation in robustness due to the diverse sampling of candidates. To deal with this issue, we propose a discriminative reverse sparse tracker via weighted multitask learning. Our positive and negative templates are retained from the target observations and the background, respectively. Here, the templates are reversely represented via the candidates, and the representation of each positive template is viewed as a single task. Compared with existing multitask trackers, the proposed algorithm has the following advantages. First, we regularize the target representations with the ℓ2,1-norm to exploit the similarity shared by the positive templates, which is reasonable because of the target appearance consistency in the tracking process. Second, the valuable prior relationship between the candidates and the templates is introduced into the representation model by a weighted multitask learning scheme. Third, both target information and background information are integrated to generate discriminative scores for enhancing the proposed tracker. The experimental results on challenging sequences show that the proposed algorithm is effective and performs favorably against 12 state-of-the-art trackers. Yehui Yang, Wenrui Hu, Wensheng Zhang 0002, Tianzhu Zhang 0001, Yuan Xie 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Temporal Restricted Visual Tracking Via Reverse-Low-Rank Sparse LearningabstractAn effective representation model, which aims to mine the most meaningful information in the data, plays an important role in visual tracking. Some recent particle-filter-based trackers achieve promising results by introducing the low-rank assumption into the representation model. However, their assumed low-rank structure of candidates limits the robustness when facing severe challenges such as abrupt motion. To avoid the above limitation, we propose a temporal restricted reverse-low-rank learning algorithm for visual tracking with the following advantages: 1) the reverse-low-rank model jointly represents target and background templates via candidates, which exploits the low-rank structure among consecutive target observations and enforces the temporal consistency of target in a global level; 2) the appearance consistency may be broken when target suffers from sudden changes. To overcome this issue, we propose a local constraint via 11,2 mixed-norm, which can not only ensures the local consistency of target appearance, but also tolerates the sudden changes between two adjacent frames; and 3) to alleviate the inference of unreasonable representation values due to outlier candidates, an adaptive weighted scheme is designed to improve the robustness of the tracker. By evaluating on 26 challenge video sequences, the experiments show the effectiveness and favorable performance of the proposed algorithm against 12 state-of-the-art visual trackers. Yehui Yang, Wenrui Hu, Yuan Xie 0006, Wensheng Zhang 0002, Tianzhu Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Moving Object Detection Using Tensor-Based Low-Rank and Saliently Fused-Sparse DecompositionabstractIn this paper, we propose a new low-rank and sparse representation model for moving object detection. The model preserves the natural space-time structure of video sequences by representing them as three-way tensors. Then, it operates the low-rank background and sparse foreground decomposition in the tensor framework. On the one hand, we use the tensor nuclear norm to exploit the spatio-temporal redundancy of background based on the circulant algebra. On the other, we use the new designed saliently fused-sparse regularizer (SFS) to adaptively constrain the foreground with spatio-temporal smoothness. To refine the existing foreground smooth regularizers, the SFS incorporates the local spatio-temporal geometric structure information into the tensor total variation by using the 3D locally adaptive regression kernel (3D-LARK). What is more, the SFS further uses the 3D-LARK to compute the space-time motion saliency of foreground, which is combined with the l1norm and improves the robustness of foreground extraction. Finally, we solve the proposed model with globally optimal guarantee. Extensive experiments on challenging well-known data sets demonstrate that our method significantly outperforms the state-of-the-art approaches and works effectively on a wide range of complex scenarios. Wenrui Hu, Yehui Yang, Wensheng Zhang 0002, Yuan Xie 0006 |
IEEE Trans. Image Process. | 2 |
| 2017 | The Twist Tensor Nuclear Norm for Video CompletionabstractIn this paper, we propose a new low-rank tensor model based on the circulant algebra, namely, twist tensor nuclear norm (t-TNN). The twist tensor denotes a three-way tensor representation to laterally store 2-D data slices in order. On one hand, t-TNN convexly relaxes the tensor multirank of the twist tensor in the Fourier domain, which allows an efficient computation using fast Fourier transform. On the other, t-TNN is equal to the nuclear norm of block circulant matricization of the twist tensor in the original domain, which extends the traditional matrix nuclear norm in a block circulant way. We test the t-TNN model on a video completion application that aims to fill missing values and the experiment results validate its effectiveness, especially when dealing with video recorded by a nonstationary panning camera. The block circulant matricization of the twist tensor can be transformed into a circulant block representation with nuclear norm invariance. This representation, after transformation, exploits the horizontal translation relationship between the frames in a video, and endows the t-TNN model with a more powerful ability to reconstruct panning videos than the existing state-of-the-art low-rank models. Wenrui Hu, Dacheng Tao, Wensheng Zhang 0002, Yuan Xie 0006, Yehui Yang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Global Coupled Learning and Local Consistencies Ensuring for sparse-based tracking
Yehui Yang, Yuan Xie 0006, Wensheng Zhang 0002, Wenrui Hu, Yuanhua Tan |
Neurocomputing | 1 |