VLDB 2026 Research / reviewers in the wild / expert
Yiwen Luo
dblp:96/4586
· DBLP profile ↗
20ranked-venue papers
10as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lora: Towards Improved Applicability of Reconfigurable Architecture for Versatile Nonlinear Functions
Yuan Dai, Guibin Zou, Yuanda Yang, Jiahang Lou, Yiwen Luo, Xinyu Cai, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ISCA | 6 |
| 2025 | Enhancing Multimodal Chain-of-Thought Reasoning with Tree-Searched Self-TrainingabstractDespite impressive performance on general visual benchmarks, Multimodal Large Language Models (MLLMs) still struggle with generating consistent and accurate reasoning processes for complex visual reasoning tasks. The limited availability of multimodal reasoning datasets further complicates fine-tuning efforts to improve reasoning capabilities. To address these challenges, we introduce Tree-Searched Self-Training (TSST), a novel framework that enhances multimodal reasoning through self-improvement without relying on extensive manual annotations. TSST introduces a hierarchical tree-search mechanism that combines stepwise rationales generation with value-guided path selection, enabling the model to explore and identify high-quality reasoning trajectories autonomously. Training on these self-generated data significantly improves the model’s reasoning ability, boosting consistency and accuracy. Extensive experiments on ScienceQA and VCR benchmarks show that TSST outperforms existing fine-tuning baselines by 6.0% in accuracy, demonstrating its effectiveness in enhancing multimodal reasoning capabilities. The code is publicly available at: https://github.com/laaambs/tsst. Yiwen Luo, Yong Luo 0002, Zengmao Wang |
ICME | 1 |
| 2025 | Probabilistic Prior-Guided Anatomical Alignment for MRI Super-Resolution
Yiwen Luo, Xiaoying Tang 0001, Yixuan Yuan |
MICCAI (4) | 1 |
| 2025 | LLM-Guided Decoupled Probabilistic Prompt for Continual Learning in Medical Image DiagnosisabstractDeep learning-based traditional diagnostic models typically exhibit limitations when applied to dynamic clinical environments that require handling the emergence of new diseases. Continual learning (CL) offers a promising solution, aiming to learn new knowledge while preserving previously learned knowledge. Though recent rehearsal-free CL methods employing prompt tuning (PT) have shown promise, they rely on deterministic prompts that struggle to handle diverse fine-grained knowledge. Moreover, existing PT methods utilize randomly initialized prompts that are trained under standard classification constraints, impeding expert knowledge integration and optimal performance acquisition. In this paper, we propose an LLM-guided Decoupled Probabilistic Prompt (LDPP) for Continual Learning in medical image diagnosis. Specifically, we develop an Expert Knowledge Generation (EKG) module that leverages LLM to acquire decoupled expert knowledge and comprehensive category descriptions. Then, we introduce a Decoupled Probabilistic Prompt pool (DePP) to construct a shared decoupled probabilistic prompt pool, which constructs a shared prompt pool with probabilistic prompts derived from the expert knowledge set. These prompts dynamically provide diverse and flexible descriptions for input images. Finally, We design a Steering Prompt Pool (SPP) to enhance intra-class compactness and promote model performance by learning non-shared prompts. With extensive experimental validation, LDPP consistently sets state-of-the-art performance under the challenging class-incremental setting in CL. Code is available at: https://github.com/CUHK-AIM-Group/LDPP. Yiwen Luo, Wuyang Li, Xiang Li 0001, Tianming Liu 0001, Tianye Niu, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Rich Human Feedback for Text-to-Image GenerationabstractRecent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However, many generated images still suffer from issues such as artifacts/implausibility, misalignment with text descriptions, and low aesthetic quality. Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) for large language models, prior works collected human-provided scores as feedback on generated images and trained a reward model to improve the T2I generation. In this paper, we enrich the feedback signal by (i) marking image regions that are implausible or misaligned with the text, and (ii) annotating which words in the text prompt are misrepresented or missing on the image. We collect such rich human feedback on 18K generated images (RichHF-18K) and train a multimodal transformer to predict the rich feedback automatically. We show that the predicted rich human feedback can be leveraged to improve image generation, for example, by selecting high-quality training data to finetune and improve the generative models, or by creating masks with predicted heatmaps to inpaint the problematic regions. Notably, the improvements generalize to models (Muse) beyond those used to generate the images on which human feedback data were collected (Stable Diffusion variants). The RichHF-18K data set will be released in our GitHub repository: https://github.com/google-research/google-research/tree/master/richhf_18k. Youwei Liang, Junfeng He, Gang Li 0021, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang 0008, Junjie Ke, Krishnamurthy Dvijotham, Katie Collins, Yiwen Luo, Yang Li 0058, Kai Kohlhoff, Deepak Ramachandran, Vidhya Navalpakkam |
CVPR | 14 |
| 2024 | Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?abstractThis paper investigates an under-explored challenge in large language models (LLMs): chain-of-thought prompting with noisy rationales, which include irrelevant or inaccurate reasoning thoughts within examples used for in-context learning. We construct NoRa dataset that is tailored to evaluate the robustness of reasoning in the presence of noisy rationales. Our findings on NoRa dataset reveal a prevalent vulnerability to such noise among current LLMs, with existing robust methods like self-correction and self-consistency showing limited efficacy. Notably, compared to prompting with clean rationales, base LLM drops by 1.4%-19.8% in accuracy with irrelevant thoughts and more drastically by 2.2%-40.4% with inaccurate thoughts.
Addressing this challenge necessitates external supervision that should be accessible in practice. Here, we propose the method of contrastive denoising with noisy chain-of-thought (CD-CoT). It enhances LLMs' denoising-reasoning capabilities by contrasting noisy rationales with only one clean rationale, which can be the minimal requirement for denoising-purpose prompting. This method follows a principle of exploration and exploitation: (1) rephrasing and selecting rationales in the input space to achieve explicit denoising and (2) exploring diverse reasoning paths and voting on answers in the output space. Empirically, CD-CoT demonstrates an average improvement of 17.8% in accuracy over the base model and shows significantly stronger denoising capabilities than baseline methods. The source code is publicly available at: https://github.com/tmlr-group/NoisyRationales. Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo, Zengmao Wang, Bo Han 0003 |
NeurIPS | 4 |
| 2024 | Dynamic Attribute-guided Few-shot Open-set Network for medical image diagnosis
Yiwen Luo, Xiaoqing Guo, Li Liu 0017, Yixuan Yuan |
Expert Syst. Appl. | 1 |
| 2024 | Collaborative Intelligent Delivery With One Truck and Multiple Heterogeneous Drones in COVID-19 Pandemic EnvironmentabstractThe outbreak of COVID-19 has caused a serious impact on the traditional logistics industry. Considering that the truck-drone collaborative delivery system can both reduce the risk of COVID-19 propagation and deliver supplies in a cost effective and timely manner, this paper introduces the Multiple visits Travelling Salesman Problem with Multiple Heterogeneous Drones (MTSP-MHD). The model allows a truck to carry a fleet of heterogeneous multi-visit drones for cooperative deliveries, where the drones are capable of delivering to multiple customers on a single route and the flight is restricted by energy consumption and payload constraints. To solve MTSP-MHD, we develop an approach that combines K-Means$++$clustering, Nearest neighbor search and Greedy strategies (KNG) to construct feasible solutions. Meanwhile, an Improved Artificial Bee Colony algorithm combining Metropolis acceptance criterion of Simulated Annealing, Tabu list of Tabu Search, and Elite selection strategies (IABC-MTE) is proposed to enhance the quality of solutions. Particularly, three problem-specific neighborhood operators are adopted to search for new solutions. The massive experimental results indicate that IABC-MTE achieves significant improvements over other competitors, with average objective value reductions ranging from 1.81% to 29.16% and standard deviations reduced by 0.04 to 26.44. Finally, the influencing factors of the drone fleet, the performance of different drone fleets and delivery modes are evaluated in detail. Yiwen Luo, Xiaoheng Deng, Yan Ke, Shaohua Wan 0001, Yurong Qian |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Semantic-guided Unknown-aware Rare Disease Diagnosis SystemabstractThe significant challenge posed by rare disease diagnoses has recently motivated researchers to explore computer-aided solutions. While deep learning approaches have shown potential in developing automatic diagnosis systems, their effectiveness diminishes when addressing rare diseases with limited data. Moreover, existing diagnostic models typically fail to detect unknown diseases, making them inappropriate for real-world applications. In this paper, we address the above challenges by proposing a Semantic-guided unknown-aware Rare Disease Diagnosis (SRDD) model. SRDD aims to tackle the performance degradation of classification models in low-data regimes, as well as their inability to distinguish unknown diseases. Specifically, we propose a Semantic-guided Saliency Discovery (SSD) module to explore semantic saliency information within images by aligning image regions with the semantic knowledge embedded within category labels. The image content is subsequently decomposed into semantically related information (SRI) and image instance template (IIT). Then, we design a Reciprocal samples Synthetic Strategy (RSS) to create known and unknown reciprocal points using SRI and IIT. This facilitates a compact feature space for known classes while preserving space for unknown data, promoting accurate known disease diagnosis and unknown disease detection. We validate SRDD on the public skin disease dataset SD-260. SRDD achieves state-of-the-art performance in both known disease classification and unknown disease identification. Yiwen Luo, Yixuan Yuan |
BIBM | 1 |
| 2023 | PVDet: Towards pedestrian and vehicle detection on gigapixel-level images
Wanghao Mo, Hongyang Wei, Ruyi Cao, Yan Ke, Yiwen Luo |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | A Reconstruction-Based Visual-Acoustic-Semantic Embedding Method for Speech-Image RetrievalabstractSpeech-image retrieval aims at learning the relevance between image and speech.Prior approaches are mainly based on bi-modal contrastive learning, which can not alleviate the cross-modal heterogeneous issue between visual and acoustic modalities well. To address this issue, we propose a visual-acoustic-semantic embedding (VASE) method. First, we propose a tri-modal ranking loss by taking advantage of semantic information corresponding to the acoustic data, which introduces the auxiliary alignment to enhance the alignment between image and speech. Second, we introduce a cycle-consistency loss based on feature reconstruction. It can further alleviate the heterogeneous issue between different data modalities (e.g., visual-acoustic, visual-textual and acoustic-textual). Extensive experiments have demonstrated the effectiveness of our proposed method. In addition, our VASE model achieves state-of-the-art performance on the speech-image retrieval task on the Flickr8K [Harwath and Glass, 2015]s and Places [Harwathet al., 2018] datasets. Wei Tang 0016, Yan Huang 0008, Yiwen Luo, Liang Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Efficient video annotation with visual interpolation and frame selection guidanceabstractWe introduce a unified framework for generic video annotation with bounding boxes. Video annotation is a longstanding problem, as it is a tedious and time-consuming process. We tackle two important challenges of video annotation: (1) automatic temporal interpolation and extrapolation of bounding boxes provided by a human annotator on a subset of all frames, and (2) automatic selection of frames to annotate manually. Our contribution is two-fold: first, we propose a model that has both interpolating and extrapolating capabilities; second, we propose a guiding mechanism that sequentially generates suggestions for what frame to annotate next, based on the annotations made previously. We extensively evaluate our approach on several challenging datasets in simulation and demonstrate a reduction in terms of the number of manual bounding boxes drawn by 60% over linear interpolation and by 35% over an off-the-shelf tracker. Moreover, we also show 10% annotation time improvement over a state-of-the-art method for video annotation with bounding boxes [25]. Finally, we run human annotation experiments and provide extensive analysis of the results, showing that our approach reduces actual measured annotation time by 50% compared to commonly used linear interpolation. Alina Kuznetsova, Aakrati Talati, Yiwen Luo, Keith Simmons, Vittorio Ferrari |
WACV | 3 |
| 2021 | Accurate Retinal Vessel Segmentation in Color Fundus Images via Fully Attention-Based NetworksabstractAutomatic retinal vessel segmentation is important for the diagnosis and prevention of ophthalmic diseases. The existing deep learning retinal vessel segmentation models always treat each pixel equally. However, the multi-scale vessel structure is a vital factor affecting the segmentation results, especially in thin vessels. To address this crucial gap, we propose a novel Fully Attention-based Network (FANet) based on attention mechanisms to adaptively learn rich feature representation and aggregate the multi-scale information. Specifically, the framework consists of the image pre-processing procedure and the semantic segmentation networks. Green channel extraction (GE) and contrast limited adaptive histogram equalization (CLAHE) are employed as pre-processing to enhance the texture and contrast of retinal blood images. Besides, the network combines two types of attention modules with the U-Net. We propose a lightweight dual-direction attention block to model global dependencies and reduce intra-class inconsistencies, in which the weights of feature maps are updated based on the semantic correlation between pixels. The dual-direction attention block utilizes horizontal and vertical pooling operations to produce the attention map. In this way, the network aggregates global contextual information from semantic-closer regions or a series of pixels belonging to the same object category. Meanwhile, we adopt the selective kernel (SK) unit to replace the standard convolution for obtaining multi-scale features of different receptive field sizes generated by soft attention. Furthermore, we demonstrate that the proposed model can effectively identify irregular, noisy, and multi-scale retinal vessels. The abundant experiments on DRIVE, STARE, and CHASE_DB1 datasets show that our method achieves state-of-the-art performance. Kaiqi Li, Xingqun Qi, Yiwen Luo, Zeyi Yao, Xiaoguang Zhou, Muyi Sun |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Lymph Node Metastasis Classification Based on Semi-Supervised Multi-View NetworkabstractLymphatic metastasis is one of the most common proliferation pathways of thyroid carcinoma. Accurate diagnosis of lymph nodes is of great significance to surgical planning and prognosis. Due to the continuous development of deep learning recently, computer-aided diagnosis (CAD) systems for thyroid cancer have made considerable progress, but the research on the effective diagnosis of lymphatic metastasis remains insufficient. Focusing on this issue, we propose a semi-supervised multi-view network to diagnose lymph node metastasis, which combines coarse-view and fine-view to obtain a more comprehensive description. This method consists of three parts as follows: 1) joint probabilistic labels of the nodule partition information are generated by fuzzy clustering and perform semi-supervised learning on coarse-view with real labels; 2) an attention mechanism based network for fine-view is designed to capture various differentiated local features in a pyramid manner; 3) the two parts are then combined to extract global and local features more effectively to derive more accurate diagnostic reasoning. Especially, the introduction of fuzzy logic greatly reduces the impact of the uncertainty of the generated labels, thereby ensuring the effectiveness of the pseudo-labels. Extensive experiments on our collected dataset demonstrate that the proposed method is more efficient than other state-of-the-art methods. Yiwen Luo, Jingmin Xin, Junqin Feng, Litao Ruan, Nanning Zheng 0001 |
BIBM | 1 |
| 2020 | Accurate Cell Segmentation in Digital Pathology Images via Attention Enforced NetworksabstractAutomatic cell segmentation is an essential step in the pipeline of computer-aided diagnosis (CAD), such as the detection and grading of breast cancer. Accurate segmentation of cells can not only assist the pathologists to make a more precise diagnosis, but also save much time and labor. However, this task suffers from stain variation, cell inhomogeneous intensities, background clutters and cells from different tissues. To address these issues, we propose an Attention Enforced Network (AENet), which is built on spatial attention module and channel attention module, to integrate local features with global dependencies and weight effective channels adaptively. Besides, we introduce a feature fusion branch to bridge high-level and low-level features. Finally, the marker controlled watershed algorithm is applied to post-process the predicted segmentation maps for reducing the fragmented regions. In the test stage, we present an individual color normalization method to deal with the stain variation problem. We evaluate this model on the MoNuSeg dataset. The quantitative comparisons against several prior methods demonstrate the superiority of our approach. Zeyi Yao, Kaiqi Li, Yiwen Luo, Xiaoguang Zhou, Muyi Sun, Guanhong Zhang |
ICPR | 3 |
| 2018 | A content-based goods image recommendation system
Li Yu 0005, Fang-Jian Han, Shaobing Huang, Yiwen Luo |
Multim. Tools Appl. | 4 |
| 2017 | Pipeline Implementation of Polyphase PSO for Adaptive Beamforming AlgorithmabstractAdaptive beamforming is a powerful technique for anti-interference, where searching and tracking optimal solutions are a great challenge. In this paper, a partial Particle Swarm Optimization (PSO) algorithm is proposed to track the optimal solution of an adaptive beamformer due to its great global searching character. Also, due to its naturally parallel searching capabilities, a novel Field Programmable Gate Arrays (FPGA) pipeline architecture using polyphase filter bank structure is designed. In order to perform computations with large dynamic range and high precision, the proposed implementation algorithm uses an efficient user-defined floating-point arithmetic. In addition, a polyphase architecture is proposed to achieve full pipeline implementation. In the case of PSO with large population, the polyphase architecture can significantly save hardware resources while achieving high performance. Finally, the simulation results are presented by cosimulation with ModelSim and SIMULINK. Shaobing Huang, Li Yu 0005, Fang-Jian Han, Yiwen Luo |
Wirel. Commun. Mob. Comput. | 4 |
| 2014 | Intelligent control and navigation of an indoor quad-copterabstractThis paper documents the development of a quad-copter with indoor control scheme and navigation capability. A stabilized flying control system including traditional PID controller, Raspberry Pi onboard flight computer and electronic speed controller was developed to provide the basic platform for the quad-copter. PID tuning was utilized optimize the performance. In addition, RGB and depth cameras and other sensors were deployed to enable remote semi-autonomous control. Yiwen Luo, Meng Joo Er, Li Ling Yong, Chiang-Ju Chien |
ICARCV | 1 |
| 2008 | MQSearch: image search by multi-class queryabstractImage search is becoming prevalent in web search as the number of digital photos grows exponentially on the internet. For a successful image search system, removing outliers in the top ranked results is a challenging task. Typical content based image search engines take an input image from one class as a query and compute relevance between the query and images in a database. The results often contain a large number of outliers, since these outliers may be similar to the query image in some way. In this paper we present a novel search scheme using query images from multiple classes. Instead of conducting query search for one image class at a time, we conduct multi-class query search jointly. By using several query classes that are similar to each other for multi-class query, we can utilize information across similar classes to fine tune the similarity measure to remove outliers. This strategy can be used for any information search application. In this work, we use content based image search to illustrate the concept. Yiwen Luo, Wei Liu 0005, Jianzhuang Liu, Xiaoou Tang |
CHI | 1 |
| 2008 | Photo and Video Quality Evaluation: Focusing on the Subject
Yiwen Luo, Xiaoou Tang |
ECCV (3) | 1 |