Ruiyang Zhang

dblp:242/1802 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 70% Trustworthy machine learning · 19% Vision and language · 6%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.622025
Harnessing Uncertainty-Aware Bounding Boxes for Unsupervised 3D Object Detection · ICCV 2025
Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene · ECCV (11) 2024
Computer vision › 3D vision › 3d object detection › label-efficient 3d object detection
unsupervised 3d object detection
1.622025
Harnessing Uncertainty-Aware Bounding Boxes for Unsupervised 3D Object Detection · ICCV 2025
Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene · ECCV (11) 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty-aware detection
0.912025
Harnessing Uncertainty-Aware Bounding Boxes for Unsupervised 3D Object Detection · ICCV 2025
Information retrieval
retrieval-augmented generation
0.912025
MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval · EMNLP 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval · EMNLP 2025
Robotics › Autonomous driving
perception
0.312025
Harnessing Uncertainty-Aware Bounding Boxes for Unsupervised 3D Object Detection · ICCV 2025

Methods — techniques the papers use, named apart from their topics

page graph construction · 1.7graph traversal · 1.7fine-tuning · 1.7uncertainty regularization · 0.9uncertainty estimation · 0.9pseudo-labeling · 0.9self-supervised learning · 0.82d scene scaling · 0.8
YearPublicationVenuePosition
2026 Uncertainty-Aware Cross-Scale Hand-Eye Calibration of 2-D Optical Coherence Tomography Using a Plane Target
Haitian Lyu, Jiewen Lai, Ruiyang Zhang, Wu Yuan 0001, Hongliang Ren 0001
IEEE Trans. Robotics4
2025 MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval
abstract
Document Understanding is a foundational AI capability with broad applications, and Document Question Answering (DocQA) is a key evaluation task.Traditional methods convert the document into text for processing by Large Language Models (LLMs), but this process strips away critical multi-modal information like figures.While Large Vision-Language Models (LVLMs) address this limitation, their constrained input size makes multi-page document comprehension infeasible.Retrievalaugmented generation (RAG) methods mitigate this by selecting relevant pages, but they rely solely on semantic relevance, ignoring logical connections between pages and the query, which is essential for reasoning.To this end, we propose MoLoRAG, a logicaware retrieval framework for multi-modal, multi-page document understanding.By constructing a page graph that captures contextual relationships between pages, a lightweight VLM performs graph traversal to retrieve relevant pages, including those with logical connections often overlooked.This approach combines semantic and logical relevance to deliver more accurate retrieval.After retrieval, the top-K pages are fed into arbitrary LVLMs for question answering.To enhance flexibility, MoLoRAG offers two variants: a training-free solution for easy deployment and a fine-tuned version to improve logical relevance checking.Experiments on four DocQA datasets demonstrate average improvements of 9.68% in accuracy over LVLM direct inference and 7.44% in retrieval precision over baselines.Codes and datasets are released at https://github.com/WxxShirley/MoLoRAG.
Xixi Wu, Yanchao Tan, Nan Hou, Ruiyang Zhang, Hong Cheng 0001
EMNLP4
2025 Harnessing Uncertainty-Aware Bounding Boxes for Unsupervised 3D Object Detection
abstract
Unsupervised 3D object detection aims to identify objects of interest from unlabeled raw data, such as LiDAR points. Recent approaches usually adopt pseudo 3D bounding boxes (3D bboxes) from clustering algorithm to initialize the model training. However, pseudo bboxes inevitably contain noise, and such inaccuracies accumulate to the final model, compromising the performance. Therefore, in an attempt to mitigate the negative impact of inaccurate pseudo bboxes, we introduce a new uncertainty-aware framework for unsupervised 3D object detection, dubbed UA3D. In particular, our method consists of two phases: uncertainty estimation and uncertainty regularization. (1) In the uncertainty estimation phase, we incorporate an extra auxiliary detection branch alongside the original primary detector. The prediction disparity between the primary and auxiliary detectors could reflect fine-grained uncertainty at the box coordinate level. (2) Based on the assessed uncertainty, we adaptively adjust the weight of every 3D bbox coordinate via uncertainty regularization, refining the training process on pseudo bboxes. For pseudo bbox coordinate with high uncertainty, we assign a relatively low loss weight. Extensive experiments verify that the proposed method is robust against the noisy pseudo bboxes, yielding substantial improvements on nuScenes and Lyft compared to existing approaches, with increases of +6.9% AP$_{BEV}$ and +2.5% AP$_{3D}$ on nuScenes, and +4.1% AP$_{BEV}$ and +2.0% AP$_{3D}$ on Lyft.
Ruiyang Zhang, Hu Zhang 0005, Zhedong Zheng
ICCV1
2024 Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
Ruiyang Zhang, Hu Zhang 0005, Hang Yu 0006, Zhedong Zheng
ECCV (11)1
2023 Large AI Models in Health Informatics: Applications, Challenges, and the Future
abstract
Large AI models, or foundation models, are models recently emerging with massive scales both parameter-wise and data-wise, the magnitudes of which can reach beyond billions. Once pretrained, large AI models demonstrate impressive performance in various downstream tasks. A prime example is ChatGPT, whose capability has compelled people's imagination about the far-reaching influence that large AI models can have and their potential to transform different domains of our lives. In health informatics, the advent of large AI models has brought new paradigms for the design of methodologies. The scale of multi-modal data in the biomedical and health domain has been ever-expanding especially since the community embraced the era of deep learning, which provides the ground to develop, validate, and advance large AI models for breakthroughs in health-related areas. This article presents a comprehensive review of large AI models, from background to their applications. We identify seven key sectors in which large AI models are applicable and might have substantial influence, including: 1) bioinformatics; 2) medical diagnosis; 3) medical imaging; 4) medical informatics; 5) medical education; 6) public health; and 7) medical robotics. We examine their challenges, followed by a critical discussion about potential future directions and pitfalls of large AI models in transforming the field of health informatics.
Jianing Qiu, Lin Li 0070, Jiankai Sun, Jiachuan Peng, Peilun Shi, Ruiyang Zhang, Yinzhao Dong, Kyle Lam, Frank P.-W. Lo, Bo Xiao 0002, Wu Yuan 0001, Ningli Wang, Dong Xu 0002, Benny P. L. Lo
IEEE J. Biomed. Health Informatics6
2022 Dynamic hidden variable fuzzy broad neural network based batch process anomaly detection with incremental learning capabilities
Peng Chang 0001, Ruiyang Zhang, Ding ChunHao
Expert Syst. Appl.2