EDBT 2026 Demo / reviewers in the wild / expert
Yu-Cheng Chou
dblp:73/7569
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 46% Knowledge representation and reasoning · 21% Generative modeling · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning |
1.7 | 2 | 2025 | SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning · NeurIPS 2025 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmark · ICCV 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.7 | 2 | 2025 | SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning · NeurIPS 2025 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmark · ICCV 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion Models · NeurIPS 2025 |
Computer vision › Vision and language
spatial reasoning benchmark |
0.9 | 1 | 2025 | 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmark · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.9 | 1 | 2025 | Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion Models · NeurIPS 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion Models · NeurIPS 2025 |
Medical and health informatics › medical imaging
medical image analysis |
0.8 | 1 | 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation? · NeurIPS 2024 |
Medical and health informatics › medical imaging › medical image analysis
medical image segmentation |
0.8 | 1 | 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation? · NeurIPS 2024 |
Computer vision › Vision and language
image captioning |
0.3 | 1 | 2025 | Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion Models · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
out-of-distribution evaluation |
0.2 | 1 | 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation? · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
visual question answering evaluation · 0.9large vision-language model · 0.9knowledge distillation · 0.9flipeval · 0.9continuous embeddings · 0.9chain-of-thought reasoning · 0.9autoencoder · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmarkabstract3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/. Wufei Ma, Guofeng Zhang 0025, Yu-Cheng Chou, Jieneng Chen, Celso de Melo, Alan L. Yuille |
ICCV | 4 |
| 2025 | SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningabstractDespite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning. Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, Alan L. Yuille |
NeurIPS | 2 |
| 2025 | Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion ModelsabstractBuilding state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-quality image-text pairs, requiring millions of GPU hours. This paper introduces the Vision-Language-Vision **(VLV)** auto-encoder framework, which strategically leverages key pretrained components: a vision encoder, the decoder of a Text-to-Image (T2I) diffusion model, and subsequently, a Large Language Model (LLM). Specifically, we establish an information bottleneck by regularizing the language representation space, achieved through freezing the pretrained T2I diffusion decoder. Our VLV pipeline effectively distills knowledge from the text-conditioned diffusion model using continuous embeddings, demonstrating comprehensive semantic understanding via high-quality reconstructions. Furthermore, by fine-tuning a pretrained LLM to decode the intermediate language representations into detailed descriptions, we construct a state-of-the-art (SoTA) captioner comparable to leading models like GPT-4o and Gemini 2.0 Flash. Our method demonstrates exceptional cost-efficiency and significantly reduces data requirements; by primarily utilizing single-modal images for training and maximizing the utility of existing pretrained models (image encoder, T2I diffusion model, and LLM), it circumvents the need for massive paired image-text datasets, keeping the total training expenditure under $1,000 USD. Tiezheng Zhang, Yu-Cheng Chou, Jieneng Chen, Alan L. Yuille, Junfei Xiao |
NeurIPS | 3 |
| 2024 | Embracing Massive Medical Data
Yu-Cheng Chou, Zongwei Zhou, Alan L. Yuille |
MICCAI (1) | 1 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 7 |
| 2024 | TRCC: Transferable Congestion Control With Reinforcement LearningabstractThe breathtaking progress in machine learning has motivated studies in learning-based congestion control algorithms, which are expected to adjust congestion window choices according to the dynamic network environment. Nevertheless, existing learning-based congestion control protocols are mostly designed and trained for a specific network environment. When applied to a different network environment, the previously trained model may see a considerable degradation in performance. As the rapid development of communication technologies has given rise to the emergence of a diversity of new networks, it is desirable for learning-based congestion control models to quickly transfer to different network environments. Driven by this motivation, we propose a novel TRansferable Congestion Control (TRCC) protocol, which takes full advantage of both reinforcement learning and transfer learning to intelligently cope with network congestion in scenarios. The key idea to enable a fast transfer from the source network environment to the target network environment is to fine-tune a well-trained model in the source network to suit the target network in a time-effective way. We theoretically prove the transferability and quick convergence of our proposed transfer reinforcement learning-based congestion control algorithm by deriving the Markov transition matrix and the similarity of the reward function. Our experiments validate that TRCC can converge in a new network environment in a short time while achieving comparable performance with baseline algorithms. Zhicong Zheng, Zhenchang Xia, Yu-Cheng Chou, Yanjiao Chen |
IEEE Internet Things J. | 3 |
| 2021 | Guidance And Teaching Network For Video Salient Object DetectionabstractOwing to the difficulties of mining spatial-temporal cues, the existing approaches for video salient object detection (VSOD) are limited in understanding complex and noisy scenarios, and often fail in inferring prominent objects. To alleviate such shortcomings, we propose a simple yet efficient architecture, termed Guidance and Teaching Network (GTNet), to independently distill effective spatial and temporal cues with implicit guidance and explicit teaching at feature- and decision-level, respectively. To be specific, we (a) introduce a temporal modulator to implicitly bridge features from motion into appearance branch, which is capable of fusing cross-modal features collaboratively, and (b) utilise motion-guided mask to propagate the explicit cues during the feature aggregation. This novel learning strategy achieves satisfactory results via decoupling the complex spatial-temporal cues and mapping informative cues across different modalities. Extensive experiments on three challenging benchmarks show that the proposed method can run at$\sim$28 fps on a single TITAN Xp GPU and perform competitively against 14 cutting-edge baselines. Yingxia Jiao, Yu-Cheng Chou, Shouyuan Yang, Ge-Peng Ji |
ICIP | 3 |
| 2021 | A Multi-objective Reinforcement Learning Perspective on Internet Congestion ControlabstractThe advent of new network architectures has resulted in the rise of network applications with different network performance requirements: live video streaming applications require low latency. In contrast, file transfer applications require high throughput. Existing congestion control protocols may fail to simultaneously meet the performance requirements of these different types of applications since their designed objective function is fixed and difficult to readjust according to the needs of the application. In this paper, we develop MOCC (Multi-Objective Congestion Control), a novel multi-objective congestion control protocol that can meet the performance requirements of different applications without the need to redesign the objective function. MOCC leverages multi-objective reinforcement learning with preferences in order to adapt to different types of applications. By addressing challenges such as slow convergence speed and the difficulty of designing the end of the episode, MOCC can quickly converge to the equilibrium point and adapt multi-objective reinforcement learning to congestion control. Through an extensive array of experiments, we discover that MOCC outperforms the most recent state-of-the-art congestion control protocols and can achieve a trade-off between throughput, latency, and packet loss, meeting the performance requirements of different types of applications by setting preferences. Zhenchang Xia, Yanjiao Chen, Yu-Cheng Chou, Zhicong Zheng, Baochun Li |
IWQoS | 4 |
| 2021 | Progressively Normalized Self-Attention Network for Video Polyp Segmentation
Ge-Peng Ji, Yu-Cheng Chou, Deng-Ping Fan, Geng Chen 0001, Huazhu Fu, Debesh Jha, Ling Shao 0001 |
MICCAI (1) | 2 |
| 2018 | A Clonal Selection Algorithm for Energy-Efficient Mobile Agent Itinerary Planning in Wireless Sensor Networks
Yu-Cheng Chou, Madoka Nakajima |
Mob. Networks Appl. | 1 |
| 2014 | Machine Learning Approaches for Interactive Verification
Yu-Cheng Chou, Hsuan-Tien Lin |
PAKDD (2) | 1 |
| 2010 | An embeddable mobile agent platform supporting runtime code mobility, interaction and coordination of mobile agents and host systems
Yu-Cheng Chou, David Ko, Harry H. Cheng |
Inf. Softw. Technol. | 1 |
| 2009 | Mobile agent-based computational steering for distributed applicationsabstractAbstract The mobile agent‐based computational steering (MACS) for distributed applications is presented in this article. In the MACS, a mobile agent platform, Mobile‐C, is embedded in a program through the Mobile‐C library to support C/C++ mobile agent code. Runtime replaceable algorithms of a program are represented as agent services in C/C++ source code and can be replaced with new ones through mobile agents. In the MACS, a mobile agent created and deployed by a user from the steering host migrates to computing hosts successively to replace algorithms of running programs that constitute a distributed application without the need of stopping the execution and recompiling the programs. The methodology of dynamic algorithm alteration in the MACS is described in detail with an example of matrix operation. The Mobile‐C library enables the integration of Mobile‐C into any C/C++ programs to carry out computational steering through mobile agents. The source code level execution of mobile agent code facilitates handling issues such as portability and secure execution of mobile agent code. In the MACS, the network load between the steering and computing hosts can be reduced, and the successive operations of a mobile agent on multiple computing hosts are not affected whether the steering host stays online or not. The employment of the middle‐level language C/C++ enables the MACS to accommodate the diversity of scientific and engineering fields to allow for runtime interaction and steering of distributed applications to match the dynamic requirements imposed by the user or the execution environment. An experiment is used to validate the feasibility of the MACS in real‐world mobile robot applications. The experiment replaces a mobile robot's behavioral algorithm with a mobile agent at runtime. Copyright © 2009 John Wiley & Sons, Ltd. Yu-Cheng Chou, David Ko, Harry H. Cheng |
Concurr. Comput. Pract. Exp. | 1 |