VLDB 2026 Research / reviewers in the wild / expert
Bo-Kai Ruan
dblp:254/3905
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-9847-3628ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 45% Language models and text generation · 16% Reinforcement learning · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.9 | 4 | 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback · NeurIPS 2025 Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025 TF-TI2I: Training-Free Text-And-Image-To-Image Generation via Multi-Modal Implicit-Context Learning in Text-To-Image Models · ICCV 2025 |
Machine learning › Graph learning
graph-based retrieval |
1.0 | 1 | 2026 | HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation · WWW 2026 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.0 | 1 | 2026 | HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation · WWW 2026 |
Knowledge graphs
knowledge graph reasoning |
1.0 | 1 | 2026 | HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation · WWW 2026 |
Knowledge graphs › knowledge graph reasoning
multi-hop reasoning |
1.0 | 1 | 2026 | HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation · WWW 2026 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.9 | 1 | 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback · NeurIPS 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback · NeurIPS 2025 |
Machine learning › Reinforcement learning
preference learning |
0.9 | 1 | 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.9 | 1 | 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025 |
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance |
0.9 | 1 | 2025 | Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025 |
Visual content generation and editing › image editing › image compositing
image harmonization |
0.9 | 1 | 2025 | Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References · AAAI 2025 |
Visual content generation and editing › image editing › image compositing › image harmonization
painterly image harmonization |
0.9 | 1 | 2025 | Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References · AAAI 2025 |
Computer vision › Face, body and person analysis
facial action unit recognition |
0.6 | 1 | 2022 | Mimicking the Annotation Process for Recognizing the Micro Expressions · ACM Multimedia 2022 |
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition |
0.6 | 1 | 2022 | Mimicking the Annotation Process for Recognizing the Micro Expressions · ACM Multimedia 2022 |
Computer vision › Face, body and person analysis › facial expression analysis › facial expression recognition
micro-expression recognition |
0.6 | 1 | 2022 | Mimicking the Annotation Process for Recognizing the Micro Expressions · ACM Multimedia 2022 |
Machine learning › Representation and self-supervised learning
text embedding |
0.3 | 1 | 2025 | Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.9beam search · 2.0zero-shot disentanglement · 1.7winner-takes-all · 1.7similarity reweighting · 1.7similarity disentangle mask · 1.7reference contextual masking · 1.7ranking optimization · 0.9inverse reinforcement learning · 0.9direct preference optimization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeCond: Free Lunch in the Input Conditions of Text-Guided InpaintingabstractText-to-image inpainting models often exhibit an unpredictable balance among image coherence and prompt adherence. This rigidity limits their adaptability across diverse scenarios, including coarse masks, non-object, and interaction prompts. Recognizing this instability as an indicator of learned generation diversity, we aim to control model behavior for given objective. We propose Empirical Feature Intervention (EFI), a metric-agnostic framework that precomputes how feature interventions influence evaluation metrics—such as CLIP, Human Preference Score (HPS), and Image Reward (IR). Building on EFI, we introduce FreeCond, a free-of-cost framework that applies two simple input interventions (Image Frequency and Mask Value Modulation), these interventions can be further optimized via Surrogate Intervention Optimization (SIO) based on a surrogate model regressed with precomputed EFI data. FreeCond enables real-time, user-interactive control of pretrained models without retraining or architectural modifications. Also, to benchmark performance on challenging settings, we present FCIBench. Experiments on EditBench, BrushBench, and FCIBench demonstrate that FreeCond substantially improves CLIP, HPS, and IR metrics by up to 22%, 8%, and 54%, respectively. Teng-Fang Hsiao, Bo-Kai Ruan, Sung-Lin Tsai, Yi-Lun Wu, Hong-Han Shuai |
WACV | 2 |
| 2026 | HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented GenerationabstractGraph-based Retrieval-Augmented Generation (RAG) typically operates on binary Knowledge Graphs (KGs). However, decomposing complex facts into binary triples often leads to semantic fragmentation and longer reasoning paths, increasing the risk of retrieval drift and computational overhead. In contrast, n-ary hypergraphs preserve high-order relational integrity, enabling shallower and more semantically cohesive inference. To exploit this topology, we propose HyperRAG, a framework tailored for n-ary hypergraphs featuring two complementary retrieval paradigms: (i) HyperRetriever learns structural-semantic reasoning over n-ary facts to construct query-conditioned relational chains. It enables accurate factual tracking, adaptive high-order traversal, and interpretable multi-hop reasoning under context constraints. (ii) HyperMemory leverages the LLM's parametric memory to guide beam search, dynamically scoring n-ary facts and entities for query-aware path expansion. Extensive evaluations on WikiTopics (11 closed-domain datasets) and three open-domain QA benchmarks (HotpotQA, MuSiQue, and 2WikiMultiHopQA) validate HyperRAG's effectiveness. HyperRetriever achieves the highest answer accuracy overall, with average gains of 2.95% in MRR and 1.23% in Hits@10 over the strongest baseline. Qualitative analysis further shows that HyperRetriever bridges reasoning gaps through adaptive and interpretable n-ary chain construction, benefiting both open and closed-domain QA. Our codes are publicly available at https://github.com/Vincent-Lien/HyperRAG.git. Wen-Sheng Lien, Yu-Kai Chan, Hao-Lung Hsiao, Bo-Kai Ruan, Meng-Fen Chiang, Chien-An Chen, Yi-Ren Yeh, Hong-Han Shuai |
WWW | 4 |
| 2025 | Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content ReferencesabstractPainterly image harmonization aims at seamlessly blending disparate visual elements within a single image. However, previous approaches often struggle due to limitations in training data or reliance on additional prompts, leading to inharmonious and content-disrupted output. To surmount these hurdles, we design a Training-and-prompt-Free General Painterly Harmonization method (TF-GPH). TF-GPH incorporates a novel “Similarity Disentangle Mask”, which disentangles the foreground content and background image by redirecting their attention to corresponding reference images, enhancing the attention mechanism for multi-image inputs. Additionally, we propose a “Similarity Reweighting” mechanism to balance harmonization between stylization and content preservation. This mechanism minimizes content disruption by prioritizing the content-similar features within the given background style reference. Finally, we address the deficiencies in existing benchmarks by proposing novel range-based evaluation metrics and a new benchmark to better reflect real-world applications. Extensive experiments demonstrate the efficacy of our method across benchmarks. Teng-Fang Hsiao, Bo-Kai Ruan, Hong-Han Shuai |
AAAI | 2 |
| 2025 | TF-TI2I: Training-Free Text-And-Image-To-Image Generation via Multi-Modal Implicit-Context Learning in Text-To-Image ModelsabstractText-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation. Existing methods often partially utilize image inputs, focusing on specific elements like objects or styles, or they experience a decline in generation quality with complex, multi-image instructions. To overcome these challenges, we introduce Training-Free Text-and-Image-to-Image (TF-TI2I), which adapts cutting-edge T2I models such as SD3 without the need for additional training. Our method capitalizes on the MM-DiT architecture, in which we point out that textual tokens can implicitly learn visual information from vision tokens. We enhance this interaction by extracting a condensed visual representation from reference images, facilitating selective information sharing through Reference Contextual Masking -- this technique confines the usage of contextual tokens to instruction-relevant visual information. Additionally, our Winner-Takes-All module mitigates distribution shifts by prioritizing the most pertinent references for each vision token. Addressing the gap in TI2I evaluation, we also introduce the FG-TI2I Bench, a comprehensive benchmark tailored for TI2I and compatible with existing T2I methods. Our approach shows robust performance across various benchmarks, confirming its effectiveness in handling complex image-generation tasks. Teng-Fang Hsiao, Bo-Kai Ruan, Yi-Lun Wu, Tzu-Ling Lin, Hong-Han Shuai |
ICCV | 2 |
| 2025 | Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion GenerationabstractAccurate color alignment in text-to-image (T2I) generation is critical for applications such as fashion, product visualization, and interior design, yet current diffusion models struggle with nuanced and compound color terms (e.g., Tiffany blue, baby pink), often producing images that are misaligned with human intent. Existing approaches rely on cross-attention manipulation, reference images, or fine-tuning but fail to systematically resolve ambiguous color descriptions. To precisely render colors under prompt ambiguity, we propose a training-free framework that enhances color fidelity by leveraging a large language model (LLM) to disambiguate color-related prompts and guiding color blending operations directly in the text embedding space. Our method first employs a large language model (LLM) to resolve ambiguous color terms in the text prompt, and then refines the text embeddings based on the spatial relationships of the resulting color terms in the CIELab color space. Unlike prior methods, our approach improves color accuracy without requiring additional training or external reference images. Experimental results demonstrate that our framework improves color alignment without compromising image quality, bridging the gap between text semantics and visual generation. All supplementary materials are available at https://Sung-Lin.github.io/TintBench/. Sung-Lin Tsai, Bo-Lun Huang, Yu-Ting Shen, Cheng-Yu Yeo, Chiang Tseng, Bo-Kai Ruan, Wen-Sheng Lien, Hong-Han Shuai |
ACM Multimedia | 6 |
| 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User FeedbackabstractDirect preference optimization (DPO) methods have shown strong potential in aligning text-to-image diffusion models with human preferences by training on paired comparisons. These methods improve training stability by avoiding the REINFORCE algorithm but still struggle with challenges such as accurately estimating image probabilities due to the non-linear nature of the sigmoid function and the limited diversity of offline datasets. In this paper, we introduce Diffusion Denoising Ranking Optimization (Diffusion-DRO), a new preference learning framework grounded in inverse reinforcement learning. Diffusion-DRO removes the dependency on a reward model by casting preference learning as a ranking problem, thereby simplifying the training objective into a denoising formulation and overcoming the non-linear estimation issues found in prior methods. Moreover, Diffusion-DRO uniquely integrates offline expert demonstrations with online policy-generated negative samples, enabling it to effectively capture human preferences while addressing the limitations of offline data. Comprehensive experiments show that Diffusion-DRO delivers improved generation quality across a range of challenging and unseen prompts, outperforming state-of-the-art baselines in both both quantitative metrics and user studies. Our source code and pre-trained models are available at https://github.com/basiclab/DiffusionDRO. Yi-Lun Wu, Bo-Kai Ruan, Chiang Tseng, Hong-Han Shuai |
NeurIPS | 2 |
| 2024 | Modeling Uncertainty for Low-Resolution Facial Expression RecognitionabstractRecently, facial expression recognition techniques have made significant progress on high-resolution web images. However, in real-world applications, the obtained images are often with low resolution since they are mostly captured in a wide range of public spaces. As a result, the ambiguity of the expression labels hinders recognition performance due to not only subjective emotion annotations but also ambiguous images. Existing approaches tend to perform poorly when the resolution of face images decreases. In this work, we aim to model the aleatoric uncertainty induced by low-image-resolution and label ambiguity for robust facial expression recognition. We propose probabilistic data uncertainty learning to capture the ambiguity induced by poor image resolution. Additionally, we introduce the emotion wheel to learn the label-uncertainty-aware embedding. Moreover, we exploit the ambiguous nature of neutrality and propose a neutral expression constraint to learn more robust features for facial expression recognition. To the best of our knowledge, this is the first work utilizing the intrinsic nature of neutrality as a regularization to benefit model training. Extensive experimental results show the effectiveness and robustness of our approach. Under low-resolution conditions, our proposed method outperforms the state-of-the-art approaches by 3.02% and 3.16% in terms of accuracy on RAF-DB and FERPlus, respectively. Ling Lo, Bo-Kai Ruan, Hong-Han Shuai, Hao-Wen Cheng |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Temporal Difference-Aware Graph Convolutional Reinforcement Learning for Multi-Intersection Traffic Signal ControlabstractTraffic light control plays a crucial role in intelligent transportation systems. This paper introduces Temporal Difference-Aware Graph Convolutional Reinforcement Learning (TeDA-GCRL), a decentralized RL-based method for efficient multi-intersection traffic signal control. Specifically, we put forward a new graph architecture using each lane as a node for considering intersection relations. Additionally, we propose two new rewards by considering temporal information, namely Temporal-Aware Pressure on Incoming Lanes (TAPIL) and Temporal-Aware Action Consistency (TAAC), which enhance learning efficiency and time-interval sensitivity. Experimental results on five datasets show the superiority of TeDA-GCRL over state-of-the-art methods by at least 9.5% in average travel time. Wei-Yu Lin, Yun-Zhu Song, Bo-Kai Ruan, Hong-Han Shuai, Li-Chun Wang 0001, Yung-Hui Li |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Mimicking the Annotation Process for Recognizing the Micro ExpressionsabstractMicro-expression recognition (MER) has recently become a popular research topic due to its wide applications, e.g., movie rating and recognizing the neurological disorder. By virtue of deep learning techniques, the performance of MER has been significantly improved and reached unprecedented results. This paper proposes a novel architecture to mimic how the expressions are annotated. Specifically, during the annotation process in several datasets, the AU labels are first obtained with FACS, and the expression labels are then decided based on the combinations of the AU labels. Meanwhile, these AU labels describe either the eyes or mouth movements (mutually-exclusive). Following this idea, we design a dual-branch structure with a new augmentation method to separately capture the eyes and mouth features and teach the model what the general expressions should be. Moreover, to adaptively fuse the area features for different expressions, we propose Area Weighted Module to assign different weights to each region. Additionally, we set up an auxiliary task to align the AU similarity scores to help our model capture facial patterns further with AU labels. The proposed approach outperforms other state-of-the-art methods in terms of accuracy on the CASME II and SAMM datasets. Moreover, we provide a new visualization approach to show the relationship between the facial regions and AU features. Bo-Kai Ruan, Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 1 |