VLDB 2026 Research / reviewers in the wild / expert
Yiyang Hu
dblp:93/10785
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning-Grounded Intent Injection for Generative RecommendationabstractIndustrial generative recommendation systems operating over discrete Semantic IDs (SIDs) are largely behavior-driven, and thus struggle to proactively activate latent demand before explicit user signals emerge, leading to intent cold-start. To address this, we propose RIGER (Reasoning-grounded Intent injection for GE nerative Recommendation), a deployable two-stage framework that integrates offline large language model (LLM) reasoning into an online generative recommender under strict latency constraints. Offline, to ensure scalable deployment, we distill the latent-intent inference capability of a strong LLM into a lightweight forecasting model using an automated data curation pipeline---leveraging judge-guided prompt calibration and future-query-guided rejection filtering. Online, to bridge the representation mismatch between free-form textual intents and the discrete SID token space, predicted intents are converted into SID-native tokens through a behavior-grounded mapping and injected into the deployed decoder-only retrieval backbone. We further fine-tune the model with beam-aware GRPO, introducing a hierarchical intent-alignment exploration reward in SID space while preserving exploitation behavior through KL regularization. Offline evaluations demonstrate a substantial increase in intent-aligned density and diversity with only a marginal reduction in hindsight recall, indicating that RIGER effectively enhances proactive intent exploration while preserving its capability to exploit historical behaviors. In a large-scale e-commerce display advertising system, RIGER improves clicks by 1.6% and advertiser spend by 1.3%. Xusong Chen, Peini Guo, Yiyang Hu, Mengqin Que, Zhiwei Fang, Changping Peng, Ching Law |
SIGIR | 5 |
| 2025 | NCL-CIR: Noise-aware Contrastive Learning for Composed Image RetrievalabstractComposed Image Retrieval (CIR) seeks to find a target image using a multi-modal query, which combines an image with modification text to pinpoint the target. While recent CIR methods have shown promise, they mainly focus on exploring relationships between the query pairs (image and text) through data augmentation or model design. These methods often assume perfect alignment between queries and target images, an idealized scenario rarely encountered in practice. In reality, pairs are often partially or completely mismatched due to issues like inaccurate modification texts, low-quality target images, and annotation errors. Ignoring these mismatches leads to numerous False Positive Pair (FFPs) denoted as noise pairs in the dataset, causing the model to overfit and ultimately reducing its performance. To address this problem, we propose the Noise-aware Contrastive Learning for CIR (NCL-CIR), comprising two key components: the Weight Compensation Block (WCB) and the Noise-pair Filter Block (NFB). The WCB coupled with diverse weight maps can ensure more stable token representations of multi-modal queries and target images. Meanwhile, the NFB, in conjunction with the Gaussian Mixture Model (GMM) predicts noise pairs by evaluating loss distributions, and generates soft labels correspondingly, allowing for the design of the soft-label based Noise Contrastive Estimation (NCE) loss function. Consequently, the overall architecture helps to mitigate the influence of mismatched and partially matched samples, with experimental results demonstrating that NCL-CIR achieves exceptional performance on the benchmark datasets. Yujian Lee, Zailong Chen, Yiyang Hu, Guquan Jing |
ICASSP | 6 |
| 2025 | Face Relighting with Ratio Function for Explicit Geometric RepresentationabstractThis paper addresses the problem of face relighting under varying illumination conditions. Lighting is a fundamental element in portrait photography that shapes the mood, geometry, and overall realism of the captured characters. Most previous studies have mainly treated relighting as a 2D generation task without incorporating the geometric features of the characters. In contrast, inspired by ratio image-based methods, this paper proposes to disentangle shadow and brightness variations through geometric information and utilizes generative adversarial networks (GANs) to obtain relighted images with brightness consistency. We design a novel relighting-ratio function that integrates the Cook-Torrance reflectance model to more explicitly represent the face geometry than previous ratio image-based methods. This relighting-ratio function is derived from an image rendering formula that quantizes variables such as albedo that are affected by the lighting direction, while systematically excluding variables such as normal and viewpoint that are not affected by lighting. We conduct quantitative and qualitative experiments on the Multi-PIE and CelebA-HQ datasets and show that the proposed method outperforms existing SOTA methods using lighting directions. Yiyang Hu, Zequn Zhang, Hui Zhang 0062, Guquan Jing |
ICASSP | 1 |
| 2025 | ESTI: An Efficient Spatial-Temporal Interaction Network For Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to identify the target pedestrian from video sequences. However, redundant information exist in input frames. Extracting spatial-temporal features in whole adjacent frames can introduce additional computational overhead. Furthermore, this process leads to the loss of critical spatial and temporal details, causing suboptimal representations. To mitigate these issues, we propose an Efficient Spatial-Temporal Interaction (ESTI) network, which processes half of the input sequence separately through spatial and temporal branches, extracting high-level discriminative features across multiple layers and avoiding redundancy computations. In particular, we propose a Feature Enhancement Module (FEM) for the spatial branch to focus on enhancing spatial dependencies adaptively, and a Temporal Interaction Module (TIM) for temporal branch to capture temporal correlations effectively. Spatial-temporal interaction is performed at the final layer to generate distinctive representations. Extensive experiments on three challenging video Re-ID datasets show that our ESTI achieves competitive results while maintaining low computational complexity. Guquan Jing, Yiyang Hu, Yujian Lee, Hui Zhang 0062 |
ICME | 3 |
| 2025 | Boosting Audio-Visual Segmentation via Triple-Modalities AlignmentabstractThe Audio-Visual Segmentation (AVS) task aims to identify sound-producing objects in the visual domain using auditory cues. Enhancing segmentation efficiency by incorporating prior knowledge, such as object locations and textual prompts, has proven to be crucial. However, existing methods suffer from feature misalignment during model training, leading to ineffective integration and reduced performance. To address this, we propose Triple-modalities alignment (TM-align), which combines audio signals, visual images, and textual prompts. By leveraging prompts from a frozen multi-modal large language model (MLLM), we extract two types of semantic information: contextual semantic description (C.S.D) and prompt specific summary (P.S.S). TM-align yields three pairs of aligned features: visual and C.S.D, visual and P.S.S, visual and audio, within two of our proposed cross-modalities alignment (CMA) models. To further enhance the alignment, we employ Jensen-Shannon Divergence (JSD) to regulate the domain distribution of the latter two features. By effectively aligning the three modalities, TM-align reduces redundancy and improves the overall AVS performance. Experimental results demonstrate that TM-align outperforms the mainstream AVS models.1 Yujian Lee, Zailong Chen, Wentao Fan 0001, Guquan Jing, Yiyang Hu |
ICME | 6 |
| 2025 | Contextual Reasoning for Robust Composed Image Retrieval with Vision-Language ModelsabstractComposed Image Retrieval (CIR) combines a reference image with modification text for precise and flexible searches. However, existing methods face two key challenges: first, the limited information in modification text hampers the model's ability to understand user intent, leading to reduced accuracy and diversity; second, reliance on unidirectional constraints overlooks the complementary role of reference and target captions. In this paper, we propose CR-CIR a novel framework that leverages Contextual Reasoning and vision-language models to enhance CIR. Specifically, we use a VLM (e.g., BLIP2) to address the scarcity of textual annotations in existing datasets by generating descriptive captions for both reference and target images. In addition, we enhance the modification text with contextual information using a VLM (e.g., MiniCPM), enriching the model's understanding of user intent. Then our method incorporates a Dual Reasoning Modification Module, which imposes bidirectional constraints by integrating both image and text modalities. Additionally, we introduce a Modality Shift Regularization Loss that assumes symmetry and correlation between text and image domain transformations in the latent space. This new loss function enforces consistent modality shifts, significantly enhancing the model's interpretative and generalization abilities. Experimental results on benchmark CIR datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. Our code and dataset will be available at https://github.com/kola1124/CR-CIR.git. Yujian Lee, Xubo Liu 0001, Hui Zhang 0062, Zailong Chen, Yiyang Hu, Guquan Jing, Yunting Lai |
ICMR | 6 |
| 2025 | Text-Guided Realistic Single Image Relighting with Wavelet Mamba Diffusion Network
Yunting Lai, Hui Zhang 0062, Yiyang Hu, Guquan Jing |
ICMR | 4 |
| 2025 | 3D-Aided Pedestrian Representation Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to match the target pedestrian from video sequences. Recent methods perform frame-level feature extraction followed by temporal aggregation to obtain video representations. However, they pay insufficient attention to the quality of frame-level features, which suffer from issues including multi-frame misalignment, partial occlusion and appearance confusion. People live in a 3D space. 3D pedestrian representations can provide rich geometric information and shape cues that offer promising solutions to these challenges in video-based Re-ID. To mitigate these issues, this paper proposes a 3D-Aid Pedestrian Representation Learning (3DAPRL) network, which introduces 3D modality to video-based Re-ID. Specifically, two novel modules are designed,i.e., the Cross-Modal Fusion (CMF) module and the Shape-aware Spatial-Temporal Interaction (SSTI) module, to enhance pedestrian representation learning. The CMF module generates discriminative fusion representations by utilizing 3D pedestrian data, while the SSTI module learns spatial-temporal 3D shape representation which are distinguishable for finding the target pedestrian in video scenarios. Both features generated from the CMF and SSTI modules contribute to the final video representation. Extensive experiments on four challenging video-based Re-ID datasets demonstrate that our 3DAPRL network reaches better performance than state-of-the-arts methods. Guquan Jing, Yujian Lee, Yiyang Hu, Hui Zhang 0062 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | V2ICooper: Toward Vehicle-to-Infrastructure Cooperative Perception with Spatiotemporal Asynchronous Fusion
Hao Zhang 0065, Feiyu Jin, Yiyang Hu, Rongzhen Li, Kai Liu 0001 |
WASA (3) | 4 |
| 2024 | PM2.5 prediction based on dynamic spatiotemporal graph neural network
Haibin Liao, Mou Wu, Yiyang Hu, Haowei Gong |
Appl. Intell. | 4 |
| 2023 | DEdgeNet: Extrinsic Calibration of Camera and LiDAR with Depth-discontinuous EdgesabstractThis paper addresses the problem of calibrating extrinsic parameter matrix between an RGB camera and a LiDAR. Multimodal sensing systems are essential for fully autonomous navigation platforms. A key pre-requisite for such a system is calibration between different sensors. As the two most widely equipped sensors, calibration between RGB cameras and LiDARs remains challenging. Existing methods address this problem without using explicit geometric priors. In this paper, we propose a novel real-time network that utilizes depth-discontinuous edges extracted from a single image to calibrate cameras and LiDARs. Our network consists of two key components: (1) a self-supervised edge extraction network named DEdgeNet, which detects depth-discontinuous edges from a single image and extracts corresponding features; (2) prediction of the extrinsic parameter matrix between the camera and the LiDAR by matching fixed features in RGB images and updating depth features in a coarse-to-fine frame. Specifically, considering that edges are rich and common in natural scenes, DEdgeNet simplifies RGB image encoding and extracts fixed edges for feature matching. We conducted extensive experiments on the KITTI-odometry dataset. The results show that our method achieves an average rotation error of 0.028° and an average translation error of 0.247 cm, which demonstrates the superiority of our method. Yiyang Hu, Leiping Jie, Hui Zhang 0062 |
ICRA | 1 |