VLDB 2026 Research / reviewers in the wild / expert
Yanhua Cheng
dblp:42/8495
· DBLP profile ↗
21ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fairness-Aware Design for Contextual Experiments: Guaranteeing Reliability and Equity in Heterogeneous SubgroupsabstractExperimental design is critical for evidence-based decision-making in healthcare, marketing, and public policy. However, designing efficient experiments across heterogeneous subgroups presents significant challenges. Existing methods often optimize for statistical power or overall sample efficiency, overlooking crucial fairness considerations across these different subgroups. To address this gap, we introduce a Fairness-Aware Contextual Track-and-Stop Design (F-CTSD) algorithm. The proposed F-CTSD algorithm provides statistical guarantees on subgroup fairness while minimizing required sample sizes. We quantify the fairness-efficiency trade-off and derive the sample complexity bound for the proposed F-CTSD algorithm under its fairness constraints. We further theoretically prove that the proposed F-CTSD algorithm consistently produces accurate treatment effect estimates even under fairness requirements, enhancing statistical reliability. Numerical experiments show that the proposed F-CTSD algorithm outperforms existing methods, achieving higher sample efficiency while reducing subgroup fairness violations by 4.95%. Guangyan Gan, Yanhua Cheng, Yongxiang Tang 0001, Xialong Liu, Peng Jiang 0002 |
AAAI | 3 |
| 2026 | OPS: An Order-Preserving Sorting Network for Information RetrievalabstractLearning-to-rank (LTR) is a fundamental component of modern large-scale information retrieval (IR) systems, playing an essential role across various stages of the ranking pipeline. Recently, differentiable sorting networks have attracted increasing attention for LTR as a permutation-level learning paradigm, enabling end-to-end optimization directly on ranking structure. However, existing approaches suffer from two critical limitations: (i) permutation-matrix fidelity, i.e., the predicted soft permutation matrix may deviate from the exact hard permutation matrix required by permutation-level objectives; and (ii) uncertainty in target ordering arising from coarse or tied relevance labels, where the ground-truth order is set-valued rather than unique. Yongxiang Tang 0001, Guikai Luan, Yanhua Cheng, Xialong Liu, Peng Jiang 0002 |
SIGIR | 5 |
| 2025 | Learning Monotonic Probabilities with a Generative Cost ModelabstractIn many machine learning tasks, it is often necessary for the relationship between input and output variables to be monotonic, including both strictly monotonic and implicitly monotonic relationships. Traditional methods for maintaining monotonicity mainly rely on construction or regularization techniques, whereas this paper shows that the issue of strict monotonic probability can be viewed as a partial order between an observable revenue variable and a latent cost variable. This perspective enables us to reformulate the monotonicity challenge into modeling the latent cost variable. To tackle this, we introduce a generative network for the latent cost variable, termed the Generative Cost Model (GCM), which inherently addresses the strict monotonic problem, and propose the Implicit Generative Cost Model (IGCM) to address the implicit monotonic problem. We further validate our approach with a numerical simulation of quantile regression and conduct multiple experiments on public datasets, showing that our method significantly outperforms existing monotonic modeling techniques. The code for our experiments can be found at https://github.com/tyxaaron/GCM. Yongxiang Tang 0001, Yanhua Cheng, Xiaocheng Liu, Jiaochen Chen, Yanxiang Zeng, Ning Luo 0004, Pengjia Yuan, Xialong Liu, Peng Jiang 0002 |
ICML | 2 |
| 2025 | S-Diff: An Anisotropic Diffusion Model for Collaborative Filtering in Spectral DomainabstractRecovering potential user preferences from user-item interaction matrices is a key challenge in recommender systems. While diffusion models can sample and reconstruct preferences from latent distributions, they often fail to capture similar users' collective preferences effectively. Additionally, latent variables degrade into pure Gaussian noise during the forward process, lowering the signal-to-noise ratio, which in turn degrades performance. To address this, we propose S-Diff, inspired by graph-based collaborative filtering, better to utilize low-frequency components in the graph spectral domain. S-Diff maps user interaction vectors into the spectral domain and parameterizes diffusion noise to align with graph frequency. As a result, this anisotropic diffusion retains significant low-frequency components, preserving a high signal-to-noise ratio. S-Diff further employs a conditional denoising network to encode user interactions, recovering true preferences from noisy data. This method achieves promising results across multiple datasets. Yanhua Cheng, Yongxiang Tang 0001, Xiaocheng Liu, Xialong Liu, Lisong Wang, Peng Jiang 0002 |
WSDM | 2 |
| 2024 | Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningabstractIn recent years, text-to-video retrieval methods based on CLIP have experienced rapid development. The primary direction of evolution is to exploit the much wider gamut of visual and textual cues to achieve alignment. Concretely, those methods with impressive performance often design a heavy fusion block for sentence (words)-video (frames) interaction, regardless of the prohibitive computation complexity. Nevertheless, these approaches are not optimal in terms of feature utilization and retrieval efficiency. To address this issue, we adopt multi-granularity visual feature learning, ensuring the model's comprehensiveness in capturing visual content features spanning from abstract to detailed levels during the training phase. To better leverage the multi-granularity features, we devise a two-stage retrieval architecture in the retrieval phase. This solution ingeniously balances the coarse and fine granularity of retrieval content. Moreover, it also strikes a harmonious equilibrium between retrieval effectiveness and efficiency. Specifically, in training phase, we design a parameter-free text-gated interaction block (TIB) for fine-grained video representation learning and embed an extra Pearson Constraint to optimize cross-modal representation learning. In retrieval phase, we use coarse-grained video representations for fast recall of top-k candidates, which are then reranked by fine-grained video representations. Extensive experiments on four benchmarks demonstrate the efficiency and effectiveness. Notably, our method achieves comparable performance with the current state-of-the-art methods while being nearly 50 times faster. Kaibin Tian, Yanhua Cheng, Xinglin Hou, Quan Chen 0006, Han Li 0005 |
AAAI | 2 |
| 2024 | Spatiotemporal Fine-grained Video Description for Short VideosabstractIn the mobile internet era, short videos are inundating people's lives. However, research on visual language models specifically designed for short videos has not yet received sufficient attention. Short videos are not just videos of limited duration. The prominent visual details and high information density of short videos differentiate them to long videos. In this paper, we propose the SpatioTemporal Fine-grained Description (STFVD) emphasizing on the uniqueness of short videos, which entails capturing the intricate details of the main subject and fine-grained movements. To this end, we create a comprehensive Short Video Advertisements Description (SVAD) dataset, comprising 34,930 clips from 5,046 videos. The dataset covers a range of topics, including 191 sub-industries, 649 popular products, and 470 trending games. Various efforts have been made in the data annotation process to ensure the inclusion of fine-grained spatiotemporal information, resulting in 34,930 high-quality annotations. Compared to existing datasets, samples in SVAD exhibit a superior text information density, suggesting that SVAD is more appropriate for the analysis of short videos. Based on the SVAD dataset, we develop a visual language model (SVAD-VLM) to generate spatiotemporal fine-grained description for short videos. We use a prompt-guided keyword generation task to efficiently learn key visual information. Moreover, we also utilize dual visual alignment to exploit the advantage of mixed-datasets training. Experiments on SVAD dataset demonstrate the challenge of STFVD and the competitive performance of proposed method compared to previous ones. Te Yang, Jian Jia, Bo Wang 0071, Yanhua Cheng, Yan Li 0043, Dongze Hao, Xipeng Cao, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xiangyu Zhu 0001, Zhen Lei 0001 |
ACM Multimedia | 4 |
| 2023 | Cross-Domain Product Representation Learning for Rich-Content E-CommerceabstractThe proliferation of short video and live-streaming platforms has revolutionized how consumers engage in online shopping. Instead of browsing product pages, consumers are now turning to rich-content e-commerce, where they can purchase products through dynamic and interactive media like short videos and live streams. This emerging form of online shopping has introduced technical challenges, as products may be presented differently across various media domains. Therefore, a unified product representation is essential for achieving cross-domain product recognition to ensure an optimal user search experience and effective product recommendations. Despite the urgent industrial need for a unified cross-domain product representation, previous studies have predominantly focused only on product pages without taking into account short videos and live streams. To fill the gap in the rich-content e-commerce area, in this paper, we introduce a large-scale cRoss-dOmain Product rEcognition dataset, called ROPE. ROPE covers a wide range of product categories and contains over 180,000 products, corresponding to millions of short videos and live streams. It is the first dataset to cover product pages, short videos, and live streams simultaneously, providing the basis for establishing a unified product representation across different media domains. Furthermore, we propose a Cross-dOmain Product rEpresentation framework, namely COPE, which unifies product representations in different domains through multimodal learning including text and vision. Extensive experiments on downstream tasks demonstrate the effectiveness of COPE in learning a joint feature space for all product domains. Xuehan Bai, Yan Li 0043, Yanhua Cheng, Quan Chen 0006, Han Li 0005 |
ICCV | 3 |
| 2023 | Cross-view Semantic Alignment for Livestreaming Product RecognitionabstractLive commerce is the act of selling products online through live streaming. The customer’s diverse demands for online products introduce more challenges to Livestreaming Product Recognition. Previous works have primarily focused on fashion clothing data or utilize single-modal input, which does not reflect the real-world scenario where multimodal data from various categories are present. In this paper, we present LPR4M, a large-scale multimodal dataset that covers 34 categories, comprises 3 modalities (image, video, and text), and is 50× larger than the largest publicly available dataset. LPR4M contains diverse videos and noise modality pairs while exhibiting a long-tailed distribution, resembling real-world problems. Moreover, a cRoss-vIew semantiC alignmEnt (RICE) model is proposed to learn discriminative instance features from the image and video views of the products. This is achieved through instance-level contrastive learning and cross-view patch-level feature propagation. A novel Patch Feature Reconstruction loss is proposed to penalize the semantic misalignment between cross-view patches. Extensive experiments demonstrate the effectiveness of RICE and provide insights into the importance of dataset diversity and expressivity. The dataset and code are available at https://github.com/adxcreative/RICE. Yan Li 0043, Yanhua Cheng, Quan Chen 0006, Han Li 0005 |
ICCV | 4 |
| 2019 | MAPNet: Multi-modal attentive pooling network for RGB-D indoor scene classification
Yabei Li, Zhang Zhang 0001, Yanhua Cheng, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 3 |
| 2019 | Corrigendum to "MAPNet: Multi-modal attentive pooling network for RGB-D indoor scene classification" [Pattern Recognition 90 (2019) 436-449]
Yabei Li, Zhang Zhang 0001, Yanhua Cheng, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 3 |
| 2018 | DF2Net: Discriminative Feature Learning and Fusion Network for RGB-D Indoor Scene ClassificationabstractThis paper focuses on the task of RGB-D indoor scene classification. It is a very challenging task due to two folds. 1) Learning robust representation for indoor scene is difficult because of various objects and layouts. 2) Fusing the complementary cues in RGB and Depth is nontrivial since there are large semantic gaps between the two modalities. Most existing works learn representation for classification by training a deep network with softmax loss and fuse the two modalities by simply concatenating the features of them. However, these pipelines do not explicitly consider intra-class and inter-class similarity as well as inter-modal intrinsic relationships. To address these problems, this paper proposes a Discriminative Feature Learning and Fusion Network (DF2Net) with two-stage training. In the first stage, to better represent scene in each modality, a deep multi-task network is constructed to simultaneously minimize the structured loss and the softmax loss. In the second stage, we design a novel discriminative fusion network which is able to learn correlative features of multiple modalities and distinctive features of each modality. Extensive analysis and experiments on SUN RGB-D Dataset and NYU Depth Dataset V2 show the superiority of DF2Net over other state-of-the-art methods in RGB-D indoor scene classification task. Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan |
AAAI | 3 |
| 2017 | Locality-Sensitive Deconvolution Networks with Gated Fusion for RGB-D Indoor Semantic SegmentationabstractThis paper focuses on indoor semantic segmentation using RGB-D data. Although the commonly used deconvolution networks (DeconvNet) have achieved impressive results on this task, we find there is still room for improvements in two aspects. One is about the boundary segmentation. DeconvNet aggregates large context to predict the label of each pixel, inherently limiting the segmentation precision of object boundaries. The other is about RGB-D fusion. Recent state-of-the-art methods generally fuse RGB and depth networks with equal-weight score fusion, regardless of the varying contributions of the two modalities on delineating different categories in different scenes. To address the two problems, we first propose a locality-sensitive DeconvNet (LS-DeconvNet) to refine the boundary segmentation over each modality. LS-DeconvNet incorporates locally visual and geometric cues from the raw RGB-D data into each DeconvNet, which is able to learn to upsample the coarse convolutional maps with large context whilst recovering sharp object boundaries. Towards RGB-D fusion, we introduce a gated fusion layer to effectively combine the two LS-DeconvNets. This layer can learn to adjust the contributions of RGB and depth over each pixel for high-performance object recognition. Experiments on the large-scale SUN RGB-D dataset and the popular NYU-Depth v2 dataset show that our approach achieves new state-of-the-art results for RGB-D indoor semantic segmentation. Yanhua Cheng, Rui Cai 0002, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang |
CVPR | 1 |
| 2017 | Semantics-guided multi-level RGB-D feature fusion for indoor semantic segmentationabstractIndoor RGB-D semantic segmentation is a new and challenging problem. Traditional methods usually apply two-stream convolutional neural networks (CNNs) to represent RGB and depth images respectively, and fuse the two streams on a specific layer. In this paper, we explore several fusion strategies based on this two-stream-CNN framework and point out such a single-layer fusion method cannot exploit the complementary RGB and depth cues well for semantic segmentation. To address this problem, we propose a novel Semantics-guided Multi-level feature fusion approach, which first learns deep feature representation from bottom to up, and then gradually fuses the RGB and depth features from high level to low level under the guidance of the semantic cues. Experimental results on SUN RGB-D dataset demonstrate the advantages of the proposed method over the state of the arts. Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2016 | Semi-Supervised Multimodal Deep Learning for RGB-D Object Recognition
Yanhua Cheng, Xin Zhao 0012, Rui Cai 0002, Zhiwei Li 0006, Kaiqi Huang, Yong Rui |
IJCAI | 1 |
| 2015 | Convolutional Fisher Kernels for RGB-D Object RecognitionabstractThis paper studies the problem of improving object recognition using the novel RGB-D data. To address the problem, a new convolutional Fisher Kernels (CFK) method is proposed to represent RGB-D objects powerfully yet efficiently. The core idea of our approach is to integrate the both advantages of the convolutional neural networks (CNN) and Fisher Kernel encoding (FK): CNN model is flexible to adapt to new data sources, but requires for large amounts of training data with significant computational resources for good generalization, In comparison, FK encoding is able to represent objects powerfully and efficiently with small training data, however, its success highly depends on the well-designed SIFT features in literature, which may not be suitable for the new depth data. CFK can be interpreted as a two-layer feature learning structure to bridge the two models. The first layer employs a single-layer CNN to learn low-level translation ally invariant features for both RGB and depth data efficiently. The second layer aggregates the convolutional responses by FK encoding. Here 2D and 3D spatial pyramids are applied to further improve the Fisher vector representation of each modality. Experiments on RGB-D object recognition benchmarks demonstrate that our approach can achieve the state-of-the-art results. Yanhua Cheng, Rui Cai 0002, Xin Zhao 0012, Kaiqi Huang |
3DV | 1 |
| 2015 | Query Adaptive Similarity Measure for RGB-D Object RecognitionabstractThis paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object pose and scale changes, as well as intra-class variations, and (2) effectively fusing RGB and depth cues is still an open problem. To address these problems, this paper first proposes a new similarity measure based on dense matching, through which objects in comparison are warped and aligned, to better tolerate variations. Towards RGB and depth fusion, we argue that a constant and golden weight doesn't exist. The two modalities have varying contributions when comparing objects from different categories. To capture such a dynamic characteristic, a group of matchers equipped with various fusion weights is constructed, to explore the responses of dense matching under different fusion configurations. All the response scores are finally merged following a learning-to-combination way, which provides quite good generalization ability in practice. The proposed approach win the best results on several public benchmarks, e.g., achieves 92.7% top-1 test accuracy on the Washington RGB-D object dataset, with a 5.1% improvement over the state-of-the-art. Yanhua Cheng, Rui Cai 0002, Chi Zhang 0069, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang, Yong Rui |
ICCV | 1 |
| 2015 | MeshStereo: A Global Stereo Model with Mesh Alignment Regularization for View InterpolationabstractWe present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D triangles with shared vertices. Lifting the 2D triangulation to 3D naturally generates a corresponding mesh. A technical difficulty is to properly split vertices to multiple copies when they appear at depth discontinuous boundaries. To deal with this problem, we formulate our objective as a two-layer MRF, with the upper layer modeling the splitting properties of the vertices and the lower layer optimizing a region-based stereo matching. Experiments on the Middlebury and the Herodion datasets demonstrate that our model is able to synthesize visually coherent new view angles with high PSNR, as well as outputting high quality disparity maps which rank at the first place on the new challenging high resolution Middlebury 3.0 benchmark. Chi Zhang 0069, Zhiwei Li 0006, Yanhua Cheng, Rui Cai 0002, Hongyang Chao, Yong Rui |
ICCV | 3 |
| 2015 | Semi-supervised learning and feature evaluation for RGB-D object recognition
Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
Comput. Vis. Image Underst. | 1 |
| 2014 | Semi-supervised Learning for RGB-D Object RecognitionabstractConventional supervised object recognition methods have been investigated for many years. Despite their successes, there are still two suffering limitations: (1) various information of an object is represented by artificial features only derived from RGB images, (2) lots of manually labeled data is required by supervised learning. To address those limitations, we propose a new semi-supervised learning framework based on RGB and depth (RGB-D) images to improve object recognition. In particular, our framework has two modules: (1) RGB and depth images are represented by convolutional-recursive neural networks to construct high level features, respectively, (2) co-training is exploited to make full use of unlabeled RGB-D instances due to the existing two independent views. Experiments on the standard RGB-D object dataset demonstrate that our method can compete against with other state-of-the-art methods with only 20% labeled data. Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
ICPR | 1 |
| 2005 | A new secure M+1st price auction schemeabstractSecurity and privacy are the crucial conditions in the seal-auction design. For anonymous bid, many electronic auctions are presented recently. But up to now, the study of secure M+1/sup st/ price auction is still very weak. In this paper, based on signcryption scheme, Chinese Remainder Theory (CRT) and broadcast technology, an efficient and secure M+1/sup st/ price auction scheme is presented. Discussion about security and communication cost of our proposed scheme was given at the end of this paper. Xueying Ma, Chengqing Ye, Yanhua Cheng |
SMC | 4 |
| 2005 | Time-optimal controller for AQM router supporting TCP flows with small RTTabstractActive queue management (AQM) is an effective method to enhance congestion control, and to achieve tradeoff between link utilization and delay. Recently it has shown that the active queue management schemes implemented in the routers of communication networks supporting transmission control protocol (TCP) flows can be modeled as a feedback control system. In this paper, based on Lyapunov function we developed an optimal controller for small RTT TCP flows to improve active queue management router's stability and response time, which are often in conflict with each other in system performance. Using ns simulations, it is shown that optimal controller outperform PI controller significantly. Chengqing Ye, Yanhua Cheng |
SMC | 4 |