VLDB 2026 Research / reviewers in the wild / expert
Cheng Da
dblp:208/4178
· DBLP profile ↗
14ranked-venue papers
8as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Granularity Prediction with Learnable Fusion for Scene Text Recognition
Cheng Da, Cong Yao |
Int. J. Comput. Vis. | 1 |
| 2025 | Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationabstractPreference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the **Latent Reward Model (LRM)**, which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce **Latent Preference Optimization (LPO)**, a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods. Cheng Da, Kun Ding 0001, Huan Yang 0005, Yan Li 0043, Tingting Gao, Di Zhang 0026, Shiming Xiang, Chunhong Pan |
NeurIPS | 2 |
| 2025 | ST-PPO: a spatio-temporal attention enhanced proximal policy optimization algorithm for autonomous driving in complex traffic scenarios
Cheng Da, Yongsheng Qian, Junwei Zeng, Xunting Wei, Futao Zhang |
Mach. Learn. | 1 |
| 2023 | LISTER: Neighbor Decoding for Length-Insensitive Scene Text RecognitionabstractThe diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR) only work well on short or seen-length text, lacking the capability of recognizing longer text or performing length extrapolation. This is a crucial issue, since the lengths of the text to be recognized are usually not given in advance in real-world applications, but it has not been adequately investigated in previous works. Therefore, we propose in this paper a method called Length-Insensitive Scene TExt Recognizer (LISTER), which remedies the limitation regarding the robustness to various text lengths. Specifically, a Neighbor Decoder is proposed to obtain accurate character attention maps with the assistance of a novel neighbor matrix regardless of the text lengths. Besides, a Feature Enhancement Module is devised to model the long-range dependency with low computation cost, which is able to perform iterations with the neighbor decoder to enhance the feature map progressively. To the best of our knowledge, we are the first to achieve effective length-insensitive scene text recognition. Extensive experiments demonstrate that the proposed LISTER algorithm exhibits obvious superiority on long text recognition and the ability for length extrapolation, while comparing favourably with the previous state-of-the-art methods on standard benchmarks for STR (mainly short text)1. Changxu Cheng, Peng Wang 0103, Cheng Da, Qi Zheng 0002, Cong Yao |
ICCV | 3 |
| 2023 | Vision Grid Transformer for Document Layout AnalysisabstractDocument pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pretrained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a two-stream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D4LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet (95.7%→96.2%), DocBank (79.6%→84.1%), and D4LA (67.7%→68.8%). The code and models as well as the D4LA dataset will be made publicly available1. Cheng Da, Chuwei Luo, Qi Zheng 0002, Cong Yao |
ICCV | 1 |
| 2022 | Levenshtein OCR
Cheng Da, Peng Wang 0103, Cong Yao |
ECCV (28) | 1 |
| 2022 | Multi-granularity Prediction for Scene Text Recognition
Peng Wang 0103, Cheng Da, Cong Yao |
ECCV (28) | 2 |
| 2021 | Fashion Focus: Multi-modal Retrieval System for Video Commodity Localization in E-commerceabstractNowadays, live-stream and short video shopping in E-commerce have grown exponentially. However, the sellers are required to manually match images of the selling products to the timestamp of exhibition in the untrimmed video, resulting in a complicated process. To solve the problem, we present an innovative demonstration of multi-modal retrieval system called ``Fashion Focus'', which enables to exactly localize the product images in the online video as the focuses. Different modality contributes to the community localization, including visual content, linguistic features and interaction context are jointly investigated via presented multi-modal learning. Our system employs two procedures for analysis, including video content structuring and multi-modal retrieval, to automatically achieve accurate video-to-shop matching. Fashion Focus presents a unified framework that can orientate the consumers towards relevant product exhibitions during watching videos and help the sellers to effectively deliver the products over search and recommendation. Yanhao Zhang 0002, Qiang Wang 0054, Cheng Da, Siyang Sun |
AAAI | 5 |
| 2021 | AsyNCE: Disentangling False-Positives for Weakly-Supervised Video GroundingabstractWeakly-supervised video grounding has been investigated to ground textual phases in video content with only video-sentence pairs provided during training, for the lack of prohibitively costly bounding box annotations. Existing methods cast this task into a frame-level multiple instance learning (MIL) problem with the ranking loss. While an object might appear sparsely across multiple frames, causing uncertain false-positive frames. Thus, directly computing the average loss of all frames is inadequate in video domain. Moreover, the positive and negative pairs are equally coupling in ranking loss, so that it is impossible to handle false-positive frames individually. Additionally, naive inner production is suboptimal for the similarity measure of cross domains. To solve these issues, we propose a novel AsyNCE loss to flexibly disentangle the positive pairs from negative ones in frame-level MIL, which allows for mitigating the uncertainty of false-positive frames effectively. Besides, a cross-modal transformer block is introduced to purify the text feature by image frame context, generating a visual-guided text feature for better similarity measure. Extensive experiments on YouCook2, RoboWatch and WAB datasets demonstrate the superiority and robustness of our method over state-of-the-art methods. Cheng Da, Yanhao Zhang 0002, Chunhong Pan |
ACM Multimedia | 1 |
| 2019 | No-Reference Image Quality Assessment with Reinforcement Recursive List-Wise RankingabstractOpinion-unaware no-reference image quality assessment (NR-IQA) methods have received many interests recently because they do not require images with subjective scores for training. Unfortunately, it is a challenging task, and thus far no opinion-unaware methods have shown consistently better performance than the opinion-aware ones. In this paper, we propose an effective opinion-unaware NR-IQA method based on reinforcement recursive list-wise ranking. We formulate the NR-IQA as a recursive list-wise ranking problem which aims to optimize the whole quality ordering directly. During training, the recursive ranking process can be modeled as a Markov decision process (MDP). The ranking list of images can be constructed by taking a sequence of actions, and each of them refers to selecting an image for a specific position of the ranking list. Reinforcement learning is adopted to train the model parameters, in which no ground-truth quality scores or ranking lists are necessary for learning. Experimental results demonstrate the superior performance of our approach compared with existing opinion-unaware NR-IQA methods. Furthermore, our approach can compete with the most effective opinion-aware methods. It improves the state-of-the-art by over 2% on the CSIQ benchmark and outperforms most compared opinion-aware models on TID2013. Jie Gu 0002, Gaofeng Meng, Cheng Da, Shiming Xiang, Chunhong Pan |
AAAI | 3 |
| 2019 | Nonlinear Asymmetric Multi-Valued HashingabstractMost existing hashing methods resort to binary codes for large scale similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose Nonlinear Asymmetric Multi-Valued Hashing (NAMVH) supported by two distinct non-binary embeddings. Specifically, a real-valued embedding is used for representing the newly-coming query by an ideally nonlinear transformation. Besides, a multi-integer-embedding is employed for compressing the whole database, which is modeled by Binary Sparse Representation (BSR) with fixed sparsity. With these two non-binary embeddings, NAMVH preserves more precise similarities between data points and enables access to the incremental extension with database samples evolving dynamically. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the pairwise label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by a well-designed alternative optimization method. Extensive experiments on seven large scale datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency. Cheng Da, Gaofeng Meng, Shiming Xiang, Kun Ding 0001, Shibiao Xu, Qing Yang 0002, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Weakly Semantic Guided Action RecognitionabstractAction recognition plays a fundamental role in computer vision and video analysis. Nevertheless, extracting effective spatial-temporal features remains a challenging task. This paper proposes three simple but effective weakly semantic guided modules (SGMs) for both environment-constrained and cross-domain action recognition. The SGMs are composed of total 3-D convolution and element-wise gated operations; thus, they are efficient and easy to implement. The semantic guidance is obtained in a weakly supervised manner, in which each video clip is labeled with only an action class instead of pixel-level semantics. Benefitting from the semantic guidance, the network [called semantic guided network (SGN)] can focus on the salient parts of the video clips. Consequently, the redundant information can be reduced and the model is more robust to noise. Besides, benefitting from the intrinsic property of SGMs, SGN is totally end-to-end trainable. Quantities of experiments on both environment-constrained (e.g., Penn, HMDB-51, and UCF101) and cross-domain (e.g., ODAR) action recognition datasets demonstrate its effectiveness. Specifically, SGN gets improvements of 3.7%, 2.1%, and 5.2% for Penn, HMDB-51, and UCF-101 than the baseline ResNet3D, respectively, and SGN ranked third place in the ODAR 2017 challenge. Tingzhao Yu, Lingfeng Wang 0002, Cheng Da, Huxiang Gu, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 3 |
| 2017 | AMVH: Asymmetric Multi-Valued hashingabstractMost existing hashing methods resort to binary codes for similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose an asymmetric multi-valued hashing method supported by two different non-binary embeddings. (1) A real-valued embedding is used for representing the newly-coming query. (2) A multi-integer-embedding is employed for compressing the whole database, which is modeled by binary sparse representation with fixed sparsity. With these two non-binary embeddings, the similarities between data points can be preserved precisely. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by alternative optimization. Extensive experiments on three multilabel datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency. Cheng Da, Shibiao Xu, Kun Ding 0001, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
CVPR | 1 |
| 2017 | Efficient similarity learning for asymmetric hashingabstractHashing techniques with asymmetric schemes (e.g., only bi-narizing the database points) have recently attracted wide attention in the circle of image retrieval. In comparison with those methods which binarize simultaneously both of the query and database points, they not only enjoy the storage and search efficiencies, but also provide higher accuracy. Gearing to this line, this paper proposes a metric-embedded asymmetric hashing (MEAH) that learns jointly a bilinear similarity measure and binary codes of database points in an unsupervised manner. Technically, the learned similarity measure is able to bridge the gap between the binary codes and the real-valued codes, which are represented possibly with different dimensions. What is more, this measure is capable of preserving the global structure hidden in the database. Extensive experiments on two public image benchmarks demonstrate the superiority of our approach over the several state-of-the-art unsupervised hashing methods. Cheng Da, Yang Yang 0062, Chunlei Huo, Shiming Xiang, Chunhong Pan |
ICIP | 1 |