VLDB 2026 Research / reviewers in the wild / expert
Wentian Zhao
dblp:189/4664
· DBLP profile ↗
23ranked-venue papers
6as first author
17since 2021 · last 2025
0009-0006-7645-8263ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cobweb: Enhanced Generation Diversity for Black-box Fairness TestingabstractBlack-box fairness testing aims to reveal potential discriminatory behaviors in deployed AI models by generating individual discriminatory instances, thereby safeguarding trustworthiness in socially critical domains such as smart education and talent recruitment. However, existing approaches often suffer from limited instance space coverage due to their reliance on local neighborhood searches around predefined seed instances, leading to poor diversity and suboptimal exploration. To address these limitations, we propose Cobweb, a novel black-box fairness testing framework that combines genetic algorithms with explicit region-guided generation and multiobjective evolutionary search. First, the initial population is constructed via spatial uniform initialization to maximize instance diversity through pairwise distance optimization. Second, we introduce an explicit region-guided strategy and design a dualobjective optimization mechanism, which concurrently optimizes for both individual discrimination and spatial sparsity, enabling outward search into low-density regions. These mechanisms jointly improve Cobweb’s ability to generate diverse and spatially representative discriminatory instances. Extensive experiments on seven benchmark tabular datasets demonstrate that Cobweb achieves significant improvements over state-of-the-art baselines, with gains in effectiveness ($\sim 1.6 \times$), efficiency ($\sim 2.9 \times$) on average under fixed query budgets. In particular, Cobweb consistently achieves higher scores in three complementary diversity metrics, confirming its superior exploratory capability. Retraining target models with Cobwebgenerated instances leads to an average reduction of 57% in individual fairness violations, confirming the practical value of our approach. Yingqian Guo, Wentian Zhao |
APSEC | 2 |
| 2025 | Latent Search-Based Boundary Aware Fairness Testing for Deployed Deep ModelsabstractThe rapid adoption of deep neural networks (DNNs) in safety and social-critical applications has intensified concerns about discriminatory behaviour, especially when models are accessible only as black-box services. Existing individual-fairness testing methods either depend on gradient information or are tailored to low-dimensional, structured data, leaving the high-dimensional image domain largely unexplored. To bridge this gap, we introduce a fairness testing framework that is aware of the classifier’s decision boundary and uncovers discriminatory behaviour. The proposed method first constructs a linear surrogate of the classifier’s decision boundary within the latent space of a generative adversarial network (GAN) by leveraging an auxiliary dataset. This method follows a two-phase search paradigm. The first phase performs a coarse and wide-ranging sweep, steering the candidate samples towards a region of greatest predictive ambiguity, which lies near the decision boundary of the target model. The second phase refines the samples by thoroughly examining the neighborhoods of the candidate samples. Together, these two phases enable the rapid generation of extensive collections of inputs that expose discriminatory behaviour. Extensive experiments on public benchmarks reveal that the proposed approach outperforms state-of-the-art methods in both effectiveness and efficiency. Moreover, we empirically show that retraining models with the instances synthesized by our framework yields marked improvements in group fairness. We also extend the framework to tabular domains, where it exhibits comparably strong empirical performance. Qingyuan Sun, Wentian Zhao |
APSEC | 2 |
| 2025 | TPOTI: A Triplet-Network-based Obfuscated Tor Traffic Identification
Menglei Li, Wentian Zhao |
Networking | 3 |
| 2025 | An Embedded Covert Channel Construction using HLS Protocol
Dingming Liu, Wentian Zhao |
Networking | 3 |
| 2025 | Understanding and Mitigating Numerical Sources of Nondeterminism in LLM InferenceabstractLarge Language Models (LLMs) are now integral across various domains and have demonstrated impressive performance. Progress, however, rests on the premise that benchmark scores are both accurate and reproducible. We demonstrate that the reproducibility of LLM performance is fragile: changing system configuration, such as evaluation batch size, GPU count, and GPU version, can introduce significant differences in the generated responses.
This issue is especially pronounced in reasoning models, where minor rounding differences in early tokens can cascade into divergent chains of thought, ultimately affecting accuracy. For instance, under bfloat16 precision with greedy decoding, a reasoning model like DeepSeek-R1-Distill-Qwen-7B can exhibit up to 9\% variation in accuracy and 9,000 tokens difference in response length due to differences in GPU count, type, and evaluation batch size.
We trace the root cause of this variability to the non-associative nature of floating-point arithmetic under limited numerical precision.
This work presents the first systematic investigation into how numerical precision affects reproducibility in LLM inference. Through carefully controlled experiments across various hardware, software, and precision settings, we quantify when and how model outputs diverge.
Our analysis reveals that floating-point precision—while critical for reproducibility—is often neglected in evaluation practices.
Inspired by this, we develop a lightweight inference pipeline, dubbed LayerCast, that stores weights in 16-bit precision but performs all computations in FP32, balancing memory efficiency with numerical stability. Code is available at https://github.com/nanomaoli/llm_reproducibility. Jiayi Yuan 0001, Xinheng Ding, Wenya Xie, Yu-Jhe Li, Wentian Zhao, Kun Wan 0001, Xia Ben Hu, Zirui Liu 0001 |
NeurIPS | 6 |
| 2024 | Relational Distant Supervision for Image Captioning without Image-Text PairsabstractUnsupervised image captioning aims to generate descriptions of images without relying on any image-sentence pairs for training. Most existing works use detected visual objects or concepts as bridge to connect images and texts. Considering that the relationship between objects carries more information, we use the object relationship as a more accurate connection between images and texts. In this paper, we adapt the idea of distant supervision that extracts the knowledge about object relationships from an external corpus and imparts them to images to facilitate inferring visual object relationships, without introducing any extra pre-trained relationship detectors. Based on these learned informative relationships, we construct pseudo image-sentence pairs for captioning model training. Specifically, our method consists of three modules: (1) a relationship learning module that learns to infer relationships from images under the distant supervision; (2) a relationship-to-sentence module that transforms the inferred relationships into sentences to generate pseudo image-sentence pairs; (3) an image captioning module that is trained by using the generated image-sentence pairs. Promising results on three datasets show that our method outperforms the state-of-the-art methods of unsupervised image captioning. Yayun Qi, Wentian Zhao, Xinxiao Wu |
AAAI | 2 |
| 2024 | Boundary-Guided Black-Box Fairness TestingabstractAlthough deep learning models have achieved outstanding performance in many applications, there are still concerns about their fairness. A series of fairness testing methods, which evaluate the fairness of deep learning models by generating discriminatory samples, have been proposed. However, these methods either neglect the naturalness of discriminatory samples or roughly select natural discriminatory samples, leading to a decrease in efficiency. In this paper, we introduce a boundary-guided black-box fairness testing method to effectively generate individual discriminatory samples with high efficiency and enhanced naturalness. Our boundary-guided method involves a global exploration phase, which explores multiple paths from the initial samples to the surrogate decision boundary of the target model, imitated from the semantic latent space of a generative adversarial network (GAN). Then, a local perturbation phase explores the nearby space around a given sample for identifying potential discriminatory samples. Extensive experiments on various datasets demonstrate that our approach outperforms state-of-the-art methods in terms of efficiency and effectiveness while maintaining high naturalness. Ziqiang Yin, Wentian Zhao |
COMPSAC | 2 |
| 2024 | DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D VisionabstractWe have witnessed significant progress in deep learning-based 3D vision, ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However, existing scene-level datasets for deep learning-based 3D vision, limited to ei-ther synthetic environments or a narrow selection of real-world scenes, are quite insufficient. This insufficiency not only hinders a comprehensive benchmark of existing methods but also caps what could be explored in deep learning-based 3D analysis. To address this critical gap, we present DL3DV-10K, a large-scale scene dataset, featuring 51.2 million frames from 10,510 videos captured from 65 types of point- of-interest (POI) locations, covering both bounded and unbounded scenes, with different levels of reflection, transparency, and lighting. We conducted a comprehensive benchmark of recent NVS methods on DL3DV-10K, which revealed valuable insights for future research in NVS. In addition, we have obtained encouraging results in a pilot study to learn generalizable NeRF from DL3DV-10K, which manifests the necessity of a large-scale scene-level dataset to forge a path toward a foundation model for learning 3D representation. Our DL3DV-10K dataset, benchmark results, and models will be publicly accessible. Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan 0001, Lantao Yu, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua 0001, Tianyi Zhang 0001, Bedrich Benes, Aniket Bera |
CVPR | 4 |
| 2024 | Boosting Entity-Aware Image Captioning With Multi-Modal Knowledge GraphabstractEntity-aware image captioning aims to describe named entities and events related to the image by utilizing the background knowledge in the associated article. This task remains challenging as it is difficult to learn the association between named entities and visual cues due to the long-tail distribution of named entities. Furthermore, the complexity of the article brings difficulty in extracting fine-grained relationships between entities to generate informative event descriptions about the image. To tackle these challenges, we propose a novel approach that constructs a multi-modal knowledge graph (MMKG) to associate the visual objects with named entities and capture the relationship between entities simultaneously with the help of external knowledge collected from the web. Specifically, we build a text sub-graph by extracting named entities and their relationships from the article, and build an image sub-graph by detecting the objects in the image. To connect these two sub-graphs, we propose a cross-modal entity matching module trained using a knowledge base that contains Wikipedia entries and the corresponding images. Finally, the MMKG is integrated into the captioning model via a graph attention mechanism. Extensive experiments on both GoodNews and NYTimes800 k datasets demonstrate the effectiveness of our method. Wentian Zhao, Xinxiao Wu |
IEEE Trans. Multim. | 1 |
| 2023 | Topic-aware video summarization using multimodal transformer
Yubo Zhu, Wentian Zhao, Xinxiao Wu |
Pattern Recognit. | 2 |
| 2022 | Adaptive Recursive Circle Framework for Fine-Grained Action RecognitionabstractIntuitively, distinguishing fine-grained actions in videos requires recursively capturing subtle visual cues and learning abstract features. However, existing deep neural network based methods are counter-intuitive in that their network layers do not explicitly model the recursive feature abstraction. Therefore, we are motivated to propose an Adaptive Recursive Circle (ARC) framework that equips common neural network layers with recursive attention and recursive fusion. ARC layer inherits the same operators and parameters as the original layer, but, most critically, it treats the layer input as an evolving state, thus explicitly achieving recursive feature abstraction by alternating the state update and the feature generation. Specifically, at each recursive step, the input state is firstly updated via both recursive attention and recursive fusion from the previously generated features, and then the feature abstraction is performed with the newly updated input state. Significant improvements are observed on multiple datasets. For example, an ARC-equipped TSM-ResNet-18 outperforms TSM-ResNet-50 on the Something-Something V1 and Diving48 datasets with only half over-heads. Code will be available at: https://github.com/0HaNC/ARC-ActionRecog. Hanxi Lin, Wentian Zhao, Xinxiao Wu |
ICME | 2 |
| 2022 | A robust context attention network for human hand detection
Zhihuai Xie, Wentian Zhao, Zhenhua Guo 0001 |
Expert Syst. Appl. | 3 |
| 2022 | Learning Cooperative Neural Modules for Stylized Image Captioning
Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Multi-modal Dependency Tree for Video CaptioningabstractGenerating fluent and relevant language to describe visual content is critical for the video captioning task. Many existing methods generate captions using sequence models that predict words in a left-to-right order. In this paper, we investigate a graph-structured model for caption generation by explicitly modeling the hierarchical structure in the sentences to further improve the fluency and relevance of sentences. To this end, we propose a novel video captioning method that generates a sentence by first constructing a multi-modal dependency tree and then traversing the constructed tree, where the syntactic structure and semantic relationship in the sentence are represented by the tree topology. To take full advantage of the information from both vision and language, both the visual and textual representation features are encoded into each tree node. Different from existing dependency parsing methods that generate uni-modal dependency trees for language understanding, our method construct s multi-modal dependency trees for language generation of images and videos. We also propose a tree-structured reinforcement learning algorithm to effectively optimize the captioning model where a novel reward is designed by evaluating the semantic consistency between the generated sub-tree and the ground-truth tree. Extensive experiments on several video captioning datasets demonstrate the effectiveness of the proposed method. Wentian Zhao, Xinxiao Wu, Jiebo Luo 0001 |
NeurIPS | 1 |
| 2021 | Improve CAM with Auto-adapted Segmentation and Co-supervised AugmentationabstractWeakly Supervised Object Localization (WSOL) methods generate both classification and localization results by learning from only image category labels. Previous methods usually utilize class activation map (CAM) to obtain target object regions. However, most of them only focus on improving foreground object parts in CAM, but ignore the important effect of its background contents. In this paper, we propose a confidence segmentation (ConfSeg) module that builds confidence score for each pixel in CAM without introducing additional hyper-parameters. The generated sample-specific confidence mask is able to indicate the extent of determination for each pixel in CAM, and further supervises additional CAM extended from internal feature maps. Besides, we introduce Co-supervised Augmentation (CoAug) module to capture feature-level representation for foreground and background parts in CAM separately. Then a metric loss is applied at batch sample level to augment distinguish ability of our model, which helps a lot to localize more related object parts. Our final model, CSoA, combines the two modules and achieves superior performance, e.g. 37.69% and 48.81% Top-1 localization error on CUB-200 and ILSVRC datasets, respectively, which outperforms all previous methods and becomes the new state-of-the-art. Ziyi Kou, Guofeng Cui, Wentian Zhao, Chenliang Xu |
WACV | 4 |
| 2021 | How to Make a BLT Sandwich? Learning VQA towards Understanding Web Instructional VideosabstractUnderstanding web instructional videos is an essential branch of video understanding in two aspects. First, most existing video methods focus on short-term actions for a-few-second-long video clips; these methods are not directly applicable to long videos. Second, unlike unconstrained long videos, e.g., movies, instructional videos are more structured in that they have step-by-step procedures constraining the understanding task. In this work, we study problem-solving on instructional videos via Visual Question Answering (VQA). Surprisingly, it has not been an emphasis for the video community despite its rich applications. We thereby introduce YouCookQA, an annotated QA dataset for instructional videos based on YouCook2 [27]. The questions in YouCookQA are not limited to cues on a single frame but relations among multiple frames in the temporal dimension. Observing the lack of effective representations for modeling long videos, we propose a set of carefully designed models including a Recurrent Graph Convolutional Network (RGCN) that captures both temporal order and relational information. Furthermore, we study multiple modalities including descriptions and transcripts for the purpose of boosting video understanding. Extensive experiments on YouCookQA suggest that RGCN performs the best in terms of QA accuracy and better performance is gained by introducing human-annotated descriptions. YouCookQA dataset is available at https://github.com/Jossome/YoucookQA. Wentian Zhao, Ziyi Kou, Jing Shi 0005, Chenliang Xu |
WACV | 2 |
| 2021 | Cross-Domain Image Captioning via Cross-Modal Retrieval and Model AdaptationabstractIn recent years, large scale datasets of paired images and sentences have enabled the remarkable success in automatically generating descriptions for images, namely image captioning. However, it is labour-intensive and time-consuming to collect a sufficient number of paired images and sentences in each domain. It may be beneficial to transfer the image captioning model trained in an existing domain with pairs of images and sentences (i.e., source domain) to a new domain with only unpaired data (i.e., target domain). In this paper, we propose a cross-modal retrieval aided approach to cross-domain image captioning that leverages a cross-modal retrieval model to generate pseudo pairs of images and sentences in the target domain to facilitate the adaptation of the captioning model. To learn the correlation between images and sentences in the target domain, we propose an iterative cross-modal retrieval process where a cross-modal retrieval model is first pre-trained using the source domain data and then applied to the target domain data to acquire an initial set of pseudo image-sentence pairs. The pseudo image-sentence pairs are further refined by iteratively fine-tuning the retrieval model with the pseudo image-sentence pairs and updating the pseudo image-sentence pairs using the retrieval model. To make the linguistic patterns of the sentences learned in the source domain adapt well to the target domain, we propose an adaptive image captioning model with a self-attention mechanism fine-tuned using the refined pseudo image-sentence pairs. Experimental results on several settings where MSCOCO is used as the source domain and five different datasets (Flickr30k, TGIF, CUB-200, Oxford-102 and Conceptual) are used as the target domains demonstrate that our method achieves mostly better or comparable performance against the state-of-the-art methods. We also extend our method to cross-domain video captioning where MSR-VTT is used as the source domain and two other datasets (MSVD and Charades Captions) are used as the target domains to further demonstrate the effectiveness of our method. Wentian Zhao, Xinxiao Wu, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | MemCap: Memorizing Style Knowledge for Image CaptioningabstractGenerating stylized captions for images is a challenging task since it requires not only describing the content of the image accurately but also expressing the desired linguistic style appropriately. In this paper, we propose MemCap, a novel stylized image captioning method that explicitly encodes the knowledge about linguistic styles with memory mechanism. Rather than relying heavily on a language model to capture style factors in existing methods, our method resorts to memorizing stylized elements learned from training corpus. Particularly, we design a memory module that comprises a set of embedding vectors for encoding style-related phrases in training corpus. To acquire the style-related phrases, we develop a sentence decomposing algorithm that splits a stylized sentence into a style-related part that reflects the linguistic style and a content-related part that contains the visual content. When generating captions, our MemCap first extracts content-relevant style knowledge from the memory module via an attention mechanism and then incorporates the extracted knowledge into a language model. Extensive experiments on two stylized image captioning datasets (SentiCap and FlickrStyle10K) demonstrate the effectiveness of our method. Wentian Zhao, Xinxiao Wu, Xiaoxun Zhang |
AAAI | 1 |
| 2020 | Video Question Answering on Screencast TutorialsabstractThis paper presents a new video question answering task on screencast tutorials. We introduce a dataset including question, answer and context triples from the tutorial videos for a software. Unlike other video question answering works, all the answers in our dataset are grounded to the domain knowledge base. An one-shot recognition algorithm is designed to extract the visual cues, which helps enhance the performance of video question answering. We also propose several baseline neural network architectures based on various aspects of video contexts from the dataset. The experimental results demonstrate that our proposed models significantly improve the question answering performances by incorporating multi-modal contexts and domain knowledge. Wentian Zhao, Seokhwan Kim, Hailin Jin |
IJCAI | 1 |
| 2019 | Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningabstractVideo captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning approach that takes into account both visual perception and syntax representation learning to generate accurate descriptions of videos. Specifically, we use sentence templates composed of Part-of-Speech (POS) tags to represent the syntax structure of captions, and accordingly, syntax representation learning is performed by directly inferring POS tags from videos. The visual perception is implemented by a mixture model which translates visual cues into lexical words that are conditional on the learned syntactic structure of sentences. Thus, a video captioning task consists of two sub-tasks: video POS tagging and visual cue translation, which are jointly modeled and trained in an end-to-end fashion. Evaluations on three public benchmark datasets demonstrate that our proposed method achieves substantially better performance than the state-of-the-art methods, which validates the superiority of joint modeling of syntax representation learning and visual perception for video captioning. Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001, Yunde Jia |
ICCV | 3 |
| 2019 | GAN-EM: GAN Based EM Learning FrameworkabstractExpectation maximization (EM) algorithm is to find maximum likelihood solution for models having latent variables. A typical example is Gaussian Mixture Model (GMM) which requires Gaussian assumption, however, natural images are highly non-Gaussian so that GMM cannot be applied to perform image clustering task on pixel space. To overcome such limitation, we propose a GAN based EM learning framework that can maximize the likelihood of images and estimate the latent variables. We call this model GAN-EM, which is a framework for image clustering, semi-supervised classification and dimensionality reduction. In M-step, we design a novel loss function for discriminator of GAN to perform maximum likelihood estimation (MLE) on data with soft class label assignments. Specifically, a conditional generator captures data distribution for K classes, and a discriminator tells whether a sample is real or fake for each class. Since our model is unsupervised, the class label of real data is regarded as latent variable, which is estimated by an additional network (E-net) in E-step. The proposed GAN-EM achieves state-of-the-art clustering and semi-supervised classification results on MNIST, SVHN and CelebA, as well as comparable quality of generated images to other recently developed generative models. Wentian Zhao, Zhihuai Xie, Jing Shi 0005, Chenliang Xu |
IJCAI | 1 |
| 2018 | Performance Guaranteed Traffic Signal Control with Frame-Based AlgorithmabstractIn urban area, fast growth in the number of vehicles has led to a series of traffic problems, including traffic jams,high traffic accident rates, etc. Efficient traffic signal control methods has been shown to be an essential way to significantly mitigate traffic problems. In this poster, different from previous online methods, we propose a frame-based model and an efficient algorithm to solve the drawbacks of online algorithm by scheduling the vehicles that have accumulated at the intersection over a period of time. Preliminary experiments exhibit that the proposed algorithm could greatly improve the throughput of the intersection. Xili Wan, Wentian Zhao, Xinjie Guan, Feng Ye 0002, Guangwei Bai |
SECON | 2 |
| 2016 | Fault diagnosis network design for vehicle on-board equipments of high-speed railway: A deep learning approach
Jiateng Yin, Wentian Zhao |
Eng. Appl. Artif. Intell. | 2 |