VLDB 2026 Research / reviewers in the wild / expert
Ailin Deng
dblp:70/3580
· DBLP profile ↗
13ranked-venue papers
5as first author
10since 2021 · last 2025
0009-0009-2859-3665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Words or Vision: Do Vision-Language Models Have Blind Faith in Text?abstractVision-Language Models (VLMs) excel in integrating visual and textual information for vision-centric tasks, but their handling of inconsistencies between modalities is underexplored. We investigate VLMs’ modality preferences when faced with visual data and varied textual inputs in vision-centered settings. By introducing textual variations to four vision-centric tasks and evaluating ten Vision-Language Models (VLMs), we discover a "blind faith in text" phenomenon: VLMs disproportionately trust textual data over visual data when inconsistencies arise, leading to significant performance drops under corrupted text and raising safety concerns. We analyze factors influencing this text bias, including instruction prompts, language model size, text relevance, token order, and the interplay between visual and textual certainty. While certain factors, such as scaling up the language model size, slightly mitigate text bias, others like token order can exacerbate it due to positional biases inherited from language models. To address this issue, we explore supervised fine-tuning with text augmentation and demonstrate its effectiveness in reducing text bias. Additionally, we provide a theoretical analysis suggesting that the blind faith in text phenomenon may stem from an imbalance of pure text and multi-modal data during training. Our findings highlight the need for balanced training and careful consideration of modality interactions in VLMs to enhance their robustness and reliability in handling multi-modal data inconsistencies. Ailin Deng, Tri Cao, Bryan Hooi |
CVPR | 1 |
| 2025 | MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning ResearchabstractRecent advancements in AI agents have demonstrated their growing potential to drive and support scientific discovery. In this work, we introduce MLR-Bench, a comprehensive benchmark for evaluating AI agents on open-ended machine learning research. MLR-Bench includes three key components: (1) 201 research tasks sourced from NeurIPS, ICLR, and ICML workshops covering diverse ML topics; (2) MLR-Judge, an automated evaluation framework combining LLM-based reviewers with carefully designed review rubrics to assess research quality; and (3) MLR-Agent, a modular agent scaffold capable of completing research tasks through four stages: idea generation, proposal formulation, experimentation, and paper writing. Our framework supports both stepwise assessment across these distinct research stages, and end-to-end evaluation of the final research paper. We then use MLR-Bench to evaluate six frontier LLMs and an advanced coding agent, finding that while LLMs are effective at generating coherent ideas and well-structured papers, current coding agents frequently (e.g., in 80\% of the cases) produce fabricated or invalidated experimental results—posing a major barrier to scientific reliability. We validate MLR-Judge through human evaluation, showing high agreement with expert reviewers, supporting its potential as a scalable tool for research evaluation. We open-source MLR-Bench to help the community benchmark, diagnose, and improve AI research agents toward trustworthy and transparent scientific discovery. Miao Xiong, Ailin Deng, Yue Liu 0008, Bryan Hooi |
NeurIPS | 5 |
| 2024 | ID3: Identity-Preserving-yet-Diversified Diffusion Models for Synthetic Face Recognition
Jianqing Xu, Shen Li 0004, Miao Xiong, Ailin Deng, Jiazhen Ji, Yuge Huang, Guodong Mu, Wenjie Feng 0001, Shouhong Ding, Bryan Hooi |
NeurIPS | 5 |
| 2023 | Prompt-and-Align: Prompt-Based Social Alignment for Few-Shot Fake News DetectionabstractDespite considerable advances in automated fake news detection, due to the timely nature of news, it remains a critical open question how to effectively predict the veracity of news articles based on limited fact-checks. Existing approaches typically follow a "Train-from-Scratch" paradigm, which is fundamentally bounded by the availability of large-scale annotated data. While expressive pre-trained language models (PLMs) have been adapted in a "Pre-Train-and-Fine-Tune" manner, the inconsistency between pre-training and downstream objectives also requires costly task-specific supervision. In this paper, we propose "Prompt-and-Align" (P&A), a novel prompt-based paradigm for few-shot fake news detection that jointly leverages the pre-trained knowledge in PLMs and the social context topology. Our approach mitigates label scarcity by wrapping the news article in a task-related textual prompt, which is then processed by the PLM to directly elicit task-specific knowledge. To supplement the PLM with social context without inducing additional training overheads, motivated by empirical observation on user veracity consistency (i.e., social users tend to consume news of the same veracity type), we further construct a news proximity graph among news articles to capture the veracity-consistent signals in shared readerships, and align the prompting predictions along the graph edges in a confidence-informed manner. Extensive experiments on three real-world benchmarks demonstrate that P&A sets new states-of-the-art for few-shot fake news detection performance by significant margins. Shen Li 0004, Ailin Deng, Miao Xiong, Bryan Hooi |
CIKM | 3 |
| 2023 | Probabilistic Knowledge Distillation of Face EnsemblesabstractMean ensemble (i.e. averaging predictions from multiple models) is a commonly-used technique in machine learning that improves the performance of each individual model. We formalize it as feature alignment for ensemble in open-set face recognition and generalize it into Bayesian Ensemble Averaging (BEA) through the lens of probabilistic modeling. This generalization brings up two practical benefits that existing methods could not provide: (1) the uncertainty of a face image can be evaluated and further decomposed into aleatoric uncertainty and epistemic uncertainty, the latter of which can be used as a measure for out-of-distribution detection of faceness; (2) a BEA statistic provably reflects the aleatoric uncertainty of a face image, acting as a measure for face image quality to improve recognition performance. To inherit the uncertainty estimation capability from BEA without the loss of inference efficiency, we propose BEA-KD, a student model to distill knowledge from BEA. BEA-KD mimics the overall behavior of ensemble members and consistently outperforms SOTA knowledge distillation methods on various challenging benchmarks. Jianqing Xu, Shen Li 0004, Ailin Deng, Miao Xiong, Jiaxiang Wu 0002, Shouhong Ding, Bryan Hooi |
CVPR | 3 |
| 2023 | Great Models Think Alike: Improving Model Reliability via Inter-Model Latent AgreementabstractReliable application of machine learning is of primary importance to the practical deployment of deep learning methods. A fundamental challenge is that models are often unreliable due to overconfidence. In this paper, we estimate a model’s reliability by measuring the agreement between its latent space, and the latent space of a foundation model. However, it is challenging to measure the agreement between two different latent spaces due to their incoherence, e.g., arbitrary rotations and different dimensionality. To overcome this incoherence issue, we design a neighborhood agreement measure between latent spaces and find that this agreement is surprisingly well-correlated with the reliability of a model’s predictions. Further, we show that fusing neighborhood agreement into a model’s predictive confidence in a post-hoc way significantly improves its reliability. Theoretical analysis and extensive experiments on failure detection across various datasets verify the effectiveness of our method on both in-distribution and out-of-distribution settings. Ailin Deng, Miao Xiong, Bryan Hooi |
ICML | 1 |
| 2023 | Proximity-Informed Calibration for Deep Neural NetworksabstractConfidence calibration is central to providing accurate and interpretable uncertainty estimates, especially under safety-critical scenarios. However, we find that existing calibration algorithms often overlook the issue of proximity bias, a phenomenon where models tend to be more overconfident in low proximity data (i.e., data lying in the sparse region of the data distribution) compared to high proximity samples, and thus suffer from inconsistent miscalibration across different proximity samples. We examine the problem over $504$ pretrained ImageNet models and observe that: 1) Proximity bias exists across a wide variety of model architectures and sizes; 2) Transformer-based models are relatively more susceptible to proximity bias than CNN-based models; 3) Proximity bias persists even after performing popular calibration algorithms like temperature scaling; 4) Models tend to overfit more heavily on low proximity samples than on high proximity samples. Motivated by the empirical findings, we propose ProCal, a plug-and-play algorithm with a theoretical guarantee to adjust sample confidence based on proximity. To further quantify the effectiveness of calibration algorithms in mitigating proximity bias, we introduce proximity-informed expected calibration error (PIECE) with theoretical analysis. We show that ProCal is effective in addressing proximity bias and improving calibration on balanced, long-tail, and distribution-shift settings under four metrics over various model architectures. We believe our findings on proximity bias will guide the development of fairer and better-calibrated} models, contributing to the broader pursuit of trustworthy AI. Miao Xiong, Ailin Deng, Pang Wei W. Koh, Shen Li 0004, Jianqing Xu, Bryan Hooi |
NeurIPS | 2 |
| 2022 | Trust, but Verify: Using Self-supervised Probing to Improve Trustworthiness
Ailin Deng, Shen Li 0004, Miao Xiong, Bryan Hooi |
ECCV (13) | 1 |
| 2022 | CADET: Calibrated Anomaly Detection for Mitigating Hardness BiasabstractThe detection of anomalous samples in large, high-dimensional datasets is a challenging task with numerous practical applications. Recently, state-of-the-art performance is achieved with deep learning methods: for example, using the reconstruction error from an autoencoder as anomaly scores. However, the scores are uncalibrated: that is, they follow an unknown distribution and lack a clear interpretation. Furthermore, the reconstruction error is highly influenced by the `hardness' of a given sample, which leads to false negative and false positive errors. In this paper, we empirically show the significance of this hardness bias present in a range of recent deep anomaly detection methods. To mitigate this, we propose an efficient and plug-and-play error calibration method which mitigates this hardness bias in the anomaly scoring without the need to retrain the model. We verify the effectiveness of our method on a range of image, time-series, and tabular datasets and against several baseline methods. Ailin Deng, Adam Goodge, Lang Yi Ang, Bryan Hooi |
IJCAI | 1 |
| 2021 | Graph Neural Network-Based Anomaly Detection in Multivariate Time SeriesabstractGiven high-dimensional time series data (e.g., sensor data), how can we detect anomalous events, such as system faults and attacks? More challengingly, how can we do this in a way that captures complex inter-sensor relationships, and detects and explains anomalies which deviate from these relationships? Recently, deep learning approaches have enabled improvements in anomaly detection in high-dimensional datasets; however, existing methods do not explicitly learn the structure of existing relationships between variables, or use them to predict the expected behavior of time series. Our approach combines a structure learning approach with graph neural networks, additionally using attention weights to provide explainability for the detected anomalies. Experiments on two real-world sensor datasets with ground truth anomalies show that our method detects anomalies more accurately than baseline approaches, accurately captures correlations between sensors, and allows users to deduce the root cause of a detected anomaly. Ailin Deng, Bryan Hooi |
AAAI | 1 |
| 2017 | 3-D-MIMO With Massive Antennas Paves the Way to 5G Enhanced Mobile Broadband: From System Design to Field TrialsabstractThree-dimensional (3D) multiple input and multiple output (3D-MIMO) with massive antennas is a key technology to achieve high spectral efficiency and user experienced data rate for the fifth generation (5G) mobile communication system. To implement 3D-MIMO in 5G system, practical constraints on the product design should be considered. This paper proposes a systematic design for the 3D-MIMO product by considering the restrictions of both base band and the hardware, including cost, size, weight, and heat dissipation. The design has been implemented for 2.6-GHz time-division duplex band, and field trials have been conducted for performance validation with practical intercell interference in commercial network. The trial results show that this 3D-MIMO design can meet the spectral efficiency requirement of the 5G enhanced mobile broadband services. The performance gain of 3D-MIMO varies with the traffic load. When the traffic load is heavy, 3D-MIMO can enhance the cell throughput by 4~6.7 times. When the traffic load is low, the performance gain of this 3D-MIMO design decreases. The results from field trial also show that the performance of 3D-MIMO degrades in mobility scenarios, where further enhancement on acquiring instant channel status information are necessary to improve the robustness of 3D-MIMO to mobility. Guangyi Liu 0001, Xueying Hou, Jing Jin 0007, Fei Wang 0004, Qixing Wang, Yue Hao 0006, Yuhong Huang, Xiaoyun Wang 0005, Ailin Deng |
IEEE J. Sel. Areas Commun. | 10 |
| 2003 | Image retrieval by fuzzy clustering of relevance feedback recordsabstractWe present an image retrieval method based on the accumulated user relevance feedback records. Our method conducts the semi-supervised fuzzy clustering on the records, and the subsequent information filtering within the target cluster is performed to guide the refinement of query parameters. During information filtering, both the user's relevance evaluation and the corresponding query image of the records are used to predict the semantic correlation between the current retrieval query sample and the database images. Experiment results show that our method outperforms the traditional ones in both efficiency and effectiveness. Ailin Deng |
ICME | 4 |
| 2003 | An Image Retrieval Method Based on Information Filtering of User Relevance Feedback Records
Qi Zhang 0025, Ailin Deng, Liang Zhang 0019, Baile Shi |
WAIM | 4 |