Chao Wan

dblp:64/8352 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 62% Recommender systems · 27% Indexing and storage engines · 10%
Artificial intelligence
2 papers
Language models and text generation · 46% Generative modeling · 40% Trustworthy machine learning · 8%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Information retrieval › evaluation › benchmark
benchmark construction
0.912025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Recommender systems › interactive recommendation
conversational recommendation
0.912025
Collaborative Retrieval for Large Language Model-based Conversational Recommender Systems · WWW 2025
Recommender systems
large language model-based recommendation
0.912025
Collaborative Retrieval for Large Language Model-based Conversational Recommender Systems · WWW 2025
Information retrieval
retrieval evaluation
0.912025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Machine learning › Generative modeling
diffusion model
0.712023
Latent Diffusion for Language Generation · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.712023
Latent Diffusion for Language Generation · NeurIPS 2023
Information retrieval › indexing
differentiable search index
0.712023
IncDSI: Incrementally Updatable Document Retrieval · ICML 2023
Information retrieval
document retrieval
0.712023
IncDSI: Incrementally Updatable Document Retrieval · ICML 2023
Indexing and storage engines › index maintenance
incremental index update
0.712023
IncDSI: Incrementally Updatable Document Retrieval · ICML 2023
Information retrieval › retrieval models
neural retrieval
0.712023
IncDSI: Incrementally Updatable Document Retrieval · ICML 2023
Machine learning › Trustworthy machine learning
data leakage
0.312025
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation · ICML 2025
Information retrieval
retrieval-augmented generation
0.312025
Collaborative Retrieval for Large Language Model-based Conversational Recommender Systems · WWW 2025

Methods — techniques the papers use, named apart from their topics

question-answer pair generation · 1.7document corpus generation · 1.7retrieval-augmented generation · 0.9large language model · 0.9collaborative filtering · 0.9encoder-decoder language model · 0.7diffusion model · 0.7constrained optimization · 0.7autoencoder · 0.7
YearPublicationVenuePosition
2026 Event-triggered stochastic model predictive control based on distributionally robust optimization approach for network control systems under DoS attacks
Chao Wan
Signal Process.2
2026 Networked Filtering With Self Triggered Communication
abstract
In practical applications, the topology of a sensor network is commonly time-varying and even unknown, but almost all existing networked filtering methods rely on this unavailable topology information. In view of this, we investigate a topology ignorant communication scheme that does not need to know the topology of the network, where nodes use only neighboring information to trigger their activations. The triggering condition considers the estimation accuracy which benefits the purpose of networked filtering (i.e., distributed filtering in sensor network). With this communication mode, we propose a filter: self triggered activation based distributed filter (STA-DF). Its convergence to the centralized estimation along with iteration length and stability are analyzed. Simulation results are provided to verify its superiority to existing methods.
Chao Wan, Zhansheng Duan
IEEE Signal Process. Lett.2
2025 PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
abstract
High-quality benchmarks are essential for evaluating reasoning and retrieval capabilities of large language models (LLMs). However, curating datasets for this purpose is not a permanent solution as they are prone to data leakage and inflated performance results. To address these challenges, we propose PhantomWiki: a pipeline to generate unique, factually consistent document corpora with diverse question-answer pairs. Unlike prior work, PhantomWiki is neither a fixed dataset, nor is it based on any existing data. Instead, a new PhantomWiki instance is generated on demand for each evaluation. We vary the question difficulty and corpus size to disentangle reasoning and retrieval capabilities, respectively, and find that PhantomWiki datasets are surprisingly challenging for frontier LLMs. Thus, we contribute a scalable and data leakage-resistant framework for disentangled evaluation of reasoning, retrieval, and tool-use abilities.
Albert Gong, Kamile Stankeviciute, Chao Wan, Anmol Kabra, Raphael Thesmar, Johann Lee, Julius Klenke, Carla P. Gomes, Kilian Q. Weinberger
ICML3
2025 Collaborative Retrieval for Large Language Model-based Conversational Recommender Systems
abstract
Conversational recommender systems (CRS) aim to provide personalized recommendations via interactive dialogues with users. While large language models (LLMs) enhance CRS with their superior understanding of context-aware user preferences, they typically struggle to leverage behavioral data, which have proven to be important for classical collaborative filtering (CF)-based approaches. For this reason, we propose CRAG-Collaborative Retrieval Augmented Generation for LLM-based CRS. To the best of our knowledge, CRAG is the first approach that combines state-of-the-art LLMs with CF for conversational recommendations. Our experiments on two publicly available movie conversational recommendation datasets, i.e., a refined Reddit dataset (which we name Reddit-v2) as well as the Redial dataset, demonstrate the superior item coverage and recommendation performance of CRAG, compared to several CRS baselines. Moreover, we observe that the improvements are mainly due to better recommendation accuracy on recently released movies. The code and data are available at https://github.com/yaochenzhu/CRAG.
Yaochen Zhu, Chao Wan, Harald Steck, Dawen Liang, Yesu Feng, Nathan Kallus, Jundong Li
WWW2
2025 Deep learning-enhanced Koopman-operator-based self-triggered distributionally robust optimization SMPC strategy for satellite attitude control
Chao Wan
Knowl. Based Syst.3
2023 IncDSI: Incrementally Updatable Document Retrieval
abstract
Differentiable Search Index is a recently proposed paradigm for document retrieval, that encodes information about a corpus of documents within the parameters of a neural network and directly maps queries to corresponding documents. These models have achieved state-of-the-art performances for document retrieval across many benchmarks. These kinds of models have a significant limitation: it is not easy to add new documents after a model is trained. We propose IncDSI, a method to add documents in real time (about 20-50ms per document), without retraining the model on the entire dataset (or even parts thereof). Instead we formulate the addition of documents as a constrained optimization problem that makes minimal changes to the network parameters. Although orders of magnitude faster, our approach is competitive with re-training the model on the whole dataset and enables the development of document retrieval systems that can be updated with new information in real-time. Our code for IncDSI is available at https://github.com/varshakishore/IncDSI.
Varsha Kishore, Chao Wan, Justin Lovelace, Yoav Artzi, Kilian Q. Weinberger
ICML2
2023 Latent Diffusion for Language Generation
abstract
Diffusion models have achieved great success in modeling continuous data modalities such as images, audio, and video, but have seen limited use in discrete domains such as language. Recent attempts to adapt diffusion to language have presented diffusion as an alternative to existing pretrained language models. We view diffusion and existing language models as complementary. We demonstrate that encoder-decoder language models can be utilized to efficiently learn high-quality language autoencoders. We then demonstrate that continuous diffusion models can be learned in the latent space of the language autoencoder, enabling us to sample continuous latent representations that can be decoded into natural language with the pretrained decoder. We validate the effectiveness of our approach for unconditional, class-conditional, and sequence-to-sequence language generation. We demonstrate across multiple diverse data sets that our latent language diffusion models are significantly more effective than previous diffusion language models. Our code is available at \url{https://github.com/justinlovelace/latent-diffusion-for-language}.
Justin Lovelace, Varsha Kishore, Chao Wan, Eliot Shekhtman, Kilian Q. Weinberger
NeurIPS3
2023 Multi-time Scale Attention Network for WEEE reverse logistics return prediction
Jia Zhang 0029, Min Gao 0001, Jinyong Gao, Meiling Deng, Chao Wan, Yanyan Yang 0002
Expert Syst. Appl.7
2021 A fast DC-based dictionary learning algorithm with the SCAD penalty
Zhenni Li, Chao Wan, Benying Tan, Zuyuan Yang, Shengli Xie 0001
Neurocomputing2
2021 Wideband cryogenic amplifier for a superconducting nanowire single-photon detector
abstract
We present a low-power inductorless wideband differential cryogenic amplifier using a 0.13-µm SiGe BiCMOS process for a superconducting nanowire single-photon detector (SNSPD). With a shunt-shunt feedback and capacitive coupling structure, theoretical analysis and simulations were undertaken, highlighting the relationship of the amplifier gain with the tunable design parameters of the circuit. In this way, the design and optimization flexibility can be increased, and a required gain can be achieved even without an accurate cryogenic device model. To realize a flat terminal impedance over the frequency of interest, an RC shunt compensation structure was employed, improving the amplifier’s closed-loop stability and suppressing the amplifier overshoot. The S -parameters and transient performance were measured at room temperature (300 K) and cryogenic temperature (4.2 K). With good input and output matching, the measurement results showed that the amplifier achieved a 21-dB gain with a 3-dB bandwidth of 1.13 GHz at 300 K. At 4.2 K, the gain of the amplifier can be tuned from 15 to 24 dB, achieving a 3-dB bandwidth spanning from 120 kHz to 1.3 GHz and consuming only 3.1 mW. Excluding the chip pads, the amplifier chip core area was only about 0.073 mm 2 .
Lianming Li, Xu Wu 0002, Xiaokang Niu, Chao Wan, Lin Kang, Xiaoqing Jia, Labao Zhang, Qingyuan Zhao, Xuecou Tu
Frontiers Inf. Technol. Electron. Eng.5
2019 Gossip-based Distributed Filtering Over Networks Using Projection
Chao Wan, Yongxin Gao, X. Rong Li
FUSION1
2018 Distributed Filtering Over Networks Using Greedy Gossip
abstract
This paper studies the problem of distributed filtering for state estimation of a dynamic system by using observations from sensors in a network, and proposes a greedy gossip-based distributed filtering (GG-DF) algorithm. The sensor nodes can make estimation and work collaboratively. The information transmission across the network abides by the asynchronous gossip strategy that only two neighboring nodes are selected to communicate and exchange information with each other in each communication round. First, we propose a cost function of the estimation error of the entire network. Then, we derive our algorithm by making a greedy selection to minimize the cost. Finally, we provide performance and convergence analysis of the proposed algorithm, along with simulation results compared with existing methods.
Chao Wan, Yongxin Gao, X. Rong Li, Enbin Song
FUSION1
2018 A computer assisted automatic grenade throw training system with simple digital cameras
Bin Liu 0040, Yubo Ma, Chao Wan
Multim. Tools Appl.5
2017 Distributed filtering over networks based on diffusion strategy
abstract
This paper studies and formulates the problem of distributed filtering with a diffusion strategy for state estimation of a dynamic system by using observations from sensors in a network. The sensor-nodes have estimation ability and work in a collaborative manner. The information transmission across the network abides by the diffusion strategy that each node communicates only with its neighbors. First, we propose a cost function for a trade-off between accuracy and consensus. Then, we derive our algorithm based on this cost and analyze its mean-square performance. Illustrative numerical examples are provided to verify the good performance of our method.
Chao Wan, Yongxin Gao, X. Rong Li
FUSION1