VLDB 2026 Research / reviewers in the wild / expert
Chenyou Fan
dblp:180/5779 · also Chengyou Fan
· DBLP profile ↗
48ranked-venue papers
14as first author
39since 2021 · last 2026
0000-0002-9835-8507ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Fine-Tuned CLIP-OOD Detection Method with Double Loss Constraint Through Optimal Transport Semantic AlignmentabstractDetecting Out-Of-Distribution (OOD) samples in image classification is crucial for model reliability. With the rise of Vision-Language Models (VLMs), CLIP-OOD has become a research hotspot. However, we observe the Low Focus Attention phenomenon from the image encoders of CLIP, which means the attention of image encoders often spreads to non-in-distribution regions. This phenomenon comes from the semantic mismalignment and inter-class feature confusion. To address these issues, we propose a novel fine-tuned OOD detection method with the Double loss constraint based on Optimal Transport (DOT-OOD). DOT-OOD integrates the Double Loss Constraint (DLC) module and Optimal Transport (OT) module. The DLC module comprises the Aligned Image-Text Concept Matching Loss and the Negative Sample Repulsion Loss, which respectively (1) focus on the core semantics of ID images and achieve cross-modal semantic alignment, (2) expand inter-class distances and enhance discriminative. While the OT module is introduced to obtain enhanced image feature representations. Extensive experimental results show that in the 16-shot scenario of the ImageNet-1k benchmark, DOT-OOD reduces the FPR95 by over 10% and improves the AUROC from 94.48% to 96.57% compared with SOTAs. Hengyang Lu, Yuntao Du 0001, Chenyou Fan |
AAAI | 7 |
| 2026 | SF-QC: Explore the Selection Bias of Large Language Models in Zero-Shot Out-of-Distribution Intent DetectionabstractOut-of-distribution (OOD) intent detection aims to identify user queries that exceed predefined intent categories and is a crucial technology for ensuring interaction reliability and user experience in practical applications such as dialogue systems and intelligent customer service. Although large language models (LLMs) perform excellently in natural language processing tasks, they encounter difficulties with zero-shot OOD intent detection. This article reveals the dual biases of LLMs in zero-shot OOD intent detection: 1) LLMs tend to prefer the intents that are presented earlier in the intent list; and 2) LLMs frequently misclassify unknown intent as in-distribution (ID) intent. These biases severely limit the deployment of LLMs in intent-recognition systems that require high reliability. To address this problem, we propose a novel semantics-based sorting-filtering and querying-confirming (SF-QC) method: the semantic sorting-filtering (SF) module dynamically adjusts the intent list to mitigate position preference, and then the querying-confirming (QC) module deeply validates the initial ID response to reduce OOD misclassification, starting from the root cause of biases to guide the LLM to make unbiased judgments. Experiments on two mainstream intent datasets show that our proposed SF-QC approach has up to 7.82% (Macro-F1) and 9.76% (ACC) improvement in overall performance, and up to 14.9% (Macro-F1) and 20.66% (ACC) improvement in OOD performance over the chain-of-thought (CoT) method. The excellent detection performance of SF-QC provides key technical support for the robust deployment of LLMs in practical applications such as dialogue systems, helping to reduce system risks and enhance user experience. Hengyang Lu, Xin-Yi Liu, Jiaming Zhang 0006, Chenyou Fan, Wei Fang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Don't Judge From a Single Perspective: LVLM-Based Multiview Multimodal Fake News DetectionabstractThe prevalence of social media platforms has accelerated the propagation of fake news, with carefully crafted multimodal fake news often gaining public trust and causing severe impacts. Methods based on multimodal small language models (MSLMs) struggle in scenarios with scarce labeled data. Large vision-language model (LVLM)-based approaches, though alleviating this limitation, often face two major problems: They rely too much on text-image consistency when predicting the authenticity of news. They judge which samples need to introduce external knowledge based on LVLMs’ direct outputs, which is not accurate because of LVLMs’ overconfidence. This article proposes a novel LVLM-based multiview multimodal fake news detection (MV-MFND) framework to address these problems. MV-MFND leverages LVLMs to analyze text, image, and text-image pairs separately, thus predicting authenticity from multiple perspectives. MV-MFND determines which samples need to introduce external knowledge by analyzing the softmax probability of specific tokens output by the LVLM. Our framework MV-MFND achieves SOTA performance on three real-world datasets Fakeddit, Twitter, and MR2-en on Qwen. Hengyang Lu, Xinnan Liu, Qianyi Zhan, Chenyou Fan, Wei Fang 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Forward KL Regularized Preference Optimization for Aligning Diffusion PoliciesabstractDiffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous methods conduct return-conditioned policy generation or Reinforcement Learning (RL)-based policy optimization, while they both rely on pre-defined reward functions. In this work, we propose a novel framework, Forward KL regularized Preference optimization for aligning Diffusion policies, to align the diffusion policy with preferences directly. We first train a diffusion policy from the offline dataset without considering the preference, and then align the policy to the preference data via direct preference optimization. During the alignment phase, we formulate direct preference learning in a diffusion policy, where the forward KL regularization is employed in preference optimization to avoid generating out-of-distribution actions. We conduct extensive experiments for MetaWorld manipulation and D4RL tasks. The results show our method exhibits superior alignment with preferences and outperforms previous state-of-the-art algorithms. Zhao Shan, Chenyou Fan, Jiyuan Shi, Chenjia Bai |
AAAI | 2 |
| 2025 | Who Is Undercover? Guiding LLMs to Explore Multi-perspective Team Tactic in the Game
Ruiqi Dong, Zhixuan Liao 0001, Danni Ma, Chenyou Fan |
DASFAA (6) | 5 |
| 2025 | Hybrid Feature Fusion for Enhancing Medical Document EmbeddingabstractDespite the strong capabilities of large language models in generative tasks, issues related to information unreliability and hallucinations pose significant challenges in high-precision fields, such as drug analysis and recommendations in the medical domain. In this work, we introduce the HFFN model, a retrieval-augmented framework designed for medical document retrieval tasks, which combines an embedding backbone with a Hybrid Feature Fusion module to enhance retrieval quality. HFFN improves model representation by learning to control the weight parameters of nonlinear features, thereby avoiding the instability associated with using the same activation function across different datasets. Experimental results demonstrate that, compared to a single-hidden-layer MLP, HFFN improves NDCG score across various baseline embedding models by 2.3%-15.1%. Tianqi Pang, Xuncan Xiao, Xiaofan Zhang 0006, Chenyou Fan |
ICASSP | 6 |
| 2025 | Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsabstractRecently, learning-based stereo matching networks have advanced significantly. However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets. Leveraging Vision Foundation Models (VFMs) can intuitively enhance the model's robustness, but integrating such a model into stereo matching cost-effectively to fully realize their robustness remains a key challenge. To address this, we propose SMoEStereo, a novel framework that adapts VFMs for stereo matching through a tailored, scene-specific fusion of Low-Rank Adaptation (LoRA) and Mixture-of-Experts (MoE) modules. SMoEStereo introduces MoE-LoRA with adaptive ranks and MoE-Adapter with adaptive kernel sizes. The former dynamically selects optimal experts within MoE to adapt varying scenes across domains, while the latter injects inductive bias into frozen VFMs to improve geometric feature extraction. Importantly, to mitigate computational overhead, we further propose a lightweight decision network that selectively activates MoE modules based on input complexity, balancing efficiency with accuracy. Extensive experiments demonstrate that our method exhibits state-of-the-art cross-domain and joint generalization across multiple benchmarks without dataset-specific adaptation. The code is available at \textcolor{red}{https://github.com/cocowy1/SMoE-Stereo}. Yun Wang 0053, Longguang Wang, Chenghao Zhang 0003, Zhanjie Zhang, Ao Ma 0005, Chenyou Fan, Tin Lun Lam, Junjie Hu 0003 |
ICCV | 7 |
| 2025 | Task-Agnostic Pre-training and Task-Guided Fine-tuning for Versatile Diffusion PlannerabstractDiffusion models have demonstrated their capabilities in modeling trajectories of multi-tasks. However, existing multi-task planners or policies typically rely on task-specific demonstrations via multi-task imitation, or require task-specific reward labels to facilitate policy optimization via Reinforcement Learning (RL). They are costly due to the substantial human efforts required to collect expert data or design reward functions. To address these challenges, we aim to develop a versatile diffusion planner capable of leveraging large-scale inferior data that contains task-agnostic sub-optimal trajectories, with the ability to fast adapt to specific tasks. In this paper, we propose SODP, a two-stage framework that leverages Sub-Optimal data to learn a Diffusion Planner, which is generalizable for various downstream tasks. Specifically, in the pre-training stage, we train a foundation diffusion planner that extracts general planning capabilities by modeling the versatile distribution of multi-task trajectories, which can be sub-optimal and has wide data coverage. Then for downstream tasks, we adopt RL-based fine-tuning with task-specific rewards to quickly refine the diffusion planner, which aims to generate action sequences with higher task-specific returns. Experimental results from multi-task domains including Meta-World and Adroit demonstrate that SODP outperforms state-of-the-art methods with only a small amount of data for reward-guided fine-tuning. Chenyou Fan, Chenjia Bai, Zhao Shan, Haoran He, Yang Zhang 0072, Zhen Wang 0004 |
ICML | 1 |
| 2025 | Heterogeneous Federated Learning with Scalable Server Mixture-of-ExpertsabstractClassical Federated Learning (FL) encounters significant challenges when deploying large models on power-constrained clients. To tackle this, we propose an asymmetric FL mechanism that enables the aggregation of compact client models into a comprehensive server model. We design the server model as a Mixture-of-Experts (MoE), where each expert has the same architecture as each client model. This uniformity allows for efficient fusion of the most pertinent client models to update each server expert, based on the measured relevance between each client and server expert. To address the Non-IID data issue, we further optimize the server-side MoE architecture by incorporating a main expert that always activates alongside a set of selectively activated routed experts. This configuration ensures a balance between learning general knowledge and specific data distribution. Our Fed-MoE framework is model-agnostic and has demonstrated notable improvements on vision FL tasks with million-scale ResNet backbones, and language tasks with billion-scale BERT and GPT-2 backbones. Jingang Jiang 0003, Yanzhao Chen, Haiqi Jiang 0003, Chenyou Fan |
IJCAI | 5 |
| 2025 | MiDream: Multi-Domain Retrieval Enhancement with Mixture-of-ExpertsabstractWe study the multi-domain retrieval task, where user queries require retrieving relevant documents spanning multiple domains such as medicine, economics, and films. An underlying challenge is domain confusion, in which a query pertains to any of the K domains, complicating the routing to the appropriate domain expert and hindering the generation of precise domain-specific embeddings. One straightforward approach is to train a unified expert that embeds all categories into a shared space. However, this practice suffers from low retrieval precision, as embeddings from different domains may overlap, resulting in ambiguous retrieval due to insufficient domain-specific knowledge. To address this issue, we propose MiDream, MultI-Domain Retrieval EnhAnceMent, by leveraging the strengths of Mixture-of-Experts (MoE) techniques. MiDream mitigates the domain-confusion through a Domain Routing Network (DRN) that learns to route queries to a relevant domain-specific expert dynamically. This routing enables optimization of each domain-specific expert, establishing more discriminative embedding space. Each domain-specific expert is itself a MoE structure, consisting of a Sub-Domain Routing Network (SDRN) and two lightweight experts with Gated Linear layers, ensuring rapid adaptation across domains with efficient training and swift inference. We validate the effectiveness of MiDream on multi-domain datasets, demonstrating that our approach outperforms existing retrieval methods with the equivalent parameter counts, achieving an 8.7%~14.6% improvement in NDCG and a 10.4%~14.9% improvement in MRR. Kehui Tan, Wenjing Pang, Jiayang Yao, Tianqi Pang, Chenyou Fan |
IJCNN | 6 |
| 2025 | CMA: A Unified Contextual Meta-Adaptation Methodology for Time-Series Denoising and Prediction
Haiqi Jiang 0003, Ying Ding 0007, Chenjie Pan, Aimin Huang, Chenyou Fan |
KDD (2) | 6 |
| 2025 | StoryCrafter: Instance-Aligned Multi-Character Storytelling with Diffusion Policy LearningabstractOpen-ended visual storytelling presents a formidable challenge for current text-to-image models, which frequently struggle to preserve both narrative coherence and consistent character depictions across generated sequences. To address this, we introduce StoryCrafter, a multi-character diffusion model that leverages a novel instance-level cross-attention module with supervised fine-tuning to ensure precise text-character alignment and consistent multi-character interactions throughout the narrative. Further, we propose Direct-Diffusion Group Relative Policy Optimization (D2GRPO), a novel RLHF stage that optimizes denoising strategies using automated story-aligned rewards, selecting the best candidate frames from a generated group. We evaluate our approach through human assessments and vision language model (VLM) scoring, measuring text-to-image alignment, style and character consistency, and fine-grained detail quality. Experiments on three benchmarks demonstrate that StoryCrafter outperforms existing methods, achieving 7% improvements in storytelling consistency and 10% in character accuracy, while outperforming baselines in both human and VLM evaluations. Ruiqi Dong, Wenjing Pang, Chenjie Pan, Hengyang Lu, Chenyou Fan |
ACM Multimedia | 5 |
| 2025 | Enhancing Human Trajectory Prediction with Reinforcement Learning from Quantified Human Preferences
Chenyou Fan, Kehui Tan, Yanzhao Chen, Tianqi Pang, Haiqi Jiang 0003, Junjie Hu 0003 |
PRCV (7) | 1 |
| 2025 | Knowledge-guided prompt-based continual learning: Aligning task-prompts through contrastive hard negatives
Hengyang Lu, Long-kang Lin, Chenyou Fan, Chong-Jun Wang, Wei Fang 0001, Xiaojun Wu 0001 |
Knowl. Based Syst. | 3 |
| 2025 | MuSIA: Exploiting multi-source information fusion with abnormal activations for out-of-distribution detection
Hengyang Lu, Chenyou Fan, Yuntao Du 0001, Zhenhao Shao, Wei Fang 0001, Xiaojun Wu 0001 |
Neural Networks | 4 |
| 2025 | Robust Depth Estimation Under Sensor Degradations: A Multi-Sensor Fusion PerspectiveabstractThe significance of depth estimation has spurred recent endeavors to enhance it through Multi-Sensor Fusion (MSF). However, prevailing MSF methods exhibit limitations concerning accuracy and resilience when confronted with sensor degradations. While certain forms of degradation, such as suboptimal lighting and adverse weather conditions, can be mitigated by collecting pertinent data in data-driven learning, this approach proves ineffective for Out-of-Distribution (OOD) sensor degradations. In this paper, we propose a novel approach termed Combinable and Separable Multi-Sensor Fusion (CSMSF) designed to bolster depth estimation robustness against multiple sensor degradations. CSMSF hinges on four core principles: i) improved performance is achieved with an increased number of valid sensors, ii) a single valid sensor can independently enable its own depth estimation, iii) maintaining a judicious equilibrium between accuracy and model complexity, and iv) autonomous diagnosis of sensor observation failure. Leveraging these advantages, CSMSF identifies and rejects degraded sensors, allowing autonomous selection of valid sensors for scene depth estimation. The experimental results demonstrate the superior robustness of the proposed CSMSF, underscoring its efficacy in addressing challenges associated with sensor degradations across diverse environmental conditions. Junjie Hu 0003, Chenyou Fan, Mete Ozay, Qing Gao 0002, Yulan Guo, Tin Lun Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Unlocking Drone Perception in Low AGL Heights: Progressive Semi-Supervised Learning for Ground-to-Aerial Perception Knowledge TransferabstractWe explore the novel challenge of drone perception across varying low AGL (above ground level) heights, a task essential for dynamic tasks, unlike the fixed ground viewpoint in autonomous driving. Supervised learning for this incurs high annotation costs, and current semi-supervised methods struggle with viewpoint differences. In this paper, we introduce ground-to-aerial perception knowledge transfer and propose a progressive semi-supervised learning framework for drone perception using only labeled data from the ground viewpoint and unlabeled data from flying viewpoints. The framework hinges on four key components: 1) a dense viewpoint sampling strategy, segmenting the vertical flight height range into evenly distributed intervals; 2) nearest neighbor pseudo-labeling, inferring labels of the nearest neighbor viewpoint using a model learned on the preceding viewpoint; 3) MixView, generating augmented images among different viewpoints to mitigate viewpoint differences; and 4) a progressive distillation strategy, gradually learning until reaching the maximum flying height. To validate our approach, we create both synthesized and real-world datasets. Extensive experimental analyses reveal a remarkable relative accuracy improvement of 25.7% and 16.9% for the synthesized dataset and the real world, respectively. Code and datasets are available on https://github.com/FreeformRobotics/Progressive-Self-Distillation-for-Ground-to-Aerial-Perception-Knowledge-Transfer. Junjie Hu 0003, Chenyou Fan, Mete Ozay, Yuan Gao 0024, Tin Lun Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Lifelong-MonoDepth: Lifelong Learning for Multidomain Monocular Metric Depth EstimationabstractWith the rapid advancements in autonomous driving and robot navigation, there is a growing demand for lifelong learning (LL) models capable of estimating metric (absolute) depth. LL approaches potentially offer significant cost savings in terms of model training, data storage, and collection. However, the quality of RGB images and depth maps is sensor-dependent, and depth maps in the real world exhibit domain-specific characteristics, leading to variations in depth ranges. These challenges limit existing methods to LL scenarios with small domain gaps and relative depth map estimation. To facilitate lifelong metric depth learning, we identify three crucial technical challenges that require attention: 1) developing a model capable of addressing the depth scale variation through scale-aware depth learning; 2) devising an effective learning strategy to handle significant domain gaps; and 3) creating an automated solution for domain-aware depth inference in practical applications. Based on the aforementioned considerations, in this article, we present 1) a lightweight multihead framework that effectively tackles the depth scale imbalance; 2) an uncertainty-aware LL solution that adeptly handles significant domain gaps; and 3) an online domain-specific predictor selection method for real-time inference. Through extensive numerical studies, we show that the proposed method can achieve good efficiency, stability, and plasticity, leading the benchmarks by 8%-15%. The code is available at https://github.com/FreeformRobotics/Lifelong-MonoDepth. Junjie Hu 0003, Chenyou Fan, Liguang Zhou, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Machine Learning Model Trading With Verification Under Information AsymmetryabstractMachine learning (ML) model trading, known for its role in protecting data privacy, faces a major challenge: information asymmetry. This issue can lead to model deception, a problem that current literature has not fully solved, where the seller misrepresents model performance to earn more. We propose a game-theoretic approach, adding a verification step in the ML model market that lets buyers check model quality before buying. However, this method can be expensive and offers imperfect information, making it harder for buyers to decide. Our analysis reveals that a seller might probabilistically conduct model deception considering the chance of model verification. This deception probability decreases with the verification accuracy and increases with the verification cost. To maximize seller payoff, we further design optimal pricing schemes accounting for heterogeneous buyers’ strategic behaviors. Interestingly, we find that reducing information asymmetry benefits both the seller and buyer. Meanwhile, protecting buyer order information doesn’t improve the payoff for the buyer or the seller. These findings highlight the importance of reducing information asymmetry in ML model trading and open new directions for future research. Xiang Li 0148, Jianwei Huang 0001, Kai Yang 0001, Chenyou Fan |
IEEE Trans. Netw. | 4 |
| 2024 | Low-Parameter Federated Learning with Large Language Models
Jingang Jiang 0003, Haiqi Jiang 0003, Chenyou Fan |
WISA | 5 |
| 2024 | Medical Document Embedding Enhancement with Heterogeneous Mixture-of-ExpertsabstractRetrieval-Augmented Generation (RAG) has emerged as a crucial technique to enhance the accuracy and reliability of large language models, particularly in specialized domains like medicine. However, the effectiveness of RAG heavily depends on the quality of text embeddings used for retrieval. In this paper, we introduce Med-MoE-Embed, a novel approach to improve medical text embeddings tasks. Med-MoE-Embed leverages a pretrained embedding backbone augmented with a trainable Mixture of Experts (MoE) network, allowing for efficient adaptation to specific medical subdomains and tasks. We design each expert to be a compact KANs or a MLP with heterogeneous activation functions such as GELU and SWIGLU. Furthermore, we propose a two-step fine-tuning process that optimizes expert training and selection, enhancing the model’s adaptability across various medical datasets. Our extensive evaluation focuses on RAG tasks in the medical domain, demonstrating significant improvements in retrieval accuracy and generation quality. Med-MoE-Embed mitigates the challenges of limited data accessibility and domain-specific requirements in the medical field, offering a versatile and efficient solution for enhancing embedding quality in medical natural language processing applications. Tianqi Pang, Kui Xue, Xiaofan Zhang 0006, Chenyou Fan |
BIBM | 6 |
| 2024 | Carbon Price Forecasting with LLM-Based Refinement and Transfer-Learning
Haiqi Jiang 0003, Ying Ding 0007, Chenyou Fan |
ICANN (9) | 4 |
| 2024 | REMED: Retrieval-Augmented Medical Document Query Responding with Embedding Fine-TuningabstractWhile advanced Large Language Models (LLMs) exhibit considerable promise, their tendency to generate unreliable information poses significant challenges, particularly in high-risk domains like healthcare. However, the advent of Retrieval-Augmented Generation (RAG) offers a novel solution tailored for the medical realm. This study further enhances retrieval accuracy by introducing REMED, a specialized medical document retrieval framework designed to address the hallucination problem prevalent in LLMs. The REMED framework integrates dataset construction, an efficient embedding fine-tuning EM-FT model, retrieval-augmented generation, and human evaluation of LLM responses. The EM-FT model can end-to-end fine-tune the medical sentence representations in large pre-trained models through an efficient embedding fine-tuning method, thereby enhancing the performance of medical retrieval. We adopt contrastive learning as the loss function to optimize the performance of the EM-FT model, enabling it to accurately capture the similarity between query and relevant documents. This approach not only improves the retrieval accuracy of positively related contents but also effectively reduces the matching with negatively related contents. Compared to direct dense vector retrieval, fine-tuning query and content vectors first and then performing dense retrieval tasks significantly improved the performance. Through validation on two datasets, we demonstrate that our EM-FT method improves recall and precision on MMD by 3.2%-6.0% and on MPD by 14.4%-42.6% compared to using the embedding model directly for retrieval. Furthermore, through human evaluation on the PULSE-7Bv5 model, we further confirm the effectiveness of our retrieval results in improving the quality of generated text. Tianqi Pang, Kehui Tan, Yujun Yao, Chenyou Fan, Xiaofan Zhang 0006 |
IJCNN | 6 |
| 2024 | TeMME: Temporal Knowledge Graph Completion Using Multi-grade Multivector Embeddings
Hengyang Lu, Hao-Kun Yu, Chenyou Fan, Qianyi Zhan, Wei Fang 0001, Xiaojun Wu 0001 |
PRICAI (4) | 3 |
| 2024 | Dense depth distillation with out-of-distribution simulated images
Junjie Hu 0003, Chenyou Fan, Mete Ozay, Hualie Jiang, Tin Lun Lam |
Knowl. Based Syst. | 2 |
| 2023 | RL-Based CEP Operator Placement Method on Edge Networks Using Response Time Feedback
Yuyou Wang, Hao Hu 0001, Hongyu Kuang, Chenyou Fan, Liang Wang 0006, XianPing Tao |
WISA | 4 |
| 2023 | A Novel Differentiable Rank Learning Method Towards Stock Movement Quantile ForecastingabstractWe focus on Stock Movement Forecasting (SMF) using AI techniques to develop modern automated trading systems. Previous studies with deep-learning-based methodology have only considered binary up-or-down trends, ignoring the importance of fine-grained categorization of the stock movements to facilitate decision-making. However, the challenges of SMF arise from the randomness of the global market impacting cross-sectional stocks and the volatility of internal dynamics in each time series. To address these challenges, we present a novel end-to-end learning-to-rank framework that incorporates both market-level and stock-level dynamics. Specifically, we aim to identify cross-sectional stocks that exhibit notable movements at every time step and learn to rank steps with the most significant movements in the temporal dimension. We conduct extensive evaluations of our multi-task learning framework utilizing real-world market data, which demonstrate superior performance when compared to state-of-the-art methods, with improvements in the Gain and Sharpe Ratio by 5–15%. Chenyou Fan, Hengyang Lu, Aimin Huang |
ECAI | 1 |
| 2023 | Machine Learning Model Trading with Information AsymmetryabstractMachine learning (ML) model trading prevents data breaches in privacy-sensitive data-driven applications. Departing from commonly assumed complete information scenarios, we consider the more practical trading scenario where model deception may emerge under information asymmetry. More specifically, the model seller may provide false information on model quality to maximize her payoff. This paper takes the first step in tackling information asymmetry through the lens of model verification. We propose an ML model market that allows buyers to verify model quality before purchasing. Such verification can be costly and often imperfect, which makes the buyer's decision highly nontrivial. We first formulate the ML model trading process as a three-stage sequential game with imperfect information, where the seller determines the model delivery strategy after observing the buyer's order decision. Our analysis reveals that at the equilibrium, the seller will probabilistically conduct model deception, considering the possibility of model verification. The equilibrium deception probability increases with the buyer's verification cost and decreases with verification accuracy. Interestingly, we also show that reducing information asymmetry through verification benefits both the buyer and seller. We further consider a second market model with buyer order information protection, where the buyer's order information is unobservable before the seller makes the delivery strategy. Our analysis shows a surprising result under this market model: protecting buyer's order information will not increase the payoff of either the buyer or seller. Xiang Li 0148, Jianwei Huang 0001, Kai Yang 0001, Chenyou Fan |
ICC | 4 |
| 2023 | Trajectory Prediction with Contrastive Pre-training and Social Rank Fine-Tuning
Chenyou Fan, Haiqi Jiang 0003, Aimin Huang, Junjie Hu 0003 |
ICONIP (10) | 1 |
| 2023 | Pre-trained Financial Model for Price Movement Forecasting
Chenyou Fan, Tianqi Pang, Aimin Huang |
ICONIP (15) | 1 |
| 2023 | Boosting LightWeight Depth Estimation via Knowledge Distillation
Junjie Hu 0003, Chenyou Fan, Hualie Jiang, Xiyue Guo, Yuan Gao 0024, Xiangyong Lu, Tin Lun Lam |
KSEM (1) | 2 |
| 2023 | Federated Prompting and Chain-of-Thought Reasoning for Improving LLMs Answering
Tianqi Pang, Chenyou Fan |
KSEM (4) | 3 |
| 2023 | Emergent leader-follower relationship in networked multiagent systems
Jinzhuo Liu, Chenyou Fan, Yunchen Peng, Jinze Du, Zhen Wang 0004, Chen Chu |
Sci. China Inf. Sci. | 2 |
| 2023 | Few-Shot Multi-Agent Perception With Ranking-Based Feature LearningabstractIn this article, we focus on performing few-shot learning (FSL) under multi-agent scenarios in which participating agents only have scarce labeled data and need to collaborate to predict labels of query observations. We aim at designing a coordination and learning framework in which multiple agents, such as drones and robots, can collectively perceive the environment accurately and efficiently under limited communication and computation conditions. We propose a metric-based multi-agent FSL framework which has three main components: an efficient communication mechanism that propagates compact and fine-grained query feature maps from query agents to support agents; an asymmetric attention mechanism that computes region-level attention weights between query and support feature maps; and a metric-learning module which calculates the image-level relevance between query and support data fast and accurately. Furthermore, we propose a specially designed ranking-based feature learning module, which can fully utilize the order information of training data by maximizing the inter-class distance, while minimizing the intra-class distance explicitly. We perform extensive numerical studies and demonstrate that our approach can achieve significantly improved accuracy in visual and acoustic perception tasks such as face identification, semantic segmentation, and sound genre recognition, consistently outperforming the state-of-the-art baselines by 5%-20%. Chenyou Fan, Junjie Hu 0003, Jianwei Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Deep Depth Completion From Extremely Sparse Data: A SurveyabstractDepth completion aims at predicting dense pixel-wise depth from an extremely sparse map captured from a depth sensor, e.g., LiDARs. It plays an essential role in various applications such as autonomous driving, 3D reconstruction, augmented reality, and robot navigation. Recent successes on the task have been demonstrated and dominated by deep learning based solutions. In this article, for the first time, we provide a comprehensive literature review that helps readers better grasp the research trends and clearly understand the current advances. We investigate the related studies from the design aspects of network architectures, loss functions, benchmark datasets, and learning strategies with a proposal of a novel taxonomy that categorizes existing methods. Besides, we present a quantitative comparison of model performance on three widely used benchmarks, including indoor and outdoor datasets. Finally, we discuss the challenges of prior works and provide readers with some insights for future research directions. Junjie Hu 0003, Chenyu Bao, Mete Ozay, Chenyou Fan, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Where to Attack: A Dynamic Locator Model for Backdoor Attack in Text ClassificationsabstractNowadays, deep-learning based NLP models are usually trained with large-scale third-party data which can be easily injected with malicious backdoors. Thus, BackDoor Attack (BDA) study has become a trending research to help promote the robustness of an NLP system. Text-based BDA aims to train a poisoned model with both clean and poisoned texts to perform normally on clean inputs while being misled to predict those trigger-embedded texts as target labels set by attackers. Previous works usually choose fixed Positions-to-Poison (P2P) first, then add triggers upon those positions such as letter insertion or deletion. However, considering the positions of words with important semantics may vary in different contexts, fixed P2P models are severely limited in flexibility and performance. We study the text-based BDA from the perspective of automatically and dynamically selecting P2P from contexts. We design a novel Locator model which can predict P2P dynamically without human intervention. Based on the predicted P2P, four effective strategies are introduced to show the BDA performance. Experiments on two public datasets show both tinier test accuracy gap on clean data and higher attack success rate on poisoned ones. Human evaluation with volunteers also shows the P2P predicted by our model are important for classification. Source code is available at https://github.com/jncsnlp/LocatorModel Hengyang Lu, Chenyou Fan, Jun Yang 0038, Wei Fang 0001, Xiaojun Wu 0001 |
COLING | 2 |
| 2022 | Private Semi-Supervised Federated LearningabstractWe study a federated learning (FL) framework to effectively train models from scarce and skewly distributed labeled data. We consider a challenging yet practical scenario: a few data sources own a small amount of labeled data, while the rest mass sources own purely unlabeled data. Classical FL requires each client to have enough labeled data for local training, thus is not applicable in this scenario. In this work, we design an effective federated semi-supervised learning framework (FedSSL) to fully leverage both labeled and unlabeled data sources. We establish a unified data space across all participating agents, so that each agent can generate mixed data samples to boost semi-supervised learning (SSL), while keeping data locality. We further show that FedSSL can integrate differential privacy protection techniques to prevent labeled data leakage at the cost of minimum performance degradation. On SSL tasks with as small as 0.17% and 1% of MNIST and CIFAR-10 datasets as labeled data, respectively, our approach can achieve 5-20% performance boost over the state-of-the-art methods. Chenyou Fan, Junjie Hu 0003, Jianwei Huang 0001 |
IJCAI | 1 |
| 2021 | Few-Shot Multi-Agent PerceptionabstractWe study few-shot learning (FSL) under multi-agent scenarios, in which participating agents only have local scarce labeled data and need to collaborate to predict query data labels. Though each of the agents, such as drones and robots, has minimal communication and computation capability, we aim at designing coordination schemes such that they can collectively perceive the environment accurately and efficiently. We propose a novel metric-based multi-agent FSL framework which has three main components: an efficient communication mechanism that propagates compact and fine-grained query feature maps from query agents to support agents; an asymmetric attention mechanism that computes region-level attention weights between query and support feature maps; and a metric-learning module which calculates the image-level relevance between query and support data fast and accurately. Through analysis and extensive numerical studies, we demonstrate that our approach can save communication and computation costs and significantly improve performance in both visual and acoustic perception tasks such as face identification, semantic segmentation, and sound genre recognition. Chenyou Fan, Junjie Hu 0003, Jianwei Huang 0001 |
ACM Multimedia | 1 |
| 2021 | Federated Few-Shot Learning with Adversarial LearningabstractWe are interested in developing a unified machine learning framework for effectively training machine learning models from many small data sources such as mobile devices. This is a commonly encountered situation in mobile computing scenarios, where data is scarce and distributed while the tasks are distinct. In this paper, we propose a federated few-shot learning (FedFSL) framework to learn a few-shot classification model that can classify unseen data classes with only a few labeled samples. With the federated learning strategy, FedFSL can utilize many data sources while keeping data privacy and communication efficiency. To tackle the issue of obtaining misaligned decision boundaries produced by client models, we propose to regularize local updates by minimizing the divergence of client models. We also formulate the training in an adversarial fashion and optimize the client models to produce a discriminative feature space that can better represent unseen data samples. We demonstrate the intuitions and conduct experiments to show our approaches outperform baselines by more than 10% in learning benchmark vision tasks and 5% in language tasks. Chenyou Fan, Jianwei Huang 0001 |
WiOpt | 1 |
| 2020 | Projection Robust Wasserstein Distance and Riemannian OptimizationabstractProjection robust Wasserstein (PRW) distance, or Wasserstein projection pursuit (WPP), is a robust variant of the Wasserstein distance. Recent work suggests that this quantity is more robust than the standard Wasserstein distance, in particular when comparing probability measures in high-dimensions. However, it is ruled out for practical application because the optimization model is essentially non-convex and non-smooth which makes the computation intractable. Our contribution in this paper is to revisit the original motivation behind WPP/PRW, but take the hard route of showing that, despite its non-convexity and lack of nonsmoothness, and even despite some hardness results proved by~\citet{Niles-2019-Estimation} in a minimax sense, the original formulation for PRW/WPP \textit{can} be efficiently computed in practice using Riemannian optimization, yielding in relevant cases better behavior than its convex relaxation. More specifically, we provide three simple algorithms with solid theoretical guarantee on their complexity bound (one in the appendix), and demonstrate their effectiveness and efficiency by conducing extensive experiments on synthetic and real data. This paper provides a first step into a computational theory of the PRW distance and provides the links between optimal transport and Riemannian optimization. Tianyi Lin, Chenyou Fan, Nhat Ho, Marco Cuturi, Michael I. Jordan |
NeurIPS | 2 |
| 2020 | Federated Generative Adversarial Learning
Chenyou Fan |
PRCV (3) | 1 |
| 2019 | Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question AnsweringabstractIn this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion features; 2) a redesigned question memory which helps understand the complex semantics of question and highlights queried subjects; and 3) a new multimodal fusion layer which performs multi-step reasoning by attending to relevant visual and textual hints with self-updated attention. Our VideoQA model firstly generates the global context-aware visual and textual features respectively by interacting current inputs with memory contents. After that, it makes the attentional fusion of the multimodal visual and textual representations to infer the correct answer. Multiple cycles of reasoning can be made to iteratively refine attention weights of the multimodal data and improve the final representation of the QA pair. Experimental results demonstrate our approach achieves state-of-the-art performance on four VideoQA benchmark datasets. Chenyou Fan, Xiaofan Zhang 0006, Shu Zhang 0001, Chi Zhang 0012, Heng Huang 0001 |
CVPR | 1 |
| 2019 | Multi-Horizon Time Series Forecasting with Temporal Attention LearningabstractWe propose a novel data-driven approach for solving multi-horizon probabilistic forecasting tasks that predicts the full distribution of a time series on future horizons. We illustrate that temporal patterns hidden in historical information play an important role in accurate forecasting of long time series. Traditional methods rely on setting up temporal dependencies manually to explore related patterns in historical data, which is unrealistic in forecasting long-term series on real-world data. Instead, we propose to explicitly learn constructing hidden patterns' representations with deep neural networks and attending to different parts of the history for forecasting the future. Chenyou Fan, Chi Zhang 0012, Rong Yuan, Jian Pei 0001, Heng Huang 0001 |
KDD | 1 |
| 2018 | Joint Person Segmentation and Identification in Synchronized First- and Third-Person Videos
Chenyou Fan, Michael S. Ryoo, David Crandall |
ECCV (1) | 2 |
| 2018 | Multi-task Spatiotemporal Neural Networks for Structured Surface ReconstructionabstractDeep learning methods have surpassed the performance of traditional techniques on a wide range of problems in computer vision, but nearly all of this work has studied consumer photos, where precisely correct output is often not critical. It is less clear how well these techniques may apply on structured prediction problems where fine-grained output with high precision is required, such as in scientific imaging domains. Here we consider the problem of segmenting echogram radar data collected from the polar ice sheets, which is challenging because segmentation boundaries are often very weak and there is a high degree of noise. We propose a multi-task spatiotemporal neural network that combines 3D ConvNets and Recurrent Neural Networks (RNNs) to estimate ice surface boundaries from sequences of tomographic radar images. We show that our model outperforms the state-of-the-art on this problem by (1) avoiding the need for hand-tuned parameters, (2) extracting multiple surfaces (ice-air and ice-bed) simultaneously, (3) requiring less non-visual metadata, and (4) being about 6 times faster. Chenyou Fan, John Paden, Geoffrey C. Fox, David Crandall |
WACV | 2 |
| 2018 | Deepdiary: Lifelogging image captioning and summarization
Chenyou Fan, David Crandall |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Title Learning Latent Subevents in Activity Videos Using Temporal Attention FiltersabstractIn this paper, we newly introduce the concept of temporal attention filters, and describe how they can be used for human activity recognition from videos. Many high-level activities are often composed of multiple temporal parts (e.g., sub-events) with different duration/speed, and our objective is to make the model explicitly learn such temporal structure using multiple attention filters and benefit from them. Our temporal filters are designed to be fully differentiable, allowing end-of-end training of the temporal filters together with the underlying frame-based or segment-based convolutional neural network architectures. This paper presents an approach of learning a set of optimal static temporal attention filters to be shared across different videos, and extends this approach to dynamically adjust attention filters per testing video using recurrent long short-term memory networks (LSTMs). This allows our temporal attention filters to learn latent sub-events specific to each activity. We experimentally confirm that the proposed concept of temporal attention filters benefits the activity recognition, and we visualize the learned latent sub-events. A. J. Piergiovanni, Chenyou Fan, Michael S. Ryoo |
AAAI | 2 |
| 2017 | Identifying First-Person Camera Wearers in Third-Person VideosabstractWe consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establish person-level correspondences across first-and third-person videos, which is challenging because the camera wearer is not visible from his/her own egocentric video, preventing the use of direct feature matching. In this paper, we propose a new semi-Siamese Convolutional Neural Network architecture to address this novel challenge. We formulate the problem as learning a joint embedding space for first-and third-person videos that considers both spatial-and motion-domain cues. A new triplet loss function is designed to minimize the distance between correct first-and third-person matches while maximizing the distance between incorrect ones. This end-to-end approach performs significantly better than several baselines, in part by learning the first-and third-person features optimized for matching jointly with the distance measure itself. Chenyou Fan, Jangwon Lee 0002, Krishna Kumar Singh, Yong Jae Lee, David Crandall, Michael S. Ryoo |
CVPR | 1 |