Yanchen Wang

dblp:16/6169 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 MMA++: Effective Multi-Modal Adaptation for Vision-Language Models
abstract
Large scale pre-trained Vision-Language Models (VLMs) have shown good generalization capabilities across diverse downstream tasks. However, adapting such large-scale models to few-shot generalization scenarios remains challenging due to the trade-off between preserving general knowledge and incorporating task-specific information. In this paper, we propose MMA++, an advanced and effective Multi-Modal Adapter framework for parameter-efficient VLM adaptation. Unlike prior works that independently inject adapters into each modality or uniformly across layers, MMA++ performs a dataset-level analysis to identify discriminative and generalizable features, and selectively applies adapters to the higher layers of both vision and text encoders. To bridge the modality gap, we further propose a shared feature projection space that enhances alignment between modalities. Beyond architecture design, we identify the fusion scale $\alpha$α-which controls the strength of adapter integration-as a key factor in few-shot generalization. We empirically and theoretically demonstrate that $\alpha$α should not be static, but adapted based on training data size. To reduce the effort of tuning this value across different datasets, we propose the $\alpha$α-consistency framework, consisting of: (1) a consistency training strategy under varying fusion scales; and (2) an $\alpha$α-decoupling strategy that uses a larger fusion scale during training and a smaller one at inference to account for sample size mismatch. We evaluate MMA++ on a wide range of few-shot generalization tasks, including base-to-novel generalization, cross-dataset transfer, and domain generalization. Our method consistently achieves leading performance.
Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua Xie
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Neural Encoding and Decoding at Scale
abstract
Recent work has demonstrated that large-scale, multi-animal models are powerful tools for characterizing the relationship between neural activity and behavior. Current large-scale approaches, however, focus exclusively on either predicting neural activity from behavior (encoding) or predicting behavior from neural activity (decoding), limiting their ability to capture the bidirectional relationship between neural activity and behavior. To bridge this gap, we introduce a multimodal, multi-task model that enables simultaneous Neural Encoding and Decoding at Scale (NEDS). Central to our approach is a novel multi-task-masking strategy, which alternates between neural, behavioral, within-modality, and cross-modality masking. We pretrain our method on the International Brain Laboratory (IBL) repeated site dataset, which includes recordings from 83 animals performing the visual decision-making task. In comparison to other large-scale modeling approaches, we demonstrate that NEDS achieves state-of-the-art performance for both encoding and decoding when pretrained on multi-animal data and then fine-tuned on new animals. Surprisingly, NEDS’s learned embeddings exhibit emergent properties: even without explicit training, they are highly predictive of the brain regions in each recording. Altogether, our approach is a step towards a foundation model of the brain that enables seamless translation between neural activity and behavior.
Yizi Zhang, Yanchen Wang, Mehdi Azabou, Alexandre Andre, Hanrui Lyu, Eva L. Dyer, Liam Paninski, Cole L. Hurwitz
ICML2
2025 Not All Impressions Are Created Equal: Psychology-Informed Retention Optimization for Short-Form Video Recommendation
Zhaohui Guo, Chuanqi Wei, Yanchen Wang, Zellux Wang
RecSys6
2024 MMA: Multi-Modal Adapter for Vision-Language Models
abstract
Pretrained Vision-Language Models (VLMs) have served as excellent foundation models for transfer learning in diverse downstream tasks. However, tuning VLMs for few-shot generalization tasks faces a discrimination - generalization dilemma, i.e., general knowledge should be preserved and task-specific knowledge should be fine-tuned. How to precisely identify these two types of representations remains a challenge. In this paper, we propose a Multi-Modal Adapter (MMA) for VLMs to improve the alignment between representations from text and vision branches. MMA aggregates features from different branches into a shared feature space so that gradients can be communicated across branches. To determine how to incorporate MMA, we systematically analyze the discriminability and generalizability of features across diverse datasets in both the vision and language branches, and find that (1) higher lay-ers contain discriminable dataset-specific knowledge, while lower layers contain more generalizable knowledge, and (2) language features are more discriminable than visual features, and there are large semantic gaps between the features of the two modalities, especially in the lower layers. Therefore, we only incorporate MMA to a few higher lay-ers of transformers to achieve an optimal balance between discrimination and generalization. We evaluate the effectiveness of our approach on three tasks: generalization to novel classes, novel target datasets, and domain generalization. Compared to many state-of-the-art methods, our MMA achieves leading performance in all evaluations. Code is at https://github.com/ZjjConan/Multi-Modal-Adapter
Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua Xie
CVPR3
2024 It is Time to Develop an Auditing Framework to Promote Value Aware Chatbots
abstract
The launch of ChatGPT in November 2022 marked the beginning of a new era in AI, the availability of generative AI tools for everyone to use. ChatGPT and other similar chatbots boast a wide range of capabilities from answering student homework questions to creating music and art. Given the large amounts of human data chatbots are built on, it is inevitable that they will inherit human errors and biases. These biases have the potential to inflict significant harm or increase inequity on different subpopulations. Because chatbots do not have an inherent understanding of societal values, they may create new content that is contrary to established norms. Examples of concerning generated content includes child pornography, inaccurate facts, and discriminatory posts. In this position paper, we argue that the speed of advancement of this technology requires us, as computer and data scientists, to mobilize and develop a values-based auditing framework containing a community established standard set of measurements to monitor the health of different chatbots and LLMs. To support our argument, we use a simple audit template to share the results of basic audits we conduct that are focused on measuring potential bias in search engine style tasks, code generation, and story generation. We identify responses from GPT 3.5 and GPT 4 that are both consistent and not consistent with values derived from existing law. While the findings come as no surprise, they do underscore the urgency of developing a robust auditing framework for openly sharing results in a consistent way so that mitigation strategies can be developed by the academic community, government agencies, and companies when our values are not being adhered to. We conclude this paper with recommendations for value-based strategies for improving the technologies.
Yanchen Wang, Lisa Singh
DATA1
2024 Towards a "Universal Translator" for Neural Dynamics at Single-Cell, Single-Spike Resolution
abstract
Neuroscience research has made immense progress over the last decade, but our understanding of the brain remains fragmented and piecemeal: the dream of probing an arbitrary brain region and automatically reading out the information encoded in its neural activity remains out of reach. In this work, we build towards a first foundation model for neural spiking data that can solve a diverse set of tasks across multiple brain areas. We introduce a novel self-supervised modeling approach for population activity in which the model alternates between masking out and reconstructing neural activity across different time steps, neurons, and brain regions. To evaluate our approach, we design unsupervised and supervised prediction tasks using the International Brain Laboratory repeated site dataset, which is comprised of Neuropixels recordings targeting the same brain locations across 48 animals and experimental sessions. The prediction tasks include single-neuron and region-level activity prediction, forward prediction, and behavior decoding. We demonstrate that our multi-task-masking (MtM) approach significantly improves the performance of current state-of-the-art population models and enables multi-task learning. We also show that by training on multiple animals, we can improve the generalization ability of the model to unseen animals, paving the way for a foundation model of the brain at single-cell, single-spike resolution.
Yizi Zhang, Yanchen Wang, Donato Jiménez-Benetó, Mehdi Azabou, Blake A. Richards, Renee Tung, Olivier Winter, Eva L. Dyer, Liam Paninski, Cole L. Hurwitz
NeurIPS2
2024 Vision-based estimation of fatigue and engagement in cognitive training sessions
Yanchen Wang, Adam Turnbull, Yunlong Xu 0001, Kathi L. Heffner, Feng Lin 0008, Ehsan Adeli-Mosabbeb
Artif. Intell. Medicine1
2023 Multi-Device Rate Maximization for IRS-Based Smart Home
abstract
Intelligent reflecting surface (IRS) consisting of a plane including many passive components can control even-t signals and create a programmable wireless environment by reconfigurable passive components. Passive components are typically composed of electronic components such as positive intrinsic-negative (PIN) diodes, field effect transistors, and micro electromechanical system switches, whose inherent characteristics show without active circuitry. Given that intelligent reflecting surfaces can strengthen the property of wireless transmission channels by adjusting the phase of incident signals, we investigate a multi-input single-output smart home model for downlink multi-users in this paper. An alternating optimization (AO) algorithm is proposed to cope with the challenge of maximizing the weighted sum of objectives of multiple users. To increase the sum rate, the beamforming vector at the access point and the phase shift matrix at the IRS are jointly tuned. In perfect channel state information, the AO approach is employed to achieve the effect of user maximization. Typically, fractional programming is utilized to optimize the beamforming vector at AP, while the Riemannian conjugate gradient approach is used to design the phase shift at IRS. Compared with the baseline methods, our proposed method can significantly improve convergence of the system.
Yanchen Wang, Yejun He
IWCMC1
2022 Students or Mechanical Turk: Who Are the More Reliable Social Media Data Labelers?
Lisa Singh, Rebecca Vanarsdall, Yanchen Wang, Carole Roan Gresenz
DATA3
2019 Blending Noisy Social Media Signals with Traditional Movement Variables to Predict Forced Migration
abstract
Worldwide displacement due to war and conflict is at all-time high. Unfortunately, determining if, when, and where people will move is a complex problem. This paper proposes integrating both publicly available organic data from social media and newspapers with more traditional indicators of forced migration to determine when and where people will move. We combine movement and organic variables with spatial and temporal variation within different Bayesian models and show the viability of our method using a case study involving displacement in Iraq. Our analysis shows that incorporating open-source generated conversation and event variables maintains or improves predictive accuracy over traditional variables alone. This work is an important step toward understanding how to leverage organic big data for societal--scale problems.
Lisa Singh, Laila Wahedi, Yanchen Wang, Yifang Wei, Christo Kirov, Susan Martin, Katharine M. Donato, Yaguang Liu, Kornraphop Kawintiranon
KDD3
2015 Innovative mobile tool for engineering embedded design and security educations
abstract
In the current electronic industry, embedded systems have become essential components in most electronic devices. The fast growth mobile devices have been used in every sector and brought wide applications in many aspects of our society. These devices involve in internet connections, wireless communications and even cloud computing, which increases the necessity of both hardware and software securities. Therefore there is an urgent need of efficient education in EE, CE and CS in embedded systems, especially in hardware and software co-design with the concepts of modern security methods. In this paper, we present a work-in-progress development of an effective learning tool which benefits from active learning concepts and modern mobile technologies. The contributions of this tool on promoting engineering students' learning of embedded system designs and security are also investigated.
Kuosheng Ma, Yingnan Ma, Yanchen Wang, Qichang Zheng
FIE3