Wen-Yen Chen

dblp:06/1416 · also WenYen Chen · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0004-9371-5642ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 1
YearPublicationVenuePosition
2026 Multi-Objective Bilevel Learning
abstract
As machine learning (ML) applications grow increasingly complex in recent years, modern ML frameworks often need to address multiple potentially conflicting objectives with coupled decision variables across different layers. This creates a compelling need for multi-objective bilevel learning (MOBL). So far, however, the field of MOBL remains in its infancy and many important problems remain under-explored. This motivates us to fill this gap and systematically investigate the theoretical and algorithmic foundation of MOBL. Specifically, we consider MOBL problems with multiple conflicting objectives guided by preferences at the upper-level subproblem, where part of the inputs depend on the optimal solution of the lower-level subproblem. Our goal is to develop efficient MOBL optimization algorithms to (1) identify a preference-guided Pareto-stationary solution with low oracle complexity; and (2) enable systematic Pareto front exploration. To this end, we propose a unifying algorithmic framework called weighted-Chebyshev multi-hyper-gradient-descent (WC-MHGD) for both deterministic and stochastic settings with finite-time Pareto-stationarity convergence rate guarantees, which not only implies low oracle complexity but also induces systematic Pareto front exploration. We further conduct extensive experiments to confirm our theoretical results.
Zhuqing Liu, Xin Zhang 0054, Wen-Yen Chen, Jiyan Yang, Jia Liu 0002
AAAI4
2026 Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence Recommendation
Xiao Lin 0016, Zhicheng Tang, Weilin Cong, Mengyue Hang, Zhichen Zeng 0001, Ting-Wei Li, Hyunsik Yoo, Zhining Liu 0002, Xuying Ning, Ruizhong Qiu, Wen-Yen Chen, Shuo Chang, Rong Jin 0001, Hanghang Tong
WWW13
2025 The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit
abstract
Huixue Zhou, Hengrui Gu, Zaifu Zhan, Xi Liu, Kaixiong Zhou, Yongkang Xiao, Mingfu Liang, Srinivas Prasad Govindan, Piyush Chawla, Jiyan Yang, Xiangfei Meng, Huayu Li, Buyun Zhang, Liang Luo, Wen-Yen Chen, Yiping Han, Bo Long, Rui Zhang, Tianlong Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Huixue Zhou, Hengrui Gu 0002, Zaifu Zhan, Xi Liu 0011, Kaixiong Zhou, Yongkang Xiao, Mingfu Liang, Srinivas Govindan, Piyush Chawla, Jiyan Yang, Xiangfei Meng, Buyun Zhang, Wen-Yen Chen, Yiping Han, Bo Long, Rui Zhang 0028, Tianlong Chen 0001
ACL (1)15
2025 InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction
Zhichen Zeng 0001, Xiaolong Liu 0012, Mengyue Hang, Qinghai Zhou, Chaofei Yang, Yichen Ruan, Laming Chen, Yuxin Chen 0001, Yujia Hao, Jade Nie, Xi Liu 0011, Buyun Zhang, Wei Wen 0003, Siyang Yuan, Hang Yin 0005, Xin Zhang 0054, Wen-Yen Chen, Yiping Han, Chunzhi Yang, Bo Long, Philip S. Yu, Hanghang Tong, Jiyan Yang
CIKM21
2024 DistDNAS: Search Efficient Feature Interactions within 2 Hours
abstract
Search efficiency and serving efficiency are two major axes in building feature interactions and expediting the model development process in recommender systems. Searching for the optimal feature interaction design on large-scale benchmarks requires extensive cost due to the sequential workflow on the large volume of data. In addition, fusing interactions of various sources, orders, and mathematical operations introduces potential conflicts and additional redundancy toward recommender models, leading to sub-optimal trade-offs in performance and serving cost. This paper presents DistDNAS as a neat solution to brew swift and efficient feature interaction design. DistDNAS proposes a supernet incorporating interaction modules of varying orders and types as a search space. To optimize search efficiency, DistDNAS distributes the search and aggregates the choice of optimal interaction modules on varying data dates, achieving a speed-up of over 25× and reducing the search cost from 2 days to 2 hours. To optimize serving efficiency, DistDNAS introduces a differentiable cost-aware loss to penalize the selection of redundant interaction modules, enhancing the efficiency of discovered feature interactions in serving. We extensively evaluate the best models crafted by DistDNAS on a 1TB Criteo Terabyte dataset. Experimental evaluations demonstrate 0.001 AUC improvement and 60% FLOPs saving over current state-of-the-art CTR models.
Tunhou Zhang, Wei Wen 0003, Igor Fedorov, Xi Liu 0011, Buyun Zhang, Fangqiu Han, Wen-Yen Chen, Yiping Han, Feng Yan 0001, Hai Li 0001, Yiran Chen 0001
IEEE Big Data7
2024 SiGeo: Sub-One-Shot NAS via Geometry of Loss Landscape
abstract
Neural Architecture Search (NAS) has become a widely used tool for automating neural network design. While one-shot NAS methods have successfully reduced computational requirements, they often require extensive training. On the other hand, zero-shot NAS utilizes training-free proxies to evaluate a candidate architecture's test performance but has two limitations: (1) inability to use the information gained as a network improves with training and (2) unreliable performance, particularly in complex domains like RecSys, due to the multi-modal data inputs and complex architecture configurations. To synthesize the benefits of both methods, we introduce a "sub-one-shot" paradigm that serves as a bridge between zero-shot and one-shot NAS. In sub-one-shot NAS, the supernet is trained using only a small subset of the training data, a phase we refer to as "warm-up." Within this framework, we present SiGeo, a proxy founded on a novel theoretical framework that connects the supernet warm-up with the efficacy of the proxy. Extensive experiments have consistently shown that SiGeo, when properly warmed up, surpasses state-of-the-art NAS proxies in many established NAS benchmarks in the computer vision domain. Furthermore, when tested on recommendation system benchmarks, SiGeo demonstrates its ability to match the performance of state-of-the-art weight-sharing one-shot NAS methods while significantly reducing computational costs by approximately 60%.
Kuang-Hung Liu, Igor Fedorov, Xin Zhang 0054, Wen-Yen Chen, Wei Wen 0003
KDD5
2012 Learning to blend vitality rankings from heterogeneous social networks
Jiang Bian 0002, Yi Chang 0001, Yun Fu 0003, Wen-Yen Chen
Neurocomputing4
2011 Parallel Spectral Clustering in Distributed Systems
abstract
Spectral clustering algorithms have been shown to be more effective in finding clusters than some traditional algorithms, such as k-means. However, spectral clustering suffers from a scalability problem in both memory use and computational time when the size of a data set is large. To perform clustering on large data sets, we investigate two representative ways of approximating the dense similarity matrix. We compare one approach by sparsifying the matrix with another by the Nyström method. We then pick the strategy of sparsifying the matrix via retaining nearest neighbors and investigate its parallelization. We parallelize both memory use and computation on distributed computers. Through an empirical study on a document data set of 193,844 instances and a photo data set of 2,121,863, we show that our parallel algorithm can effectively handle large problems.
Wen-Yen Chen, Yangqiu Song, Hongjie Bai, Chih-Jen Lin, Edward Y. Chang
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 PLDA: Parallel Latent Dirichlet Allocation for Large-Scale Applications
Hongjie Bai, Matt Stanton, Wen-Yen Chen, Edward Y. Chang
AAIM4
2009 Collaborative filtering for orkut communities: discovery of user latent behavior
abstract
Users of social networking services can connect with each other by forming communities for online interaction. Yet as the number of communities hosted by such websites grows over time, users have even greater need for effective community recommendations in order to meet more users. In this paper, we investigate two algorithms from very different domains and evaluate their effectiveness for personalized community recommendation. First is association rule mining (ARM), which discovers associations between sets of communities that are shared across many users. Second is latent Dirichlet allocation (LDA), which models user-community co-occurrences using latent aspects. In comparing LDA with ARM, we are interested in discovering whether modeling low-rank latent structure is more effective for recommendations than directly mining rules from the observed data. We experiment on an Orkut data set consisting of 492,104 users and 118,002 communities. Our empirical comparisons using the top-k recommendations metric show that LDA performs consistently better than ARM for the community recommendation task when recommending a list of 4 or more communities. However, for recommendation lists of up to 3 communities, ARM is still a bit better. We analyze examples of the latent information learned by LDA to explain this finding. To efficiently handle the large-scale data set, we parallelize LDA on distributed computers and demonstrate our parallel implementation's scalability with varying numbers of machines.
Wen-Yen Chen, Jon-Chyuan Chu, Junyi Luan, Hongjie Bai, Edward Y. Chang
WWW1
2008 Combinational collaborative filtering for personalized community recommendation
abstract
Rapid growth in the amount of data available on social networking sites has made information retrieval increasingly challenging for users. In this paper, we propose a collaborative filtering method, Combinational Collaborative Filtering (CCF), to perform personalized community recommendations by considering multiple types of co-occurrences in social data at the same time. This filtering method fuses semantic and user information, then applies a hybrid training strategy that combines Gibbs sampling and Expectation-Maximization algorithm. To handle the large-scale dataset, parallel computing is used to speed up the model training. Through an empirical study on the Orkut dataset, we show CCF to be both effective and scalable.
Wen-Yen Chen, Edward Y. Chang
KDD1
2008 Parallel Spectral Clustering
Yangqiu Song, Wen-Yen Chen, Hongjie Bai, Chih-Jen Lin, Edward Y. Chang
ECML/PKDD (2)2
2008 An agent-based metric for quality of services over wireless networks
Yaw-Chung Chen, Wen-Yen Chen
J. Syst. Softw.2
2006 An Agent-Based Metric for Quality of Services over Wireless Networks
abstract
In a wireless LAN environment, wireless stations with the strongest received signal can not be guaranteed to have the best quality of service if the population sharing the network capacity was not considered. In other words, within the same access point, the more the population, the less the shared bandwidth, and the worse the quality of service will be. In this paper, we proposed an anticipative agent assistance which is an agent-based metric for evaluating and managing the resource of the wireless access points, computing the potential AP list, and providing clients with resource information of APs. We also propose a novel QoS feedback mechanism which allows users to promptly adjust the service quality with AAA according to the throughput and delay requirements. We evaluate the performance of our proposed method using the ns-2 simulator. Numerical results show that AAA help reduce the transmission delay, increase the throughput, improve the network utilization, accommodate more users, and provide load-balancing
Yaw-Chung Chen, Wen-Yen Chen
COMPSAC (1)2
2006 Fotowiki: distributed map enhancement service
abstract
Fotowiki (FW) is a wiki-based map service that integrates visual and textual information with map. FW divides a geographical area into sub-areas. An individual responsible for providing information about a sub-area enters collected data into a wiki page. FW uploads distributed wiki-pages, and overlays the information on the map. This demonstration shows FW's architecture and functionalities.
Wen-Yen Chen, Benjamin N. Lee, Edward Y. Chang
ACM Multimedia1
2006 Fotofiti: web service for photo management
abstract
In this work, we present Fotofiti(FF), a web-based personal photo organizer with automatic image annotation, event management and social network integration. We describe our technique for real-time online semantic annotation of user photos. Additionally, a landmark recognition system which utilizes local features is discussed.
Benjamin N. Lee, Wen-Yen Chen, Edward Y. Chang
ACM Multimedia2
2006 A scalable service for photo annotation, sharing, and search
abstract
In this work we present the details of the implementation of Fotofiti(FF), a website that provides automatic semantic annotation of digital photographs, event management and social network integration. We describe our technique for real-time online semantic annotation using global features from both content and context. Classification experiments using various learning techniques were performed on a realworld data-set. Additionally, a scalable landmark recognition system which utilizes local features is discussed.
Benjamin N. Lee, Wen-Yen Chen, Edward Y. Chang
ACM Multimedia2