Yichao Lu

dblp:68/10509 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 TriSeRec: A Tri-view Representation Learning Framework for Sequential/Session-based Recommendation
abstract
Sequential/session-based recommendation models aim to learn evolving user preferences from historical user behaviors. State-of-the-art sequential/session-based recommendation models often use graph neural networks or self-attention as their building blocks. Graph neural networks excel at learning local patterns encoded in graph-structured data and have therefore shown great performance on session-based recommendation datasets, where user interactions are usually relatively short. Self-attentive models, on the other hand, are much more powerful in capturing long-range dependencies and are able to outperform graph neural network-based approaches on sequential recommendation, where longer user interactions are more frequent. As such, the recommender systems community has noted a lack of a unified framework that can simultaneously achieve great performance on both sequential and session-based recommendation. In an effort to fill this gap, this paper presents TriSeRec, a Tri-view representation learning framework for Sequential/session-based Recommendation. By converting interaction sequences into two graphical views and one sequential view, three view-specific user representations are learned by TriSeRec using graph neural networks and self-attention. The tri-view representation learning module, which is built upon the recently proposed generalized Cauchy-Schwarz divergence, disentangles and then fuses consistent and complementary information in all three views to form the final user representations for next-item predictions. Experiments on popular large-scale, real-world benchmark datasets show that TriSeRec achieves state-of-the-art performance on both sequential recommendation and session-based recommendation.
Xinchen Yuan, Yichao Lu
CIKM2
2024 Lumos: Empowering Multimodal LLMs with Scene Text Recognition
abstract
We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component that extracts text from first person point-of-view images, the output of which is used to augment input to a Multimodal Large Language Model (MM-LLM). While building Lumos, we encountered numerous challenges related to STR quality, overall latency, and model inference. In this paper, we delve into those challenges, and discuss the system architecture, design choices, and modeling techniques employed to overcome these obstacles. We also provide a comprehensive evaluation for each component, showcasing high quality and efficiency.
Ashish Shenoy, Yichao Lu, Srihari Jayakumar, Debojeet Chatterjee, Mohsen Moslehpour, Pierce Chuang, Abhay Harpale, Vikas Bhardwaj, Di Xu 0009, Shicong Zhao, Longfang Zhao, Ankit Ramchandani, Xin Dong 0001
KDD2
2021 Predicting Research Trends in Artificial Intelligence with Gradient Boosting Decision Trees and Time-aware Graph Neural Networks
abstract
The Science4cast 2021 competition focuses on predicting future edges in an evolving semantic network, where each vertex represents an artificial intelligence concept, and an edge between a pair of vertices denotes that the two concepts have been investigated together in a scientific paper. In this paper, we describe our solution to this competition. We present two distinct approaches: a tree-based gradient boosting approach and a deep learning approach, and demonstrate that both approaches achieve competitive performance. Our final solution, which is based on a blend of the two approaches, achieved the 1st place among all the participating teams. The source code for this paper is available at https://github.com/YichaoLu/Science4cast2021.
Yichao Lu
IEEE BigData1
2021 Context-aware Scene Graph Generation with Seq2Seq Transformers
abstract
Scene graph generation is an important task in computer vision aimed at improving the semantic understanding of the visual world. In this task, the model needs to detect objects and predict visual relationships between them. Most of the existing models predict relationships in parallel assuming their independence. While there are different ways to capture these dependencies, we explore a conditional approach motivated by the sequence-to-sequence (Seq2Seq) formalism. Different from the previous research, our proposed model predicts visual relationships one at a time in an autoregressive manner by explicitly conditioning on the already predicted relationships. Drawing from translation models in NLP, we propose an encoder-decoder model built using Transformers where the encoder captures global context and long range interactions. The decoder then makes sequential predictions by conditioning on the scene graph constructed so far. In addition, we introduce a novel reinforcement learning-based training strategy tailored to Seq2Seq scene graph generation. By using a self-critical policy gradient training approach with Monte Carlo search we directly optimize for the (mean) recall metrics and bridge the gap between training and evaluation. Experimental results on two public benchmark datasets demonstrate that our Seq2Seq learning approach achieves strong empirical performance, outperforming previous state-of-the-art, while remaining efficient in terms of training and inference time. Full code for this work is available here: https://github.com/layer6ai-labs/SGG-Seq2Seq.
Yichao Lu, Himanshu Rai, Boris Knyazev 0001, Guangwei Yu, Shashank Shekhar 0005, Graham W. Taylor, Maksims Volkovs
ICCV1
2021 Weakly Supervised Extractive Summarization with Attention
abstract
Automatic summarization aims to extract important information from large amounts of textual data in order to create a shorter version of the original texts while preserving its information.Training traditional extractive summarization models relies heavily on humanengineered labels such as sentence-level annotations of summary-worthiness.However, in many use cases, such human-engineered labels do not exist and manually annotating thousands of documents for the purpose of training models may not be feasible.On the other hand, indirect signals for summarization are often available, such as agent actions for customer service dialogues, headlines for news articles, diagnosis for Electronic Health Records, etc.In this paper, we develop a general framework that generates extractive summarization as a byproduct of supervised learning tasks for indirect signals via the help of attention mechanism.We test our models on customer service dialogues and experimental results demonstrated that our models can reliably select informative sentences and words for automatic summarization.
Yingying Zhuang, Yichao Lu
SIGDIAL2
2020 Don't Use English Dev: On the Zero-Shot Cross-Lingual Evaluation of Contextual Embeddings
abstract
Multilingual contextual embeddings have demonstrated state-of-the-art performance in zero-shot cross-lingual transfer learning, where multilingual BERT is fine-tuned on one source language and evaluated on a different target language.However, published results for mBERT zero-shot accuracy vary as much as 17 points on the MLDoc classification task across four papers.We show that the standard practice of using English dev accuracy for model selection in the zero-shot setting makes it difficult to obtain reproducible results on the MLDoc and XNLI tasks.English dev accuracy is often uncorrelated (or even anti-correlated) with target language accuracy, and zero-shot performance varies greatly at different points in the same fine-tuning run and between different fine-tuning runs.These reproducibility issues are also present for other tasks with different pre-trained embeddings (e.g., MLQA with XLM-R).We recommend providing oracle scores alongside zero-shot results: still fine-tune using English data, but choose a checkpoint with the target dev set.Reporting this upper bound makes results more consistent by avoiding arbitrarily bad checkpoints.
Phillip Keung, Yichao Lu, Julian Salazar, Vikas Bhardwaj
EMNLP (1)2
2020 The Multilingual Amazon Reviews Corpus
abstract
We present the Multilingual Amazon Reviews Corpus (MARC), a large-scale collection of Amazon reviews for multilingual text classification.The corpus contains reviews in English, Japanese, German, French, Spanish, and Chinese, which were collected between 2015 and 2019.Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID, and the coarse-grained product category (e.g., 'books', 'appliances', etc.)The corpus is balanced across the 5 possible star ratings, so each rating constitutes 20% of the reviews in each language.For each language, there are 200,000, 5,000, and 5,000 reviews in the training, development, and test sets, respectively.We report baseline results for supervised text classification and zero-shot crosslingual transfer learning by fine-tuning a multilingual BERT model on reviews data.We propose the use of mean absolute error (MAE) instead of classification accuracy for this task, since MAE accounts for the ordinal nature of the ratings.
Phillip Keung, Yichao Lu, György Szarvas, Noah A. Smith
EMNLP (1)2
2020 Efficient and Information-Preserving Future Frame Prediction and Beyond
Wei Yu 0001, Yichao Lu, Steve M. Easterbrook, Sanja Fidler
ICLR2
2020 Unsupervised Bitext Mining and Translation via Self-trained Contextual Embeddings
abstract
We describe an unsupervised method to create pseudo-parallel corpora for machine translation (MT) from unaligned text. We use multilingual BERT to create source and target sentence embeddings for nearest-neighbor search and adapt the model via self-training. We validate our technique by extracting parallel sentence pairs on the BUCC 2017 bitext mining task and observe up to a 24.5 point increase (absolute) in F1 scores over previous unsupervised methods. We then improve an XLM-based unsupervised neural MT system pre-trained on Wikipedia by supplementing it with pseudo-parallel text mined from the same corpus, boosting unsupervised translation performance by up to 3.5 BLEU on the WMT’14 French-English and WMT’16 German-English tasks and outperforming the previous state-of-the-art. Finally, we enrich the IWSLT’15 English-Vietnamese corpus with pseudo-parallel Wikipedia sentence pairs, yielding a 1.2 BLEU improvement on the low-resource MT task. We demonstrate that unsupervised bitext mining is an effective way of augmenting MT datasets and complements existing techniques like initializing with pre-trained contextual embeddings.
Phillip Keung, Julian Salazar, Yichao Lu, Noah A. Smith
Trans. Assoc. Comput. Linguistics3
2019 Adversarial Learning with Contextual Embeddings for Zero-resource Cross-lingual Classification and NER
abstract
Phillip Keung, Yichao Lu, Vikas Bhardwaj. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Phillip Keung, Yichao Lu, Vikas Bhardwaj
EMNLP/IJCNLP (1)2
2018 Why I like it: multi-task learning for recommendation and explanation
abstract
We describe a novel, multi-task recommendation model, which jointly learns to perform rating prediction and recommendation explanation by combining matrix factorization, for rating prediction, and adversarial sequence to sequence learning for explanation generation. The result is evaluated using real-world datasets to demonstrate improved rating prediction performance, compared to state-of-the-art alternatives, while producing effective, personalized explanations.
Yichao Lu, Ruihai Dong, Barry Smyth
RecSys1
2018 Coevolutionary Recommendation Model: Mutual Learning between Ratings and Reviews
abstract
Collaborative filtering (CF) is a common recommendation approach that relies on user-item ratings. However, the natural sparsity of user-item rating data can be problematic in many domains and settings, limiting the ability to generate accurate predictions and effective recommendations. Moreover, in some CF approaches latent features are often used to represent users and items, which can lead to a lack of recommendation transparency and explainability. User-generated, customer reviews are now commonplace on many websites, providing users with an opportunity to convey their experiences and opinions of products and services. As such, these reviews have the potential to serve as a useful source of recommendation data, through capturing valuable sentiment information about particular product features. In this paper, we present a novel deep learning recommendation model, which co-learns user and item information from ratings and customer reviews, by optimizing matrix factorization and an attention-based GRU network. Using real-world datasets we show a significant improvement in recommendation performance, compared to a variety of alternatives. Furthermore, the approach is useful when it comes to assigning intuitive meanings to latent features to improve the transparency and explainability of recommender systems.
Yichao Lu, Ruihai Dong, Barry Smyth
WWW1
2015 Finding Linear Structure in Large Datasets with Scalable Canonical Correlation Analysis
abstract
Canonical Correlation Analysis (CCA) is a widely used spectral technique for finding correlation structures in multi-view datasets. In this paper, we tackle the problem of large scale CCA, where classical algorithms, usually requiring computing the product of two huge matrices and huge matrix decomposition, are computationally and storage expensive. We recast CCA from a novel perspective and propose a scalable and memory efficient \textitAugmented Approximate Gradient (AppGrad) scheme for finding top k dimensional canonical subspace which only involves large matrix multiplying a thin matrix of width k and small matrix decomposition of dimension k\times k. Further, \textitAppGrad achieves optimal storage complexity O(k(p_1+p_2)), compared with classical algorithms which usually require O(p_1^2+p_2^2) space to store two dense whitening matrices. The proposed scheme naturally generalizes to stochastic optimization regime, especially efficient for huge datasets where batch algorithms are prohibitive. The online property of stochastic \textitAppGrad is also well suited to the streaming scenario, where data comes sequentially. To the best of our knowledge, it is the first stochastic algorithm for CCA. Experiments on four real data sets are provided to show the effectiveness of the proposed methods.
Yichao Lu, Dean P. Foster
ICML2
2014 An efficient decoder architecture for cyclic non-binary LDPC codes
abstract
This paper proposes a hybrid message-passing decoding algorithm which consumes very low computational complexity, while achieving competitive error performance compared with conventional min-max algorithm. Simulation result on a (255,175) cyclic code shows that this algorithm obtains at least 0.5dB coding gain over other state-of-the-art low-complexity non-binary LDPC (NB-LDPC) decoding algorithms. Based on this algorithm, a partial-parallel decoder architecture is implemented for cyclic NB-LDPC codes, where the variable node units are redesigned and the routing network is optimized for the proposed algorithm. Synthesis results demonstrate that about 24.3% gates and 12% memories can be saved over previous works.
Yichao Lu, Guifen Tian, Satoshi Goto
ISCAS1
2014 large scale canonical correlation analysis with iterative least squares
Yichao Lu, Dean P. Foster
NIPS1
2014 Fast Ridge Regression with Randomized Principal Component Analysis and Gradient Descent
Yichao Lu, Dean P. Foster
UAI1
2013 Informed dynamic scheduling for majority-logic decoding of non-binary LDPC codes
abstract
Recent studies show that dynamic scheduling using the latest available information is superior to the flooding scheduling. While the existing works on informed dynamic scheduling focus on belief propagation (BP) based algorithm for binary LDPC codes, in this paper we devise a novel method which is appropriate for majority-logic decoding of non-binary LDPC codes. We firstly clarify that previous methods are not suitable for majority-logic decoding of non-binary LDPC codes. Then we give a practical scheduling strategy which utilizes the stability of check nodes to select the messages for propagation. Furthermore, we discuss the computational complexity of this proposed scheme in detail. Simulation results verify that our approach can achieve better performance compared with layered and flooding schemes.
Nanfan Qiu, Wen Chen 0001, Yichao Lu, Yang Yu 0041
GLOBECOM3
2013 New Subsampling Algorithms for Fast Least Squares Regression
abstract
We address the problem of fast estimation of ordinary least squares (OLS) from large amounts of data ($n \gg p$). We propose three methods which solve the big data problem by subsampling the covariance matrix using either a single or two stage estimation. All three run in the order of size of input i.e. O($np$) and our best method, {\it Uluru}, gives an error bound of $O(\sqrt{p/n})$ which is independent of the amount of subsampling as long as it is above a threshold. We provide theoretical bounds for our algorithms in the fixed design (with Randomized Hadamard preconditioning) as well as sub-Gaussian random design setting. We also compare the performance of our methods on synthetic and real-world datasets and show that if observations are i.i.d., sub-Gaussian then one can directly subsample without the expensive Randomized Hadamard preconditioning without loss of accuracy.
Paramveer S. Dhillon, Yichao Lu, Dean P. Foster, Lyle H. Ungar
NIPS2
2013 Faster Ridge Regression via the Subsampled Randomized Hadamard Transform
abstract
We propose a fast algorithm for ridge regression when the number of features is much larger than the number of observations ($p \gg n$). The standard way to solve ridge regression in this setting works in the dual space and gives a running time of $O(n^2p)$. Our algorithm (SRHT-DRR) runs in time $O(np\log(n))$ and works by preconditioning the design matrix by a Randomized Walsh-Hadamard Transform with a subsequent subsampling of features. We provide risk bounds for our SRHT-DRR algorithm in the fixed design setting and show experimental results on synthetic and real datasets.
Yichao Lu, Paramveer S. Dhillon, Dean P. Foster, Lyle H. Ungar
NIPS1