VLDB 2026 Research / reviewers in the wild / expert
Ko-Jen Hsiao
dblp:81/3155
· DBLP profile ↗
9ranked-venue papers
5as first author
3since 2021 · last 2025
0009-0003-6784-7556ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Orthogonal Low Rank Embedding StabilizationabstractThe instability of embedding spaces across model retraining cycles presents significant challenges to downstream applications using user or item embeddings derived from recommendation systems as input features. This paper introduces a novel orthogonal low-rank transformation methodology designed to stabilize the user/item embedding space, ensuring consistent embedding dimensions across retraining sessions. Our approach leverages a combination of efficient low-rank singular value decomposition and orthogonal Procrustes transformation to map embeddings into a standardized space. This transformation is computationally efficient, lossless, and lightweight, preserving the dot product and inference quality while reducing operational burdens. Unlike existing methods that modify training objectives or embedding structures, our approach maintains the integrity of the primary model application and can be seamlessly integrated with other stabilization techniques. Kevin Zielnicki, Ko-Jen Hsiao |
RecSys | 2 |
| 2024 | Sliding Window Training - Utilizing Historical Recommender Systems Data for Foundation ModelsabstractLong-lived recommender systems (RecSys) often encounter lengthy user-item interaction histories that span many years. To effectively learn long term user preferences, Large RecSys foundation models (FM) need to encode this information in pretraining. Usually, this is done by either generating a long enough sequence length to take all history sequences as input at the cost of large model input dimension or by dropping some parts of the user history to accommodate model size and latency requirements on the production serving side. In this paper, we introduce a sliding window training technique to incorporate long user history sequences during training time without increasing the model input dimension. We show the quantitative & qualitative improvements this technique brings to the RecSys FM in learning user long term preferences. We additionally show that the average quality of items in the catalog learnt in pretraining also improves. Swanand Joshi, Yesu Feng, Ko-Jen Hsiao, Sudarshan Lamkhede |
RecSys | 3 |
| 2024 | VideoRecSys + LargeRecSys 2024abstractWith the exponential growth of video and other content across various domains including entertainment, e-commerce, education and social media, there is a growing need for personalized content recommendations that are relevant to users’ interests. However, building effective and scalable content recommender systems is challenging due to factors such as the vast volume of content, diversity of user preferences, inherent noise and bias in data, and the need for real-time recommendations. Khushhall Chandra Mahajan, Amey Porobo Dharwadker, Brad Schumitsch, Arnab Bhadury, Ding Tong, Ko-Jen Hsiao, Liang Liu 0017 |
RecSys | 7 |
| 2016 | Multicriteria Similarity-Based Anomaly Detection Using Pareto Depth AnalysisabstractWe consider the problem of identifying patterns in a data set that exhibits anomalous behavior, often referred to as anomaly detection. Similarity-based anomaly detection algorithms detect abnormally large amounts of similarity or dissimilarity, e.g., as measured by the nearest neighbor Euclidean distances between a test sample and the training samples. In many application domains, there may not exist a single dissimilarity measure that captures all possible anomalous patterns. In such cases, multiple dissimilarity measures can be defined, including nonmetric measures, and one can test for anomalies by scalarizing using a nonnegative linear combination of them. If the relative importance of the different dissimilarity measures are not known in advance, as in many anomaly detection applications, the anomaly detection algorithm may need to be executed multiple times with different choices of weights in the linear combination. In this paper, we propose a method for similarity-based anomaly detection using a novel multicriteria dissimilarity measure, the Pareto depth. The proposed Pareto depth analysis (PDA) anomaly detection algorithm uses the concept of Pareto optimality to detect anomalies under multiple criteria without having to run an algorithm multiple times with different choices of weights. The proposed PDA approach is provably better than using linear combinations of the criteria, and shows superior performance on experiments with synthetic and real data sets. Ko-Jen Hsiao, Kevin S. Xu 0001, Jeff Calder, Alfred O. Hero III |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Pareto-Depth for Multiple-Query Image RetrievalabstractMost content-based image retrieval systems consider either one single query, or multiple queries that include the same object or represent the same semantic information. In this paper, we consider the content-based image retrieval problem for multiple query images corresponding to different image semantics. We propose a novel multiple-query information retrieval algorithm that combines the Pareto front method with efficient manifold ranking. We show that our proposed algorithm outperforms state of the art multiple-query retrieval algorithms on real-world image databases. We attribute this performance improvement to concavity properties of the Pareto fronts, and prove a theoretical result that characterizes the asymptotic concavity of the fronts. Ko-Jen Hsiao, Jeff Calder, Alfred O. Hero III |
IEEE Trans. Image Process. | 1 |
| 2014 | Luxapose: indoor positioning with mobile phones and visible lightabstractWe explore the indoor positioning problem with unmodified smartphones and slightly-modified commercial LED luminaires. The luminaires-modified to allow rapid, on-off keying-transmit their identifiers and/or locations encoded in human-imperceptible optical pulses. A camera-equipped smartphone, using just a single image frame capture, can detect the presence of the luminaires in the image, decode their transmitted identifiers and/or locations, and determine the smartphone's location and orientation relative to the luminaires. Continuous image capture and processing enables continuous position updates. The key insights underlying this work are (i) the driver circuits of emerging LED lighting systems can be easily modified to transmit data through on-off keying; (ii) the rolling shutter effect of CMOS imagers can be leveraged to receive many bits of data encoded in the optical transmissions with just a single frame capture, (iii) a camera is intrinsically an angle-of-arrival sensor, so the projection of multiple nearby light sources with known positions onto a camera's image plane can be framed as an instance of a sufficiently-constrained angle-of-arrival localization problem, and (iv) this problem can be solved with optimization techniques. We explore the feasibility of the design through an analytical model, demonstrate the viability of the design through a prototype system, discuss the challenges to a practical deployment including usability and scalability, and demonstrate decimeter-level accuracy in both carefully controlled and more realistic human mobility scenarios. Ye-Sheng Kuo, Pat Pannuto, Ko-Jen Hsiao, Prabal Dutta |
MobiCom | 3 |
| 2014 | Social collaborative retrievalabstractSocially-based recommendation systems have recently attracted significant interest, and a number of studies have shown that social information can dramatically improve a system's predictions of user interests. Meanwhile, there are now many potential applications that involve aspects of both recommendation and information retrieval, and the task of collaborative retrieval---a combination of these two traditional problems---has recently been introduced. Successful collaborative retrieval requires overcoming severe data sparsity, making additional sources of information, such as social graphs, particularly valuable. In this paper we propose a new model for collaborative retrieval, and show that our algorithm outperforms current state-of-the-art approaches by incorporating information from social networks. We also provide empirical analyses of the ways in which cultural interests propagate along a social graph using a real-world music dataset. Ko-Jen Hsiao, Alex Kulesza, Alfred O. Hero III |
WSDM | 1 |
| 2012 | Multi-criteria Anomaly Detection using Pareto Depth AnalysisabstractWe consider the problem of identifying patterns in a data set that exhibit anomalous behavior, often referred to as anomaly detection. In most anomaly detection algorithms, the dissimilarity between data samples is calculated by a single criterion, such as Euclidean distance. However, in many cases there may not exist a single dissimilarity measure that captures all possible anomalous patterns. In such a case, multiple criteria can be defined, and one can test for anomalies by scalarizing the multiple criteria by taking some linear combination of them. If the importance of the different criteria are not known in advance, the algorithm may need to be executed multiple times with different choices of weights in the linear combination. In this paper, we introduce a novel non-parametric multi-criteria anomaly detection method using Pareto depth analysis (PDA). PDA uses the concept of Pareto optimality to detect anomalies under multiple criteria without having to run an algorithm multiple times with different choices of weights. The proposed PDA approach scales linearly in the number of criteria and is provably better than linear combinations of the criteria. Ko-Jen Hsiao, Kevin S. Xu 0001, Jeff Calder, Alfred O. Hero III |
NIPS | 1 |
| 2008 | Fast fingertip positioning by combining particle filtering with particle random diffusionabstractA new and efficient algorithm to find the fingertips for human computer interface is proposed in this paper.The Fingertip Positioning algorithm of this paper combines particle filtering with Particle Random Diffusion to find the different fingertips quickly and robustly. There are two special methods used in this algorithm, which are Particle Random Diffusion and Fingertip Particle Selection. Without checking every pixel in the image of video sequence, this algorithm works well with low computational cost. This algorithm comes out with good performance in cluttered backgrounds and processes in real time. Based on the accurate positions of fingertips, hand gesture recognition can be done efficiently and robustly. Ko-Jen Hsiao, Tse-Wei Chen 0001, Shao-Yi Chien |
ICME | 1 |