EDBT 2026 Demo / reviewers in the wild / expert
Winston H. Hsu
dblp:16/5668
· DBLP profile ↗
14ranked-venue papers in the field
0as first author
5since 2021 · last 2023
0000-0002-3330-0638ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 14
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | STAMINA (Spatial-Temporal Aligned Meteorological INformation Attention) and FPL (Focal Precip Loss): Advancements in Precipitation Nowcasting for Heavy Rainfall EventsabstractPrecipitation nowcasting is crucial for weather-dependent decision-making in various sectors, providing accurate and high-resolution predictions of precipitation within a typical two-hour timeframe. Deep learning techniques have shown promise in improving nowcasting accuracy by leveraging large radar datasets. However, accurately predicting heavy rainfall events remains challenging due to several persistent problems in previous work. These include spatial-temporal misalignment between meteorological information and precipitation data, as well as the performance gap between different rainfall levels. To address these challenges, we propose two innovative modules: Spatial-Temporal Aligned Meteorological INformation Attention (STAMINA) and Focal Precip Loss (FPL). STAMINA integrates meteorological information using spatial-temporal embedding and pixelwise linear attention mechanisms to overcome spatial-temporal misalignment. FPL addresses event imbalance through event weighting and a penalty mechanism. Through extensive experiments, we demonstrate significant performance improvements achieved by STAMINA and FPL, with an 8% improvement in predicting light rainfall and, more significantly, a 30% improvement in heavy rainfall compared to the state-of-the-art DGMR model. These modules offer practical and effective solutions for enhancing nowcasting accuracy, with a specific focus on improving predictions for heavy rainfall events. By tackling the persistent problems in previous work, our proposed approach represents a significant advancement in the field of precipitation nowcasting. Ping-Chia Huang, Yueh-Li Chen, Yi-Syuan Liou, Bing-Chen Tsai, Chun-Chieh Wu, Winston H. Hsu |
CIKM | 6 |
| 2023 | CTCam: Enhancing Transportation Evaluation through Fusion of Cellular Traffic and Camera-Based Vehicle FlowsabstractTraffic prediction utility often faces infrastructural limitations, which restrict its coverage. To overcome this challenge, we present Geographical Cellular Traffic (GCT) flow that leverages cellular network data as a new source for transportation evaluation. The broad coverage of cellular networks allows GCT flow to capture various mobile user activities across regions, aiding city authorities in resource management through precise predictions. Acknowledging the complexity arising from the diversity of mobile users in GCT flow, we supplement it with camera-based vehicle flow data from limited deployments and verify their spatio-temporal attributes and correlations through extensive data analysis. Our two-stage fusion approach integrates these multi-source data, addressing their coverage and magnitude discrepancies, thereby enhancing the prediction of GCT flow for accurate transportation evaluation. Overall, we propose novel uses of telecom data in transportation and verify its effectiveness in multi-source fusion with vision-based data. Shen-Lung Tung, Hung-Ting Su, Winston H. Hsu |
CIKM | 4 |
| 2021 | Multivariate and Propagation Graph Attention Network for Spatial-Temporal Prediction with Outdoor Cellular TrafficabstractSpatial-temporal prediction is a critical problem for intelligent transportation, which is helpful for tasks such as traffic control and accident prevention. Previous studies rely on large-scale traffic data collected from sensors. However, it is unlikely to deploy sensors in all regions due to the device and maintenance costs. This paper addresses the problem via outdoor cellular traffic distilled from over two billion records per day in a telecom company, because outdoor cellular traffic induced by user mobility is highly related to transportation traffic. We study road intersections in urban and aim to predict future outdoor cellular traffic of all intersections given historic outdoor cellular traffic. Furthermore, we propose a new model for multivariate spatial-temporal prediction, mainly consisting of two extending graph attention networks (GAT). First GAT is used to explore correlations among multivariate cellular traffic. Another GAT leverages the attention mechanism into graph propagation to increase the efficiency of capturing spatial dependency. Experiments show that the proposed model significantly outperforms the state-of-the-art methods on our dataset. Hung-Ting Su, Shen-Lung Tung, Winston H. Hsu |
CIKM | 4 |
| 2021 | TrUMAn: Trope Understanding in Movies and AnimationsabstractUnderstanding and comprehending video content is crucial for many real-world applications such as search and recommendation systems. While recent progress of deep learning has boosted performance on various tasks using visual cues, deep cognition to reason intentions, motivation, or causality remains challenging. Existing datasets that aim to examine video reasoning capability focus on visual signals such as actions, objects, relations, or could be answered utilizing text bias. Observing this, we propose a novel task, along with a new dataset: Trope Understanding in Movies and Animations (TrUMAn), with 2423 videos associated with 132 tropes, intending to evaluate and develop learning systems beyond visual signals. Tropes are frequently used storytelling devices for creative works. By coping with the trope understanding task and enabling the deep cognition skills of machines, data mining applications and algorithms could be taken to the next level. To tackle the challenging TrUMAn dataset, we present a Trope Understanding and Storytelling (TrUSt) with a new Conceptual Storyteller module, which guides the video encoder by performing video storytelling on a latent space. Experimental results demonstrate that state-of-the-art learning systems on existing tasks reach only 12.01% of accuracy with raw input signals. Also, even in the oracle case with human-annotated descriptions, BERT contextual embedding achieves at most 28% of accuracy. Our proposed TrUSt boosts the model performance and reaches 13.94% performance. We also provide detailed analysis to pave the way for future research. TrUMAn is publicly available at:https://www.cmlab.csie.ntu.edu.tw/project/trope Hung-Ting Su, Po-Wei Shen, Bing-Chen Tsai, Wen-Feng Cheng, Ke-Jyun Wang, Winston H. Hsu |
CIKM | 6 |
| 2021 | Situation and Behavior Understanding by Trope Detection on FilmsabstractThe human ability of deep cognitive skills is crucial for the development of various real-world applications that process diverse and abundant user generated input. While recent progress of deep learning and natural language processing have enabled learning system to reach human performance on some benchmarks requiring shallow semantics, such human ability still remains challenging for even modern contextual embedding models, as pointed out by many recent studies [9, 10, 22, 24, 32]. Existing machine comprehension datasets assume sentence-level input, lack of casual or motivational inferences, or can be answered with question-answer bias. Here, we present a challenging novel task, trope detection on films, in an effort to create a situation and behavior understanding for machines. Tropes are frequently used storytelling devices for creative works. Comparing to existing movie tag prediction tasks, tropes are more sophisticated as they can vary widely, from a moral concept to a series of circumstances, and embedded with motivations and cause-and-effects. We introduce a new dataset, Tropes in Movie Synopses (TiMoS), with 5623 movie synopses and 95 different tropes collecting from a Wikipedia-style database, TVTropes. We present a multi-stream comprehension network (MulCom) leveraging multi-level attention of words, sentences, and role relations. Experimental result demonstrates that modern models including BERT contextual embedding, movie tag prediction systems, and relational networks, perform at most 37% of human performance (23.97/64.87) in terms of F1 score. Our MulCom outperforms all modern baselines, by 1.5 to 5.0 F1 score and 1.5 to 3.0 mean of average precision (mAP) score. We also provide a detailed analysis and human evaluation to pave ways for future research. Chen-Hsi Chang, Hung-Ting Su, Juiheng Hsu, Yu-Siang Wang, Zhe Yu Liu, Ya-Liang Chang, Wen-Feng Cheng, Ke-Jyun Wang, Winston H. Hsu |
WWW | 10 |
| 2015 | Approximating Weighted Hamming Distance by Probabilistic Selection for Multiple Hash Tables
Chiang-Yu Tsai, Yin-Hsi Kuo, Winston H. Hsu |
ECIR | 3 |
| 2014 | Predicting Viewer Affective Comments Based on Image Content in Social MediaabstractVisual sentiment analysis is getting increasing attention because of the rapidly growing amount of images in online social interactions and several emerging applications such as online propaganda and advertisement. Recent studies have shown promising progress in analyzing visual affect concepts intended by the media content publisher. In contrast, this paper focuses on predicting what viewer affect concepts will be triggered when the image is perceived by the viewers. For example, given an image tagged with "yummy food," the viewers are likely to comment "delicious" and "hungry," which we refer to as viewer affect concepts (VAC) in this paper. To the best of our knowledge, this is the first work explicitly distinguishing intended publisher affect concepts and induced viewer affect concepts associated with social visual content, and aiming at understanding their correlations. We present around 400 VACs automatically mined from million-scale real user comments associated with images in social media. Furthermore, we propose an automatic visual based approach to predict VACs by first detecting publisher affect concepts in image content and then applying statistical correlations between such publisher affect concepts and the VACs. We demonstrate major benefits of the proposed methods in several real-world tasks - recommending images to invoke certain target VACs among viewers, increasing the accuracy of predicting VACs by 20.1% and finally developing a social assistant tool that may suggest plausible, content-specific and desirable comments when users view new images. Yan-Ying Chen, Tao Chen 0015, Winston H. Hsu, Hong-Yuan Mark Liao, Shih-Fu Chang |
ICMR | 3 |
| 2014 | Rank-Preserving and Unsupervised Hash Learning from Auxiliary Contextual CuesabstractDue to the rapid growth of multimedia capture devices, the needs of efficient large-scale multimedia retrieval systems have attracted much attention in different research communities. Recent work shows that hash-based approaches (e.g., LSH - locality-sensitive hashing) can provide efficient and effective retrieval results for approximate nearest neighbor (ANN) problem. Several hash-based learning methods have been proposed to generate more compact binary codes; however, most of them attempt to preserve the similarity score or class (label) information. In the content-based image retrieval (CBIR) tasks, the rank of the search results is also an important information. For web-scale image search, supervised annotations are generally not available; nevertheless, we can easily derive auxiliary semantic cues from the ranking results by different modalities such as keyword search, GPS, tags, where such rankings are easily available. In this work, we propose to preserve the ranking information by utilizing the rank order to learn the hashing functions. Experimental results show that the proposed method is comparable to the state-of-the-art learning-based approaches and can be extended to other existing methods. Moreover, we can further improve the retrieval accuracy by incorporating auxiliary ranking information. Yin-Hsi Kuo, Winston H. Hsu |
ICMR | 2 |
| 2014 | Efficient Face Detection by Leveraging Knowledge from Large-Scale PhotosabstractThe explosive growth of uploaded photos requires service providers to improve the overall efficiency for analyzing and manipulating uploaded photos. Seeing the needs, we investigate the feasibility to speed up photo-related analysis (e.g., face detection) by leveraging the observations from large-scale consumer photos. To deal with the huge computations of face detection on the large-scale photos, we aim to reduce search hypotheses (i.e., subwindows) for face detectors by thresholding (or ranking) on the confidence of hypothesis that indicates the possibility to be a face. Unlike the traditional sliding window which views all the hypotheses equally, prior knowledge of face sizes and locations in images is learned from the large-scale consumer photos, and used to model the confidence of each hypothesis. By thresholding on the ranked hypotheses, less likely hypotheses can be reduced for better efficiency. Our experiment demonstrates that 40% subwindows can be skipped while maintaining 95% face recall, compared to traditional sliding-window-based methods. Yi-De Lin, Yen-Liang Lin, Winston H. Hsu |
ICMR | 3 |
| 2014 | Learning to personalize trending image search suggestionabstractTrending search suggestion is leading a new paradigm of image search, where user's exploratory search experience is facilitated with the automatic suggestion of trending queries. Existing image search engines, however, only provide general suggestions and hence cannot capture user's personal interest. In this paper, we move one step forward to investigate personalized suggestion of trending image searches according to users' search behaviors. To this end, we propose a learning-based framework including two novel components. The first component, i.e., trending-aware weight-regularized matrix factorization (TA-WRMF), is able to suggest personalized trending search queries by learning user preference from many users as well as auxiliary common searches. The second component associates the most representative and trending image with each suggested query. The personalized suggestion of image search consists of a trending textual query and its associated trending image. The combined textual-visual queries not only are trending (bursty) and personalized to user's search preference, but also provide the compelling visual aspect of these queries. We evaluate our proposed learning-based framework on a large-scale search logs with 21 million users and 41 million queries in two weeks from a commercial image search engine. The evaluations demonstrate that our system achieve about 50% gain compared with state-of-the-art in terms of query prediction accuracy. Chun-Che Wu, Tao Mei 0001, Winston H. Hsu, Yong Rui |
SIGIR | 3 |
| 2012 | Where is who: large-scale photo retrieval by facial attributes and canvas layoutabstractThe ubiquitous availability of digital cameras has made it easier than ever to capture moments of life, especially the ones accompanied with friends and family. It is generally believed that most family photos are with faces that are sparsely tagged. Therefore, a better solution to manage and search in the tremendously growing personal or group photos is highly anticipated. In this paper, we propose a novel way to search for face photos by simultaneously considering attributes (e.g., gender, age, and race), positions, and sizes of the target faces. To better match the content and layout of the multiple faces in mind, our system allows the user to graphically specify the face positions and sizes on a query "canvas," where each attribute combination is defined as an icon for easier representation. As a secondary feature, the user can even place specific faces from the previous search results for appearance-based retrieval. The scenario has been realized on a tablet device with an intuitive touch interface. Experimenting with a large-scale Flickr dataset of more than 200k faces, the proposed formulation and joint ranking have made us achieve a hit rate of 0.420 at rank 100, significantly improving from 0.036 of the prior search scheme using attributes alone. We have also achieved an average running time of 0.0558 second by the proposed block-based indexing approach. Yu-Heng Lei, Yan-Ying Chen, Bor-Chun Chen, Lime Iida, Winston H. Hsu |
SIGIR | 5 |
| 2011 | Region-based landmark discovery by crowdsourcing geo-referenced photosabstractWe propose a novel model for landmark discovery that locates region-based landmarks on map in contrast to the traditional point-based landmarks. The proposed method preserves more information and automatically identifies candidate regions on map by crowdsourcing geo-referenced photos. Gaussian kernel convolution is applied to remove noises and generate detected region. We adopt F1 measure to evaluate discovered landmarks and manually check the association between tags and regions. The experiment results show that more than 90% of attractions in the selected city can be correctly located by this method. Yen-Ta Huang, An-Jung Cheng, Liang-Chi Hsieh, Winston H. Hsu, Kuo-Wei Chang |
SIGIR | 4 |
| 2011 | Multi-layer graph-based semi-supervised learning for large-scale image datasets using mapreduceabstractSemi-supervised learning is to exploit the vast amount of unlabeled data in the world. This paper proposes a scalable graph-based technique leveraging the distributed computing power of the MapReduce programming model. For a higher quality of learning, the paper also presents a multi-layer learning structure to unify both visual and textual information of image data during the learning process. Experimental results show the effectiveness of the proposed methods. Wen-Yu Lee, Liang-Chi Hsieh, Guan-Long Wu, Winston H. Hsu, Ya-Fan Su |
SIGIR | 4 |
| 2008 | AdImage: video advertising by image matching and ad scheduling optimizationabstractWith the prevalence of recording devices and the ease of media sharing, consumers are embracing huge amounts of Internet videos. There arise the needs for effective video advertisement systems following their phenomenal success in text. We propose a novel advertising system, AdImage, which automatically associates relevant ads by matching characteristic images, referred to as adImages (analogous to adWords) here. The proposed image matching method is invariant to certain distortions commonly observed in shared videos. AdImage also avoids the pitfalls of poor tagging qualities in shared videos and provides a brand-new venue to specify ad targets by image objects. Moreover, we formulate the image matching scores and the parameterized bidding information as a nonlinear optimization problem for maximizing the system revenues and user perception. Wei-Shing Liao, Kuan-Ting Chen, Winston H. Hsu |
SIGIR | 3 |