VLDB 2026 Research / reviewers in the wild / expert
Chih-Fan Hsu
dblp:120/8997
· DBLP profile ↗
23ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0002-4180-8255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 5 since 2021Computer networks · 4 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Comprehensive Evaluation of Encrypted DNN Inference MethodsabstractFully Homomorphic Encryption (FHE) in the realm of deep neural network (DNN) encrypted inference represents a pivotal advancement in privacy-preserving machine learning. This technology allows users to securely access DNN inference services hosted on remote servers without compromising their personal privacy. Given its wide range of potential applications, FHE has garnered significant research attention. Despite its rapid progress, FHE still faces considerable challenges, particularly the high computational resource demands. Moreover, varying configurations of FHE schemes can lead to notable differences in performance, whether in terms of efficiency or inference accuracy, making it difficult to strike an optimal balance tailored to specific application requirements.To tackle this challenge, we introduce a new approach that simulates FHE-induced errors to assess the impact of different FHE architectures and parameter configurations on encrypted inference during the model testing phase. Our simulation framework enables efficient approximation of homomorphic computation outcomes on a DNN model, specifically for third-generation FHEs, without the need for executing the complete homomorphic process. This method significantly streamlines the research process for optimizing parameters based on specific application needs. We validate the effectiveness of our approach through extensive performance benchmarking across a variety of experimental settings. Yu-Te Ku, Ming-Chien Ho, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, Wei-Chao Chen |
ISCAS | 5 |
| 2025 | Optimizing Encrypted Neural Networks: Model Design, Quantization and Fine-Tuning Using FHEW/TFHEabstractThird-generation Fully Homomorphic Encryption (FHE), particularly the FHEW/TFHE schemes, is recognized for its balanced security requirements, small parameters, and low memory usage, though the current methods in the scenarios of Deep Neural Network (DNN) inference still have high computational costs, limiting the practical applicability. This work demonstrates how to improve practicality of the third-generation technologies for DNN tasks while preserving its key advantages. Our work focuses on two main contributions. First, we developed a computational architecture called FHE-Neuron, which reconfigures the parameters and bootstrapping structure of traditional FHEW/TFHE Boolean operations. This architecture significantly reducing the cost of encrypted DNN inference by dynamically switching the precision of encrypted data during computation—using high precision for cost-effective linear operations and low precision for computationally expensive nonlinear operations. Second, we introduced an FHE-aware Quantization and Fine-tuning framework that optimizes model parameters to align with FHE-Neuron’s constraints, ensuring high accuracy in encrypted inference. We validate our approach on various neural network models across several computing platforms. In our experiments, our method achieves one-image inference time on average 4.5 milliseconds for MNIST and 17 milliseconds for Fashion MNIST, achieving accuracy rates of 96.52% and 88.57% respectively. For the CIFAR-10 dataset, our system completes one image inference in 30 seconds with a 90.5% accuracy rate. Yu-Te Ku, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, I-Ping Tu, Wei-Chao Chen |
Proc. Priv. Enhancing Technol. | 3 |
| 2024 | Invited Paper: Efficient Design of FHEW/TFHE Bootstrapping Implementation with Scalable ParametersabstractFully Homomorphic Encryption (FHE) is vital for computing over encrypted data, thereby enabling numerous privacy-preserving applications. This work focuses on the third generation FHE schemes (e.g., FHEW and TFHE), known for their fast bootstrapping, small FHE parameters, and robust security built on milder assumptions. Ming-Chien Ho, Yu-Te Ku, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, Wei-Chao Chen |
ICCAD | 5 |
| 2024 | Learning With Instance-Dependent Noisy Labels By Anchor Hallucination And Hard Sample Label CorrectionabstractLearning from noisy-labeled data is crucial for real-world applications. Traditional Noisy-Label Learning (NLL) methods categorize training data into clean and noisy sets based on the loss distribution of training samples. However, they often neglect that clean samples, especially those with intricate visual patterns, may also yield substantial losses. This oversight is particularly significant in datasets with Instance-Dependent Noise (IDN), where mislabeling probabilities correlate with visual appearance. Our approach explicitly distinguishes between clean $v s$. noisy and easy $v s$. hard samples. We identify training samples with small losses, assuming they have simple patterns and correct labels. Utilizing these easy samples, we hallucinate multiple anchors to select hard samples for label correction. Corrected hard samples, along with the easy samples, are used as labeled data in subsequent semi-supervised training. Experiments on synthetic and real-world IDN datasets demonstrate the superior performance of our method over other state-of-the-art NLL methods. Po-Hsuan Huang, Chia-Ching Lin, Chih-Fan Hsu, Ming-Ching Chang, Wei-Chao Chen |
ICIP | 3 |
| 2024 | MAFS: Modality-Aware Federated Semi-Supervised Learning with Selective Data Sharing Specified by Individual Clients
Yi-Chen Li 0005, Chih-Fan Hsu, Jian-Kai Wang, Chung-Chi Tsai, Cheng-Hsin Hsu |
MMAsia | 2 |
| 2024 | Federated Learning Using Multi-Modal Sensors with Heterogeneous Privacy Sensitivity LevelsabstractData from multi-modal sensors, such as Red-Green-Blue (RGB) cameras, thermal cameras, microphones, and mmWave radars, have gradually been adopted in various classification problems for better accuracy. Some sensors, like RGB cameras and microphones, however, capture privacy-invasive data, which are less likely to be used in centralized learning. Although the Federated Learning (FL) paradigm frees clients from sharing their sensor data, doing so results in reduced classification accuracy and increased training time. In this article, we introduce a novel Heterogeneous Privacy Federated Learning (HPFL) paradigm to better capitalize on the less privacy-invasive sensor data, such as thermal images and mmWave point clouds, by uploading them to the server for closing the performance gap between FL and centralized learning. HPFL not only allows clients to keep the more privacy-invasive sensor data private, such as RGB images and human voices, but also gives each client total freedom to define the levels of their privacy concern on individual sensor modalities. For example, more sensitive users may prefer to keep their thermal images private, while others do not mind sharing these images. We carry out extensive experiments to evaluate the HPFL paradigm using two representative classification problems: semantic segmentation and emotion recognition. Several key findings demonstrate the merits of HPFL: (i) compared to FedAvg, it improves foreground accuracy by 18.20% in semantic segmentation and boosts the F1-score by 4.20% in emotion recognition, (ii) with heterogeneous privacy concern levels, it achieves an even larger F1-score improvement of 6.17–16.05% in emotion recognition, and (iii) it also outperforms the state-of-the-art FL approaches by 12.04–17.70% in foreground accuracy and 2.54–4.10% in F1-score. Chih-Fan Hsu, Yi-Chen Li 0005, Chung-Chi Tsai, Jian-Kai Wang, Cheng-Hsin Hsu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Quantitative Comparison of Point Cloud Compression Algorithms With PCC ArenaabstractWith the growth of Extended Reality (XR) and capturing devices, point cloud representation has become attractive to academics and industry. Point Cloud Compression (PCC) algorithms further promote numerous XR applications that may change our daily life. However, in the literature, PCC algorithms are often evaluated with heterogeneous datasets, metrics, and parameters, making the results hard to interpret. In this article, we propose an open-source benchmark platform called PCC Arena. Our platform is modularized in three aspects: PCC algorithms, point cloud datasets, and performance metrics. Users can easily extend PCC Arena in each aspect to fulfill the requirements of their experiments. To show the effectiveness of PCC Arena, we integrate seven PCC algorithms into PCC Arena along with six point cloud datasets. We then compare the algorithms on ten carefully selected metrics to evaluate the quality of the output point clouds. We further conduct a user study to quantify the user-perceived quality of rendered images that are produced by different PCC algorithms. Several novel insights are revealed in our comparison: (i) Signal Processing (SP)-based PCC algorithms are stable for different usage scenarios, but the trade-offs between coding efficiency and quality should be carefully addressed, (ii) Neural Network (NN)-based PCC algorithms have the potential to consume lower bitrates yet provide similar results to SP-based algorithms, (iii) NN-based PCC algorithms may generate artifacts and suffer from long running time, and (iv) NN-based PCC algorithms are worth more in-depth studies as the recently proposed NN-based PCC algorithms improve the quality and running time. We believe that PCC Arena can play an essential role in allowing engineers and researchers to better interpret and compare the performance of future PCC algorithms. Cheng-Hao Wu, Chih-Fan Hsu, Tzu-Kuan Hung, Carsten Griwodz, Wei Tsang Ooi, Cheng-Hsin Hsu |
IEEE Trans. Multim. | 2 |
| 2022 | Towards a Utopia of Dataset Sharing: A Case Study on Machine Learning-based Malware Detection AlgorithmsabstractWorking with a high-quality (complete and up-to-date) dataset is the key to building a good machine learning model, especially in security research areas. However, it is not easy to collect a good quality dataset for security research communities because of the sensitive property of most security datasets. We believe that having more contributors to share up-to-date samples would increase the quality of datasets. Therefore, this study aims to increase security dataset sharing for research communities by eliminating possible information leakage. We propose a dataset sharing model and the core algorithm, FeatureTransformer, which guarantees no sensitive information leakage from a shared dataset. FeatureTransformer transforms extracted raw features into intermediate features that conceal sensitive information. Meanwhile, models built from transformed features maintain similar performance compared to models built from the original raw features. We show the effectiveness of our model by evaluating FeatureTransformer with typical malware classification problems using (1) traditional machine learning classifiers and (2) neural network-based classifiers. The experiment results show that the models trained with transformed features merely suffer from 2.56% and 1.48% accuracy degradation on the investigated problems. It indicates that models validated by datasets processed by FeatureTransformer work well with the original raw (untransformed) datasets. We believe that our privacy-preserving model can stimulate dataset sharing and advance the development of machine learning approaches in solving security problems. Ping-Jui Chuang, Chih-Fan Hsu, Yung-Tien Chu, Szu-Chun Huang, Chun-Ying Huang |
AsiaCCS | 2 |
| 2022 | Optimizing Immersive Video Coding Configurations Using Deep Learning: A Case Study on TMIVabstractImmersive video streaming technologies improve Virtual Reality (VR) user experience by providing users more intuitive ways to move in simulated worlds, e.g., with 6 Degree-of-Freedom (6DoF) interaction mode. A naive method to achieve 6DoF is deploying cameras at numerous different positions and orientations that may be required based on users’ movement, which unfortunately is expensive, tedious, and inefficient. A better solution for realizing 6DoF interactions is to synthesize target views on-the-fly from a limited number of source views. While such view synthesis is enabled by the recent Test Model for Immersive Video (TMIV) codec, TMIV dictates manually-composed configurations, which cannot exercise the tradeoff among video quality, decoding time, and bandwidth consumption. In this article, we study the limitation of TMIV and solve its configuration optimization problem by searching for the optimal configuration in a huge configuration space. We first identify the critical parameters in the TMIV configurations. Then, we introduce two Neural Network (NN) -based algorithms from two heterogeneous aspects: (i) a Convolutional Neural Network (CNN) algorithm solving a regression problem and (ii) a Deep Reinforcement Learning (DRL) algorithm solving a decision making problem, respectively. We conduct both objective and subjective experiments to evaluate the CNN and DRL algorithms on two diverse datasets: an equirectangular and a perspective projection dataset. The objective evaluations reveal that both algorithms significantly outperform the default configurations. In particular, with the equirectangular (perspective) projection dataset, the proposed algorithms only require 95% (23%) decoding time, stream 79% (23%) views, and improve the utility by 6% (73%) on average. The subjective evaluations confirm the proposed algorithms consume fewer resources while achieving comparable Quality of Experience (QoE) than the default and the optimal TMIV configurations. Chih-Fan Hsu, Tse-Hou Hung, Cheng-Hsin Hsu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Look at Me! Correcting Eye Gaze in Live Video CommunicationabstractAlthough live video communication is widely used, it is generally less engaging than face-to-face communication because of limitations on social, emotional, and haptic feedback. Missing eye contact is one such problem caused by the physical deviation between the screen and camera on a device. Manipulating video frames to correct eye gaze is a solution to this problem. In this article, we introduce a system to rotate the eyeball of a local participant before the video frame is sent to the remote side. It adopts a warping-based convolutional neural network to relocate pixels in eye regions. To improve visual quality, we minimize the L2 distance between the ground truths and warped eyes. We also present several newly designed loss functions to help network training. These new loss functions are designed to preserve the shape of eye structures and minimize color changes around the periphery of eye regions. To evaluate the presented network and loss functions, we objectively and subjectively compared results generated by our system and the state-of-the-art, DeepWarp, in relation to two datasets. The experimental results demonstrated the effectiveness of our system. In addition, we showed that our system can perform eye-gaze correction in real time on a consumer-level laptop. Because of the quality and efficiency of the system, gaze correction by postprocessing through this system is a feasible solution to the problem of missing eye contact in video communication. Chih-Fan Hsu, Yu-Shuen Wang, Chin-Laung Lei, Kuan-Ta Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Realizing the real-time gaze redirection system with convolutional neural networkabstractRetaining eye contact of remote users is a critical issue in video conferencing systems because of parallax caused by the physical distance between a screen and a camera. To achieve this objective, we present a real-time gaze redirection system called Flx-gaze to post-process each video frame before sending it to the remote end. Specifically, we relocate and relight the pixels representing eyes by using a convolutional neural network (CNN). To prevent visual artifacts during manipulation, we minimize not only the L2 loss function but also four novel loss functions when training the network. Two of them retain the rigidity of eyeballs and eyelids; and the other two prevent color discontinuity on the eye peripheries. By leveraging the CPU and the GPU resources, our implementation achieves real-time performance (i.e., 31 frames per second). Experimental results show that the gazes redirected by our system are of high quality under this restrict time constraint. We also conducted an objective evaluation of our system by measuring the peak signal-to-noise ratio (PSNR) between the real and the synthesized images. Chih-Fan Hsu, Yu-Shuen Wang, Chin-Laung Lei, Kuan-Ta Chen |
MMSys | 1 |
| 2018 | Active Learning for Crowdsourced QoE ModelingabstractQuality of experience (QoE) models predict the subjective quality of multimedia based on the relevant quality of service (QoS) factors. Due to the large space of QoS factors and the high costs of conducting subjective tests, efficient sampling strategies are required to determine which QoS configurations are to be queried, that is, evaluated by subjects. In this study, we extend the IQX model proposed by [M. Fiedler, T. Hoßfeld, and P. Tran-Gia, “A generic quantitative relationship between quality of experience and quality of service,”IEEE Netw., vol. 24, no. 2, pp. 36–41, Mar./Apr.2010.] toward a multidimensional QoS–QoE model (MIQX). To explore the complicated interaction between QoS factors more efficiently, we develop active learning algorithms for the multidimensional QoE model. Then, we conduct comprehensive experiments to compare the effectiveness of applying different sampling methods to crowdsourced video quality assessment tasks. In offline experiments that assume annotators give the same scores after changing the querying order, we demonstrate that active learning performs best and that a space-filling algorithm performs significantly better than random sampling. However, when we analyze the performance of the active sampling approaches more deeply using a novel field experiment, we observe that the active learning algorithms, which have been shown to be effective in the offline setting, can fail due to the habituation effect and individual differences of annotators. The active learning methods can also succeed when these issues are mitigated. These findings suggest that simply simulating the sample acquisition order, which is widely adopted in previous active learning literature[2]–[5], is not sufficient for multimedia quality assessment tasks. Haw-Shiuan Chang, Chih-Fan Hsu, Tobias Hoßfeld, Kuan-Ta Chen |
IEEE Trans. Multim. | 2 |
| 2017 | Is Foveated Rendering Perceivable in Virtual Reality?: Exploring the Efficiency and Consistency of Quality Assessment MethodsabstractFoveated rendering leverages human visual system to increase video quality under limited computing resources for Virtual Reality (VR). More specifically, it increases the frame rate and the video quality of the foveal vision via lowering the resolution of the peripheral vision. Optimizing foveated rendering systems is, however, not an easy task, because there are numerous parameters that need to be carefully chosen, such as the number of layers, the eccentricity degrees, and the resolution of the peripheral region. Furthermore, there is no standard and efficient way to evaluate the Quality of Experiment (QoE) of foveated rendering systems. In this paper, we propose a framework to compare the performance of different subjective assessment methods on foveated rendering systems. We consider two performance metrics: efficiency and consistency, using the perceptual ratio, which is the probability of the foveated rendering is perceivable by users. A regression model is proposed to model the relationship between the human perceived quality and foveated rendering parameters. Our comprehensive study and analysis reveal several insights: 1) there is no absolute superior subjective assessment method, 2) subjects need to make more observations to confirm the foveated rendering is imperceptible than perceptible, 3) subjects barely notice the foveated rendering with an eccentricity degree of 7.5 degrees+ and peripheral region of a resolution of 540p+, and 4) QoE levels are highly dependent on the individuals and scenes. Our findings are crucial for optimizing the foveated rendering systems for future VR applications. Chih-Fan Hsu, Anthony Chen, Cheng-Hsin Hsu, Chun-Ying Huang, Chin-Laung Lei, Kuan-Ta Chen |
ACM Multimedia | 1 |
| 2016 | Performance Measurements of Virtual Reality Systems: Quantifying the Timing and Positioning AccuracyabstractWe propose the very first non-intrusive measurement methodology for quantifying the performance of commodity Virtual Reality (VR) systems. Our methodology considers the VR system under test as a black-box and works with any VR applications. Multiple performance metrics on timing and positioning accuracy are considered, and detailed testbed setup and measurement steps are presented. We also apply our methodology to several VR systems in the market, and carefully analyze the experiment results. We make several observations: (i) 3D scene complexity affects the timing accuracy the most, (ii) most VR systems implement the dead reckoning algorithm, which incurs a non-trivial correction latency after incorrect predictions, and (iii) there exists an inherent trade-off between two positioning accuracy metrics: precision and sensitivity. Chun-Ming Chang, Cheng-Hsin Hsu, Chih-Fan Hsu, Kuan-Ta Chen |
ACM Multimedia | 3 |
| 2016 | Smart Beholder: An Extensible Smart Lens PlatformabstractSmart Lenses refer to detachable, orientable and zoomable lenses that stream live videos over wireless networks to heterogeneous computing devices, including tablets and smartphones. Various novel applications are made possible by smart lenses, including mobile photography, smart surveillance cameras, and Unmanned Aerial Vehicle (UAV) cameras. However, to our best knowledge, existing smart lenses are closed and proprietary, and thus we initiate an open-source project called Smart Beholder for end-to-end solutions of smart lenses. The code and documents of Smart Beholder can be found at our website http://www.smartbeholder.org. Our Smart Beholder platform are useful to researchers for fast prototyping, developers for rapid development, and amateurs for hobbies. We have implemented Smart Beholder server (camera) using a popular embedded Linux platform, called Raspberry Pi. We have also realized Smart Beholder client (controller) on various OS's, including Android. Our experimental results show the practicality and efficiency of our proposed Smart Beholder: we outperform commercial products in the market in terms of both objective and subjective metrics. We believe the release of Smart Beholder will stimulate future studies on novel multimedia applications enabled by smart lenses. Chun-Ying Huang, Ching-Ling Fan, Chih-Fan Hsu, Hsin-Yu Chang, Tsung-Han Tsai 0005, Kuan-Ta Chen, Cheng-Hsin Hsu |
ACM Multimedia | 3 |
| 2016 | Toward an Adaptive Screencast Platform: Measurement and OptimizationabstractThe binding between computing devices and displays is becoming dynamic and adaptive, and screencast technologies enable such binding over wireless networks. In this article, we design and conduct the first detailed measurement study on the performance of the state-of-the-art screencast technologies. Several commercial and one open-source screencast technologies are considered in our detailed analysis, which leads to several insights: (1) there is no single winning screencast technology, indicating room to further enhance the screencast technologies; (2) hardware video encoders significantly reduce the CPU usage at the expense of slightly higher GPU usage and end-to-end delay, and should be adopted in future screencast technologies; (3) comprehensive error resilience tools are needed as wireless communication is vulnerable to packet loss; (4) emerging video codecs designed for screen contents lead to a better Quality of Experience (QoE) of screencast; and (5) rate adaptation mechanisms are critical to avoiding degraded QoE due to network dynamics. As a case study, we propose a nonintrusive yet accurate available bandwidth estimation mechanism. Real experiments demonstrate the practicality and efficiency of our proposed solution. Our measurement methodology, open-source screencast platform, and case study allow researchers and developers to quantitatively evaluate other design considerations, which will lead to optimized screencast technologies. Chih-Fan Hsu, Ching-Ling Fan, Tsung-Han Tsai 0005, Chun-Ying Huang, Cheng-Hsin Hsu, Kuan-Ta Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2015 | Smart Beholder: An Open-Source Smart Lens for Mobile PhotographyabstractSmart lenses are detachable lenses connected to mobile devices via wireless networks, which are not constrained by the small form factor of mobile devices, and have potential to deliver better photo (video) quality. However, the viewfinder previews of smart lenses on mobile devices are difficult to optimize, due to the strict resource constraints on smart lenses and fluctuating wireless network conditions. In this paper, we design, implement, and evaluate an open-source smart lens, called Smart Beholder. It achieves three design goals: (i) cost effectiveness, (ii) low interaction latency, and (iii) high preview quality by: (i) selecting an embedded system board that is just powerful enough, (ii) minimizing per-component latency, and (iii) dynamically adapting the video coding parameters to maximizing Quality of Experience (QoE), respectively. Several optimization techniques, such as anti-drifting mechanism for video frames and QoE-driven resolution/frame rate adaptation algorithm, are proposed in this paper. Our measurement study shows that Smart Beholder outperforms Altek Cubic and Sony QX100 in terms of lower bitrate, lower latency, slightly higher frame rate, and better preview quality. We also demonstrate that sys adapts to network dynamics. Smart Beholder has been made public at http://www.smartbeholder.org as an experimental platform for researchers and developers to optimize smart lenses and other embedded real-time video streaming systems. Chun-Ying Huang, Chih-Fan Hsu, Tsung-Han Tsai 0005, Ching-Ling Fan, Cheng-Hsin Hsu, Kuan-Ta Chen |
ACM Multimedia | 2 |
| 2015 | Screencast dissected: performance measurements and design considerationsabstractDynamic and adaptive binding between computing devices and displays is increasingly more popular, and screencast technologies enable such binding over wireless networks. In this paper, we design and conduct the first detailed measurement study on the performance of the state-of-the-art screencast technologies. Several commercial and one open-source screencast technologies are considered in our detailed analysis, which leads to several insights: (i) there is no single winning screencast technology, indicating rooms to further enhance the screencast technologies, (ii) hardware video encoders significantly reduce the CPU usage at the expense of slightly higher GPU usage and end-to-end delay, and should be adopted in future screencast technologies, (iii) comprehensive error resilience tools are needed as wireless communication is vulnerable to packet loss, (iv) emerging video codecs designed for screen contents lead to better Quality of Experience (QoE) of screencast, and (v) rate adaptation mechanisms are critical to avoiding degraded QoE due to network dynamics. Furthermore, our measurement methodology and open-source screencast platform allow researchers and developers to quantitatively evaluate other design considerations, which will lead to optimized screencast technologies. Chih-Fan Hsu, Tsung-Han Tsai 0005, Chun-Ying Huang, Cheng-Hsin Hsu, Kuan-Ta Chen |
MMSys | 1 |
| 2015 | Enabling Adaptive Cloud Gaming in an Open-Source Cloud Gaming PlatformabstractWe study the problem of optimally adapting ongoing cloud gaming sessions to maximize the gamer experience in dynamic environments. The considered problem is quite challenging because: 1) gamer experience is subjective and hard to quantify; 2) the existing open-source cloud gaming platform does not support dynamic reconfigurations of video codecs; and 3) the resource allocation among concurrent gamers leaves a huge room to optimize. We rigorously address these three challenges by: 1) conducting a crowdsourced user study over the live Internet for an empirical gaming experience model; 2) enhancing the cloud gaming platform to support frame rate and bitrate adaptation on-the-fly; and 3) proposing optimal yet efficient algorithms to maximize the overall gaming experience or ensure the fairness among gamers. We conduct extensive trace-driven simulations to demonstrate the merits of our algorithms and implementation. Our simulation results show that the proposed efficient algorithms: 1) outperform the baseline algorithms by up to 46% and 30%; 2) run fast and scale to large (≤8000 gamers) problems; and 3) achieve the user-specified optimization criteria, such as maximizing average gamer experience or maximizing the minimum gamer experience. The resulting cloud gaming platform can be leveraged by many researchers, developers, and gamers. Hua-Jun Hong, Chih-Fan Hsu, Tsung-Han Tsai 0005, Chun-Ying Huang, Kuan-Ta Chen, Cheng-Hsin Hsu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Screencast in the Wild: Performance and LimitationsabstractDisplays without associated computing devices are increasingly more popular, and the binding between computing devices and displays is no longer one-to-one but more dynamic and adaptive. Screencast technologies enable such dynamic binding over ad hoc one-hop networks or Wi-Fi access points. In this paper, we design and conduct the first detailed measurement study on the performance of state-of-the-art screencast technologies. By varying the user demands and network conditions, we find that Splashtop and Miracast outperform other screencast technologies under typical setups. Our experiments also show that the screencast technologies either: (i) do not dynamically adjust bitrate or (ii) employ a suboptimal adaptation strategy. The developers of future screencast technologies are suggested to pay more attentions on the bitrate adaptation strategy, e.g., by leveraging cross-layer optimization paradigm. Chih-Fan Hsu, De-Yu Chen, Chun-Ying Huang, Cheng-Hsin Hsu, Kuan-Ta Chen |
ACM Multimedia | 1 |
| 2014 | Fast salient object detection through efficient subwindow search
Mei-Chen Yeh, Chih-Fan Hsu, Chia-Ju Lu |
Pattern Recognit. Lett. | 2 |
| 2013 | Real-time salient object detectionabstractSalient object detection techniques have a variety of multimedia applications of broad interest. However, the detection must be fast to truly aid in these processes. There exist many robust algorithms tackling the salient object detection problem but most of them are computationally demanding. In this demonstration we show a fast salient object detection system implemented in a conventional PC environment. We examine the challenges faced in the design and development of a practical system that can achieve accurate detection in real-time. Chia-Ju Lu, Chih-Fan Hsu, Mei-Chen Yeh |
ACM Multimedia | 2 |
| 2012 | Generation of Environmental Representation of a Large Indoor Parking Lot
Jung Ming Wang, Chih-Fan Hsu, Sei-Wang Chen, Chiou-Shann Fuh |
ICONIP (2) | 2 |