VLDB 2026 Research / reviewers in the wild / expert
Yi-Min Tsai
dblp:12/3621
· DBLP profile ↗
17ranked-venue papers
4as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 50% 3D vision · 50% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
implicit neural representation |
0.7 | 1 | 2023 | Cascaded Local Implicit Transformer for Arbitrary-Scale Super-Resolution · CVPR 2023 |
Machine learning › Generative modeling › image reconstruction
super-resolution |
0.7 | 1 | 2023 | Cascaded Local Implicit Transformer for Arbitrary-Scale Super-Resolution · CVPR 2023 |
Image and video processing › super-resolution
image super-resolution |
0.4 | 1 | 2020 | Unified Dynamic Convolutional Network for Super-Resolution With Variational Degradations · CVPR 2020 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.7frequency encoding · 0.7attention · 0.7dynamic convolution · 0.4convolutional neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Cascaded Local Implicit Transformer for Arbitrary-Scale Super-ResolutionabstractImplicit neural representation has recently shown a promising ability in representing images with arbitrary resolutions. In this paper, we present a Local Implicit Transformer (LIT), which integrates the attention mechanism and frequency encoding technique into a local implicit image function. We design a cross-scale local attention block to effectively aggregate local features and a local frequency encoding block to combine positional encoding with Fourier domain information for constructing high-resolution images. To further improve representative power, we propose a Cascaded LIT (CLIT) that exploits multi-scale features, along with a cumulative training strategy that gradually increases the upsampling scales during training. We have conducted extensive experiments to validate the effectiveness of these components and analyze various training strategies. The qualitative and quantitative results demonstrate that LIT and CLIT achieve favorable results and outperform the prior works in arbitrary super-resolution tasks. Hao-Wei Chen, Yu-Syuan Xu, Min-Fong Hong, Yi-Min Tsai, Hsien-Kai Kuo, Chun-Yi Lee |
CVPR | 4 |
| 2022 | Self-Supervised Robustifying Guidance for Monocular 3D Face Reconstruction
Hitika Tiwari, Min-Hung Chen, Yi-Min Tsai, Hsien-Kai Kuo, Hung-Jen Chen 0004, Kevin Jou, K. S. Venkatesh, Yong-Sheng Chen |
BMVC | 3 |
| 2021 | Learning to Compensate: A Deep Neural Network Framework for 5G Power Amplifier CompensationabstractOwing to the complicated characteristics of 5G communication system, designing RF components through math-ematical modeling becomes a challenging obstacle. Moreover, such mathematical models need numerous manual adjustments for various specification requirements. In this paper, we present a learning-based framework to model and compensate Power Amplifiers (PAs) in 5G communication. In the proposed frame-work, Deep Neural Networks (DNNs) are used to learn the characteristics of the PAs, while, correspondent Digital Pre-Distortions (DPDs) are also learned to compensate for the nonlinear and memory effects of PAs. On top of the framework, we further propose two frequency domain losses to guide the learning process to better optimize the target, compared to naive time domain Mean Square Error (MSE). The proposed framework serves as a drop-in replacement for the conventional approach. The proposed approach achieves an average of 56.7% reduction of nonlinear and memory effects, which converts to an average of 16.3% improvement over a carefully-designed mathematical model, and even reaches 34% enhancement in severe distortion scenarios. Yi-Min Tsai, Hsien-Kai Kuo, Hantao Huang, Hsin-Hung Chen, Sheng-Hong Yan, Wei-Lun Ou, Chia-Ming Cheng |
ICC | 3 |
| 2020 | Unified Dynamic Convolutional Network for Super-Resolution With Variational DegradationsabstractDeep Convolutional Neural Networks (CNNs) have achieved remarkable results on Single Image Super-Resolution (SISR). Despite considering only a single degradation, recent studies also include multiple degrading effects to better reflect real-world cases. However, most of the works assume a fixed combination of degrading effects, or even train an individual network for different combinations. Instead, a more practical approach is to train a single network for wide-ranging and variational degradations. To fulfill this requirement, this paper proposes a unified network to accommodate the variations from inter-image (cross-image variations) and intra-image (spatial variations). Different from the existing works, we incorporate dynamic convolution which is a far more flexible alternative to handle different variations. In SISR with non-blind setting, our Unified Dynamic Convolutional Network for Variational Degradations (UDVD) is evaluated on both synthetic and real images with an extensive set of variations. The qualitative results demonstrate the effectiveness of UDVD over various existing works. Extensive experiments show that our UDVD achieves favorable or comparable performance on both synthetic and real images. Yu-Syuan Xu, Shou-Yao Roy Tseng, Yu Tseng, Hsien-Kai Kuo, Yi-Min Tsai |
CVPR | 5 |
| 2019 | Architecture-Aware Network Pruning for Vision Quality ApplicationsabstractConvolutional neural network (CNN) delivers impressive achievements in computer vision and machine learning field. However, CNN incurs high computational complexity, especially for vision quality applications because of large image resolution. In this paper, we propose an iterative architecture-aware pruning algorithm with adaptive magnitude threshold while cooperating with quality-metric measurement simultaneously. We show the performance improvement applied on vision quality applications and provide comprehensive analysis with flexible pruning configuration. With the proposed method, the Multiply-Accumulate (MAC) of state-of-the-art low-light imaging (SID) and super-resolution (EDSR) are reduced by 58% and 37% without quality drop, respectively. The memory bandwidth (BW) requirements of convolutional layer can be also reduced by 20% to 40%. Wei-Ting Wang, Wei-Shiang Lin, Cheng-Ming Chiang, Yi-Min Tsai |
ICIP | 5 |
| 2012 | Statistical screening for IC Trojan detectionabstractWe present statistical screening of test vectors for detecting a Trojan, malicious circuitry hidden inside an integrated circuit (IC). When applied a test vector, a Trojan-embedded chip draws extra leakage current that is unfortunately too small for the detector in most cases and concealed by process variation related to chip fabrication. To remedy the problem, we formulate a statistical approach that can screen and select test vectors in detecting Trojans. We validate our approach analytically and with gate-level simulations and show that our screening method leads to a substantial reduction in false positives and false negatives when detecting IC Trojans of various sizes. Youngjune Gwon, H. T. Kung 0001, Dario Vlah, Keng-Yen Huang, Yi-Min Tsai |
ISCAS | 5 |
| 2012 | A high speed feature matching architecture for real-time video stabilizationabstractAn efficient feature matching architecture targets at real-time video stabilization is revealed in this paper. For some applications, such as vehicular application, real-time video stabilization is needed to provide instant stable video input. However, feature matching is usually the bottleneck to achieve high performance. High speed feature matching architecture is proposed to accelerate the performance of video stabilization. Locality sensitive hashing (LSH) helps us realize the feature matching procedure in hardware implementation. By applying the proposed dynamic table allocation and on-chip cache mechanism, this work achieves 422K queries/s and real-time feature matching with 90% in memory reduction and more than 50% in relieving the feature bus burden of the system. Keng-Yen Huang, Yi-Min Tsai, Tien-Ju Yang, Liang-Gee Chen |
ISCAS | 2 |
| 2012 | WarmL1: A warm-start homotopy-based reconstruction algorithm for sparse signalsabstractA sparse signal can be reconstructed from a small amount of random and linear measurements by solving a system of underdetermined equations. In this paper, we study the reconstruction problem while the system undergoes dynamic modifications. Resolving this problem from scratch requires high computational efforts. Therefore, we propose an efficient homotopy-based reconstruction algorithm with warmstart, named WarmL1. WarmL1 quickly updates the previous solution to the desired one. Based on the concept of homotopy, WarmL1 breaks the reconstruction procedure into simple steps, and solves the problem iteratively. Four possible applications are presented and discussed to demonstrate the usage of WarmL1 for different warm-start situations. Experiments on these applications are performed. The results show that WarmL1 achieves 3.2× to 37.5× speeding up or up to 1/5100 l2-error at the same computational cost compared to related works. Tien-Ju Yang, Yi-Min Tsai, Chung-Te Li, Liang-Gee Chen |
ISIT | 2 |
| 2012 | Visual Vocabulary Processor Based on Binary Tree Architecture for Real-Time Object Recognition in Full-HD ResolutionabstractFeature matching is an indispensable process for object recognition, which is an important issue for wearable devices with video analysis functionalities. To implement a low-power SoC for object recognition, the proposed visual vocabulary processor (VVP) is employed to accelerate the speed of feature matching. The VVP can transform hundreds of 128-D SIFT vectors into a 64-D histogram for object matching by using the binary-tree-based architecture, and 16 calculators for the computations of the Euclidean distances are designed for each of the two processors in each level. A total of 126 visual words can be saved in the six-level hierarchical memory, which instantly offers the data required for the matching process, and more than 5 times of bandwidth can be saved compared with the non-binary-tree-based architecture. As a part of the recognition SoC, the VVP is implemented with the 65-nm CMOS technology, and the experimental results show that the gate count and the average power consumption are 280 K and 5.6 mW, respectively. Tse-Wei Chen 0001, Yu-Chi Su, Keng-Yen Huang, Yi-Min Tsai, Shao-Yi Chien, Liang-Gee Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Algorithm and implementation of multi-channel spike sorting using GPU in a home-care surveillance systemabstractIntensive home-care surveillance programs are associated with a marked decrease in the need for hospitalization. They can improve the functional statuses of elderly patients with severe congestive diseases. The GPU-based home-care surveillance system is effective and has a major impact on health expenditure than traditional surveillance equipments. In this work, we propose a spike sorting technique as a specific case for the GPU-based home surveillance system. Spike sorting is the procedure of classifying spikes corresponding to the firing neurons. In neuroscience research, spike sorting is adopted to analyze neural activities, brain functions and sensation. It is also a key component in cortically-controlled neuro-prosthetics for patients. In order to efficiently distinguish different neural spike activities, a robust spike sorting algorithm is required for above applications. To improve accuracy, multi-channel spike sorting is necessary. In addition, real-time monitoring for a home-care system is required. Therefore, we exploit a CUDA implementation using GPU for acceleration. Yun-Yu Chen, Yi-Min Tsai, Liang-Gee Chen |
ICME | 2 |
| 2011 | Algorithm and architecture design of a knowledge-based vehicle tracking for intelligent cruise controlabstractThe paper exploits a vision-based intelligent vehicle cruise control system from the application level to the architecture level. Firstly, design considerations of the system are addressed in both computing power and accuracy aspects. Secondly, we present an efficient knowledge-based front-vehicle tracking algorithm. The algorithm yields below 5% error rate that outperforms the state-of-the-arts. Thirdly, a run-length-based algorithm optimization flow is introduced. Finally, specific hardware architecture is developed. It achieves 1280×960/80FPS and 4096×2160/10FPS requirements for multi-vehicle tracking tasks. Yi-Min Tsai, Chih-Chung Tsai, Keng-Yen Huang, Liang-Gee Chen |
ICME | 1 |
| 2011 | Smart display: A mobile self-adaptive projector-camera systemabstractOwing to the diversity of projection surfaces, an effective mobile display system must be adaptive to the surface to avoid introducing a clipped scene. In this paper, we propose a smart mobile display system which automatically adapts to the location and motion of a surface. Firstly, an imperceptible structured light technique is adopted and continuous adaptation is accomplished. Secondly, a specifically designed code image with high distortion tolerance is proposed. Thirdly, we present a priority-based correction method to revise previous decoding results. Finally, a matching procedure resisting noise interruption is introduced. The system achieves 95% correct rate under the common indoor illuminance. In addition, the system performance is independent of the projected content and the surface shapes. Tien-Ju Yang, Yi-Min Tsai, Liang-Gee Chen |
ICME | 2 |
| 2010 | Video stabilization for vehicular applications using SURF-like descriptor and KD-treeabstractThis paper describes a method to stabilize video for vehicular applications based on feature analysis. An investigation on camera motion model is conducted. Harris features are extracted under the proposed resolution adaptation scheme. Besides, features are described with SURF-like descriptor. For feature matching, KD-tree with best-bin-first search significantly reduces the matching time. A damping filer is utilized to model and predict the unwanted oscillation. 93.1% correct rate in average is achieved in divergent driving conditions. Only 0.114 second is required to process a frame at resolution 1280×960. The provided benchmark shows outperformance of the proposed method. Keng-Yen Huang, Yi-Min Tsai, Chih-Chung Tsai, Liang-Gee Chen |
ICIP | 2 |
| 2010 | An exploration of on-road vehicle detection using hierarchical scaling schemesabstractThis paper targets at detecting preceding vehicles in a wide range of distance. We propose an Adaboost-based approach combined with hierarchical image and sub-window scaling schemes. The relationship is investigated among object characteristics, image structures and image scales. A parameter set is developed to easily adjust overall performance, which benefits researchers to establish a vehicle detection system. It achieves 96.6% detection rate with 2.0% false alarm rate along proposed methodology. The benchmark of several learning-based vehicle detection approaches is also provided. The results show the outperformance of the proposed method. Yi-Min Tsai, Keng-Yen Huang, Chih-Chung Tsai, Liang-Gee Chen |
ICIP | 1 |
| 2010 | Learning-Based Vehicle Detection Using Up-Scaling Schemes and Predictive Frame Pipeline StructuresabstractThis paper aims at detecting preceding vehicles in a variety of distance. A sub-region up-scaling scheme significantly raises far distance detection capability. Three frame pipeline structures involving object predictors are explored to further enhance accuracy and efficiency. It claims a 140-meter detecting distance along proposed methodology. 97.1% detection rate with 4.2% false alarm rate is achieved. At last, the benchmark of several learning-based vehicle detection approaches is provided. Yi-Min Tsai, Keng-Yen Huang, Chih-Chung Tsai, Liang-Gee Chen |
ICPR | 1 |
| 2008 | A real-time augmented view synthesis system for transparent car pillarsabstractIn this paper, a real-time augmented view synthesis system is proposed. With real-time consideration and augmented reality property, the proposed system provides a novel application for making car pillars transparent to enlarge the eyesight of the drivers. Thanks to the proposed trinocular depth estimation, online depth generation becomes possible through trinocular fast dense disparity estimation. With the proposed texture mapping free viewpoint depth image based rendering on the GPU, the processing speed for view interpolation achieves real-time. The computational power doesn’t cost much so that this real-time system is achievable on common computers. The experimental results show that the proposed system is able to project view synthesized video which is integrated with the background to make the user perceive seamless outside scene inside their car. Yu-Lin Chang, Yi-Min Tsai, Liang-Gee Chen |
ICIP | 2 |
| 2007 | Symmetric trinocular dense disparity estimation for car surrounding camera arrayabstractThis paper presented a novel dense disparity estimation method which is called as symmetric trinocular dense disparity estimation. Also a car surrounding camera array application is proposed to improve the driving safety by the proposed symmetric trinocular dense disparity estimation algorithm. The symmetric trinocular property is conducted to show the benefit of doing disparity estimation with three cameras. A 1D fast search algorithm is described to speed up the slowness of the original full search algorithms. And the 1D fast search algorithm utilizes the horizontal displacement property of the cameras to further check the correctness of the disparity vector. The experimental results show that the symmetric trinocular property improves the quality and smoothness of the disparity vector. Yi-Min Tsai, Yu-Lin Chang, Liang-Gee Chen |
VCIP | 1 |