EDBT 2026 Demo / reviewers in the wild / expert
He-Yuan Lin
dblp:86/1277
· DBLP profile ↗
17ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 46% Performance modeling and evaluation · 23% Hardware accelerators and domain-specific architectures · 19% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 62% Image and video coding · 38% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
degree of parallelism |
0.1 | 1 | 2012 | Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration · IEEE Trans. Parallel Distributed Syst. 2012 |
Parallel and multicore computing › parallel algorithms
parallel algorithm design |
0.1 | 1 | 2012 | Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration · IEEE Trans. Parallel Distributed Syst. 2012 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2012 | Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration · IEEE Trans. Parallel Distributed Syst. 2012 |
Image and video processing
motion estimation |
0.1 | 1 | 2007 | Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007 |
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design |
0.1 | 1 | 2007 | Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007 |
Integrated circuit design › digital circuit design
VLSI architecture |
0.1 | 1 | 2007 | Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007 |
Image and video coding › video coding standards
H.264/AVC |
0.0 | 1 | 2007 | Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007 |
Image and video coding
video coding standards |
0.0 | 1 | 2007 | Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007 |
Methods — techniques the papers use, named apart from their topics
rank theorem · 0.1optimization theory · 0.1one-at-a-time search · 0.1linear algebra · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture CoexplorationabstractDegree of parallelism (DoP) is an essential complexity metric that characterizes the number of independent operation sets (IOSs) that can be concurrently executed within an algorithm. This paper presents a generic framework to identify IOSs and to quantify the DoP based on rank theorem in linear algebra. This framework is applied to extract algorithmic parallelisms at various granularities, namely, multigrain parallelism. Our parallelism is intrinsic and platform independent and can provide insights into architectural information, thus facilitating mapping onto generic platforms and early back annotation for modifying algorithms. It plays a significant role in the concurrent optimization of both algorithms and architectures, referred to as Algorithm/Architecture Coexploration (AAC), by trading off between the DoP and the number of operations (NoO). This paper reports three case studies for AAC. The case study on an IDCT reveals that our framework accurately quantifies the parallelism for mapping the algorithm onto generic platforms, including FPGA and multicore systems. The IDCT parallelized by our technique surpasses a conventional spectral parallelization. By exploiting fine-grain parallelism, this paper presents a better porting of a discrete wavelet transform (DWT) onto single instruction multiple data (SIMD) machines compared with a commercial compiler. A high-quality deinterlacer is implemented on a low-cost multicore platform for real-time high-definition applications by analyzing the multigrain parallelism. These case studies reveal the effectiveness of our parallel analysis framework which is applicable to generic systems. Compared with traditional graph traversal techniques, our linear algebraic approach impressively features low complexity and is practical for complicated algorithms. Gwo Giun Lee, He-Yuan Lin, Chun-Fu Chen 0001, Tsung-Yuan Huang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | A high throughput parallel AVC/H.264 context-based adaptive binary arithmetic decoderabstractIn this paper, based on the proposed parallelization scheme of binary arithmetic decoding, a parallel AVC/H.264 con text-based adaptive binary arithmetic coding (CABAC) de coder with high throughput is proposed. Following the top down design methodology, algorithm analyzing and data flow modeling in both high and low granularities are per formed to achieve the proposed architecture. According to the analysis for algorithm, the similarity between CABAC decoder and Viterbi decoder is found to extend the degree of parallelism for binary arithmetic decoding. The application of proposed design is specified to support AVC/H.264 High Profile, 4.2 Level, and 1920 × 1088 resolution at 64 frames per second. By increasing the degree of parallelism of bin decoding, the throughput of the proposed architecture is shown by the experiments to have improved 3.5 times as compared to the sequential bin decoding, and the decoded bin per second can reach 378M at clock speed 108MHz. Jia-Wei Liang, He-Yuan Lin, Gwo Giun Lee |
ICASSP | 2 |
| 2011 | Reconfigurable inverse transform architecture for multiple purpose video codingabstractIn this paper, an area efficient reconfigurable inverse transformation architecture for multiple standards is proposed. We present a top-down design methodology with complexity analysis, commonalities extraction, and dataflow modeling to systematically design reconfigurable architecture. By exporting and sharing the commonalities, the adder usage of the proposed reconfigurable inverse transform processing element can be reduced 44% compared with the total amount of adders in performing target inverse transform types. Then, the reconfigurable architecture is synthesized using TSMC 0.18 urn library. The working frequency is 108Mhz, which is derived from the dataflow scheduling. The area synthesis result is 32k gates, which indicates that the proposed design has more efficient area than other documented design in VLSI implementation. In addition, the proposed architecture also satisfies the accuracy requirement. Therefore, the proposed design have lower cost and enough flexibility for multi-standard purposes with 1920×1088 resolution and 64 frames per second and the color format is 4:2:0 for real time processing. Tsung-Yuan Huang, He-Yuan Lin, Chun-Fu Chen 0001, Gwo Giun Lee |
ISCAS | 2 |
| 2010 | Reconfigurable architecture design of motion compensation for multi-standard video codingabstractThis paper proposes a reconfigurable video decoder architecture of motion compensation for multi-standard video coding including MPEG-2, MPEG-4 and H.264. Through top-down design methodology, we analyze the motion compensation algorithm of the targeted applications and extract the commonality of motion compensation algorithms among the three different standards. To design a reconfigurable processing element to perform the integer-sample and fractional-sample interpolation operations to simultaneously support the main video standards, a regular data flow is arranged during the design space exploration. In addition, the bandwidth reduction strategies are also adopted to reduce the memory access times and power consumption of motion compensation operations for high bandwidth requirement, especially in H.264. The design implementation of the proposed architecture is synthesized using TMSC 0.18um technology library and can operate at 108HMz to achieve the real time motion compensation coding of 1920×1088 at 30 frames per second in the three video standards. Gwo Giun Lee, Wei-Chiao Yang, Min-Shan Wu, He-Yuan Lin |
ISCAS | 4 |
| 2009 | An RVC dataflow description of the AVC Constrained Baseline Profile decoderabstractVideo codec applications become more and more complex to design. To ease the description of such applications, MPEG creates a Framework called Reconfigurable Video Coding (RVC). All existing codecs in MPEG have a similar structure, they are based on a hybrid decoding structure and some of their part can be reused on other design. In RVC, the dataflow is expressed using a network of components also called Functional Units (FUs) interconnected by FIFOs. An FU, programmed in CAL Language, includes the processing and the internal states. This paper puts the focus on a parallel dataflow description of the most complex MPEG RVC decoder available called MPEG4-AVC Constrained Baseline Profile (CBP) decoder. Jérôme Gorin, Mickaël Raulet, Yuan-Long Cheng, He-Yuan Lin, Nicolas Siret, Kazuo Sugimoto, Gwo Giun Lee |
ICIP | 4 |
| 2009 | Algorithm and Architecture Design for Wide Range ELA DeinterlacerabstractIn this paper we introduce a novel algorithm that can detect local features and choose a proper interpolation method for de-interlacing. An edge is a high frequency pattern with certain direction which is a noticeable feature in video sequences. We proposed a wide range ELA (WRELA) algorithm capable of accurately detecting edge directions. The edge direction can be acquired from an optimized procedure. Finally, we can interpolate the missing pixel in the edge with the direction which has the highest correspondence. We also implement the architecture of the proposed de-interlacing algorithm by UMC 0.18 mum technology. This design is capable of real-time de-interlacing for high definition 720i sequences with the clock speed running at 54 MHz, and the gate count is acceptable. Experimental results show that our proposed de-interlacer provides not only high objective performance in terms of PSNR but also impressive visual quality especially for edges. Rong-Lai Lai, Bo-Han Chen, Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang, Yuan-Long Cheng, Jia-Wei Liang |
ISCAS | 4 |
| 2009 | Low Complexity and High Throughput VLSI Architecture for AVC/H.264 CAVLC DecodingabstractThis paper introduces a low complexity VLSI hardware architecture for entropy coding with increased throughput, based on the study of the statistical properties of the context-based adaptive variable length coding (CAVLC) in AVC/H.264. These enhanced designs are due to the results of the statistical analyses, in which better symbol length prediction was achieved by breaking the recursive dependency among codewords for multi-symbol decoder implementation. The proposed CAVLC decoder can also easily meet real-time requirements for high definition (HD) (1920times1080) applications, while the clock speed is operated only at 13 MHz under the best case scenario. Gwo Giun Lee, Chia-Cheng Lo, Yuan-Ching Chen, Sheau-Fang Lei, He-Yuan Lin, Ming-Jiun Wang |
ISCAS | 5 |
| 2009 | A Motion-compensated Spectrum-adaptive Deinterlacing AlgorithmabstractThis paper presents a motion-compensated deinterlacing algorithm featuring spectrum-adaptive interpolation of interlaced field. Using motion-compensated reference pictures, the proposed spectrum-adaptive filter tactically identifies the baseband via spectrum analysis and removes the replicas of interlaced sampling, making overall algorithm adapt to versatile video scene and different degree of motion compensation scenarios. The experimental results indicate that our proposed algorithm has better objective performance than other motion-compensated and non-motion-compensated algorithms do especially in complex moving textures. The subjective results also support the benefits of our spectrum-adaptive filter. Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Ching-Jui Hsiao |
ISCAS | 3 |
| 2009 | Rate control algorithm based on intra-picture complexity for H.264/AVCabstractAn efficient rate control algorithm based on the content-adaptive initial quantisation parameter (QP) setting scheme and the peak signal-to-noise ratio (PSNR) variation-limited bit-allocation strategy for low-complexity mobile applications is presented. This algorithm can efficiently measure the residual complexity of intra-pictures without performing the computation-intensive intra-prediction and mode decision in H.264/AVC, based on the structural and statistical features of local textures. This can adaptively set proper initial QP values for versatile video contents. In addition, this bit-allocation strategy can effectively distribute bit-rate budgets based on the monotonic property to enhance overall coding efficiency while maintaining the consistency of visual quality by limiting the variation of quantisation distortion. The experimental results reveal that the proposed algorithm surpasses the conventional rate control approaches in terms of the average PSNR from 0.34 to 0.95 dB. Moreover, this algorithm provides more impressive visual quality and more robust buffer controllability when compared with other algorithms. Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang |
IET Image Process. | 2 |
| 2008 | Textural complexity-based rate control algorithmabstractThis paper presents an efficient rate control algorithm based on our content-adaptive initial quantization parameter setting scheme for H.264/AVC. For versatile video scenes, our algorithm can adaptively set an appropriate initial QP based on the textural complexity estimated from the first picture. In addition, our bit-allocation strategy effectively distributes the bit-rate budget based on the monotonic property to enhance the coding efficiency. Our proposed algorithm surpasses JVT-H014 rate control algorithm and Cauchy-density-based bit-allocation scheme in terms of average PSNR for about 0.47 dB and 0.81 dB respectively. Besides, our algorithm provides more impressive visual quality and more robust buffer controllability. Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang |
ICME | 2 |
| 2008 | On the efficient algorithm/architecture co-exploration for complex video processingabstractTargeted for highly sophisticated visual signal processing, we introduce in this paper complexity metrics or measures of algorithms which featuring architectural information are feedback or back annotated in early design stages to facilitate concurrent exploration of both algorithmic and architectural optimizations. With application to 3D spatio-temporal motion estimation for video coding, we have demonstrated significant reduction in design cost while the algorithmic performance still surpasses recent published works and even full search under many circumstances. Moreover, we have also shown the importance and substantiality of this complexity analysis technique in the extraction of features common to various different de-interlacing algorithms adapted for versatile video content, in designing highly efficient reconfigurable video architectures. As such this novel algorithm/architecture co-exploration methodology forms the basis for dataflow models with more accurate software/hardware partitioning resulting in multi-million gate and/or instruction software simulation platforms and fast prototyping hardware platforms for the next generation electronic system level design of SoC’s. Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Ron-Lai Lai |
ICME | 3 |
| 2008 | A high-quality spatial-temporal content-adaptive deinterlacing algorithmabstractThis paper introduced a spatial-temporal content-adaptive algorithm, which can precisely select an appropriate interpolation technique for high-quality deinterlacing according to the spectral, edge-oriented and statistical features of local video content. Our algorithm employs a linear-phase statistical-adaptive vertical-temporal filter to deal with generic video scenes and adopts a modified edge-based line-averaging interpolation to efficiently recover moving edges. In addition, annoying flickering artifacts are efficiently suppressed by a flickering detection and a field-averaging filter. As a result, our algorithm outperforms other non-motion compensated methods in terms of objective PSNR and reveals more impressive subjective visual quality. Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang, Rong-Lai Lai, Chih Wen Jhuo |
ISCAS | 2 |
| 2007 | Motion Adaptive Deinterlacing via Edge Pattern RecognitionabstractIn this paper, a novel edge pattern recognition (EPR) deinterlacing algorithm with successive 4-field enhanced motion detection is introduced. The EPR algorithm surpasses the performance of ELA-based and other conventional methods especially at textural scenes. In addition, the current 4-field enhanced motion detection scheme overcomes conventional motion missing artifacts by gaining good motion detection accuracies and suppression of "motion missing" detection errors efficiently. Furthermore, with the incorporation of our new successive 4-field enhanced motion detection, the interpolation technique of EPR algorithm is capable of flexible adaptation in achieving better performance on textural scenes in generic video sequences. Gwo Giun Lee, Hsin-Te Li, Ming-Jiun Wang, He-Yuan Lin |
ISCAS | 4 |
| 2007 | Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video CodingabstractThis paper presents a new spatio-temporal motion estimation algorithm and its VLSI architecture for video coding based on algorithm and architecture co-design methodology. The algorithm consists of the new strategies of spatio-temporal motion vector prediction, modified one-at-a-time search scheme, and multiple update paths derived from optimization theory. The hardware specification is for high-definition video coding. We applied the ME algorithm to H.264 reference software. Our algorithm surpasses recently published research and achieves close performance to full search. The VLSI implementation proves the low cost feature of our algorithm. The algorithm and architecture co-design concept is highly emphasized in this paper. We provide some quantitative example to show the necessity of algorithm and architecture co-design Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Drew Wei-Chi Su, Bo-Yun Lin |
IEEE Trans. Multim. | 3 |
| 2006 | A 3D Spatio-Temporal Motion Estimation Algorithm for Video CodingabstractThis paper presents a new spatio-temporal motion estimation algorithm for video coding. The algorithm is based on optimization theory and consists of the strategies including 3D spatio-temporal motion vector prediction, modified one-at-a-time search scheme, and multiple update paths. The simulation results indicate our algorithm is better than other recently proposed ones under the same computational budget and is very close to full search. The low-cost feature and regular demand of computational resource make our algorithm suitable for VLSI implementation. The algorithm also makes single chip solution for high-definition coding feasible Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Drew Wei-Chi Su, Bo-Yun Lin |
ICME | 3 |
| 2006 | Multiresolution-based texture adaptive motion detection for de-interlacingabstractMotion-adaptive de-interlacing algorithm selects from inter-field and intra-field interpolations according to motion. Correct determination of motion information is essential for this purpose. Fine textures, having high local pixel variation, tend to cause false detection of motion. This paper proposed a texture detection mechanism utilizing multiresolution technique to improve the correctness of detection. A recursive 2-field algorithm is also proposed to reduce the memory cost in 4-field method. This algorithm provides better perceptual visual quality than other motion adaptive de-interlacing as shown by the experimental results. Gwo Giun Lee, Drew Wei-Chi Su, He-Yuan Lin, Ming-Jiun Wang |
ISCAS | 3 |
| 2006 | Model-based optimal rate control algorithm for real-time hybrid video encoderabstractThis paper presents a frame level optimal rate control scheme based on the proposed rate and distortion functions. A linear rate-quantizer model and a linear distortion-quantizer model are proposed to perform rate-distortion optimization. The coefficients of rate distortion models are estimated by linear regression with the past rate distortion characteristics. We propose the two-stage strategy for rate-distortion optimization. First we use Lagrange multiplier optimization approach to obtain the closed-form solution based on our rate distortion models. Then we take the inter-frame dependency into account to further adjust bit rate allocation. We apply our optimal rate control algorithm to H.264, and the proposed algorithm outperforms JVT-G012 rate control scheme in terms of average PSNR. He-Yuan Lin, Gwo Giun Lee, Ming-Jiun Wang, Drew Wei-Chi Su, Bo-Yun Lin |
ISCAS | 1 |