Yicheng Jin

dblp:55/3693 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 GaniX: Advancing GANimation through ConvNeXt features
abstract
Recent advancements in Generative Adversarial Networks (GANs) have revolutionized facial expression synthesis, pushing the boundaries of what is achievable in this field. One of the most notable breakthrough architectures is GANimation, which leverages a conditioning mechanism using images of Action Units (AUs) annotations to control the intensity of individual AUs and combine multiple AUs seamlessly. However, existing methods for generating realistic expressions are limited, often resulting in flaws and blurring in regions with intense expressions. Additionally, transitioning from one emotion to another, like grief to indignation, may lead to unwanted overlapping artifacts. In contrast, ConvNeXts have displayed exceptional performance across a wide array of vision tasks, making integrating Transformer-inspired training techniques a practical avenue for improving convolutional networks. In this paper, we reevaluate GANimation in light of ConvNeXt's design principles and present GaniX, a novel approach that significantly augments the network's capacity to extract features and diversify expressions. This enhancement not only mitigates artifacts but also makes generated facial images more lifelike and expressive. GaniX, composed entirely of standard convolutional layers, exhibits superior performance in facial expression editing, as demonstrated through experiments on widely used public datasets and real-world images.
Yicheng Jin
IJCNN1
2023 Calabash: Accelerating Attention Using a Systolic Array Chain on FPGAs
abstract
In recent years, attention mechanism has achieved remarkable performance in natural language processing and computer vision applications, at the expense of high computation cost. FPGAs have been demonstrated to be an effective hardware platform for various AI applications. However, the attention mechanism involves complex data dependency, which makes FPGA acceleration difficult. In this paper, we propose Calabash, an FPGA accelerator for attention-based applications. We design a chain of two systolic arrays, applying the same dataflow. Then, we design two scheduling techniques for different matrices to ensure the intermediate matrix can be cached in the on-chip memory. Finally, we develop analytical models for resource utilization estimation, workload balancing, and latency prediction to guide design space exploration. Experiments show that Calabash achieves 1.76 TOP/s, 1.06 TOP/s on Xilinx VU9P and ZCU102 platforms, yielding an average 50.1X and 3.94X energy-efficiency improvement compared with CPU and GPU, respectively.
Zizhang Luo, Liqiang Lu, Yicheng Jin, Liancheng Jia, Yun Liang 0001
FPL3
2022 AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction
abstract
Hardware specialization is a promising trend to sustain performance growth. Spatial hardware accelerators that employ specialized and hierarchical computation and memory resources have recently shown high performance gains for tensor applications such as deep learning, scientific computing, and data mining. To harness the power of these hardware accelerators, programmers have to use specialized instructions with certain hardware constraints. However, these hardware accelerators and instructions are quite new and there is a lack of understanding of the hardware abstraction, performance optimization space, and automatic methodologies to explore the space. Existing compilers use hand-tuned computation implementations and optimization templates, resulting in sub-optimal performance and heavy development costs.
Size Zheng 0001, Renze Chen, Anjiang Wei, Yicheng Jin, Qin Han, Liqiang Lu, Bingyang Wu, Shengen Yan, Yun Liang 0001
ISCA4
2022 An Efficient Hardware Design for Accelerating Sparse CNNs With NAS-Based Models
abstract
Deep convolutional neural networks (CNNs) have achieved remarkable performance at the cost of huge computation. As the CNN models become more complex and deeper, compressing CNNs to sparse by pruning the redundant connection in the networks has emerged as an attractive approach to reduce the amount of computation and memory requirement. On the other hand, FPGAs have been demonstrated to be an effective hardware platform to accelerate CNN inference. However, most existing FPGA accelerators focus on dense CNN models, which are inefficient when executing sparse models as most of the arithmetic operations involve addition and multiplication with zero operands. In this work, we propose an accelerator with software–hardware co-design for sparse CNNs on FPGAs. To efficiently deal with the irregular connections in the sparse convolutional layers, we propose a weight-oriented dataflow that exploits element–matrix multiplication as the key operation. Each weight is processed individually, which yields low decoding overhead. Then, we design an FPGA accelerator that features a tile look-up table (TLUT) and a channel multiplexer (CMUX). The TLUT is designed to match the index between sparse weights and input pixels. Using TLUT, the runtime decoding overhead is mitigated by using an efficient indexing operation. Moreover, we propose a weight layout to enable efficient on-chip memory access without conflicts. To cooperate with the weight layout, a CMUX is inserted to locate the address. Finally, we build a neural architecture search (NAS) engine that leverages the reconfigurability of FPGAs to generate an efficient CNN model and choose the optimal hardware design parameters. The experiments demonstrate that our accelerator can achieve 223.4-309.0 GOP/s for the modern CNNs on Xilinx ZCU102, which provides a$2.4\times $–$12.9\times $speedup over previous dense CNN accelerators on FPGAs. Our FPGA-aware NAS approach shows$2\times $speedup over MobileNetV2 with 1.5% accuracy loss.
Yun Liang 0001, Liqiang Lu, Yicheng Jin, Jiaming Xie, Ruirui Huang, Jiansong Zhang 0001, Wei Lin 0016
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 NeoFlow: A Flexible Framework for Enabling Efficient Compilation for High Performance DNN Training
abstract
Deep neural networks (DNNs) are increasingly deployed in various image recognition and natural language processing applications. The continuous demand for accuracy and high performance has led to innovations in DNN design and a proliferation of new operators. However, existing DNN training frameworks such as PyTorch and TensorFlow only support a limited range of operators and rely on hand-optimized libraries to provide efficient implementations for these operators. To evaluate novel neural networks with new operators, the programmers have to either replace the holistic new operators with existing operators or provide low-level implementations manually. Therefore, a critical requirement for DNN training frameworks is to provide high-performance implementations for the neural networks containing new operators automatically in the absence of efficient library support. In this paper, we introduce NeoFlow, which is a flexible framework for enabling efficient compilation for high-performance DNN training. NeoFlow allows the programmers to directly write customized expressions as new operators to be mapped to graph representation and low-level implementations automatically, providing both high programming productivity and high performance. First, NeoFlow provides expression-based automatic differentiation to support customized model definitions with new operators. Then, NeoFlow proposes an efficient compilation system that partitions the neural network graph into subgraphs, explores optimized schedules, and generates high-performance libraries for subgraphs automatically. Finally, NeoFlow develops an efficient runtime system to combine the compilation and training as a whole by overlapping their execution. In the experiments, we examine the numerical accuracy and performance of NeoFlow. The results show that NeoFlow can achieve similar or even better performance at the operator and whole graph level for DNNs compared to deep learning frameworks. Especially, for novel networks training, the geometric mean speedups of NeoFlow to PyTorch, TensorFlow, and CuDNN are 3.16X, 2.43X, and 1.92X, respectively.
Size Zheng 0001, Renze Chen, Yicheng Jin, Anjiang Wei, Bingyang Wu, Shengen Yan, Yun Liang 0001
IEEE Trans. Parallel Distributed Syst.3
2021 Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture
abstract
In recent years, attention-based models have achieved impressive performance in natural language processing and computer vision applications by effectively capturing contextual knowledge from the entire sequence. However, the attention mechanism inherently contains a large number of redundant connections, imposing a heavy computational burden on model deployment. To this end, sparse attention has emerged as an attractive approach to reduce the computation and memory footprint, which involves the sampled dense-dense matrix multiplication (SDDMM) and sparse-dense matrix multiplication (SpMM) at the same time, thus requiring the hardware to eliminate zero-valued operations effectively. Existing techniques based on irregular sparse patterns or regular but coarse-grained patterns lead to low hardware efficiency or less computation saving.
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li 0031, Tao Wang 0004, Yun Liang 0001
MICRO2
2009 Realistic rendering of fire scene
abstract
The realistic rendering of fire scene is one of the most challenging tasks in computer graphics. Based on GPU, the rendering of fire and smoke is studied by using particle system. The fire is treated as three parts: fireball, flame and ember, and every part is simulated by using the respective particle system. The light reflection and scattering are realized to improve the realism of smoke simulation. The method developed in the paper has been applied successfully in the scene of maritime search and rescue simulator.
Hongxiang Ren, Yicheng Jin, Lining Chen
CAD/Graphics2
2009 Sea Surface Simulation in Large Coastal Region for Maritime Simulators
abstract
A novel method for simulating large-scale, nearshore sea surface is presented. The whole simulating process can be divided into two phases: pre-computing phase and real-time computing phase. In pre-computing phase, wave parameter distribution of a simulation area is worked out using the SWAN (Simulating WAves Nearshore) model on a coarse grid. In real-time computing phase, a small height field surrounding the viewpoint is generated using the Fast Fourier Transformation (FFT) based method every frame. The wave parameters used in the FFT-based method are dynamically changed according to viewpoint’s projective position on the coarse grid. Then, the height field is saved as a vertex texture, which can be tiled in horizontal directions seamlessly. At last, the vertex texture is sampled by a sector sea surface, which is viewpoint-dependent and has Level Of Detail (LOD) effect. Several shading trips are used to generate optical effects of sea surface, such as reflection, refraction etc. Experimental results show this method can be applied to simulating large-scale sea surface in real-time with realistic effect. It is especially suitable for the application in maritime simulators.
Yongjin Li, Yicheng Jin, Helong Shen, Xinyu Zhang 0020
ICIG2
2008 3D Real-Time Visualization of Oil Spill on Sea
abstract
Instantaneous oil spill of the static point source isstudied. The trajectories of the different stages arecalculated by using the respective mathematical models. The fast Fourier transform method based on wave spectrum is developed to simulate the large-scale ocean scene. The particle system is used to implement the oil particle model. The technique of planar refraction map is adopted to render the spilled oil on the sea surface. The method developed in the paper has been successfully applied to simulate an accidental spill in the marine simulator.
Hongxiang Ren, Yicheng Jin
CW2
2006 GPU Based Real-time Shadow Research in Large Ship-handling Simulator
abstract
Shadow can strengthen the reality of 3D virtual scene and provide important visual information for the objects’ spacial relationship in the scene. This paper presents a GPU based method for all the process of shadow volume algorithm. It can achieve greater shadow performance in Ship-handling simulator than previous methods by migrating the silhouette extraction and shadow volume rending to GPU, and applying a series of techniques for culling, clipping, and simplifying shadow volume geometry. We make a series of optimization for the visual system of the ship-handling simulator such as limit the range of illumination, the alternation of z-pass and z-fail algorithm, using the ultra shadow technology and two sided stencil buffer to accelerate the rending of shadow.
Yicheng Jin
IV2