Yilei Li

dblp:40/10850 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2025 Garment3DGen: 3D Garment Stylization and Texture Generation
abstract
We introduce Garment3DGen a new method to synthesize 3D garment assets from a base mesh given a single input image as guidance. Our proposed approach allows users to generate 3D textured clothes based on both real and synthetic images, such as those generated by text prompts. The generated assets can be directly draped and simulated on human bodies. We leverage the recent progress of image-to-3D diffusion methods to generate 3D garment geometries. However, since these geometries cannot be utilized directly for downstream tasks, we propose to use them as pseudo ground-truth and set up a mesh deformation optimization procedure that deforms a base template mesh to match the generated 3D target. Carefully designed losses allow the base mesh to freely deform towards the desired target, yet preserve mesh quality and topology such that they can be simulated. Finally, we generate high-fidelity texture maps that are globally and locally consistent and faithfully capture the input guidance, allowing us to render the generated 3D assets. With Garment3DGen users can generate the simulation-ready 3D garment of their choice without the need of artist intervention. We present a plethora of quantitative and qualitative
Nikolaos Sarafianos, Tuur Stuyck, Xiaoyu Xiang, Yilei Li, Jovan Popovic
3DV4
2025 SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein Identity
abstract
Score distillation has emerged as one of the most prevalent approaches for text-to-3D asset synthesis. Essentially, score distillation updates 3D parameters by lifting and back-propagating scores averaged over different views. In this paper, we reveal that the gradient estimation in score distillation is inherent to high variance. Through the lens of variance reduction, the effectiveness of SDS and VSD can be interpreted as applications of various control variates to the Monte Carlo estimator of the distilled score. Motivated by this rethinking and based on Stein’s identity, we propose a more general solution to reduce variance for score distillation, termed \emph{Stein Score Distillation (SSD)}. SSD incorporates control variates constructed by Stein identity, allowing for arbitrary baseline functions. This enables us to include flexible guidance priors and network architectures to explicitly optimize for variance reduction. In our experiments, the overall pipeline, dubbed \emph{SteinDreamer}, is implemented by instantiating the control variate with a monocular depth estimator. The results show that SSD can effectively reduce the distillation variance and consistently improve visual quality for both object- and scene-level generation.
Peihao Wang, Zhiwen Fan, Dejia Xu, Dilin Wang, Sreyas Mohan, Forrest N. Iandola, Yilei Li, Qiang Liu 0001, Zhangyang Wang, Vikas Chandra
AISTATS8
2025 Make-A-Texture: Fast Shape-Aware Texture Generation in 3 Seconds
abstract
We present Make-A-Texture, a new framework that efficiently synthesizes high-resolution texture maps from textual prompts for given 3D geometries. Our approach progressively generates textures that are consistent across multiple viewpoints with a depth-aware inpainting diffusion model, in an optimized sequence of viewpoints determined by an automatic view selection algorithm. A significant feature of our method is its remarkable efficiency, achieving a full texture generation within an end-to-end runtime of just 3.07 seconds on a single NVIDIA H100 GPU, significantly outperforming existing methods. Such an acceleration is achieved by optimizations in the diffusion model and a specialized backprojection method. Moreover, our method reduces the artifacts in the backprojection phase, by selectively masking out non-frontal faces, and internal faces of open-surfaced objects. Experimental results demonstrate that Make-A-Texture matches or exceeds the quality of other state-of-the-art methods. Our work significantly improves the applicability and practicality of texture generation models for real-world 3D content creation, including interactive creation and text-guided texture editing.
Xiaoyu Xiang, Liat Sless Gorelik, Yuchen Fan 0001, Omri Armstrong, Forrest N. Iandola, Yilei Li, Ita Lifshitz
WACV6
2024 Taming Mode Collapse in Score Distillation for Text-to-3D Generation
abstract
Despite the remarkable performance of score distillation in text-to-3D generation, such techniques notoriously suf-fer from view inconsistency issues, also known as “Janus” artifact, where the generated objects fake each view with multiple front faces. Although empirically effective methods have approached this problem via score debiasing or prompt engineering, a more rigorous perspective to explain and tackle this problem remains elusive. In this paper, we reveal that the existing score distillation-based text-to-3D generation frameworks degenerate to maximal likelihood seeking on each view independently and thus suffer from the mode collapse problem, manifesting as the Janus artifact in practice. To tame mode collapse, we improve score distillation by re-establishing the entropy term in the corresponding variational objective, which is applied to the distribution of rendered images. Maximizing the entropy encourages diversity among different views in generated 3D assets, thereby mitigating the Janus problem. Based on this new objective, we derive a new update rule for 3D score distillation, dubbed Entropic Score Distillation (ESD). We theoretically reveal that ESD can be simplified and implemented by just adopting the classifier-free guidance trick upon variational score distillation. Although embarrassingly straightforward, our extensive experiments demonstrate that ESD can be an effective treatment for Janus artifacts in score distillation.
Peihao Wang, Dejia Xu, Zhiwen Fan, Dilin Wang, Sreyas Mohan, Forrest N. Iandola, Yilei Li, Qiang Liu 0001, Zhangyang Wang, Vikas Chandra
CVPR8
2024 WaSt-3D: Wasserstein-2 Distance for Scene-to-Scene Stylization on 3D Gaussians
Dmytro Kotovenko, Olga Grebenkova, Nikolaos Sarafianos, Avinash Paliwal, Pingchuan Ma 0006, Omid Poursaeed, Sreyas Mohan, Yuchen Fan 0001, Yilei Li, Björn Ommer
ECCV (21)9
2024 Semi-supervised fault diagnosis of wheelset bearings in high-speed trains using autocorrelation and improved flow Gaussian mixture model
Jiayi Wu 0006, Yilei Li, Limin Jia 0002, Guoping An, Yan-Fu Li, Jérôme Antoni, Ge Xin
Eng. Appl. Artif. Intell.2
2023 MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition
abstract
In this paper, we study a novel problem in egocentric action recognition, which we term as “Multimodal Generalization“ (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely missing. We thoroughly investigate MMG in the context of standard supervised action recognition and the more challenging few-shot setting for learning new action categories. MMG consists of two novel scenarios, designed to support security, and efficiency considerations in real-world applications: (1) missing modality generalization where some modalities that were present during the train time are missing during the inference time, and (2) cross-modal zero-shot generalization, where the modalities present during the inference time and the training time are disjoint. To enable this investigation, we construct a new dataset MMG-Ego4D containing data points with video, audio, and inertial motion sensor (IMU) modalities. Our dataset is derived from Ego4D [27] dataset, but processed and thoroughly re-annotated by human experts to facilitate research in the MMG problem. We evaluate a diverse array of models on MMG-Ego4D and propose new methods with improved generalization ability. In particular, we introduce a new fusion module with modality dropout training, contrastive-based alignment training, and a novel cross-modal prototypical loss for better few-shot performance. We hope this study will serve as a benchmark and guide future research in multimodal generalization problems. The benchmark and code are available at https://github.com/facebookresearch/MMG_Ego4D
Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang
CVPR5
2021 Improving Efficiency in Neural Network Accelerator using Operands Hamming Distance Optimization
abstract
Neural network accelerator is a key enabler for the on-device AI inference, for which energy efficiency is an important metric. The datapath energy, including the computation energy and the data movement energy among the arithmetic units, claims a significant part of the total accelerator energy. By revisiting the basic physics of the arithmetic logic circuits, we show that the datapath energy is highly correlated with the bit flips when streaming the input operands into the arithmetic units, defined as the hamming distance (HD) of the input operand matrices. Based on the insight, we propose a post-training optimization algorithm and a HD-aware training algorithm to co-design and co-optimize the accelerator and the network synergistically. The experimental results based on post-layout simulation with MobileNetV2 demonstrate on average 2.85x datapath energy reduction and up to 8.51x datapath energy reduction for certain layers.
Meng Li 0004, Yilei Li, Vikas Chandra
ASP-DAC2
2021 PyTorchVideo: A Deep Learning Library for Video Understanding
abstract
We introduce PyTorchVideo, an open-source deep-learning library that provides a rich set of modular, efficient, and reproducible components for a variety of video understanding tasks, including classification, detection, self-supervised learning, and low-level processing. The library covers a full stack of video understanding tools including multimodal data loading, transformations, and models that reproduce state-of-the-art performance. PyTorchVideo further supports hardware acceleration that enables real-time inference on mobile devices. The library is based on PyTorch and can be used by any training framework; for example, PyTorchLightning, PySlowFast, or Classy Vision. PyTorchVideo is available at https://pytorchvideo.org/.
Haoqi Fan 0001, Tullie Murrell, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Nikhila Ravi, Meng Li 0004, Haichuan Yang, Jitendra Malik, Ross B. Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, Christoph Feichtenhofer
ACM Multimedia6
2020 Difference-Frequency Ultrasound Imaging With Non-Linear Contrast
abstract
Conventional ultrasound imaging is based on the scattering of sound from inhomogeneities in the density and the speed of sound and is often used in medicine to resolve pathologic compared to normal tissue. Here we demonstrate a difference-frequency ultrasound (dfUS) imaging method that is based on the interaction of two sound pulses that propagate non-collinearly and intersect in space and time. The dfUS signal arises primarily from the second-order non-linear coefficient, a contrast mechanism that differs from linear and harmonic US imaging. The distinct contrast mechanism allows dfUS to image anatomic features that are not identifiable in conventional US images of salmon and pig kidney tissue. Further, dfUS produces enhanced contrast of glioblastoma tumor implanted in the mouse brain, revealing its potential for improving medical diagnosis. Progress towards a real-time system is discussed.
Yilei Li, Dina Polyak, Eli Johnson, Derek W. Yecies, Saba Shevidi, Adam de la Zerda, Melanie Hayden Gephart, Steven Chu
IEEE Trans. Medical Imaging1
2019 A 7.5-mW 10-Gb/s 16-QAM wireline transceiver with carrier synchronization and threshold calibration for mobile inter-chip communications in 16-nm FinFET
abstract
A compact energy-efficient 16-QAM wireline transceiver with carrier synchronization and threshold calibration is proposed to leverage high-density fine-pitch interconnects. Utilizing frequency-division multiplexing, the transceiver transfers four-bit data through one RF band to reduce intersymbol interferences. A forwarded clock is also transmitted through the same interconnect with the data simultaneously to enable low-power PVT-insensitive symbol clock recovery. A carrier synchronization algorithm is proposed to overcome nontrivial current and phase mismatches by including DC offset calibration and dedicated I/Q phase adjustments. Along with this carrier synchronization, a threshold calibration process is used for the transceiver to tolerate channel and circuit variations. The transceiver implemented in 16-nm FinFET occupies only 0.006-mm2 and achieves 10 Gb/s with 0.75-pJ/bit efficiency and <2.5-ns latency.
Jieqiong Du, Chien-Heng Wong, Yo-Hao Tu, Wei-Han Cho, Yilei Li, Yuan Du, Po-Tsang Huang, Sheau Jiung Lee, Mau-Chung Frank Chang
NOCS5
2019 Optimization of the Trade-Off Between Speckle Reduction and Axial Resolution in Frequency Compounding
abstract
We measured the reduction of speckle by frequency compounding using Gaussian pulses, which have the least time-bandwidth product. The experimental results obtained from a tissue mimicking phantom agree quantitatively with numerical simulations of randomly distributed point scatterers. For a fixed axial resolution, the amount of speckle reduction is found to approach a maximum as the number of bands increases while the total spectral range that they cover is kept constant. An analytical solution of the maximal speckle reduction is derived and shows that the maximum improves approximately as the inverse square root of the Gaussian pulse bandwidth. Since the axial resolution is proportional to the inverse of the pulse bandwidth, an optimized trade-off between speckle reduction and axial resolution is obtained. Considerations for the applications of the optimized trade-off are discussed.
Yilei Li, Yonatan Winetraub, Orly Liba, Adam de la Zerda, Steven Chu
IEEE Trans. Medical Imaging1
2018 A Single Layer 3-D Touch Sensing System for Mobile Devices Application
abstract
Touch sensing has been widely implemented as a main methodology to bridge human and machine interactions. The traditional touch sensing range is 2-D and therefore limits the user experience. To overcome these limitations, we propose a novel 3-D contactless touch sensing called Airtouch system, which improves user experience by remotely detecting single/multi-finger position. A single layer touch panel with triangle-shaped electrodes is proposed to achieve multitouch detection capability as well as manufacturing cost reduction. Moreover, an oscillator-based-capacitive touch sensing circuit is implemented as the sensing hardware with the bootstrapping technique to eliminate the interchannel coupling effects. To further improve the system accuracy, a grouping algorithm is proposed to group the useful channels' data and filter out hardware noise impact. Finally, improved algorithms are proposed to eliminate the fringing capacitance effect and achieve accurate finger position estimation. EM simulation proved that the proposed algorithm reduced the maximum systematic error by 11 dB in the horizontal position detection. The proposed system consumes 2.3 mW and is fully compatible with existing mobile device environments. A prototype is built to demonstrate that the system can successfully detect finger movement in a vertical direction up to 6 cm and achieve a horizontal resolution up to 0.6 cm at 1 cm finger-height. As a new interface for human and machine interactions, this system offers great potential in finger movement detection and gesture recognition for small-sized electronics and advanced human interactive games for mobile device.
Yan Zhang 0050, Yilei Li, Yuan Du, Yen-Cheng Kuan, Mau-Chung Frank Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 A Novel Fully Synthesizable All-Digital RF Transmitter for IoT Applications
abstract
In this paper, a fully synthesizable all-digital transmitter (ADTX) is first proposed. This transmitter (TX) uses Cartesian architecture and supports wide-band quadratic-amplitude modulation with wide carrier frequency range. Furthermore, the design methodology for ADTX and corresponding bandpass filter is discussed. This TX is synthesized with digital register transfer level-graphic database system flow, and can be easily implemented in any standard CMOS technology. An exemplary TX is synthesized by TSMC 28-nm standard cell library with extremely small area (0.0009 mm2) and supports carrier frequency as high as 6 GHz with excellent error vector magnitude (<;-30 dB). To the best of the authors' knowledge, this is the first work on a fully synthesizable design of RF transistors, allowing easy technology migration and portability.
Yilei Li, Kirti Dhwaj, Chien-Heng Wong, Yuan Du, Yiwu Tang, Yiyu Shi 0001, Tatsuo Itoh, Mau-Chung Frank Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Invited - Airtouch: a novel single layer 3D touch sensing system for human/mobile devices interactions
abstract
Touchscreen technology plays an important role in the booming mobile devices market. Traditional touchscreen only provides 2D interactions with limited user experience. To overcome these limitations, we propose a novel 3D touch sensing system called the Airtouch system, which can recognize the movement of the finger in a 3-dimensional space. Half of the manufacturing cost is reduced by applying only single layer electrodes in the touch panel design. Moreover, an oscillator based correlated double sampling circuit is implemented as the self-capacitive sensor with bootstrapping technique to reduce inter-channel-coupling effect. Additionally, new algorithm for finger positioning is created with grouping filter invented to reduce system background noise. The demonstrated setup can successfully detect finger movement within a vertical range of 6cm and achieve a horizontal resolution up to 1cm. This system offers great potential in both gesture recognition for small-sized electronics, and advanced human interactive games for TV and mobile device.
Adrian Tang 0002, Yan Zhang 0050, Yilei Li, Kye Cheung, Mau-Chung Frank Chang
DAC5
2016 Invited - A 2.2 GHz SRAM with high temperature variation immunity for deep learning application under 28nm
abstract
With the coming era of Big Data, hardware implementation of machine learning has become attractive for many applications, such as real-time object recognition and face recognition. The implementation of machine learning algorithms needs intensive memory access, and SRAM is critical for the overall performance. This paper proposes a new design of high speed SRAM for machine learning purposes. With fast access time (cycle time: 650 ps, access time: 350 ps), low sensitivity to temperature variation and high configurability (less than 10% performance difference between 125_rcw_tt vs 0_rcw_tt), the proposed SRAM is a better candidate for hardware machine learning system than the conventional SRAM. Compared with Samsung HL 152, our design has smaller size (121×43 um2 vs 127×44 um2) with half the number of pins ports (12 vs 25) and higher speed (2.2GHz vs 0.8GHz).
Yen-Hsiang Wang, Yilei Li, Chien-Heng Wong, Tien Pei Chou, Young-Kai Chen, Mau-Chung Frank Chang
DAC3
2012 A triple-band flexible low-noise transmitter with linearity enhancement
abstract
A low-noise triple band transmitter for GSM triple-band and WCDMA is presented. By programming parameters, analog baseband and RF frontend can handle signals of different protocols and frequency bands. Novel linearization method is used in LPF and driver amplifier to meet the demand of 3G systems. Noise optimization is also adopted so that our transmitter achieves −157 dBc/Hz noise at 190 MHz frequency offset in WCDMA band, which eliminates external SAW filter.
Yilei Li, Chuansheng Dong, Kefeng Han, Yongchang Yu, Xi Tan 0001, Na Yan 0004, Hao Min
ISCAS1
2012 Cognitive model based fashion style decision making
Jiyun Li, Yilei Li
Expert Syst. Appl.2