VLDB 2026 Research / reviewers in the wild / expert
Yilei Li
dblp:40/10850
· DBLP profile ↗
18ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Garment3DGen: 3D Garment Stylization and Texture GenerationabstractWe introduce Garment3DGen a new method to synthesize 3D garment assets from a base mesh given a single input image as guidance. Our proposed approach allows users to generate 3D textured clothes based on both real and synthetic images, such as those generated by text prompts. The generated assets can be directly draped and simulated on human bodies. We leverage the recent progress of image-to-3D diffusion methods to generate 3D garment geometries. However, since these geometries cannot be utilized directly for downstream tasks, we propose to use them as pseudo ground-truth and set up a mesh deformation optimization procedure that deforms a base template mesh to match the generated 3D target. Carefully designed losses allow the base mesh to freely deform towards the desired target, yet preserve mesh quality and topology such that they can be simulated. Finally, we generate high-fidelity texture maps that are globally and locally consistent and faithfully capture the input guidance, allowing us to render the generated 3D assets. With Garment3DGen users can generate the simulation-ready 3D garment of their choice without the need of artist intervention. We present a plethora of quantitative and qualitative Nikolaos Sarafianos, Tuur Stuyck, Xiaoyu Xiang, Yilei Li, Jovan Popovic |
3DV | 4 |
| 2025 | SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein IdentityabstractScore distillation has emerged as one of the most prevalent approaches for text-to-3D asset synthesis. Essentially, score distillation updates 3D parameters by lifting and back-propagating scores averaged over different views. In this paper, we reveal that the gradient estimation in score distillation is inherent to high variance. Through the lens of variance reduction, the effectiveness of SDS and VSD can be interpreted as applications of various control variates to the Monte Carlo estimator of the distilled score. Motivated by this rethinking and based on Stein’s identity, we propose a more general solution to reduce variance for score distillation, termed \emph{Stein Score Distillation (SSD)}. SSD incorporates control variates constructed by Stein identity, allowing for arbitrary baseline functions. This enables us to include flexible guidance priors and network architectures to explicitly optimize for variance reduction. In our experiments, the overall pipeline, dubbed \emph{SteinDreamer}, is implemented by instantiating the control variate with a monocular depth estimator. The results show that SSD can effectively reduce the distillation variance and consistently improve visual quality for both object- and scene-level generation. Peihao Wang, Zhiwen Fan, Dejia Xu, Dilin Wang, Sreyas Mohan, Forrest N. Iandola, Yilei Li, Qiang Liu 0001, Zhangyang Wang, Vikas Chandra |
AISTATS | 8 |
| 2025 | Make-A-Texture: Fast Shape-Aware Texture Generation in 3 SecondsabstractWe present Make-A-Texture, a new framework that efficiently synthesizes high-resolution texture maps from textual prompts for given 3D geometries. Our approach progressively generates textures that are consistent across multiple viewpoints with a depth-aware inpainting diffusion model, in an optimized sequence of viewpoints determined by an automatic view selection algorithm. A significant feature of our method is its remarkable efficiency, achieving a full texture generation within an end-to-end runtime of just 3.07 seconds on a single NVIDIA H100 GPU, significantly outperforming existing methods. Such an acceleration is achieved by optimizations in the diffusion model and a specialized backprojection method. Moreover, our method reduces the artifacts in the backprojection phase, by selectively masking out non-frontal faces, and internal faces of open-surfaced objects. Experimental results demonstrate that Make-A-Texture matches or exceeds the quality of other state-of-the-art methods. Our work significantly improves the applicability and practicality of texture generation models for real-world 3D content creation, including interactive creation and text-guided texture editing. Xiaoyu Xiang, Liat Sless Gorelik, Yuchen Fan 0001, Omri Armstrong, Forrest N. Iandola, Yilei Li, Ita Lifshitz |
WACV | 6 |
| 2024 | Taming Mode Collapse in Score Distillation for Text-to-3D GenerationabstractDespite the remarkable performance of score distillation in text-to-3D generation, such techniques notoriously suf-fer from view inconsistency issues, also known as “Janus” artifact, where the generated objects fake each view with multiple front faces. Although empirically effective methods have approached this problem via score debiasing or prompt engineering, a more rigorous perspective to explain and tackle this problem remains elusive. In this paper, we reveal that the existing score distillation-based text-to-3D generation frameworks degenerate to maximal likelihood seeking on each view independently and thus suffer from the mode collapse problem, manifesting as the Janus artifact in practice. To tame mode collapse, we improve score distillation by re-establishing the entropy term in the corresponding variational objective, which is applied to the distribution of rendered images. Maximizing the entropy encourages diversity among different views in generated 3D assets, thereby mitigating the Janus problem. Based on this new objective, we derive a new update rule for 3D score distillation, dubbed Entropic Score Distillation (ESD). We theoretically reveal that ESD can be simplified and implemented by just adopting the classifier-free guidance trick upon variational score distillation. Although embarrassingly straightforward, our extensive experiments demonstrate that ESD can be an effective treatment for Janus artifacts in score distillation. Peihao Wang, Dejia Xu, Zhiwen Fan, Dilin Wang, Sreyas Mohan, Forrest N. Iandola, Yilei Li, Qiang Liu 0001, Zhangyang Wang, Vikas Chandra |
CVPR | 8 |
| 2024 | WaSt-3D: Wasserstein-2 Distance for Scene-to-Scene Stylization on 3D Gaussians
Dmytro Kotovenko, Olga Grebenkova, Nikolaos Sarafianos, Avinash Paliwal, Pingchuan Ma 0006, Omid Poursaeed, Sreyas Mohan, Yuchen Fan 0001, Yilei Li, Björn Ommer |
ECCV (21) | 9 |
| 2024 | Semi-supervised fault diagnosis of wheelset bearings in high-speed trains using autocorrelation and improved flow Gaussian mixture model
Jiayi Wu 0006, Yilei Li, Limin Jia 0002, Guoping An, Yan-Fu Li, Jérôme Antoni, Ge Xin |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | MMG-Ego4D: Multi-Modal Generalization in Egocentric Action RecognitionabstractIn this paper, we study a novel problem in egocentric action recognition, which we term as “Multimodal Generalization“ (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely missing. We thoroughly investigate MMG in the context of standard supervised action recognition and the more challenging few-shot setting for learning new action categories. MMG consists of two novel scenarios, designed to support security, and efficiency considerations in real-world applications: (1) missing modality generalization where some modalities that were present during the train time are missing during the inference time, and (2) cross-modal zero-shot generalization, where the modalities present during the inference time and the training time are disjoint. To enable this investigation, we construct a new dataset MMG-Ego4D containing data points with video, audio, and inertial motion sensor (IMU) modalities. Our dataset is derived from Ego4D [27] dataset, but processed and thoroughly re-annotated by human experts to facilitate research in the MMG problem. We evaluate a diverse array of models on MMG-Ego4D and propose new methods with improved generalization ability. In particular, we introduce a new fusion module with modality dropout training, contrastive-based alignment training, and a novel cross-modal prototypical loss for better few-shot performance. We hope this study will serve as a benchmark and guide future research in multimodal generalization problems. The benchmark and code are available at https://github.com/facebookresearch/MMG_Ego4D Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang |
CVPR | 5 |
| 2021 | Improving Efficiency in Neural Network Accelerator using Operands Hamming Distance OptimizationabstractNeural network accelerator is a key enabler for the on-device AI inference, for which energy efficiency is an important metric. The datapath energy, including the computation energy and the data movement energy among the arithmetic units, claims a significant part of the total accelerator energy. By revisiting the basic physics of the arithmetic logic circuits, we show that the datapath energy is highly correlated with the bit flips when streaming the input operands into the arithmetic units, defined as the hamming distance (HD) of the input operand matrices. Based on the insight, we propose a post-training optimization algorithm and a HD-aware training algorithm to co-design and co-optimize the accelerator and the network synergistically. The experimental results based on post-layout simulation with MobileNetV2 demonstrate on average 2.85x datapath energy reduction and up to 8.51x datapath energy reduction for certain layers. Meng Li 0004, Yilei Li, Vikas Chandra |
ASP-DAC | 2 |
| 2021 | PyTorchVideo: A Deep Learning Library for Video UnderstandingabstractWe introduce PyTorchVideo, an open-source deep-learning library that provides a rich set of modular, efficient, and reproducible components for a variety of video understanding tasks, including classification, detection, self-supervised learning, and low-level processing. The library covers a full stack of video understanding tools including multimodal data loading, transformations, and models that reproduce state-of-the-art performance. PyTorchVideo further supports hardware acceleration that enables real-time inference on mobile devices. The library is based on PyTorch and can be used by any training framework; for example, PyTorchLightning, PySlowFast, or Classy Vision. PyTorchVideo is available at https://pytorchvideo.org/. Haoqi Fan 0001, Tullie Murrell, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Nikhila Ravi, Meng Li 0004, Haichuan Yang, Jitendra Malik, Ross B. Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, Christoph Feichtenhofer |
ACM Multimedia | 6 |
| 2020 | Difference-Frequency Ultrasound Imaging With Non-Linear ContrastabstractConventional ultrasound imaging is based on the scattering of sound from inhomogeneities in the density and the speed of sound and is often used in medicine to resolve pathologic compared to normal tissue. Here we demonstrate a difference-frequency ultrasound (dfUS) imaging method that is based on the interaction of two sound pulses that propagate non-collinearly and intersect in space and time. The dfUS signal arises primarily from the second-order non-linear coefficient, a contrast mechanism that differs from linear and harmonic US imaging. The distinct contrast mechanism allows dfUS to image anatomic features that are not identifiable in conventional US images of salmon and pig kidney tissue. Further, dfUS produces enhanced contrast of glioblastoma tumor implanted in the mouse brain, revealing its potential for improving medical diagnosis. Progress towards a real-time system is discussed. Yilei Li, Dina Polyak, Eli Johnson, Derek W. Yecies, Saba Shevidi, Adam de la Zerda, Melanie Hayden Gephart, Steven Chu |
IEEE Trans. Medical Imaging | 1 |
| 2019 | A 7.5-mW 10-Gb/s 16-QAM wireline transceiver with carrier synchronization and threshold calibration for mobile inter-chip communications in 16-nm FinFETabstractA compact energy-efficient 16-QAM wireline transceiver with carrier synchronization and threshold calibration is proposed to leverage high-density fine-pitch interconnects. Utilizing frequency-division multiplexing, the transceiver transfers four-bit data through one RF band to reduce intersymbol interferences. A forwarded clock is also transmitted through the same interconnect with the data simultaneously to enable low-power PVT-insensitive symbol clock recovery. A carrier synchronization algorithm is proposed to overcome nontrivial current and phase mismatches by including DC offset calibration and dedicated I/Q phase adjustments. Along with this carrier synchronization, a threshold calibration process is used for the transceiver to tolerate channel and circuit variations. The transceiver implemented in 16-nm FinFET occupies only 0.006-mm2 and achieves 10 Gb/s with 0.75-pJ/bit efficiency and <2.5-ns latency. Jieqiong Du, Chien-Heng Wong, Yo-Hao Tu, Wei-Han Cho, Yilei Li, Yuan Du, Po-Tsang Huang, Sheau Jiung Lee, Mau-Chung Frank Chang |
NOCS | 5 |
| 2019 | Optimization of the Trade-Off Between Speckle Reduction and Axial Resolution in Frequency CompoundingabstractWe measured the reduction of speckle by frequency compounding using Gaussian pulses, which have the least time-bandwidth product. The experimental results obtained from a tissue mimicking phantom agree quantitatively with numerical simulations of randomly distributed point scatterers. For a fixed axial resolution, the amount of speckle reduction is found to approach a maximum as the number of bands increases while the total spectral range that they cover is kept constant. An analytical solution of the maximal speckle reduction is derived and shows that the maximum improves approximately as the inverse square root of the Gaussian pulse bandwidth. Since the axial resolution is proportional to the inverse of the pulse bandwidth, an optimized trade-off between speckle reduction and axial resolution is obtained. Considerations for the applications of the optimized trade-off are discussed. Yilei Li, Yonatan Winetraub, Orly Liba, Adam de la Zerda, Steven Chu |
IEEE Trans. Medical Imaging | 1 |
| 2018 | A Single Layer 3-D Touch Sensing System for Mobile Devices ApplicationabstractTouch sensing has been widely implemented as a main methodology to bridge human and machine interactions. The traditional touch sensing range is 2-D and therefore limits the user experience. To overcome these limitations, we propose a novel 3-D contactless touch sensing called Airtouch system, which improves user experience by remotely detecting single/multi-finger position. A single layer touch panel with triangle-shaped electrodes is proposed to achieve multitouch detection capability as well as manufacturing cost reduction. Moreover, an oscillator-based-capacitive touch sensing circuit is implemented as the sensing hardware with the bootstrapping technique to eliminate the interchannel coupling effects. To further improve the system accuracy, a grouping algorithm is proposed to group the useful channels' data and filter out hardware noise impact. Finally, improved algorithms are proposed to eliminate the fringing capacitance effect and achieve accurate finger position estimation. EM simulation proved that the proposed algorithm reduced the maximum systematic error by 11 dB in the horizontal position detection. The proposed system consumes 2.3 mW and is fully compatible with existing mobile device environments. A prototype is built to demonstrate that the system can successfully detect finger movement in a vertical direction up to 6 cm and achieve a horizontal resolution up to 0.6 cm at 1 cm finger-height. As a new interface for human and machine interactions, this system offers great potential in finger movement detection and gesture recognition for small-sized electronics and advanced human interactive games for mobile device. Yan Zhang 0050, Yilei Li, Yuan Du, Yen-Cheng Kuan, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | A Novel Fully Synthesizable All-Digital RF Transmitter for IoT ApplicationsabstractIn this paper, a fully synthesizable all-digital transmitter (ADTX) is first proposed. This transmitter (TX) uses Cartesian architecture and supports wide-band quadratic-amplitude modulation with wide carrier frequency range. Furthermore, the design methodology for ADTX and corresponding bandpass filter is discussed. This TX is synthesized with digital register transfer level-graphic database system flow, and can be easily implemented in any standard CMOS technology. An exemplary TX is synthesized by TSMC 28-nm standard cell library with extremely small area (0.0009 mm2) and supports carrier frequency as high as 6 GHz with excellent error vector magnitude (<;-30 dB). To the best of the authors' knowledge, this is the first work on a fully synthesizable design of RF transistors, allowing easy technology migration and portability. Yilei Li, Kirti Dhwaj, Chien-Heng Wong, Yuan Du, Yiwu Tang, Yiyu Shi 0001, Tatsuo Itoh, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Invited - Airtouch: a novel single layer 3D touch sensing system for human/mobile devices interactionsabstractTouchscreen technology plays an important role in the booming mobile devices market. Traditional touchscreen only provides 2D interactions with limited user experience. To overcome these limitations, we propose a novel 3D touch sensing system called the Airtouch system, which can recognize the movement of the finger in a 3-dimensional space. Half of the manufacturing cost is reduced by applying only single layer electrodes in the touch panel design. Moreover, an oscillator based correlated double sampling circuit is implemented as the self-capacitive sensor with bootstrapping technique to reduce inter-channel-coupling effect. Additionally, new algorithm for finger positioning is created with grouping filter invented to reduce system background noise. The demonstrated setup can successfully detect finger movement within a vertical range of 6cm and achieve a horizontal resolution up to 1cm. This system offers great potential in both gesture recognition for small-sized electronics, and advanced human interactive games for TV and mobile device. Adrian Tang 0002, Yan Zhang 0050, Yilei Li, Kye Cheung, Mau-Chung Frank Chang |
DAC | 5 |
| 2016 | Invited - A 2.2 GHz SRAM with high temperature variation immunity for deep learning application under 28nmabstractWith the coming era of Big Data, hardware implementation of machine learning has become attractive for many applications, such as real-time object recognition and face recognition. The implementation of machine learning algorithms needs intensive memory access, and SRAM is critical for the overall performance. This paper proposes a new design of high speed SRAM for machine learning purposes. With fast access time (cycle time: 650 ps, access time: 350 ps), low sensitivity to temperature variation and high configurability (less than 10% performance difference between 125_rcw_tt vs 0_rcw_tt), the proposed SRAM is a better candidate for hardware machine learning system than the conventional SRAM. Compared with Samsung HL 152, our design has smaller size (121×43 um2 vs 127×44 um2) with half the number of pins ports (12 vs 25) and higher speed (2.2GHz vs 0.8GHz). Yen-Hsiang Wang, Yilei Li, Chien-Heng Wong, Tien Pei Chou, Young-Kai Chen, Mau-Chung Frank Chang |
DAC | 3 |
| 2012 | A triple-band flexible low-noise transmitter with linearity enhancementabstractA low-noise triple band transmitter for GSM triple-band and WCDMA is presented. By programming parameters, analog baseband and RF frontend can handle signals of different protocols and frequency bands. Novel linearization method is used in LPF and driver amplifier to meet the demand of 3G systems. Noise optimization is also adopted so that our transmitter achieves −157 dBc/Hz noise at 190 MHz frequency offset in WCDMA band, which eliminates external SAW filter. Yilei Li, Chuansheng Dong, Kefeng Han, Yongchang Yu, Xi Tan 0001, Na Yan 0004, Hao Min |
ISCAS | 1 |
| 2012 | Cognitive model based fashion style decision making
Jiyun Li, Yilei Li |
Expert Syst. Appl. | 2 |