Yuanqi Li

dblp:228/9753 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A multi-module optimized UAV-YOLOv11 model and 3D coordinate reconstruction method for UAV-based monitoring of cable dome nodes
Lingqian Wen, Xing Xie 0001, Yuanqi Li, Jiayang Lv
Adv. Eng. Informatics5
2026 Weakly-Supervised Shape Multi-Completion of Point Clouds by Structural Decomposition
abstract
The challenge of transforming partial point clouds into complete meshes still persists, with current methods facing issues like data accessibility constraint, shape preservation failure and poor robustness on real-scan data. Drawing inspiration from the structural information of objects to enhance the completion, we introduce an innovative weakly-supervised shape completion method leveraging structural decomposition without the necessity of SDFs during training. By representing objects as abstract structural frameworks and part details, our method initiates by forecasting the structure of the input partial point clouds, and individually restore each component through part decomposition completion and generation. Extracted part details are represented in images, which are porous and incomplete. Hence, we utilize a completion network to complete such details. For multiple results generation, a diffusion-based generation network is employed to generate a variety of details for the missing areas. The predicted structure and details are subsequently converted back into meshes, yielding the complete results. Since the details are depicted in images, our approach eliminates the need for SDFs during the training phase, achieving weakly-supervision. We conduct extensive comparisons on both artificial and real-scan datasets, demonstrating an average improvement of over 38.1% compared to the prior method, and achieving SOTA performance.
Changfeng Ma, Pengxiao Guo, Shuangyu Yang, Yuanqi Li, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.4
2026 Surface Reconstruction From Point Clouds via Image-Free Point-to-Gaussian Inference
abstract
The task of surface reconstruction from point clouds is to produce high-quality meshes using sampled 3D points (no images available). Traditional methods primarily focus on geometric accuracy but often produce meshes without texture colors. In this paper, we present a brand-new perspective in point cloud reconstruction task-Imagining points as more informative Gaussian splats and obtaining colored surfaces through free-form Gaussian-rendering reconstruction. We train a universal Point-to-Gaussian model to infer the attributes of Gaussian splats for any given pointcloud with merely point coordinates (and color optionally) as input, without requiring any image. Significant technical designs are applied on initialization, regularization and loss functions, making the whole learning process stable. The inferred Gaussian splats can faithfully recover the original appearance of objects or scenes (capable of quick rendering from any viewpoint, like human's imagination ability), meanwhile closely adhering to the input shape. After obtaining a sufficient number of virtually rendered images and depth maps, we employ the truncated signed distance function (TSDF) fusion to get the reconstruction results, producing high-quality and colored meshes. Extensive experiments demonstrate that our approach surpasses state-of-the-art methods in surface reconstruction metrics while maintaining high efficiency and simplicity.
Xinran Yang, Donghao Ji, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.3
2025 360-GS: Layout-Guided Panoramic Gaussian Splatting for Indoor Roaming
abstract
3D Gaussian Splatting (3D-GS) has recently attracted great attention with real-time and photo-realistic renderings. This technique typically takes perspective images as input and optimizes a set of 3D elliptical Gaussians by splatting them onto the image planes, resulting in$2 D$Gaussians. However, applying 3D-GS to panoramic inputs presents challenges in effectively modeling the projection onto the spherical surface of 360° images using 2D Gaussians. In practical applications, input panoramas are often sparse, leading to unreliable initialization of 3D Gaussians and subsequent degradation of 3D-GS quality. In addition, due to the under-constrained geometry of texture-less planes (e.g., walls and floors), 3D-GS struggles to model these flat regions with elliptical Gaussians, resulting in significant floaters in novel views. To address these issues, we propose 360-GS, a novel layout-guided 360° Gaussian splatting for a limited set of panoramic inputs. Instead of splatting 3D Gaussians directly onto the spherical surface, 360-GS projects them onto the tangent plane of the unit sphere and then maps them to the spherical projections.
Jiayang Bai, Letian Huang, Jie Guo 0001, Wen Gong, Yuanqi Li, Yanwen Guo 0001
3DV5
2025 High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance Encoding
abstract
The proliferation of Light Detection and Ranging (LiDAR) technology has facilitated the acquisition of three-dimensional point clouds, which are integral to applications in VR, AR, and Digital Twin. Oriented normals, critical for 3D reconstruction and scene analysis, cannot be directly extracted from scenes using LiDAR due to its operational principles. Previous traditional or learning-based methods are prone to inaccuracies due to uneven distribution and noise due to the dependence on local geometry features. This paper addresses the challenge of estimating oriented point normals by introducing a point cloud normal estimation framework via hybrid angular and Euclidean distance encoding (HAE). Our method overcomes the limitations of local geometric information by combining angular and Euclidean spaces to extract features from both point cloud coordinates and light rays, leading to more accurate normal estimation. The core of our network consists of an angular distance encoding module, which leverages both ray directions and point coordinates for unoriented normal refinement, and a ray feature fusion module for normal orientation, that is robust to noise. We also provide a point cloud dataset with ground truth normals, generated a virtual scanner, which reflects real scanning distributions and noise profiles.
Yuanqi Li, Jingcheng Huang, Hongshen Wang, Peiyuan Lv, Jiuming Zheng, Jie Guo 0001, Yanwen Guo 0001
CVPR1
2025 SGCR: Spherical Gaussians for Efficient 3D Curve Reconstruction
abstract
Neural rendering techniques have made substantial progress in generating photo-realistic 3D scenes. The latest 3D Gaussian Splatting technique has achieved high quality novel view synthesis as well as fast rendering speed. However, 3D Gaussians lack proficiency in defining accurate 3D geometric structures despite their explicit primitive representations. This is due to the fact that Gaussian’s attributes are primarily tailored and fine-tuned for rendering diverse 2D images by their anisotropic nature. To pave the way for efficient 3D reconstruction, we present Spherical Gaussians, a simple and effective representation for 3D geometric boundaries, from which we can directly reconstruct 3D feature curves from a set of calibrated multi-view images. Spherical Gaussians is optimized from grid initialization with a view-based rendering loss, where a 2D edge map is rendered at a specific view and then compared to the ground-truth edge map extracted from the corresponding image, without the need for any 3D guidance or supervision. Given Spherical Gaussians serve as intermedia for the robust edge representation, we further introduce a novel optimization-based algorithm called SGCR to directly extract accurate parametric curves from aligned Spherical Gaussians. We demonstrate that SGCR outperforms existing state-of-the-art methods in 3D edge reconstruction while enjoying great efficiency. Code is available at https://github.com/Martinyxr/SGCR.
Xinran Yang, Donghao Ji, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Junyuan Xie
CVPR3
2025 EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry Features
abstract
Point cloud reconstruction is a critical process in 3D representation and reverse engineering. When it comes to CAD models, edges are significant features that play a crucial role in characterizing the geometry of 3D shapes. However, few points are exactly sampled on edges during acquisition, resulting in apparent artifacts for the reconstruction task. Upsampling point cloud is a direct technical route, but there is a main challenge that the upsampled points may not align with the model edge accurately. To overcome this, we develop an integrated framework to estimate edges by joint regression of three geometry features—point-to-edge direction, point-to-edge distance and point normal. Benefiting these features, we implement a novel refinement process to move and produce more points which lie accurately on edges of the model, allowing for high-quality edge-preserving reconstruction. Experiments and comparisons against previous methods demonstrate our method’s effectiveness and superiority.
Xinran Yang, Donghao Ji, Yuanqi Li, Junyuan Xie, Jie Guo 0001, Yanwen Guo 0001
CVPR3
2025 GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections
Haiyang Bai, Songru Jiang, Tao Lu 0005, Yuanqi Li, Jie Guo 0001, Runze Fu, Yanwen Guo 0001, Lijun Chen 0006
ICCV6
2025 Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
abstract
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding.
Xiaoyu Zhan, Wenxuan Huang 0001, Xinyu Fu 0009, Changfeng Ma, Shaosheng Cao, Bohan Jia, Shaohui Lin, Zhenfei Yin, Lei Bai 0001, Wanli Ouyang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001
NeurIPS12
2025 Spectral-GS: Taming 3D Gaussian Splatting with Spectral Entropy
abstract
Recently, 3D Gaussian Splatting (3DGS) has achieved impressive results in novel view synthesis, demonstrating high fidelity and efficiency. However, it easily exhibits needle-like artifacts, especially when increasing the sampling rate. Mip-Splatting tries to remove these artifacts with a 3D smoothing filter for frequency constraints and a 2D Mip filter for approximated supersampling. Unfortunately, it tends to produce over-blurred results, and sometimes needle-like Gaussians still persist. Our spectral analysis of the covariance matrix during optimization and densification reveals that current 3DGS lacks shape awareness, relying instead on spectral radius and view positional gradients to determine splitting. As a result, needle-like Gaussians with small positional gradients and low spectral entropy fail to split and overfit high-frequency details. Furthermore, both the filters used in 3DGS and Mip-Splatting reduce the spectral entropy and increase the condition number during zooming in to synthesize novel view, causing view inconsistencies and more pronounced artifacts. Our Spectral-GS, based on spectral analysis, introduces 3D shape-aware splitting and 2D view-consistent filtering strategies, effectively addressing these issues, enhancing 3DGS’s capability to represent high-frequency details without noticeable artifacts, and achieving high-quality realistic rendering.
Letian Huang, Jie Guo 0001, Jialin Dan, Ruoyu Fu, Yuanqi Li, Yanwen Guo 0001
SIGGRAPH Asia5
2025 Realistic Simulation of Underwater Scene for Image Enhancement
abstract
In recent years, learning-based methods have performed remarkably well in underwater image enhancement, but their performance is limited by the lack of high-quality, diverse training datasets. Current underwater image datasets are unable to address the following three issues: intra-domain gaps in underwater environments, inter-domain gaps between synthetic and real data, and domain inaccuracies. To overcome these limitations, we construct a realistic underwater scene using 3D graphics engine through a three-step approach: 1) integrate a simulation-specific underwater light propagation models to create volumetric fog; 2) employ physical model-based rendering for accurate light field simulation; 3) configure scenes with parameters extracted from real underwater images. Based on this framework, we develop an underwater image enhancement dataset (MUSE). Experiments demonstrate that models trained on MUSE outperform those trained on conventional datasets, highlighting the effectiveness of our approach.
Tingyu Liu, Qunyan Jiang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Zhonghua Ni
IEEE Trans. Geosci. Remote. Sens.4
2025 TransparentGS: Fast Inverse Rendering of Transparent Objects with Gaussians
abstract
The emergence of neural and Gaussian-based radiance field methods has led to considerable advancements in novel view synthesis and 3D object reconstruction. Nonetheless, specular reflection and refraction continue to pose significant challenges due to the instability and incorrect overfitting of radiance fields to high-frequency light variations. Currently, even 3D Gaussian Splatting (3D-GS), as a powerful and efficient tool, falls short in recovering transparent objects with nearby contents due to the existence of apparent secondary ray effects. To address this issue, we propose TransparentGS, a fast inverse rendering pipeline for transparent objects based on 3D-GS. The main contributions are three-fold. Firstly, an efficient representation of transparent objects, transparent Gaussian primitives, is designed to enable specular refraction through a deferred refraction strategy. Secondly, we leverage Gaussian light field probes (GaussProbe) to encode both ambient light and nearby contents in a unified framework. Thirdly, a depth-based iterative probes query (IterQuery) algorithm is proposed to reduce the parallax errors in our probe-based framework. Experiments demonstrate the speed and accuracy of our approach in recovering transparent objects from complex environments, as well as several applications in computer graphics and vision.
Letian Huang, Dongwei Ye, Jialin Dan, Chengzhi Tao, Kun Zhou 0001, Bo Ren 0003, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.8
2025 Deep Point Cloud Edge Reconstruction via Surface Patch Segmentation
abstract
Parametric edge reconstruction for point cloud data is a fundamental problem in computer graphics. Existing methods first classify points as either edge points (including corners) or non-edge points, and then fit parametric edges to the edge points. However, few points are exactly sampled on edges in practical scenarios, leading to significant fitting errors in the reconstructed edges. Prominent deep learning-based methods also primarily emphasize edge points, overlooking the potential of non-edge areas. Given that sparse and non-uniform edge points cannot provide adequate information, we address this challenge by leveraging neighboring segmented patches to supply additional cues. We introduce a novel two-stage framework that reconstructs edges precisely and completely via surface patch segmentation. First, we propose PCER-Net, a Point Cloud Edge Reconstruction Network that segments surface patches, detects edge points, and predicts normals simultaneously. Second, a joint optimization module is designed to reconstruct a complete and precise 3D wireframe by fully utilizing the predicted results of the network. Concretely, the segmented patches enable accurate fitting of parametric edges, even when sparse points are not precisely distributed along the model's edges. Corners can also be naturally detected from the segmented patches. Benefiting from fitted edges and detected corners, a complete and precise 3D wireframe model with topology connections can be reconstructed by geometric optimization. Finally, we present a versatile patch-edge dataset, including CAD and everyday models (furniture), to generalize our method. Extensive experiments and comparisons against previous methods demonstrate our effectiveness and superiority. We will release the code and dataset to facilitate future research.
Yuanqi Li, Hongshen Wang, Jingcheng Huang, Jianwei Guo 0003, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.1
2025 Make PBR Materials Tileable With Latent Diffusion Inpainting
abstract
Physically-based-rendering (PBR) materials are crucial in modern rendering pipelines, and many studies have focused on acquiring these materials from reality or images. However, existing methods may result in non-tileable results, since the realistic inputs usually have seams. Compared to non-tileable materials, tileable PBR materials have more universal application scenarios. To address this issue, we introduce MaTi, a novel pipeline that converts non-tileable PBR materials into tileable ones with minimal distortion. MaTi rearranges material patches to align boundaries at the center of the image, and then uses a diffusion model to inpaint the seams. We use scaled gamma correction to reduce the occurrence of collapse when processing special material maps. The color correction and triangular blending are adopt to preserve the original material information. Additionally, we design a division and blending strategy to efficiently handle high resolution materials. Our experiments demonstrate that MaTi can seamlessly modify PBR materials while preserving the original information, outperforming existing synthesis methods.
Xiaoyu Zhan, Jianxin Yang, Jun Wang 0039, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.4
2024 LiDAR-Net: A Real-Scanned 3D Point Cloud Dataset for Indoor Scenes
abstract
In this paper, we present LiDAR-Net, a new real-scanned indoor point cloud dataset, containing nearly 3.6 billion precisely point-level annotated points, covering an expansive area of 30,000m2. It encompasses three prevalent daily environments, including learning scenes, working scenes, and living scenes. LiDAR-Net is characterized by its non-uniform point distribution, e.g., scanning holes and scanning lines. Additionally, it meticulously records and an-notates scanning anomalies, including reflection noise and ghost. These anomalies stem from specular reflections on glass or metal, as well as distortions due to moving persons. LiDAR-Net's realistic representation of non-uniform distribution and anomalies significantly enhances the training of deep learning models, leading to improved generalization in practical applications. We thoroughly evaluate the performance of state-of-the-art algorithms on LiDAR-Net and provide a detailed analysis of the results. Crucially, our research identifies several fundamental challenges in understanding indoor point clouds, contributing essential insights to future explorations in this field. Our dataset can be found online: http://lidar-net.njumeta.com.
Yanwen Guo 0001, Yuanqi Li, Dayong Ren, Xiaohong Zhang 0009, Liang Pu, Changfeng Ma, Xiaoyu Zhan, Jie Guo 0001, Mingqiang Wei, Yan Zhang 0057, Piaopiao Yu, Shuangyu Yang, Donghao Ji, Huisheng Ye
CVPR2
2024 Semantic Human Mesh Reconstruction with Textures
abstract
The field of 3D detailed human mesh reconstruction has made significant progress in recent years. However, current methods still face challenges when used in industrial applications due to unstable results, low-quality meshes, and a lack of UV unwrapping and skinning weights. In this paper, we present SHERT, a novel pipeline that can reconstruct semantic human meshes with textures and high-precision details. SHERT applies semantic- and normal-based sampling between the detailed surface (e.g. mesh and SDF) and the corresponding SMPL-X model to obtain a partially sampled semantic mesh and then generates the complete semantic mesh by our specifically designed self-supervised completion and refinement networks. Using the complete semantic mesh as a basis, we employ a texture diffusion model to create human textures that are driven by both images and texts. Our reconstructed meshes have stable UV unwrapping, high-quality triangle meshes, and consistent semantic information. The given SMPL-X model provides semantic information and shape priors, allowing SHERT to perform well even with incorrect and incomplete inputs. The semantic information also makes it easy to substitute and animate different body parts such as the face, body, and hands. Quantitative and qualitative experiments demonstrate that SHERT is capable of producing high-fidelity and robust semantic meshes that outperform state-of-the-art methods.
Xiaoyu Zhan, Jianxin Yang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Wenping Wang 0001
CVPR3
2024 Prompt3D: Random Prompt Assisted Weakly-Supervised 3D Object Detection
abstract
The prohibitive cost of annotations for fully supervised 3D indoor object detection limits its practicality. In this work, we propose Random Prompt Assisted Weakly-supervised 3D Object Detection, termed as Prompt3D, a weakly-supervised approach that leverages position-level labels to overcome this challenge. Explicitly, our method focuses on enhancing labeling using synthetic scenes crafted from 3D shapes generated via random prompts. First, a Synthetic Scene Generation (SSG) module is introduced to assemble synthetic scenes with a curated collection of 3D shapes, created via random prompts for each category. These scenes are enriched with automatically generated point-level annotations, providing a robust supervisory frame-work for training the detection algorithm. To enhance the transfer of knowledge from virtual to real datasets, we then introduce a Prototypical Proposal Feature Alignment (PPFA) module. This module effectively alleviates the domain gap by directly minimizing the distance between feature prototypes of the same class proposals across two domains. Compared with sota BR, our method improves by 5.4% and 8.7% on mAP with VoteNet and GroupFree3D serving as detectors respectively, demonstrating the effectiveness of our proposed method. Code is available at: https://github.com/huishengye/prompt3d.
Xiaohong Zhang 0009, Huisheng Ye, Qinyu Tang, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001
CVPR5
2024 On the Error Analysis of 3D Gaussian Splatting and an Optimal Projection Strategy
Letian Huang, Jiayang Bai, Jie Guo 0001, Yuanqi Li, Yanwen Guo 0001
ECCV (17)4
2024 A novel percussion-based approach for pipeline leakage detection with improved MobileNetV2
Longguang Peng, Yuanqi Li, Guofeng Du
Eng. Appl. Artif. Intell.3
2024 Parameter-Estimate-First False Data Injection Attacks in AC State Estimation Deployed With Moving Target Defense
abstract
Enabled by the widely deployed distributed flexible alternating current transmission system (D-FACTS) devices in practical systems, moving target defense (MTD) has been considered as an effective way to detect stealthy false data injection (FDI) attacks by actively changing branch parameters. However, existing MTD methods heavily depend on the assumption that opponents can not timely obtain the newly changed branch parameters. In this paper, a parameter-estimate-first FDI (PEF-FDI) attack is proposed to reveal vulnerabilities of MTD methods in AC state estimation, which can bypass bad data detectors in the existence of MTD. Specifically, a PEF-FDI attack model is proposed to timely construct attack vector and stealthily misguide the results of alternating current (AC) state estimation in the presence of MTD. Requirements of constructing PEF-FDI attacks on eavesdropped measurements are deduced to reveal the limitation on capability of attackers. Simulations in the IEEE 118-bus system verify the performance of the proposed PEF-FDI attacks.
Chensheng Liu, Yuanqi Li, Hongcheng Zhu, Yang Tang 0001, Wenli Du
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 Rendering discrete participating media using geometrical optics approximation
abstract
We consider the scattering of light in participating media composed of sparsely and randomly distributed discrete particles. The particle size is expected to range from the scale of the wavelength to several orders of magnitude greater, resulting in an appearance with distinct graininess as opposed to the smooth appearance of continuous media. One fundamental issue in the physically-based synthesis of such appearance is to determine the necessary optical properties in every local region. Since these properties vary spatially, we resort to geometrical optics approximation (GOA), a highly efficient alternative to rigorous Lorenz—Mie theory, to quantitatively represent the scattering of a single particle. This enables us to quickly compute bulk optical properties for any particle size distribution. We then use a practical Monte Carlo rendering solution to solve energy transfer in the discrete participating media. Our proposed framework is the first to simulate a wide range of discrete participating media with different levels of graininess, converging to the continuous media case as the particle concentration increases.
Jie Guo 0001, Bingyang Hu, Yuanqi Li, Yanwen Guo 0001, Lingqi Yan 0001
Comput. Vis. Media4
2022 Video Vectorization via Bipartite Diffusion Curves Propagation and Optimization
abstract
We propose a new video vectorization approach for converting videos in the raster format to vector representation with the benefits of resolution independence and compact storage. Through classifying extracted curves in each video frame into salient ones and non-salient ones, we introduce a novel bipartite diffusion curves (BDCs) representation in order to preserve both important image features such as sharp boundaries and regions with smooth color variation. This bipartite representation allows us to propagate non-salient curves across frames such that the propagation, in conjunction with geometry optimization and color optimization of salient curves, ensures the preservation of fine details within each frame and across different frames, and meanwhile, achieves good spatial-temporal coherence. Thorough experiments on a variety of videos show that our method is capable of converting videos to the vector representation with low reconstruction errors, low computational cost, and fine details, demonstrating our superior performance over the state of the art. We also show that, when used for video upsampling, our method produces results comparable to video super-resolution.
Yuanqi Li, Chuan Wang 0001, Jie Guo 0001, Jue Wang 0001, Yanwen Guo 0001, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2020 Policy Search by Target Distribution Learning for Continuous Control
abstract
It is known that existing policy gradient methods (such as vanilla policy gradient, PPO, A2C) may suffer from overly large gradients when the current policy is close to deterministic, leading to an unstable training process. We show that such instability can happen even in a very simple environment. To address this issue, we propose a new method, called target distribution learning (TDL), for policy improvement in reinforcement learning. TDL alternates between proposing a target distribution and training the policy network to approach the target distribution. TDL is more effective in constraining the KL divergence between updated policies, and hence leads to more stable policy improvements over iterations. Our experiments show that TDL algorithms perform comparably to (or better than) state-of-the-art algorithms for most continuous control tasks in the MuJoCo environment while being more stable in training.
Chuheng Zhang, Yuanqi Li, Jian Li 0015
AAAI2
2020 DoubleEnsemble: A New Ensemble Method Based on Sample Reweighting and Feature Selection for Financial Data Analysis
abstract
Modern machine learning models (such as deep neural networks and boosting decision tree models) have become increasingly popular in financial market prediction, due to their superior capacity to extract complex non-linear patterns. However, since financial datasets have very low signal-to-noise ratio and are non-stationary, complex models are often very prone to overfitting and suffer from instability issues. Moreover, as various machine learning and data mining tools become more widely used in quantitative trading, many trading firms have been producing an increasing number of features (aka factors). Therefore, how to automatically select effective features becomes an imminent problem. To address these issues, we propose DoubleEnsemble, an ensemble framework leveraging learning trajectory based sample reweighting and shuffling based feature selection. Specifically, we identify the key samples based on the training dynamics on each sample and elicit key features based on the ablation impact of each feature via shuffling. Our model is applicable to a wide range of base models, capable of extracting complex patterns, while mitigating the overfitting and instability issues for financial market prediction. We conduct extensive experiments, including price prediction for cryptocurrencies and stock trading, using both DNN and gradient boosting decision tree as base models. Our experiment results demonstrate that DoubleEnsemble achieves a superior performance compared with several baseline methods.
Chuheng Zhang, Yuanqi Li, Xi Chen 0010, Yifei Jin, Pingzhong Tang, Jian Li 0015
ICDM2
2020 Semeru: A Memory-Disaggregated Managed Runtime
Chenxi Wang 0005, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen 0001, Michael D. Bond, Ravi Netravali, Miryung Kim, Guoqing Harry Xu
OSDI4
2020 Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video Analytics
abstract
To cope with the high resource (network and compute) demands of real-time video analytics pipelines, recent systems have relied on frame filtering. However, filtering has typically been done with neural networks running on edge/backend servers that are expensive to operate. This paper investigates on-camera filtering, which moves filtering to the beginning of the pipeline. Unfortunately, we find that commodity cameras have limited compute resources that only permit filtering via frame differencing based on low-level video features. Used incorrectly, such techniques can lead to unacceptable drops in query accuracy. To overcome this, we built Reducto, a system that dynamically adapts filtering decisions according to the time-varying correlation between feature type, filtering threshold, query accuracy, and video content. Experiments with a variety of videos and queries show that Reducto achieves significant (51-97% of frames) filtering benefits, while consistently meeting the desired accuracy.
Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Guoqing Harry Xu, Ravi Netravali
SIGCOMM1