Wei Hua 0002

dblp:04/2514-2 · DBLP profile ↗
← Back
52ranked-venue papers
4as first author
20since 2021 · last 2025
0000-0003-2868-1920ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Improving generative trajectory prediction via collision-free modeling and goal scene reconstruction
Zhaoxin Su, Gang Huang 0004, Zhou Zhou 0003, Yongfu Li 0001, Sanyuan Zhang, Wei Hua 0002
Pattern Recognit. Lett.6
2025 Adaptive Depth-Converted-Scale Convolution for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object size and relationship among different objects, are the main clues to extract the scene structure. However, previous works lack the explicit handling of the changing sizes of the object due to the change of its depth. Especially in a monocular video, the size of the same object is continuously changed, resulting in size and depth ambiguity. To address this problem, we propose a Depth-converted-Scale Convolution (DcSConv) enhanced monocular depth estimation framework, by incorporating the prior relationship between the object depth and object scale to extract features from appropriate scales of the convolution receptive field. The proposed DcSConv focuses on the adaptive scale of the convolution filter instead of the local deformation of its shape. It establishes that the scale of the convolution filter matters no less (or even more in the evaluated task) than its local deformation. Moreover, a Depth-converted-Scale aware Fusion (DcS-F) is developed to adaptively fuse the DcSConv features and the conventional convolution features. Our DcSConv enhanced monocular depth estimation framework can be applied on top of existing CNN based methods as a plug-and-play module to enhance the conventional convolution block. Extensive experiments with different baselines have been conducted on the KITTI benchmark and our method achieves the best results with an improvement up to 11.6% in terms of SqRel reduction. Ablation study also validates the effectiveness of each proposed module.
Yanbo Gao, Huibin Bai, Huasong Zhou, Xingyu Gao 0001, Shuai Li 0005, Hui Yuan 0001, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Circuits Syst. Video Technol.8
2025 A Hierarchical Controller for Connected Truck Platoon: Analysis and Verification
abstract
This paper proposes a novel hierarchical controller for connected truck platoons. To this end, the predecessor following topology is used to characterize the communication connectivity between connected trucks. Then, a longitudinal efficient controller consisting of upper-level and lower-level controllers is proposed. In particular, the upper-level controller is designed based on the kinematic model to handle the car-following interactions between connected trucks and delays in communication and input. The lower-level controller comprises a feedforward and a feedback control law. The feedforward control law converts the desired acceleration from the upper-level controller into the vehicle throttle or braking pressure using the inverse dynamic model, while the feedback control law compensates for the control error caused by unknown vehicle parameters. In addition, in the linear region, the internal stability is analyzed based on the second-order kinematic model using s-domain analysis and linearization method, respectively. Then, the string stability is proved. The influence of parameters on the stability performance is extensively discussed using the stability diagram. Finally, the feasibility of the proposed controller is verified via co-simulations in PreScan and TruckSim, in terms of acceleration, velocity, and spacing error profiles.
Yongfu Li 0001, Junhong Fan, Longwang Huang, Gang Huang 0004, Wei Hua 0002, Wei Wu 0009, Shuyou Yu 0001, Shuming Shi 0002, Xinbo Gao 0001
IEEE Trans. Intell. Transp. Syst.5
2025 LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace Representation
abstract
Monocular depth estimation (MDE) has attracted increasing interest in the past few years, owing to its important role in 3D vision. MDE is the estimation of a depth map from a monocular image/video to represent the 3D structure of a scene, which is a highly ill-posed problem. To solve this problem, in this paper, we propose a LiftFormer based on lifting theory topology, for constructing an intermediate subspace that bridges the image color features and depth values, and a subspace that enhances the depth prediction around edges. MDE is formulated by transforming the depth value prediction problem into depth-oriented geometric representation (DGR) subspace feature representation, thus bridging the learning from color values to geometric depth values. A DGR subspace is constructed based on frame theory by using linearly dependent vectors in accordance with depth bins to provide a redundant and robust representation. The image spatial features are transformed into the DGR subspace, where these features correspond directly to the depth values. Moreover, considering that edges usually present sharp changes in a depth map and tend to be erroneously predicted, an edge-aware representation (ER) subspace is constructed, where depth features are transformed and further used to enhance the local features around edges. The experimental results demonstrate that our LiftFormer achieves state-of-the-art performance on widely used datasets, and an ablation study validates the effectiveness of both proposed lifting modules in our LiftFormer.
Shuai Li 0005, Huibin Bai, Yanbo Gao, Chong Lv, Hui Yuan 0001, Chuankun Li, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Multim.7
2024 3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
abstract
Text-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly at-tributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However, these methods heavily rely on the out-puts of existing models, leading to error accumulation in geometry and appearance that prevent the models from being used in various scenarios (e.g., outdoor and unreal sce-narios). To address this limitation, we generatively refine the newly generated local views by querying and aggregating global 3D information, and then progressively generate the 3D scene. Specifically, we employ a tri-plane features-based NeRF as a unified representation of the 3D scene to constrain global 3D consistency, and propose a generative refinement network to synthesize new contents with higher quality by exploiting the natural image prior from 2D dif-fusion model as well as the global 3D information of the current scene. Our extensive experiments demonstrate that, in comparison to previous methods, our approach supports wide variety of scene generation and arbitrary camera tra-jectories with improved visual quality and 3D consistency.
Songchun Zhang, Quan Zheng 0004, Rui Ma 0011, Wei Hua 0002, Hujun Bao, Weiwei Xu 0003, Changqing Zou
CVPR5
2024 Multimodal-XAD: Explainable Autonomous Driving Based on Multimodal Environment Descriptions
abstract
In recent years, deep learning-based end-to-end autonomous driving has become increasingly popular. However, deep neural networks are like black boxes. Their outputs are generally not explainable, making them not reliable to be used in real-world environments. To provide a solution to this problem, we propose an explainable deep neural network that jointly predicts driving actions and multimodal environment descriptions of traffic scenes, including bird-eye-view (BEV) maps and natural-language environment descriptions. In this network, both the context information from BEV perception and the local information from semantic perception are considered before producing the driving actions and natural-language environment descriptions. To evaluate our network, we build a new dataset with hand-labelled ground truth for driving actions and multimodal environment descriptions. Experimental results show that the combination of context information and local information enhances the prediction performance of driving action and environment description, thereby improving the safety and explainability of our end-to-end autonomous driving network.
Yuchao Feng, Wei Hua 0002, Yuxiang Sun 0002
IEEE Trans. Intell. Transp. Syst.3
2024 Finite-Time Cooperative Control for Vehicle Platoon With Sliding-Mode Controller and Disturbance Observer
abstract
This article proposes a finite-time based sliding-mode controller (FTSMC) and disturbance observer (FTDO) for connected vehicle (CV) platoon with uncertain dynamics. In particular, a recursive structure consisting of first-level and second-level sliding mode surfaces (SMSs) is developed for the chattering of the conventional SMC. Herein, the first-level SMS is designed based on a proportional-integral-derivative SMS considering the spacing error and interactive behaviors of vehicles, and the second-level SMS is based on an integral terminal SMS. Simultaneously, the disturbances suffered from the ego-vehicle uncertainty and nonlinearity are estimated by the FTDO. The FTSMC and FTDO are proposed to regulate the CV platoon in finite time. Meanwhile, the CV platoon can ensure the finite stability and string stability via rigorous analysis. Finally, the feasibility of the proposed controller is verified by extensive simulations and co-simulations, and small-scaled experiments.
Yongxin Zhu 0004, Yongfu Li 0001, Keyue Zeng, Longwang Huang, Gang Huang 0004, Wei Hua 0002, Xinbo Gao 0001
IEEE Trans. Intell. Transp. Syst.6
2024 A Structure-Preserving and Illumination-Consistent Cycle Framework for Image Harmonization
abstract
Ina composite image, the foreground and background are filmed under different scenarios, such as different lighting conditions, causing inconsistency and reducing the overall realism of the image. Image harmonization aims to generate visually realistic composite images by adjusting the foreground to the background conditions while maintaining the structure. Existing methods focus on adjusting the foreground object by directly training the foreground generation network with the ground truth, neglecting the different roles of the illumination and structure of the foreground in image harmonization. Moreover, the use of background, except for providing illumination, is not thoroughly investigated in this task. In this paper, we propose a structure-preserving and illumination-consistent cycle (SP-IC cycle) framework for image harmonization by exploring the illumination and structure of both the foreground and background. It achieves image harmonization by specifically changing the illumination and keeping the structure instead of ambiguously changing the foreground. Then, an illumination-consistent foreground harmonization cycle is developed to change the foreground illumination, while a structure-preserving cycle is designed to keep the foreground structure. Background information is explored in both cycles to assist in decomposing the illumination and structure of the foreground. In addition, the proposed SP-IC cycle framework can be applied to any image harmonization method to further boost its performance. Experimental results demonstrate that our method achieves better harmonious image quality than state-of-the-art methods, especially on an illumination-varying dataset.
Qingjie Shi, Yanbo Gao, Shuai Li 0005, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Multim.5
2024 Neural Kernel Regression for Consistent Monte Carlo Denoising
abstract
Unbiased Monte Carlo path tracing that is extensively used in realistic rendering produces undesirable noise, especially with low samples per pixel (spp). Recently, several methods have coped with this problem by importing unbiased noisy images and auxiliary features to neural networks to either predict a fixed-sized kernel for convolution or directly predict the denoised result. Since it is impossible to produce arbitrarily high spp images as the training dataset, the network-based denoising fails to produce high-quality images under high spp. More specifically, network-based denoising is inconsistent and does not converge to the ground truth as the sampling rate increases. On the other hand, the post-correction estimators yield a blending coefficient for a pair of biased and unbiased images influenced by image errors or variances to ensure the consistency of the denoised image. As the sampling rate increases, the blending coefficient of the unbiased image converges to 1, that is, using the unbiased image as the denoised results. However, these estimators usually produce artifacts due to the difficulty of accurately predicting image errors or variances with low spp. To address the above problems, we take advantage of both kernel-predicting methods and post-correction denoisers. A novel kernel-based denoiser is proposed based on distribution-free kernel regression consistency theory, which does not explicitly combine the biased and unbiased results but constrain the kernel bandwidth to produce consistent results under high spp. Meanwhile, our kernel regression method explores bandwidth optimization in the robust auxiliary feature space instead of the noisy image space. This leads to consistent high-quality denoising at both low and high spp. Experiment results demonstrate that our method outperforms existing denoisers in accuracy and consistency.
Pengju Qiao, Qi Wang 0111, Yuchi Huo, Shiji Zhai, Wei Hua 0002, Hujun Bao, Tao Liu 0016
ACM Trans. Graph.6
2023 I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFs
abstract
In this work, we present I2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based frame-work jointly recovers the underlying shapes, incident radiance and materials from multi-view images. We introduce a novel bubble loss for fine-grained small objects and error-guided adaptive sampling scheme to largely improve the reconstruction quality on large-scale indoor scenes. Further, we propose to decompose the neural radiance field into spatially-varying material of the scene as a neural field through surface-based, differentiable Monte Carlo raytracing and emitter semantic segmentations, which enables physically based and photorealistic scene relighting and editing applications. Through a number of qualitative and quantitative experiments, we demonstrate the superior quality of our method on indoor scene reconstruction, novel view synthesis, and scene editing compared to state-of-the-art baselines. Our project page is at https://jingsenzhu.github.io/i2-sdf.
Jingsen Zhu, Yuchi Huo, Qi Ye 0001, Fujun Luan, Jifan Li, Dianbing Xi, Lisha Wang, Rui Tang 0015, Wei Hua 0002, Hujun Bao, Rui Wang 0004
CVPR9
2023 PriorLane: A Prior Knowledge Enhanced Lane Detection Approach Based on Transformer
abstract
Lane detection is one of the fundamental modules in self-driving. In this paper we employ a transformer-only method for lane detection, thus it could benefit from the blooming development of fully vision transformer and achieve the state-of-the-art (SOTA) performance on both CULane and TuSimple benchmarks, by fine-tuning the weight fully pre-trained on large datasets. More importantly, this paper proposes a novel and general framework called PriorLane, which is used to enhance the segmentation performance of the fully vision transformer by introducing the low-cost local prior knowledge. Specifically, PriorLane utilizes an encoder-only transformer to fuse the feature extracted by a pre-trained segmentation model with prior knowledge embeddings. Note that a Knowledge Embedding Alignment (KEA) module is adapted to enhance the fusion performance by aligning the knowledge embedding. Extensive experiments on our Zjlab dataset show that PriorLane outperforms SOTA lane detection methods by a 2.82% mIoU when prior knowledge is employed, and the code will be released at: https://github.com/vincentqqb/PriorLane.
Qibo Qiu, Haiming Gao, Wei Hua 0002, Gang Huang 0004, Xiaofei He 0001
ICRA3
2023 OPUPO: Defending Against Membership Inference Attacks With Order-Preserving and Utility-Preserving Obfuscation
abstract
In this work, we present OPUPO to protect machine learning classifiers against black-box membership inference attacks by alleviating the prediction difference between training and non-training samples. Specifically, we apply order-preserving and utility-preserving obfuscation to prediction vectors. The order-preserving constraint strictly maintains the order of confidence scores in the prediction vectors, guaranteeing that the model's classification accuracy is not affected. The utility-preserving constraint, on the other hand, enables adaptive distortions to the prediction vectors in order to protect their utility. Moreover, OPUPO is proved to be adversary resistant that even well-informed defense-aware adversaries cannot restore the original prediction vectors to bypass the defense. We evaluate OPUPO on machine learning and deep learning classifiers trained with four popular datasets. Experiments verify that OPUPO can effectively defend against state-of-the-art attack techniques with negligible computation overhead. In specific, the inference accuracy could be reduced from as high as 87.66% to around 50%, i.e., random guess, and the prediction time will increase by only 0.44% on average. The experiments also show that OPUPO could achieve better privacy-utility trade-off than existing defenses.
Gang Huang 0004, Wei Hua 0002
IEEE Trans. Dependable Secur. Comput.4
2023 Physics-Informed Time-Aware Neural Networks for Industrial Nonintrusive Load Monitoring
abstract
Nonintrusive load monitoring enables the situational awareness of appliance-level energy consumption without installing appliance-specific sensors. It has been researched for over 30 years, with deep learning methods being the state-of-the-art solutions. However, current works mainly focus on the residential scenario, and industrial load disaggregation as a more challenging problem from the appliance-type perspective is much less investigated. Nevertheless, industrial loads play an important role in energy savings and climate change mitigation and adaptation. Therefore, this article focuses on the industrial nonintrusive load monitoring problem and proposes a physics-informed time-aware neural network method for it. Herein, multiple features of industrial loads are considered, and the physics relationship among them is leveraged to improve the learning process explicitly. In addition, a 2-D convolutional layer is further proposed to encode the timestamp for feature enhancement. Experiments on real-world industrial data from ten appliances will verify the effectiveness of the proposed method.
Gang Huang 0004, Zhou Zhou 0003, Fei Wu 0001, Wei Hua 0002
IEEE Trans. Ind. Informatics4
2023 NLE-DM: Natural-Language Explanations for Decision Making of Autonomous Driving Based on Semantic Scene Understanding
abstract
In recent years, the advancement of deep-learning technologies has greatly promoted the research progress of autonomous driving. However, deep neural network is like a black box. Given a specific input, it is difficult to explain the output of the network. Without explainable results, it would be unsafe to deploy deep networks in unseen environments or environments with potential unexpected situations. Especially for decision-making networks, inappropriate outputs could lead to severe traffic accidents. To provide a solution to this problem, we propose a deep neural network that jointly predicts the decision-making actions and corresponding natural-language explanations based on semantic scene understanding. Two types of explanations, the reasons of driving actions and the surrounding environment descriptions of the ego-vehicle, are designed. Both the reasons and descriptions are in the form of natural language. The decision-making actions could be explained by the corresponding reasons or the environment descriptions. We also release a large-scale dataset with hand-labelled ground truth including driving actions and environment descriptions. The superiority of our network over other methods is demonstrated on both our dataset and a public dataset.
Yuchao Feng, Wei Hua 0002, Yuxiang Sun 0002
IEEE Trans. Intell. Transp. Syst.2
2023 A Generic Approach to Eco-Driving of Connected Automated Vehicles in Mixed Urban Traffic and Heterogeneous Power Conditions
abstract
The connected automated vehicles (CAVs) are envisioned to be implemented most likely on electric vehicles, while traditional fuel-powered manually-driven vehicles (MVs) would probably still dominate the automobile market in the next decade. In this context, this paper addresses urban eco-driving of CAVs in mixed traffic and heterogeneous power conditions. The paper aims to develop a practical and deployable eco-driving strategy for CAVs in mixed traffic flow of CAVs and MVs under realistic and complex traffic conditions. Several typical eco-driving scenarios were studied in detail. In a nutshell, the eco-driving strategy for each CAV was determined by solving a typical two-point boundary value problem with minimum electric energy consumption in urban traffic conditions with small market penetration rates (MPRs) of CAVs. A rolling-horizon scheme was applied to implement the eco-driving strategy to handle uncertain/unpredictable disturbances of preceding MVs and the interference of junction queues to the eco-driving maneuvers of CAVs. The paper also studied how eco-driving for electrified CAVs would affect MVs’ fuel consumptions. Simulation studies were carried out on urban arterial roads of multiple signalized intersections in various scenarios of demand and MPR to verify the energy savings effect of the proposed eco-driving strategy. The results showed that via eco-driving electrified CAVs each had a potential of reducing energy consumption by 40%-61%, meanwhile leading to 5%-34% fuel savings on average for each following MV. Further issues concerning the energy saving mechanism of electrified CAVs, impacts of MVs cut-in from adjacent lanes, and passenger comfort were also examined.
Yonghui Hu, Daofei Li, Lihui Zhang, Simon Hu 0001, Wei Hua 0002, Jingqiu Guo
IEEE Trans. Intell. Transp. Syst.7
2023 NeLT: Object-Oriented Neural Light Transfer
abstract
This article presents object-oriented neural light transfer (NeLT), a novel neural representation of the dynamic light transportation between an object and the environment. Our method disentangles the global illumination of a scene into individual objects’ light transportation represented via neural networks, then composes them explicitly. It therefore enables flexible rendering with dynamic lighting, cameras, materials, and objects. Our rendering features various important global illumination effects, such as diffuse illumination, glossy illumination, dynamic shadowing, and indirect illumination, which completes the capability of existing neural object representation. Experiments show that NeLT does not require path tracing or shading results as input but achieves rendering quality comparable to state-of-the-art rendering frameworks, including the recent deep learning based denoisers.
Chuankun Zheng, Yuchi Huo, Shaohua Mo, Zhizhen Wu, Wei Hua 0002, Rui Wang 0004, Hujun Bao
ACM Trans. Graph.6
2022 Crossmodal Transformer Based Generative Framework for Pedestrian Trajectory Prediction
abstract
Providing guidance about collision avoidance, pedestrian trajectory prediction is an important task for autonomous driving. In this paper, to produce plausible trajectory predictions in the first-person view circumstance, we propose a crossmodal transformer based generative framework which could leverage sequences of cues from multiple modalities as well as pedestrian attributes. For the encoder, crossmodal transformers are exploited during the past stage to explore the cross-relation features of four modality-modality pairs, which are then fused with the help of a branch assigning operation and a modality attention module. For the decoder, we employ a bézier curve interpolation based method to project encoder features into trajectory results. Our training process not only considers the pedestrian's intention of crossing road but also optimizes our model to achieve more accurate predictions at the terminal time steps. Experimental results demonstrate that our framework outperforms state-of-the-art methods on both JAAD and PIE datasets. Especially, compared with the best baseline, our method could achieve 15.1%/14.3% and 14.3%/22.2% improvement for deterministic/multimodal prediction in the metric of box center final displacement error on JAAD and PIE, respectively.
Zhaoxin Su, Gang Huang 0004, Sanyuan Zhang, Wei Hua 0002
ICRA4
2022 PowerNet: Learning-Based Real-Time Power-Budget Rendering
abstract
With the prevalence of embedded GPUs on mobile devices, power-efficient rendering has become a widespread concern for graphics applications. Reducing the power consumption of rendering applications is critical for extending battery life. In this paper, we present a new real-time power-budget rendering system to meet this need by selecting the optimal rendering settings that maximize visual quality for each frame under a given power budget. Our method utilizes two independent neural networks trained entirely by synthesized datasets to predict power consumption and image quality under various workloads. This approach spares time-consuming precomputation or runtime periodic refitting and additional error computation. We evaluate the performance of the proposed framework on different platforms, two desktop PCs and two smartphones. Results show that compared to the previous state of the art, our system has less overhead and better flexibility. Existing rendering engines can integrate our system with negligible costs.
Yunjin Zhang, Rui Wang 0004, Yuchi Huo, Wei Hua 0002, Hujun Bao
IEEE Trans. Vis. Comput. Graph.4
2022 Color Contrast Enhanced Rendering for Optical See-Through Head-Mounted Displays
abstract
Most commercially available optical see-through head-mounted displays (OST-HMDs) utilize optical combiners to simultaneously visualize the physical background and virtual objects. The displayed images perceived by users are a blend of rendered pixels and background colors. Enabling high fidelity color perception in mixed reality (MR) scenarios using OST-HMDs is an important but challenging task. We propose a real-time rendering scheme to enhance the color contrast between virtual objects and the surrounding background for OST-HMDs. Inspired by the discovery of color perception in psychophysics, we first formulate the color contrast enhancement as a constrained optimization problem. We then design an end-to-end algorithm to search the optimal complementary shift in both chromaticity and luminance of the displayed color. This aims at enhancing the contrast between virtual objects and the real background as well as keeping the consistency with the original displayed color. We assess the performance of our approach using a simulated OST-HMD environment and an off-the-shelf OST-HMD. Experimental results from objective evaluations and subjective user studies demonstrate that the proposed approach makes rendered virtual objects more distinguishable from the surrounding background, thereby bringing a better visual experience.
Yunjin Zhang, Rui Wang 0004, Yifan Peng 0001, Wei Hua 0002, Hujun Bao
IEEE Trans. Vis. Comput. Graph.4
2021 CR-LSTM: Collision-prior Guided Social Refinement for Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction is a challenge because of the complex social interactions in context and the elusive intention of each pedestrian. Collision avoidance is one of the most common social interactions in real world, while existing data-driven works have not handled it well yet. In order to address this issue, we propose a framework that considers the theory about the minimum distance between each pedestrian-pedestrian pair and the corresponding time as the collision related prior knowledge. With the prior, our social refinement module, called Collision-prior Guided Refinement, can be guided to understand the collision situations of a crowd through a message passing mechanism. To focus on more useful information from context, we also introduce pedestrian-wise attention and collision gate to jointly judge collision potential for all pedestrian-pedestrian pairs. Experimental results demonstrate that our framework can achieve competitive results on ETH and UCY datasets by comparing with existing works. In addition, it indicates the superiority of our framework in the aspect of collision avoidance.
Zhaoxin Su, Sanyuan Zhang, Wei Hua 0002
IROS3
2020 Spherical Gaussian-based Lightcuts for Glossy Interreflections
abstract
Abstract It is still challenging to render directional but non‐specular reflections in complex scenes. The SG‐based (Spherical Gaussian) many‐light framework provides a scalable solution but still requires a large number of glossy virtual lights to avoid spikes as well as reduce clamping errors. Directly gathering contributions from these glossy virtual lights to each pixel in a pairwise way is very inefficient. In this paper, we propose an adaptive algorithm with tighter error bounds to efficiently compute glossy interreflections from glossy virtual lights. This approach is an extension of the Lightcuts that builds hierarchies on both lights and pixels with new error bounds and new GPU‐based traversal methods between light and pixel hierarchies. Results demonstrate that our method is able to faithfully and efficiently compute glossy interreflections in scenes with highly glossy and spatial varying reflectance. Compared with the conventional Lightcuts method, our approach generates lightcuts with only one‐fourth to one‐fifth light nodes therefore exhibits better scalability. Additionally, after being implemented on GPU, our algorithms achieve a magnitude of faster performance than the previous method.
Yuchi Huo, Shihao Jin, Tao Liu 0016, Wei Hua 0002, Rui Wang 0004, Hujun Bao
Comput. Graph. Forum4
2020 Automatic Band-Limited Approximation of Shaders Using Mean-Variance Statistics in Clamped Domain
abstract
Abstract In this paper, we present a new shader smoothing method to improve the quality and generality of band‐limiting shader programs. Previous work [YB18] treats intermediate values in the program as random variables, and utilizes mean and variance statistics to smooth shader programs. In this work, we extend such a band‐limiting framework by exploring the observation that one intermediate value in the program is usually computed by a complex composition of functions, where the domain and range of composited functions heavily impact the statistics of smoothed programs. Accordingly, we propose three new shader smoothing rules for specific composition of functions by considering the domain and range, enabling better mean and variance statistics of approximations. Aside from continuous functions, the texture, such as color texture or normal map, is treated as a discrete function with limited domain and range, thereby can be processed similarly in the newly proposed framework. Experiments show that compared with previous work, our method is capable of generating better smoothness of shader programs as well as handling a broader set of shader programs.
Rui Wang 0004, Yuchi Huo, Wenting Zheng, Wei Hua 0002, Hujun Bao
Comput. Graph. Forum5
2019 Human Sensitivity to Slopes of Slanted Paths
abstract
Redirected walking allows users to walk naturally through a large immersive virtual environment while the physical space is limited. Previous studies have analyzed human sensitivity to redirected walking in a horizontal direction, but users also need to walk on slopes to change their height. In this work, we expand the vertical movement space by positioning users on virtual paths with slopes that are different from those of real paths. We conduct psychological experiments to explore human sensitivity to slope gains that describe the discrepancies between the slopes of paths in virtual and real environments. The investigation shows that humans can walk on virtual slopes that are higher or lower than the real position without detecting the slopes and establishes corresponding detection thresholds.
Luyao Hu, Yaorui Zhang, Rui Wang 0004, Zaifeng Gao, Hujun Bao, Wei Hua 0002
VR6
2016 Adaptive matrix column sampling and completion for rendering participating media
abstract
Several scalable many-light rendering methods have been proposed recently for the efficient computation of global illumination. However, gathering contributions of virtual lights in participating media remains an inefficient and time-consuming task. In this paper, we present a novel sparse sampling and reconstruction method to accelerate the gathering step of the many-light rendering for participating media. Our technique explores the observation that the scattered lightings are usually locally coherent and of low rank even in heterogeneous media. In particular, we first introduce a matrix formation with light segments as columns and eye ray segments as rows, and formulate the gathering step into a matrix sampling and reconstruction problem. We then propose an adaptive matrix column sampling and completion algorithm to efficiently reconstruct the matrix by only sampling a small number of elements. Experimental results show that our approach greatly improves the performance, and obtains up to one order of magnitude speedup compared with other state-of-the-art methods of many-light rendering for participating media.
Yuchi Huo, Rui Wang 0004, Tianlei Hu, Wei Hua 0002, Hujun Bao
ACM Trans. Graph.4
2013 Shadow geometry maps for alias-free shadows
Rui Wang 0004, Yingqing Wu, Minghao Pan, Wei Chen 0001, Wei Hua 0002
Sci. China Inf. Sci.5
2013 GPU-based out-of-core many-lights rendering
abstract
In this paper, we present a GPU-based out-of-core rendering approach under the many-lights rendering framework. Many-lights rendering is an efficient and scalable rendering framework for a large number of lights. But when the data sizes of lights and geometry are both beyond the in-core memory storage size, the data management of these two out-of-core data becomes critical and challenging. In our approach, we formulate such a data management as a graph traversal optimization problem that first builds out-of-core lights and geometry data into a graph, and then guides shading computations by finding a shortest path to visit all vertices in the graph. Based on the proposed data management, we develop a GPU-based out-of-GPU-core rendering algorithm that manages data between the CPU host memory and the GPU device memory. Two main steps are taken in the algorithm: the out-of-core data preparation to pack data into optimal data layouts for the many-lights rendering, and the out-of-core shading using graph-based data management. We demonstrate our algorithm on scenes with out-of-core detailed geometry and out-of-core lights. Results show that our approach generates complex global illumination effects with increased data access coherence and has one order of magnitude performance gain over the CPU-based approach.
Rui Wang 0004, Yuchi Huo, Yazhen Yuan, Kun Zhou 0001, Wei Hua 0002, Hujun Bao
ACM Trans. Graph.5
2013 Analytic Double Product Integrals for All-Frequency Relighting
abstract
This paper presents a new technique for real-time relighting of static scenes with all-frequency shadows from complex lighting and highly specular reflections from spatially varying BRDFs. The key idea is to depict the boundaries of visible regions using piecewise linear functions, and convert the shading computation into double product integrals—the integral of the product of lighting and BRDF on visible regions. By representing lighting and BRDF with spherical Gaussians and approximating their product using Legendre polynomials locally in visible regions, we show that such double product integrals can be evaluated in an analytic form. Given the precomputed visibility, our technique computes the visibility boundaries on the fly at each shading point, and performs the analytic integral to evaluate the shading color. The result is a real-time all-frequency relighting technique for static scenes with dynamic, spatially varying BRDFs, which can generate more accurate shadows than the state-of-the-art real-time PRT methods.
Rui Wang 0004, Minghao Pan, Weifeng Chen 0002, Zhong Ren 0001, Kun Zhou 0001, Wei Hua 0002, Hujun Bao
IEEE Trans. Vis. Comput. Graph.6
2012 Compressing repeated content within large-scale remote sensing images
Wei Hua 0002, Rui Wang 0004, Xusheng Zeng, Ying Tang 0004, Huamin Wang 0001, Hujun Bao
Vis. Comput.1
2011 Controllable highly regular triangulation
Jin Huang 0001, Muyang Zhang, Wenjie Pei, Wei Hua 0002, Hujun Bao
Sci. China Inf. Sci.4
2011 Discriminative concept factorization for data representation
Wei Hua 0002, Xiaofei He 0001
Neurocomputing1
2011 GPU-friendly shape interpolation based on trajectory warping
abstract
Abstract In this paper, we propose a GPU‐friendly shape interpolation method. In contrast with state‐of‐the‐art interpolation algorithms, our method computes the trajectory of each vertex independently instead of solving large linear systems in every interpolation step. Given two poses being interpolated, we find trajectory parameters for each vertex by optimization with the consideration of the key pose reconstruction and as‐rigid‐as‐possible deformation in the pre‐computing stage. During run‐time, the vertices coordinates on the intermediate shape can be computed in parallel according to a close form formulation. In the results we demonstrate that our method achieves extremely high performance on modern GPU and can be extended easily to multi‐pose interpolation. Copyright © 2011 John Wiley & Sons, Ltd.
Lu Chen 0001, Jin Huang 0001, Hongxin Zhang 0001, Wei Hua 0002
Comput. Animat. Virtual Worlds4
2011 Robust Bilayer Segmentation and Motion/Depth Estimation with a Handheld Camera
abstract
Extracting high-quality dynamic foreground layers from a video sequence is a challenging problem due to the coupling of color, motion, and occlusion. Many approaches assume that the background scene is static or undergoes the planar perspective transformation. In this paper, we relax these restrictions and present a comprehensive system for accurately computing object motion, layer, and depth information. A novel algorithm that combines different clues to extract the foreground layer is proposed, where a voting-like scheme robust to outliers is employed in optimization. The system is capable of handling difficult examples in which the background is nonplanar and the camera freely moves during video capturing. Our work finds several applications, such as high-quality view interpolation and video editing.
Guofeng Zhang 0001, Jiaya Jia, Wei Hua 0002, Hujun Bao
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 GPU-based dynamic quad stream for forest rendering
Wei Hua 0002, Hujun Bao
Sci. China Inf. Sci.2
2010 Interactive hair rendering under environment lighting
abstract
We present an algorithm for interactive hair rendering with both single and multiple scattering effects under complex environment lighting. The outgoing radiance due to single scattering is determined by the integral of the product of the environment lighting, the scattering function, and the transmittance that accounts for self-shadowing among hair fibers. We approximate the environment light by a set of spherical radial basis functions (SRBFs) and thus convert the outgoing radiance integral into the sum of radiance contributions of all SRBF lights. For each SRBF light, we factor out the effective transmittance to represent the radiance integral as the product of two terms: the transmittance and the convolution of the SRBF light and the scattering function. Observing that the convolution term is independent of the hair geometry, we precompute it for commonly-used scattering models, and reduce the run-time computation to table lookups. We further propose a technique, called the convolution optical depth map , to efficiently approximate the effective transmittance by filtering the optical depth maps generated at the center of the SRBF using a depth-dependent kernel. As for the multiple scattering computation, we handle SRBF lights by using similar factorization and precomputation schemes, and adopt sparse sampling and interpolation to speed up the computation. Compared to off-line algorithms, our algorithm can generate images of comparable quality, but at interactive frame rates.
Zhong Ren 0001, Kun Zhou 0001, Wei Hua 0002, Baining Guo
ACM Trans. Graph.4
2010 Adaptive voxels: interactive rendering of massive 3D models
Fenglin Tian, Wei Hua 0002, Zilong Dong, Hujun Bao
Vis. Comput.2
2009 Confidence-Based Color Modeling for Online Video Segmentation
Fan Zhong 0001, Xueying Qin, Jiazhou Chen 0002, Wei Hua 0002, Qunsheng Peng 0001
ACCV (2)4
2009 Video stabilization based on a 3D perspective camera model
Guofeng Zhang 0001, Wei Hua 0002, Xueying Qin, Yuanlong Shao, Hujun Bao
Vis. Comput.2
2008 Procedural modeling of urban zone by optimization
abstract
Abstract Procedural modeling technology may be applied for constructing a large‐scale urban scene. Most of the previous studies have exploited a grammar‐based modeling method to generate models. Nevertheless, we formulate the urban planning as a constrained layout optimization problem, propose an algorithm to solve the problem, and procedurally generate models of the urban zone. It produces an extensive urban virtual environment for computer games and simulations at low cost. We optimize a cost function to distribute buildings and roads subject to some urban planning constraints. We employ particle swarm optimization and two‐step path planning to find the optimal solution, which is further interpreted as the 2D blueprint of the urban zone. During the optimization, we adopt the spatial pattern tree structure to reduce the combinational search space greatly. 3D city models in large scale are then assembled according to the 2D blueprints. Experimental results prove that our method can efficiently produce the virtual urban scene similar to that designed by urban planners. Copyright © 2008 John Wiley & Sons, Ltd.
Wei Hua 0002, Hujun Bao
Comput. Animat. Virtual Worlds2
2008 Real-time editing and relighting of homogeneous translucent materials
Rui Wang 0004, Ewen Cheslack-Postava, Rui Wang 0003, David P. Luebke, Qianyong Chen, Wei Hua 0002, Qunsheng Peng 0001, Hujun Bao
Vis. Comput.6
2007 Robust Metric Reconstruction from Challenging Video Sequences
abstract
Although camera self-calibration and metric reconstruction have been extensively studied during the past decades, automatic metric reconstruction from long video sequences with varying focal length is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to select the initial frames for initializing the projective reconstruction? What criteria should be used? How to handle the large zooming problem? How to choose an appropriate moment for upgrading the projective reconstruction to a metric one? This paper gives a careful investigation of all these issues. Practical and effective approaches are proposed. In particular, we show that existing image-based distance is not an adequate measurement for selecting the initial frames. We propose a novel measurement to take into account the zoom degree, the self-calibration quality, as well as image-based distance. We then introduce a new strategy to decide when to upgrade the projective reconstruction to a metric one. Finally, to alleviate the heavy computational cost in the bundle adjustment, a local on-demand approach is proposed. Our method is also extensively compared with the state-of-the-art commercial software to evidence its robustness and stability.
Guofeng Zhang 0001, Xueying Qin, Wei Hua 0002, Tien-Tsin Wong, Pheng-Ann Heng, Hujun Bao
CVPR3
2007 Procedural Modeling of Residential Zone Subject to Urban Planning Constraints
Wei Hua 0002, Hujun Bao
ICEC2
2007 Stereoscopic Video Synthesis from a Monocular Video
abstract
This paper presents an automatic and robust approach to synthesize stereoscopic videos from ordinary monocular videos acquired by commodity video cameras. Instead of recovering the depth map, the proposed method synthesizes the binocular parallax in stereoscopic video directly from the motion parallax in monocular video. The synthesis is formulated as an optimization problem via introducing a cost function of the stereoscopic effects, the similarity, and the smoothness constraints. The optimization selects the most suitable frames in the input video for generating the stereoscopic video frames. With the optimized selection, convincing and smooth stereoscopic video can be synthesized even by simple constant-depth warping. No user interaction is required. We demonstrate the visually plausible results obtained given the input clips acquired by ordinary handheld video camera.
Guofeng Zhang 0001, Wei Hua 0002, Xueying Qin, Tien-Tsin Wong, Hujun Bao
IEEE Trans. Vis. Comput. Graph.2
2006 Fast display of large-scale forest with fidelity
abstract
Abstract We propose a new hierarchical representation for a forest model, namely hierarchical layered depth mosaics (HLDM). Each node in the HLDM comprises a number of discrete textured quadrilaterals, called depth mosaics (DMs). The DMs are generated from the sampled depth images of the polygonal tree models. Meanwhile, their textures are compressed by a new approach accounting for occlusion. Our rendering procedure traverses the HLDM and renders the appropriate nodes according to a view‐dependent selection criterion. A blending scheme is adopted to mitigate the visual ‘popping’ caused by the transition of levels of detail. The experiment demonstrates that the viewer could interactively walk or fly above the forest with fidelity. Copyright © 2006 John Wiley & Sons, Ltd.
Huaisheng Zhang, Wei Hua 0002, Qing Wang 0042, Hujun Bao
Comput. Animat. Virtual Worlds2
2006 Synthesizing trees by plantons
Rui Wang 0004, Wei Hua 0002, Zilong Dong, Qunsheng Peng 0001, Hujun Bao
Vis. Comput.2
2005 Interactive 3D Editing on Tiled Display Wall
Xiuhui Wang, Wei Hua 0002, Hujun Bao
ICCSA (3)2
2005 Intersection fields for interactive global illumination
Zhong Ren 0001, Wei Hua 0002, Lu Chen 0001, Hujun Bao
Vis. Comput.2
2004 Generalized NURBS Curves and Surfaces
abstract
A representation, the generalized NURBS (G-NURBS), is proposed for modeling parametric curves and surfaces. G-NURBS provides a unified framework for traditional parametric curve and surface. G-NURBS surface based on arbitrary irregular mesh can represent closed surface or trimmed surface with only one surface patch.
Qing Wang 0042, Wei Hua 0002, Guiqing Li, Hujun Bao
GMP2
2004 Huge texture mapping for real-time visualization of large-scale terrain
abstract
Texture mapping greatly influences the performance of visualization in many 3D applications. Sometimes the texture data is so large that it has to be stored in slower external storage, rather than fast texture memory or host memory. In these circumstances, texture mapping becomes the performance bottleneck. In this paper, we present a compact multiresolution model, Texture Mipmap Quadtree (TMQ), to represent large-scale textures. It facilitates fast loading and pre-filtering of textures from slower external storage. Integrating continuous LOD model of terrain geometry, we present a criterion to select proper textures from TMQ according to viewing parameters during rendering stage. By exploiting temporal coherence, a dynamic texture management scheme is devised based on two-level cache hierarchy to further increase the performance of texture mapping.
Wei Hua 0002, Huaisheng Zhang, Yanqing Lu, Hujun Bao, Qunsheng Peng 0001
VRST1
2004 Anti-Aliased Rendering of Water Surface
Xueying Qin, Eihachiro Nakamae, Wei Hua 0002, Yasuo Nagai, Qunsheng Peng 0001
J. Comput. Sci. Technol.3
2003 Real-Time Ray Casting Rendering of Volume Clipping in Medical Visualization
Wei Chen 0001, Wei Hua 0002, Hujun Bao, Qunsheng Peng 0001
J. Comput. Sci. Technol.2
2002 The global occlusion map: a new occlusion culling approach
abstract
Occlusion culling is an important technique to speed up the rendering process for walkthroughs in a complex environment. In this paper, we present a new approach for occlusion culling with respect to a view cell. A compact representation, the Global Occlusion Map (GOM), is proposed for storing the global visibility information of general 3D models with respect to the view cell. The GOM provides a collection of Directional Visibility Barriers (DVB), which are virtual occluding planes aligned with the main axes of the world coordinates that act as occluders to reject invisible objects lying behind them in every direction from a view cell. Since the GOM is a two-dimensional array, its size is bounded, depending only on the number of the sampled viewing directions. Furthermore, it is easy to conservatively compress the GOM by treating it as a depth image. Due to the axial orientations of the DVBs, both the computational and storage costs for occlusion culling based on the GOM is minimized. Our implementation shows the Global Occlusion Map is effective and efficient in urban walkthrough applications.
Wei Hua 0002, Hujun Bao, Qunsheng Peng 0001, A. Robin Forrest
VRST1
2001 A New Approach of Point-Based Rendering
abstract
A new approach of point-based rendering is presented for objects with complex shape. By this approach, an object is represented by a set of discrete sample points, each point stands for a sampled surface area whose boundary is assumed to be a circle when viewed along its normal. An algorithm is proposed to build a point-based model with level of details. Efficiency of the rendering process is ensured by performing adaptive sampling of the model without calculating the sample points on the fly. To get rid of the aliasing effects due to point sampling, we have developed a delta-z-buffer and a flag buffer to smooth the boundary of two adjacent sampled areas and to fill the holes that may appear inside the image of an object. Experimental results show that our new approach is capable of rendering complex objects consisting of only 50,000 points with acceptable image quality, and high rendering efficiency.
Qunsheng Peng 0001, Wei Hua 0002, Xuehui Yang
Computer Graphics International2