VLDB 2026 Research / reviewers in the wild / expert
Boxi Wu 0001
dblp:23/1091-1
· DBLP profile ↗
24ranked-venue papers
3as first author
24since 2021 · last 2026
0000-0003-4494-193XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Trackingabstract3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters, to predict trajectories. While effective in simple scenarios, this approach often struggles in crowded environments or with inaccurate detections, as it overlooks the rich geometric relationships between objects. This highlights the need to leverage spatial cues. However, existing geometry-aware methods can be susceptible to interference from irrelevant objects, leading to ambiguous features and incorrect associations. To address this, we propose focusing on cue-consistency: identifying and matching stable spatial patterns over time. We introduce the Dynamic Scene Cue-Consistency Tracker (DSC-Track) to implement this principle. Firstly, we design a unified spatiotemporal encoder using Point Pair Features (PPF) to learn discriminative trajectory embeddings while suppressing interference. Secondly, our cue-consistency transformer module explicitly aligns consistent feature representations between historical tracks and current detections. Finally, a dynamic update mechanism preserves salient spatiotemporal information for stable online tracking. Extensive experiments on the nuScenes and Waymo Open Datasets validate the effectiveness and robustness of our approach. On the nuScenes benchmark, for instance, our method achieves state-of-the-art performance, reaching 73.2% and 70.3% AMOTA on the validation and test sets, respectively. Boxi Wu 0001, Tu Zheng, Wang Yunhua, Zheng Yang 0008 |
AAAI | 3 |
| 2026 | Vidsketch: Hand-drawn sketch-driven video generation with diffusion control
Lifan Jiang, Boxi Wu 0001, Deng Cai 0001 |
Neural Networks | 3 |
| 2026 | ConsistencyTrack: A robust multi-object tracker with a generation strategy of consistency model
Lifan Jiang, Zhihui Wang 0003, Siqi Yin, Guangxiao Ma, Peng Zhang 0057, Boxi Wu 0001 |
Pattern Recognit. | 6 |
| 2026 | Beyond Fidelity: Diverse Image Synthesis via Retrieval-Augmented DiffusionabstractImage synthesis is a key application of generative AI. It can help reduce overfitting and the high cost of collecting real-world data for downstream discriminative models. However, current methods mainly focus on making images look realistic and ignore their true goal: improving downstream model generalization and robustness. We find that existing approaches tend to produce high-fidelity synthetic images that closely resemble the original data. This limits their value for improving downstream task performance. To overcome this, we argue that diversity, not just fidelity, must guide synthetic data generation if it is to truly complement human-collected datasets. In this paper, we introduce a Retrieval-Augmented Generation framework for diverse diffusion-based image synthesis. At each generation step, we retrieve the top-K most similar samples in feature space from both real and previously generated images. We then apply an Anti-Attention mechanism that actively pushes the new image away from these retrieved samples in feature space, maximizing dissimilarity. We propose novel evaluation metrics to assess image synthesis diversity and demonstrate significant improvements over existing benchmarks. Moreover, downstream models trained with our synthetic data achieved a 1.9% absolute accuracy gain on standard benchmarks, outperforming existing synthesis techniques. Linxuan Xia, Boxi Wu 0001, Tianrun Wu, Deng Cai 0001, Wei Liu 0005 |
IEEE Trans. Image Process. | 2 |
| 2025 | Local Conditional Controlling for Text-to-Image Diffusion ModelsabstractDiffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the generation process together with text prompts to obtain desired images. This controlling process is globally operated on the entire image, which limits the flexibility of control regions. In this paper, we explore a novel and practical task setting: local control. It focuses on controlling specific local region according to user-defined image conditions, while the remaining regions are only conditioned by the original text prompt. However, it is non-trivial to achieve it. The naive manner of directly adding local conditions may lead to the local control dominance problem, which forces the model to focus on the controlled region and neglect object generation in other regions. To mitigate this problem, we propose Regional Discriminate Loss to update the noised latents, aiming at enhanced object generation in non-control regions. Furthermore, the proposed Focused Token Response suppresses weaker attention scores which lack the strongest response to enhance object distinction and reduce duplication. Lastly, we adopt Feature Mask Constraint to reduce quality degradation in images caused by information differences across the local control region. All proposed strategies are operated at the inference stage. Extensive experiments demonstrate that our method can synthesize high-quality images aligned with the text prompt under local control conditions. Yang Yang 0002, Zekai Luo, Hengjia Li, Zheng Yang 0008, Xiaofei He 0001, Wei Zhao 0019, Qinglin Lu, Wei Liu 0005, Boxi Wu 0001 |
AAAI | 12 |
| 2025 | Object-level Data Augmentation for Visual 3D Object Detection in Autonomous DrivingabstractData augmentation plays an important role in visual-based 3D object detection. Existing detectors typically employ image/BEV-level data augmentation techniques, failing to utilize flexible object-level augmentations because of 2D-3D inconsistencies. This limitation hinders us from increasing the diversity of training data. To alleviate this issue, we propose an object-level data augmentation approach that incorporates scene reconstruction and neural scene rendering. Specifically, we reconstruct the scene and objects by extracting image features from sequences and aligning them with associated LiDAR point clouds. This approach is intended to conduct the editing process within a 3D space, allowing for flexible object manipulation. Additionally, we introduce a neural scene renderer to project the edited 3D scene onto a specified camera plane and render it onto a 2D image. Combined with scene reconstruction, it overcomes the challenges stemming from 2D/3D inconsistencies, enabling the generation of object-level augmented images with corresponding labels for model training. To validate the proposed method, we apply our method to various multi-camera 3D object detectors, consistently boosting the performance. Junkai Xu, Zheng Yang 0008, Xiaofei He 0001, Boxi Wu 0001 |
ICASSP | 6 |
| 2025 | Magicid: Hybrid Preference Optimization for Id-Consistent and Dynamic-Preserved Video Customization
Hengjia Li, Lifan Jiang, Hongwei Yi, Boxi Wu 0001, Deng Cai 0001 |
ICCV | 6 |
| 2025 | PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic DegradationabstractThe current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies in maintaining high ID fidelity consistently while preserving the original motion dynamic and semantic following after the identity injection. Current video identity customization methods mainly rely on reconstructing given identity images on text-to-image models, which have a divergent distribution with the T2V model. This process introduces a tuning-inference gap, leading to dynamic and semantic degradation. To tackle this problem, we propose a novel framework, dubbed $\textbf{PersonalVideo}$, that applies a mixture of reward supervision on synthesized videos instead of the simple reconstruction objective on images. Specifically, we first incorporate identity consistency reward to effectively inject the reference's identity without the tuning-inference gap. Then we propose a novel semantic consistency reward to align the semantic distribution of the generated videos with the original T2V model, which preserves its dynamic and semantic following capability during the identity injection. With the non-reconstructive reward training, we further employ simulated prompt augmentation to reduce overfitting by supervising generated results in more semantic scenarios, gaining good robustness even with only a single reference image. Extensive experiments demonstrate our method's superiority in delivering high identity faithfulness while preserving the inherent video generation qualities of the original T2V model, outshining prior methods. Hengjia Li, Haonan Qiu, Shiwei Zhang 0001, Xiang Wang 0012, Yujie Wei 0001, Yingya Zhang, Boxi Wu 0001, Deng Cai 0001 |
ICCV | 8 |
| 2025 | Vividportraits: Face Parsing Guided Portrait AnimationabstractPortrait animation aims to transfer the facial expressions and movements of a target character onto a reference character. This task presents two main challenges: accurately transferring motion and expressions while fully preserving the identity features of the reference portrait. We introduce Vividportraits, a diffusion-based model designed to effectively meet these objectives. In contrast to existing methods that rely on sparse representations such as facial landmarks, our approach leverages facial parsing maps for motion guidance, enabling a more precise conveyance of subtle expressions. A random scaling technique is applied during training to prevent the model from internalizing identity-specific features from the driving images. Furthermore, we perform foreground-background segmentation on the reference portrait to reduce data redundancy. The long-video generation process is refined to improve consistency across sequences. Our model, exclusively trained on public datasets, demonstrates superior performance relative to current state-of-the-art methods, achieving a notable 8% improvement in expression metric. More visual results are available on the anonymous website https://www.vividportraits.cn. Xuze Tian, Jinshan Zhang 0001, Boxi Wu 0001, Meng Xi 0002, Zejian Li, Jianwei Yin |
ICMR | 4 |
| 2025 | Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsabstractDirect preference optimization (DPO) is an effective method for aligning generative models with human preferences and has been successfully applied to fine‑tune text‑to‑image diffusion models. Its practical adoption, however, is hindered by a labor‑intensive pipeline that first produces a large set of candidate images and then requires humans to rank them pairwise. We address this bottleneck with self‑supervised direct preference optimization, a new paradigm that removes the need for any pre‑generated images or manual ranking. During training, we create preference pairs on the fly through self‑supervised image transformations, allowing the model to learn from fresh and diverse comparisons at every iteration. This online strategy eliminates costly data collection and annotation while remaining plug‑and‑play for any text‑to‑image diffusion method. Surprisingly, the on‑the‑fly pairs produced by the proposed method not only match but exceed the effectiveness of conventional DPO, which we attribute to the greater diversity of preferences sampled during training. Extensive experiments with Stable Diffusion 1.5 and Stable Diffusion XL confirm that our method delivers substantial gains. Boxi Wu 0001, Xiaofei He 0001 |
NeurIPS | 2 |
| 2025 | LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion ModelsabstractCustomization generation techniques have significantly advanced the synthesis of specific concepts across varied contexts. Multi-concept customization emerges as the challenging task within this domain. Existing approaches often rely on training a fusion matrix of multiple Low-Rank Adaptations (LoRAs) to merge various concepts into a single image. However, we identify this straightforward method faces two major challenges: 1) concept confusion, where the model struggles to preserve distinct individual characteristics, and 2) concept vanishing, where the model fails to generate the intended subjects. To address these issues, we introduce LoRA-Composer, a training-free framework designed for seamlessly integrating multiple LoRAs, thereby enhancing the harmony among different concepts within generated images. LoRA-Composer addresses concept vanishing through concept injection constraints, enhancing visibility via an expanded cross-attention mechanism. To combat concept confusion, concept isolation constraints are introduced, refining the self-attention computation. Furthermore, we propose two inference techniques to accelerate inference speed without performance degradation and enhance the accuracy of the generated region, respectively. Extensive experiments demonstrate that LoRA-Composer significantly outperforms standard baselines, especially in scenarios without image-based conditions such as canny edge or pose estimation. Yang Yang 0002, Chaotian Song, Hengjia Li, Qinglin Lu, Deng Cai 0001, Xiaofei He 0001, Boxi Wu 0001, Wei Liu 0005 |
IEEE Trans. Image Process. | 11 |
| 2024 | Towards Fine-Grained HBOE with Rendered Orientation Set and Laplace SmoothingabstractHuman body orientation estimation (HBOE) aims to estimate the orientation of a human body relative to the camera’s frontal view. Despite recent advancements in this field, there still exist limitations in achieving fine-grained results. We identify certain defects and propose corresponding approaches as follows: 1). Existing datasets suffer from non-uniform angle distributions, resulting in sparse image data for certain angles. To provide comprehensive and high-quality data, we introduce RMOS (Rendered Model Orientation Set), a rendered dataset comprising 150K accurately labeled human instances with a wide range of orientations. 2). Directly using one-hot vector as labels may overlook the similarity between angle labels, leading to poor supervision. And converting the predictions from radians to degrees enlarges the regression error. To enhance supervision, we employ Laplace smoothing to vectorize the label, which contains more information. For fine-grained predictions, we adopt weighted Smooth-L1-loss to align predictions with the smoothed-label, thus providing robust supervision. 3). Previous works ignore body-part-specific information, resulting in coarse predictions. By employing local-window self-attention, our model could utilize different body part information for more precise orientation estimations. We validate the effectiveness of our method in the benchmarks with extensive experiments and show that our method outperforms state-of-the-art. Project is available at: https://github.com/Whalesong-zrs/Towards-Fine-grained-HBOE. Ruisi Zhao, Zheng Yang 0008, Binbin Lin 0001, Xiaohui Zhong, Xiaobo Ren, Deng Cai 0001, Boxi Wu 0001 |
AAAI | 8 |
| 2024 | Learning Occupancy for Monocular 3D Object DetectionabstractMonocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth esti-mates to facilitate the learning, but often fail to fully ex-ploit the benefits of three-dimensional feature extraction in frustum and 3D space. In this paper, we propose Occu- pancyM3D, a method of learning occupancy for monocu-lar 3D detection. It directly learns occupancy in frustum and 3D space, leading to more discriminative and informative 3D features and representations. Specifically, by using synchronized raw sparse LiDAR point clouds, we define the space status and generate voxel-based occupancy labels. We formulate occupancy prediction as a simple classification problem and design associated occupancy losses. Re-sulting occupancy estimates are employed to enhance orig-inal frustum/3D features. As a result, experiments on KITTI and Waymo open datasets demonstrate that the proposed method achieves a new state of the art and surpasses other methods by a significant margin. Junkai Xu, Zheng Yang 0008, Xiaopei Wu, Wei Qian 0003, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001 |
CVPR | 8 |
| 2024 | TASeg: Temporal Aggregation Network for LiDAR Semantic SegmentationabstractTraining deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the spar-sity problem as it makes the input signal denser. However, previous multi-frame fusion algorithms fall short in utilizing sufficient temporal information due to the memory constraint, and they also ignore the informative temporal images. To fully exploit rich information hidden in long-term temporal point clouds and images, we present the Temporal Aggre-gation Network, termed TASeg. Specifically, we propose a Temporal LiDAR Aggregation and Distillation (TLAD) algorithm, which leverages historical priors to assign dif-ferent aggregation steps for different classes. It can largely reduce memory and time overhead while achieving higher accuracy. Besides, TLAD trains a teacher injected with gt priors to distill the model, further boosting the performance. To make full use of temporal images, we design a Temporal Image Aggregation and Fusion (TIAF) module, which can greatly expand the camera FOVand enhance the present features. Temporal LiDAR points in the camera FOV are used as mediums to transform temporal image features to the present coordinate for temporal multi-modal fusion. Moreover, we develop a Static-Moving Switch Augmentation (SMSA) algorithm, which utilizes sufficient temporal information to enable objects to switch their motion states freely, thus greatly increasing static and moving training samples. Our TASeg ranks 1st††the date of CVPR deadline, i.e., 2023-11-18 07:59 AM UTC. on three challenging tracks, i.e., SemanticKITTI single-scan track, multi-scan track and nuScenes LiDAR segmentation track, strongly demonstrating the superiority of our method. Codes are available at https://github.com/LittlePey/TASeg. Xiaopei Wu, Yuenan Hou, Xiaoshui Huang, Binbin Lin 0001, Tong He 0001, Xinge Zhu, Yuexin Ma, Boxi Wu 0001, Haifeng Liu 0001, Deng Cai 0001, Wanli Ouyang |
CVPR | 8 |
| 2024 | Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-Dataset 3D Object DetectionabstractRecent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels, i.e., 3D boxes, to supervise models for the target domain. However, this selection process inevitably introduces unreliable 3D boxes, in which 3D points cannot be definitively assigned as foreground or background. Previous techniques mitigate this by reweighting these boxes as pseudo labels, but these boxes can still poison the training process. To resolve this problem, in this paper, we propose a novel pseudo label refinery framework. Specifically, in the selection process, to improve the reliability of pseudo boxes, we propose a complementary augmentation strategy. This strategy involves either removing all points within an unre-liable box or replacing it with a high-confidence box. More-over, the point numbers of instances in high-beam datasets are considerably higher than those in low-beam datasets, also degrading the quality of pseudo labels during the training process. We alleviate this issue by generating additional proposals and aligning RoI features across different domains. Experimental results demonstrate that our method effectively enhances the quality of pseudo labels and consistently surpasses the state-of-the-art methods on six autonomous driving benchmarks. Code will be available at https://github.com/Zhanwei-Z/PERE. Zhanwei Zhang, Minghao Chen 0001, Hengjia Li, Binbin Lin 0001, Ping Li 0006, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001 |
CVPR | 9 |
| 2024 | CrossFormer++: A Versatile Vision Transformer Hinging on Cross-Scale AttentionabstractWhile features of different scales are perceptually important to visual inputs, existing vision transformers do not yet take advantage of them explicitly. To this end, we first propose a cross-scale vision transformer, CrossFormer. It introduces a cross-scale embedding layer (CEL) and a long-short distance attention (LSDA). On the one hand, CEL blends each token with multiple patches of different scales, providing the self-attention module itself with cross-scale features. On the other hand, LSDA splits the self-attention module into a short-distance one and a long-distance counterpart, which not only reduces the computational burden but also keeps both small-scale and large-scale features in the tokens. Moreover, through experiments on CrossFormer, we observe another two issues that affect vision transformers' performance, i.e., the enlarging self-attention maps and amplitude explosion. Thus, we further propose a progressive group size (PGS) paradigm and an amplitude cooling layer (ACL) to alleviate the two issues, respectively. The CrossFormer incorporating with PGS and ACL is called CrossFormer++. Extensive experiments show that CrossFormer++ outperforms the other vision transformers on image classification, object detection, instance segmentation, and semantic segmentation tasks. Wenxiao Wang 0001, Wei Chen 0005, Qibo Qiu, Long Chen 0016, Boxi Wu 0001, Binbin Lin 0001, Xiaofei He 0001, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Temporal Feature Fusion for 3D Detection in Monocular VideoabstractPrevious monocular 3D detection works focus on the single frame input in both training and inference. In real-world applications, temporal and motion information naturally exists in monocular video. It is valuable for 3D detection but under-explored in monocular works. In this paper, we propose a straightforward and effective method for temporal feature fusion, which exhibits low computation cost and excellent transferability, making it conveniently applicable to various monocular models. Specifically, with the help of optical flow, we transform the backbone features produced by prior frames and fuse them into the current frame. We introduce the scene feature propagating mechanism, which accumulates history scene features without extra time-consuming. In this process, occluded areas are removed via forward-backward scene consistency. Our method naturally introduces valuable temporal features, facilitating 3D reasoning in monocular 3D detection. Furthermore, accumulated history scene features via scene propagating mitigate heavy computation overheads for video processing. Experiments are conducted on variant baselines, which demonstrate that the proposed method is model-agonistic and can bring significant improvement to multiple types of single-frame methods. Zheng Yang 0008, Binbin Lin 0001, Xiaofei He 0001, Boxi Wu 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Towards In-Distribution Compatible Out-of-Distribution DetectionabstractDeep neural network, despite its remarkable capability of discriminating targeted in-distribution samples, shows poor performance on detecting anomalous out-of-distribution data. To address this defect, state-of-the-art solutions choose to train deep networks on an auxiliary dataset of outliers. Various training criteria for these auxiliary outliers are proposed based on heuristic intuitions. However, we find that these intuitively designed outlier training criteria can hurt in-distribution learning and eventually lead to inferior performance. To this end, we identify three causes of the in-distribution incompatibility: contradictory gradient, false likelihood, and distribution shift. Based on our new understandings, we propose a new out-of-distribution detection method by adapting both the top-design of deep models and the loss function. Our method achieves in-distribution compatibility by pursuing less interference with the probabilistic characteristic of in-distribution features. On several benchmarks, our method not only achieves the state-of-the-art out-of-distribution detection performance but also improves the in-distribution accuracy. Boxi Wu 0001, Jie Jiang 0015, Haidong Ren, Zifan Du, Wenxiao Wang 0001, Zhifeng Li 0001, Deng Cai 0001, Xiaofei He 0001, Binbin Lin 0001, Wei Liu 0005 |
AAAI | 1 |
| 2023 | CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training costs. In this paper, we explore the potential of Contrastive Language-Image Pre-training models (CLIP) to localize different categories with only image-level labels and without further training. To efficiently generate high-quality segmentation masks from CLIP, we propose a novel WSSS framework called CLIP-ES. Our framework improves all three stages of WSSS with special designs for CLIP: 1) We introduce the softmax function into GradCAM and exploit the zero-shot ability of CLIP to suppress the confusion caused by non-target classes and backgrounds. Mean-while, to take full advantage of CLIP, we re-explore text inputs under the WSSS setting and customize two text-driven strategies: sharpness-based prompt selection and synonym fusion. 2) To simplify the stage of CAM refinement, we propose a real-time class-aware attention-based affinity (CAA) module based on the inherent multi-head self-attention (MHSA) in CLIP- ViTs. 3) When training the final segmentation model with the masks generated by CLIP, we introduced a confidence-guided loss (CGL) focus on confident regions. Our CLIP-ES achieves SOTA performance on Pascal VOC 2012 and MS COCO 2014 while only taking 10% time of previous methods for the pseudo mask generation. Code is available at https://github.com/linyq2117/CLIP-ES. Yuqi Lin, Minghao Chen 0001, Wenxiao Wang 0001, Boxi Wu 0001, Binbin Lin 0001, Haifeng Liu 0001, Xiaofei He 0001 |
CVPR | 4 |
| 2023 | GD-MAE: Generative Decoder for MAE Pre-Training on LiDAR Point CloudsabstractDespite the tremendous progress of Masked Autoencoders (MAE) in developing vision tasks such as image and video, exploring MAE in large-scale 3D point clouds remains challenging due to the inherent irregularity. In contrast to previous 3D MAE frameworks, which either design a complex decoder to infer masked information from maintained regions or adopt sophisticated masking strategies, we instead propose a much simpler paradigm. The core idea is to apply a Generative Decoder for MAE (GD-MAE) to automatically merges the surrounding context to restore the masked geometric knowledge in a hierarchical fusion manner. In doing so, our approach is free from introducing the heuristic design of decoders and enjoys the flexibility of exploring various masking strategies. The corresponding part costs less than 12% latency compared with conventional methods, while achieving better performance. We demonstrate the efficacy of the proposed method on several large-scale benchmarks: Waymo, KITTI, and ONCE. Consistent improvement on downstream detection tasks illustrates strong robustness and generalization capability. Not only our method reveals state-of-the-art results, but remarkably, we achieve comparable accuracy even with 20% of the labeled data on the Waymo dataset. Code will be released. Honghui Yang, Tong He 0001, Boxi Wu 0001, Binbin Lin 0001, Xiaofei He 0001, Wanli Ouyang |
CVPR | 5 |
| 2023 | One-shot Implicit Animatable Avatars with Model-based PriorsabstractExisting neural rendering methods for creating human avatars typically either require dense input signals such as video or multi-view images, or leverage a learned prior from large-scale specific 3D human datasets such that reconstruction can be performed with sparse-view inputs. Most of these methods fail to achieve realistic reconstruction when only a single image is available. To enable the data-efficient creation of realistic anima table 3D humans, we propose ELICIT, a novel method for learning human-specific neural radiance fields from a single image. Inspired by the fact that humans can effortlessly estimate the body geometry and imagine full-body clothing from a single image, we leverage two priors in ELICIT: 3D geometry prior and visual semantic prior. Specifically, ELICIT utilizes the 3D body shape geometry prior from a skinned vertex-based template model (i.e., SMPL) and implements the visual clothing semantic prior with the CLIP-based pre-trained models. Both priors are used to jointly guide the optimization for creating plausible content in the invisible areas. Taking advantage of the CLIP models, ELICIT can use text descriptions to generate text-conditioned unseen regions. In order to further improve visual details, we propose a segmentation-based sampling strategy that locally refines different parts of the avatar. Comprehensive evaluations on multiple popular benchmarks, including ZJU-MoCAP, Human3.6M, and DeepFashion, show that ELICIT outperforms strong baseline methods of avatar creation when only a single image is available. The code is public for research purposes at https://huangyangyi.github.io/ELICIT Yangyi Huang, Hongwei Yi, Weiyang Liu, Boxi Wu 0001, Wenxiao Wang 0001, Binbin Lin 0001, Debing Zhang, Deng Cai 0001 |
ICCV | 5 |
| 2022 | Towards Efficient Adversarial Training on Vision Transformers
Boxi Wu 0001, Jindong Gu, Zhifeng Li 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005 |
ECCV (13) | 1 |
| 2022 | WeakM3D: Towards Weakly Supervised Monocular 3D Object Detection
Senbo Yan, Boxi Wu 0001, Zheng Yang 0008, Xiaofei He 0001, Deng Cai 0001 |
ICLR | 3 |
| 2021 | Do Wider Neural Networks Really Help Adversarial Robustness?abstractAdversarial training is a powerful type of defense against adversarial examples. Previous empirical results suggest that adversarial training requires wider networks for better performances. However, it remains elusive how does neural network width affect model robustness. In this paper, we carefully examine the relationship between network width and model robustness. Specifically, we show that the model robustness is closely related to the tradeoff between natural accuracy and perturbation stability, which is controlled by the robust regularization parameter λ. With the same λ, wider networks can achieve better natural accuracy but worse perturbation stability, leading to a potentially worse overall model robustness. To understand the origin of this phenomenon, we further relate the perturbation stability with the network's local Lipschitzness. By leveraging recent results on neural tangent kernels, we theoretically show that wider networks tend to have worse perturbation stability. Our analyses suggest that: 1) the common strategy of first fine-tuning λ on small networks and then directly use it for wide model training could lead to deteriorated model robustness; 2) one needs to properly enlarge λ to unleash the robustness potential of wider models fully. Finally, we propose a new Width Adjusted Regularization (WAR) method that adaptively enlarges λ on wide models and significantly saves the tuning time. Boxi Wu 0001, Deng Cai 0001, Xiaofei He 0001, Quanquan Gu |
NeurIPS | 1 |