EDBT 2026 Demo / reviewers in the wild / expert
Weixuan Sun
dblp:186/6724
· DBLP profile ↗
18ranked-venue papers
4as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio-Visual Segmentation with Semantics
Jinxing Zhou, Xuyang Shen, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong |
Int. J. Comput. Vis. | 5 |
| 2024 | NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation
Ruikai Cui, Weizhe Liu, Weixuan Sun, Senbo Wang, Taizhang Shang, Yang Li 0193, Xibin Song, Han Yan 0004, Zhennan Wu, Shenzhou Chen, Hongdong Li, Pan Ji |
ECCV (19) | 3 |
| 2024 | CO2: Efficient Distributed Training with Full Communication-Computation OverlapabstractThe fundamental success of large language models hinges upon the efficacious implementation of large-scale distributed training techniques. Nevertheless, building a vast, high-performance cluster featuring high-speed communication interconnectivity is prohibitively costly, and accessible only to prominent entities. In this work, we aim to lower this barrier and democratize large-scale training with limited bandwidth clusters. We propose a new approach called CO2 that introduces local-updating and asynchronous communication to the distributed data-parallel training, thereby facilitating the full overlap of COmmunication with COmputation. CO2 is able to attain a high scalability even on extensive multi-node clusters constrained by very limited communication bandwidth. We further propose the staleness gap penalty and outer momentum clipping techniques together with CO2 to bolster its convergence and training stability. Besides, CO2 exhibits seamless integration with well-established ZeRO-series optimizers which mitigate memory consumption of model states with large model training. We also provide a mathematical proof of convergence, accompanied by the establishment of a stringent upper bound. Furthermore, we validate our findings through an extensive set of practical experiments encompassing a wide range of tasks in the fields of computer vision and natural language processing. These experiments serve to demonstrate the capabilities of CO2 in terms of convergence, generalization, and scalability when deployed across configurations comprising up to 128 A100 GPUs. The outcomes emphasize the outstanding capacity of CO2 to hugely improve scalability, no matter on clusters with 800Gbps RDMA or 80Gbps TCP/IP inter-node connections. Weigao Sun, Zhen Qin 0003, Weixuan Sun, Shidi Li, Dong Li 0033, Xuyang Shen, Yu Qiao 0001, Yiran Zhong |
ICLR | 3 |
| 2024 | Various Lengths, Constant Speed: Efficient Language Modeling with Lightning AttentionabstractWe present Lightning Attention, the first linear attention implementation that maintains a constant training speed for various sequence lengths under fixed memory consumption. Due to the issue with cumulative summation operations (cumsum), previous linear attention implementations cannot achieve their theoretical advantage in a casual setting. However, this issue can be effectively solved by utilizing different attention calculation strategies to compute the different parts of attention. Specifically, we split the attention calculation into intra-blocks and inter-blocks and use conventional attention computation for intra-blocks and linear attention kernel tricks for inter-blocks. This eliminates the need for cumsum in the linear attention calculation. Furthermore, a tiling technique is adopted through both forward and backward procedures to take full advantage of the GPU hardware. To enhance accuracy while preserving efficacy, we introduce TransNormerLLM (TNL), a new architecture that is tailored to our lightning attention. We conduct rigorous testing on standard and self-collected datasets with varying model sizes and sequence lengths. TNL is notably more efficient than other language models. In addition, benchmark results indicate that TNL performs on par with state-of-the-art LLMs utilizing conventional transformer structures. The source code is released at github.com/OpenNLPLab/TransnormerLLM. Zhen Qin 0003, Weigao Sun, Dong Li 0033, Xuyang Shen, Weixuan Sun, Yiran Zhong |
ICML | 5 |
| 2024 | LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single ImageabstractLarge Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies, stemming from the inherent challenges of deducing 3D shapes solely from image data. In this work, we introduce a novel framework, the Large Image and Point Cloud Alignment Model (LAM3D), which utilizes 3D point cloud data to enhance the fidelity of generated 3D meshes. Our methodology begins with the development of a point-cloud-based network that effectively generates precise and meaningful latent tri-planes, laying the groundwork for accurate 3D mesh reconstruction. Building upon this, our Image-Point-Cloud Feature Alignment technique processes a single input image, aligning to the latent tri-planes to imbue image features with robust 3D information. This process not only enriches the image features but also facilitates the production of high-fidelity 3D meshes without the need for multi-view input, significantly reducing geometric distortions. Our approach achieves state-of-the-art high-fidelity 3D mesh reconstruction from a single image in just 6 seconds, and experiments on various datasets demonstrate its effectiveness. Ruikai Cui, Xibin Song, Weixuan Sun, Senbo Wang, Weizhe Liu, Shenzhou Chen, Taizhang Shang, Yang Li 0193, Nick Barnes, Hongdong Li, Pan Ji |
NeurIPS | 3 |
| 2024 | Frankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-PlaneabstractWe present Frankenstein, a diffusion-based framework that can generate semantic-compositional 3D scenes in a single pass. Unlike existing methods that output a single, unified 3D shape, Frankenstein simultaneously generates multiple separated shapes, each corresponding to a semantically meaningful part. The 3D scene information is encoded in one single triplane tensor, from which multiple Signed Distance Function (SDF) fields can be decoded to represent the compositional shapes. During training, an auto-encoder compresses tri-planes into a latent space, and then the denoising diffusion process is employed to approximate the distribution of the compositional scenes. Frankenstein demonstrates promising results in generating room interiors as well as human avatars with automatically separated parts. The generated scenes facilitate many downstream applications, such as part-wise re-texturing, object rearrangement in the room or avatar cloth re-targeting. Han Yan 0004, Yang Li 0193, Zhennan Wu, Shenzhou Chen, Weixuan Sun, Taizhang Shang, Weizhe Liu, Xiaqiang Dai, Chao Ma 0004, Hongdong Li, Pan Ji |
SIGGRAPH Asia | 5 |
| 2024 | Bi-directional Training for Composed Image Retrieval via Text Prompt LearningabstractComposed image retrieval searches for a target image based on a multi-modal user query comprised of a reference image and modification text describing the desired changes. Existing approaches to solving this challenging task learn a mapping from the (reference image, modification text)-pair to an image embedding that is then matched against a large image corpus. One area that has not yet been explored is the reverse direction, which asks the question, what reference image when modified as described by the text would produce the given target image? In this work we propose a bi-directional training scheme that leverages such reversed queries and can be applied to existing composed image retrieval architectures with minimum changes, which improves the performance of the model. To encode the bi-directional query we prepend a learnable token to the modification text that designates the direction of the query and then finetune the parameters of the text embedding module. We make no other changes to the network architecture. Experiments on two standard datasets show that our novel approach achieves improved performance over a baseline BLIP-based model that itself already achieves competitive performance. Our code is released at https://github.com/Cuberick-Orion/Bi-Blip4CIR Zheyuan Liu 0002, Weixuan Sun, Yicong Hong, Damien Teney, Stephen Gould |
WACV | 2 |
| 2024 | BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane ExtrapolationabstractWe present BlockFusion, a diffusion-based model that generates 3D scenes as unit blocks and seamlessly incorporates new blocks to extend the scene. BlockFusion is trained using datasets of 3D blocks that are randomly cropped from complete 3D scene meshes. Through per-block fitting, all training blocks are converted into the hybrid neural fields: with a tri-plane containing the geometry features, followed by a Multi-layer Perceptron (MLP) for decoding the signed distance values. A variational auto-encoder is employed to compress the tri-planes into the latent tri-plane space, on which the denoising diffusion process is performed. Diffusion applied to the latent representations allows for high-quality and diverse 3D scene generation. To expand a scene during generation, one needs only to append empty blocks to overlap with the current scene and extrapolate existing latent tri-planes to populate new blocks. The extrapolation is done by conditioning the generation process with the feature samples from the overlapping tri-planes during the denoising iterations. Latent tri-plane extrapolation produces semantically and geometrically meaningful transitions that harmoniously blend with the existing scene. A 2D layout conditioning mechanism is used to control the placement and arrangement of scene elements. Experimental results indicate that BlockFusion is capable of generating diverse, geometrically consistent and unbounded large 3D scenes with unprecedented high-quality shapes in both indoor and outdoor scenarios. Zhennan Wu, Yang Li 0193, Han Yan 0004, Taizhang Shang, Weixuan Sun, Senbo Wang, Ruikai Cui, Weizhe Liu, Hiroyuki Sato 0002, Hongdong Li, Pan Ji |
ACM Trans. Graph. | 5 |
| 2023 | Learning Audio-Visual Source Localization via False Negative Aware Contrastive LearningabstractSelf-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the help of contrastive learning, which assumes only the audio and visual contents from the same video are positive samples for each other. However, this assumption would suffer from false negative samples in real-world training. For example, for an audio sample, treating the frames from the same audio class as negative samples may mislead the model and therefore harm the learned representations (e.g., the audio of a siren wailing may reasonably correspond to the ambulances in multiple images). Based on this observation, we propose a new learning strategy named False Negative Aware Contrastive (FNAC) to mitigate the problem of misleading the training with such false negative samples. Specifically, we utilize the intra-modal similarities to identify potentially similar samples and construct corresponding adjacency matrices to guide contrastive learning. Further, we propose to strengthen the role of true negative samples by explicitly leveraging the visual features of sound sources to facilitate the differentiation of authentic sounding source regions. FNAC achieves state-of-the-art performances on Flickr-SoundNet, VGG-Sound, and AVSBench, which demonstrates the effectiveness of our method in mitigating the false negative issue. The code is available at https://github.com/OpenNLPLab/FNAC_AVL. Weixuan Sun, Zheyuan Liu 0002, Yiran Zhong, Tianpeng Feng, Yandong Guo, Nick Barnes |
CVPR | 1 |
| 2023 | Toeplitz Neural Network for Sequence Modeling
Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dong Li 0033, Dongxu Li 0003, Yuchao Dai, Lingpeng Kong, Yiran Zhong |
ICLR | 3 |
| 2023 | Matting Moments: A Unified Data-Driven Matting Engine for Mobile AIGC in Photo GalleryabstractImage matting is a fundamental technique in visual understanding and has become one of the most significant capabilities in mobile phones. Despite the development of mobile storage and computing power, achieving diverse mobile Artificial Intelligence Generated Content (AIGC) applications remains a great challenge. To address this issue, we present an innovative demonstration of an automatic system called "Matting Moments" that enables automatic image editing based on matting models in different scenarios. Coupled with accurate and refined matting subjects, our system provides visual element editing abilities and backend services for distribution and recommendation that respond to emotional expressions. Our system comprises three components: 1) photo content structuring, 2) data-driven matting engine, and 3) AIGC functions for generation, which automatically achieve diverse photo beautification in the gallery. This system offers a unified framework that guides consumers to obtain intelligent recommendations with beautifully generated contents, helping them enjoy the moments and memories of their present life. Fanyi Wang, Weixuan Sun, Jingwen Su, Xinjie Feng, Zhengxia Zou |
IJCAI | 3 |
| 2023 | Vicinity Vision TransformerabstractVision transformers have shown great success on numerous computer vision tasks. However, their central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory footprint being quadratic. Linear attention was introduced in natural language processing (NLP) which reorders the self-attention mechanism to mitigate a similar issue, but directly applying existing linear attention to vision may not lead to satisfactory results. We investigate this problem and point out that existing linear attention methods ignore an inductive bias in vision tasks, i.e., 2D locality. In this article, we propose Vicinity Attention, which is a type of linear attention that integrates 2D locality. Specifically, for each image patch, we adjust its attention weight based on its 2D Manhattan distance from its neighbouring patches. In this case, we achieve 2D locality in a linear complexity where the neighbouring image patches receive stronger attention than far away patches. In addition, we propose a novel Vicinity Attention Block that is comprised of Feature Reduction Attention (FRA) and Feature Preserving Connection (FPC) in order to address the computational bottleneck of linear attention approaches, including our Vicinity Attention, whose complexity grows quadratically with respect to the feature dimension. The Vicinity Attention Block computes attention in a compressed feature space with an extra skip connection to retrieve the original feature distribution. We experimentally validate that the block further reduces computation without degenerating the accuracy. Finally, to validate the proposed methods, we build a linear vision transformer backbone named Vicinity Vision Transformer (VVT). Targeting general vision tasks, we build VVT in a pyramid structure with progressively reduced sequence length. We perform extensive experiments on CIFAR-100, ImageNet-1 k, and ADE20 K datasets to validate the effectiveness of our method. Our method has a slower growth rate in terms of computational overhead than previous transformer-based and convolution-based networks when the input resolution increases. In particular, our approach achieves state-of-the-art image classification accuracy with 50% fewer parameters than previous approaches. Weixuan Sun, Zhen Qin 0003, Yi Zhang 0137, Kaihao Zhang, Nick Barnes, Stanley T. Birchfield, Lingpeng Kong, Yiran Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Audio-Visual Segmentation
Jinxing Zhou, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong |
ECCV (37) | 4 |
| 2022 | The Devil in Linear TransformerabstractLinear transformers aim to reduce the quadratic space-time complexity of vanilla transformers.However, they usually suffer from degraded performances on various tasks and corpora.In this paper, we examine existing kernel-based linear transformers and identify two key issues that lead to such performance gaps: 1) unbounded gradients in the attention computation adversely impact the convergence of linear transformer models; 2) attention dilution which trivially distributes attention scores over long sequences while neglecting neighbouring structures.To address these issues, we first identify that the scaling of attention matrices is the devil in unbounded gradients, which turns out unnecessary in linear attention as we show theoretically and empirically.To this end, we propose a new linear attention that replaces the scaling operation with a normalization to stabilize gradients.For the issue of attention dilution, we leverage a diagonal attention to confine attention to only neighbouring tokens in early layers.Benefiting from the stable gradients and improved attention, our new linear transformer model, TRANSNORMER, demonstrates superior performance on text classification and language modeling tasks, as well as on the challenging Long-Range Arena benchmark, surpassing vanilla transformer and existing linear variants by a clear margin while being significantly more space-time efficient.The code is available at TRANSNORMER. Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dongxu Li 0003, Lingpeng Kong, Nick Barnes, Yiran Zhong |
EMNLP | 3 |
| 2022 | cosFormer: Rethinking Softmax In Attention
Zhen Qin 0003, Weixuan Sun, Dongxu Li 0003, Yunshen Wei, Baohong Lv, Lingpeng Kong, Yiran Zhong |
ICLR | 2 |
| 2022 | Inferring the Class Conditional Response Map for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation (WSSS) relies on class activation maps (CAMs) for pseudo labels generation. As CAMs only highlight the most discriminative regions of objects, the generated pseudo labels are usually unsatisfactory to serve directly as supervision. To solve this, most existing approaches follow a multi-training pipeline to refine CAMs for better pseudo-labels, which includes: 1) re-training the classification model to generate CAMs; 2) post-processing CAMs to obtain pseudo labels; and 3) training a semantic segmentation model with the obtained pseudo labels. However, this multi-training pipeline requires complicated adjustment and additional time. To address this, we propose a class-conditional inference strategy and an activation aware mask refinement loss function to generate better pseudo labels without retraining the classifier. The class conditional inference-time approach is presented to separately and iteratively reveal the classification network’s hidden object activation to generate more complete response maps. Further, our activation aware mask refinement loss function introduces a novel way to exploit saliency maps during segmentation training and refine the foreground object masks without suppressing background objects. Our method achieves superior WSSS results without requiring re-training of the classifier. https://github.com/weixuansun/InferCam Weixuan Sun, Jing Zhang 0052, Nick Barnes |
WACV | 1 |
| 2022 | Self-Supervised Pansharpening Based on a Cycle-Consistent Generative Adversarial NetworkabstractIn the field of remote sensing image pansharpening, deep learning-based methods have shown impressive performances recently. However, most deep learning-based pansharpening methods are based on supervised learning, which requires a large number of training images. In addition, obtaining large amounts of images with a high spatial and spectral resolution for training may be difficult in practice. In this letter, a novel self-supervised learning method based on a cycle-consistent generative adversarial network (CycleGAN) is proposed for remote sensing image pansharpening, without requiring large volumes of data for training. The framework contains two generators and two discriminators, and applies a residual neural network to the first generator. The panchromatic (PAN) image and multispectral (MS) image are input into the first generator to obtain the fused image, and then the fused image is input into the second generator to obtain a PAN image, which should be consistent with the input PAN image. The experimental results show that the proposed method performs better than the state-of-the-art unsupervised pansharpening method, and also achieves a competitive performance when compared with a supervised method. Jie Li 0022, Weixuan Sun, Menghui Jiang, Qiangqiang Yuan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | 3D Guided Weakly Supervised Semantic Segmentation
Weixuan Sun, Jing Zhang 0052, Nick Barnes |
ACCV (1) | 1 |