EDBT 2026 Demo / reviewers in the wild / expert
Xinyu Gong
dblp:215/5405
· DBLP profile ↗
19ranked-venue papers
6as first author
13since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Grounded-Instruct-Pix2Pix: Improving Instruction Based Image Editing with Automatic Target GroundingabstractText-guided Image Editing has recently attracted significant attention due to advances in the denoising diffusion models field. Current methods make it possible to execute complex image editing operations with simple text prompts. But despite impressive results, they often fail to restrict the edit area to only the object of interest, specified in the text prompt. To this end, we propose a novel framework we name Grounded-Instruct-Pix2Pix, which is capable of localized instruction-guided image editing in various scenarios including multi-object cases and complex backgrounds. Our experiments on a diverse set of images clearly showcase its advantage over the recent state-of-the-art approaches, especially at restricting the editing effect to the area of interest only. Grounded-Instruct-Pix2Pix implementation will be available at https://github.com/arthur-71/Grounded-Instruct-Pix2Pix. Artur Shagidanov, Hayk Poghosyan, Xinyu Gong, Zhangyang Wang, Shant Navasardyan, Humphrey Shi |
ICASSP | 3 |
| 2024 | Stroke-Seg: A deep learning-based framework for chinese stroke segmentationabstractAbstract Chinese stroke segmentation is a crucial and challenging task for various downstream applications such as font generation, aesthetic evaluation etc. Conventional semantic segmentation techniques typically face difficulties in accurately segmenting strokes, as intersection regions in Chinese characters can belong to multiple strokes simultaneously, and these approaches often lack a holistic understanding of character composition. This paper proposes a character stroke segmentation framework named Stroke‐Seg that integrates with various semantic segmentation architectures, demonstrating adaptability to different backbone networks to tackle the above tasks. A multi‐label output strategy is proposed to effectively classify strokes in intersection areas, overcoming the limitations of traditional semantic segmentation approaches. Additionally, a prior knowledge vector is incorporated into the input layer to provide character‐specific information on stroke composition, enhancing the ability of the framework to precisely identify and segment strokes. The effectiveness of the proposed framework is demonstrated through evaluations of a comprehensive dataset (brush calligraphy stroke segmentation dataset). Evaluations show that the proposed framework significantly improves the capability of semantic segmentation networks, achieving a remarkable improvement of up to 19.9% in the true stroke rate, compared to traditional stroke segmentation techniques. The Stroke‐Seg framework integrated with TransUNet demonstrates its high performance with an impressive 98.2% true stroke rate. Furthermore, the proposed framework combined with FCN still achieves good performance while consuming the least amount of computational resources and memory, demonstrating its potential for lightweight design. Xinyu Gong, Zeyang Bai, Haitao Nie |
IET Image Process. | 1 |
| 2024 | Understanding and Accelerating Neural Architecture Search With Training-Free and Theory-Grounded MetricsabstractThis work targets designing a principled and unified training-free framework for Neural Architecture Search (NAS), with high performance, low cost, and in-depth interpretation. NAS has been explosively studied to automate the discovery of top-performer neural networks, but suffers from heavy resource consumption and often incurs search bias due to truncated training or approximations. Recent NAS works Mellor et al. 2021, Chen et al. 2021, Abdelfattah et al. 2021 start to explore indicators that can predict a network's performance without training. However, they either leveraged limited properties of deep networks, or the benefits of their training-free indicators were not applied to more extensive search methods. By rigorous correlation analysis, we present a unified framework to understand and accelerate NAS, by disentangling “TEG” characteristics of searched networks –Trainability,Expressivity,Generalization– all assessed in a training-free manner. The TEG indicators could be scaled up and integrated with various NAS search methods, including both supernet and single-path NAS approaches. Extensive studies validate the effective and efficient guidance from our TEG-NAS framework, leading to both improved search accuracy and over 56% reduction in search time cost. Moreover, we visualize search trajectories on three landscapes of “TEG” characteristics, observing that a good local minimum is easier to find on NAS-Bench-201 given its simple topology, whereas balancing “TEG” characteristics is much harder on the DARTS space due to its complex landscape geometry. Wuyang Chen 0001, Xinyu Gong, Yunchao Wei, Humphrey Shi, Zhicheng Yan 0001, Yi Yang 0001, Zhangyang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | MMG-Ego4D: Multi-Modal Generalization in Egocentric Action RecognitionabstractIn this paper, we study a novel problem in egocentric action recognition, which we term as “Multimodal Generalization“ (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely missing. We thoroughly investigate MMG in the context of standard supervised action recognition and the more challenging few-shot setting for learning new action categories. MMG consists of two novel scenarios, designed to support security, and efficiency considerations in real-world applications: (1) missing modality generalization where some modalities that were present during the train time are missing during the inference time, and (2) cross-modal zero-shot generalization, where the modalities present during the inference time and the training time are disjoint. To enable this investigation, we construct a new dataset MMG-Ego4D containing data points with video, audio, and inertial motion sensor (IMU) modalities. Our dataset is derived from Ego4D [27] dataset, but processed and thoroughly re-annotated by human experts to facilitate research in the MMG problem. We evaluate a diverse array of models on MMG-Ego4D and propose new methods with improved generalization ability. In particular, we introduce a new fusion module with modality dropout training, contrastive-based alignment training, and a novel cross-modal prototypical loss for better few-shot performance. We hope this study will serve as a benchmark and guide future research in multimodal generalization problems. The benchmark and code are available at https://github.com/facebookresearch/MMG_Ego4D Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang |
CVPR | 1 |
| 2023 | NeRF-SOS: Any-View Self-supervised Object Segmentation on Complex Scenes
Zhiwen Fan, Peihao Wang, Yifan Jiang 0001, Xinyu Gong, Dejia Xu, Zhangyang Wang |
ICLR | 4 |
| 2022 | Unified Implicit Neural Stylization
Zhiwen Fan, Yifan Jiang 0001, Peihao Wang, Xinyu Gong, Dejia Xu, Zhangyang Wang |
ECCV (15) | 4 |
| 2022 | Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained AnalysisabstractAdvanced deep neural networks (DNNs), designed by either human or AutoML algorithms, are growing increasingly complex. Diverse operations are connected by complicated connectivity patterns, e.g., various types of skip connections. Those topological compositions are empirically effective and observed to smooth the loss landscape and facilitate the gradient flow in general. However, it remains elusive to derive any principled understanding of their effects on the DNN capacity or trainability, and to understand why or in which aspect one specific connectivity pattern is better than another. In this work, we theoretically characterize the impact of connectivity patterns on the convergence of DNNs under gradient descent training in fine granularity. By analyzing a wide network's Neural Network Gaussian Process (NNGP), we are able to depict how the spectrum of an NNGP kernel propagates through a particular connectivity pattern, and how that affects the bound of convergence rates. As one practical implication of our results, we show that by a simple filtration of "unpromising" connectivity patterns, we can trim down the number of models to evaluate, and significantly accelerate the large-scale neural architecture search without any overhead. Wuyang Chen 0001, Wei Huang 0034, Xinyu Gong, Boris Hanin, Zhangyang Wang |
NeurIPS | 3 |
| 2022 | Sandwich Batch Normalization: A Drop-In Replacement for Feature Distribution HeterogeneityabstractWe present Sandwich Batch Normalization (SaBN), a frustratingly easy improvement of Batch Normalization (BN) with only a few lines of code changes. SaBN is motivated by addressing the inherent feature distribution heterogeneity that one can be identified in many tasks, which can arise from data heterogeneity (multiple input domains) or model heterogeneity (dynamic architectures, model conditioning, etc.). Our SaBN factorizes the BN affine layer into one shared sandwich affine layer, cascaded by several parallel independent affine layers. Concrete analysis reveals that, during optimization, SaBN promotes balanced gradient norms while still preserving diverse gradient directions – a property that many application tasks seem to favor. We demonstrate the prevailing effectiveness of SaBN as a drop-in replacement in four tasks: conditional image generation, neural architecture search (NAS), adversarial training, and arbitrary style transfer. Leveraging SaBN immediately achieves better Inception Score and FID on CIFAR-10 and ImageNet conditional image generation with three state-of-the-art GANs; boosts the performance of a state-of-the-art weight-sharing NAS algorithm significantly on NAS-Bench-201; substantially improves the robust and standard accuracies for adversarial defense; and produces superior arbitrary styl-ized results. We also provide visualizations and analysis to help understand why SaBN works. Codes are available at: https://github.com/VITA-Group/Sandwich-Batch-Normalization. Xinyu Gong, Wuyang Chen 0001, Tianlong Chen 0001, Zhangyang Wang |
WACV | 1 |
| 2022 | Auto-X3D: Ultra-Efficient Video Understanding via Finer-Grained Neural Architecture SearchabstractEfficient video architecture is the key to deploying video recognition systems on devices with limited computing resources. Unfortunately, existing video architectures are often computationally intensive and not suitable for such applications. The recent X3D work presents a new family of efficient video models by expanding a hand-crafted image architecture along multiple axes, such as space, time, width, and depth. Although operating in a conceptually large space, X3D searches one axis at a time, and merely explored a small set of 30 architectures in total, which does not sufficiently explore the space. This paper bypasses existing 2D architectures, and directly searched for 3D architectures in a fine-grained space, where block type, filter number, expansion ratio and attention block are jointly searched. A probabilistic neural architecture search method is adopted to efficiently search in such a large space. Evaluations on Kinetics and Something-Something-V2 benchmarks confirm our AutoX3D models outperform existing ones in accuracy up to 1.3% under similar FLOPs, and reduce the computational cost up to ×1.74 when reaching similar performance. Yifan Jiang 0001, Xinyu Gong, Humphrey Shi, Zhicheng Yan 0001, Zhangyang Wang |
WACV | 2 |
| 2021 | Searching for Two-Stream Models in Multivariate Space for Video RecognitionabstractConventional video models rely on a single stream to capture the complex spatial-temporal features. Recent work on two-stream video models, such as SlowFast network and AssembleNet, prescribe separate streams to learn complementary features, and achieve stronger performance. However, manually designing both streams as well as the in-between fusion blocks is a daunting task, requiring to explore a tremendously large design space. Such manual exploration is time-consuming and often ends up with suboptimal architectures when computational resources are limited and the exploration is insufficient. In this work, we present a pragmatic neural architecture search approach, which is able to search for two-stream video models in giant spaces efficiently. We design a multivariate search space, including 6 search variables to capture a wide variety of choices in designing two-stream models. Furthermore, we propose a progressive search procedure, by searching for the architecture of individual streams, fusion blocks and attention blocks one after the other. We demonstrate two-stream models with significantly better performance can be automatically discovered in our design space. Our searched two-stream models, namely Auto-TSNet, consistently outperform other models on standard benchmarks. On Kinetics, compared with the SlowFast model, our Auto-TSNet-L model reduces FLOPS by nearly 11× while achieving the same accuracy 78.9%. On Something-Something-V2, Auto- TSNet-M improves the accuracy by at least 2% over other methods which use less than 50 GFLOPS per video. Xinyu Gong, Zheng Shou 0001, Matt Feiszli, Zhangyang Wang, Zhicheng Yan 0001 |
ICCV | 1 |
| 2021 | Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective
Wuyang Chen 0001, Xinyu Gong, Zhangyang Wang |
ICLR | 2 |
| 2021 | SAFIN: Arbitrary Style Transfer with Self-Attentive Factorized Instance NormalizationabstractArtistic style transfer aims to transfer the style characteristics of one image onto another image while retaining its content. Existing approaches commonly leverage various normalization techniques, although these face limitations in adequately transferring diverse textures to different spatial locations. Self-Attention based approaches have tackled this issue with partial success but suffer from unwanted artifacts. Motivated by these observations, this paper aims to combine the best of both worlds: self-attention and normalization. That yields a new plug-and-play module that we name Self-Attentive Factorized Instance Normalization (SAFIN). SAFIN is essentially a spatially-adaptive normalization module whose parameters are inferred through attention on the content and style image. We demonstrate that plugging SAFIN into the base network of another state-of-the-art method results in enhanced stylization. We also develop a novel base network composed of Wavelet Transform for multi-scale style transfer, which when combined with SAFIN, produces visually appealing results with lesser unwanted textures. Aaditya Singh, Shreeshail Hingane, Xinyu Gong, Zhangyang Wang |
ICME | 3 |
| 2021 | EnlightenGAN: Deep Light Enhancement Without Paired SupervisionabstractDeep learning-based methods have achieved remarkable success in image restoration and enhancement, but are they still competitive when there is a lack of paired training data? As one such example, this paper explores the low-light image enhancement problem, where in practice it is extremely challenging to simultaneously take a low-light and a normal-light photo of the same visual scene. We propose a highly effective unsupervised generative adversarial network, dubbed EnlightenGAN, that can be trained without low/normal-light image pairs, yet proves to generalize very well on various real-world test images. Instead of supervising the learning using ground truth data, we propose to regularize the unpaired training using the information extracted from the input itself, and benchmark a series of innovations for the low-light image enhancement problem, including a global-local discriminator structure, a self-regularized perceptual loss fusion, and the attention mechanism. Through extensive experiments, our proposed approach outperforms recent methods under a variety of metrics in terms of visual quality and subjective user study. Thanks to the great flexibility brought by unpaired training, EnlightenGAN is demonstrated to be easily adaptable to enhancing real-world images from various domains. Our codes and pre-trained models are available at: https://github.com/VITA-Group/EnlightenGAN. Yifan Jiang 0001, Xinyu Gong, Ding Liu 0001, Yu Cheng 0001, Xiaohui Shen, Jianchao Yang, Pan Zhou 0001, Zhangyang Wang |
IEEE Trans. Image Process. | 2 |
| 2020 | FasterSeg: Searching for Faster Real-time Semantic Segmentation
Wuyang Chen 0001, Xinyu Gong, Xianming Liu 0005, Zhangyang Wang |
ICLR | 2 |
| 2020 | NADS: Neural Architecture Distribution Search for Uncertainty AwarenessabstractMachine learning (ML) systems often encounter Out-of-Distribution (OoD) errors when dealing with testing data coming from a distribution different from training data. It becomes important for ML systems in critical applications to accurately quantify its predictive uncertainty and screen out these anomalous inputs. However, existing OoD detection approaches are prone to errors and even sometimes assign higher likelihoods to OoD samples. Unlike standard learning tasks, there is currently no well established guiding principle for designing OoD detection architectures that can accurately quantify uncertainty. To address these problems, we first seek to identify guiding principles for designing uncertainty-aware architectures, by proposing Neural Architecture Distribution Search (NADS). NADS searches for a distribution of architectures that perform well on a given task, allowing us to identify common building blocks among all uncertainty-aware architectures. With this formulation, we are able to optimize a stochastic OoD detection objective and construct an ensemble of models to perform OoD detection. We perform multiple OoD detection experiments and observe that our NADS performs favorably, with up to 57% improvement in accuracy compared to state-of-the-art methods among 15 different testing configurations. Randy Ardywibowo, Shahin Boluki, Xinyu Gong, Zhangyang Wang, Xiaoning Qian |
ICML | 3 |
| 2020 | AutoSpeech: Neural Architecture Search for Speaker RecognitionabstractSpeaker recognition systems based on Convolutional Neural Networks (CNNs) are often built with off-the-shelf backbones such as VGG-Net or ResNet.However, these backbones were originally proposed for image classification, and therefore may not be naturally fit for speaker recognition.Due to the prohibitive complexity of manually exploring the design space, we propose the first neural architecture search approach for the speaker recognition tasks, named as AutoSpeech.Our algorithm first identifies the optimal operation combination in a neural cell and then derives a CNN model by stacking the neural cell for multiple times.The final speaker recognition model can be obtained by training the derived CNN model through the standard scheme.To evaluate the proposed approach, we conduct experiments on both speaker identification and speaker verification tasks using the VoxCeleb1 dataset.Results demonstrate that the derived CNN architectures from the proposed approach significantly outperform current speaker recognition systems based on VGG-M, ResNet-18, and ResNet-34 backbones, while enjoying lower model complexity. Shaojin Ding, Tianlong Chen 0001, Xinyu Gong, Weiwei Zha, Zhangyang Wang |
INTERSPEECH | 3 |
| 2019 | Conditional Adversarial Generative Flow for Controllable Image SynthesisabstractFlow-based generative models show great potential in image synthesis due to its reversible pipeline and exact log-likelihood target, yet it suffers from weak ability for conditional image synthesis, especially for multi-label or unaware conditions. This is because the potential distribution of image conditions is hard to measure precisely from its latent variable $z$. In this paper, based on modeling a joint probabilistic density of an image and its conditions, we propose a novel flow-based generative model named conditional adversarial generative flow (CAGlow). Instead of disentangling attributes from latent space, we blaze a new trail for learning an encoder to estimate the mapping from condition space to latent space in an adversarial manner. Given a specific condition $c$, CAGlow can encode it to a sampled $z$, and then enable robust conditional image synthesis in complex situations like combining person identity with multiple attributes. The proposed CAGlow can be implemented in both supervised and unsupervised manners, thus can synthesize images with conditional information like categories, attributes, and even some unknown properties. Extensive experiments show that CAGlow ensures the independence of different conditions and outperforms regular Glow to a significant extent. Rui Liu 0019, Yu Liu 0015, Xinyu Gong, Xiaogang Wang 0001, Hongsheng Li 0001 |
CVPR | 3 |
| 2019 | AutoGAN: Neural Architecture Search for Generative Adversarial NetworksabstractNeural architecture search (NAS) has witnessed prevailing success in image classification and (very recently) segmentation tasks. In this paper, we present the first preliminary study on introducing the NAS algorithm to generative adversarial networks (GANs), dubbed AutoGAN. The marriage of NAS and GANs faces its unique challenges. We define the search space for the generator architectural variations and use an RNN controller to guide the search, with parameter sharing and dynamic-resetting to accelerate the process. Inception score is adopted as the reward, and a multi-level search strategy is introduced to perform NAS in a progressive way. Experiments validate the effectiveness of AutoGAN on the task of unconditional image generation. Specifically, our discovered architectures achieve highly competitive performance compared to current state-of-the-art hand-crafted GANs, e.g., setting new state-of-the-art FID scores of 12.42 on CIFAR-10, and 31.01 on STL-10, respectively. We also conclude with a discussion of the current limitations and future potential of AutoGAN. The code is available at https://github.com/TAMU-VITA/AutoGAN. Xinyu Gong, Shiyu Chang, Yifan Jiang 0001, Zhangyang Wang |
ICCV | 1 |
| 2018 | Neural Stereoscopic Image Style Transfer
Xinyu Gong, Hao-Zhi Huang 0001, Lin Ma 0002, Fumin Shen, Wei Liu 0005, Tong Zhang 0001 |
ECCV (5) | 1 |