EDBT 2026 Demo / reviewers in the wild / expert
Bin Chen 0011
dblp:22/5523-11
· DBLP profile ↗
134ranked-venue papers
10as first author
113since 2021 · last 2026
0000-0002-4798-230XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 1 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 49 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 7 since 2021Computer networks · 11 · 2 first-author · 9 since 2021Theory of computation · 8 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior RefinementabstractRecent advancements in diffusion-based generative priors have enabled visually plausible image compression at extremely low bit rates. However, existing approaches suffer from slow sampling processes and suboptimal bit allocation due to fragmented training paradigms. In this work, we propose Accelerate Diffusion-based Image Compression via Consistency Prior Refinement (DiffCR), a novel compression framework for efficient and high-fidelity image reconstruction. At the heart of DiffCR is a Frequency-aware Skip Estimation (FaSE) module that refines the epsilon-prediction prior from a pre-trained latent diffusion model and aligns it with compressed latents at different timesteps via Frequency Decoupling Attention (FDA). Furthermore, a lightweight consistency estimator enables fast two-step decoding by preserving the semantic trajectory of diffusion sampling. Without updating the backbone diffusion model, DiffCR achieves substantial bitrate savings (27.2% BD-rate(LPIPS) and 65.1% BD-rate(PSNR)) and over 10 times speed-up compared to SOTA diffusion-based compression baselines. Yichong Xia, Yimin Zhou 0011, Jinpeng Wang 0002, Bin Chen 0011 |
AAAI | 4 |
| 2026 | Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion ModelsabstractHao Fang, Xiaohang Sui, Hongyao Yu, Kuofeng Gao, Jiawei Kong, Sijin Yu, Bin Chen, Shu-Tao Xia. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Fang 0011, Xiaohang Sui, Hongyao Yu, Kuofeng Gao, Jiawei Kong 0001, Sijin Yu, Bin Chen 0011, Shutao Xia |
ACL (1) | 7 |
| 2026 | From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video AgentsabstractNiu Lian, Yuting Wang, Hanshu Yao, Jinpeng Wang, Bin Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Niu Lian, Hanshu Yao, Jinpeng Wang 0002, Bin Chen 0011, Yaowei Wang 0001, Min Zhang 0005, Shutao Xia |
ACL (1) | 5 |
| 2026 | When Efficiency Meets Safety: A Benchmark Security Analysis of KV Cache Compression in Large Language ModelsabstractXiaoxiao Ma, Kuofeng Gao, Zeyi Lu, Wenxi Jiang, Hao Fang, Hao Wu, Bin Chen, Shu-Tao Xia. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kuofeng Gao, Zeyi Lu, Wenxi Jiang, Hao Fang 0011, Bin Chen 0011, Shutao Xia |
ACL (1) | 7 |
| 2026 | Generalization Bounds for Transformer Channel DecodersabstractTransformer channel decoders, such as the Error Correction Code Transformer (ECCT), have shown strong empirical performance in channel decoding, yet their generalization behavior remains theoretically unclear. This paper studies the generalization performance of ECCT from a learning-theoretic perspective. By establishing a connection between multiplicative noise estimation errors and bit-error-rate (BER), we derive an upper bound on the generalization gap via bit-wise Rademacher complexity. The resulting bound characterizes the dependence on code length, model parameters, and training set size, and applies to both single-layer and multi-layer ECCTs. We further show that parity-check-based masked attention induces sparsity that reduces the covering number, leading to a tighter generalization bound. To the best of our knowledge, this work provides the first theoretical generalization guarantees for this class of decoders. Qinshan Zhang, Bin Chen 0011, Yong Jiang 0001, Shutao Xia |
ISIT | 2 |
| 2026 | Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature FilteringabstractModel Inversion Attacks (MIAs) pose a significant threat to data privacy by reconstructing sensitive training samples from the knowledge embedded in trained machine learning models. Despite recent progress in enhancing the effectiveness of MIAs across diverse settings, defense strategies have lagged behind—struggling to balance model utility with robustness against increasingly sophisticated attacks. In this work, we propose the ideal inversion error to measure the privacy leakage, and our theoretical and empirical investigations reveals that higher-rank features are inherently more prone to privacy leakage. Motivated by this insight, we propose a lightweight and effective defense strategy based on low-rank feature filtering, which explicitly reduces the attack surface by constraining the dimension of intermediate representations. Extensive experiments across various model architectures and datasets demonstrate that our method consistently outperforms existing defenses, achieving state-of-the-art performance against a wide range of MIAs. Notably, our approach remains effective even in challenging regimes involving high-resolution data and high-capacity models, where prior defenses fail to provide adequate protection. The code is available at https://github.com/Chrisqcwx/LoFt. Hongyao Yu, Yixiang Qiu, Hao Fang 0011, Tianqu Zhuang, Bin Chen 0011, Sijin Yu, Bin Wang 0034, Shutao Xia, Ke Xu 0002 |
KDD (1) | 5 |
| 2026 | MoE-LC: General-Purpose Lossless Compression for Multi-modal Data via Entropy-Aware Multi-ExpertsabstractThe web-scale surge of multimodal content, including short-video feeds and autonomous sensing streams, has made web-native lossless compression a prerequisite for delivery and storage across browsers and edge–cloud pipelines. However, existing methods often fail to adapt to shifting distributions across different batches and struggle to balance computational resources in the face of large conditional entropy disparities among diverse modalities. To address these limitations, we propose MoE-LC, a new mixture-of-experts framework for multi-modal lossless compression that dynamically accommodates heterogeneous data distributions and varying complexity levels. First, the Batch-Adaptive Experts (BAE) module introduces batch-specific parameters with a residual gating mechanism, ensuring stable modeling under non-stationary distributions. Second, the Entropy-Aware Multi-Expert Selection (MES) strategy adaptively allocates the number of experts according to the data's estimated compression difficulty (entropy), thereby improving resource utilization and computational efficiency. Finally, the Precision-Aware Expert Routing (PER) component applies high-precision computation solely to the most critical experts, significantly reducing overhead without sacrificing compression accuracy. Experimental results across multiple real-world datasets demonstrate that MoE-LC achieves 5.33%--70.89% improvements in compression ratio and 37.25%--1532.41% gains in throughput compared to advanced baselines, offering a scalable solution for real-time, large-scale multi-modal data compression. Our code is available at https://github.com/Magie0/MoE_LC. Zeyi Lu, Yujun Huang, Minxiao Chen, Bin Chen 0011, Shutao Xia |
WWW | 5 |
| 2026 | A temporal-aware generative network for cross-modal video universal adversarial perturbation generation
Kai-Wen Zhang, Shuo-Yang Sun, Hao Fang 0011, Changle Zhou, Bin Chen 0011, Shutao Xia |
Knowl. Based Syst. | 5 |
| 2026 | Learning gated experts for segment anything in the wild
Yizhen Guo, Hang Guo 0002, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia |
Pattern Recognit. | 5 |
| 2026 | Perceptual image compression with textual side information
Shiyu Qin, Bin Chen 0011, Yujun Huang, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
Pattern Recognit. | 2 |
| 2026 | Universal image restoration via task-adaptive diffusion degradation oriented model
Junxi Wu, Sicheng Pan, Naiqi Li, Bin Chen 0011, Baoyi An 0002, Zhi Wang 0001, Yaowei Wang 0001, Shutao Xia |
Pattern Recognit. | 4 |
| 2026 | Leveraging Neural Architecture Search for improved downstream-agnostic adversarial attack
Haodong Xiao, Bin Chen 0011, Hao Fang 0011, Yulin Wu 0001, Xuan Wang 0002, Zhi Wang 0001, Shutao Xia |
Pattern Recognit. | 4 |
| 2026 | Revolving NBP-Like Decoders for Cyclic and Quasi-Cyclic LDPC CodesabstractRecently, deep learning has demonstrated improvements over classical decoding algorithms in various families of error-correcting codes with short to moderate block lengths. Neural decoders like neural belief propagation (NBP), constructed based on the Tanner graphs of linear codes, have been extensively studied, but their application to longer codes is limited due to increased network complexity with longer block lengths. This complexity leads to higher computational costs and deployment challenges, such as GPU memory usage. To address this, we propose a novel revolving framework for NBP-like decoders tailored to cyclic and QC-LDPC codes, a commonly used class of error correction codes. Our approach leverages the cyclic structure and section-wise cyclic structure inherent in cyclic and QC-LDPC codes respectively, significantly simplifying the complexity of the network. Experimental results demonstrate the effectiveness of our method across various cyclic and QC-LDPC codes. Especially, compared to the traditional decoding scheme based on the section-wise cyclic structure, our proposed decoder shows substantial improvements on 5G LDPC codes, exhibits better performance than the traditional min-sum decoder, and approaches the sum-product decoder without a noticeable error floor within the investigated noise levels. Qinshan Zhang, Bin Chen 0011, Tianqu Zhuang, Yong Jiang 0001, Shutao Xia |
IEEE Trans. Commun. | 2 |
| 2026 | Dual Feature Fusion for Incomplete Multi-View Multi-Label LearningabstractMulti-view Multi-label Learning (MVML) aims to leverage multi-view information from input samples to achieve accurate predictions of multiple labels. Unfortunately, most existing MVML methods operate under the assumption of data completeness, which makes them ineffective in practical scenarios involving missing views or uncertain labels. Recent methods address incomplete data, but few approaches handle scenarios where both views and labels are missing. To address this challenge, we propose a Dual-view Feature-guided Fusion Learning (DFFL) framework. DFFL considers both view-specific unique features and inter-view consistent features. Specifically, DFFL constructs view-uniqueness contrastive learning to ensure that features within the same view maintain high semantic relevance under the condition of view missing, while the semantics between different views are different. Unlike previous methods, DFFL assumes that label relevance can be reversely mapped to high-dimensional features. By establishing View-consistency learning, the mutual information in the shared embedding space is maximized to achieve consistent feature alignment. In particular, DFFL minimizes the conditional entropy of the marginal distribution of multi-view features through dual prediction, thereby deriving the maximum joint distribution of feature fusion and combining the missing view index matrix to achieve feature fusion. This process can effectively alleviate the fusion feature suppression existing in previous methods. Finally, the missing label index matrix is combined with the fusion feature to complete the classification task. We validate the framework on five widely used datasets, and experimental results demonstrate that our approach achieves superior performance compared to state-of-the-art methods. Ablation studies further validated the effectiveness of each component in DFFL. Xinyu Xiao, Shuhan Qi, Yulin Wu 0001, Bin Chen 0011, Xuan Wang 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | Efficient Self-Supervised Video Hashing with Selective State SpacesabstractSelf-supervised video hashing (SSVH) is a practical task in video indexing and retrieval. Although Transformers are predominant in SSVH for their impressive temporal modeling capabilities, they often suffer from computational and memory inefficiencies. Drawing inspiration from Mamba, an advanced state-space model, we explore its potential in SSVH to achieve a better balance between efficacy and efficiency. We introduce S5VH, a Mamba-based video hashing model with an improved self-supervised learning paradigm. Specifically, we design bidirectional Mamba layers for both the encoder and decoder, which are effective and efficient in capturing temporal relationships thanks to the data-dependent selective scanning mechanism with linear complexity. In our learning strategy, we transform global semantics in the feature space into semantically consistent and discriminative hash centers, followed by a center alignment loss as a global learning signal. Our self-local-global (SLG) paradigm significantly improves learning efficiency, leading to faster and better convergence. Extensive experiments demonstrate S5VH's improvements over state-of-the-art methods, superior transferability, and scalable advantages in inference efficiency. Jinpeng Wang 0002, Niu Lian, Jun Li 0131, Bin Chen 0011, Yongbing Zhang 0002, Shutao Xia |
AAAI | 6 |
| 2025 | Modeling Uncertainty in Composed Image Retrieval via Probabilistic EmbeddingsabstractHaomiao Tang, Jinpeng Wang, Yuang Peng, GuangHao Meng, Ruisheng Luo, Bin Chen, Long Chen, Yaowei Wang, Shu-Tao Xia. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haomiao Tang, Jinpeng Wang 0002, Yuang Peng, Guanghao Meng, Ruisheng Luo, Bin Chen 0011, Long Chen 0016, Yaowei Wang 0001, Shutao Xia |
ACL (1) | 6 |
| 2025 | Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context LearningabstractVisual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which in practice may not hold true. Multiple suitable prompts may exist, but individually they often fall short, leading to difficulties in selection and the exclusion of useful context. To address this, we propose a new perspective: prompt condensation. Rather than relying on a single prompt, candidate prompts collaborate to efficiently integrate informative contexts without sacrificing resolution. We devise Condenser, a lightweight external plugin that compresses relevant fine-grained context across multiple prompts. Optimized end-to-end with the backbone, Condenser ensures accurate integration of contextual cues. Experiments demonstrate Condenser outperforms state-of-the-arts across benchmark tasks, showing superior context compression, scalability with more prompts, and enhanced computational efficiency compared to ensemble methods, positioning it as a highly competitive solution for VICL. Code is open-sourced at https://github.com/gimpong/CVPR25-Condenser. Jinpeng Wang 0002, Tianci Luo, Yaohua Zha, Ruisheng Luo, Bin Chen 0011, Tao Dai 0001, Long Chen 0016, Yaowei Wang 0001, Shutao Xia |
CVPR | 6 |
| 2025 | AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingabstractSelf-Supervised Video Hashing (SSVH) compresses videos into hash codes for efficient indexing and retrieval using unlabeled training videos. Existing approaches rely on random frame sampling to learn video features and treat all frames equally. This results in suboptimal hash codes, as it ignores frame-specific information density and reconstruction difficulty. To address this limitation, we propose a new framework, termed AutoSSVH, that employs adversarial frame sampling with hash-based contrastive learning. Our adversarial sampling strategy automatically identifies and selects challenging frames with richer information for reconstruction, enhancing encoding capability. Additionally, we introduce a hash component voting strategy and a point-to-set (P2Set) hash-based contrastive objective, which help capture complex inter-video semantic relationships in the Hamming space and improve the discriminability of learned hash codes. Extensive experiments demonstrate that Au-toSSVH achieves superior retrieval efficacy and efficiency compared to state-of-the-art approaches. Code is available at https://github.com/EliSpectre/CVPR25-AutoSSVH. Niu Lian, Jun Li 0131, Jinpeng Wang 0002, Ruisheng Luo, Yaowei Wang 0001, Shutao Xia, Bin Chen 0011 |
CVPR | 7 |
| 2025 | PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba AdapterabstractApplying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementary information in the intermediate layer, thereby failing to fully unlock the potential of pre-trained models. To overcome this limitation, we propose an orthogonal solution: Point Mamba Adapter (PMA), which constructs an ordered feature sequence from all layers of the pre-trained model and leverages Mamba to fuse all complementary semantics, thereby promoting comprehensive point cloud understanding. Constructing this ordered sequence is non-trivial due to the inherent isotropy of 3D space. Therefore, we further propose a geometry-constrained gate prompt generator (G2PG) shared across different layers, which applies shared geometric constraints to the output gates of the Mamba and dynamically optimizes the spatial order, thus enabling more effective integration of multi-layer information. Extensive experiments conducted on challenging point cloud datasets across various tasks demonstrate that our PMA elevates the capability for point cloud understanding to a new level by fusing diverse complementary intermediate features. Code is available at https://github.com/zyh16143998882/PMA. Yaohua Zha, Yanzi Wang, Hang Guo 0002, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhihao Ouyang, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 6 |
| 2025 | Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset DistillationabstractDataset distillation is an emerging dataset reduction method, which condenses large-scale datasets while maintaining task accuracy. Current parameterization methods achieve enhanced performance under extremely high compression ratio by optimizing determined synthetic dataset in informative feature domain. However, they limit themselves to a fixed optimization space for distillation, neglecting the diverse guidance across different informative latent spaces. To overcome this limitation, we propose a novel parameterization method dubbed Hierarchical Parameterization Distillation (H-PD), to systematically explore hierarchical feature within provided feature space (e.g., layers within pre-trained generative adversarial networks). We verify the correctness of our insights by applying the hierarchical optimization strategy on GAN-based parameterization method. In addition, we introduce a novel class-relevant feature distance metric to alleviate the computational burden associated with synthetic dataset evaluation, bridging the gap between synthetic and original datasets. Experimental results demonstrate that the proposed H-PD achieves a significant performance improvement under various settings with equivalent time consumption, and even surpasses current generative distillation using diffusion models under extreme compression ratios IPC=1 and IPC=10. Our code is available at https://github.com/ndhg1213/H-PD Xinhao Zhong, Hao Fang 0011, Bin Chen 0011, Xulin Gu, Meikang Qiu, Shuhan Qi, Shutao Xia |
CVPR | 3 |
| 2025 | DCDiff: Enhancing JPEG Compression via Diffusion-based DC Coefficients EstimationabstractJPEG is the most widely-used image compression method on low-cost cameras which cannot support learning-based compressors. One promising approach to enhance JPEG aims to drop DC coefficients at the cameras’ ends (without extra computation) and reconstruct those DC coefficients after receiving them. They all face the challenge that their DC reconstruction relies on a statistical property, which will cause deviationintroduced errors and propagate. In this paper, we propose DCDiff, a novel end-to-end DC estimation method to tackle the above challenge. Instead of using statistical methods to recover DC coefficients and then fix errors, we directly leverage a generative model to estimate DC coefficients in an end-to-end manner. In the meantime, we generate masks to correct certain image locations that do not satisfy the statistical distribution to suppress error propagation. Extensive experiments show that DCDiff not only outperforms all baselines on compression performance but also introduces a tiny impact on downstream tasks and is fully compatible with 2 typical low-cost processors with JPEG support. Han Qiu 0001, Tianwei Zhang 0004, Bin Chen 0011, Chao Zhang 0008 |
DAC | 4 |
| 2025 | Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text DetectorsabstractHao Fang, Jiawei Kong, Tianqu Zhuang, Yixiang Qiu, Kuofeng Gao, Bin Chen, Shu-Tao Xia, Yaowei Wang, Min Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hao Fang 0011, Jiawei Kong 0001, Tianqu Zhuang, Yixiang Qiu, Kuofeng Gao, Bin Chen 0011, Shutao Xia, Yaowei Wang 0001, Min Zhang 0005 |
EMNLP | 6 |
| 2025 | MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional ThresholdsabstractThe rapid advancement of large language models has intensified public concerns about the potential misuse.Therefore, it is important to build trustworthy AI-generated text detection systems.Existing methods neglect stylistic modeling and mostly rely on static thresholds, which greatly limits the detection performance.In this paper, we propose the Mixture of Stylistic Experts (MoSEs) framework that enables stylistics-aware uncertainty quantification through conditional threshold estimation.MoSEs contain three core components, namely, the Stylistics Reference Repository (SRR), the Stylistics-Aware Router (SAR), and the Conditional Threshold Estimator (CTE).For input text, SRR can activate the appropriate reference data in SRR and provide them to CTE.Subsequently, CTE jointly models the linguistic statistical properties and semantic features to dynamically determine the optimal threshold.With a discrimination score, MoSEs yields prediction labels with the corresponding confidence level.Our framework achieves an average improvement 11.34% in detection performance compared to baselines.More inspiringly, MoSEs shows a more evident improvement 39.15% in the low-resource case.Our code is available at https://github.com/ creator-xi/MoSEs. Junxi Wu, Jinpeng Wang 0002, Bin Chen 0011, Dongjian Hu, Shutao Xia |
EMNLP | 4 |
| 2025 | LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video CodecabstractExisting Implicit Neural Representation (INR) video compression techniques have opened up new avenues in the field of video compression. NeRV maps the temporal coordinates to high-resolution images using neural networks, providing a more flexible and efficient encoding method for video data. However, NeRV implicitly stores all video information in the network, requiring post network compression techniques such as pruning. To integrate explicit compression and implicit representation into an end-to-end framework, this study proposes a novel neural representation-based video compression paradigm called Latent code based Neural representation video compression (LNeRV). Specifically, LNeRV consists of hierarchy feature grids, synthesis network and entropy coding network. With a single-stage training process, LNeRV achieves video compression and dynamically allocates bits according to video complexity, better fitting dynamic videos. We provide a comprehensive compression-to-decompression workflow for our approach. Extensive experimental results verify the effectiveness of our LNeRV. Jiahong Chen, Bin Chen 0011, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICASSP | 3 |
| 2025 | RobNAS: Robust Neural Architecture Search for Point Cloud Adversarial DefenseabstractAs point clouds gain widespread application in fields such as autonomous driving and scene modeling, an increasing number of point cloud learning networks have emerged. As a result, research on 3D adversarial attacks and defenses has rapidly advanced. To the best of our knowledge, existing 3D defense methods primarily focus on enhancing the robustness of networks through point cloud processing or adversarial training, without attention given to network architecture. In this paper, we propose RobNAS, enhancing the adversarial robustness of point cloud classification networks by integrating adversarial training with Neural Architecture Search (NAS) from an architectural robustness perspective. Specifically, we first incorporate PGD-based adversarial training during the architecture search phase of RobNAS to obtain the most robust architecture. Subsequently, during the adversarial training phase, we introduce various adversarial examples to enhance the robustness of the model weights. Our experimental results demonstrate that our method achieves State-Of-The-Art (SOTA) performance. Furthermore, we aim to shed light on the promising potential of architectural robustness for learning robust point cloud representation. Shuoyang Sun, Hao Fang 0011, Bin Chen 0011, Jiawei Li 0006, Enze Huo, Shutao Xia |
ICASSP | 4 |
| 2025 | One Perturbation is Enough: On Generating Universal Adversarial Perturbations Against Vision-Language Pre-Training ModelsabstractVision-Language Pre-training (VLP) models have exhibited unprecedented capability in many applications by taking full advantage of the multimodal alignment. However, previous studies have shown they are vulnerable to maliciously crafted adversarial samples. Despite recent success, these methods are generally instance-specific and require generating perturbations for each input sample. In this paper, we reveal that VLP models are also vulnerable to the instance-agnostic universal adversarial perturbation (UAP). Specifically, we design a novel Contrastive-training Perturbation Generator with Cross-modal conditions (C-PGC) to achieve the attack. In light that the pivotal multimodal alignment is achieved through the advanced contrastive learning technique, we devise to turn this powerful weapon against themselves, i.e., employ a malicious version of contrastive learning to train the C-PGC based on our carefully crafted positive and negative image-text pairs for essentially destroying the alignment relationship learned by VLP models. Besides, C-PGC fully utilizes the characteristics of Vision-and-Language (V+L) scenarios by incorporating both unimodal and cross-modal information as effective guidance. Extensive experiments show that C-PGC successfully forces adversarial samples to move away from their original area in the VLP model's feature space, thus essentially enhancing attacks across various victim models and V+L tasks. The GitHub repository is available at https://github.com/ffhibnese/CPGC_VLP_Universal_Attacks. Hao Fang 0011, Jiawei Kong 0001, Bin Chen 0011, Jiawei Li 0006, Shutao Xia, Ke Xu 0002 |
ICCV | 4 |
| 2025 | Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
Jun Li 0131, Jinpeng Wang 0002, Chaolei Tan, Niu Lian, Long Chen 0016, Yaowei Wang 0001, Min Zhang 0005, Shutao Xia, Bin Chen 0011 |
ICCV | 9 |
| 2025 | Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression
Shiyu Qin, Jinpeng Wang 0002, Yimin Zhou 0011, Bin Chen 0011, Tianci Luo, Baoyi An 0002, Tao Dai 0001, Shutao Xia, Yaowei Wang 0001 |
ICCV | 4 |
| 2025 | An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compressionabstract3D Gaussian Splatting (3DGS) has witnessed its rapid development in novel view synthesis, which attains high quality reconstruction and real-time rendering. At the same time, there is still a gap before implicit neural representation (INR) can become a practical compressor due to the lack of stream decoding and real-time frame reconstruction on consumer-grade hardware. It remains a question whether the fast rendering and partial parameter decoding characteristics of 3DGS are applicable to video compression. To address these challenges, we propose a Toast-like Sliding Window (TSW) orthographic projection for converting any 3D Gaussian model into a video representation model. This method efficiently represents video by leveraging temporal redundancy through a sliding window approach. Additionally, the converted model is inherently stream-decodable and offers a higher rendering frame rate compared to INR methods. Building on TSW, we introduce an end-to-end trainable video compression method, GSVC, which employs deformable Gaussian representation and optical flow guidance to capture dynamic content in videos. Experimental results demonstrate that our method effectively transforms a 3D Gaussian model into a practical video compressor. GSVC further achieves better rate-distortion performance than NeRV on the UVG dataset, while achieving higher frame reconstruction speed (+30%~40% fps) and stream decoding. Code is available at [Github](https://github.com/actcwlf/GSVC) Bin Chen 0011, Zimo Liu, Yaowei Wang 0001, Shutao Xia |
ICLR | 2 |
| 2025 | DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementabstractReconstructing high-quality images under low bitrates conditions presents a challenge, and previous methods have made this task feasible by leveraging the priors of diffusion models. However, the effective exploration of pre-trained latent diffusion models and semantic information integration in image compression tasks still needs further study. To address this issue, we introduce Diffusion-based High Perceptual Fidelity Image Compression with Semantic Refinement (DiffPC), a two-stage image compression framework based on stable diffusion. DiffPC efficiently encodes low-level image information, enabling the highly realistic reconstruction of the original image by leveraging high-level semantic features and the prior knowledge inherent in diffusion models. Specifically, DiffPC utilizes a multi-feature compressor to represent crucial low-level information with minimal bitrates and employs pre-embedding to acquire more robust hybrid semantics, thereby providing additional context for the decoding end. Furthermore, we have devised a control module tailored for image compression tasks, ensuring structural and textural consistency in reconstruction even at low bitrates and preventing decoding collapses induced by condition leakage. Extensive experiments demonstrate that our method achieves state-of-the-art perceptual fidelity and surpasses previous perceptual image compression methods by a significant margin in statistical fidelity. Yichong Xia, Yimin Zhou 0011, Jinpeng Wang 0002, Baoyi An 0002, Haoqian Wang, Yaowei Wang 0001, Bin Chen 0011 |
ICLR | 7 |
| 2025 | Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual InformationabstractDataset distillation (DD) aims to minimize the time and memory consumption needed for training deep neural networks on large datasets, by creating a smaller synthetic dataset that has similar performance to that of the full real dataset. However, current dataset distillation methods often result in synthetic datasets that are excessively difficult for networks to learn from, due to the compression of a substantial amount of information from the original data through metrics measuring feature similarity, e,g., distribution matching (DM). In this work, we introduce conditional mutual information (CMI) to assess the class-aware complexity of a dataset and propose a novel method by minimizing CMI. Specifically, we minimize the distillation loss while constraining the class-aware complexity of the synthetic dataset by minimizing its empirical CMI from the feature space of pre-trained networks, simultaneously. Conducting on a thorough set of experiments, we show that our method can serve as a general regularization method to existing DD methods and improve the performance and training efficiency. Xinhao Zhong, Bin Chen 0011, Hao Fang 0011, Xulin Gu, Shutao Xia, En-Hui Yang |
ICLR | 2 |
| 2025 | Stealthy Shield Defense: A Conditional Mutual Information-Based Approach against Black-Box Model Inversion AttacksabstractModel inversion attacks (MIAs) aim to reconstruct the private training data by accessing the public model, raising concerns about privacy leakage. Black-box MIAs, where attackers can only query the model and obtain outputs, are closer to real-world scenarios. The latest black-box attacks have outperformed state-of-the-art white-box attacks, and existing defenses cannot resist them effectively. To fill this gap, we propose Stealthy Shield Defense (SSD), a post-processing algorithm against black-box MIAs. Our idea is to modify the model's outputs to minimize the conditional mutual information (CMI). We mathematically prove that CMI is a special case of Information Bottleneck (IB), and thus inherits the benefits of IB---making predictions less dependent on inputs and more dependent on ground truths. This theoretically guarantees our effectiveness, both in resisting MIAs and preserving utility. To minimize CMI, we formulate a convex optimization problem and solve it via the water-filling method. Without the need to retrain the model, our defense is plug-and-play and easy to deploy. Experimental results indicate that SSD outperforms existing defenses, in terms of MIA resistance and model's utility, across various attack algorithms, private datasets, and model architectures. Our code is available at https://github.com/ZhuangQu/Stealthy-Shield-Defense. Tianqu Zhuang, Hongyao Yu, Yixiang Qiu, Hao Fang 0011, Bin Chen 0011, Shutao Xia |
ICLR | 5 |
| 2025 | 3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric PriorsabstractExisting multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camera systems, commonly found in virtual reality and autonomous driving applications. To address this limitation, we propose 3D-LMVIC, a novel learning-based multi-view image compression framework that leverages 3D Gaussian Splatting to derive geometric priors for accurate disparity estimation. Furthermore, we introduce a depth map compression model to minimize geometric redundancy across views, along with a multi-view sequence ordering strategy based on a defined distance measure between views to enhance correlations between adjacent views. Experimental results demonstrate that 3D-LMVIC achieves superior performance compared to both traditional and learning-based methods. Additionally, it significantly improves disparity estimation accuracy over existing two-view approaches. Yujun Huang, Bin Chen 0011, Niu Lian, Xin Wang 0001, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICML | 2 |
| 2025 | Clients Collaborate: Flexible Differentially Private Federated Learning with Guaranteed Improvement of Utility-Privacy Trade-offabstractTo defend against privacy leakage of user data, differential privacy is widely used in federated learning, but it is not free. The addition of noise randomly disrupts the semantic integrity of the model and this disturbance accumulates with increased communication rounds. In this paper, we introduce a novel federated learning framework with rigorous privacy guarantees, named FedCEO, designed to strike a trade-off between model utility and user privacy by letting clients "C*ollaborate with Each Other". Specifically, we perform efficient tensor low-rank proximal optimization on stacked local model parameters at the server, demonstrating its capability to flexibly truncate high-frequency components in spectral space. This capability implies that our FedCEO can effectively recover the disrupted semantic information by smoothing the global semantic space for different privacy settings and continuous training processes. Moreover, we improve the SOTA utility-privacy trade-off bound by order of $\sqrt{d}$, where $d$ is the input dimension. We illustrate our theoretical results with experiments on representative datasets and observe significant performance improvements and strict privacy guarantees under different privacy settings. The *code is available at https://github.com/6lyc/FedCEO_Collaborate-with-Each-Other. Yuecheng Li, Lele Fu, Jian Lou 0001, Bin Chen 0011, Lei Yang 0030, Zibin Zheng, Chuan Chen 0001 |
ICML | 5 |
| 2025 | Expert-Enhanced Masked Point Modeling for Point Cloud Self-Supervised LearningabstractRecently, learning-based point cloud analysis has played a crucial role in robotic perception. Masked Point Modeling (MPM), owing to its powerful representational capabilities, has become the mainstream point cloud self-supervised learning method. However, existing MPM-based methods often suffer from the problem of negative transfer, due to the disparity in semantic distribution between upstream data and downstream data. To address this issue, we propose an expert enhancement strategy for existing MPM-based methods. Specifically, we insert a Sparse Mixture of Experts (SMoE) layer after each block of the backbone network, which utilizes a multi-branch expert architecture with routers that allocate data of different semantics to the appropriate experts for analysis. During the pre-training phase, our expert-enhanced model not only learns universal 3D representations for the backbone network but also acquires powerful semantic routing capabilities for all expert layers. In the fine-tuning phase, we freeze all backbones and conduct end-to-end fine-tuning solely on our expert layers to adaptively select multiple experts most relevant to the semantics of each downstream data for analysis. Extensive downstream experiments demonstrate the superiority of our method, especially outperforming baseline (Point-MAE) by 5.16%, 5.86%, and 4.62% in three variants of ScanObjectNN while utilizing only 12% of its trainable parameters. Our code is released at https://github.com/chenchen1104/point_e2mae. Yaohua Zha, Naiqi Li, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
ICRA | 5 |
| 2025 | Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised LearningabstractPoint clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on learning domain-specific 3D representations within a single domain, neglecting the complementary nature of cross-domain knowledge, which limits the learning of 3D representations. In this paper, we propose to learn a comprehensive Point cloud Mixture-of-Domain-Experts model (Point-MoDE) via a block-to-scene pre-training strategy. Specifically, We first propose a mixture-of-domain-expert model consisting of scene domain experts and multiple shared object domain experts. Furthermore, we propose a block-to-scene pretraining strategy, which leverages the features of point blocks in the object domain to regress their initial positions in the scene domain through object-level block mask reconstruction and scene-level block position regression. By integrating the complementary knowledge between object and scene, this strategy simultaneously facilitates the learning of both object-domain and scene-domain representations, leading to a more comprehensive 3D representation. Extensive experiments in downstream tasks demonstrate the superiority of our model. Yaohua Zha, Tao Dai 0001, Hang Guo 0002, Yanzi Wang, Bin Chen 0011, Ke Chen 0004, Shutao Xia |
IJCAI | 5 |
| 2025 | γ-FedHT: Stepsize-Aware Hard-Threshold Gradient Compression in Federated Learning
Rongwei Lu, Yifei Zhu 0001, Bin Chen 0011, Zhi Wang 0001 |
INFOCOM | 6 |
| 2025 | EDPC: Accelerating Lossless Compression via Lightweight Probability Models and Decoupled Parallel DataflowabstractThe explosive growth of multi-source multimedia data has significantly increased the demands for transmission and storage, placing substantial pressure on bandwidth and storage infrastructures. While Autoregressive Compression Models (ACMs) have markedly improved compression efficiency through probabilistic prediction, current approaches remain constrained by two critical limitations: suboptimal compression ratios due to insufficient fine-grained feature extraction during probability modeling, and real-time processing bottlenecks caused by high resource consumption and low compression speeds. To address these challenges, we propose Efficient Dual-path Parallel Compression (EDPC), a hierarchically optimized compression framework that synergistically enhances modeling capability and execution efficiency via coordinated dual-path operations. At the modeling level, we introduce the Information Flow Refinement (IFR) metric grounded in mutual information theory, and design a Multi-path Byte Refinement Block (MBRB) to strengthen cross-byte dependency modeling via heterogeneous feature propagation. At the system level, we develop a Latent Transformation Engine (LTE) for compact high-dimensional feature representation and a Decoupled Pipeline Compression Architecture (DPCA) to eliminate encoding-decoding latency through pipelined parallelization. Experimental results demonstrate that EDPC achieves comprehensive improvements over state-of-the-art methods, including a 2.7× faster compression speed, and a 3.2% higher compression ratio. These advancements establish EDPC as an efficient solution for real-time processing of large-scale multimedia data in bandwidth-constrained scenarios. Our code is available at https://github.com/Magie0/EDPC. Zeyi Lu, Yujun Huang, Minxiao Chen, Bin Chen 0011, Baoyi An 0002, Shutao Xia |
ACM Multimedia | 5 |
| 2025 | ICAS: Detecting Training Data from Autoregressive Image Generative Models
Hongyao Yu, Yixiang Qiu, Hao Fang 0011, Tianqu Zhuang, Jiaxin Hong, Bin Chen 0011, Shutao Xia |
ACM Multimedia | 7 |
| 2025 | Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMsabstractLarge Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding the visual information during decoding. To alleviate this issue, we introduce a novel Conditional Pointwise Mutual Information (C-PMI) calibrated decoding strategy, which adaptively strengthens the mutual dependency between generated texts and input images to mitigate hallucinations. Unlike existing methods solely focusing on text token sampling, we propose to jointly model the contributions of visual and textual tokens to C-PMI, formulating hallucination mitigation as a bi-level optimization problem aimed at maximizing mutual information. To solve it, we design a token purification mechanism that dynamically regulates the decoding process by sampling text tokens remaining maximally relevant to the given image, while simultaneously refining image tokens most pertinent to the generated response. Extensive experiments across various benchmarks reveal that the proposed method significantly reduces hallucinations in LVLMs while preserving decoding efficiency. Hao Fang 0011, Changle Zhou, Jiawei Kong 0001, Kuofeng Gao, Bin Chen 0011, Guojun Ma, Shutao Xia |
NeurIPS | 5 |
| 2025 | Information-Theoretic Point Cloud Defense: Harnessing Conditional Mutual Information Against Adversarial Attacks
Xinhao Zhong, Shuoyang Sun, Jiaxin Hong, Bin Chen 0011, Teko Ranoka, Xuan Wang 0002, Shutao Xia |
PRCV (10) | 5 |
| 2025 | Generative Adversarial CLIPs for Unsupervised Backlit Image Enhancement
Suoyang Sun, Yujun Huang, Bin Chen 0011 |
WASA (1) | 4 |
| 2025 | MB-RACS: Measurement-Bounds-Based Rate-Adaptive Image Compressed Sensing NetworkabstractConventional compressed sensing (CS) algorithms typically apply a uniform sampling rate to different image blocks. A more strategic approach could be to allocate the number of measurements adaptively, based on each image block's complexity. In this paper, we propose a Measurement-Bounds-based Rate-Adaptive Image Compressed Sensing Network (MB-RACS) framework, which aims to adaptively determine the sampling rate for each image block in accordance with traditional measurement bounds theory. Moreover, since in real-world scenarios statistical information about the original image cannot be directly obtained, we suggest a multi-stage rate-adaptive sampling strategy. This strategy sequentially adjusts the sampling ratio allocation based on the information gathered from previous samplings. We formulate the multi-stage rate-adaptive sampling as a convex optimization problem and address it using a combination of Newton's method and binary search techniques. Our experiments demonstrate that the proposed MB-RACS method surpasses current leading methods, with experimental evidence also underscoring the effectiveness of each module within our proposed framework. Yujun Huang, Bin Chen 0011, Naiqi Li, Baoyi An 0002, Shutao Xia, Yaowei Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | FCA-Net: Accelerating stereo image compression through cascade alignment of side information
Yichong Xia, Yujun Huang, Bin Chen 0011, Genping Wang, Haoqian Wang, Yaowei Wang 0001 |
Pattern Recognit. | 3 |
| 2025 | Transformer-Based Decoders for Cyclic Codes: A Tanner Cycle-Equivalent ApproachabstractTransformers have recently emerged as effective neural decoders, capable of capturing complex interdependencies and delivering strong performance. Cyclic codes, including BCH codes, Reed-Solomon codes, and Reed-Muller codes, are fundamental in practical channel coding applications. In this work, we propose a novel approach for Transformer-based neural decoders specifically designed for cyclic codes. By leveraging the relational properties among the parity-check matrix, Tanner graph, and mask matrix, we integrate these with the inherent algebraic structure of cyclic codes. We then identify a cyclic shift reuse property that can be effectively applied to both parameter matrices and self-attention matrices. Extensive simulations on BCH codes demonstrate that our method reduces the total number of parameters by up to 64.4% compared to prior Transformer-based decoders while achieving comparable decoding performance. Additionally, our method decreases the computational cost of self-attention by up to 10.6%. Finally, through experiments across various dimensions, we explore the effect of embedding dimensions on performance and provide insights into the relationship between embedding dimensions and code positions. Weijun Fang, Bin Chen 0011 |
IEEE Trans. Commun. | 3 |
| 2025 | GI-NAS: Boosting Gradient Inversion Attacks Through Adaptive Neural Architecture SearchabstractGradient Inversion Attacks invert the transmitted gradients in Federated Learning (FL) systems to reconstruct the sensitive data of local clients and have raised considerable privacy concerns. A majority of gradient inversion methods rely heavily on explicit prior knowledge (e.g., a well pre-trained generative model), which is often unavailable in realistic scenarios. This is because real-world client data distributions are often highly heterogeneous, domain-specific, and unavailable to attackers, making it impractical for attackers to obtain perfectly matched pre-trained models, which inevitably suffer from fundamental distribution shifts relative to target private data. To alleviate this issue, researchers have proposed to leverage the implicit prior knowledge of an over-parameterized network. However, they only utilize a fixed neural architecture for all the attack settings. This would hinder the adaptive use of implicit architectural priors and consequently limit the generalizability. In this paper, we further exploit such implicit prior knowledge by proposing Gradient Inversion via Neural Architecture Search (GI-NAS), which adaptively searches the network and captures the implicit priors behind neural architectures. Extensive experiments verify that our proposed GI-NAS can achieve superior attack performance compared to state-of-the-art gradient inversion methods, even under more practical settings with high-resolution images, large-sized batches, and advanced defense strategies. To the best of our knowledge, we are the first to successfully introduce NAS to the gradient inversion community. We believe that this work exposes critical vulnerabilities in real-world federated learning by demonstrating high-fidelity reconstruction of sensitive data without requiring domain-specific priors, forcing urgent reassessment of FL privacy safeguards. Hao Fang 0011, Bin Chen 0011, Xiaohang Sui, Chuan Chen 0001, Shutao Xia, Ke Xu 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Data-Aware Gradient Compression for FL in Communication-Constrained Mobile ComputingabstractFederated Learning (FL) in mobile environments faces significant communication bottlenecks. Gradient compression has proven as an effective solution to this issue, offering substantial benefits in environments with limited bandwidth and metered data. Yet, it encounters severe performance drops in non-IID environments due to a one-size-fits-all compression approach, which does not account for the varying data volumes across workers. Assigning varying compression ratios to workers with distinct data distributions and volumes is therefore a promising solution. This work derives the convergence rate of distributed SGD with non-uniform compression, which reveals the intricate relationship between model convergence and the compression ratios applied to individual workers. Accordingly, we frame the relative compression ratio assignment as an$n$-variable chi-squared nonlinear optimization problem, constrained by a limited communication budget. We propose DAGC-R, which assigns conservative compression to workers handling larger data volumes. Recognizing the computational limitations of mobile devices, we propose the DAGC-A, which is computationally less demanding and enhances the robustness of compression in non-IID scenarios. Our experiments confirm that the DAGC-R and DAGC-A can speed up the training speed by up to 25.43% and 16.65% compared to the uniform compression respectively, when dealing with highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Yinan Mao, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | An Efficient Implicit Neural Representation Image Codec Based on Mixed Autoregressive Model for Low-Complexity DecodingabstractDisplaying high-quality images on edge devices, such as augmented reality devices, is essential for enhancing the user experience. However, these devices often face power consumption and computing resource limitations, making it challenging to apply many deep learning-based image compression algorithms in this field. Implicit Neural Representation (INR) for image compression is an emerging technology that offers two key benefits compared to cutting-edge autoencoder models: low computational complexity and parameter-free decoding. It also outperforms many traditional and early neural compression methods in terms of quality. In this study, we introduce a new Mixed AutoRegressive Model (MARM) to significantly reduce the decoding time for the current INR codec, along with a new synthesis network to enhance reconstruction quality. MARM includes our proposed AutoRegressive Upsampler (ARU) blocks, which are highly computationally efficient, and ARM from previous work to balance decoding time and reconstruction quality. We also propose enhancing ARU's performance using a checkerboard two-stage decoding strategy. Moreover, the ratio of different modules can be adjusted to maintain a balance between quality and speed. Comprehensive experiments demonstrate that our method significantly improves computational efficiency while preserving image quality. With different parameter settings, our method can achieve over a magnitude acceleration in decoding time without industrial level optimization or achieve state-of-the-art reconstruction quality compared with other INR codecs. To the best of our knowledge, our method is the first INR-based codec comparable with Ballé et al. [1] in both decoding speed and quality while maintaining low complexity. Jiahong Chen, Bin Chen 0011, Zimo Liu, Baoyi An 0002, Shutao Xia, Zhi Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Image Compression for Resource-Constrained AIoT System With Compressed SensingabstractIn today’s big data era, a key requirement is to implement intelligent semantic analysis (such as image recognition) on data gathered from an extensive array of smart devices in Artificial Intelligence IoT (AIoT) scenarios, all of which is processed at central cloud service providers. Recent advancements in deep-learning-based image compression have fostered semantic compression between machines. However, the deployment of an overparameterized encoder on Internet of Things (IoT) devices remains a challenge due to their restricted computing and storage capabilities. To tackle this issue, we propose a novel approach named compressed sensing (CS)-based asymmetric semantic image compression (CS-ASIC), explicitly designed for resource-constrained AIoT systems. This asymmetric semantic compression scheme intends to surpass the limitations of IoT devices, thereby facilitating efficient semantic compression for machine vision tasks. CS-ASIC notably includes a lightweight front encoder founded on deep image CS techniques, which utilizes rich image priors to learn measurement matrices for sampling. In tandem, a deep iterative decoder is designed cooperatively with the linear encoder offloaded at the server to enhance image reconstruction and semantic analysis across various semantic analysis tasks. Furthermore, we introduce a groundbreaking lossy CS semantic rate-distortion theoretical framework that justifies a compromise in rate for extended semantic distortion. Extensive experimental results underscore the superiority of the proposed CS-ASIC concerning the signal-semantic rate-distortion tradeoff, and its lower encoding complexity over existing codecs in an AIoT simulation environment. Bin Chen 0011, Yujun Huang, Han Qiu 0001, Shutao Xia, Wei Fei, Xuan Wang 0002, Meikang Qiu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video RetrievalabstractGiven a text query, partially relevant video retrieval (PRVR) seeks to find untrimmed videos containing pertinent moments in a database. For PRVR, clip modeling is essential to capture the partial relationship between texts and videos. Current PRVR methods adopt scanning-based clip construction to achieve explicit clip modeling, which is information-redundant and requires a large storage overhead. To solve the efficiency problem of PRVR methods, this paper proposes GMMFormer, a Gaussian-Mixture-Model based Transformer which models clip representations implicitly. During frame interactions, we incorporate Gaussian-Mixture-Model constraints to focus each frame on its adjacent frames instead of the whole video. Then generated representations will contain multi-scale clip information, achieving implicit clip modeling. In addition, PRVR methods ignore semantic differences between text queries relevant to the same video, leading to a sparse embedding space. We propose a query diverse loss to distinguish these text queries, making the embedding space more intensive and contain more semantic information. Extensive experiments on three large-scale video datasets (i.e., TVR, ActivityNet Captions, and Charades-STA) demonstrate the superiority and efficiency of GMMFormer. Jinpeng Wang 0002, Bin Chen 0011, Ziyun Zeng, Shutao Xia |
AAAI | 3 |
| 2024 | Towards Compact 3D Representations via Point Feature Enhancement Masked AutoencodersabstractLearning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they often suffer from heavy computational burdens and heavily rely on massive cross-modal data pairs that are often unavailable, which hinders their applications in practice. Instead, single-modal methods with solely point clouds as input are preferred in real applications due to their simplicity and efficiency. However, such methods easily suffer from limited 3D representations with global random mask input. To learn compact 3D representations, we propose a simple yet effective Point Feature Enhancement Masked Autoencoders (Point-FEMAE), which mainly consists of a global branch and a local branch to capture latent semantic features. Specifically, to learn more compact features, a share-parameter Transformer encoder is introduced to extract point features from the global and local unmasked patches obtained by global random and local block mask strategies, followed by a specific decoder to reconstruct. Meanwhile, to further enhance features in the local branch, we propose a Local Enhancement Module with local patch convolution to perceive fine-grained local context at larger scales. Our method significantly improves the pre-training efficiency compared to cross-modal alternatives, and extensive downstream experiments underscore the state-of-the-art effectiveness, particularly outperforming our baseline (Point-MAE) by 5.16%, 5.00%, and 5.04% in three variants of ScanObjectNN, respectively. Code is available at https://github.com/zyh16143998882/AAAI24-PointFEMAE. Yaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 6 |
| 2024 | Vision-Language Pre-training with Object Contrastive Learning for 3D Scene UnderstandingabstractIn recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building task-specific models, and fail to extract universal 3D vision-language embedding that generalize well. We carefully investigate three common tasks in semantic 3D scene understanding, and derive key insights into the development of a pre-training model. Motivated by these observations, we propose a vision-language pre-training framework 3DVLP (3D vision-language pre-training with object contrastive learning), which transfers flexibly on 3D vision-language downstream tasks. 3DVLP takes visual grounding as the proxy task and introduces Object-level IoU-guided Detection (OID) loss to obtain high-quality proposals in the scene. Moreover, we design Object-level Cross-Contrastive alignment (OCC) task and Object-level Self-Contrastive learning (OSC) task to align the objects with descriptions and distinguish different objects in the scene, respectively. Extensive experiments verify the excellent performance of 3DVLP on three 3D vision-language tasks, reflecting its superiority in semantic 3D scene understanding. Code is available at https://github.com/iridescentttt/3DVLP. Taolin Zhang 0003, Sunan He, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia |
AAAI | 5 |
| 2024 | CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
Hao Fang 0011, Jiawei Kong 0001, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
ECCV (28) | 3 |
| 2024 | A Closer Look at GAN Priors: Exploiting Intermediate Features for Enhanced Model Inversion Attacks
Yixiang Qiu, Hao Fang 0011, Hongyao Yu, Bin Chen 0011, Meikang Qiu, Shutao Xia |
ECCV (32) | 4 |
| 2024 | A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
Jiajun Song, Jiajun Luo, Rongwei Lu, Shuzhao Xie, Bin Chen 0011, Zhi Wang 0001 |
Euro-Par (3) | 5 |
| 2024 | WaterDiff: Perceptual Image Watermarks Via Diffusion ModelabstractRecent studies have demonstrated that diffusion probabilistic models (DPMs) have numerous advantages in image generation through learning a decodable latent representation. This characteristic makes DPMs an appropriate reversible model for encoding and decoding of image watermarking. We present WaterDiff, which leverages pretrained DPMs for perceptual image watermarking problem. Specifically, WaterDiff embeds the watermark into the decomposed stochastic feature, then the stochastic features is combined with the corresponding semantic latent vector to produce a watermarked image via DPMs. This process balances the perceptual quality (stealthiness) and watermarking capacity by fully exploiting the latent diffusion prior. Extensive experiments indicate that WaterDiff guarantee both perceptual imperceptibility and robustness against state-of-the-art watermarking attacks. Yuqi Tan, Yuang Peng, Hao Fang 0011, Bin Chen 0011, Shutao Xia |
ICASSP | 4 |
| 2024 | Progressive Learning with Visual Prompt Tuning for Variable-Rate Image CompressionabstractIn this paper, we propose a progressive learning paradigm for transformer-based variable-rate image compression. Our approach covers a wide range of compression rates with the assistance of the Layer-adaptive Prompt Module (LPM). Inspired by visual prompt tuning, we use LPM to extract prompts for input images and hidden features at the encoder side and decoder side, respectively, which are fed as additional information into the swin transformer layer of a pre-trained transformer-based image compression model to affect the allocation of attention region and the bits, which in turn changes the target compression ratio of the model. To ensure the network is more lightweight, we involves the integration of prompt networks with less convolutional layers. Exhaustive experiments show that compared to methods based on multiple models, which are optimized separately for different target rates, the proposed method arrives at the same performance with 80% savings in parameter storage and 90% savings in datasets. Meanwhile, our model outperforms all current variable bitrate image methods in terms of rate-distortion performance and approaches the state-of-the-art fixed bitrate image compression methods trained from scratch. Shiyu Qin, Yimin Zhou 0011, Jin-Peng Wang, Bin Chen 0011, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICIP | 4 |
| 2024 | Invertible Residual Rescaling Models
Jinmin Li, Tao Dai 0001, Yaohua Zha, Yilu Luo, Longfei Lu, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
IJCAI | 6 |
| 2024 | GladCoder: Stylized QR Code Generation with Grayscale-Aware Denoising Process
Yuqiu Xie, Bolin Jiang, Jiawei Li 0006, Naiqi Li, Bin Chen 0011, Tao Dai 0001, Yuang Peng, Shutao Xia |
IJCAI | 5 |
| 2024 | Section-Wise Revolving NBP-Like Decoders for QC-LDPC CodesabstractRecently, deep learning has demonstrated improvements over classical decoding algorithms in various families of error-correcting codes with short to moderate block lengths. Neural decoders like neural belief propagation (NBP), constructed based on the Tanner graphs of linear codes, have been extensively studied, but their application to longer codes is limited due to increased network complexity with longer block lengths. This complexity leads to higher computational costs and deployment challenges, such as GPU memory usage. To address this, we propose a novel revolving framework for NBP-like decoders tailored to QC-LDPC codes, a commonly used class of error correction codes. Our approach leverages the section-wise cyclic structure inherent in QC-LDPC codes, significantly simplifying the complexity of the network. Experimental results demonstrate the effectiveness of our method across various QC-LDPC codes. Especially, compared to the traditional decoding scheme based on the section-wise cyclic structure, our proposed decoder shows substantial improvements on 5G LDPC codes, exhibits better performance than the traditional min-sum decoder, and approaches the sum-product decoder without a noticeable error floor within the investigated noise levels. Qinshan Zhang, Bin Chen 0011, Tianqu Zhuang, Yong Jiang 0001, Shutao Xia |
ISIT | 2 |
| 2024 | IGSPAD: Inverting 3D Gaussian Splatting for Pose-agnostic Anomaly DetectionabstractPose-agnostic anomaly detection refers to the situation where the pose of test samples is inconsistent with the training dataset, allowing anomalies to appear at any position in any pose. We propose a novel method IGSPAD to address this challenge. Specifically, we employ 3D Gaussian splatting to represent the normal information from the training dataset. To accurately determine the pose of the test sample, we introduce an approach termed Inverting 3D Gaussian Splatting (IGS) to address the challenge of 6D pose estimation for anomalous images. The pose derived from IGS is utilized to render a normal image well-aligned with the test sample. Subsequently, the image encoder of the Segment Anything Model is employed to identify discrepancies between the rendered image and the test sample, predicting the location of anomalies. Experimental results on the MAD dataset demonstrate that the proposed method significantly surpasses the existing state-of-the-art method in terms of precision (from 97.8% to 99.7% at pixel level and from 90.9% to 98.0% at image level) and efficiency. Bolin Jiang, Yuqiu Xie, Jiawei Li 0006, Naiqi Li, Bin Chen 0011, Shutao Xia |
ACM Multimedia | 5 |
| 2024 | BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional BootstrappingabstractAdaptation of
pretrained vision-language models such as CLIP to various downstream tasks have raised great interest in recent researches.
Previous works have proposed a variety of test-time adaptation (TTA) methods to achieve strong generalization without any knowledge of the target domain.
However, existing training-required TTA approaches like TPT necessitate entropy minimization that involves large computational overhead, while training-free methods like TDA overlook the potential for information mining from the test samples themselves.
In this paper, we break down the design of existing popular training-required and training-free TTA methods and bridge the gap between them within our framework.
Specifically, we maintain a light-weight key-value memory for feature retrieval from instance-agnostic historical samples and instance-aware boosting samples.
The historical samples are filtered from the testing data stream and serve to extract useful information from the target distribution, while the boosting samples are drawn from regional bootstrapping and capture the knowledge of the test sample itself.
We theoretically justify the rationality behind our method and empirically verify its effectiveness on both the out-of-distribution and the cross-domain datasets, showcasing its applicability in real-world situations. Taolin Zhang 0003, Jinpeng Wang 0002, Hang Guo 0002, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
NeurIPS | 5 |
| 2024 | Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-ExpertsabstractDesigning single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising results, the existing all-in-one paradigm still suffers from high computational costs as well as limited generalization on unseen degradations. In this work, we introduce an alternative solution to improve the generalization of image restoration models. Drawing inspiration from recent advancements in Parameter Efficient Transfer Learning (PETL), we aim to tune only a small number of parameters to adapt pre-trained restoration models to various tasks. However, current PETL methods fail to generalize across varied restoration tasks due to their homogeneous representation nature. To this end, we propose AdaptIR, a Mixture-of-Experts (MoE) with orthogonal multi-branch design to capture local spatial, global spatial, and channel representation bases, followed by adaptive base combination to obtain heterogeneous representation for different degradations. Extensive experiments demonstrate that our AdaptIR achieves stable performance on single-degradation tasks, and excels in hybrid-degradation tasks, with training only 0.6% parameters for 8 hours. Hang Guo 0002, Tao Dai 0001, Yuanchao Bai, Bin Chen 0011, Xudong Ren, Zexuan Zhu 0001, Shutao Xia |
NeurIPS | 4 |
| 2024 | ReFIR: Grounding Large Restoration Models with Retrieval AugmentationabstractRecent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents or textures when dealing with severe degradations, due to their heavy reliance on limited internal knowledge. In this paper, we propose an orthogonal solution called the Retrieval-augmented Framework for Image Restoration (ReFIR), which incorporates retrieved images as external knowledge to extend the knowledge boundary of existing LRMs in generating details faithful to the original scene. Specifically, we first introduce the nearest neighbor lookup to retrieve content-relevant high-quality images as reference, after which we propose the cross-image injection to modify existing LRMs to utilize high-quality textures from retrieved images. Thanks to the additional external knowledge, our ReFIR can well handle the hallucination challenge and facilitate faithfully results. Extensive experiments demonstrate that ReFIR can achieve not only high-fidelity but also realistic restoration results. Importantly, our ReFIR requires no training and is adaptable to various LRMs. Hang Guo 0002, Tao Dai 0001, Zhihao Ouyang, Taolin Zhang 0003, Yaohua Zha, Bin Chen 0011, Shutao Xia |
NeurIPS | 6 |
| 2024 | LCM: Locally Constrained Compact Point Cloud Model for Masked Point ModelingabstractThe pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limitation, we first conduct a comprehensive analysis of existing Transformer-based MPM, emphasizing the idea that redundancy reduction is crucial for point cloud analysis. To this end, we propose a Locally constrained Compact point cloud Model (LCM) consisting of a locally constrained compact encoder and a locally constrained Mamba-based decoder. Our encoder replaces self-attention with our local aggregation layers to achieve an elegant balance between performance and efficiency. Considering the varying information density between masked and unmasked patches in the decoder inputs of MPM, we introduce a locally constrained Mamba-based decoder. This decoder ensures linear complexity while maximizing the perception of point cloud geometry information from unmasked patches with higher information density. Extensive experimental results show that our compact model significantly surpasses existing Transformer-based models in both performance and efficiency, especially our LCM-based Point-MAE model, compared to the Transformer-based model, achieved an improvement of 1.84%, 0.67%, and 0.60% in performance on the three variants of ScanObjectNN while reducing parameters by 88% and computation by 73%. The code is available at https://github.com/zyh16143998882/LCM. Yaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai 0001, Hang Guo 0002, Bin Chen 0011, Zhi Wang 0001, Zhihao Ouyang, Shutao Xia |
NeurIPS | 6 |
| 2024 | COSMIC: Compress Satellite Image Efficiently via Diffusion CompensationabstractWith the rapidly increasing number of satellites in space and their enhanced capabilities, the amount of earth observation images collected by satellites is exceeding the transmission limits of satellite-to-ground links. Although existing learned image compression solutions achieve remarkable performance by using a sophisticated encoder to extract fruitful features as compression and using a decoder to reconstruct. It is still hard to directly deploy those complex encoders on current satellites' embedded GPUs with limited computing capability and power supply to compress images in orbit. In this paper, we propose COSMIC, a simple yet effective learned compression solution to transmit satellite images. We first design a lightweight encoder (i.e. reducing FLOPs by 2.5~5X) on satellite to achieve a high image compression ratio to save satellite-to-ground links. Then, for reconstructions on the ground, to deal with the feature extraction ability degradation due to simplifying encoders, we propose a diffusion-based model to compensate image details when decoding. Our insight is that satellite's earth observation photos are not just images but indeed multi-modal data with a nature of Text-to-Image pairing since they are collected with rich sensor data (e.g. coordinates, timestep, etc.) that can be used as the condition for diffusion generation. Extensive experiments show that COSMIC outperforms state-of-the-art baselines on both perceptual and distortion metrics. Han Qiu 0001, Maosen Zhang, Jun Liu 0063, Bin Chen 0011, Tianwei Zhang 0004, Hewu Li |
NeurIPS | 5 |
| 2024 | Editable-DeepSC: Cross-Modal Editable Semantic Communication SystemsabstractDifferent from data-oriented communication systems that primarily focus on how to accurately transmit every bit of data, task-oriented semantic communication systems only transmit the specific semantic information required by downstream tasks, strive to minimize the communication overhead and maintain competitive tasks execution performance in the presence of channel noise. However, it is worth noting that in many scenarios, the transmitted semantic information needs to be dynamically modified according to the users' preferences in a conversational and interactive way, which few existing works take into consideration. In this paper, we propose a novel cross-modal editable semantic communication system, named Editable-DeepSC, to tackle this challenge. By utilizing inversion methods based on StyleGAN priors, Editable-DeepSC takes cross-modal text-image pairs as the inputs and transmits the edited information of images based on textual instructions. Extensive numerical results demonstrate that our proposed Editable-DeepSC can achieve remarkable editing effects and transmission efficiency under the perturbations of channel noise, outperforming existing data-oriented communication methods. Bin Chen 0011, Qinshan Zhang, Shutao Xia |
VTC Spring | 2 |
| 2024 | Optimal ternary locally repairable codes
Jie Hao 0001, Shutao Xia, Kenneth W. Shum, Bin Chen 0011, Fang-Wei Fu 0001, Yixian Yang |
Des. Codes Cryptogr. | 4 |
| 2024 | Hugs Bring Double Benefits: Unsupervised Cross-Modal Hashing with Multi-granularity Aligned Transformers
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Dongliang Liao, Gongfu Li, Shutao Xia |
Int. J. Comput. Vis. | 3 |
| 2024 | CMCL: Cross-Modal Compressive Learning for Resource-Constrained Intelligent IoT SystemsabstractCompressive Learning (CL) has proven to be highly successful in executing joint signal sampling and inference for intricate vision tasks through resource-limited Internet of Things (IoT) devices. Recent studies have turned their attention towards utilizing the deep neural networks (DNNs) methodology, also known as DeepCL, to enhance performance in unimodal vision tasks. This approach incorporates learnable compressed sensing in a comprehensive, end-to-end manner. Current DeepCL techniques typically employ initial signal reconstruction as the input for subsequent DNNs for inference. However, this practice presents potential risks such as privacy breaches and reduced performance due to information processing inequality. To address these issues, this paper introduces the first cross-modal compressive learning (CMCL) approach that enables image captioning directly on compressed measurements. When compared to previous DeepCL strategies, the proposed CMCL offers significant improvements in computational efficiency and privacy protection. Extensive experiments demonstrate that CMCL performance is nearly on par with leading image captioning methods, showcasing a metric value that is merely 2.75% lower than the uncompressed method when the data is compressed eightfold. Bin Chen 0011, Yujun Huang, Baoyi An 0002, Yaowei Wang 0001, Xuan Wang 0002 |
IEEE Internet Things J. | 1 |
| 2024 | Multi-scale architectures matter: Examining the adversarial robustness of flow-based lossless compression
Yichong Xia, Bin Chen 0011, Tianshuo Ge, Yujun Huang, Haoqian Wang, Yaowei Wang 0001 |
Pattern Recognit. | 2 |
| 2024 | Pyramid hybrid pooling quantization for efficient fine-grained image retrieval
Ziyun Zeng, Jinpeng Wang 0002, Bin Chen 0011, Tao Dai 0001, Shutao Xia, Zhi Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2024 | New Constructions of MDS Array Codes and Optimal Locally Repairable Array CodesabstractMDS array codes have been extensively studied due to their applications in storage systems. In this paper, we first propose a novel method of constructing MDS array codes by deleting one row and one column from the circulant matrices associated to some polynomials. Several new classes of MDS array codes with flexible parameters are constructed. In particular, we give a new algebraic presentation of the Blaum-Roth codes with sparser parity-check matrices. We also obtain a family of MDS array codes over finite fields with even characteristics whose parity-check matrices have the lowest density. Furthermore, based on these new MDS array codes, we give a general construction of optimal locally repairable array codes (LRACs) achieving the Singleton-type bound. Additionally, we obtain some new optimal LRACs of long lengths. Finally, we present a scheduled algorithm for syndrome computations of binary optimal LRACs with redundancy 4, which can tolerate three failures. The number of XORs per data bit required in our algorithm approaches 2 as the length approaches infinity, which is the same as the MDS codes tolerating three failures. However, the number of nodes required during the repair of a failed node in our optimal LRACs is only about half of that in MDS array codes. Weijun Fang, Jingjie Lv, Bin Chen 0011, Shutao Xia, Xiangyu Chen 0004 |
IEEE Trans. Inf. Theory | 3 |
| 2024 | Bounds and Constructions of Singleton-Optimal Locally Repairable Codes With Small LocalitiesabstractAn$(n, k, d; r)_{q}$-locally repairable code (LRC) is called a Singleton-optimal LRC if it achieves the Singleton-type bound. Analogous to the classical MDS conjecture, the maximal length problem of Singleton-optimal LRCs has attracted a lot of attention in recent years. In this paper, we give an improved upper bound for the length of q-ary Singleton-optimal LRCs with disjoint repair groups such that$(r+1)\mid n$based on the parity-check matrix approach. In particular, for any Singleton-optimal$(n, k, d; r)_{q}$-LRCs, we show that: 1)$n\le q+d-4$, when$r=2$and$d=3e+8$with$e\ge 0$; 2)$n\leq (r+1)\left \lfloor {{\frac {2(q^{2}+q+1)}{r(r+1)} +e+1}}\right \rfloor $, when$d\ge 8$and$\max \left \{{{3,\frac {d-e-6}{e+1}}}\right \}\le r\le \frac {d-e-3}{e+1}$for any$0\le e\le \left \lfloor {{\frac {d-6}{4} }}\right \rfloor $. Furthermore, we establish equivalent connections between the existence of Singleton-optimal$(n,k,d;r)_{q}$-LRCs for$d=6, r=3$and$d=7, r=2$with disjoint repair groups and some subsets of lines in finite projective space with certain properties. Consequently, we prove that the length of q-ary Singleton-optimal LRCs with minimum distance$d=6$and locality$r=3$is upper bounded by$O(q^{1.5})$. We construct Singleton-optimal$(8\le n\le q+1,k,d=6,r=3)_{q}$-LRC with disjoint repair groups such that$4\mid n$and determine the exact value of the maximum code length for some specific q. We also prove the existence of$(n, k, d=7; r=2)_{q}$-Singleton-optimal LRCs for$n \approx \sqrt {2}q$. Weijun Fang, Ran Tao 0010, Fang-Wei Fu 0001, Bin Chen 0011, Shutao Xia |
IEEE Trans. Inf. Theory | 4 |
| 2024 | FlexNF: Flexible Network Function Orchestration for Scalable On-Path Service Chain ServingabstractProgrammable Data Plane (PDP) has been leveraged to offload Network Functions (NFs). Due to its high processing capability, the PDP improves the performance of NFs by more than one order of magnitude. However, the coarse-grained NF orchestration on the PDP makes it hard to fulfill the dynamic service chain demands and unreasonable network function deployment causes long end-to-end delays. In this paper, we propose the Flexible Network Function (FlexNF) deployment on the PDP. First, we design an NF Selection Framework, leveraging the service selection label and re-entering operations for flexible NF orchestration. Second, to support runtime NF reconfiguration to meet the dynamic flow demands, we propose the Per-Flow On-Demand servicing mechanism, where one Match-Action Table with multiple mixed NFs works as different NFs for different flows. Third, to ensure the QoS of flows, on the one hand, we design an SP-aware NF Placement Algorithm to find a near-optimal placement solution that accommodates peak traffic volume while minimizing the overall routing path lengths of all the requests, on the other hand, we design a Two-Stage Service Path Construction Algorithm to provide on-path service while considering load balancing. We implement 15 types of network functions on the P4 switch, based on which we construct the comprehensive experiments. FlexNF reduces the traffic delay by 42.6% while increasing the service chain acceptance rate by five times compared with current solutions. Besides, when switching functions, the FlexNF improves the throughput by 2.04Gbps and reduces the packet loss by 8.269% compared with current solutions. Jingyu Xiao, Xudong Zuo, Qing Li 0006, Dan Zhao 0003, Yong Jiang 0001, Jiyong Sun, Bin Chen 0011 |
IEEE/ACM Trans. Netw. | 8 |
| 2023 | Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature DomainabstractBeyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed compression scenario, the existing method only implements patch matching at the image domain to solve the parallax problem caused by the difference in viewing points. However, the patch matching at the image domain is not robust to the variance of scale, shape, and illumination caused by the different viewing angles, and can not make full use of the rich texture information of the side information image. To resolve this issue, we propose Multi-Scale Feature Domain Patch Matching (MSFDPM) to fully utilizes side information at the decoder of the distributed image compression model. Specifically, MSFDPM consists of a side information feature extractor, a multi-scale feature domain patch matching module, and a multi-scale feature fusion network. Furthermore, we reuse inter-patch correlation from the shallow layer to accelerate the patch matching of the deep layer. Finally, we find that our patch matching in a multi-scale feature domain further improves compression rate by about 20% compared with the patch matching method at image domain. Yujun Huang, Bin Chen 0011, Shiyu Qin, Jiawei Li 0006, Yaowei Wang 0001, Tao Dai 0001, Shutao Xia |
AAAI | 2 |
| 2023 | FSR: A General Frequency-Oriented Framework to Accelerate Image Super-resolution NetworksabstractDeep neural networks (DNNs) have witnessed remarkable achievement in image super-resolution (SR), and plenty of DNN-based SR models with elaborated network designs have recently been proposed. However, existing methods usually require substantial computations by operating in spatial domain. To address this issue, we propose a general frequency-oriented framework (FSR) to accelerate SR networks by considering data characteristics in frequency domain. Our FSR mainly contains dual feature aggregation module (DFAM) to extract informative features in both spatial and transform domains, followed by a four-path SR-Module with different capacities to super-resolve in the frequency domain. Specifically, DFAM further consists of a transform attention block (TABlock) and a spatial context block (SCBlock) to extract global spectral information and local spatial information, respectively, while SR-Module is a parallel network container that contains four to-be-accelerated branches. Furthermore, we propose an adaptive weight strategy for a trade-off between image details recovery and visual quality. Extensive experiments show that our FSR can save FLOPs by almost 40% while reducing inference time by 50% for other SR methods (e.g., FSRCNN, CARN, SRResNet and RCAN). Code is available at https://github.com/THU-Kingmin/FSR. Jinmin Li, Tao Dai 0001, Mingyan Zhu 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 4 |
| 2023 | Contrastive Masked Autoencoders for Self-Supervised Video HashingabstractSelf-Supervised Video Hashing (SSVH) models learn to generate short binary representations for videos without ground-truth supervision, facilitating large-scale video retrieval efficiency and attracting increasing research attention. The success of SSVH lies in the understanding of video content and the ability to capture the semantic relation among unlabeled videos. Typically, state-of-the-art SSVH methods consider these two points in a two-stage training pipeline, where they firstly train an auxiliary network by instance-wise mask-and-predict tasks and secondly train a hashing model to preserve the pseudo-neighborhood structure transferred from the auxiliary network. This consecutive training strategy is inflexible and also unnecessary. In this paper, we propose a simple yet effective one-stage SSVH method called ConMH, which incorporates video semantic information and video similarity relationship understanding in a single stage. To capture video semantic information for better hashing learning, we adopt an encoder-decoder structure to reconstruct the video from its temporal-masked frames. Particularly, we find that a higher masking ratio helps video understanding. Besides, we fully exploit the similarity relationship between videos by maximizing agreement between two augmented views of a video, which contributes to more discriminative and robust hash codes. Extensive experiments on three large-scale video datasets (i.e., FCVID, ActivityNet and YFCC) indicate that ConMH achieves state-of-the-art results. Code is available at https://github.com/huangmozhi9527/ConMH. Jinpeng Wang 0002, Bin Chen 0011, Ziyun Zeng, Shutao Xia |
AAAI | 3 |
| 2023 | Backdoor Attack on Hash-based Image Retrieval via Clean-label Data Poisoning
Kuofeng Gao, Jiawang Bai, Bin Chen 0011, Dongxian Wu, Shutao Xia |
BMVC | 3 |
| 2023 | Learning Transferable Spatiotemporal Representations from Natural Script KnowledgeabstractPre-training on large-scale video data has become a common recipe for learning transferable spatiotemporal representations in recent years. Despite some progress, existing methods are mostly limited to highly curated datasets (e.g., K400) and exhibit unsatisfactory out-of-the-box representations. We argue that it is due to the fact that they only capture pixel-level knowledge rather than spatiotemporal semantics, which hinders further progress in video understanding. Inspired by the great success of image-text pre-training (e.g., CLIP), we take the first step to exploit language semantics to boost transferable spatiotemporal representation learning. We introduce a new pre-text task, Turning to Video for Transcript Sorting (TVTS), which sorts shuffled ASR scripts by attending to learned video representations. We do not rely on descriptive captions and learn purely from video, i.e., leveraging the natural transcribed speech knowledge to provide noisy but useful semantics over time. Our method enforces the vision model to contextualize what is happening over time so that it can re-organize the narrative transcripts, and can seamlessly apply to large-scale uncurated video data in the real world. Our method demonstrates strong out-of-the-box spatiotemporal representations on diverse benchmarks, e.g., +13.6% gains over VideoMAE on SSV2 via linear probing. The code is available at https://github.com/TencentARC/TVTS. Ziyun Zeng, Yuying Ge, Xihui Liu, Bin Chen 0011, Ping Luo 0002, Shutao Xia, Yixiao Ge |
CVPR | 4 |
| 2023 | Difficulty-Aware Data Augmentor for Scene Text RecognitionabstractDeep neural network (DNN) based scene text recognition (STR) methods usually require a large amount of annotated data for training, which is time-consuming and cost-expensive in practice. To address this issue, many data augmentation methods have been developed to train recognizers by improving the diversity of training samples. However, most existing methods neglect the difficulty inherent in samples, and easily suffer from the problem of over-diversity, i.e., the distribution of the augmented data significantly deviates from that of clean data. In this paper, we propose a novel difficulty-aware data augmentation framework for scene text recognition, which jointly considers the difficulty of samples and the strength of augmentations. Specifically, our framework first predicts the sample difficulty, followed by an adaptive data augmentation strategy. Furthermore, we build a more diverse set of augmentation methods for STR and integrate it into our augmentation framework. Extensive experiments on scene text recognition benchmarks show that our augmentation framework significantly improves the performance of recognizers. Guanghao Meng, Tao Dai 0001, Bin Chen 0011, Naiqi Li, Yong Jiang 0001, Shutao Xia |
ICASSP | 3 |
| 2023 | Adaptive ProductAE: SNR-Aware Adaptive Decoding of Neural Product CodesabstractAs a pioneering and deep-learning driven neural channel code, Product Autoencoder (ProductAE) shows significant superiority over Turbo Autoencoder (TurboAE) and other classical codes. However, received noisy codewords with different noise levels have different decoding difficulties for the neural decoder of ProductAE. The existing neural decoder processes all noisy codewords equally without discrimination, and neglects the prior knowledge of signal-to-noise ratio (SNR), thus leading to high decoding complexity. In this paper, we propose a new approach to speed up the decoding process of ProductAE. Intuitively, noisy codewords with different SNR can be recovered by decoders of different complexity, and our proposed novel pipeline, Adaptive ProductAE, an SNR-aware adaptive decoding strategy, adopts deep decoders to different SNR. Specifically, adaptive ProductAE combines an encoding module, an additional classification module, and two independent decoding branches in a unified framework. After receiving noisy codewords, it first judges the decoding difficulty of each signal and assigns them to different branches to get the decoded messages. Furthermore, we introduce a novel training strategy with a gap loss to maintain the classification and decoding performance. It can thus switch to a simpler branch of decoding networks automatically when it comes to signals with a higher SNR, and the overall computational cost can be reduced. Experiments show that our adaptive ProductAE saves up to 30.4% FLOPs for a moderate-length code of parameters (225, 100) and 25.5% FLOPs for a code of parameters (441, 196) in a higher SNR range. Qinshan Zhang, Bin Chen 0011, Yujun Huang, Shutao Xia |
ICC | 2 |
| 2023 | GIFD: A Generative Gradient Inversion Method with Feature Domain OptimizationabstractFederated Learning (FL) has recently emerged as a promising distributed machine learning framework to preserve clients' privacy, by allowing multiple clients to upload the gradients calculated from their local data to a central server. Recent studies find that the exchanged gradients also take the risk of privacy leakage, e.g., an attacker can invert the shared gradients and recover sensitive data against an FL system by leveraging pre-trained generative adversarial networks (GAN) as prior knowledge. However, performing gradient inversion attacks in the latent space of the GAN model limits their expression ability and generalizability. To tackle these challenges, we propose Gradient Inversion over Feature Domains (GIFD), which disassembles the GAN model and searches the feature domains of the intermediate layers. Instead of optimizing only over the initial latent code, we progressively change the optimized layer, from the initial latent space to intermediate layers closer to the output images. In addition, we design a regularizer to avoid unreal image generation by adding a small l1ball constraint to the searching range. We also extend GIFD to the out-of-distribution (OOD) setting, which weakens the assumption that the training sets of GANs and FL tasks obey the same data distribution. Extensive experiments demonstrate that our method can achieve pixel-level reconstruction and is superior to the existing methods. Notably, GIFD also shows great generalizability under different defense strategy settings and batch sizes. Hao Fang 0011, Bin Chen 0011, Xuan Wang 0002, Zhi Wang 0001, Shutao Xia |
ICCV | 2 |
| 2023 | Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud ModelsabstractPre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the efficiency when applying large-scale pre-trained models. Inspired by the recent success of visual prompt tuning (VPT), this paper attempts to explore prompt tuning on pre-trained point cloud models, to pursue an elegant balance between performance and parameter efficiency. We find while instance-agnostic static prompting, e.g. VPT, shows some efficacy in downstream transfer, it is vulnerable to the distribution diversity caused by various types of noises in real-world point cloud data. To conquer this limitation, we propose a novel Instance-aware Dynamic Prompt Tuning (IDPT) strategy for pre-trained point cloud models. The essence of IDPT is to develop a dynamic prompt generation module to perceive semantic prior features of each point cloud instance and generate adaptive prompt tokens to enhance the model's robustness. Notably, extensive experiments demonstrate that IDPT outperforms full finetuning in most tasks with a mere 7% of the trainable parameters, providing a promising solution to parameter-efficient learning for pre-trained point cloud models. Code is available at https://github.com/zyh16143998882/ICCV23-IDPT. Yaohua Zha, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ICCV | 4 |
| 2023 | LKBQ: Pushing the Limit of Post-Training Quantization to Extreme 1 bitabstractRecent advances have shown the potential for post-training quantization (PTQ) to reduce excessive hardware resources and quantize deep models to low bits in a short time, compared with Quantization-Aware Training (QAT). However, existing PTQ approaches lose a lot of accuracies when quantizing the model to extremely low bits, e.g., 1 bit. In this work, we propose layer-by-layer self-knowledge distillation binary post-training quantization (LKBQ), the first method capable of quantizing the weights of neural networks to 1 bit in PTQ domain. We show that careful use of layer-by-layer self-distillation within the LKBQ can provide a significant performance boost. Furthermore, our evaluation results show that the initialization of quantized network weights can have a huge impact on the results. Then we propose three methods for weight initialization. Finally, in light of the characteristics of the binarized network, we propose a method named gradient scaling to further improve efficiency. Our experiments show that LKBQ pushes the limit of PTQ to extreme 1-bit for the first time. Bin Chen 0011, Qian-Wei Wang, Yujun Huang, Shutao Xia |
ICIP | 2 |
| 2023 | Unsupervised Anomaly Detection with Local-Sensitive VQVAE and Global-Sensitive TransformersabstractUnsupervised anomaly detection (UAD) has been widely implemented in industrial and medical applications, which reduces the cost of manual annotation and improves efficiency in disease diagnosis. Recently, deep auto-encoder with its variants has demonstrated its advantages in many UAD scenarios. Training on the normal data, these models are expected to locate anomalies by producing higher reconstruction error for the abnormal areas than the normal ones. However, this assumption does not always hold because of the uncontrollable generalization capability. To solve this problem, we present LSGS, a method that builds on Vector Quantised-Variational Autoencoder (VQVAE) with a novel aggregated codebook and transformers with global attention. In this work, the VQVAE focus on feature extraction and reconstruction of images, and the transformers fit the manifold and locate anomalies in the latent space. Then, leveraging the generated encoding sequences that conform to a normal distribution, we can reconstruct a more accurate image for locating the anomalies. Experiments on various datasets demonstrate the effectiveness of the proposed method. Mingqing Wang, Jiawei Li 0006, Chengxiao Luo, Bin Chen 0011, Shutao Xia, Zhi Wang 0001 |
ICIP | 5 |
| 2023 | DDA: A Dynamic Difficulty-aware Data Augmenter for Image Super-resolutionabstractDeep neural networks (DNNs) have been recently widely used in image super-resolution (SR) and have achieved remarkable performance. However, most existing methods focus on elaborate network design, while rarely considering the training strategy, which affects the model performance and training efficiency. In practice, most SR methods still train the networks with the commonly-used data augmentation (e.g., random crop and sampling), which is shown to converge slowly for deep SR networks. To address this issue, in this paper, we propose a dynamic difficulty-aware data augmenter, named DDA, by considering the restoration difficulty and distribution of input patches. Our DDA mainly consists of difficulty-aware divider, dynamic sampler, and adaptive re-weighter. Specifically, our DDA first uses the difficulty-aware divider to divide the input image into small over-lapping patches, followed by classification into$N$different classes based on the restoration difficulty. Next, dynamic sampler samples the training patches from each class with probability based on training loss. Furthermore, to remedy the imbalance of training patches between different classes, adaptive re-weighter updates the weight of each training patch according to the accumulated training loss. Extensive experiments demonstrate the effectiveness of our DDA on different SR methods by improving training efficiency and model performance across a wide range of scenarios. Xinyi Zhang 0008, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
IJCNN | 3 |
| 2023 | DAGC: Data-Aware Adaptive Gradient CompressionabstractGradient compression algorithms are widely used to alleviate the communication bottleneck in distributed ML. However, existing gradient compression algorithms suffer from accuracy degradation in Non-IID scenarios, because a uniform compression scheme is used to compress gradients at workers with different data distributions and volumes, since workers with larger volumes of data are forced to adapt to the same aggressive compression ratios as others. Assigning different compression ratios to workers with different data distributions and volumes is thus a promising solution. In this study, we first derive a function from capturing the correlation between the number of training iterations for a model to converge to the same accuracy, and the compression ratios at different workers; This function particularly shows that workers with larger data volumes should be assigned with higher compression ratios1to guarantee better accuracy. Then, we formulate the assignment of compression ratios to the workers as an n-variables chi-square nonlinear optimization problem under fixed and limited total communication constrain. We propose an adaptive gradient compression strategy called DAGC, which assigns each worker a different compression ratio according to their data volumes. Our experiments confirm that DAGC can achieve better performance facing highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Jiajun Song, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
INFOCOM | 3 |
| 2023 | Binary MDS Array Codes with Flexible Array Dimensions and Their Fast EncodingabstractIn this short paper, we will provide a new explicit construction of binary MDS array codes with triple parities from their parity-check matrices, which contains array codes with array number 8 (8 bits=1 byte). In addition, to demonstrate the applicability of our MDS array codes, we present an effective decoding method aimed at the erased errors. Furthermore, a fast encoding algorithm of our extended MDS array codes is also explored, whose computational complexity is 2 XORs per bit when their code lengths approach infinity. Jingjie Lv, Weijun Fang, Bin Chen 0011, Shutao Xia, Xiangyu Chen 0004 |
ISIT | 3 |
| 2023 | One-stage Low-resolution Text Recognition with High-resolution Knowledge TransferabstractRecognizing characters from low-resolution (LR) text images poses a significant challenge due to the information deficiency as well as the noise and blur in low-quality images. Current solutions for low-resolution text recognition (LTR) typically rely on a two-stage pipeline that involves super-resolution as the first stage followed by the second-stage recognition. Although this pipeline is straightforward and intuitive, it has to use an additional super-resolution network, which causes inefficiencies during training and testing. Moreover, the recognition accuracy of the second stage heavily depends on the reconstruction quality of the first stage, causing ineffectiveness.In this work, we attempt to address these challenges from a novel perspective: adapting the recognizer to low-resolution inputs by transferring the knowledge from the high-resolution. Guided by this idea, we propose an efficient and effective knowledge distillation framework to achieve multi-level knowledge transfer.Specifically, the visual focus loss is proposed to extract the character position knowledge with resolution gap reduction and character region focus, the semantic contrastive loss is employed to exploit the contextual semantic knowledge with contrastive learning, and the soft logits loss facilitates both local word-level and global sequence-level learning from the soft teacher label.Extensive experiments show that the proposed one-stage pipeline significantly outperforms super-resolution based two-stage frameworks in terms of effectiveness and efficiency, accompanied by favorable robustness.Code is available at https://github.com/csguoh/KD-LTR. Hang Guo 0002, Tao Dai 0001, Mingyan Zhu 0001, Guanghao Meng, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ACM Multimedia | 5 |
| 2023 | Perfect LRCs and k-optimal LRCs
Weijun Fang, Bin Chen 0011, Shutao Xia, Fang-Wei Fu 0001, Xiangyu Chen 0004 |
Des. Codes Cryptogr. | 2 |
| 2023 | Adversarial Examples Generation for Deep Product Quantization Networks on Image RetrievalabstractDeep product quantization networks (DPQNs) have been successfully used in image retrieval tasks, due to their powerful feature extraction ability and high efficiency of encoding high-dimensional visual features. Recent studies show that deep neural networks (DNNs) are vulnerable to input with small and maliciously designed perturbations (a.k.a., adversarial examples) for classification. However, little effort has been devoted to investigating how adversarial examples affect DPQNs, which raises the potential safety hazard when deploying DPQNs in a commercial search engine. To this end, we propose an adversarial example generation framework by generating adversarial query images for DPQN-based retrieval systems. Unlike the adversarial generation for the classic image classification task that heavily relies on ground-truth labels, we alternatively perturb the probability distribution of centroids assignments for a clean query, then we can induce effective non-targeted attacks on DPQNs in white-box and black-box settings. Moreover, we further extend the non-targeted attack to a targeted attack by a novel sample space averaging scheme ([Formula: see text]AS), whose theoretical guarantee is also obtained. Extensive experiments show that our methods can create adversarial examples to successfully mislead the target DPQNs. Besides, we found that our methods both significantly degrade the retrieval performance under a wide variety of experimental settings. The source code is available at https://github.com/Kira0096/PQAG. Bin Chen 0011, Tao Dai 0001, Jiawang Bai, Yong Jiang 0001, Shutao Xia, Xuan Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Contrastive Quantization with Code Memory for Unsupervised Image RetrievalabstractThe high efficiency in computation and storage makes hashing (including binary hashing and quantization) a common strategy in large-scale retrieval systems. To alleviate the reliance on expensive annotations, unsupervised deep hashing becomes an important research problem. This paper provides a novel solution to unsupervised deep quantization, namely Contrastive Quantization with Code Memory (MeCoQ). Different from existing reconstruction-based strategies, we learn unsupervised binary descriptors by contrastive learning, which can better capture discriminative visual semantics. Besides, we uncover that codeword diversity regularization is critical to prevent contrastive learning-based quantization from model degeneration. Moreover, we introduce a novel quantization code memory module that boosts contrastive learning with lower feature drift than conventional feature memories. Extensive experiments on benchmark datasets show that MeCoQ outperforms state-of-the-art methods. Code and configurations are publicly released. Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
AAAI | 3 |
| 2022 | Hugs Are Better Than Handshakes: Unsupervised Cross-Modal Transformer Hashing with Multi-granularity Alignment
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Dongliang Liao, Gongfu Li, Shutao Xia |
BMVC | 3 |
| 2022 | Motion-Aware Graph Reasoning Hashing for Self-supervised Video Retrieval
Ziyun Zeng, Jinpeng Wang 0002, Bin Chen 0011, Shutao Xia |
BMVC | 3 |
| 2022 | Compressive sensing based asymmetric semantic image compression for resource-constrained IoT systemabstractThe widespread application of Internet-of-Things (IoT) and deep learning have made machine-to-machine semantic communication possible. However, it remains challenging to deploy DNN model on IoT devices, due to their limited computing and storage capacity. In this paper, we propose Compressed Sensing based Asymmetric Semantic Image Compression (CS-ASIC) for resource-constrained IoT systems, which consists of a lightweight front encoder and a deep iterative decoder offloaded at the server. We further consider a task-oriented scenario and optimize CS-ASIC for the semantic recognition tasks. The experiment results demonstrate that CS-ASIC achieves considerable data-semantic rate-distortion trade-off, and low encoding complexity over prevailing codecs. Yujun Huang, Bin Chen 0011, Jianghui Zhang, Han Qiu 0001, Shutao Xia |
DAC | 2 |
| 2022 | Improved DC Estimation for JPEG Compression Via Convex RelaxationabstractMass image transmission has undergone an explosion of growth with the development of the internet, DCT-based lossy image compression like JPEG is pervasively conducted to save the transmission bandwidth. Recently, DCT-domain coefficient estimation approaches have been proposed to further improve the compression ratio by discarding DC coefficients at the sender’s end while recovering them at the receiver’s end via DC estimation. However, known DC estimation needs to enumerate all possible DC coefficients. Consequently, they are limited and resource-consuming due to the low delay requirements in real-time transmission. In this paper, we propose an improved DC estimation method via convex relaxation, which achieves state-of-the-art performance in terms of both recovery image quality and time complexity. Extensive experiments across various data sets demonstrate the advantages of our method. Jianghui Zhang, Bin Chen 0011, Yujun Huang, Han Qiu 0001, Zhi Wang 0001, Shutao Xia |
ICIP | 2 |
| 2022 | New constructions of binary MDS array codes and locally repairable array codesabstractIn this paper, we firstly present a new construction of binary maximum distance separable (MDS) array codes, from which some types of new MDS array codes of minimum distance 4 with array dimension (p−1)×(ℓ+2) can be deduced. Based on the construction, binary locally repairable array codes (LRACs) of minimum distance 4 are also explored, whose array dimension is (p−1)×2ℓ and column locality is ℓ − 1. Particularly, when 2 is a primitive root module p, a scheduled algorithm for syndrome computation of the LRACs is proposed, which converges to 2 XORs per data bit when ℓ approaches infinity. Jingjie Lv, Weijun Fang, Bin Chen 0011, Shutao Xia, Xiangyu Chen 0004 |
ISIT | 3 |
| 2022 | Hybrid Contrastive Quantization for Efficient Cross-View Video RetrievalabstractWith the recent boom of video-based social platforms (e.g., YouTube and TikTok), video retrieval using sentence queries has become an important demand and attracts increasing research attention. Despite the decent performance, existing text-video retrieval models in vision and language communities are impractical for large-scale Web search because they adopt brute-force search based on high-dimensional embeddings. To improve efficiency, Web search engines widely apply vector compression libraries (e.g., FAISS [26]) to post-process the learned embeddings. Unfortunately, separate compression from feature encoding degrades the robustness of representations and incurs performance decay. To pursue a better balance between performance and efficiency, we propose the first quantized representation learning method for cross-view video retrieval, namely Hybrid Contrastive Quantization (HCQ). Specifically, HCQ learns both coarse-grained and fine-grained quantizations with transformers, which provide complementary understandings for texts and videos and preserve comprehensive semantic information. By performing Asymmetric-Quantized Contrastive Learning (AQ-CL) across views, HCQ aligns texts and videos at coarse-grained and multiple fine-grained levels. This hybrid-grained learning strategy serves as strong supervision on the cross-view video quantization model, where contrastive learning at different levels can be mutually promoted. Extensive experiments on three Web video benchmark datasets demonstrate that HCQ achieves competitive performance with state-of-the-art non-compressed retrieval methods while showing high efficiency in storage and computation. Code and configurations are available at https://github.com/gimpong/WWW22-HCQ. Jinpeng Wang 0002, Bin Chen 0011, Dongliang Liao, Ziyun Zeng, Gongfu Li, Shutao Xia, Jin Xu 0014 |
WWW | 2 |
| 2022 | Practical protection against video data leakage via universal adversarial head
Jiawang Bai, Bin Chen 0011, Kuofeng Gao, Xuan Wang 0002, Shutao Xia |
Pattern Recognit. | 2 |
| 2022 | Deep image prior based defense against adversarial examples
Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia |
Pattern Recognit. | 3 |
| 2021 | Weakly Supervised Deep Hyperspherical Quantization for Image RetrievalabstractDeep quantization methods have shown high efficiency on large-scale image retrieval. However, current models heavily rely on ground-truth information, hindering the application of quantization in label-hungry scenarios. A more realistic demand is to learn from inexhaustible uploaded images that are associated with informal tags provided by amateur users. Though such sketchy tags do not obviously reveal the labels, they actually contain useful semantic information for supervising deep quantization. To this end, we propose Weakly-Supervised Deep Hyperspherical Quantization (WSDHQ), which is the first work to learn deep quantization from weakly tagged images. Specifically, 1) we use word embeddings to represent the tags and enhance their semantic information based on a tag correlation graph. 2) To better preserve semantic information in quantization codes and reduce quantization error, we jointly learn semantics-preserving embeddings and supervised quantizer on hypersphere by employing a well-designed fusion layer and tailor-made loss functions. Extensive experiments show that WSDHQ can achieve state-of-art performance in weakly-supervised compact coding. Jinpeng Wang 0002, Bin Chen 0011, Qiang Zhang 0026, Zaiqiao Meng, Shangsong Liang, Shutao Xia |
AAAI | 2 |
| 2021 | SwinFGHash: Fine-grained Image Retrieval via Transformer-based Hashing Network
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Shudeng Wu, Shutao Xia |
BMVC | 4 |
| 2021 | Efficient Face Manipulation Via Deep Feature Disentanglement And Reintegration NetabstractDeep neural networks (DNNs) have been widely used in facial manipulation. Existing methods focus on training deeper networks in indirect supervision ways (e.g., feature constraint), or in unsupervised ways (e.g., cycle-consistency loss) due to the lack of ground-truth face images for manipulated outputs. However, such methods can not synthesize realistic face images well and suffer from very high training overhead. To address this issue, we propose a novel Feature Disentanglement and Reintegraion network (FDRNet), which employs ground-truth images as informative supervision and dynamically adapts the fusion of informative features of the ground-truth images effectively and efficiently. FDRNet consists of a Feature Disentanglement (FD) Network and a Feature Reintegration (FR) Network, which encodes informative disentangled representations from the ground-truth images and fuses the disentangled representations to reconstruct the face images. By learning disentangled representations, our method can generate plausible faces conditioned on both landmarks and identities, which can be used for a variety of face manipulation tasks. Experiments on the CelebA-HQ and FFHQ datasets are conducted to demonstrate the superiority of our method over state-of-the-art methods in terms of effectiveness and efficiency. Tao Dai 0001, Bin Chen 0011, Shutao Xia, Xiu Li 0001 |
ICASSP | 3 |
| 2021 | HOCA: Higher-Order Channel Attention for Single Image Super-ResolutionabstractConvolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich feature statistics higher than second-order, thus hindering the representation ability of CNNs. To address this issue, we propose a higher-order channel attention (HOCA) module to enhance the representation ability of CNNs. In our HOCA module, to capture different types of semantic information, we first compute k-order of feature statistics, followed by channel attention to learn the feature interdependencies. Considering the diversity of input contents, we design a gate mechanism to adaptively select a specific k-order channel attention. Besides, our HOCA module serves as a plug-and-play module and can be easily plugged into existing state-of-art CNN-based SR methods. Extensive experiments on public benchmarks show that our HOCA module effectively improves the performance of various CNN-based SR methods. Yalei Lv, Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia, Jingchao Cao |
ICASSP | 3 |
| 2021 | Webly Supervised Deep Attentive QuantizationabstractLearning to hash has been widely applied in large-scale image retrieval. Although current deep hashing methods yield state-of-the-art performance, their heavy dependence on groundtruth information actually makes it difficult to deploy in practical applications such as social media. To solve this problem, we propose a novel method termed Webly Supervised Deep Attentive Quantization (WSDAQ), where deep quantization is trained on web images associated with some userprovided weak tags, without consulting any ground-truth labels. Specifically, we design a tag processing module to leverage semantic information of tags so as to better supervised quantization learning. Besides, we propose an end-to-end trainable Attentive Product Quantization Module (APQM) to quantize deep features of images. Furthermore, we use a noise-contrastive estimation loss to train the model from the perspective of contrastive learning. Experiments validate that WSDAQ is superior to state-of-the-art baselines in compact coding trained on weakly-tagged web images. Jinpeng Wang 0002, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
ICASSP | 2 |
| 2021 | Class Aware Robust TrainingabstractAdversarial training (AT) has been one of the most effective ways to defend adversarial attack. However, existing AT variants exhibit a large imbalanced robust accuracies among different classes, which might harm the robustness of some important class(es) in some real-world applications. For instance, diseased cells are much more important than healthy ones in medical image recognition. Given a certain task, the important class is often a priori. To improve robust accuracy of the important class(es), we are the first to propose a novel adversarial training method with class imbalance taken into account. We term it Class-Aware Robust Training (CART). CART can significantly increase the robustness of the important class(es) by an optional weighted combination of original adversarial example generation and that of the important class. Extensive experiments on three benchmark datasets verify the efficacy of CART for enhancing the robust accuracy of important classes while keeping competitive average robust accuracy. Zhikang Xia, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
ICASSP | 2 |
| 2021 | Attribute Structured Knowledge DistillationabstractKnowledge distillation aims at transferring sufficient knowledge from one cumbersome teacher network to another compressed student network. Most previous knowledge distillation methods mainly focus on mimicking the teacher's output of each instance or inter-instance relations from the teacher to the student, while neglecting the attribute structured relations at an instance level. In this paper, we propose a novel Attribute Structured Knowledge Distillation (ASKD) to transfer attribute-level structured relations. It models two types of structured relations, including inter-region structure and inter-class structure, which captures informative knowledge from the teacher model. Inter-region structure captures local feature relations, while interclass structure focuses on cross-class relations. Transferring such attribute-level structure from the teacher to the student would make the student focus on more informative knowledge of the teacher. Extensive experiments on public datasets, including CIFAR-100 and TinyImageNet, demonstrate the superiority of our method over the state-of-the-art methods under similar architecture and different architecture teacher-student pairs. Tao Dai 0001, Bin Chen 0011, Shutao Xia |
IJCNN | 3 |
| 2021 | Singleton-Optimal LRCs and Perfect LRCs via Cyclic CodesabstractLocally repairable codes (LRCs) have emerged as an important coding scheme in distributed storage systems (DSSs) with relatively low repair cost by accessing fewer non-failure nodes. Theoretical bounds and optimal constructions of LRCs have been widely investigated. Optimal LRCs via cyclic codes provide significant benefit of elegant algebraic structure and efficient encoding procedure. In this paper, we continue to consider the constructions of optimal LRCs via cyclic codes with longer code length. Specifically, we first obtain two classes of Singleton-optimal cyclic LRCs with length$n=3(q+1)$when$3\vert (q-1)$and$q$is even, and length$n=\frac{3}{2}(q+1)$when$3\vert (q-1)$and$q$is odd, respectively. To the best of our knowledge, this is the first construction of q-ary cyclic Singleton-optimal LRCs with length$n > q+1$and minimum distance$d\geq 5$. By using cyclic codes as well, we construct a new family of perfect LRCs with$d=5$, which generalize the result of Goparaju and Calderbank. Weijun Fang, Bin Chen 0011, Shutao Xia, Fang-Wei Fu 0001 |
ISIT | 2 |
| 2021 | Knowledge Distillation via Channel Correlation Structure
Bin Chen 0011, Tao Dai 0001, Maowei Hu, Yong Jiang 0001, Shutao Xia |
KSEM | 2 |
| 2021 | Mix-order Attention Networks for Image RestorationabstractConvolutional neural networks (CNNs) have obtained great success in image restoration tasks, like single image denoising, demosaicing, and super-resolution. However, most existing CNN-based methods neglect the diversity of image contents and degradations in the corrupted images and treat channel-wise features equally, thus hindering the representation ability of CNNs. To address this issue, we propose deep mix-order attention networks (MAN) to extract features that capture rich feature statistics within networks. Our MAN is mainly built on simple residual blocks and our mix-order channel attention (MOCA) module, which further consists of feature gating and feature pooling blocks to capture different types of semantic information. With our MOCA, our MAN can be flexible to handle various types of image contents and degradations. Besides, our MAN can be generalized to different image restoration tasks, like image denoising, super-resolution, and demosaicing. Extensive experiments demonstrate that our method obtains favorably against state-of-the-art methods in terms of quantitative and qualitative metrics. Tao Dai 0001, Yalei Lv, Bin Chen 0011, Zhi Wang 0001, Zexuan Zhu 0001, Shutao Xia |
ACM Multimedia | 3 |
| 2021 | Correlation-based structural dropout for convolutional neural networks
Yuyuan Zeng, Tao Dai 0001, Bin Chen 0011, Shutao Xia, Jian Lu 0002 |
Pattern Recognit. | 3 |
| 2021 | Improved Bounds and Singleton-Optimal Constructions of Locally Repairable Codes With Minimum Distance 5 and 6abstractRepair locality has been an important metric in a distributed storage system (DSS). Erasure codes with small locality are more popular in a DSS, which means fewer available nodes participating in the repair process of failed nodes. Locally repairable codes (LRCs) as a new coding scheme have given more rise to the system performance and attracted a lot of interest in the theoretical research in coding theory. The particular concern among the research problems is the bounds and optimal constructions of LRCs. The problem of optimal constructions of LRCs includes the most important case of Singleton-optimal LRCs whose minimum distance achieves the Singleton-like bound, which is the core consideration in this paper. In this work, we first of all derive an improved and general upper bound on the code length of Singleton-optimal LRCs with minimum distance d = 5, 6, some known constructions are shown to exactly achieve our new bound, which verifies its tightness. For locality r = 2 and distance d = 6, we construct three newSingleton-optimal LRCs whose code length n = 3(q + 1), n = 3(q + √q + 1) and n = 3(2q - 4), respectively. Moreover, we obtain a complete characterization for Singletonoptimal LRCs with r = 2 and d = 6. Such characterization has established an important connection between the existence of Singleton-optimal LRCs and that of a special subset of lines of finite projective plane P G(2, q), thus provides a methodology for constructing LRCs with longer length based on any advance on finite projective plane P G(2, q). In the end, we employ the well-known line-point incidence matrix and Johnson bounds for constant weight codes to derive tighter upper bounds on the code length. These new bounds further help us to prove that some of the previous Singleton-optimal constructions or their extensions achieve the longest possible code length for q = 3, 4, 5, 7. It's worth noting that all of our Singleton-optimal constructions possess small locality r = 2, which are attractive in a DSS. Bin Chen 0011, Weijun Fang, Shutao Xia, Jie Hao 0001, Fang-Wei Fu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2020 | Adversarial Attack on Deep Product Quantization Network for Image RetrievalabstractDeep product quantization network (DPQN) has recently received much attention in fast image retrieval tasks due to its efficiency of encoding high-dimensional visual features especially when dealing with large-scale datasets. Recent studies show that deep neural networks (DNNs) are vulnerable to input with small and maliciously designed perturbations (a.k.a., adversarial examples). This phenomenon raises the concern of security issues for DPQN in the testing/deploying stage as well. However, little effort has been devoted to investigating how adversarial examples affect DPQN. To this end, we propose product quantization adversarial generation (PQ-AG), a simple yet effective method to generate adversarial examples for product quantization based retrieval systems. PQ-AG aims to generate imperceptible adversarial perturbations for query images to form adversarial queries, whose nearest neighbors from a targeted product quantizaiton model are not semantically related to those from the original queries. Extensive experiments show that our PQ-AQ successfully creates adversarial examples to mislead targeted product quantization retrieval models. Besides, we found that our PQ-AG significantly degrades retrieval performance in both white-box and black-box settings. Bin Chen 0011, Tao Dai 0001, Shutao Xia |
AAAI | 2 |
| 2020 | Targeted Attack for Deep Hashing Based Retrieval
Jiawang Bai, Bin Chen 0011, Yiming Li 0004, Dongxian Wu, Weiwei Guo, Shutao Xia, En-Hui Yang |
ECCV (1) | 2 |
| 2020 | Sample-aware Data Augmentor for Scene Text RecognitionabstractDeep neural networks (DNNs) have been widely used in scene text recognition, and achieved remarkable performance. Such DNN-based scene text recognizers usually require plenty of training data for training, but data collection and annotation is usually cost-expensive in practice. To alleviate this issue, data augmentation is often applied to train the scene text recognizers. However, existing data augmentation methods including affine transformation and elastic transformation methods suffer from the problems of under- and over-diversity, due to the complexity of text contents and shapes. In this paper, we propose a sample-aware data augmentor to transform samples adaptively based on the contents of samples. Specifically, our data augmentor consists of three parts: gated module, affine transformation module, and elastic transformation module. In our data augmentor, affine transformation module focuses on keeping the affinity of samples, while elastic transformation module aims to improve the diversity of samples. With the gated module, our data augmentor determines transformation type adaptively based on the properties of training samples and the recognizer capability during the training process. Besides, our framework introduces an adversarial learning strategy to optimize the augmentor and the recognizer jointly. Extensive experiments on scene text recognition benchmarks show that our sample-aware data augmentor significantly improves the performance of state-of-the-art scene text recognizer. Guanghao Meng, Tao Dai 0001, Shudeng Wu, Bin Chen 0011, Jian Lu 0002, Yong Jiang 0001, Shutao Xia |
ICPR | 4 |
| 2020 | Transferable Adversarial Attacks for Deep Scene Text DetectionabstractScene text detection (STD) aims to locate text in images and plays an important role in many computer vision tasks including automatic driving and text recognition systems. Recently, deep neural networks (DNNs) have been widely and successfully used in scene text detection, leading to plenty of DNN-based STD methods including regression-based and segmentation-based STD methods. However, recent studies have also shown that DNN is vulnerable to adversarial attacks, which can significantly degrade the performance of DNN models. In this paper, we investigate the robustness of DNN-based STD methods against adversarial attacks. To this end, we propose a generic and efficient attack method to generate adversarial examples, which are produced by adding small but imperceptible adversarial perturbation to the input images. Experiments on attacking four various models and a real-world STD engine of Google optical character recognition (OCR) show that the state-of-the-art DNN-based STD methods including regression-based and segmentation-based methods are vulnerable to adversarial attacks. Shudeng Wu, Tao Dai 0001, Guanghao Meng, Bin Chen 0011, Jian Lu 0002, Shutao Xia |
ICPR | 4 |
| 2020 | Complete Characterization of Optimal LRCs with Minimum Distance 6 and Locality 2: Improved Bounds and ConstructionsabstractLocally repairable codes (LRCs) with locality r were introduced to recover an erased code symbol by accessing at most r other code symbols. An LRC achieving the well-known Singleton-type bound is called an optimal LRC. Constructing optimal LRCs has been a hot topic of coding theory in recent years. Similar to the famous MDS conjecture, the maximum code length of an optimal LRC has been investigated by Guruswami et al. (TIT2019) and some constructions of optimal LRCs with large code length are also presented by Jin (TIT2019) and Xing and Yuan (arXiv2018). In this paper, we consider the maximum code length of optimal LRCs with minimum distance 6 and locality 2. Firstly, we give a complete characterization for optimal LRCs with d = 6 and r = 2, which shows that the existence of such an LRC is equivalent to the existence of a special subset of lines of finite projective plane PG(2, q). Based on this characterization, we generalize the results of Chen et al. (ISIT2018) and obtain two new constructions of optimal (n, k, d = 6; r = 2)-LRCs with n = 3(q + √q + 1) and n = 3(2q -4), respectively. By using the techniques of line-point incidence matrix and Johnson bound, we show that the code length of any q-ary optimal LRCs with d = 6 and r = 2 must be bounded by O(q1.5). To the best of our knowledge, both of the code length of our new constructions and upper bounds are better than previously known ones. Moreover, we also determine the exact value of the maximum code length of q-ary optimal LRCs with d = 6 and r = 2 for q = 4, 5. Weijun Fang, Bin Chen 0011, Shutao Xia, Fang-Wei Fu 0001 |
ISIT | 2 |
| 2020 | Perfect LRCs and k-Optimal LRCsabstractLinear codes with locality, called locally repairable codes (LRCs), have been applied in distributed storage systems (DSSs) to minimize the number of storage nodes to be downloaded during repairing a failed node. A linear code has locality r if one can recover an erased code symbol by accessing at most r other code symbols. Bounds and constructions of LRCs have been widely investigated in recent years. In this paper, we first propose the definition of perfect LRCs, whose dimension k achieves the Hamming-type bound proposed by Wang et al. (TIT2019). Then we establish important connections of the existence of LRCs with finite geometry and finite fields, and two systematic constructions of perfect LRCs are obtained. Rewriting the Hamming-type bound by the property of integers, we present a new construction of k-optimal LRCs achieving this bound, which have longer code length than the previously known ones. Weijun Fang, Bin Chen 0011, Shutao Xia, Fang-Wei Fu 0001 |
ISIT | 2 |
| 2020 | DIPDefend: Deep Image Prior Driven Defense against Adversarial ExamplesabstractDeep neural networks (DNNs) have shown serious vulnerability to adversarial examples with imperceptible perturbation to clean images. Most existing input-transformation based defense methods (e.g., ComDefend) rely heavily on the learned external priors from an external large training dataset, while neglecting the rich image internal priors of the input itself, thus limiting the generalization of the defense models against the adversarial examples with biased image statistics from the external training dataset. Motivated by deep image prior that can capture rich image statistics from a single image, we propose an effective Deep Image Prior Driven Defense (DIPDefend) method against adversarial examples. With a DIP generator to fit the target/adversarial input, we find that our image reconstruction exhibits quite interesting learning preference from a feature learning perspectives, i.e., the early stage primarily learns the robust features resistant to adversarial perturbation, followed by learning non-robust features that are sensitive to adversarial perturbation. Besides, we develop an adaptive stopping strategy that adapts our method to diverse images. In this way, the proposed model obtains a unique defender for each individual adversarial input, thus being robust to various attackers. Experimental results demonstrate the superiority of our method over the state-of-the-art defense methods against white-box and black-box adversarial attacks. Tao Dai 0001, Dongxian Wu, Bin Chen 0011, Jian Lu 0002, Yong Jiang 0001, Shutao Xia |
ACM Multimedia | 4 |
| 2020 | Mean-removed product quantization for large-scale image retrieval
Bin Chen 0011, Shutao Xia |
Neurocomputing | 2 |
| 2020 | Bounds and Constructions of Locally Repairable Codes: Parity-Check Matrix ApproachabstractA locally repairable code (LRC) is a linear code such that every code symbol can be recovered by accessing a small number of other code symbols. In this paper, we study bounds and constructions of LRCs from the viewpoint of parity-check matrices. Firstly, a simple and unified framework based on parity-check matrix to analyze the bounds of LRCs is proposed, and several new explicit bounds on the minimum distance of LRCs in terms of the field size are presented. In particular, we give an alternate proof of the Singleton-like bound for LRCs first proved by Gopalan et al. Some structural properties on optimal LRCs that achieve the Singleton-like bound are given. Then, we focus on constructions of optimal LRCs over the binary field. It is proved that there are only five classes of possible parameters with which optimal binary LRCs exist. Moreover, by employing the proposed parity-check matrix approach, we completely enumerate all these five classes of optimal binary LRCs attaining the Singleton-like bound in the sense of equivalence of linear codes. Jie Hao 0001, Shutao Xia, Kenneth W. Shum, Bin Chen 0011, Fang-Wei Fu 0001, Yixian Yang |
IEEE Trans. Inf. Theory | 4 |
| 2019 | Improved bounds and Optimal Constructions of Locally Repairable Codes with distance 5 and 6abstractRepair locality has been an important metric in distributed storage systems (DSS). Erasure codes with small locality are more popular in DSS, which means fewer available nodes participating in the repair of failed nodes. Locally repairable codes (LRCs) peoposed as a new coding scheme, give more rise to the system performance and attract a lot of interest in the theoretical research in coding theory. The particular concern among the research problems is the bounds and optimal constructions of LRCs. In this direction, we first of all derive an improved upper bound on the code length of optimal LRCs with minimum distance d = 5, 6, some known constructions are shown to exactly achieve our new bound, which verifies its tightness. Then we construct a class of distance-optimal LRCs based on the structure of a sunflower, whose code length n = 3(q + 1), locality r = 2 and distance d = 6. Note that the code length of this class of LRCs outperforms all known optimal constructions with the same parameters. Moreover, by employing the combinatorial structure of the sunflower and the q-Steiner system, we obtain two classes of k-optimal (dimension-optimal) LRCs with respect to our new bound for d = 6. It's worth noting that all of our optimal constructions possess small locality r = 2, which are attractive in DSS. Bin Chen 0011, Shutao Xia, Jie Hao 0001 |
ISIT | 1 |
| 2019 | Constructions of Optimal $(r, \delta)$ Locally Repairable Codes via Constacyclic CodesabstractLocally repairable codes (LRCs) are introduced in distributed storage systems due to their low repair overhead. An LRC is called optimal if its minimum distance attains the Singleton-like upper bound. Chen et al. (2018) recently studied the constructions of optimal (r, δ)-LRCs with length n | (q+1) and (r + δ - 1) | n, where many classes of optimal cyclic constructions were obtained. In this paper, by employing constacyclic MDS codes, we construct seven classes of optimal (r, δ)-LRCs with new parameters. After adding these new optimal LRCs via constacyclic codes, we have completely obtained all optimal (r, δ)-LRCs with length n | (q + 1) and (r + δ - 1) | n for all possible parameters for the completeness in the coding theory. It is worth noting that the optimal constacyclic LRCs with new parameters provide more alternatives to cyclic LRCs in the practical demands of distributed storage systems, where specific values of n, k, r, and δ are required. Moreover, constacyclic LRCs also possess the encoding and decoding efficiency as cyclic LRCs. Bin Chen 0011, Weijun Fang, Shutao Xia, Fang-Wei Fu 0001 |
IEEE Trans. Commun. | 1 |
| 2018 | On Optimal Pseudo-cyclic ($r, \delta$) Locally Repairable CodesabstractPseudo-cyclic codes is a generalization of cyclic codes and provides a way to obtain MDS codes with more parameters in coding theory. Specially, if a is not a quadratic residue in Fq, xn-a only has quadratic factors over Fqfor even n and k. Based on these facts, we consider the constructions of optimal q-ary pseudo-cyclic (r, δ) locally repairable codes (LRCs) with length n | q+1 in this paper. To be specific, we obtain four classes of optimal pseudo-cyclic (r, δ) -LRCs with new parameters. Bin Chen 0011, Shutao Xia, Jie Hao 0001, Fang-Wei Fu 0001 |
ISIT | 1 |
| 2018 | Bandwidth Efficiency of Distance-optimal Scalar Locally Repairable CodesabstractLocally reparable codes (LRCs) are introduced in distributed storage due to their lower repair degree compared with Reed-Solomon codes. However, Reed-Solomon (RS) codes have recently attracted renewed attention due to the lowbandwidth linear repair scheme proposed by Guruswami and Wootters. Based on the linear repair scheme for RS codes, we consider the bandwidth-efficient repair schemes for distanceoptimal scalar (r, δ) LRCs in this paper. More precisely, we obtain bandwidth-efficient repair schemes for two classes of known distance-optimal scalar (r, δ) LRCs based on polynomial evaluations, whose local sub-codes are proved to be RS codes or GRS codes. Bin Chen 0011, Shutao Xia |
ITW | 1 |
| 2018 | On Optimal (r, δ)-LRCs with Length n | (q+1)abstractOptimal (r, δ) locally repairable codes ((r, δ)-LRCs for short) with length n | (q+1) have been studied in [5] and [6]. In this paper, along with their ideas, by using cyclic or constacyclic codes, we construct three classes of such LRCs with new parameters which are not obtained in [5] and [6]. Thus, optimal (r, δ)-LRCs with length n | (q+1) and (r + δ - 1) | n are completely determined for all possible parameters. Weijun Fang, Fang-Wei Fu 0001, Bin Chen 0011, Shutao Xia |
ITW | 3 |
| 2018 | Constructions of Optimal Cyclic (r, δ) Locally Repairable CodesabstractA code is said to be an r-local locally repairable code (LRC) if each of its coordinates can be repaired by accessing at most r other coordinates. When some of the r coordinates are also erased, the r-local LRC cannot accomplish the local repair, which leads to the concept of (r, δ)-locality. A q-ary [n, k] linear code C is said to have (r, δ)-locality (δ ≥ 2) if for each coordinate i, there exists a punctured subcode of C with support containing i, whose length is at most r+δ-1, and whose minimum distance is at least δ. The (r, δ)-LRC can tolerate δ-1 erasures in every local code (i.e., punctured subcode), which degenerates to an r-local LRC when δ = 2. A q-ary (r, δ) LRC is called optimal if it meets the singleton-like bound for (r, δ)-LRCs. A class of optimal q-ary cyclic r-local LRCs with lengths n | q - 1 were constructed by Tamo, Barg, Goparaju, and Calderbank based on the q-ary Reed-Solomon codes. In this paper, we construct a class of optimal q-ary cyclic (r, δ)-LRCs (δ ≥ 2) with length n | q - 1, which generalizes the results of Tamo et al. Moreover, we construct a new class of optimal q-ary cyclic r-local LRCs with lengths n | q + 1 and a new class of optimal q-ary cyclic (r, δ)-LRCs (δ ≥ 2) with lengths n | q + 1. The constructed optimal LRCs with length n = q + 1 have the best-known length for a given finite field with size q when the minimum distance is larger than 4. Bin Chen 0011, Shutao Xia, Jie Hao 0001, Fang-Wei Fu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2017 | On the linear codes with (r, δ)-locality for distributed storageabstractRecently linear codes with locality properties have attracted a lot of interest due to their desirable applications in distributed storage systems. An [n, k, d] linear code with (r, δ)-locality can enable the local recovery of a failed node in case of more than one node failures. In this paper, we study the theoretical bounds and constructions of linear codes with (r, δ)-locality for all code symbols. A parity-check matrix approach is employed to present an alternate simple proof of the Singleton-like bound for linear codes with all symbol (r, δ)-locality. A refined Singleton-like bound is given for the case that r | k and r + δ - 1 † n. Base on the new proof technique, we enumerate all the possible two classes of optimal binary linear codes meeting the Singleton-like bound. In other words, except the proposed two classes of optimal binary linear codes, there is no other binary linear codes with minimum distance d = n - k - ([k/r] - 1)(δ- 1) + 1. Jie Hao 0001, Shutao Xia, Bin Chen 0011 |
ICC | 3 |
| 2017 | Locally repairable codes with multiple (ri, δi)-localitiesabstractIn distributed storage systems, locally repairable codes (LRCs) are introduced to realize low disk I/O and repair cost. In order to tolerate multiple node failures, the LRCs with (r, δ)-localitty are further proposed. Since hot data is not uncommon in a distributed storage system, both Zeh et al. and Kadhe et al. focus on the LRCs with multiple localities or unequal localities (ML-LRCs) recently, which said that the localities among the code symbols can be different. ML-LRCs are attractive and useful in reducing repair cost for hot data. In this paper, we generalize the ML-LRCs to the (r, δ)-locality case of multiple node failures, and define an LRC with multiple (ri, δi)i ϵ [s]localities (s > 2), where r12s. Such codes ensure that some hot data could be repaired more quickly and have better failure-tolerance in certain cases because of relatively smaller riand larger δι. Then, we derive a Singleton-like upper bound on the minimum distance for the proposed LRCs by employing the regenerating-set technique. Finally, we obtain a class of explicit and structured constructions of optimal ML-LRCs, and further extend them to the cases of multiple (ri, δi)i ϵ [s]localities. Bin Chen 0011, Shutao Xia, Jie Hao 0001 |
ISIT | 1 |
| 2017 | On optimal ternary locally repairable codesabstractIn an [n, k, d] linear code, a code symbol is said to have locality r if it can be repaired by accessing at most r other code symbols. For an (n, k, r) locally repairable code (LRC), the minimum distance satisfies the well-known Singleton-like bound d ≤ n - k - [k/r] + 2. In this paper, we study optimal ternary LRCs meeting this Singleton-like bound by employing a parity-check matrix approach. It is proved that there are only 8 classes of possible parameters with which optimal ternary LRCs exist. Moreover, we obtain explicit constructions of optimal ternary LRCs for all these 8 classes of parameters, where the minimum distance could only be 2, 3, 4, 5 and 6. Jie Hao 0001, Shutao Xia, Bin Chen 0011 |
ISIT | 3 |
| 2017 | On the weight hierarchy of locally repairable codesabstractAn (n, k, r) locally repairable code (LRC) is an [n, k, d] linear code where every code symbol can be repaired from at most r other code symbols. An LRC is said to be optimal if the minimum distance attains the Singleton-like bound d ≤ n - k - ⌈k/r⌉ + 2. The generalized Hamming weights (GHWs) of linear codes are fundamental parameters which have many useful applications. In this paper, we study the GHWs of LRCs. Firstly, we obtain a generalized Singleton-like bound on the i-th (1 ≤ i ≤ k) GHWs of (n, k, r) LRCs. The proposed bound can give the Singleton-like bound when i = 1 and reduce to the classical generalized Singleton bound when there is no locality constraint. Then, it is shown that for optimal (n, k, r) LRCs with r | k, the weight hierarchy can be completely determined. For optimal (n, k, r) LRCs with r | k, some lower bounds on GHWs of LRCs and their dual codes are given. Finally, two general bounds on linear codes in terms of GHWs are presented. Jie Hao 0001, Shutao Xia, Bin Chen 0011, Fang-Wei Fu 0001 |
ITW | 3 |
| 2016 | Some results on optimal locally repairable codesabstractIn a linear code, a code symbol is said to have locality r if it can be repaired by accessing at most r other code symbols. For an (n, k, r) locally repairable codes (LRC), the most important bounds on minimum distances might be the well-known Singleton-like bound and the Cadambe-Mazumdar bound which takes the field size into account. In this paper, we study the constructions of optimal LRCs from the view of parity-check matrices. Firstly, all the optimal binary LRCs meeting the Singleton-like bound are found in the sense of equivalence of linear codes, i.e., except the proposed 4 classes of LRCs, there is no other binary (n, k, r) LRC with minimum distance d = n - k - ⌈k/r⌉+2. Then a class of binary LRCs with distance 4 and arbitrary locality is proposed and shown to be optimal with respect to the Cadambe-Mazumdar bound. Moreover, we give a class of high rate optimal q-ary LRCs meeting the Singleton-like bound with minimum distance 4 while the required field size is only q ≥ r - 1. Finally, several methods to obtain short optimal LRCs from long optimal LRCs are proposed at the end of this paper. Jie Hao 0001, Shutao Xia, Bin Chen 0011 |
ISIT | 3 |
| 2016 | Recursive bounds for locally repairable codes with multiple repair groupsabstractRecently, codes with locality have been widely studied to deal with the node repair problem in distributed storage systems. Locally repairable codes are linear codes with locality properties for code symbols. If a code symbol can be repaired respectively by t disjoint groups of other symbols, each of which has size at most r, this code symbol is said to have (r, t)-locality. In this paper, we present recursive bounds for LRCs with (r, t)-locality for all code symbols. The recursive bounds have simple forms and can be used to derive various bounds for LRCs. Moreover, it is shown that many previous well known bounds of LRCs can be derived by using our recursive bounds. Besides the recursive bounds, we also propose a linear programming bound for LRCs with (r, t)-locality for all code symbols. Jie Hao 0001, Shutao Xia, Bin Chen 0011 |
ISIT | 3 |