VLDB 2026 Research / reviewers in the wild / expert
Jiajun Liang
dblp:184/6428
· DBLP profile ↗
35ranked-venue papers
4as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 2 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 18 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dorado: Scaling SmartNIC Session Tables on Commodity DDRs
Heng Yu 0005, Jiajun Liang, Baozeng Zhang, Guozhi Lin, Xinyi Zhang 0004, Jian Zhao 0006, Ziyue Zhai, Chao Pei, Jilong Wang 0001, Gaogang Xie, Ang Chen 0001, Congcong Miao |
SIGCOMM | 3 |
| 2026 | Pegasus: A Data Center Network for Bare-Metal AI CloudabstractToday, AI cloud is key to serving diverse users with AI services, where cloud networking forms the basis. In this paper, we share our experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment. The key designs of Pegasus include: 1) Network virtualization: a DPU-RNIC decoupled collaborative hardware architecture to enable a single DPU to virtualize multiple RNICs while reducing the power consumption. We design two-level flow tables on both DPU and RNICs to support underlay-overlay IP address translation and ensure isolation. For DPU-RNIC communication, we introduce a per-RNIC communication state machine to reduce communication overhead. 2) Network transport: customized and transparent transport offloading in the RNIC for low-latency and high-throughput communication performance for various AI workloads. We carefully offload per-packet load balancing and credit-based congestion control in RNICs, optimizing reorder delay and eliminating the impacts of hardware jitter. Pegasus has been deployed in production for over two years, currently covering 8K GPUs and supporting a wide range of tenants' AI applications. Xianneng Zou, Zhaoxun Zhou, Xingda Wei, Zhaohe Chen, Yinben Xia, Lizhou Gao, Jiajun Liang, Chunxu Zhao, Jiewei Yang, Yunpeng Guan, Dongbo Gu, Chao Pei, Zekun He, Yachen Wang |
SIGCOMM | 16 |
| 2026 | From surrogate-assisted prediction to optimization: A hybrid intelligence framework for synergistic balance of efficiency, energy, and safety in tunnel boring machine operation
Kangping Gao, Jiajun Liang, Changsheng Wang, Xiaoyin Nie |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | An explainable artificial intelligence-Based approach for intelligent prediction and decision mechanism analysis of tunnel boring machine excavation performance
Jiajun Liang, Kangping Gao, Jingjing Feng |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Partial multi-label learning via adaptive bipartite graph embedding
Jiajun Liang, Dawei Zhao 0002 |
Neurocomputing | 2 |
| 2025 | MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion TransformerabstractDiffusion models have demonstrated superior performance in portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control. This challenge arises from the difficulty in balancing the weak control strength of audio modality and the strong control strength of visual modality. To address this issue, we introduce MegActor-Sigma: a mixed-modal conditional diffusion transformer (DiT), which can flexibly inject audio and visual modality control signals into portrait animation. Specifically, we make substantial advancements over its predecessor, MegActor, by leveraging the promising model structure of DiT and integrating audio and visual conditions through advanced modules within the DiT framework. To further achieve flexible combinations of mixed-modal control signals, we propose a ``Modality Decoupling Control" training strategy to balance the control strength between visual and audio modalities, along with the ``Amplitude Adjustment" inference strategy to freely regulate the motion amplitude of each modality. Finally, to facilitate extensive studies in this field, we design several dataset evaluation metrics to filter out public datasets and solely use this filtered dataset for training. Extensive experiments demonstrate the superiority of our approach in generating vivid portrait animations. Shurong Yang, Juhao Wu, Minhao Jing, Linze Li 0001, Renhe Ji, Jiajun Liang, Haoqiang Fan |
AAAI | 7 |
| 2025 | Asymmetric Decision-Making in Online Knowledge Distillation: Unifying Consensus and DivergenceabstractOnline Knowledge Distillation (OKD) methods represent a streamlined, one-stage distillation training process that obviates the necessity of transferring knowledge from a pretrained teacher network to a more compact student network. In contrast to existing logits-based OKD methods, this paper presents an innovative approach to leverage intermediate spatial representations. Our analysis of the intermediate features from both teacher and student models reveals two pivotal insights: (1) the similar features between students and teachers are predominantly focused on the foreground objects. (2) teacher models emphasize foreground objects more than students. Building on these findings, we propose Asymmetric Decision-Making (ADM) to enhance feature consensus learning for student models while continuously promoting feature diversity in teacher models. Specifically, Consensus Learning for student models prioritizes spatial features with high consensus relative to teacher models. Conversely, Divergence Learning for teacher models highlights spatial features with lower similarity compared to student models, indicating superior performance by teacher models in these regions. Consequently, ADM facilitates the student models to catch up with the feature learning process of the teacher models. Extensive experiments demonstrate that ADM consistently surpasses existing OKD methods across various online knowledge distillation settings and also achieves superior results when transferred to offline knowledge distillation, semantic segmentation and diffusion distillation tasks. Zhaowei Chen, Borui Zhao, Yuchen Ge, Renjie Song, Jiajun Liang |
ICML | 6 |
| 2025 | Flow-GRPO: Training Flow Matching Models via Online RLabstractWe propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation. Jie Liu 0047, Gongye Liu, Jiajun Liang, Yangguang Li 0001, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Wanli Ouyang |
NeurIPS | 3 |
| 2025 | Improving Video Generation with Human FeedbackabstractVideo generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs. Jie Liu 0047, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Menghan Xia, Xintao Wang 0002, Xiaohong Liu 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Yujiu Yang 0001, Wanli Ouyang |
NeurIPS | 3 |
| 2025 | LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional EncodingabstractDiffusion transformers (DiTs) struggle to generate images at resolutions higher than their training resolutions.
The primary obstacle is that the explicit positional encodings (PE), such as RoPE, need extrapolating to unseen positions which degrades performance when the inference resolution differs from training. In this paper, We propose a Length-Extrapolatable Diffusion Transformer (LEDiT) to overcome this limitation. LEDiT needs no explicit PEs, thereby avoiding PE extrapolation. The key innovation of LEDiT lies in the use of causal attention. We demonstrate that causal attention can implicitly encode global positional information and show that such information facilitates extrapolation. We further introduce a locality enhancement module, which captures fine-grained local information to complement the global coarse-grained position information encoded by causal attention. Experimental results on both conditional and text-to-image generation tasks demonstrate that LEDiT supports up to 4× resolution scaling (e.g., from 256$\times$256 to 512$\times$512), achieving better image quality compared to the state-of-the-art
length extrapolation methods. We believe that LEDiT marks a departure from the standard RoPE-based methods and offers a promising insight into length extrapolation. Project page: https://shenzhang2145.github.io/ledit/ Yaning Tan, Zhaowei Chen, Shuheng Li, Caihua Chen, Jiajun Liang |
NeurIPS | 11 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 10 |
| 2025 | Enhancing knowledge distillation for semantic segmentation through text-assisted modular plugins
Letian Wu, Chuankai Zhang, Jiajun Liang, Wankou Yang |
Pattern Recognit. | 5 |
| 2024 | A Simple Baseline for Efficient Hand Mesh ReconstructionabstractHand mesh reconstruction has attracted considerable attention in recent years, with various approaches and techniques being proposed. Some of these methods in-corporate complex components and designs, which, while effective, may complicate the model and hinder efficiency. In this paper, we decompose the mesh decoder into token generator and mesh regressor. Through extensive ablation experiments, we found that the token generator should select discriminating and representative points, while the mesh regressor needs to upsample sparse keypoints into dense meshes in multiple stages. Given these function-alities, we can achieve high performance with minimal computational resources. Based on this observation, we propose a simple yet effective baseline that outperforms state-of-the-art methods by a large margin, while maintaining real-time efficiency. Our method outperforms existing solutions, achieving state-of-the-art (SOTA) results across multiple datasets. On the FreiHAND dataset, our approach produced a PA-MPJPE of 5.8mm and a PA-MPVPE of 6.1mm. Similarly, on the DexYCB dataset, we observed a PA-MPJPE of 5.5mm and a PA-MPVPE of 5.5mm. As for performance speed, our method reached up to 33 frames per second (fps) when using HRNet and up to 70 fps when employing FastViT-MA36. Code will be made available. Zhishan Zhou, Zhi Lv, Minqiang Zou, Jiajun Liang |
CVPR | 6 |
| 2024 | Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang 0001, Nie Lin, Zhishan Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Feng Lu 0005, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Saurabh Gupta 0001, Yoichi Sato 0001, Otmar Hilliges, Hyung Jin Chang, Angela Yao |
ECCV (25) | 7 |
| 2024 | Sparse Beats Dense: Rethinking Supervision in Radar-Camera Depth Completion
Minhao Jing, Jin Wang 0039, Shichao Dong 0001, Jiajun Liang, Haoqiang Fan, Renhe Ji |
ECCV (49) | 5 |
| 2024 | Cascade Prompt Learning for Vision-Language Model Adaptation
Xin Zhang 0170, Zheng Li 0028, Zhaowei Chen, Jiajun Liang, Jian Yang 0003, Xiang Li 0041 |
ECCV (50) | 5 |
| 2024 | HiDiffusion: Unlocking Higher-Resolution Creativity and Efficiency in Pretrained Diffusion Models
Zhaowei Chen, Jiajun Liang |
ECCV (51) | 6 |
| 2024 | Echocardiographic segmentation based on semi-supervised deep learning with attention mechanism
Jiajun Liang, Huijuan Pan, Zhuo Xiang, Harry Qin, Yali Qiu, Libao Guo, Tianfu Wang 0001, Bai Ying Lei |
Multim. Tools Appl. | 1 |
| 2024 | MCSDNet: Mesoscale Convective System Detection Network via Multiscale Spatiotemporal InformationabstractThe accurate detection of mesoscale convective systems (MCSs) is crucial for meteorological monitoring due to their potential to cause significant destruction through severe weather phenomena, such as hail, thunderstorms, and heavy rainfall. However, the existing methods for MCS detection mostly targets on single-frame detection, which just considers the static characteristics and ignores the temporal evolution in the life cycle of MCS. In this article, we propose a novel encoder-decoder neural network named mesoscale convective system detection network (MCSDNet) to detect MCS regions. MCSDNet has a simple architecture and is easy to expand. Different from the previous models, MCSDNet targets on multiframes detection and leverages multiscale spatiotemporal information in remote sensing imagery (RSI). As far as we know, it is the first work to utilize multiscale spatiotemporal information to detect MCS regions. First, we design a multiscale spatiotemporal information module to extract multilevel semantic from different encoder levels, which makes our models can extract more detail spatiotemporal features. Second, spatiotemporal mix unit (STMU), a dual spatiotemporal attention, is introduced to MCSDNet to capture both intraframe features and interframe. Finally, we present MCS remote sensing image (MCSRSI) the first publicly available dataset for multiframes MCS detection based on FY-4A satellite. We also conduct several experiments on MCSRSI and find that our proposed MCSDNet achieves the best performance on MCS detection task when comparing with other baseline methods. We hope that the combination of our open-access dataset and promising results will encourage the future research for MCS detection task and provide a robust framework for related tasks in atmospheric science. Our code is available at:https://github.com/250HandsomeLiang/MCSDNet.git Baoquan Zhang, Jiajun Liang, Rui Ye 0002, Chuyao Luo, Xutao Li 0003, Yunming Ye, Xukai Fu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Boosting Semi-Supervised Learning by Exploiting All Unlabeled DataabstractSemi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch) use a combination of consistency regularization and pseudo-labeling to achieve remarkable successes. However, these methods all suffer from the waste of complicated examples since all pseudo-labels have to be selected by a high threshold to filter out noisy ones. Hence, the examples with ambiguous predictions will not contribute to the training phase. For better leveraging all unlabeled examples, we propose two novel techniques: Entropy Meaning Loss (EML) and Adaptive Negative Learning (ANL). EML incorporates the prediction distribution of non-target classes into the optimization objective to avoid competition with target class, and thus generating more high-confidence predictions for selecting pseudo-label. ANL introduces the additional negative pseudo-label for all unlabeled data to leverage low-confidence examples. It adaptively allocates this label by dynamically evaluating the top-k performance of the model. EML and ANL do not introduce any additional parameter and hyperparameter. We integrate these techniques with FixMatch, and develop a simple yet powerful framework called FullMatch. Extensive experiments on several common SSL benchmarks (CIFAR-10/100, SVHN, STL-10 and ImageNet) demonstrate that FullMatch exceeds FixMatch by a large margin. Integrated with FlexMatch (an advanced FixMatch-based framework), we achieve state-of-the-art performance. Source code is available at https://github.com/megvii-research/FullMatch. Xin Tan 0002, Borui Zhao, Zhaowei Chen, Renjie Song, Jiajun Liang, Xuequan Lu |
CVPR | 6 |
| 2023 | Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection GeneralizationabstractIn this paper, we analyse the generalization ability of binary classifiers for the task of deepfake detection. We find that the stumbling block to their generalization is caused by the unexpected learned identity representation on images. Termed as the Implicit Identity Leakage, this phenomenon has been qualitatively and quantitatively verified among various DNNs. Furthermore, based on such understanding, we propose a simple yet effective method named the ID-unaware Deepfake Detection Model to reduce the influence of this phenomenon. Extensive experimental results demonstrate that our method outperforms the state-of-the-art in both in-dataset and cross-dataset evaluation. The code is available at https://github.com/megvii-research/CADDM. Shichao Dong 0001, Jin Wang 0039, Renhe Ji, Jiajun Liang, Haoqiang Fan, Zheng Ge |
CVPR | 4 |
| 2023 | Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision TransformersabstractAlthough vision transformers (ViTs) have shown promising results in various computer vision tasks recently, their high computational cost limits their practical applications. Previous approaches that prune redundant tokens have demonstrated a good trade-off between performance and computation costs. Nevertheless, errors caused by pruning strategies can lead to significant information loss. Our quantitative experiments reveal that the impact of pruned tokens on performance should be noticeable. To address this issue, we propose a novel joint Token Pruning & Squeezing module (TPS) for compressing vision transformers with higher efficiency. Firstly, TPS adopts pruning to get the reserved and pruned subsets. Secondly, TPS squeezes the information of pruned tokens into partial reserved tokens via the unidirectional nearest-neighbor matching and similarity-based fusing steps. Compared to state-of-the-art methods, our approach outperforms them under all token pruning intensities. Especially while shrinking DeiT-tiny&small computational budgets to 35%, it improves the accuracy by 1%-6% compared with baselines on ImageNet classification. The proposed method can accelerate the throughput of DeiT-small beyond DeiT-tiny, while its accuracy surpasses DeiT-tiny by 4.78%. Experiments on various transformers demonstrate the effectiveness of our method, while analysis experiments prove our higher robustness to the errors of the token pruning policy. Code is available at https://github.com/megvii-research/TPS-CVPR2023. Siyuan Wei, Tianzhu Ye, Jiajun Liang |
CVPR | 5 |
| 2023 | DOT: A Distillation-Oriented TrainerabstractKnowledge distillation transfers knowledge from a large model to a small one via task and distillation losses. In this paper, we observe a trade-off between task and distillation losses, i.e., introducing distillation loss limits the convergence of task loss. We believe that the trade-off results from the insufficient optimization of distillation loss. The reason is: The teacher has a lower task loss than the student, and a lower distillation loss drives the student more similar to the teacher, then a better-converged task loss could be obtained. To break the trade-off, we propose the Distillation-Oriented Trainer (DOT). DOT separately considers gradients of task and distillation losses, then applies a larger momentum to distillation loss to accelerate its optimization. We empirically prove that DOT breaks the trade-off, i.e., both losses are sufficiently optimized. Extensive experiments validate the superiority of DOT. Notably, DOT achieves a +2.59% accuracy improvement on ImageNet-1k for the ResNet50-MobileNetV1 pair. Conclusively, DOT greatly benefits the student’s optimization properties in terms of loss convergence and model generalization. https://github.com/megvii-research/mdistiller. Borui Zhao, Quan Cui, Renjie Song, Jiajun Liang |
ICCV | 4 |
| 2023 | Cumulative Spatial Knowledge Distillation for Vision TransformersabstractDistilling knowledge from convolutional neural networks (CNNs) is a double-edged sword for vision transformers (ViTs). It boosts the performance since the image-friendly local-inductive bias of CNN helps ViT learn faster and better, but leading to two problems: (1) Network designs of CNN and ViT are completely different, which leads to different semantic levels of intermediate features, making spatial-wise knowledge transfer methods (e.g., feature mimicking) inefficient. (2) Distilling knowledge from CNN limits the network convergence in the later training period since ViT’s capability of integrating global information is suppressed by CNN’s local-inductive-bias supervision.To this end, we present Cumulative Spatial Knowledge Distillation (CSKD). CSKD distills spatial-wise knowledge to all patch tokens of ViT from the corresponding spatial responses of CNN, without introducing intermediate features. Furthermore, CSKD exploits a Cumulative Knowledge Fusion (CKF) module, which introduces the global response of CNN and increasingly emphasizes its importance during the training. Applying CKF leverages CNN’s local inductive bias in the early training period and gives full play to ViT’s global capability in the later one. Extensive experiments and analysis on ImageNet-1k and downstream datasets demonstrate the superiority of our CSKD. Code: https://github.com/Zzzzz1/CSKD Borui Zhao, Renjie Song, Jiajun Liang |
ICCV | 3 |
| 2022 | DarkVisionNet: Low-Light Imaging via RGB-NIR Fusion with Deep Inconsistency PriorabstractRGB-NIR fusion is a promising method for low-light imaging. However, high-intensity noise in low-light images amplifies the effect of structure inconsistency between RGB-NIR images, which fails existing algorithms. To handle this, we propose a new RGB-NIR fusion algorithm called Dark Vision Net (DVN) with two technical novelties: Deep Structure and Deep Inconsistency Prior (DIP). The Deep Structure extracts clear structure details in deep multiscale feature space rather than raw input space, which is more robust to noisy inputs. Based on the deep structures from both RGB and NIR domains, we introduce the DIP to leverage the structure inconsistency to guide the fusion of RGB-NIR. Benefits from this, the proposed DVN obtains high-quality low-light images without the visual artifacts. We also propose a new dataset called Dark Vision Dataset (DVD), consisting of aligned RGB-NIR image pairs, as the first public RGB-NIR fusion benchmark. Quantitative and qualitative results on the proposed benchmark show that DVN significantly outperforms other comparison algorithms in PSNR and SSIM, especially in extremely low light conditions. Shuangping Jin, Bingbing Yu, Minhao Jing, Jiajun Liang, Renhe Ji |
AAAI | 5 |
| 2022 | Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal InformationabstractFine-grained image classification is a challenging computer vision task where various species share similar visual appearances, resulting in misclassification if merely based on visual clues. Therefore, it is helpful to leverage additional information, e.g., the locations and dates for data shooting, which can be easily accessible but rarely exploited. In this paper, we first demonstrate that existing multimodal methods fuse multiple features only on a single dimension, which essentially has insufficient help in feature discrimination. To fully explore the potential of multimodal information, we propose a dynamic MLP on top of the image representation, which interacts with multimodal features at a higher and broader dimension. The dynamic MLP is an efficient structure parameterized by the learned embeddings of variable locations and dates. It can be regarded as an adaptive nonlinear projection for generating more discriminative image representations in visual tasks. To our best knowledge, it is the first attempt to explore the idea of dynamic networks to exploit multimodal information in fine-grained image classification tasks. Extensive experiments demonstrate the effectiveness of our method. The t-SNE algorithm visually indicates that our technique improves the recognizability of image representations that are visually similar but with different categories. Furthermore, among published works across multiple fine-grained datasets, dynamic MLP consistently achieves SOTA results11https://paperswithcode.com/dataset/inaturalist and takes third place in the iNaturalist challenge at FGVC822https://www.kaggle.com/c/inaturalist-2021/leaderboard. Code is available at httpsr//glthub.com/megvii-research/DynamicMLPForFinegrained. Lingfeng Yang, Xiang Li 0041, Renjie Song, Borui Zhao, Juntian Tao, Jiajun Liang, Jian Yang 0003 |
CVPR | 7 |
| 2022 | Decoupled Knowledge DistillationabstractState-of-the-art distillation methods are mainly based on distilling deep features from intermediate layers, while the significance of logit distillation is greatly overlooked. To provide a novel viewpoint to study logit distillation, we re-formulate the classical KD loss into two parts, i.e., target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD). We empirically investigate and prove the effects of the two parts: TCKD transfers knowledge concerning the “difficulty” of training samples, while NCKD is the prominent reason why logit distillation works. More importantly, we reveal that the classical KD loss is a coupled formulation, which (1) suppresses the effectiveness of NCKD and (2) limits the flexibility to balance these two parts. To address these issues, we present Decoupled Knowledge Distillation (DKD), enabling TCKD and NCKD to play their roles more efficiently and flexibly. Compared with complex feature-based methods, our DKD achieves comparable or even better results and has better training efficiency on CIFAR-100, ImageNet, and MS-COCO datasets for image classification and object detection tasks. This paper proves the great potential of logit distillation, and we hope it will be helpful for future research. The code is available at https://github.com/megviiresearch/mdistiller. Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, Jiajun Liang |
CVPR | 5 |
| 2022 | Discriminability-Transferability Trade-Off: An Information-Theoretic Perspective
Quan Cui, Bingchen Zhao, Borui Zhao, Renjie Song, Boyan Zhou, Jiajun Liang, Osamu Yoshie |
ECCV (26) | 7 |
| 2022 | Explaining Deepfake Detection by Analysing Image Matching
Shichao Dong 0001, Jin Wang 0039, Jiajun Liang, Haoqiang Fan, Renhe Ji |
ECCV (14) | 3 |
| 2022 | Efficient One Pass Self-distillation with Zipf's Label Smoothing
Jiajun Liang, Linze Li 0001, Zhaodong Bing, Borui Zhao, Haoqiang Fan |
ECCV (11) | 1 |
| 2021 | Information Theoretic Limits of Exact Recovery in Sub-hypergraph Models for Community DetectionabstractIn this paper, we study the information theoretic bounds for exact recovery in sub-hypergraph models for community detection. We define a general model called the$m$-uniform sub-hypergraph stochastic block model (m-ShSBM). Under the$m$-ShSBM, we use Fano's inequality to identify the region of model parameters where any algorithm fails to exactly recover the planted communities with a large probability. We also identify the region where a Maximum Likelihood Estimation (MLE) algorithm succeeds to exactly recover the communities with high probability. Our bounds are tight up to a log($k$) term and pertain to the community detection problems in various models such as the planted hypergraph stochastic block model, the planted densest sub-hypergraph model, and the planted multipartite hypergraph model. Jiajun Liang, Chuyang Ke, Jean Honorio |
ISIT | 1 |
| 2021 | Sharp Impossibility Results for Hyper-graph TestingabstractIn a broad Degree-Corrected Mixed-Membership (DCMM) setting, we test whether a non-uniform hypergraph has only one community or has multiple communities. Since both the null and alternative hypotheses have many unknown parameters, the challenge is, given an alternative, how to identify the null that is hardest to separate from the alternative. We approach this by proposing a degree matching strategy where the main idea is leveraging the theory for tensor scaling to create a least favorable pair of hypotheses. We present a result on standard minimax lower bound theory and a result on Region of Impossibility (which is more informative than the minimax lower bound). We show that our lower bounds are tight by introducing a new test that attains the lower bound up to a logarithmic factor. We also discuss the case where the hypergraphs may have mixed-memberships. Jiashun Jin, Zheng Tracy Ke, Jiajun Liang |
NeurIPS | 3 |
| 2020 | Augmentation Data Synthesis Via Gans: Boosting Latent Fingerprint ReconstructionabstractLatent fingerprint reconstruction is a vital preprocessing step for its identification. This task is very challenging due to not only existing complicated degradation patterns but also its scarcity of paired training data. To address these challenges, we propose a novel generative adversarial network (GAN) based data augmentation scheme to improve such reconstruction. It translates the abundant clean fingerprints to their corresponding latent ones, only exploiting a small-scale latent dataset and an unpaired large-scale clean dataset, from which a large-scale paired clean-latent augmentation set is built for the reconstruction task. Specifically, our method models the distribution of the latent degradation patterns into a Gaussian one and generates latent fingerprints based on the sampled degradation patterns and clean fingerprints. Besides, we develop an auxiliary training procedure to stabilize training and further disentangle ridge structures and degradation patterns by regressing a latent fingerprint from its latent representation and its corresponding binarized fingerprint. Boosted by the proposed data augmentation, our reconstruction shows significant improvements in visual evaluation and fingerprint identification performance. Yi Wang 0004, Jiajun Liang, Yong Jiang 0001 |
ICASSP | 3 |
| 2019 | Scene Text Recognition from Two-Dimensional PerspectiveabstractInspired by speech recognition, recent state-of-the-art algorithms mostly consider scene text recognition as a sequence prediction problem. Though achieving excellent performance, these methods usually neglect an important fact that text in images are actually distributed in two-dimensional space. It is a nature quite different from that of speech, which is essentially a one-dimensional signal. In principle, directly compressing features of text into a one-dimensional form may lose useful information and introduce extra noise. In this paper, we approach scene text recognition from a two-dimensional perspective. A simple yet effective model, called Character Attention Fully Convolutional Network (CA-FCN), is devised for recognizing the text of arbitrary shapes. Scene text recognition is realized with a semantic segmentation network, where an attention mechanism for characters is adopted. Combined with a word formation module, CA-FCN can simultaneously recognize the script and predict the position of each character. Experiments demonstrate that the proposed algorithm outperforms previous methods on both regular and irregular text datasets. Moreover, it is proven to be more robust to imprecise localizations in the text detection phase, which are very common in practice. Minghui Liao, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lv, Cong Yao, Xiang Bai |
AAAI | 5 |
| 2017 | EAST: An Efficient and Accurate Scene Text DetectorabstractPrevious approaches for scene text detection have already achieved promising performances across various benchmarks. However, they usually fall short when dealing with challenging scenarios, even when equipped with deep neural network models, because the overall performance is determined by the interplay of multiple stages and components in the pipelines. In this work, we propose a simple yet powerful pipeline that yields fast and accurate text detection in natural scenes. The pipeline directly predicts words or text lines of arbitrary orientations and quadrilateral shapes in full images, eliminating unnecessary intermediate steps (e.g., candidate aggregation and word partitioning), with a single neural network. The simplicity of our pipeline allows concentrating efforts on designing loss functions and neural network architecture. Experiments on standard datasets including ICDAR 2015, COCO-Text and MSRA-TD500 demonstrate that the proposed algorithm significantly outperforms state-of-the-art methods in terms of both accuracy and efficiency. On the ICDAR 2015 dataset, the proposed algorithm achieves an F-score of 0.7820 at 13.2fps at 720p resolution. Xinyu Zhou 0004, Cong Yao, Yuzhi Wang, Shuchang Zhou 0001, Weiran He, Jiajun Liang |
CVPR | 7 |