EDBT 2026 Demo / reviewers in the wild / expert
Linlin Yang 0001
dblp:84/6484-1
· DBLP profile ↗
46ranked-venue papers
3as first author
40since 2021 · last 2026
0000-0001-6752-0252ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 3 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficiently Seeking Flat Minima for Better Generalization in Fine-Tuning Large Language Models and BeyondabstractLittle research explores the correlation between the expressive ability and generalization ability of the low-rank adaptation (LoRA). Sharpness-Aware Minimization (SAM) improves model generalization for both Convolutional Neural Networks (CNNs) and Transformers by encouraging convergence to locally flat minima. However, the connection between sharpness and generalization has not been fully explored for LoRA due to the lack of tools to either empirically seek flat minima or develop theoretical methods. In this work, we propose Flat Minima LoRA (FMLoRA) and its efficient version i.e., EFMLoRA, to seek flat minima for LoRA. Concretely, we theoretically demonstrate that perturbations in the full parameter space can be transferred to the low-rank subspace. This approach eliminates the potential interference introduced by perturbations across multiple matrices in the low-rank subspace. Our extensive experiments on large language models and vision-language models demonstrate that EFMLoRA achieves optimization efficiency comparable to that of LoRA while simultaneously attaining comparable or even better performance. For example, on the GLUE dataset with RoBERTa-large, EFMLoRA outperforms LoRA and full fine-tuning by 1.0% and 0.5% on average, respectively. On vision-language models e.g., Qwen-VL-Chat, there are performance improvements of 1.5% and 1.0% on the SQA and VizWiz datasets, respectively. These empirical results also verify that the generalization of LoRA is closely related to sharpness, which is omitted by previous methods. Jiaxin Deng, Qingcheng Zhu, Junbiao Pang, Linlin Yang 0001, Zhongqian Fu, Baochang Zhang 0001 |
AAAI | 4 |
| 2026 | AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D GenerationabstractOptimization‐based text‑to‑3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "semantic over-smoothing" artifacts. As such, we reformulate text‑to‑3D optimization as mapping a *dynamically evolving source* distribution to a fixed target distribution. We cast the problem into a dual‑conditioned latent space, conditioned on both the text prompt and the intermediately rendered image. Given this joint setup, we observe that the image condition naturally anchors the current source distribution. Building on this insight, we introduce AnchorDS, an improved score distillation mechanism that provides state‑anchored guidance with image conditions and stabilizes generation. We further penalize erroneous source estimates and design a lightweight filter strategy and fine‑tuning strategy that refines the anchor with negligible overhead. AnchorDS produces finer-grained detail, more natural colours, and stronger semantic consistency, particularly for complex prompts, while maintaining efficiency. Extensive experiments show that our method surpasses previous methods in both quality and efficiency. Jiayin Zhu, Linlin Yang 0001, Yicong Li 0004, Angela Yao |
AAAI | 2 |
| 2026 | BinParam: Binarized human parametric modeling via distribution alignment and orthogonal residuals
Linlin Yang 0001, Ziqi Xie, Boshu Jia, Baochang Zhang 0001, Libiao Jin |
Neurocomputing | 2 |
| 2026 | Kronecker reparameterized large kernel for image compressed sensing
Jiao Xie, Lingfu Jiang, Heming Jia, Shaohui Lin, Yinqi Zhang, Linlin Yang 0001, Junjun Jiang |
Neurocomputing | 6 |
| 2026 | HFAT-HMR: Empowering ViT for human mesh recovery via high-frequency enhancement and auxiliary tokens
Linlin Yang 0001, Boshu Jia, Baochang Zhang 0001, Libiao Jin |
Neurocomputing | 3 |
| 2026 | LDFE: Laplacian Decoupled Feature Enhancement block for dual-stream CNN-based RGB-IR object detection
Xiaoyan Luo, Linlin Yang 0001, Haodong Zhu, Xiaorong Shi, Guodong Guo, Baochang Zhang 0001 |
Pattern Recognit. | 3 |
| 2026 | Security-aware post-training quantization for Mixture-of-Experts large language models
Shiran Ge, Zhiyi Zhu, Canjia Li, Linlin Yang 0001, Baochang Zhang 0001 |
Pattern Recognit. | 5 |
| 2026 | Noise-Robust tiny object localization with flows
Huixin Sun, Linlin Yang 0001, Ronyu Chen, Kerui Gu, Baochang Zhang 0001, Angela Yao, Xianbin Cao 0001 |
Pattern Recognit. | 2 |
| 2026 | Industrial Scene Gas Leakage Detection: A Cross-Attention Based Multimodal Feature Difference Network and a New BenchmarkabstractIndustrial gas leakage detection is critically important for safety and environmental protection. While infrared imaging enables detection of invisible gases, two challenges remain: existing datasets lack realistic industrial scenarios, and current methods struggle to distinguish gas plumes from background interferences or segment discontinuous gas distributions. This paper introduces a benchmark comprising an Industrial RGB-Thermal Dataset (IRTD) with gas emission and leakage data from laboratory and industrial sites. A VLM-assisted RGBThermal detection framework with a Cross-Attention based Feature Difference (CAFD) module is designed to enhance gasspecific feature differentiation by computing inter-modal feature discrepancies. Evaluations on public datasets and IRTD demonstrate state-of-the-art results. Linlin Yang 0001, Xingyu Guo, Sheng Xu 0007, Xianbin Cao 0001, Baochang Zhang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | RINGet: A Robust Watermarking Framework for Diffusion-Based Video GenerationabstractWith the rapid advancement of diffusion-based text-to-video generation, challenges surrounding video ownership verification and copyright protection have become increasingly urgent. Traditional digital watermarking techniques are typically handcrafted for specific types of distortions, while existing diffusion-based watermarking methods are primarily applied to image-level tasks. As a result, the effectiveness of these methods significantly diminishes when videos undergo complex transformations. To address this issue, we introduce RINGet, a robust watermarking framework for diffusion-based video generation. RINGet embeds user-defined keys into the initial latent variables in the Fourier domain while maintaining imperceptibility in the spatial domain through reversible Fourier transforms. The framework adopts a radius-based and ring-based segmentation strategies to improve robustness against rotation while preserve the quality of the generated videos. Moreover, to mitigate distribution shifts caused by watermark embedding, the key is partitioned into discrete segments and distributed across different initial latent variables. Extensive experiments demonstrate the superiority of the RINGet framework over traditional approaches in terms of robustness, particularly under severe perceptual-domain distortions, while preserving high video quality and inference efficiency. Wenzhu Zhang, Linlin Yang 0001, Shunyang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | InstructHumans: Editing Animated 3D Human Textures With InstructionsabstractWe present InstructHumans, a novel framework for instruction-driven animatable 3D human texture editing. Existing text-based 3D editing methods often directly apply Score Distillation Sampling (SDS). SDS, designed for generation tasks, cannot account for the defining requirement of editing – maintaining consistency with the source avatar. This work shows that naively using SDS harms editing, as it may destroy consistency. We propose a modified SDS for Editing (SDS-E) that selectively incorporates subterms of SDS across diffusion timesteps. We further enhance SDS-E with spatial smoothness regularization and gradient-based viewpoint sampling for edits with sharp and high-fidelity detailing. Incorporating SDS-E into a 3D human texture editing framework allows us to outperform existing 3D editing methods. Our avatars faithfully reflect the textual edits while remaining consistent with the original avatars. Project page: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://jyzhu.top/instruct-humans/</uri>. Jiayin Zhu, Linlin Yang 0001, Angela Yao |
IEEE Trans. Multim. | 2 |
| 2025 | Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose EstimationabstractRecent synthetic 3D human datasets for the face, body, and hands have pushed the limits on photorealism. Face recognition and body pose estimation have achieved state-of-the-art performance using synthetic training data alone, but for the hand, there is still a large synthetic-to-real gap. This paper presents the first systematic study of the synthetic-to-real gap of 3D hand pose estimation. We analyze the gap and identify key components such as the forearm, image frequency statistics, hand pose, and object occlusions. To facilitate our analysis, we propose a data synthesis pipeline to synthesize high-quality data. We demonstrate that synthetic hand data can achieve the same level of accuracy as real data when integrating our identified components, paving the path to use synthetic data alone for hand pose estimation. Code and data are available at: https://github.com/delaprada/HandSynthesis.git. Zhuoran Zhao 0003, Linlin Yang 0001, Pengzhan Sun 0001, Pan Hui 0001, Angela Yao |
CVPR | 2 |
| 2025 | SET: Spectral Enhancement for Tiny Object DetectionabstractDeep learning has significantly advanced the object detection field. However, tiny object detection (TOD) remains a challenging problem. We provide a new analysis method to examine the TOD challenge through occlusion-based attribution analysis in the frequency domain. We observe that tiny objects become less distinct after feature encoding and can benefit from the removal of high-frequency information. In this paper, we propose a novel approach named Spectral Enhancement for Tiny object detection (SET), which amplifies the frequency signatures of tiny objects in a heterogeneous architecture. SET includes two modules. The Hierarchical Background Smoothing (HBS) module suppresses high-frequency noise in the background through adaptive smoothing operations. The Adversarial Perturbation Injection (API) module leverages adversarial perturbations to increase feature saliency in critical regions and prompt the refinement of object features during training. Extensive experiments on four datasets demonstrate the effectiveness of our method. Especially, SET boosts the prior art RFLA by 3.2% AP on the AI-TOD dataset. Huixin Sun, Runqi Wang, Yanjing Li, Linlin Yang 0001, Shaohui Lin, Xianbin Cao 0001, Baochang Zhang 0001 |
CVPR | 4 |
| 2025 | Uncertainty-Aware Gradient Stabilization for Small Object Detection
Huixin Sun, Yanjing Li, Linlin Yang 0001, Xianbin Cao 0001, Baochang Zhang 0001 |
ICCV | 3 |
| 2025 | WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionabstractLeveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveMamba, a cross-modality fusion method that efficiently integrates the unique and complementary frequency features of RGB and IR decomposed by Discrete Wavelet Transform (DWT). An improved detection head incorporating the Inverse Discrete Wavelet Transform (IDWT) is also proposed to reduce information loss and produce the final detection results. The core of our approach is the introduction of WaveMamba Fusion Block (WMFB), which facilitates comprehensive fusion across low-/high-frequency sub-bands. Within WMFB, the Low-frequency Mamba Fusion Block (LMFB), built upon the Mamba framework, first performs initial low-frequency feature fusion with channel swapping, followed by deep fusion with an advanced gated attention mechanism for enhanced integration. High-frequency features are enhanced using a strategy that applies an ``absolute maximum" fusion approach. These advancements lead to significant performance gains, with our method surpassing state-of-the-art approaches and achieving average mAP improvements of 4.5% on four benchmarks. Haodong Zhu, Linlin Yang 0001, Hong Li 0016, Yuguang Yang 0007, Yangyang Ren, Qingcheng Zhu, Zichao Feng, Changbai Li, Shaohui Lin, Runqi Wang, Xiaoyan Luo, Baochang Zhang 0001 |
ICCV | 3 |
| 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detectionabstractZero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as prompts for the target disease name during the inference phase.
However, these methods typically treat prompts as equivalent context to the target name, making it difficult to assign specific disease knowledge based on visual information, leading to a coarse alignment between images and target descriptions. In this paper, we propose StructuralGLIP, which introduces an auxiliary branch to encode prompts into a latent knowledge bank layer-by-layer, enabling more context-aware and fine-grained alignment. Specifically, in each layer, we select highly similar features from both the image representation and the knowledge bank, forming structural representations that capture nuanced relationships between image patches and target descriptions. These features are then fused across modalities to further enhance detection performance.
Extensive experiments demonstrate that StructuralGLIP achieves a +4.1\% AP improvement over prior state-of-the-art methods across seven zero-shot medical detection benchmarks, and consistently improves fine-tuned models by +3.2\% AP on endoscopy image datasets. Yuguang Yang 0007, Tongfei Chen, Linlin Yang 0001, Chunyu Xie, Dawei Leng, Xianbin Cao 0001, Baochang Zhang 0001 |
ICLR | 4 |
| 2025 | Efficient Low-Bit Quantization with Adaptive Scales for Multi-Task Co-TrainingabstractCo-training can achieve parameter-efficient multi-task models but remains unexplored for quantization-aware training. Our investigation shows that directly introducing co-training into existing quantization-aware training (QAT) methods results in significant performance degradation. Our experimental study identifies that the primary issue with existing QAT methods stems from the inadequate activation quantization scales for the co-training framework. To address this issue, we propose Task-Specific Scales Quantization for Multi-Task Co-Training (TSQ-MTC) to tackle mismatched quantization scales. Specifically, a task-specific learnable multi-scale activation quantizer (TLMAQ) is incorporated to enrich the representational ability of shared features for different tasks. Additionally, we find that in the deeper layers of the Transformer model, the quantized network suffers from information distortion within the attention quantizer. A structure-based layer-by-layer distillation (SLLD) is then introduced to ensure that the quantized features effectively preserve the information from their full-precision counterparts. Our extensive experiments in two co-training scenarios demonstrate the effectiveness and versatility of TSQ-MTC. In particular, we successfully achieve a 4-bit quantized low-level visual foundation model based on IPT, which attains a PSNR comparable to the full-precision model while offering a $7.99\times$ compression ratio in the $\times4$ super-resolution task on the Set5 benchmark. Linlin Yang 0001, Yanjing Li, Guodong Guo, Xianbin Cao 0001, Baochang Zhang 0001 |
ICLR | 3 |
| 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTsabstractVision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel alignment with the original images. To address these issues, we propose a systematic framework for 3D pose estimation, called ExtPose. ExtPose extends image ViT to the challenging scenario and video setting by taking in additional 2D pose evidence and capturing temporal information in a full attention-based manner. We use 2D human skeleton images to integrate structured 2D pose information. By sharing parameters and attending across modalities and frames, we enhance the consistency between 3D poses and 2D videos without introducing additional parameters. We achieve state-of-the-art (SOTA) performance on multiple human and hand pose estimation benchmarks with substantial improvements to 34.0mm (-23%) on 3DPW and 4.9mm (-18%) on FreiHAND in PA-MPJPE over the other ViT-based methods respectively. Rongyu Chen, Lian Zhuo, Linlin Yang 0001, Qi Wang 0148, Liefeng Bo, Bang Zhang, Angela Yao |
ICML | 3 |
| 2025 | M3DP: Optimizing 2D vision tasks with minimal 3D object information
Yanjing Li, Linlin Yang 0001, Xinkai Liang, Xianbin Cao 0001, Qi Wang 0009, Baochang Zhang 0001 |
Neurocomputing | 3 |
| 2025 | Normalizing Batch Normalization for Long-Tailed RecognitionabstractIn real-world scenarios, the number of training samples across classes usually subjects to a long-tailed distribution. The conventionally trained network may achieve unexpected inferior performance on the rare class compared to the frequent class. Most previous works attempt to rectify the network bias from the data-level or from the classifier-level. Differently, in this paper, we identify that the bias towards the frequent class may be encoded into features, i.e., the rare-specific features which play a key role in discriminating the rare class are much weaker than the frequent-specific features. Based on such an observation, we introduce a simple yet effective approach, normalizing the parameters of Batch Normalization (BN) layer to explicitly rectify the feature bias. To achieve this end, we represent theWeight/Bias parameters of a BN layer as a vector, normalize it into a unit one and multiply the unit vector by a scalar learnable parameter. Through decoupling the direction and magnitude of parameters in BN layer to learn, the Weight/Bias exhibits a more balanced distribution and thus the strength of features becomes more even. Extensive experiments on various long-tailed recognition benchmarks (i.e., CIFAR-10/100-LT, ImageNet-LT and iNaturalist 2018) show that our method outperforms previous state-of-the-arts remarkably. Yuxiang Bao, Guoliang Kang, Linlin Yang 0001, Xiaoyue Duan, Baochang Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | AQ-DETR: Low-Bit Quantized Detection Transformer with Auxiliary QueriesabstractDEtection TRansformer (DETR)-based models have achieved remarkable performance. However, they are accompanied by a large computation overhead cost, which significantly prevents their applications on resource-limited devices. Prior arts attempt to reduce the computational burden of DETR using low-bit quantization, while these methods sacrifice a severe significant performance on weight-activation-attention low-bit quantization. We observe that the number of matching queries and positive samples affect much on the representation capacity of queries in DETR, while quantifying queries of DETR further reduces its representational capacity, thus leading to a severe performance drop. We introduce a new quantization strategy based on Auxiliary Queries for DETR (AQ-DETR), aiming to enhance the capacity of quantized queries. In addition, a layer-by-layer distillation is proposed to reduce the quantization error between quantized attention and full-precision counterpart. Through our extensive experiments on large-scale open datasets, the performance of the 4-bit quantization of DETR and Deformable DETR models is comparable to full-precision counterparts. Runqi Wang, Huixin Sun, Linlin Yang 0001, Shaohui Lin, Chuanjian Liu, Yan Gao 0017, Yao Hu 0002, Baochang Zhang 0001 |
AAAI | 3 |
| 2024 | Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang 0001, Nie Lin, Zhishan Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Feng Lu 0005, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Saurabh Gupta 0001, Yoichi Sato 0001, Otmar Hilliges, Hyung Jin Chang, Angela Yao |
ECCV (25) | 3 |
| 2024 | CLIP in Mirror: Disentangling text from visual images through reflectionabstractThe CLIP network excels in various tasks, but struggles with text-visual images i.e., images that contain both text and visual objects; it risks confusing textual and visual representations. To address this issue, we propose MirrorCLIP, a zero-shot framework, which disentangles the image features of CLIP by exploiting the difference in the mirror effect between visual objects and text in the images. Specifically, MirrorCLIP takes both original and flipped images as inputs, comparing their features dimension-wise in the latent space to generate disentangling masks. With disentangling masks, we further design filters to separate textual and visual factors more precisely, and then get disentangled representations. Qualitative experiments using stable diffusion models and class activation mapping (CAM) validate the effectiveness of our disentanglement. Moreover, our proposed MirrorCLIP reduces confusion when encountering text-visual images and achieves a substantial improvement on typographic defense, further demonstrating its superior ability of disentanglement. Our code is available at https://github.com/tcwangbuaa/MirrorCLIP Yuguang Yang 0007, Linlin Yang 0001, Shaohui Lin, Guodong Guo, Baochang Zhang 0001 |
NeurIPS | 3 |
| 2024 | Rethinking Visibility in Human Pose Estimation: Occluded Pose Reasoning via TransformersabstractOcclusion is a common challenge in human pose estimation. Curiously, learning from occluded keypoints hinders a model to detect visible keypoints. We speculate that the impairment is likely due to a forced correlation between keypoints and visual features of the occluders. As such, we propose a novel visibility-aware attention mechanism to eliminate unreliable occluding features. The explicit occlusion handling encourages the model to reason about occluded keypoints using evidence and contextual information from the visible keypoints. It also mitigates the damage of unreliable correlations of the occluded keypoints. Our method, when added to the strong baseline SimCC, improves by 1.3 AP and 0.7 AP with ResNet and HRNet respectively. It also surpasses the state-of-the-art I2R-Net on CrowdPose by 0.3 AP and 0.6 APhard. The improvements highlight that rethinking visibility information is critical for developing effective human pose estimation systems. Pengzhan Sun 0001, Kerui Gu, Yunsong Wang, Linlin Yang 0001, Angela Yao |
WACV | 4 |
| 2023 | Analyzing and Diagnosing Pose Estimation with AttributionsabstractWe present Pose Integrated Gradient (PoseIG), the first interpretability technique designed for pose estimation. We extend the concept of integrated gradients for pose estimation to generate pixel-level attribution maps. To enable comparison across different pose frameworks, we unify different pose outputs into a common output space, along with a likelihood approximation function for gradient back-propagation. To complement the qualitative insight from the attribution maps, we propose three indices for quantitative analysis. With these tools, we systematically compare different pose estimation frameworks to understand the impacts of network design, backbone and auxiliary tasks. Our analysis reveals an interesting shortcut of the knuckles (MCP joints) for hand pose estimation and an under-explored inversion error for keypoints in body pose estimation. Project page and code: https://qy-h00.github.io/poseig/. Qiyuan He, Linlin Yang 0001, Kerui Gu, Qiuxia Lin, Angela Yao |
CVPR | 2 |
| 2023 | Cross-Domain 3D Hand Pose Estimation with Dual ModalitiesabstractRecent advances in hand pose estimation have shed light on utilizing synthetic data to train neural networks, which however inevitably hinders generalization to real-world data due to domain gaps. To solve this problem, we present a framework for cross-domain semi-supervised hand pose estimation and target the challenging scenario of learning models from labelled multimodal synthetic data and unlabelled real-world data. To that end, we propose a dual-modality network that exploits synthetic RGB and synthetic depth images. For pre-training, our network uses multi-modal contrastive learning and attention-fused supervision to learn effective representations of the RGB images. We then integrate a novel self-distillation technique during fine-tuning to reduce pseudo-label noise. Experiments show that the proposed method significantly improves 3D hand pose estimation and 2D keypoint detection on benchmarks. Qiuxia Lin, Linlin Yang 0001, Angela Yao |
CVPR | 2 |
| 2023 | Overcoming the TradeOff between Accuracy and Plausibility in 3D Hand Shape ReconstructionabstractDirect mesh fitting for 3D hand shape reconstruction is highly accurate. However, the reconstructed meshes are prone to artifacts and do not appear as plausible hand shapes. Conversely, parametric models like MANO ensure plausible hand shapes but are not as accurate as the non-parametric methods. In this work, we introduce a novel weakly-supervised hand shape estimation framework that integrates non-parametric mesh fitting with MANO model in an end-to-end fashion. Our joint model overcomes the tradeoff in accuracy and plausibility to yield well-aligned and high-quality 3D meshes, especially in challenging two-hand and hand-object interaction scenarios. Ziwei Yu, Chen Li 0038, Linlin Yang 0001, Xiaoxu Zheng, Michael Bi Mi, Gim Hee Lee, Angela Yao |
CVPR | 3 |
| 2023 | MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape RecoveryabstractFor monocular RGB-based 3D pose and shape estimation, multiple solutions are often feasible due to factors like occlusions and truncations. This work presents a multi-hypothesis probabilistic framework by optimizing the Kullback–Leibler divergence (KLD) between the data and model distribution. Our formulation reveals a connection between the pose entropy and diversity in the multiple hypotheses that has been neglected by previous works. For a comprehensive evaluation, besides the best hypothesis (BH) metric, we factor in visibility for evaluating diversity. Additionally, our framework is label-friendly – it can be learned from only partial 2D keypoints, such as visible keypoints. Experiments on both ambiguous and real-world benchmarks demonstrate that our method outperforms other state-of-the-art multi-hypothesis methods. The project page is at https://gloryyrolg.github.io/MHEntropy. Rongyu Chen, Linlin Yang 0001, Angela Yao |
ICCV | 2 |
| 2023 | Improving Deep Regression with Ordinal Entropy
Linlin Yang 0001, Michael Bi Mi, Xiaoxu Zheng, Angela Yao |
ICLR | 2 |
| 2023 | Synthetic-to-Real Pose Estimation with Geometric ReconstructionabstractPose estimation is remarkably successful under supervised learning, but obtaining annotations, especially for new deployments, is costly and time-consuming. This work tackles adapting models trained on synthetic data to real-world target domains with only unlabelled data. A common approach is model fine-tuning with pseudo-labels from the target domain; yet many pseudo-labelling strategies cannot provide sufficient high-quality pose labels. This work proposes a reconstruction-based strategy as a complement to pseudo-labelling for synthetic-to-real domain adaptation. We generate the driving image by geometrically transforming a base image according to the predicted keypoints and enforce a reconstruction loss to refine the predictions. It provides a novel solution to effectively correct confident yet inaccurate keypoint locations through image reconstruction in domain adaptation. Our approach outperforms the previous state-of-the-arts by 8% for PCK on four large-scale hand and human real-world datasets. In particular, we excel on endpoints such as fingertips and head, with 7.2% and 29.9% improvements in PCK. Qiuxia Lin, Kerui Gu, Linlin Yang 0001, Angela Yao |
NeurIPS | 3 |
| 2023 | Anti-Bandit for Neural Architecture Search
Runqi Wang, Linlin Yang 0001, Wei Wang 0016, David S. Doermann, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Bias-Compensated Integral Regression for Human Pose EstimationabstractIn human and hand pose estimation, heatmaps are a crucial intermediate representation for a body or hand keypoint. Two popular methods to decode the heatmap into a final joint coordinate are via an argmax, as done in heatmap detection, or via softmax and expectation, as done in integral regression. Integral regression is learnable end-to-end, but has lower accuracy than detection. This paper uncovers an induced bias from integral regression that results from combining the softmax and the expectation operation. This bias often forces the network to learn degenerately localized heatmaps, obscuring the keypoint's true underlying distribution and leads to lower accuracies. Training-wise, by investigating the gradients of integral regression, we show that the implicit guidance of integral regression to update the heatmap makes it slower to converge than detection. To counter the above two limitations, we propose Bias Compensated Integral Regression (BCIR), an integral regression-based framework that compensates for the bias. BCIR also incorporates a Gaussian prior loss to speed up training and improve prediction accuracy. Experimental results on both the human body and hand benchmarks show that BCIR is faster to train and more accurate than the original integral regression, making it competitive with state-of-the-art detection methods. Kerui Gu, Linlin Yang 0001, Michael Bi Mi, Angela Yao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | MagFormer: Hybrid Video Motion Magnification Transformer from Eulerian and Lagrangian Perspectives
Sicheng Gao, Yutang Feng, Linlin Yang 0001, Xuhui Liu, David S. Doermann, Baochang Zhang 0001 |
BMVC | 3 |
| 2022 | UV-Based 3D Hand-Object Reconstruction with Grasp Optimization
Ziwei Yu, Linlin Yang 0001, You Xie, Angela Yao |
BMVC | 2 |
| 2022 | Dive Deeper Into Integral Pose Regression
Kerui Gu, Linlin Yang 0001, Angela Yao |
ICLR | 2 |
| 2022 | Filter pruning via expectation-maximization
Sheng Xu 0007, Yanjing Li, Linlin Yang 0001, Baochang Zhang 0001, Dianmin Sun |
Neural Comput. Appl. | 3 |
| 2021 | Local and Global Point Cloud Reconstruction for 3D Hand Pose Estimation
Ziwei Yu, Linlin Yang 0001, Shicheng Chen, Angela Yao |
BMVC | 2 |
| 2021 | Removing the Bias of Integral Pose RegressionabstractHeatmap-based detection methods are dominant for 2D human pose estimation even though regression is more intuitive. The introduction of the integral regression method, which, architecture-wise uses an implicit heatmap, brings the two approaches even closer together. This begs the question – does detection really outperform regression? In this paper, we investigate the difference in supervision between the heatmap-based detection and integral regression, as this is the key remaining difference between the two approaches. In the process, we discover an underlying bias behind integral pose regression that arises from taking the expectation after the softmax function. To counter the bias, we present a compensation method which we find to improve integral regression accuracy on all 2D pose estimation benchmarks. We further propose a simple combined detection and bias-compensated regression method that considerably outperforms state-of-the-art baselines with few added components. Kerui Gu, Linlin Yang 0001, Angela Yao |
ICCV | 2 |
| 2021 | SemiHand: Semi-supervised Hand Pose Estimation with ConsistencyabstractWe present SemiHand, a semi-supervised framework for 3D hand pose estimation from monocular images. We pre-train the model on labelled synthetic data and fine-tune it on unlabelled real-world data by pseudo-labeling with consistency training. By design, we introduce data augmentation of differing difficulties, consistency regularizer, label correction and sample selection for RGB-based 3D hand pose estimation. In particular, by approximating the hand masks from hand poses, we propose a cross-modal consistency and leverage semantic predictions to guide the predicted poses. Meanwhile, we introduce pose registration as label correction to guarantee the biomechanical feasibility of hand bone lengths. Experiments show that our method achieves a favorable improvement on real-world datasets after fine-tuning. Linlin Yang 0001, Shicheng Chen, Angela Yao |
ICCV | 1 |
| 2021 | Uncertainty-aware Binary Neural NetworksabstractBinary Neural Networks (BNN) are promising machine learning solutions for deployment on resource-limited devices. Recent approaches to training BNNs have produced impressive results, but minimizing the drop in accuracy from full precision networks is still challenging. One reason is that conventional BNNs ignore the uncertainty caused by weights that are near zero, resulting in the instability or frequent flip while learning. In this work, we investigate the intrinsic uncertainty of vanishing near-zero weights, making the training vulnerable to instability. We introduce an uncertainty-aware BNN (UaBNN) by leveraging a new mapping function called certainty-sign (c-sign) to reduce these weights' uncertainties. Our c-sign function is the first to train BNNs with a decreasing uncertainty for binarization. The approach leads to a controlled learning process for BNNs. We also introduce a simple but effective method to measure the uncertainty-based on a Gaussian function. Extensive experiments demonstrate that our method improves multiple BNN methods by maintaining stability of training, and achieves a higher performance over prior arts. Junhe Zhao, Linlin Yang 0001, Baochang Zhang 0001, Guodong Guo, David S. Doermann |
IJCAI | 2 |
| 2020 | Cogradient Descent for Bilinear OptimizationabstractConventional learning methods simplify the bilinear model by regarding two intrinsically coupled factors independently, which degrades the optimization procedure. One reason lies in the insufficient training due to the asynchronous gradient descent, which results in vanishing gradients for the coupled variables. In this paper, we introduce a Cogradient Descent algorithm (CoGD) to address the bilinear problem, based on a theoretical framework to coordinate the gradient of hidden variables via a projection function. We solve one variable by considering its coupling relationship with the other, leading to a synchronous gradient descent to facilitate the optimization procedure. Our algorithm is applied to solve problems with one variable under the sparsity constraint, which is widely used in the learning paradigm. We validate our CoGD considering an extensive set of applications including image reconstruction, inpainting, and network pruning. Experiments show that it improves the state-of-the-art by a significant margin. Lian Zhuo, Baochang Zhang 0001, Linlin Yang 0001, Qixiang Ye, David S. Doermann, Rongrong Ji, Guodong Guo |
CVPR | 3 |
| 2020 | Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction
Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Shipeng Xie, Mingxiu Chen, Boshen Zhang, Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001, Pengfei Ren 0001, Weiting Huang, Haifeng Sun 0001, Marek Hrúz, Jakub Kanis, Zdenek Krnoul, Qingfu Wan, Shile Li, Linlin Yang 0001, Dongheui Lee, Angela Yao, Weiguo Zhou, Sijia Mei, Adrian Spurr, Umar Iqbal 0001, Pavlo Molchanov 0001, Philippe Weinzaepfel, Romain Brégier, Grégory Rogez, Vincent Lepetit, Tae-Kyun Kim 0001 |
ECCV (23) | 22 |
| 2020 | CP-NAS: Child-Parent Neural Architecture Search for 1-bit CNNsabstractNeural architecture search (NAS) proves to be among the best approaches for many tasks by generating an application-adaptive neural architectures, which are still challenged by high computational cost and memory consumption. At the same time, 1-bit convolutional neural networks (CNNs) with binarized weights and activations show their potential for resource-limited embedded devices. One natural approach is to use 1-bit CNNs to reduce the computation and memory cost of NAS by taking advantage of the strengths of each in a unified framework. To this end, a Child-Parent model is introduced to a differentiable NAS to search the binarized architecture(Child) under the supervision of a full-precision model (Parent). In the search stage, the Child-Parent model uses an indicator generated by the parent and child model accuracy to evaluate the performance and abandon operations with less potential. In the training stage, a kernel level CP loss is introduced to optimize the binarized network. Extensive experiments demonstrate that the proposed CP-NAS achieves a comparable accuracy with traditional NAS on both the CIFAR and ImageNet databases. It achieves an accuracy of 95.27% on CIFAR-10, 64.3% on ImageNet with binarized weights and activations, and a 30% faster search than prior arts. Lian Zhuo, Baochang Zhang 0001, Linlin Yang 0001, Chen Chen 0001, Yanjun Zhu, David S. Doermann |
IJCAI | 4 |
| 2019 | Disentangling Latent Hands for Image Synthesis and Pose EstimationabstractHand image synthesis and pose estimation from RGB images are both highly challenging tasks due to the large discrepancy between factors of variation ranging from image background content to camera viewpoint. To better analyze these factors of variation, we propose the use of disentangled representations and a disentangled variational autoencoder (dVAE) that allows for specific sampling and inference of these factors. The derived objective from the variational lower bound as well as the proposed training strategy are highly flexible, allowing us to handle cross-modal encoders and decoders as well as semi-supervised learning scenarios. Experiments show that our dVAE can synthesize highly realistic images of the hand specifiable by both pose and image background content and also estimate 3D hand poses from RGB images with accuracy competitive with state-of-the-art on two public benchmarks. Linlin Yang 0001, Angela Yao |
CVPR | 1 |
| 2019 | Aligning Latent Spaces for 3D Hand Pose EstimationabstractHand pose estimation from monocular RGB inputs is a highly challenging task. Many previous works for monocular settings only used RGB information for training despite the availability of corresponding data in other modalities such as depth maps. In this work, we propose to learn a joint latent representation that leverages other modalities as weak labels to boost the RGB-based hand pose estimator. By design, our architecture is highly flexible in embedding various diverse modalities such as heat maps, depth maps and point clouds. In particular, we find that encoding and decoding the point cloud of the hand surface can improve the quality of the joint latent representation. Experiments show that with the aid of other modalities during training, our proposed method boosts the accuracy of RGB-based hand pose estimation systems and significantly outperforms state-of-the-art on two public benchmarks. Linlin Yang 0001, Shile Li, Dongheui Lee, Angela Yao |
ICCV | 1 |
| 2017 | Action Recognition Using 3D Histograms of Texture and A Multi-Class Boosting ClassifierabstractHuman action recognition is an important yet challenging task. This paper presents a low-cost descriptor called 3D histograms of texture (3DHoTs) to extract discriminant features from a sequence of depth maps. 3DHoTs are derived from projecting depth frames onto three orthogonal Cartesian planes, i.e., the frontal, side, and top planes, and thus compactly characterize the salient information of a specific action, on which texture features are calculated to represent the action. Besides this fast feature descriptor, a new multi-class boosting classifier (MBC) is also proposed to efficiently exploit different kinds of features in a unified framework for action classification. Compared with the existing boosting frameworks, we add a new multi-class constraint into the objective function, which helps to maintain a better margin distribution by maximizing the mean of margin, whereas still minimizing the variance of margin. Experiments on the MSRAction3D, MSRGesture3D, MSRActivity3D, and UTD-MHAD data sets demonstrate that the proposed system combining 3DHoTs and MBC is superior to the state of the art. Baochang Zhang 0001, Chen Chen 0001, Linlin Yang 0001, Jungong Han, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |