VLDB 2026 Research / reviewers in the wild / expert
Jun Li 0027
dblp:116/1011-27
· DBLP profile ↗
116ranked-venue papers
19as first author
84since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 80 · 16 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 11 first-author · 46 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RMLer: Synthesizing Novel Objects Across Diverse Categories via Reinforcement Mixing LearningabstractNovel object synthesis by integrating distinct textual concepts from diverse categories remains a significant challenge in text-to-image generation. Existing methods often suffer from insufficient concept mixing, lack of rigorous evaluation, and suboptimal outputs, resulting in conceptual imbalance, superficial combinations, or mere juxtapositions. To address these limitations, we propose Reinforcement Mixing Learning (RMLer), a framework that formulates cross-category concept fusion as a reinforcement learning problem: mixed features serve as states, mixing strategies as actions, and visual outcomes as rewards. Specifically, we design an MLP policy network to predict dynamic coefficients for blending cross-category text embeddings. We further introduce visual rewards based on (1) semantic similarity and (2) compositional balance between the fused object and its constituent concepts, and optimize the policy via proximal policy optimization. At inference time, a selection strategy leverages these rewards to curate the highest-quality fused objects. Extensive experiments demonstrate that RMLer synthesizes coherent, high-fidelity objects from diverse categories and consistently outperforms existing methods. Our work provides a robust framework for generating novel visual concepts, with promising applications in film, gaming, and design. Jun Li 0027, Haibo Chen 0006, Shuo Chen 0003, Jian Yang 0003 |
AAAI | 1 |
| 2026 | Curriculum adaptation for one-stream RGB-T tracking
Xiantao Hu, Fansheng Zeng, Bineng Zhong 0001, Zhangyong Tang, Wenxuan Fang 0001, Jun Li 0027, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 6 |
| 2026 | ConceptCraft: One-Shot Personalized Text-to-Image Generation via Object-Background DisentanglementabstractPersonalized text-to-image generation aims to learn new concepts from user-provided images and subsequently generate diverse scenes or styles of the concepts from input prompts. Most existing methods usually require a set of images (typically 3-5) for each concept, which can be cumbersome. Although several methods allow personalized generation with a single reference image, they often require heavy model training and suffer from many issues such as domain-specific applicability, insufficient fidelity, and limited editability. To address these problems, we propose a novel one-shot personalized text-to-image generation method called ConceptCraft, which explicitly separates the reference image into object and background regions and treats them as two distinct concepts to learn, significantly improving the personalization performance. Specifically, we incorporate two unique identifiers into the text prompts: one followed by the object’s class name and the other by the word “background”. To bind these two identifiers to the reference image’s object and background respectively, we introduce a mask-aware object preservation loss and a mask-aware background preservation loss to optimize their corresponding token embeddings under well-designed text conditions, enabling both object and background personalization. In addition, we also develop an identifier regularization scheme to enhance our editability, allowing the synthesis of personalized images across a broader range of scenes and styles without changing the identity. Extensive qualitative and quantitative experiments are conducted to verify the effectiveness and superiority of our method. Haibo Chen 0006, Zhiwen Zuo, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Physics-Guided Posterior Sampling for Diffusion-Based Real-World Dehazing and Image EnhancementabstractReal-world image dehazing is a challenging task due to the collection of aligned hazy/clear image pairs under unpredictable and complex environments. To address this limitation, we propose a Physical-Guided Posterior Sampling (PGPS) method that designs a dehazing reconstruction posterior to sample an RGB and depth from pre-trained unconditional diffusion generation process. First, we introduce a Hybrid Degradation Atmospheric Scattering Model (HD-ASM) to adapt the diffusion model, enabling the generation of high-fidelity dehazed images from posterior samples without relying on the aligned hazy/clear image pairs. Second, we propose a two-stage sampling strategy with piecewise loss to improve sampling quality and stability, along with a post-processing technique to remove JPEG compression artifacts amplified by dehazing. Extensive experiments show that our method outperforms state-of-the-art techniques in image dehazing, and in the RTTS dataset’s complex human-vehicle environment. Additionally, our approach also surpasses other benchmarks in object detection, exhibiting superior generalization performance. Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Jianjun Qian, Heyou Chang, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | WeatherCycle: Unpaired Multi-Weather Restoration via Color Space Decoupled Cycle LearningabstractUnsupervised image restoration under multi-weather conditions remains a fundamental yet underexplored challenge. While existing methods often rely on task-specific physical priors, their narrow focus limits scalability and generalization to diverse real-world weather scenarios. In this work, we propose WeatherCycle, a unified unpaired framework that reformulates weather restoration as a bidirectional degradation-content translation cycle, guided by degradation-aware curriculum regularization. At its core, WeatherCycle employs alumina-chroma decompositionstrategy to decouple degradation from content without modeling complex weather, enabling domain conversion between degraded and clean images. To model diverse and complex degradations, we propose aLumina Degradation Guidance Module(LDGM), which learns luminance degradation priors from a degraded image pool and injects them into clean images via frequency-domain amplitude modulation, enabling controllable and realistic degradation modeling. Additionally, we incorporate aDifficulty-Aware Contrastive Regularization(DACR) module that identifies hard samples via a CLIP-based classifier and enforces contrastive alignment between hard samples and restored features to enhance semantic consistency and robustness. Extensive experiments across serve multi-weather datasets, demonstrate that our method achieves state-of-the-art performance among unsupervised approaches, with strong generalization to complex weather degradations. Wenxuan Fang 0001, Jiangwei Weng, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | TDHF: Task-Driven Hierarchical Image Fusion Under Low-Light ConditionsabstractVision analysis tasks often experience substantial performance drops when processing images captured in lowlight environments. Existing approaches introduce infrared images to provide complementary information, and fusion strategies are consequently developed to combine the advantages of visible light images and infrared images. While fusion methods are typically designed to enhance visual quality with the expectation of improving task performance, many existing approaches tend to overemphasize perceptual fidelity, which can inadvertently compromise task-specific feature. To address this issue, we propose a task-driven hierarchical fusion (TDHF) framework designed to retain task-relevant information throughout the fusion process. Specifically, TDHF adopts a multi-scale hierarchical architecture to capture rich feature representations and incorporates a multi-head attention mechanism to model cross-modal interactions. In addition, we introduce a single-step denoising generation module that guides the fusion of infrared edge features and texture details from low-light images progressively. This process ultimately reconstructs the task-critical Y channel in the YCrCb color space. Extensive experiments on object detection under low-light conditions across four benchmark datasets demonstrate that TDHF effectively enhances task performance, achieving up to 1.1% higher mAP on LLVIP and 1.3% higher mAP on FLIR compared with state-of-the-art methods, without relying on excessive optimization of image quality (PSNR). Daoheng Li, Mengkai Yan, Jun Li 0027, Jian Yang 0003, Jianjun Qian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DVDPEC: Driving-Video Dehazing via Position Embedding-Based CodebookabstractDespite significant progress in real-world image dehazing, efficiently generating high-fidelity, haze-free videos (especially in driving scenarios) remains challenging. Existing methods generally extend image dehazing techniques to videos by employing pre-trained single image dehazing models for preprocessing followed by refinement stages. However, this disjointed two-stage process often leads to unrealistic textures and loss of detail, as it fails to leverage large amounts of high-quality images for prior learning and the subsequent refinement struggles to correct temporal inconsistencies across frames introduced in the first stage. To address these issues, we propose DVDPEC: a Driving Video Dehazing framework utilizing a Position Embedding-based (PE-based) Codebook and a novel Flow Selective Block (FSB). The PE-based codebook stores fine-grained, spatially aware textural information specific to driving videos and leverages implicit positional embeddings for precise, position-aware codebook matching. This enables accurate prior retrieval and improves dehazing results. The FSB aggregates information from adjacent frames by dynamically combining both image flow and prior flow, effectively mitigating flow estimation ambiguities caused by haze. It enhances information fusion across frames, leading to more coherent and visually appealing dehazed videos. Extensive experiments demonstrate that DVDPEC achieves state-of-the-art performance on real-world driving video dehazing tasks, significantly enhancing texture preservation and visual fidelity. Yu Zheng 0036, Wenxuan Fang 0001, Xiantao Hu, Junkai Fan, Jiangwei Weng, Jun Li 0027, Kai Zhang 0008, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | PanoKernel: Large Distortion-Aware Kernel for Panoramic Depth Perception
Zhiqiang Yan 0001, Zhijie Shen, Jiayi Yuan 0003, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose EstimationabstractWith the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information, multi-modal fusion, or SMPL optimization to correct biased results. In this work, we try to obtain sufficient information for 3D HPE only by modeling the intrinsic properties of low-quality point clouds. Hence, a simple yet powerful method is proposed, which provides insights both on modeling and augmentation of point clouds. Specifically, we first propose a concise and effective density-aware pose transformer (DAPT) to get stable keypoint representations. By using a set of joint anchors and a carefully designed exchange module, valid information is extracted from point clouds with different densities. Then 1D heatmaps are utilized to represent the precise locations of the keypoints. Secondly, a comprehensive LiDAR human synthesis and augmentation method is proposed to pre-train the model, enabling it to acquire a better human body prior. We increase the diversity of point clouds by randomly sampling human positions and orientations and by simulating occlusions through the addition of laser-level masks. Extensive experiments have been conducted on multiple datasets, including IMU-annotated LidarHuman26M, SLOPER4D, and manually annotated Waymo Open Dataset v2.0 (Waymo), HumanM3. Our method demonstrates SOTA performance in all scenarios. In particular, compared with LPFormer on Waymo, we reduce the average MPJPE by 10.0mm. Compared with PRN on SLOPER4D, we notably reduce the average MPJPE by 20.7mm. Xiaoqi An, Lin Zhao 0003, Chen Gong 0002, Jun Li 0027, Jian Yang 0003 |
AAAI | 4 |
| 2025 | Harmonious Music-driven Group Choreography with Trajectory-Controllable DiffusionabstractCreating group choreography from music is crucial in cultural entertainment and virtual reality, with a focus on generating harmonious movements. Despite growing interest, recent approaches often struggle with two major challenges: multi-dancer collisions and single-dancer foot sliding. To address these challenges, we propose a Trajectory-Controllable Diffusion (TCDiff) framework, which leverages non-overlapping trajectories to ensure coherent and aesthetically pleasing dance movements. To mitigate collisions, we introduce a Dance-Trajectory Navigator that generates collision-free trajectories for multiple dancers, utilizing a distance-consistency loss to maintain optimal spacing. Furthermore, to reduce foot sliding, we present a footwork adaptor that adjusts trajectory displacement between frames, supported by a relative forward-kinematic loss to further reinforce the correlation between movements and trajectories. Experiments demonstrate our method's superiority. Yuqin Dai, Wanlu Zhu, Ronghui Li, Zeping Ren, Xiangzheng Zhou, Jixuan Ying, Jun Li 0027, Jian Yang 0003 |
AAAI | 7 |
| 2025 | Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving VideoabstractIn this paper, we study the challenging problem of simultaneously removing haze and estimating depth from real monocular hazy videos. These tasks are inherently complementary: enhanced depth estimation improves dehazing via the atmospheric scattering model (ASM), while superior dehazing contributes to more accurate depth estimation through the brightness consistency constraint (BCC). To tackle these intertwined tasks, we propose a novel depth-centric learning framework that integrates the ASM model with the BCC constraint. Our key idea is that both ASM and BCC rely on a shared depth estimation network. This network simultaneously exploits adjacent dehazed frames to enhance depth estimation via BCC and uses the refined depth cues to more effectively remove haze through ASM. Additionally, we leverage a non-aligned clear video and its estimated depth to independently regularize the dehazing and depth estimation networks. This is achieved by designing two discriminator networks: D_MFIR enhances high-frequency details in dehazed videos, and D_MDR reduces the occurrence of black holes in low-texture regions. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art techniques in both video dehazing and depth estimation tasks, especially in real-world hazy scenes. Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Xiang Chen 0015, Shangbing Gao, Jun Li 0027, Jian Yang 0003 |
AAAI | 6 |
| 2025 | Guided Real Image Dehazing Using YCbCr Color SpaceabstractImage dehazing, particularly with learning-based methods, has gained significant attention due to its importance in real-world applications. However, relying solely on the RGB color space often fall short, frequently leaving residual haze. This arises from two main issues: the difficulty in obtaining clear textural features from hazy RGB images and the complexity of acquiring real haze/clean image pairs outside controlled environments like smoke-filled scenes. To address these issues, we first propose a novel Structure Guided Dehazing Network (SGDN) that leverages the superior structural properties of YCbCr features over RGB. It comprises two key modules: Bi-Color Guidance Bridge (BGB) and Color Enhancement Module (CEM). BGB integrates a phase integration module and an interactive attention module, utilizing the rich texture features of the YCbCr space to guide the RGB space, thereby recovering clearer features in both frequency and spatial domains. To maintain tonal consistency, CEM further enhances the color perception of RGB features by aggregating YCbCr channel information. Furthermore, for effective supervised learning, we introduce a Real-World Well-Aligned Haze dataset, which includes a diverse range of scenes from various geographical regions and climate conditions. Experimental results demonstrate that our method surpasses existing state-of-the-art methods across multiple real-world smoke/haze datasets. Wenxuan Fang 0001, Junkai Fan, Yu Zheng 0036, Jiangwei Weng, Ying Tai, Jun Li 0027 |
AAAI | 6 |
| 2025 | Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingabstractMultimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003 |
AAAI | 6 |
| 2025 | Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth CompletionabstractIn this paper, we introduce the Selective Image Guided Network (SigNet), a novel degradation-aware framework that transforms depth completion into depth enhancement for the first time. Moving beyond direct completion using convolutional neural networks (CNNs), SigNet initially densifies sparse depth data through non-CNN densification tools to obtain coarse yet dense depth. This approach eliminates the mismatch and ambiguity caused by direct convolution over irregularly sampled sparse data. Subsequently, SigNet redefines completion as enhancement, establishing a self-supervised degradation bridge between the coarse depth and the targeted dense depth for effective RGB-D fusion. To achieve this, SigNet leverages the implicit degradation to adaptively select high-frequency components (e.g., edges) of RGB data to compensate for the coarse depth. This degradation is further integrated into a multi-modal conditional Mamba, dynamically generating the state parameters to enable efficient global high-frequency information interaction. We conduct extensive experiments on the NYUv2, DIML, SUN RGBD, and TOFDC datasets, demonstrating the state-of-the-art (SOTA) performance of SigNet. Zhiqiang Yan 0001, Zhengxue Wang, Kun Wang 0042, Jun Li 0027, Jian Yang 0003 |
CVPR | 4 |
| 2025 | Volume-Aware Distance for Robust Similarity LearningabstractMeasuring the similarity between data points plays a vital role in lots of popular representation learning tasks such as metric learning and contrastive learning. Most existing approaches utilize point-level distances to learn the point-to-point similarity between pairwise instances. However, since the finite number of training data points cannot fully cover the whole sample space consisting of an infinite number of points, the generalizability of the learned distance is usually limited by the sample size. In this paper, we thus extend the conventional form of data point to the new form of data ball with a predictable volume, so that we can naturally generalize the existing point-level distance to a new volume-aware distance (VAD) which measures the field-to-field geometric similarity. The learned VAD not only takes into account the relationship between observed instances but also uncovers the similarity among those unsampled neighbors surrounding the training data. This practice significantly enriches the coverage of sample space and thus improves the model generalizability. Theoretically, we prove that VAD tightens the error bound of traditional similarity learning and preserves crucial topological properties. Experiments on multi-domain data demonstrate the superiority of VAD over existing approaches in both supervised and unsupervised tasks. Shuo Chen 0003, Chen Gong 0002, Jun Li 0027, Jian Yang 0003 |
ICML | 3 |
| 2025 | Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
Zeren Xiong, Zedong Zhang, Xiang Li 0041, Ying Tai, Jian Yang 0003, Jun Li 0027 |
ACM Multimedia | 7 |
| 2025 | Enhancing Contrastive Learning with Variable SimilarityabstractContrastive learning has achieved remarkable success in self-supervised learning by pretraining a generalizable feature representation based on the augmentation invariance. Most existing approaches assume that different augmented views of the same instance (i.e., the *positive pairs*) remain semantically invariant. However, the augmentation results with *varying extent* may introduce semantic discrepancies or even content distortion, and thus the conventional (pseudo) supervision from augmentation invariance may lead to misguided learning objectives. In this paper, we propose a novel method called Contrastive Learning with Variable Similarity (CLVS) to accurately characterize the intrinsic similarity relationships between different augmented views. Our method dynamically adjusts the similarity based on the augmentation extent, and it ensures that strongly augmented views are always assigned lower similarity scores than weakly augmented ones. We provide a theoretical analysis to guarantee the effectiveness of the variable similarity in improving model generalizability. Extensive experiments demonstrate the superiority of our approach, achieving gains of 2.1\% on ImageNet-100 and 1.4\% on ImageNet-1k compared with the state-of-the-art methods. Haowen Cui, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003 |
NeurIPS | 3 |
| 2025 | AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group SwappingabstractFusing cross-category objects to a single coherent object has gained increasing attention in text-to-image (T2I) generation due to its broad applications in virtual reality, digital media, film, and gaming. However, existing methods often produce biased, visually chaotic, or semantically inconsistent results due to overlapping artifacts and poor integration. Moreover, progress in this field has been limited by the absence of a comprehensive benchmark dataset. To address these problems, we propose Adaptive Group Swapping (AGSwap), a simple yet highly effective approach comprising two key components: (1) Group-wise Embedding Swapping, which fuses semantic attributes from different concepts through feature manipulation, and (2) Adaptive Group Updating, a dynamic optimization mechanism guided by a balance evaluation score to ensure coherent synthesis. Additionally, we introduce Cross-category Object Fusion (COF), a large-scale, hierarchically structured dataset built upon ImageNet-1K and WordNet. COF includes 95 superclasses, each with 10 subclasses, enabling 451,250 unique fusion pairs. Extensive experiments demonstrate that AGSwap outperforms state-of-the-art compositional T2I methods, including GPT-Image-1 using simple and complex prompts. Project Page Zedong Zhang, Ying Tai, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
SIGGRAPH Asia | 5 |
| 2025 | RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Le Hui, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
Int. J. Comput. Vis. | 5 |
| 2025 | PaintDiffusion: Towards text-driven painting variation via collaborative diffusion guidance
Haibo Chen 0006, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
Neurocomputing | 4 |
| 2025 | RagNet3D: Learning distinguishable representation for pooled grids in 3D object detection
Jiaxin Chen 0001, Yuehui Han, Zhiqiang Yan 0001, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
Neurocomputing | 5 |
| 2025 | Creative style transfer for image stylization via learning neural permutation
Zedong Zhang, Gan Sun, Li-Wei H. Lehman, Jian Yang 0003, Jun Li 0027 |
Knowl. Based Syst. | 6 |
| 2025 | Tri-Perspective View Decomposition for Geometry Aware Depth Completion and Super-ResolutionabstractDepth completion and super-resolution are crucial tasks for comprehensive RGB-D scene understanding, as they involve reconstructing the precise 3D geometry of a scene from sparse or low-resolution depth measurements. However, most existing methods either rely solely on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. In this paper, we introduce Tri-Perspective View Decomposition (TPVD) frameworks that can explicitly model 3D geometry. To this end, (1) TPVD ingeniously decomposes the original 3D point cloud into three 2D views, one of which corresponds to the sparse or low-resolution depth input. (2) For sufficient geometric interaction, TPV Fusion is designed to update the 2D TPV features through recurrent 2D-3D-2D aggregation. (3) By adaptively searching for TPV affinitive neighbors, two additional refinement heads are developed for these two tasks to further improve the geometric consistency. Meanwhile, we build novel datasets named TOFDC for depth completion and TOFDSR for depth super-resolution. Both datasets are acquired using time-of-flight (TOF) sensors and color cameras on smartphones. Extensive experiments on TOFDC, KITTI, NYUv2, SUN RGBD, VKITTI, TOFDSR, RGB-D-D, Lu, and Middlebury datasets indicate that our TPVD outperforms previous depth completion and super-resolution methods, reaching the state of the art. Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Guangwei Gao, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Non-Aligned Supervision for Real Image DehazingabstractRemoving haze from real-world images is challenging due to unpredictable weather conditions, resulting in the misalignment of hazy and clear image pairs. In this paper, we propose an innovative dehazing framework that operates under non-aligned supervision. This framework is grounded in the atmospheric scattering model, and consists of three interconnected networks: dehazing, airlight, and transmission networks. In particular, we explore a non-alignment scenario that a clear reference image, unaligned with the input hazy image, is utilized to supervise the dehazing network. To implement this, we present a multi-scale reference loss that compares the feature representations between the referred image and the dehazed output. Our scenario makes it easier to collect hazy/clear image pairs in real-world environments, even under conditions of misalignment and shift views. To showcase the effectiveness of our scenario, we have collected a new hazy dataset including 415 image pairs captured by mobile Phone in both rural and urban areas, called "Phone-Hazy". Furthermore, we introduce a self-attention network based on mean and variance for modeling real infinite airlight, using the dark channel prior as positional guidance. Experimental results demonstrate the superior performance of our framework over existing state-of-the-art techniques in the real-world image dehazing task. Phone-Hazy and code will be available at https://fanjunkai1.github.io/projectpage/NSDNet/index.html. Junkai Fan, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | TRTST: Arbitrary High-Quality Text-Guided Style Transfer With TransformersabstractText-guided style transfer aims to repaint a content image with the target style described by a text prompt, offering greater flexibility and creativity compared to traditional image-guided style transfer. Despite the potential, existing text-guided style transfer methods often suffer from many issues, including insufficient visual quality, poor generalization ability, or a reliance on large amounts of paired training data. To address these limitations, we leverage the inherent strengths of transformers in handling multimodal data and propose a novel transformer-based framework called TRTST that not only achieves unpaired arbitrary text-guided style transfer but also significantly improves the visual quality. Specifically, TRTST explores combining a text transformer encoder with an image transformer encoder to project the input text prompt and content image into a joint embedding space and extract the desired style and content features. These features are then input into a multimodal co-attention module to stylize the image sequence based on the text sequence. We also propose a new adaptive parametric positional encoding (APPE) scheme which can adaptively produce different positional encodings to optimally match different inputs with a position encoder. In addition, to further improve content preservation, we introduce a text-guided identity loss to our model. Extensive results and comparisons are conducted to demonstrate the effectiveness and superiority of our method. Haibo Chen 0006, Zhoujie Wang, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Image Process. | 4 |
| 2025 | Daytime-Mixed Non-Aligned Learning for Real Nighttime Image EnhancementabstractEnhancing real nighttime images is a significant challenge due to the deterioration of visual quality caused by limited perceptibility under adverse illumination conditions, leading to loss of details and color deviation. In this paper, we propose a novel nighttime image enhancement framework using daytime-mixed non-aligned supervision. It aims to couple the information between non-aligned daytime and nighttime image pairs. Specifically, our framework consists of a simple yet effective daytime-mixed supervised learning phase and a Retinex-based reconstruction phase. In the first phase, we employ a multi-instance with adaptive information fusion (AIF) module integrated within a UNet enhancement network called MIFUNet, which is trained via a daytime-mixed supervised loss. In the second phase, the Retinex-based reconstruction employs both a light-effect estimation network and an illumination adjustment network to restore the nighttime image, guided by physical principles. To evaluate the effectiveness of our approach, we collect a real non-aligned day-night dataset named the NANE dataset, which contains 748 non-aligned image pairs and 100 nighttime images solely for testing. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art image enhancement methods. Jiangwei Weng, Junkai Fan, Jianjun Qian, Haiyang Zou, Ying Tai, Jian Yang 0003, Jun Li 0027 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Asymptotics-Aware Multi-View Subspace ClusteringabstractRecently, multi-view subspace clustering has attracted extensive attention due to the rapid increase of multi-view data in many real-world applications. The main goal of this task is to learn a common representation of multiple subspaces from the given multi-view data, and most existing methods usually directly merge multiple groups of features by the single-step integration. However, there may exist large disparities among different views of the data, and thus the conventional single-step practice can hardly obtain a generally consistent feature representation for the multi-view data. To overcome this challenge, we present a novel approach dubbed “Asymptotics-Aware Multi-view Subspace Clustering (A$^{2}$MSC)” to pursue a consistent feature representation in a multi-step way, which iteratively conducts the data recovery to gradually reduce the differences between pairwise views. Specifically, we construct an asymptotic learning rule to update the feature representation, and the iteration result converges to a consistent feature vector for characterizing each instance of the original multi-view data. After that, we utilize such a new feature representation to learn a clustering-oriented similarity matrix via minimizing a self-expressive objective, and we also design the corresponding optimization algorithm to solve it with convergence guarantees. Theoretically, we prove that the learned asymptotic representation effectively integrates multiple views, thereby ensuring the effective handling of multi-view data. Empirically, extensive experimental results demonstrate the superiority of our proposed A$^{2}$MSC over the state-of-the-art multi-view subspace clustering approaches. Yesong Xu, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Multim. | 3 |
| 2025 | Similarity-Agnostic Contrastive Learning With Alterable Self-SupervisionabstractSelf-supervised contrastive learning (CL) seeks to learn generalizable feature representations via the self-supervision of pairwise similarities, where existing CL approaches usually build definite similarity labels (e.g., positive or negative) for model training. Yet in practice, the same pair of instances may have opposite similarity labels in different scenarios, e.g., two interclass images from CIFAR-100 can be a similar pair in CIFAR-20. Learning with definite similarities can hardly obtain an ideal representation that simultaneously characterizes the similar and dissimilar patterns (e.g., the contexts and details) between each two instances. Therefore, pairwise similarities used for CL should be agnostic, and we argue that simultaneously considering both the similarity and dissimilarity for each data pair could learn more generalizable representations. To this end, we propose similarity-agnostic CL (SACL), which generalizes the instance discrimination strategy of conventional CL to a new multiobjective programming (MOP) form. In SACL, we build multiple projection layers with corresponding regularizers to constrain the distance matrix to have different sparsity in different objectives so that we can obtain alterable pairwise distances to capture both the similarity and dissimilarity between each pair of instances. We show that SACL can be equivalently converted to a single learning objective, easily solved by stochastic optimization with convergence guarantees. Theoretically, we prove a tighter error bound than conventional CL approaches; empirically, our method improves the downstream task performance for image, text, and graph data. Shuo Chen 0003, Chen Gong 0002, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Metric Learning-Based Subspace ClusteringabstractThe self-expressive strategy has shown excellent capabilities in realizing low-dimensional representations of high-dimensional data for subspace clustering algorithms. The existing designs, however, are formulated on the linearization assumptions of the data, neglecting the precise characterization of linear relationships within samples. Considering that real-world data adheres to diverse distribution forms, it becomes impractical to first treat the samples as existing in a uniform linear space before finding an appropriate manifold space. To handle this challenge, we propose a novel self-expressive-based learning framework termed metric learning-based subspace clustering (MLSC). Particularly, we smoothly incorporate metric learning into the subspace clustering framework by introducing adaptive neighbors learning and defining a linearity-aware distance to discover the linear manifold space of the original data. We simultaneously utilize the generated representation of the linear structure as input for self-expressiveness to pursue an ideal similarity matrix, which establishes an essential connection with the linearization assumption of the self-expressive strategy. Furthermore, we theoretically demonstrate that our measure can accurately describe the level of linear correlation between instances. Finally, our tests demonstrate that the proposed MLSC attains competitive clustering results compared to state-of-the-art approaches on benchmark datasets. Yesong Xu, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose OptimizationabstractNeural Radiance Fields (NeRF) have shown promise in generating realistic novel views from sparse scene images. However, existing NeRF approaches often encounter challenges due to the lack of explicit 3D supervision and imprecise camera poses, resulting in suboptimal outcomes. To tackle these issues, we propose AltNeRF---a novel framework designed to create resilient NeRF representations using self-supervised monocular depth estimation (SMDE) from monocular videos, without relying on known camera poses. SMDE in AltNeRF masterfully learns depth and pose priors to regulate NeRF training. The depth prior enriches NeRF's capacity for precise scene geometry depiction, while the pose prior provides a robust starting point for subsequent pose refinement. Moreover, we introduce an alternating algorithm that harmoniously melds NeRF outputs into SMDE through a consistence-driven mechanism, thus enhancing the integrity of depth priors. This alternation empowers AltNeRF to progressively refine NeRF representations, yielding the synthesis of realistic novel views. Extensive experiments showcase the compelling capabilities of AltNeRF in generating high-fidelity and robust novel views that closely resemble reality. Kun Wang 0042, Zhiqiang Yan 0001, Huang Tian, Zhenyu Zhang 0005, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
AAAI | 6 |
| 2024 | Driving-Video Dehazing with Non-Aligned Regularization for Safety AssistanceabstractReal driving-video dehazing poses a significant challenge due to the inherent difficulty in acquiring precisely aligned hazy/clear video pairs for effective model training, especially in dynamic driving scenarios with unpredictable weather conditions. In this paper, we propose a pioneering approach that addresses this challenge through a nonaligned regularization strategy. Our core concept involves identifying clear frames that closely match hazy frames, serving as references to supervise a video dehazing network. Our approach comprises two key components: reference matching and video dehazing. Firstly, we introduce a non-aligned reference frame matching module, leveraging an adaptive sliding window to match high-quality reference frames from clear videos. Video dehazing incorporates flow-guided cosine attention sampler and deformable cosine attention fusion modules to enhance spatial multi-frame alignment and fuse their improved information. To validate our approach, we collect a GoProHazy dataset captured effortlessly with GoPro cameras in diverse rural and urban road environments. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art methods in the challenging task of real driving-video dehazing. Project page. Junkai Fan, Jiangwei Weng, Kun Wang 0042, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
CVPR | 6 |
| 2024 | Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionabstractDepth completion is a vital taskfor autonomous driving, as it involves reconstructing the precise 3D geometry of a scene from sparse and noisy depth measurements. How-ever, most existing methods either rely only on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. To address this chal-lenge, we introduce Tri-Perspective View Decomposition (TPVD), a novel framework that can explicitly model 3D geometry. In particular, (1) TPVD ingeniously decomposes the original point cloud into three 2D views, one of which corresponds to the sparse depth input. (2) We design TPV Fusion to update the 2D TPV features through recurrent 2D-3D-2D aggregation, where a Distance-Aware Spherical Convolution (DASC) is applied. (3) By adaptively choosing TPVaffinitive neighbors, the newly proposed Geometric Spatial Propagation Network (GSPN) further improves the geometric consistency. As a result, our TPVD outperforms existing methods on KITTI, NYUv2, and SUN RGBD. Fur-thermore, we build a novel depth completion dataset named TOFDC, which is acquired by the time-of-flight (TOF) sen-sor and the color camera on smart phones. Project page. Zhiqiang Yan 0001, Yuankai Lin, Kun Wang 0042, Yupeng Zheng, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
CVPR | 7 |
| 2024 | MFF-YOLO: Multi-scale Feature Fusion Network for Small Ship Detection in Night ScenesabstractShip detection plays a critical role in intelligent maritime applications including port management, marine monitoring and so on. Most existing ship detectors are trained on high-quality conventional-sized ship images under normal lighting conditions. However,low-quality images with poor lighting conditions often exist, where the features are difficult to be distinguished. Additionally, small ships with fewer pixels exhibit minimal appearance information and weak contour characteristics in night scenes, which are harmful to the multi-scale feature fusion. To address the above challenges, we propose an effective multi-scale feature fusion network for small ship detection in night scenes named MFF-YOLO. Specifically, we first design a night-friendly enhanced channel attention module, to better represent channel-dimensional features of small ships. In addition, we construct a multi-scale feature fusion architecture based on space and channel, to obtain richer semantic information of small ships in poor lighting conditions and further enhance the feature distinguishability. Finally, a series of experiments are implemented and corresponding results demonstrate the effectiveness and feasibility of our proposed method. Jun Li 0027, Hai Cao, Houjun Wang, Weili Guo, Chen Gong 0002 |
ECAI | 2 |
| 2024 | TP2O: Creative Text Pair-to-Object Generation Using Balance Swap-Sampling
Jun Li 0027, Zedong Zhang, Jian Yang 0003 |
ECCV (72) | 1 |
| 2024 | Diff-Reg: Diffusion Model in Doubly Stochastic Matrix Space for Registration Problem
Qianliang Wu, Haobo Jiang, Lei Luo 0001, Jun Li 0027, Yaqing Ding 0001, Jin Xie 0001, Jian Yang 0003 |
ECCV (65) | 4 |
| 2024 | Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable AttributeabstractGenerating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to feature coupling and the scarcity of realistic 3D human avatar datasets. To address these issues, we propose Text2Avatar, which can generate realistic-style 3D avatars based on the coupled text prompts. Text2Avatar leverages a discrete codebook as an intermediate feature to establish a connection between text and avatars, enabling the disentanglement of features. Furthermore, to alleviate the scarcity of realistic style 3D human avatar data, we utilize a pre-trained unconditional 3D human avatar generation model to obtain a large amount of 3D avatar pseudo data, which allows Text2Avatar to achieve realistic style generation. Experimental results demonstrate that our method can generate realistic 3D avatars from coupled textual data, which is challenging for other existing methods in this field. Chaoqun Gong, Yuqin Dai, Ronghui Li, Achun Bao, Jun Li 0027, Jian Yang 0003, Yachao Zhang 0001, Xiu Li 0001 |
ICASSP | 5 |
| 2024 | Exploring Multi-Modal Control in Music-Driven Dance GenerationabstractExisting music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal control, including genre control, semantic control, and spatial control. First, we decouple the dance generation network from the dance control network, thereby avoiding the degradation in dance quality when adding additional control information. Second, we design specific control strategies for different control information and integrate them into a unified framework. Experimental results show that the proposed dance generation framework outperforms state-of-the-art methods in terms of motion quality and controllability. Ronghui Li, Yuqin Dai, Yachao Zhang 0001, Jun Li 0027, Jian Yang 0003, Xiu Li 0001 |
ICASSP | 4 |
| 2024 | DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine DomainabstractIn this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cosine domain. This unique formulation allows for the modeling of local depth correlations within each patch. Crucially, the frequency transformation segregates the depth information into various frequency components, with low-frequency components encapsulating the core scene structure and high-frequency components detailing the finer aspects. This decomposition forms the basis of our progressive strategy, which begins with the prediction of low-frequency components to establish a global scene context, followed by successive refinement of local details through the prediction of higher-frequency components. We conduct comprehensive experiments on NYU-Depth-V2, TOFDC, and KITTI datasets, and demonstrate the state-of-the-art performance of DCDepth. Code is available at https://github.com/w2kun/DCDepth. Kun Wang 0042, Zhiqiang Yan 0001, Junkai Fan, Wanlu Zhu, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
NeurIPS | 6 |
| 2024 | MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State SpaceabstractRecent advances in low light image enhancement have been dominated by Retinex-based learning framework, leveraging convolutional neural networks (CNNs) and Transformers. However, the vanilla Retinex theory primarily addresses global illumination degradation and neglects local issues such as noise and blur in dark conditions. Moreover, CNNs and Transformers struggle to capture global degradation due to their limited receptive fields. While state space models (SSMs) have shown promise in the long-sequence modeling, they face challenges in combining local invariants and global context in visual data. In this paper, we introduce MambaLLIE, an implicit Retinex-aware low light enhancer featuring a global-then-local state space design. We first propose a Local-Enhanced State Space Module (LESSM) that incorporates an augmented local bias within a 2D selective scan mechanism, enhancing the original SSMs by preserving local 2D dependency. Additionally, an Implicit Retinex-aware Selective Kernel module (IRSK) dynamically selects features using spatially-varying operations, adapting to varying inputs through an adaptive kernel selection process. Our Global-then-Local State Space Block (GLSSB) integrates LESSM and IRSK with layer normalization (LN) as its core. This design enables MambaLLIE to achieve comprehensive global long-range modeling and flexible local feature aggregation. Extensive experiments demonstrate that MambaLLIE significantly outperforms state-of-the-art CNN and Transformer-based methods. Our code is available at https://github.com/wengjiangwei/MambaLLIE. Jiangwei Weng, Zhiqiang Yan 0001, Ying Tai, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
NeurIPS | 6 |
| 2024 | Novel Object Synthesis via Adaptive Text-Image HarmonyabstractIn this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task, \textit{i.e.}, often generating an object that predominantly reflects either the text or the image due to an imbalance between their inputs. To address this issue, we propose a simple yet effective method called Adaptive Text-Image Harmony (ATIH) to generate novel and surprising objects.
First, we introduce a scale factor and an injection step to balance text and image features in cross-attention and to preserve image information in self-attention during the text-image inversion diffusion process, respectively. Second, to better integrate object text and image, we design a balanced loss function with a noise parameter, ensuring both optimal editability and fidelity of the object image. Third, to adaptively adjust these parameters, we present a novel similarity score function that not only maximizes the similarities between the generated object image and the input text/image but also balances these similarities to harmonize text and image integration.
Extensive experiments demonstrate the effectiveness of our approach, showcasing remarkable object creations such as colobus-glass jar. https://xzr52.github.io/ATIH/ Zeren Xiong, Zedong Zhang, Shuo Chen 0003, Xiang Li 0041, Gan Sun, Jian Yang 0003, Jun Li 0027 |
NeurIPS | 8 |
| 2024 | Learning Fully Parametric Subspace Clustering
Xuanrong Chen, Jianjun Qian, Shuo Chen 0003, Jian Yang 0003, Jun Li 0027 |
PRCV (1) | 6 |
| 2024 | Neural Garment Dynamic Super-Resolutionabstractshapes, motions, and garment types not present in the training data.We demonstrate significant improvements over state-of-the-art alternatives, particularly in enhancing the quality of high-frequency, fine-grained wrinkle details.Code and data is released in https://github.com/MengZephyr/Neural- Meng Zhang 0043, Jun Li 0027 |
SIGGRAPH Asia | 2 |
| 2024 | Garment Animation NeRF with Color EditingabstractAbstract Generating high‐fidelity garment animations through traditional workflows, from modeling to rendering, is both tedious and expensive. These workflows often require repetitive steps in response to updates in character motion, rendering viewpoint changes, or appearance edits. Although recent neural rendering offers an efficient solution for computationally intensive processes, it struggles with rendering complex garment animations containing fine wrinkle details and realistic garment‐and‐body occlusions, while maintaining structural consistency across frames and dense view rendering. In this paper, we propose a novel approach to directly synthesize garment animations from body motion sequences without the need for an explicit garment proxy. Our approach infers garment dynamic features from body motion, providing a preliminary overview of garment structure. Simultaneously, we capture detailed features from synthesized reference images of the garment's front and back, generated by a pre‐trained image model. These features are then used to construct a neural radiance field that renders the garment animation video. Additionally, our technique enables garment recoloring by decomposing its visual elements. We demonstrate the generalizability of our method across unseen body motions and camera views, ensuring detailed structural consistency. Furthermore, we showcase its applicability to color editing on both real and synthetic garment data. Compared to existing neural rendering techniques, our method exhibits qualitative and quantitative improvements in garment dynamics and wrinkle detail modeling. Code is available at https://github.com/wrk226/GarmentAnimationNeRF . Renke Wang, Meng Zhang 0043, Jun Li 0027, Jian Yang 0003 |
Comput. Graph. Forum | 3 |
| 2024 | Contrastive subspace distribution learning for novel category discovery in high-dimensional visual data
Shuai Wei, Shuo Chen 0003, Jian Yang 0003, Jianjun Qian, Jun Li 0027 |
Knowl. Based Syst. | 6 |
| 2024 | Create Your World: Lifelong Text-to-Image DiffusionabstractText-to-image generative models can produce diverse high-quality images of concepts with a text prompt, which have demonstrated excellent ability in image generation, image translation, etc. We in this work study the problem of synthesizing instantiations of a user's own concepts in a never-ending manner,i.e.,create your world, where the new concepts from user are quickly learned with a few examples. To achieve this goal, we propose aLifelong text-to-imageDiffusionModel (L$^{2}$DM), which intends to overcome knowledge “catastrophic forgetting” for the past encountered concepts, and semantic “catastrophic neglecting” for one or more concepts in the text prompt. In respect of knowledge “catastrophic forgetting”, our L$^{2}$DM framework devises a task-aware memory enhancement module and an elastic-concept distillation module, which could respectively safeguard the knowledge of both prior concepts and each past personalized concept. When generating images with a user text prompt, the solution to semantic “catastrophic neglecting” is that a concept attention artist module can alleviate the semantic neglecting from concept aspect, and an orthogonal attention module can reduce the semantic binding from attribute aspect. To the end, our model can generate more faithful image across a range of continual text prompts in terms of both qualitative and quantitative metrics, when comparing with the related state-of-the-art models. The code will be released athttps://wenqiliang.github.io/. Gan Sun, Wenqi Liang, Jiahua Dong 0001, Jun Li 0027, Zhengming Ding, Yang Cong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Dual feature disentanglement for face anti-spoofing
Yimei Ma, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2024 | Pedestrian Crossing Intention Prediction Based on Cross-Modal Transformer and Uncertainty-Aware Multi-Task Learning for Autonomous DrivingabstractAccurate prediction of whether pedestrians will cross the street is prevalently recognized as an indispensable function of autonomous driving systems, especially in urban environments. How to utilize the complementary information present in different types of data (or modalities) is one of the major challenges. This paper makes the first attempt to develop a cross-modal transformer-based crossing intention prediction model merely using bounding boxes and ego-vehicle speed as input features. The cross-modal transformer can leverage self-attention and cross-modal attention to mine the modality-specific and complementary correlation. A bottleneck feature fusion is presented to obtain the compressed feature representation. To facilitate the network training, we further put forward a novel uncertainty-aware multi-task learning method that jointly predicts the future bounding box as well as crossing action such that the commonalities and differences across two tasks can be exploited. To evaluate the proposed method, extensive comparative experiments and ablation studies are performed on two benchmark datasets. The results demonstrate that by only using the bounding box and ego-vehicle speed as input features, our model is on a par with other state-of-the-art approaches that rely on more inputs, and even achieves superior performance in most cases. The source code will be released at https://github.com/xbchen82/PedCMT. Xiaobo Chen 0001, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Learning Complementary Correlations for Depth Super-Resolution With Incomplete Data in Real WorldabstractDepth information is a significant ingredient to visually perceive the physical world. However, mainstream depth sensors, e.g., time-of-flight (ToF) cameras, often measure incomplete and low-resolution depth data, resulting in low-quality visual perception. In this article, we try to address a potentially valuable task, i.e., depth super-resolution (DSR) with incomplete data, which recovers dense and high-resolution depth map from incomplete and low-resolution one. To tackle this task, we introduce a novel incomplete DSR (IDSR) framework, including a primary branch for DSR to recover high-frequency details, and an auxiliary branch for depth completion (DC) to fill missing pixels. More importantly, we propose two modules, joint correlation learning (JCL) and iterative-cross (IC), to enhance the learning of complementary information flows between the two branches. The former module aims to learn the correlative relationships of the two branches, whilst the latter module adequately fuses higher level representations for more precise predictions. Extensive experiments show that our framework is effective and achieves the state-of-the-art performance on the real-world RGB-D-D and the synthetic NYUv2 datasets. Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Curriculum Temperature for Knowledge DistillationabstractMost existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid search. In general, the temperature controls the discrepancy between two distributions and can faithfully determine the difficulty level of the distillation task. Keeping a constant temperature, i.e., a fixed level of task difficulty, is usually sub-optimal for a growing student during its progressive learning stages. In this paper, we propose a simple curriculum-based technique, termed Curriculum Temperature for Knowledge Distillation (CTKD), which controls the task difficulty level during the student's learning career through a dynamic and learnable temperature. Specifically, following an easy-to-hard curriculum, we gradually increase the distillation loss w.r.t. the temperature, leading to increased distillation difficulty in an adversarial manner. As an easy-to-use plug-in technique, CTKD can be seamlessly integrated into existing knowledge distillation frameworks and brings general improvements at a negligible additional computation cost. Extensive experiments on CIFAR-100, ImageNet-2012, and MS-COCO demonstrate the effectiveness of our method. Zheng Li 0028, Xiang Li 0041, Lingfeng Yang, Borui Zhao, Renjie Song, Lei Luo 0001, Jun Li 0027, Jian Yang 0003 |
AAAI | 7 |
| 2023 | DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth CompletionabstractUnsupervised depth completion aims to recover dense depth from the sparse one without using the ground-truth annotation. Although depth measurement obtained from LiDAR is usually sparse, it contains valid and real distance information, i.e., scale-consistent absolute depth values. Meanwhile, scale-agnostic counterparts seek to estimate relative depth and have achieved impressive performance. To leverage both the inherent characteristics, we thus suggest to model scale-consistent depth upon unsupervised scale-agnostic frameworks. Specifically, we propose the decomposed scale-consistent learning (DSCL) strategy, which disintegrates the absolute depth into relative depth prediction and global scale estimation, contributing to individual learning benefits. But unfortunately, most existing unsupervised scale-agnostic frameworks heavily suffer from depth holes due to the extremely sparse depth input and weak supervisory signal. To tackle this issue, we introduce the global depth guidance (GDG) module, which attentively propagates dense depth reference into the sparse target via novel dense-to-sparse attention. Extensive experiments show the superiority of our method on outdoor KITTI, ranking 1st and outperforming the best KBNet more than 12% in RMSE. Additionally, our approach achieves state-of-the-art performance on indoor NYUv2 benchmark as well. Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
AAAI | 5 |
| 2023 | Recurrent Structure Attention Guidance for Depth Super-resolutionabstractImage guidance is an effective strategy for depth super-resolution. Generally, most existing methods employ hand-crafted operators to decompose the high-frequency (HF) and low-frequency (LF) ingredients from low-resolution depth maps and guide the HF ingredients by directly concatenating them with image features. However, the hand-designed operators usually cause inferior HF maps (e.g., distorted or structurally missing) due to the diverse appearance of complex depth maps. Moreover, the direct concatenation often results in weak guidance because not all image features have a positive effect on the HF maps. In this paper, we develop a recurrent structure attention guided (RSAG) framework, consisting of two important parts. First, we introduce a deep contrastive network with multi-scale filters for adaptive frequency-domain separation, which adopts contrastive networks from large filters to small ones to calculate the pixel contrasts for adaptive high-quality HF predictions. Second, instead of the coarse concatenation guidance, we propose a recurrent structure attention block, which iteratively utilizes the latest depth estimation and the image features to jointly select clear patterns and boundaries, aiming at providing refined guidance for accurate depth recovery. In addition, we fuse the features of HF maps to enhance the edge structures in the decomposed LF maps. Extensive experiments show that our approach obtains superior performance compared with state-of-the-art depth super-resolution methods. Our code is available at: https://github.com/Yuanjiayii/DSR-RSAG. Jiayi Yuan 0003, Haobo Jiang, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
AAAI | 5 |
| 2023 | Structure Flow-Guided Network for Real Depth Super-resolutionabstractReal depth super-resolution (DSR), unlike synthetic settings, is a challenging task due to the structural distortion and the edge noise caused by the natural degradation in real-world low-resolution (LR) depth maps. These defeats result in significant structure inconsistency between the depth map and the RGB guidance, which potentially confuses the RGB-structure guidance and thereby degrades the DSR quality. In this paper, we propose a novel structure flow-guided DSR framework, where a cross-modality flow map is learned to guide the RGB-structure information transferring for precise depth upsampling. Specifically, our framework consists of a cross-modality flow-guided upsampling network (CFUNet) and a flow-enhanced pyramid edge attention network (PEANet). CFUNet contains a trilateral self-attention module combining both the geometric and semantic correlations for reliable cross-modality flow learning. Then, the learned flow maps are combined with the grid-sampling mechanism for coarse high-resolution (HR) depth prediction. PEANet targets at integrating the learned flow map as the edge attention into a pyramid network to hierarchically learn the edge-focused guidance feature for depth edge refinement. Extensive experiments on real and synthetic DSR datasets verify that our approach achieves excellent performance compared to state-of-the-art methods. Our code is available at: https://github.com/Yuanjiayii/DSR-SFG. Jiayi Yuan 0003, Haobo Jiang, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
AAAI | 5 |
| 2023 | Creative Birds: Self-Supervised Single-View 3D Style TransferabstractIn this paper, we propose a novel method for single-view 3D style transfer that generates a unique 3D object with both shape and texture transfer. Our focus lies primarily on birds, a popular subject in 3D reconstruction, for which no existing single-view 3D transfer methods have been developed. The method we propose seeks to generate a 3D mesh shape and texture of a bird from two single-view images. To achieve this, we introduce a novel shape transfer generator that comprises a dual residual gated network (DRGNet), and a multi-layer perceptron (MLP). DRGNet extracts the features of source and target images using a shared coordinate gate unit, while the MLP generates spatial coordinates for building a 3D mesh. We also introduce a semantic UV texture transfer module that implements textural style transfer using semantic UV segmentation, which ensures consistency in the semantic meaning of the transferred regions. This module can be widely adapted to many existing approaches. Finally, our method constructs a novel 3D bird using a differentiable renderer. Experimental results on the CUB dataset verify that our method achieves state-of-the-art performance on the single-view 3D style transfer task. Code is available at https://github.com/wrk226/creative_birds. Renke Wang, Guimin Que, Shuo Chen 0003, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
ICCV | 5 |
| 2023 | Enhanced Frequency Information for Image Dehazing
Junkai Fan, Jun Li 0027, Jian Yang 0003 |
ICIG (1) | 3 |
| 2023 | Distortion and Uncertainty Aware Loss for Panoramic Depth CompletionabstractStandard MSE or MAE loss function is commonly used in limited field-of-vision depth completion, treating each pixel equally under a basic assumption that all pixels have same contribution during optimization. Recently, with the rapid rise of panoramic photography, panoramic depth completion (PDC) has raised increasing attention in 3D computer vision. However, the assumption is inapplicable to panoramic data due to its latitude-wise distortion and high uncertainty nearby textures and edges. To handle these challenges, we propose distortion and uncertainty aware loss (DUL) that consists of a distortion-aware loss and an uncertainty-aware loss. The distortion-aware loss is designed to tackle the panoramic distortion caused by equirectangular projection, whose coordinate transformation relation is used to adaptively calculate the weight of the latitude-wise distortion, distributing uneven importance instead of the equal treatment for each pixel. The uncertainty-aware loss is presented to handle the inaccuracy in non-smooth regions. Specifically, we characterize uncertainty into PDC solutions under Bayesian deep learning framework, where a novel consistent uncertainty estimation constraint is designed to learn the consistency between multiple uncertainty maps of a single panorama. This consistency constraint allows model to produce more precise uncertainty estimation that is robust to feature deformation. Extensive experiments show the superiority of our method over standard loss functions, reaching the state of the art. Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003 |
ICML | 5 |
| 2023 | An Efficient Enhanced-YOLOv5 Algorithm for Multi-scale Ship Detection
Jun Li 0027, Haobo Jiang, Weili Guo, Chen Gong 0002 |
ICONIP (6) | 1 |
| 2023 | TSSAT: Two-Stage Statistics-Aware Transformation for Artistic Style TransferabstractArtistic style transfer aims to create new artistic images by rendering a given photograph with the target artistic style. Existing methods learn styles simply based on global statistics or local patches, lacking careful consideration of the drawing process in practice. Consequently, the stylization results either fail to capture abundant and diversified local style patterns, or contain undesired semantic information of the style image and deviate from the global style distribution. To address this issue, we imitate the drawing process of humans and propose a Two-Stage Statistics-Aware Transformation (TSSAT) module, which first builds the global style foundation by aligning the global statistics of content and style features and then further enriches local style details by swapping the local statistics (instead of local features) in a patch-wise manner, significantly improving the stylization effects. Moreover, to further enhance both content and style representations, we introduce two novel losses: an attention-based content loss and a patch-based style loss, where the former enables better content preservation by enforcing the semantic relation in the content image to be retained during stylization, and the latter focuses on increasing the local style similarity between the style and stylized images. Extensive qualitative and quantitative experiments verify the effectiveness of our method. Haibo Chen 0006, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
ACM Multimedia | 3 |
| 2023 | Effective Small Ship Detection with Enhanced-YOLOv7
Jun Li 0027, Chen Gong 0002, Zhong Jin |
PRCV (10) | 1 |
| 2023 | Quality-aware pattern diffusion for video object segmentation
Chuanwei Zhou, Chunyan Xu, Jun Li 0027, Zhen Cui 0001, Jian Yang 0003 |
Neurocomputing | 3 |
| 2023 | Fast subspace clustering by learning projective block diagonal representation
Yesong Xu, Shuo Chen 0003, Jun Li 0027, Chunyan Xu, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2023 | InOR-Net: Incremental 3-D Object Recognition Network for Point Cloud Representationabstract3-D object recognition has successfully become an appealing research topic in the real world. However, most existing recognition models unreasonably assume that the categories of 3-D objects cannot change over time in the real world. This unrealistic assumption may result in significant performance degradation for them to learn new classes of 3-D objects consecutively due to the catastrophic forgetting on old learned classes. Moreover, they cannot explore which 3-D geometric characteristics are essential to alleviate the catastrophic forgetting on old classes of 3-D objects. To tackle the above challenges, we develop a novel Incremental 3-D Object Recognition Network (i.e., InOR-Net), which could recognize new classes of 3-D objects continuously by overcoming the catastrophic forgetting on old classes. Specifically, category-guided geometric reasoning is proposed to reason local geometric structures with distinctive 3-D characteristics of each class by leveraging intrinsic category information. We then propose a novel critic-induced geometric attention mechanism to distinguish which 3-D geometric characteristics within each class are beneficial to overcome the catastrophic forgetting on old classes of 3-D objects while preventing the negative influence of useless 3-D characteristics. In addition, a dual adaptive fairness compensations' strategy is designed to overcome the forgetting brought by class imbalance by compensating biased weights and predictions of the classifier. Comparison experiments verify the state-of-the-art performance of the proposed InOR-Net model on several public point cloud datasets. Jiahua Dong 0001, Yang Cong, Gan Sun, Lixu Wang, Lingjuan Lyu, Jun Li 0027, Ender Konukoglu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | LRPRNet: Lightweight Deep Network by Low-Rank Pointwise Residual ConvolutionabstractDeep learning has become popular in recent years primarily due to powerful computing devices such as graphics processing units (GPUs). However, it is challenging to deploy these deep models to multimedia devices, smartphones, or embedded systems with limited resources. To reduce the computation and memory costs, we propose a novel lightweight deep learning module by low-rank pointwise residual (LRPR) convolution, called LRPRNet. Essentially, LRPR aims at using a low-rank approximation in pointwise convolution to further reduce the module size while keeping depthwise convolutions as the residual module to rectify the LRPR module. This is critical when the low-rankness undermines the convolution process. Moreover, our LRPR is quite general and can be directly applied to many existing network architectures such as MobileNetv1, ShuffleNetv2, MixNet, and so on. Experiments on visual recognition tasks, including image classification and face alignment on popular benchmarks, show that our LRPRNet achieves competitive performance but with a significant reduction of Flops and memory cost compared to the state-of-the-art deep lightweight models. Bin Sun 0002, Jun Li 0027, Ming Shao, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | From Ensemble Clustering to Subspace Clustering: Cluster Structure EncodingabstractIn this study, we propose a novel algorithm to encode the cluster structure by incorporating ensemble clustering (EC) into subspace clustering (SC). First, the low-rank representation (LRR) is learned from a higher order data relationship induced by ensemble K-means coding, which exploits the cluster structure in a co-association matrix of basic partitions (i.e., clustering results). Second, to provide a fast predictive coding mechanism, an encoding function parameterized by neural networks is introduced to predict the LRR derived from partitions. These two steps are jointly proceeded to seamlessly integrate partition information and original features and thus deliver better representations than the ones obtained from each single source. Moreover, an alternating optimization framework is developed to learn the LRR, train the encoding function, and fine-tune the higher order relationship. Extensive experiments on eight benchmark datasets validate the effectiveness of the proposed algorithm on several clustering tasks compared with state-of-the-art EC and SC methods. Zhiqiang Tao, Jun Li 0027, Huazhu Fu, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Linearity-Aware Subspace ClusteringabstractObtaining a good similarity matrix is extremely important in subspace clustering. Current state-of-the-art methods learn the similarity matrix through self-expressive strategy. However, these methods directly adopt original samples as a set of basis to represent itself linearly. It is difficult to accurately describe the linear relation between samples in the real-world applications, and thus is hard to find an ideal similarity matrix. To better represent the linear relation of samples, we present a subspace clustering model, Linearity-Aware Subspace Clustering (LASC), which can consciously learn the similarity matrix by employing a linearity-aware metric. This is a new subspace clustering method that combines metric learning and subspace clustering into a joint learning framework. In our model, we first utilize the self-expressive strategy to obtain an initial subspace structure and discover a low-dimensional representation of the original data. Subsequently, we use the proposed metric to learn an intrinsic similarity matrix with linearity-aware on the obtained subspace. Based on such a learned similarity matrix, the inter-cluster distance becomes larger than the intra-cluster distances, and thus successfully obtaining a good subspace cluster result. In addition, to enrich the similarity matrix with more consistent knowledge, we adopt a collaborative learning strategy for self-expressive subspace learning and linearity-aware subspace learning. Moreover, we provide detailed mathematical analysis to show that the metric can properly characterize the linear correlation between samples. Yesong Xu, Shuo Chen 0003, Jun Li 0027, Jianjun Qian |
AAAI | 3 |
| 2022 | Learning Clothes-irrelevant Cues for Clothes-Changing Person Re-identification
Jingyi Mu, Jun Li 0027, Jian Yang 0003 |
BMVC | 3 |
| 2022 | Industrial Style Transfer with Large-scale Geometric Warping and Content PreservationabstractWe propose a novel style transfer method to quickly create a new visual product with a nice appearance for industrial designers reference. Given a source product, a target product, and an art style image, our method produces a neural warping field that warps the source shape to imitate the geometric style of the target and a neural texture transformation network that transfers the artistic style to the warped source product. Our model, Industrial Style Transfer (InST), consists of large-scale geometric warping (LGW) and interest-consistency texture transfer (ICTT). LGW aims to explore an unsupervised transformation between the shape masks of the source and target products for fitting large-scale shape warping. Furthermore, we introduce a mask smoothness regularization term to prevent the abrupt changes of the details of the source product. ICTT introduces an interest regularization term to maintain important contents of the warped product when it is stylized by using the art style image. Extensive experimental results demonstrate that InST achieves state-of-the-art performance on multiple visual product design tasks, e.g., companies' snail logos and classical bottles (please see Fig. 1). To the best of our knowledge, we are the first to extend the neural style transfer method to create industrial product appearances. Code is available at https://jcyang98.github.io/InST/home.html Jinchao Yang, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003 |
CVPR | 4 |
| 2022 | Multi-modal Masked Pre-training for Monocular Panoramic Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
ECCV (1) | 5 |
| 2022 | RigNet: Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
ECCV (27) | 5 |
| 2022 | PPT: Anomaly Detection Dataset of Printed Products with TemplatesabstractVisual anomaly detection has been an active topic in industrial applications. In particular, it aims to classify anomalies and precisely locate defective areas in the printed products. To the best of our knowledge, there is no anomaly detection dataset for industrial printings. In this paper, we are the first to introduce a Printed Products with Templates (PPT) dataset, which contains large templates and sliced images collected from industry scene images. PPT is a challenging dataset with more variable surface defects and more disturbing background than existing related benchmarks. Furthermore, we propose a template matching method for anomaly detection of printed products, which consists of a fast template matching block with a convolutional operation using the test sliced image as its kernel, and a prediction network for generating an anomaly map of the test sliced image. Experimental results show that our method achieves state-of-the-art performance compared to the related anomaly detection approaches. Huang Tian, Xiang Li 0041, Lingfeng Yang, Jun Li 0027, Jian Yang 0003, Weidong Du |
ICIP | 4 |
| 2022 | Learning Contrastive Embedding in Low-Dimensional SpaceabstractContrastive learning (CL) pretrains feature embeddings to scatter instances in the feature space so that the training data can be well discriminated. Most existing CL techniques usually encourage learning such feature embeddings in the highdimensional space to maximize the instance discrimination. However, this practice may lead to undesired results where the scattering instances are sparsely distributed in the high-dimensional feature space, making it difficult to capture the underlying similarity between pairwise instances. To this end, we propose a novel framework called contrastive learning with low-dimensional reconstruction (CLLR), which adopts a regularized projection layer to reduce the dimensionality of the feature embedding. In CLLR, we build the sparse / low-rank regularizer to adaptively reconstruct a low-dimensional projection space while preserving the basic objective for instance discrimination, and thus successfully learning contrastive embeddings that alleviate the above issue. Theoretically, we prove a tighter error bound for CLLR; empirically, the superiority of CLLR is demonstrated across multiple domains. Both theoretical and experimental results emphasize the significance of learning low-dimensional contrastive embeddings. Shuo Chen 0003, Chen Gong 0002, Jun Li 0027, Jian Yang 0003, Gang Niu 0001, Masashi Sugiyama |
NeurIPS | 3 |
| 2022 | Streaming feature selection via graph diffusion
Shuo Chen 0003, Zhenyong Fu, Jun Li 0027, Jian Yang 0003 |
Inf. Sci. | 4 |
| 2022 | Fast Multi-View Outlier Detection via Deep EncoderabstractMulti-view outlier detection has a wide range of applications and has been well investigated in recent years. However, 1) most existing state-of-the-art methods cannot efficiently handle outlier detection problem for large-scale multi-view data, since exploring pairwise constraints among different views causes highly-computational cost; 2) the data collected from original heterogeneous feature spaces further increases the consistent difficulty of multi-view outlier detection. To address these issues, we present a fast multi-view outlier detection model via learning a low-rank latent subspace representation with deep encoder architecture, which can not only efficiently identify the outliers for large-scale data even with numerous data views, but also exploit a discriminative common latent subspace shared by all the views. First, we learn a set of orthogonal bases as view-specific dictionaries from a small dataset, which is randomly sampled from the original dataset. Benefitting from view-specific dictionaries, the sampled data is projected and decomposed as a shared and discriminative latent subspace representations, which correspond to the view-consistent and view-specific components across multiple views, respectively. Then, the obtained discriminative latent representations are applied to train the view-specific deep encoders, which can efficiently compute the abnormal score for the remaining instances. Our proposed model can cost-effectively identify the outliers in large-scale datasets from numerous data views with less computational complexity. Experiments conducted on eight real datasets and a synthesis dataset show that our proposed model outperforms the existing ones on effectiveness and efficiency. Dongdong Hou, Yang Cong, Gan Sun, Jiahua Dong 0001, Jun Li 0027, Kai Li 0012 |
IEEE Trans. Big Data | 5 |
| 2022 | Large-Scale Subspace Clustering by Independent Distributed and Parallel CodingabstractSubspace clustering is a popular method to discover underlying low-dimensional structures of high-dimensional multimedia data (e.g., images, videos, and texts). In this article, we consider a large-scale subspace clustering (LS2C) problem, that is, partitioning million data points with a millon dimensions. To address this, we explore an independent distributed and parallel framework by dividing big data/variable matrices and regularization by both columns and rows. Specifically, LS2C is independently decomposed into many subproblems by distributing those matrices into different machines by columns since the regularization of the code matrix is equal to a sum of that of its submatrices (e.g., square-of-Frobenius/$\ell _{1}$-norm). Consensus optimization is designed to solve these subproblems in a parallel way for saving communication costs. Moreover, we provide theoretical guarantees that LS2C can recover consensus subspace representations of high-dimensional data points under broad conditions. Compared with the state-of-the-art LS2C methods, our approach achieves better clustering results in public datasets, including a million images and videos. Jun Li 0027, Zhiqiang Tao, Yue Wu 0008, Bineng Zhong 0001, Yun Fu 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Autoencoder-Based Latent Block-Diagonal Representation for Subspace ClusteringabstractBlock-diagonal representation (BDR) is an effective subspace clustering method. The existing BDR methods usually obtain a self-expression coefficient matrix from the original features by a shallow linear model. However, the underlying structure of real-world data is often nonlinear, thus those methods cannot faithfully reflect the intrinsic relationship among samples. To address this problem, we propose a novel latent BDR (LBDR) model to perform the subspace clustering on a nonlinear structure, which jointly learns an autoencoder and a BDR matrix. The autoencoder, which consists of a nonlinear encoder and a linear decoder, plays an important role to learn features from the nonlinear samples. Meanwhile, the learned features are used as a new dictionary for a linear model with block-diagonal regularization, which can ensure good performances for spectral clustering. Moreover, we theoretically prove that the learned features are located in the linear space, thus ensuring the effectiveness of the linear model using self-expression. Extensive experiments on various real-world datasets verify the superiority of our LBDR over the state-of-the-art subspace clustering approaches. Yesong Xu, Shuo Chen 0003, Jun Li 0027, Zongyan Han, Jian Yang 0003 |
IEEE Trans. Cybern. | 3 |
| 2022 | Self-Supervised Multi-Category Counting Networks for Automatic Check-OutabstractThe practical task of Automatic Check-Out (ACO) is to accurately predict the presence and count of each product in an arbitrary product combination. Beyond the large-scale and the fine-grained nature of product categories as its main challenges, products are always continuously updated in realistic check-out scenarios, which is also required to be solved in an ACO system. Previous work in this research line almost depends on the supervisions of labor-intensive bounding boxes of products by performing a detection paradigm. While, in this paper, we propose a Self-Supervised Multi-Category Counting (S2MC2) network to leverage the point-level supervisions of products in check-out images to both lower the labeling cost and be able to return ACO predictions in a class incremental setting. Specifically, as a backbone, our S2MC2 is built upon a counting module in a class-agnostic counting fashion. Also, it consists of several crucial components including an attention module for capturing fine-grained patterns and a domain adaptation module for reducing the domain gap between single product images as training and check-out images as test. Furthermore, a self-supervised approach is utilized in S2MC2 to initialize the parameters of its backbone for better performance. By conducting comprehensive experiments on the large-scale automatic check-out dataset RPC, we demonstrate that our proposed S2MC2 achieves superior accuracy in both traditional and incremental settings of ACO tasks over the competing baselines. Hao Chen 0052, Yangzhun Zhou, Jun Li 0027, Xiu-Shen Wei, Liang Xiao 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | CBi-GNN: Cross-Scale Bilateral Graph Neural Network for 3D Object Detectionabstract3D object detection from LiDAR point clouds is a challenging task, since the point clouds are irregular and sparse. Existing one-stage methods mainly predict the 3D bounding box of 3D objects by extracting deep down-scaled features of point clouds from low-level (high-resolution, HR) feature maps to high-level (low-resolution, LR). Nonetheless, most of these methods ignore geometric context information of the down-scaled feature maps across scales, especially only using the LR feature will result in incomplete structure and less location accuracy of 3D objects. In this paper, we propose a novel cross-scale graph network-based one-stage 3D object detector to fully exploit the geometric contexts of the voxels between the down-scaled feature maps. Specifically, we first employ a 3D sparse convolution neural network to form different resolutions of feature maps of voxels. We then dynamically construct a cross-scale bilateral graph to search the neighbor non-empty voxels in the HR feature map with a fixed radius for each non-empty voxel in the LR feature map. In the constructed graph, we present a bilateral attention mechanism (i.e., self-attention and spatial attention) in the HR feature map and encode each non-empty voxel in the LR feature map by aggregating the HR features to obtain the attention features. In addition, we design a non-local part pooling operation to improve the score of the detected bounding box of 3D objects. Finally, we formulate a multi-task loss to train our network for regression of the 3D bounding box of the 3D objects. Experiments on the challenging KITTI’s 3D/BEV benchmark show that our proposed detector outperforms all one-stage 3D object detectors and is comparable to two-stage 3D object detectors. Our code is available athttps://github.com/csjxchen/CBi-GNN. Jiaxin Chen 0001, Xiang Li 0041, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Deep Learning for Fashion Style GenerationabstractIn this article, we work on generating fashion style images with deep neural network algorithms. Given a garment image, and single or multiple style images (e.g., flower, blue and white porcelain), it is a challenge to generate a synthesized clothing image with single or mix-and-match styles due to the need to preserve global clothing contents with coverable styles, to achieve high fidelity of local details, and to conform different styles with specific areas. To address this challenge, we propose a fashion style generator (FashionG) framework for the single-style generation and a spatially constrained FashionG (SC-FashionG) framework for mix-and-match style generation. Both FashionG and SC-FashionG are end-to-end feedforward neural networks that consist of a generator for image transformation and a discriminator for preserving content and style globally and locally. Specifically, a global-based loss is calculated based on full images, which can preserve the global clothing form and design. A patch-based loss is calculated based on image patches, which can preserve detailed local style patterns. We develop an alternating patch-global optimization methodology to minimize these losses. Compared with FashionG, SC-FashionG employs an additional spatial constraint to ensure that each style is blended only onto a specific area of the clothing image. Extensive experiments demonstrate the effectiveness of both single-style and mix-and-match style generations. Shuhui Jiang, Jun Li 0027, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Learnable Subspace ClusteringabstractThis article studies the large-scale subspace clustering (LS2C) problem with millions of data points. Many popular subspace clustering methods cannot directly handle the LS2C problem although they have been considered to be state-of-the-art methods for small-scale data points. A simple reason is that these methods often choose all data points as a large dictionary to build huge coding models, which results in high time and space complexity. In this article, we develop a learnable subspace clustering paradigm to efficiently solve the LS2C problem. The key concept is to learn a parametric function to partition the high-dimensional subspaces into their underlying low-dimensional subspaces instead of the computationally demanding classical coding models. Moreover, we propose a unified, robust, predictive coding machine (RPCM) to learn the parametric function, which can be solved by an alternating minimization algorithm. Besides, we provide a bounded contraction analysis of the parametric function. To the best of our knowledge, this article is the first work to efficiently cluster millions of data points among the subspace clustering methods. Experiments on million-scale data sets verify that our paradigm outperforms the related state-of-the-art methods in both efficiency and effectiveness. Jun Li 0027, Hongfu Liu 0001, Zhiqiang Tao, Handong Zhao, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object DetectionabstractLocalization Quality Estimation (LQE) is crucial and popular in the recent advancement of dense object detectors since it can provide accurate ranking scores that benefit the Non-Maximum Suppression processing and improve detection performance. As a common practice, most existing methods predict LQE scores through vanilla convolutional features shared with object classification or bounding box regression. In this paper, we explore a completely novel and different perspective to perform LQE – based on the learned distributions of the four parameters of the bounding box. The bounding box distributions are inspired and introduced as "General Distribution" in GFLV1, which describes the uncertainty of the predicted bounding boxes well. Such a property makes the distribution statistics of a bounding box highly correlated to its real localization quality. Specifically, a bounding box distribution with a sharp peak usually corresponds to high localization quality, and vice versa. By leveraging the close correlation between distribution statistics and the real localization quality, we develop a considerably lightweight Distribution-Guided Quality Predictor (DGQP) for reliable LQE based on GFLV1, thus producing GFLV2. To our best knowledge, it is the first attempt in object detection to use a highly relevant, statistical representation to facilitate LQE. Extensive experiments demonstrate the effectiveness of our method. Notably, GFLV2 (ResNet101) achieves 46.2 AP at 14.6 FPS, surpassing the previous state-of-the-art ATSS baseline (43.6 AP at 14.6 FPS) by absolute 2.6 AP on COCO test-dev, without sacrificing the efficiency both in training and inference. Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003 |
CVPR | 4 |
| 2021 | Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationabstractIn this paper, by modeling the point cloud registration task as a Markov decision process, we propose an end-to-end deep model embedded with the cross-entropy method (CEM) for unsupervised 3D registration. Our model consists of a sampling network module and a differentiable CEM module. In our sampling network module, given a pair of point clouds, the sampling network learns a prior sampling distribution over the transformation space. The learned sampling distribution can be used as a "good" initialization of the differentiable CEM module. In our differentiable CEM module, we first propose a maximum consensus criterion based alignment metric as the reward function for the point cloud registration task. Based on the reward function, for each state, we then construct a fused score function to evaluate the sampled transformations, where we weight the current and future rewards of the transformations. Particularly, the future rewards of the sampled transforms are obtained by performing the iterative closest point (ICP) algorithm on the transformed state. By selecting the top-k transformations with the highest scores, we iteratively update the sampling distribution. Furthermore, in order to make the CEM differentiable, we use the sparse-max function to replace the hard top-k selection. Finally, we formulate a Geman-McClure estimator based loss to train our end-to-end registration model. Extensive experimental results demonstrate the good registration performance of our method on benchmark datasets. Code is available at https://github.com/Jiang-HB/CEMNet. Haobo Jiang, Yaqi Shen, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003 |
ICCV | 4 |
| 2021 | Regularizing Nighttime Weirdness: Efficient Self-supervised Monocular Depth Estimation in the DarkabstractMonocular depth estimation aims at predicting depth from a single image or video. Recently, self-supervised methods draw much attention since they are free of depth annotations and achieve impressive performance on several daytime benchmarks. However, they produce weird outputs in more challenging nighttime scenarios because of low visibility and varying illuminations, which bring weak textures and break brightness-consistency assumption, respectively. To address these problems, in this paper we propose a novel framework with several improvements: (1) we introduce Priors-Based Regularization to learn distribution knowledge from unpaired depth maps and prevent model from being incorrectly trained; (2) we leverage Mapping-Consistent Image Enhancement module to enhance image visibility and contrast while maintaining brightness consistency; and (3) we present Statistics-Based Mask strategy to tune the number of removed pixels within textureless regions, using dynamic statistics. Experimental results demonstrate the effectiveness of each component. Mean-while, our framework achieves remarkable improvements and state-of-the-art results on two nighttime datasets. Code is available at https://github.com/w2kun/RNW. Kun Wang 0042, Zhenyu Zhang 0005, Zhiqiang Yan 0001, Xiang Li 0041, Baobei Xu, Jun Li 0027, Jian Yang 0003 |
ICCV | 6 |
| 2021 | Large-Margin Contrastive Learning with Distance Polarization Regularizerabstract\emph{Contrastive learning} (CL) pretrains models in a pairwise manner, where given a data point, other data points are all regarded as dissimilar, including some that are \emph{semantically} similar. The issue has been addressed by properly weighting similar and dissimilar pairs as in \emph{positive-unlabeled learning}, so that the objective of CL is \emph{unbiased} and CL is \emph{consistent}. However, in this paper, we argue that this great solution is still not enough: its weighted objective \emph{hides} the issue where the semantically similar pairs are still pushed away; as CL is pretraining, this phenomenon is not our desideratum and might affect downstream tasks. To this end, we propose \emph{large-margin contrastive learning} (LMCL) with \emph{distance polarization regularizer}, motivated by the distribution characteristic of pairwise distances in \emph{metric learning}. In LMCL, we can distinguish between \emph{intra-cluster} and \emph{inter-cluster} pairs, and then only push away inter-cluster pairs, which \emph{solves} the above issue explicitly. Theoretically, we prove a tighter error bound for LMCL; empirically, the superiority of LMCL is demonstrated across multiple domains, \emph{i.e.}, image classification, sentence representation, and reinforcement learning. Shuo Chen 0003, Gang Niu 0001, Chen Gong 0002, Jun Li 0027, Jian Yang 0003, Masashi Sugiyama |
ICML | 4 |
| 2021 | Learnable low-rank latent dictionary for subspace clustering
Yesong Xu, Shuo Chen 0003, Jun Li 0027, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2021 | Clustering With Outlier RemovalabstractCluster analysis and outlier detection are two continuously rising topics in data mining area, which in fact connect to each other deeply. Cluster structure is vulnerable to outliers; inversely, outliers are the points belonging to none of any clusters. Unfortunately, most existing studies do not notice the coupled relationship between these two tasks and handle them separately. In this article, we consider the joint cluster analysis and outlier detection problem, and propose the Clustering with Outlier Removal (COR) algorithm. Specifically, the original space is transformed into a binary space via generating basic partitions. We employ Holoentropy to measure the compactness of each cluster without involving several outlier candidates. To provide a neat and efficient solution, an auxiliary binary matrix is introduced so that COR completely and efficiently solves the challenging problem via a unified K-means— with theoretical supports. Extensive experimental results on numerous data sets in various domains demonstrate the effectiveness and efficiency of COR significantly over state-of-the-art methods in terms of cluster validity and outlier detection. Some key factors including the basic partition number and generation strategy in COR with an application on abnormal flight trajectory detection are further analyzed for practical use. Hongfu Liu 0001, Jun Li 0027, Yue Wu 0008, Yun Fu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Robust Low-Rank Discovery of Data-Driven Partial Differential EquationsabstractPartial differential equations (PDEs) are essential foundations to model dynamic processes in natural sciences. Discovering the underlying PDEs of complex data collected from real world is key to understanding the dynamic processes of natural laws or behaviors. However, both the collected data and their partial derivatives are often corrupted by noise, especially from sparse outlying entries, due to measurement/process noise in the real-world applications. Our work is motivated by the observation that the underlying data modeled by PDEs are in fact often low rank. We thus develop a robust low-rank discovery framework to recover both the low-rank data and the sparse outlying entries by integrating double low-rank and sparse recoveries with a (group) sparse regression method, which is implemented as a minimization problem using mixed nuclear norms with ℓ1 and ℓ0 norms. We propose a low-rank sequential (grouped) threshold ridge regression algorithm to solve the minimization problem. Results from several experiments on seven canonical models (i.e., four PDEs and three parametric PDEs) verify that our framework outperforms the state-of-art sparse and group sparse regression methods. Code is available at https://github.com/junli2019/Robust-Discovery-of-PDEs Jun Li 0027, Gan Sun, Guoshuai Zhao 0001, Li-Wei H. Lehman |
AAAI | 1 |
| 2020 | Lifelong Spectral ClusteringabstractIn the past decades, spectral clustering (SC) has become one of the most effective clustering algorithms. However, most previous studies focus on spectral clustering tasks with a fixed task set, which cannot incorporate with a new spectral clustering task without accessing to previously learned tasks. In this paper, we aim to explore the problem of spectral clustering in a lifelong machine learning framework, i.e., Lifelong Spectral Clustering (L2SC). Its goal is to efficiently learn a model for a new spectral clustering task by selectively transferring previously accumulated experience from knowledge library. Specifically, the knowledge library of L2SC contains two components: 1) orthogonal basis library: capturing latent cluster centers among the clusters in each pair of tasks; 2) feature embedding library: embedding the feature manifold information shared among multiple related tasks. As a new spectral clustering task arrives, L2SC firstly transfers knowledge from both basis library and feature library to obtain encoding matrix, and further redefines the library base over time to maximize performance across all the clustering tasks. Meanwhile, a general online update formulation is derived to alternatively update the basis library and feature library. Finally, the empirical experiments on several real-world benchmark datasets demonstrate that our L2SC model can effectively improve the clustering performance when comparing with other state-of-the-art spectral clustering algorithms. Gan Sun, Yang Cong, Qianqian Wang 0001, Jun Li 0027, Yun Fu 0001 |
AAAI | 4 |
| 2020 | Block Mobilenet: Align Large-Pose Faces with <1MB Model Sizeabstract3D face alignment methods based on deep models have become very popular due to their empirical success. However, high time and space complexities make these methods difficult to be applied to mobile devices and embedded devices. To decrease the time and space complexity, we propose a novel Depthwise Separable Block (DSB) which consists of a depthwise block and a pointwise block. The depthwise block is constructed by stacking depthwise convolution layers and concatenating the low layer, and the pointwise block has only pointwise convolution layers stacking together. Moreover, we develop a light-weight Block-Mobilenet by using our DSBs to reconstruct Mobilenet. It is worth noting that our Block-Mobilenet successfully reduces network parameters from MB to KB. Experiments on four popular datasets verify that Block Mobilenet has better overall performance (mean NME on 68 points: 3.81%; speed on CPU: 91 FPS; storage size: 876 KB) than the state-of-the-art methods. Bin Sun 0002, Jun Li 0027, Yun Fu 0001 |
FG | 2 |
| 2020 | Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionabstractOne-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is to introduce an \emph{individual} prediction branch to estimate the quality of localization, where the predicted quality facilitates the classification to improve detection performance. This paper delves into the \emph{representations} of the above three fundamental elements: quality estimation, classification and localization. Two problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, and (2) the inflexible Dirac delta distribution for localization. To address the problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, and use a vector to represent arbitrary distribution of box locations. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain \emph{continuous} labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFL) that generalizes Focal Loss from its discrete form to the \emph{continuous} version for successful optimization. On COCO {\tt test-dev}, GFL achieves 45.0\% AP using ResNet-101 backbone, surpassing state-of-the-art SAPD (43.5\%) and ATSS (43.6\%) with higher or comparable inference speed. Xiang Li 0041, Wenhai Wang, Shuo Chen 0003, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003 |
NeurIPS | 6 |
| 2020 | Learning Semantically Enhanced Feature for Fine-Grained Image ClassificationabstractWe aim to provide a computationally cheap yet effective approach for fine-grained image classification (FGIC) in this letter. Unlike previous methods that rely on complex part localization modules, our approach learns fine-grained features by enhancing the semantics of sub-features of a global feature. Specifically, we first achieve the sub-feature semantic by arranging feature channels of a CNN into different groups through channel permutation. Meanwhile, to enhance the discriminability of sub-features, the groups are guided to be activated on object parts with strong discriminability by a weighted combination regularization. Our approach is parameter parsimonious and can be easily integrated into the backbone model as a plug-and-play module for end-to-end training with only image-level supervision. Experiments verified the effectiveness of our approach and validated its comparable performance to the state-of-the-art methods. Code is available at https://github.com/cswluo/SEF. Wei Luo 0006, Hengmin Zhang, Jun Li 0027, Xiu-Shen Wei |
IEEE Signal Process. Lett. | 3 |
| 2020 | Line-CNN: End-to-End Traffic Line Detection With Line Proposal UnitabstractThe task of traffic line detection is a fundamental yet challenging problem. Previous approaches usually conduct traffic line detection via a two-stage way, namely the line segment detection followed by a segment clustering, which is very likely to ignore the global semantic information of an entire line. To address the problem, we propose an end-to-end system called Line-CNN (L-CNN), in which the key component is a novel line proposal unit (LPU). The LPU utilizes line proposals as references to locate accurate traffic curves, which forces the system to learn the global feature representation of the entire traffic lines. We benchmark the proposed L-CNN on two public datasets including MIKKI and TuSimple, and the results suggest that L-CNN outperforms the state-of-the-art methods. In addition, L-CNN can run at approximately 30 f/s on a Titan X GPU, which indicates the practicability and effectiveness of L-CNN for real-time intelligent self-driving systems. Xiang Li 0041, Jun Li 0027, Xiaolin Hu 0001, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Data-Adaptive Metric Learning with Scale AlignmentabstractThe central problem for most existing metric learning methods is to find a suitable projection matrix on the differences of all pairs of data points. However, a single unified projection matrix can hardly characterize all data similarities accurately as the practical data are usually very complicated, and simply adopting one global projection matrix might ignore important local patterns hidden in the dataset. To address this issue, this paper proposes a novel method dubbed “Data-Adaptive Metric Learning” (DAML), which constructs a data-adaptive projection matrix for each data pair by selectively combining a set of learned candidate matrices. As a result, every data pair can obtain a specific projection matrix, enabling the proposed DAML to flexibly fit the training data and produce discriminative projection results. The model of DAML is formulated as an optimization problem which jointly learns candidate projection matrices and their sparse combination for every data pair. Nevertheless, the over-fitting problem may occur due to the large amount of parameters to be learned. To tackle this issue, we adopt the Total Variation (TV) regularizer to align the scales of data embedding produced by all candidate projection matrices, and thus the generated metrics of these learned candidates are generally comparable. Furthermore, we extend the basic linear DAML model to the kernerlized version (denoted “KDAML”) to handle the non-linear cases, and the Iterative Shrinkage-Thresholding Algorithm (ISTA) is employed to solve the optimization model. Intensive experimental results on various applications including retrieval, classification, and verification clearly demonstrate the superiority of our algorithm to other state-of-the-art metric learning methodologies. Shuo Chen 0003, Chen Gong 0002, Jian Yang 0003, Ying Tai, Le Hui, Jun Li 0027 |
AAAI | 6 |
| 2019 | Cross-X Learning for Fine-Grained Visual CategorizationabstractRecognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features are extracted for fine-grained classification. However, these methods typically treat the part-specific features of each image in isolation while neglecting their relationships between different images. In this paper, we propose Cross-X learning, a simple yet effective approach that exploits the relationships between different images and between different network layers for robust multi-scale feature learning. Our approach involves two novel components: (i) a cross-category cross-semantic regularizer that guides the extracted features to represent semantic parts and, (ii) a cross-layer regularizer that improves the robustness of multi-scale features by matching the prediction distribution across multiple layers. Our approach can be easily trained end-to-end and is scalable to large datasets like NABirds. We empirically analyze the contributions of different components of our approach and demonstrate its robustness, effectiveness and state-of-the-art performance on five benchmark datasets. Code is available at \url{https://github.com/cswluo/CrossX}. Wei Luo 0006, Xitong Yang, Xianjie Mo, Yuheng Lu, Larry Davis 0001, Jun Li 0027, Jian Yang 0003, Ser-Nam Lim |
ICCV | 6 |
| 2019 | Adversarial Graph Embedding for Ensemble ClusteringabstractEnsemble clustering generally integrates basic partitions into a consensus one through a graph partitioning method, which, however, has two limitations: 1) it neglects to reuse original features; 2) obtaining consensus partition with learnable graph representations is still under-explored. In this paper, we propose a novel Adversarial Graph Auto-Encoders (AGAE) model to incorporate ensemble clustering into a deep graph embedding process. Specifically, graph convolutional network is adopted as probabilistic encoder to jointly integrate the information from feature content and consensus graph, and a simple inner product layer is used as decoder to reconstruct graph with the encoded latent variables (i.e., embedding representations). Moreover, we develop an adversarial regularizer to guide the network training with an adaptive partition-dependent prior. Experiments on eight real-world datasets are presented to show the effectiveness of AGAE over several state-of-the-art deep embedding and ensemble clustering methods. Zhiqiang Tao, Hongfu Liu 0001, Jun Li 0027, Yun Fu 0001 |
IJCAI | 3 |
| 2019 | Curvilinear Distance Metric LearningabstractDistance Metric Learning aims to learn an appropriate metric that faithfully measures the distance between two data points. Traditional metric learning methods usually calculate the pairwise distance with fixed distance functions (\emph{e.g.,}\ Euclidean distance) in the projected feature spaces. However, they fail to learn the underlying geometries of the sample space, and thus cannot exactly predict the intrinsic distances between data points. To address this issue, we first reveal that the traditional linear distance metric is equivalent to the cumulative arc length between the data pair's nearest points on the learned straight measurer lines. After that, by extending such straight lines to general curved forms, we propose a Curvilinear Distance Metric Learning (CDML) method, which adaptively learns the nonlinear geometries of the training data. By virtue of Weierstrass theorem, the proposed CDML is equivalently parameterized with a 3-order tensor, and the optimization algorithm is designed to learn the tensor parameter. Theoretical analysis is derived to guarantee the effectiveness and soundness of CDML. Extensive experiments on the synthetic and real-world datasets validate the superiority of our method over the state-of-the-art metric learning models. Shuo Chen 0003, Lei Luo 0001, Jian Yang 0003, Chen Gong 0002, Jun Li 0027, Heng Huang 0001 |
NeurIPS | 5 |
| 2019 | Hierarchical Tracking by Reinforcement Learning-Based Searching and Coarse-to-Fine VerifyingabstractA class-agnostic tracker typically consists of three key components, i.e., its motion model, its target appearance model, and its updating strategy. However, most recent topperforming trackers mainly focus on constructing complicated appearance models and updating strategies, while using comparatively simple and heuristic motion models that may result in an inefficient search and degrade the tracking performance. To address this issue, we propose a hierarchical tracker that learns to move and track based on the combination of data-driven search at the coarse level, and coarse-to-fine verification at the fine level. At the coarse level, a data-driven motion model learned from deep recurrent reinforcement learning provides our tracker with coarse localization of an object. By formulating motion search as an action-decision problem in reinforcement learning, our tracker utilizes a recurrent convolutional neural network based deep Q-network to effectively learn data-driven searching policies. The learned motion model cannot only significantly reduce the search space, but also provide more reliable interested regions for further verifying. At the fine level, a kernelized correlation filter (KCF) based appearance model is adopted to densely yet efficiently verify a local region centered on the predicted location from the motion model. Through using of circulant matrices and fast Fourier transformation, a large number of candidate samples in the local region can be efficiently and effectively evaluated by the KCF based appearance model. Finally, a simple yet robust estimator is designed to analyze possible tracking failure. The experiments on OTB50 and OTB100 illustrate that our tracker achieves better performance than the state-of-the-art trackers. Bineng Zhong 0001, Jun Li 0027, Yulun Zhang 0001, Yun Fu 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Alignment Network Based Multi-Person Tracking With Occlusion and Motion ReasoningabstractTracking-by-detection is one of the typical paradigms for multi-person tracking, due to the availability of automatic pedestrian detectors. However, existing multi-person trackers are greatly challenged by misalignment in the pedestrian detectors (i.e., excessive background and part missing) and occlusion. To effectively handle these problems, we propose a deep alignment network-based multi-person tracking method with occlusion and motion reasoning. Specifically, the inaccurate detections are first corrected via a deep alignment network, in which an alignment estimation module is used to automatically learn the spatial transformation of these detections. As a result, the deep features from our alignment network will have better representation power and, thus, lead to more consistent tracks. Then, a coarse-to-fine schema is designed for construing a discriminative association cost matrix with spatial, motion, and appearance information. Meanwhile, a principled approach is developed to allow our method to handle occlusion with motion reasoning and the reidentification ability of the pedestrian alignment network. Finally, the association problem is solved via a simple yet real-time Hungarian algorithm. Comprehensive experiments on MOT16, ISSIA soccer, PETS09, and TUD datasets validate the effectiveness and robustness of our proposed tracker. Qinqin Zhou 0001, Bineng Zhong 0001, Yulun Zhang 0001, Jun Li 0027, Yun Fu 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Predictive Coding Machine for Compressed Sensing and Image DenoisingabstractSparse and low rank coding has widely received much attention in machine learning, multimedia and computer vision. Unfortunately, expensive inference restricts the power of coding models in real-world applications, e.g., compressed sensing and image deblurring. In order to avoid the expensive inference, we propose a predictive coding machine (PCM) which aims to train a deep neural network (DNN) encoder to approximate the codes. By this means, a test sample can be fast approximated by the well-trained DNN. However, DNN leads PCM to be a non-convex and non-smooth optimization problem, which is extremely hard to solve. To address this challenge, we extend accelerated proximal gradient for PCM by steering gradient descent of DNN. To the best of our knowledge, we are the first to propose a gradient descent algorithm guided by accelerated proximal gradient for solving the PCM problem. Besides, a sufficient condition is provided to ensure the convergence to a critical point. Moreover, when the coding models are convex in PCM, the convergence rate O(1/(m2√t)) can be held in which m is the iteration number of accelerated proximal gradient, and t is the epoch of training DNN. Numerical results verify the promising advantages of PCM in terms of effectiveness, efficiency and robustness. Jun Li 0027, Hongfu Liu 0001, Yun Fu 0001 |
AAAI | 1 |
| 2018 | Adversarial Metric LearningabstractIn the past decades, intensive efforts have been put to design various loss functions and metric forms for metric learning problem. These improvements have shown promising results when the test data is similar to the training data. However, the trained models often fail to produce reliable distances on the ambiguous test pairs due to the different samplings between training set and test set. To address this problem, the Adversarial Metric Learning (AML) is proposed in this paper, which automatically generates adversarial pairs to remedy the sampling bias and facilitate robust metric learning. Specifically, AML consists of two adversarial stages, i.e. confusion and distinguishment. In confusion stage, the ambiguous but critical adversarial data pairs are adaptively generated to mislead the learned metric. In distinguishment stage, a metric is exhaustively learned to try its best to distinguish both adversarial pairs and original training pairs. Thanks to the challenges posed by the confusion stage in such competing process, the AML model is able to grasp plentiful difficult knowledge that has not been contained by the original training pairs, so the discriminability of AML can be significantly improved. The entire model is formulated into optimization framework, of which the global convergence is theoretically proved. The experimental results on toy data and practical datasets clearly demonstrate the superiority of AML to representative state-of-the-art metric learning models. Shuo Chen 0003, Chen Gong 0002, Jian Yang 0003, Xiang Li 0041, Yang Wei 0003, Jun Li 0027 |
IJCAI | 6 |
| 2018 | Improving face representation learning with center invariant loss
Yue Wu 0008, Hongfu Liu 0001, Jun Li 0027, Yun Fu 0001 |
Image Vis. Comput. | 3 |
| 2018 | Visual Representation and Classification by Learning Group Sparse Deep Stacking NetworkabstractDeep stacking networks (DSNs) have been successfully applied in classification tasks. Its architecture builds upon blocks of simplified neural network modules (SNNM). The hidden units are assumed to be independent in the SNNM module. However, this assumption prevents SNNM from learning the local dependencies between hidden units to better capture the information in the input data for the classification task. In addition, the hidden representations of input data in each class can be expectantly split into a group in real-world classification applications. Therefore, we propose two kinds of group sparse SNNM modules by mixing -norm and -norm. The first module learns the local dependencies among hidden units by dividing them into non-overlapping groups. The second module splits the representations of samples in different classes into separate groups to cluster the samples in each class. A group sparse DSN (GS-DSN) is constructed by stacking the group sparse SNNM modules. Experimental results further verify that our GS-DSN model outperforms the relevant classification methods. Particularly, GS-DSN achieves the state-of-the-art performance (99.1%) on 15-Scene. Jun Li 0027, Heyou Chang, Jian Yang 0003, Wei Luo 0006, Yun Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Convolutional Sparse Autoencoders for Image ClassificationabstractConvolutional sparse coding (CSC) can model local connections between image content and reduce the code redundancy when compared with patch-based sparse coding. However, CSC needs a complicated optimization procedure to infer the codes (i.e., feature maps). In this brief, we proposed a convolutional sparse auto-encoder (CSAE), which leverages the structure of the convolutional AE and incorporates the max-pooling to heuristically sparsify the feature maps for feature learning. Together with competition over feature channels, this simple sparsifying strategy makes the stochastic gradient descent algorithm work efficiently for the CSAE training; thus, no complicated optimization procedure is involved. We employed the features learned in the CSAE to initialize convolutional neural networks for classification and achieved competitive results on benchmark data sets. In addition, by building connections between the CSAE and CSC, we proposed a strategy to construct local descriptors from the CSAE for classification. Experiments on Caltech-101 and Caltech-256 clearly demonstrated the effectiveness of the proposed method and verified the CSAE as a CSC model has the ability to explore connections between neighboring image content for classification tasks. Wei Luo 0006, Jun Li 0027, Jian Yang 0003, Wei Xu 0052, Jian Zhang 0025 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Sparse Subspace Clustering by Learning Approximation ℓ0 CodesabstractSubspace clustering has been widely applied to detect meaningful clusters in high-dimensional data spaces. A main challenge in subspace clustering is to quickly calculate a "good" affinity matrix. ℓ0, ℓ1, ℓ2 or nuclear norm regularization is used to construct the affinity matrix in many subspace clustering methods because of their theoretical guarantees and empirical success. However, they suffer from the following problems: (1) ℓ2 and nuclear norm regularization require very strong assumptions to guarantee a subspace-preserving affinity; (2) although ℓ1 regularization can be guaranteed to give a subspace-preserving affinity under certain conditions, it needs more time to solve a large-scale convex optimization problem; (3) ℓ0 regularization can yield a tradeoff between computationally efficient and subspace-preserving affinity by using the orthogonal matching pursuit (OMP) algorithm, but this still takes more time to search the solution in OMP when the number of data points is large. In order to overcome these problems, we first propose a learned OMP (LOMP) algorithm to learn a single hidden neural network (SHNN) to fast approximate the ℓ0code. We then exploit a sparse subspace clustering method based on ℓ0 code which is fast computed by SHNN. Two sufficient conditions are presented to guarantee that our method can give a subspace-preserving affinity. Experiments on handwritten digit and face clustering show that our method not only quickly computes the ℓ0 code, but also outperforms the relevant subspace clustering methods in clustering results. In particular, our method achieves the state-of-the-art clustering accuracy (94.32%) on MNIST. Jun Li 0027, Yu Kong 0001, Yun Fu 0001 |
AAAI | 1 |
| 2017 | Projective Low-rank Subspace Clustering via Learning Deep EncoderabstractLow-rank subspace clustering (LRSC) has been considered as the state-of-the-art method on small datasets. LRSC constructs a desired similarity graph by low-rank representation (LRR), and employs a spectral clustering to segment the data samples. However, effectively applying LRSC into clustering big data becomes a challenge because both LRR and spectral clustering suffer from high computational cost. To address this challenge, we create a projective low-rank subspace clustering (PLrSC) scheme for large scale clustering problem. First, a small dataset is randomly sampled from big dataset. Second, our proposed predictive low-rank decomposition (PLD) is applied to train a deep encoder by using the small dataset, and the deep encoder is used to fast compute the low-rank representations of all data samples. Third, fast spectral clustering is employed to segment the representations. As a non-trivial contribution, we theoretically prove the deep encoder can universally approximate to the exact (or bounded) recovery of the row space. Experiments verify that our scheme outperforms the related methods on large scale datasets in a small amount of time. We achieve the state-of-art clustering accuracy by 95.8% on MNIST using scattering convolution features. Jun Li 0027, Hongfu Liu 0001, Handong Zhao, Yun Fu 0001 |
IJCAI | 1 |
| 2017 | Large-scale Subspace Clustering by Fast Regression CodingabstractLarge-Scale Subspace Clustering (LSSC) is an interesting and important problem in big data era. However, most existing methods (i.e., sparse or low-rank subspace clustering) cannot be directly used for solving LSSC because they suffer from the high time complexity-quadratic or cubic in n (the number of data points). To overcome this limitation, we propose a Fast Regression Coding (FRC) to optimize regression codes, and simultaneously train a non-linear function to approximate the codes. By using FRC, we develop an efficient Regression Coding Clustering (RCC) framework to solve the LSSC problem. It consists of sampling, FRC and clustering. RCC randomly samples a small number of data points, quickly calculates the codes of all data points by using the non-linear function learned from FRC, and employs a large-scale spectral clustering method to cluster the codes. Besides, we provide a theorem guarantee that the non-linear function has a first-order approximation ability and a group effect. The theorem manifests that the codes are easily used to construct a dividable similarity graph. Compared with the state-of-the-art LSSC methods, our model achieves better clustering results in large-scale datasets. Jun Li 0027, Handong Zhao, Zhiqiang Tao, Yun Fu 0001 |
IJCAI | 1 |
| 2017 | Deeply Learned View-Invariant Features for Cross-View Action RecognitionabstractClassifying human actions from varied views is challenging due to huge data variations in different views. The key to this problem is to learn discriminative view-invariant features robust to view variations. In this paper, we address this problem by learning view-specific and view-shared features using novel deep models. View-specific features capture unique dynamics of each view while view-shared features encode common patterns across views. A novel sample-affinity matrix is introduced in learning shared features, which accurately balances information transfer within the samples from multiple views and limits the transfer across samples. This allows us to learn more discriminative shared features robust to view variations. In addition, the incoherence between the two types of features is encouraged to reduce information redundancy and exploit discriminative information in them separately. The discriminative power of the learned features is further improved by encouraging features in the same categories to be geometrically closer. Robust view-invariant features are finally learned by stacking several layers of features. Experimental results on three multi-view data sets show that our approaches outperform the state-of-the-art approaches. Yu Kong 0001, Zhengming Ding, Jun Li 0027, Yun Fu 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Sparseness Analysis in the Pretraining of Deep Neural NetworksabstractA major progress in deep multilayer neural networks (DNNs) is the invention of various unsupervised pretraining methods to initialize network parameters which lead to good prediction accuracy. This paper presents the sparseness analysis on the hidden unit in the pretraining process. In particular, we use the$L_{1}$-norm to measure sparseness and provide some sufficient conditions for that pretraining leads to sparseness with respect to the popular pretraining models—such as denoising autoencoders (DAEs) and restricted Boltzmann machines (RBMs). Our experimental results demonstrate that when the sufficient conditions are satisfied, the pretraining models lead to sparseness. Our experiments also reveal that when using the sigmoid activation functions, pretraining plays an important sparseness role in DNNs with sigmoid (Dsigm), and when using the rectifier linear unit (ReLU) activation functions, pretraining becomes less effective for DNNs with ReLU (Drelu). Luckily, Drelu can reach a higher recognition accuracy than DNNs with pretraining (DAEs and RBMs), as it can capture the main benefit (such as sparseness-encouraging) of pretraining in Dsigm. However, ReLU is not adapted to the different firing rates in biological neurons, because the firing rate actually changes along with the varying membrane resistances. To address this problem, we further propose a family of rectifier piecewise linear units (RePLUs) to fit the different firing rates. The experimental results show that the performance of RePLU is better than ReLU, and is comparable with those with some pretraining techniques, such as RBMs and DAEs. Jun Li 0027, Tong Zhang 0001, Wei Luo 0006, Jian Yang 0003, Xiao-Tong Yuan, Jian Zhang 0025 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Object segmentation using low-rank representation with multiple block-diagonal priorsabstractThis paper addresses the problem of segmenting objects for natural images by leveraging multiple segmentation methods. Existing image segmentation algorithms mostly partition the image into some coherent segments instead of extracting the object entirely. We observe that basic elements (e.g., superpixels) in the common segment produced by many methods are highly-correlated - they generally belong to the same object and lead to a low-rank structure. To make use of this information, this paper presents a novel approach to learn a low-rank affinity by leveraging various algorithms for segmenting the objects. First, starting from the segments produced using various algorithms, the label information for basic elements of an image, segmentation-to-superpixel label, is generated. Second, the derived labels are used to construct the multiple block-diagonal priors which are integrated into the subsequent affinity learning process. Finally the segmentation result is achieved by applying the spectral clustering technique on the obtained affinity matrix. Comprehensive experiments on MSRC, Alpert's and Berkeley segmentation datasets validate that our proposed approach achieves superior results as compared to each individual method. Moreover, our method is shown to be competitive in comparison to the state-of-the-art methods. Lingzheng Dai, Jundi Ding, Jun Li 0027, Jian Yang 0003 |
ICPR | 4 |
| 2016 | Super Resolution of the Partial Pixelated Images With Deep Convolutional Neural NetworkabstractThe problem of super resolution of partial pixelated images is considered in this paper. Partial pixelated images are more and more common in nowadays due to public safety etc. However, in some special cases, for instance criminal investigation, some images are pixelated intentionally by criminals and partial pixelate make it hard to reconstruct images even a higher resolution images. Hence, a method is proposed to handle this problem based on the deep convolutional neural network, termed depixelate super resolution CNN(DSRCNN). Given the mathematical expression pixelates, we propose a model to reconstruct the image from the pixelation and map to a higher resolution by combining the adversarial autoencoder with two depixelate layers. This model is evaluated on standard public datasets in which images are pixelated randomly and compared to the state of arts methods, shows very exciting performance. Haiyi Mao, Yue Wu 0008, Jun Li 0027, Yun Fu 0001 |
ACM Multimedia | 3 |
| 2016 | Deep Convolutional Neural Network with Independent Softmax for Large Scale Face RecognitionabstractIn this paper, we present our solution to the MS-Celeb-1M Challenge. This challenge aims to recognize 100k celebrities at the same time. The huge number of celebrities is the bottleneck for training a deep convolutional neural network of which the output is equal to the number of celebrities. To solve this problem, an independent softmax model is proposed to split the single classifier into several small classifiers. Meanwhile, the training data are split into several partitions. This decomposes the large scale training procedure into several medium training procedures which can be solved separately. Besides, a large model is also trained and a simple strategy is introduced to merge the two models. Extensive experiments on the MSR-Celeb-1M dataset demonstrate the superiority of the proposed method. Our solution ranks the first and second in two tracks of the final evaluation. Yue Wu 0008, Jun Li 0027, Yu Kong 0001, Yun Fu 0001 |
ACM Multimedia | 2 |
| 2016 | Learning Fast Low-Rank Projection for Image ClassificationabstractRooted in a basic hypothesis that a data matrix is strictly drawn from some independent subspaces, the low-rank representation (LRR) model and its variations have been successfully applied in various image classification tasks. However, this hypothesis is very strict to the LRR model as it cannot always be guaranteed in real images. Moreover, the hypothesis also prevents the sub-dictionaries of different subspaces from collaboratively representing an image. Fortunately, in supervised image classification, low-rank signal can be extracted from the independent label subspaces (ILS) instead of the independent image subspaces (IIS). Therefore, this paper proposes a projective low-rank representation (PLR) model by directly training a projective function to approximate the LRR derived from the labels. To the best of our knowledge, PLR is the first attempt to use the ILS hypothesis to relax the rigorous IIS hypothesis in the LRR models. We further prove a low-rank effect that the representations learned by PLR have high intraclass similarities and large interclass differences, which are beneficial to the classification tasks. The effectiveness of our proposed approach is validated by the experimental results on three databases. Jun Li 0027, Yu Kong 0001, Handong Zhao, Jian Yang 0003, Yun Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Sparse Deep Stacking Network for Image ClassificationabstractSparse coding can learn good robust representation to noise and model more higher-order representation for image classification. However, the inference algorithm is computationally expensive even though the supervised signals are used to learn compact and discriminative dictionaries in sparse coding techniques. Luckily, a simplified neural network module (SNNM) has been proposed to directly learn the discriminative dictionaries for avoiding the expensive inference. But the SNNM module ignores the sparse representations. Therefore, we propose a sparse SNNM module by adding the mixed-norm regularization (l1/l2 norm). The sparse SNNM modules are further stacked to build a sparse deep stacking network (S-DSN). In the experiments, we evaluate S-DSN with four databases, including Extended YaleB, AR, 15 scene and Caltech101. Experimental results show that our model outperforms related classification methods with only a linear classifier. It is worth noting that we reach 98.8% recognition accuracy on 15 scene. Jun Li 0027, Heyou Chang, Jian Yang 0003 |
AAAI | 1 |
| 2015 | Higher-level feature combination via multiple kernel learning for image classification
Wei Luo 0006, Jian Yang 0003, Wei Xu 0052, Jun Li 0027, Jian Zhang 0025 |
Neurocomputing | 4 |
| 2014 | Learning discriminative low-rank representation for image classificationabstractLow-rank representation (LRR) efficiently performs the subspace segmentation and feature extraction from corrupted data. However, there are three disadvantages in existing LRR techniques. First, the inference algorithm of LRR (as a generative model) is computationally expensive. Second, LRR ignores the discriminative information for image classification. Third, although the robust representation is implemented by recovering the low-rank components and the sparse noises, it has been limited due to the constrained assumption that noises is sparse. To solve these problems, and inspired by Denoising Autoencoders (DAE) and Contractive Autoencoders (CAE), this paper proposes a discriminative low-rank representations framework (DLRR) for image classification. We directly learn a discriminative projection dictionary that results in fast inference. Simultaneously, DLRR can obtain a robust representation from any corrupted input. Our implementation of DLRR achieves state-of-the-art results on artificial dataset and dataset of Olivetti Face Patches. Jun Li 0027, Heyou Chang, Jian Yang 0003 |
IJCNN | 1 |
| 2014 | Continuous attractors of higher-order recurrent neural networks with infinite neurons
Jun Li 0027, Jian Yang 0003, Xiao-Tong Yuan, Zhaohua Hu |
Neurocomputing | 1 |
| 2012 | Continuous attractors of recurrent neural networks with complex-valued weightsabstractThe global exponential stability (GAS), global asymptotic stability (GES) and multi-stability (MS) continuous attractors of recurrent neural networks (RNN) with complex-valued weights are studied in this paper. As a continuous attractor is the infinite equilibria, the connected matrix needs to be nonsingular. Therefore, RNN is transformed into a lower dimensional RNN using elementary operation. Firstly, based on the foregoing results some continuous attractors of RNN with real-valued weights are obtained. Secondly, continuous attractors of RNN with complex-valued weights are obtained by studying the corresponding RNN with real-valued weights. Some simulations are finally carried out to illustrate the theory. Jun Li 0027, Jian Yang 0003, Yongfeng Diao |
IJCNN | 1 |
| 2012 | Stability and periodicity of discrete Hopfield neural networks with column arbitrary-magnitude-dominant weight matrix
Jun Li 0027, Jian Yang 0003, Weigen Wu |
Neurocomputing | 1 |