VLDB 2026 Research / reviewers in the wild / expert
Yinan Yu
dblp:49/8066
· DBLP profile ↗
33ranked-venue papers
11as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-authorSoftware engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Recognizing as-built building materials: a systematic mapping studyabstract• First systematic mapping study focused on as-built building material recognition. • Several data modalities and formats are applicable to recognizing building materials. • Multi-modal data with foundation models is an emerging approach. • A collection of multi-modal material recognition public datasets has been identified. Recognising materials in existing, as-built buildings is essential for analysing safety, energy performance, material reclamation, amongst many other use cases. As-built building material recognition is challenging because building materials often mimic others, and controlled laboratory conditions are not practical. This systematic mapping study is the first to focus on as-built building material recognition, consolidating fragmented research across disciplines to reveal data modalities, recognition techniques, research gaps, and key future directions. After reviewing over 20,000 documents, several key insights were identified: 1) many studies lack contextual information necessary to assess generalization; 2) visible light, hyperspectral, infrared, radiowave, tactile, audio, electric field, and text are potential data modalities; 3) studies are needed beyond the neighborhood scale; 4) experiments optimized for ease-of-use pair images with foundation models; 5) highly cited experiments use multi-modal data with foundation models. Additionally, a structured overview of public datasets is provided. This study supports researchers by establishing effective and scalable methods for as-built building material recognition applicable in many use cases. Josie Harrison, Yinan Yu, Alexander Hollberg |
Expert Syst. Appl. | 2 |
| 2026 | GateLens: A reasoning-enhanced LLM agent for automotive software release analyticsabstractEnsuring reliable data-driven decisions is crucial in domains where analytical accuracy directly impacts safety, compliance, or operational outcomes. Decision support in such domains relies on large tabular datasets, where manual analysis is slow, costly, and error-prone. While Large Language Models (LLMs) offer promising automation potential, they face challenges in analytical reasoning, structured data handling, and ambiguity resolution. This paper introduces GateLens, an LLM-based architecture for reliable analysis of complex tabular data. Its key innovation is the use of Relational Algebra (RA) as a formal intermediate representation between natural-language reasoning and executable code, addressing the reasoning-to-code gap that can arise in direct generation approaches. In our automotive instantiation, GateLens translates natural language queries into RA expressions and generates optimized Python code. Unlike traditional multi-agent or planning-based systems that can be slow, opaque, and costly to maintain, GateLens emphasizes speed, transparency, and reliability. We validate the architecture in automotive software release analytics, where experimental results show that GateLens outperforms the existing Chain-of-Thought (CoT) + Self-Consistency (SC) based system on real-world datasets, particularly in handling complex and ambiguous queries. Ablation studies confirm the essential role of the RA layer. Industrial deployment demonstrates over 80% reduction in analysis time while maintaining high accuracy across domain-specific tasks. GateLens operates effectively in zero-shot settings without requiring few-shot examples or agent orchestration. This work advances deployable LLM system design by identifying key architectural features—intermediate formal representations, execution efficiency, and low configuration overhead—crucial for domain-specific analytical applications where accuracy, traceability, and stakeholder trust are paramount. Arsham Gholamzadeh Khoee, Robert Feldt, Dhasarathy Parthasarathy, Yinan Yu |
J. Syst. Softw. | 5 |
| 2025 | iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question AnsweringabstractWhile Large Language Models (LLMs) excel at many natural language processing tasks, they often suffer from factual inaccuracies in knowledge-intensive scenarios.Integrating external knowledge resources, particularly knowledge graphs (KGs), provides a transparent and updatable foundation for more reliable reasoning.Knowledge Base Question Answering (KBQA), which queries and reasons over KGs, is central to this effort, especially for complex, multi-hop queries.However, multi-hop reasoning poses two key challenges: (1) maintaining coherent reasoning paths, and (2) avoiding prematurely discarding critical multi-hop connections.To address these issues, we introduce iQUEST, a question-guided KBQA framework that iteratively decomposes complex queries into simpler sub-questions, ensuring a structured and focused reasoning trajectory.Additionally, we integrate a Graph Neural Network (GNN) to look ahead and incorporate 2hop neighbor information at each reasoning step.This dual approach strengthens the reasoning process, enabling the model to explore viable paths more effectively.Detailed experiments demonstrate the consistent improvement delivered by iQUEST across four benchmark datasets and four LLMs. Yinan Yu |
ACL (1) | 2 |
| 2025 | Automating a Complete Software Test Process Using LLMs: An Automotive Case StudyabstractVehicle API testing verifies whether the interactions between a vehicle's internal systems and external applications meet expectations, ensuring that users can access and control various vehicle functions and data. However, this task is inherently complex, requiring the alignment and coordination of API systems, communication protocols, and even vehicle simulation systems to develop valid test cases. In practical industrial scenarios, inconsistencies, ambiguities, and interde-pendencies across various documents and system specifications pose significant challenges. This paper presents a system designed for the automated testing of in-vehicle APIs. By clearly defining and segmenting the testing process, we enable Large Language Models (LLMs) to focus on specific tasks, ensuring a stable and controlled testing workflow. Experiments conducted on over 100 APIs demonstrate that our system effectively automates vehicle API testing. The results also confirm that LLMs can efficiently handle mundane tasks requiring human judgment, making them suitable for complete automation in similar industrial contexts. Yinan Yu, Robert Feldt, Dhasarathy Parthasarathy |
ICSE | 2 |
| 2024 | Semantic-Aware Representation of Multi-Modal Data for Data Ingress: A Literature ReviewabstractMachine Learning (ML) is continuously permeating a growing amount of application domains. Generative AI such as Large Language Models (LLMs) also sees broad adoption to process multi-modal data such as text, images, audio, and video. While the trend is to use ever-larger datasets for training, managing this data efficiently has become a significant practical challenge in the industry-double as much data is certainly not double as good. Rather the opposite is important since getting an understanding of the inherent quality and diversity of the underlying data lakes is a growing challenge for application-specific ML as well as for fine-tuning foundation models. Furthermore, information retrieval (IR) from expanding data lakes is complicated by the temporal dimension inherent in time-series data which must be considered to determine its semantic value. This study focuses on the different semantic-aware techniques to extract embeddings from mono-modal, multi-modal, and cross-modal data to enhance IR capabilities in a growing data lake. Articles were collected to summarize information about the state-of-the-art techniques focusing on applications of embedding for three different categories of data modalities. Pierre Lamart, Yinan Yu, Christian Berger 0001 |
SEAA | 2 |
| 2024 | LLMs Can Check Their Own Results to Mitigate Hallucinations in Traffic Understanding Tasks
Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger 0001 |
ICTSS | 3 |
| 2024 | GoNoGo: An Efficient LLM-Based Multi-agent System for Streamlining Automotive Software Release Decision-Making
Arsham Gholamzadeh Khoee, Yinan Yu, Robert Feldt, Andris Freimanis, Patrick Andersson Rhodin, Dhasarathy Parthasarathy |
ICTSS | 2 |
| 2024 | climateBUG : A data-driven framework for analyzing bank reporting through a climate lensabstractThis paper applies computational linguistics learning methods to the banking industry and climate change fields. We introduce our data-driven framework, climateBUG, with the aim of detecting latent information about how banks discuss their activities related to climate change using natural language processing (NLP). This framework consists of an ingestion pipeline, a configurable database, and a set of API’s. In addition, climateBUG offers two standalone components, namely a unique annotated corpus of approximately 1.1M statements from EU banks’ annual and sustainability reporting and a deep learning model adapted to the semantics of the corpus. When benchmarking on classification performance, our model outperforms other models with similar scopes due to its stronger domain relevance. We also provide examples of how the framework can be applied from a user perspective. Yinan Yu, Samuel Scheidegger, Jasmine Elliott, Åsa Löfgren |
Expert Syst. Appl. | 1 |
| 2024 | Building efficient CNNs using Depthwise Convolutional Eigen-Filters (DeCEF)abstractDeep Convolutional Neural Networks (CNNs) have been widely used in various domains due to their impressive capabilities. These models are typically composed of a large number of 2D convolutional (Conv2D) layers with numerous trainable parameters. To manage the complexity of such networks, compression techniques can be applied, which typically rely on the analysis of trained deep learning models. However, in certain situations, training a new CNN from scratch may be infeasible due to resource limitations. In this paper, we propose an alternative parameterization to Conv2D filters with significantly fewer parameters without relying on compressing a pre-trained CNN. Our analysis reveals that the effective rank of the vectorized Conv2D filters decreases with respect to the increasing depth in the network. This leads to the development of the Depthwise Convolutional Eigen-Filter (DeCEF) layer, which is a low rank version of the Conv2D layer with significantly fewer trainable parameters and floating point operations (FLOPs). The way we define the effective rank is different from previous work, and it is easy to implement and interpret. Applying this technique is straightforward – one can simply replace any standard convolutional layer with a DeCEF layer in a CNN. To evaluate the effectiveness of DeCEF layers, experiments are conducted on the benchmark datasets CIFAR-10 and ImageNet for various network architectures. The results have shown a similar or higher accuracy using about 2/3 of the original parameters and reducing the number of FLOPs to 2/3 of the base network. Additionally, analyzing the patterns in the effective rank provides insights into the inner workings of CNNs and highlights opportunities for future research. Yinan Yu, Samuel Scheidegger, Tomas McKelvey |
Neurocomputing | 1 |
| 2023 | AirDnD - Asynchronous In-Range Dynamic and Distributed Network Orchestration FrameworkabstractThe increasing usage of IoT devices has generated an extensive volume of data which resulted in the establishment of data centers with well-structured computing infrastructure. Reducing underutilized resources of such data centers can be achieved by monitoring the tasks and offloading them across various compute units. This approach can also be used in mini mobile data ponds generated by edge devices and smart vehicles. This research aims to improve and utilize the usage of computing resources in distributed edge devices by forming a dynamic mesh network. The nodes in the mesh network shall share their computing tasks with another node that possesses unused computing resources. This proposed method ensures the minimization of data transfer between entities. The proposed AirDnD vision will be applied to a practical scenario relevant to an autonomous vehicle that approaches an intersection commonly known as “looking around the corner” in related literature, collecting essential computational results from nearby vehicles to enhance its perception. The proposed solution consists of three models that transform growing amounts of geographically distributed edge devices into a living organism. Malsha Ashani Mahawatta Dona, Christian Berger 0001, Yinan Yu |
ICDCS | 3 |
| 2022 | The causal effect of subscription video streaming on DVD sales: Evidence from a natural experiment
Yinan Yu, Chih-Hung Peng, Patrick Y. K. Chau |
Decis. Support Syst. | 1 |
| 2020 | Capacity Region and Scheduling for Non-Orthogonal DuplexabstractExisting wireless mobile networks configure the uplink (UL) and downlink (DL) resources based on the orthogonal duplex principle, which has significantly evolved from the static frequency/time division duplex (FDD/TDD) to the semi-dynamic TDD in 4G and the fully-dynamic TDD in 5G. Fueled by the successful realizations of full duplex (FD) transmission, non-orthogonal duplex (NOD) emerges as an attractive candidate technique for future networks to improve the bidirectional throughput. In this paper, we investigate the performance limit of the NOD scheme, which imports the FD mode on the basis of fully-dynamic TDD. We first focus on a single-cell scenario by assuming no cooperation among cells, and strive to characterize the bidirectional capacity region by finding the capacity-optimal transmission mode scheduling policy. We obtain the optimal policies for the systems with and without modulation and coding scheme (MCS), respectively, where the policies are based on the instantaneous channel gains and the distribution of channels. We proceed to develop a scheduling policy that only relies on the instantaneous channel and queue information, and extend it to the multicell multiuser scenario with coordinated scheduling. Numerical and simulation results demonstrate the great potential of the NOD scheme in increasing bidirectional capacity and reducing the queuing delay. Shengqian Han, Yinan Yu, Juan Liu 0013, Xiaolin Hou, Wenjia Liu |
IEEE Trans. Commun. | 3 |
| 2019 | High-Level Semantic Feature Detection: A New Perspective for Pedestrian DetectionabstractObject detection generally requires sliding-window classifiers in tradition or anchor-based predictions in modern deep learning approaches. However, either of these approaches requires tedious configurations in windows or anchors. In this paper, taking pedestrian detection as an example, we provide a new perspective where detecting objects is motivated as a high-level semantic feature detection task. Like edges, corners, blobs and other feature detectors, the proposed detector scans for feature points all over the image, for which the convolution is naturally suited. However, unlike these traditional low-level features, the proposed detector goes for a higher-level abstraction, that is, we are looking for central points where there are pedestrians, and modern deep models are already capable of such a high-level semantic abstraction. Besides, like blob detection, we also predict the scales of the pedestrian points, which is also a straightforward convolution. Therefore, in this paper, pedestrian detection is simplified as a straightforward center and scale prediction task through convolutions. This way, the proposed method enjoys an anchor-free setting. Though structurally simple, it presents competitive accuracy and good speed on challenging pedestrian detection benchmarks, and hence leading to a new attractive pedestrian detector. Code and models will be available at https://github.com/liuwei16/CSP. Wei Liu 0097, Shengcai Liao, Weiqiang Ren, Weidong Hu, Yinan Yu |
CVPR | 5 |
| 2019 | SSAP: Single-Shot Instance Segmentation With Affinity PyramidabstractRecently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. We argue that treating these two sub-tasks separately is suboptimal. In fact, employing multiple separate modules significantly reduces the potential for application. The mutual benefits between the two complementary sub-tasks are also unexplored. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on a pixel-pair affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. The affinity pyramid can also be jointly learned with the semantic class labeling and achieve mutual benefits. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition module is presented to sequentially generate instances from coarse to fine. Unlike previous time-consuming graph partition methods, this module achieves 5× speedup and 9% relative improvement on Average-Precision (AP). Our approach achieves new state of the art on the challenging Cityscapes dataset. Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Yinan Yu, Ming Yang 0007, Kaiqi Huang |
ICCV | 5 |
| 2018 | CLAss-Specific Subspace Kernel Representations and Adaptive Margin Slack Minimization for Large Scale ClassificationabstractIn kernel-based classification models, given limited computational power and storage capacity, operations over the full kernel matrix becomes prohibitive. In this paper, we propose a new supervised learning framework using kernel models for sequential data processing. The framework is based on two components that both aim at enhancing the classification capability with a subset selection scheme. The first part is a subspace projection technique in the reproducing kernel Hilbert space using a CLAss-specific Subspace Kernel representation for kernel approximation. In the second part, we propose a novel structural risk minimization algorithm called the adaptive margin slack minimization to iteratively improve the classification accuracy by an adaptive data selection. We motivate each part separately, and then integrate them into learning frameworks for large scale data. We propose two such frameworks: the memory efficient sequential processing for sequential data processing and the parallelized sequential processing for distributed computing with sequential data acquisition. We test our methods on several benchmark data sets and compared with the state-of-the-art techniques to verify the validity of the proposed techniques. Yinan Yu, Konstantinos I. Diamantaras, Tomas McKelvey, Sun-Yuan Kung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Parse geometry from a line: Monocular depth estimation with partial laser observationabstractMany standard robotic platforms are equipped with at least a fixed 2D laser range finder and a monocular camera. Although those platforms do not have sensors for 3D depth sensing capability, knowledge of depth is an essential part in many robotics activities. Therefore, recently, there is an increasing interest in depth estimation using monocular images. As this task is inherently ambiguous, the data-driven estimated depth might be unreliable in robotics applications. In this paper, we have attempted to improve the precision of monocular depth estimation by introducing 2D planar observation from the remaining laser range finder without extra cost. Specifically, we construct a dense reference map from the sparse laser range data, redefining the depth estimation task as estimating the distance between the real and the reference depth. To solve the problem, we construct a novel residual of residual neural network, and tightly combine the classification and regression losses for continuous depth estimation. Experimental results suggest that our method achieves considerable promotion compared to the state-of-the-art methods on both NYUD2 and KITTI, validating the effectiveness of our method on leveraging the additional sensory information. We further demonstrate the potential usage of our method in obstacle avoidance where our methodology provides comprehensive depth information compared to the solution using monocular camera or 2D laser range finder alone. Yiyi Liao, Lichao Huang, Yue Wang 0020, Sarath Kodagoda, Yinan Yu, Yong Liu 0007 |
ICRA | 5 |
| 2016 | Adaptive margin slack minimization in RKHS for classificationabstractIn this paper, we design a novel regularized empirical risk minimization technique for classification called Adaptive Margin Slack Minimization (AMSM). The proposed method is based on minimizing a regularized upper bound of the misclassification error. Compared to the cost function of the classical L2-SVM, AMSM can be interpreted as minimizing a tighter bound with some additional flexibilities regarding the choice of marginal hyperplane. A hyperparameter-free adaptive algorithm is presented for finding a solution to the proposed risk function. Numerical results shows that AMSM outperforms L2-SVM on the tested standard datasets. Yinan Yu, Konstantinos I. Diamantaras, Tomas McKelvey, Sun-Yuan Kung |
ICASSP | 1 |
| 2015 | Deep multiple instance learning for image classification and auto-annotationabstractThe recent development in learning deep representations has demonstrated its wide applications in traditional vision tasks like classification and detection. However, there has been little investigation on how we could build up a deep learning framework in a weakly supervised setting. In this paper, we attempt to model deep learning in a weakly supervised learning (multiple instance learning) framework. In our setting, each image follows a dual multi-instance assumption, where its object proposals and possible text annotations can be regarded as two instance sets. We thus design effective systems to exploit the MIL property with deep learning strategies from the two ends; we also try to jointly learn the relationship between object and annotation proposals. We conduct extensive experiments and prove that our weakly supervised deep learning framework not only achieves convincing performance in vision tasks including classification and image annotation, but also extracts reasonable region-keyword pairs with little supervision, on both widely used benchmarks like PASCAL VOC and MIT Indoor Scene 67, and also a dataset for image-and patch-level annotations. Jiajun Wu 0001, Yinan Yu, Chang Huang |
CVPR | 2 |
| 2015 | Object detection by labeling superpixelsabstractObject detection is often conducted by object proposal generation and classification sequentially. This paper handles object detection in a superpixel oriented manner instead of the proposal oriented. Specially, this paper takes object detection as a multi-label superpixel labeling problem by minimizing an energy function. It uses the data cost term to capture the appearance, smooth cost term to encode the spatial context and label cost term to favor compact detection. The data cost is learned through a convolutional neural network and the parameters in the labeling model are learned through a structural SVM. Compared with proposal generation and classification based methods, the proposed superpixel labeling method can naturally detect objects missed by proposal generation step and capture the global image context to infer the overlapping objects. The proposed method shows its advantage in Pascal VOC and ImageNet. Notably, it performs better than the ImageNet ILSVRC2014 winner GoogLeNet (45.0% V.S. 43.9% in mAP) with much shallower and fewer CNNs. Yinan Yu, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
CVPR | 2 |
| 2015 | Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural NetworksabstractWhile feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to note that the human visual cortex generally contains more feedback than feedforward connections. In this paper, we will briefly introduce the background of feedbacks in the human visual cortex, which motivates us to develop a computational feedback mechanism in deep neural networks. In addition to the feedforward inference in traditional neural networks, a feedback loop is introduced to infer the activation status of hidden layer neurons according to the "goal" of the network, e.g., high-level semantic labels. We analogize this mechanism as "Look and Think Twice." The feedback networks help better visualize and understand how deep neural networks work, and capture visual attention on expected objects, even in images with cluttered background and multiple objects. Experiments on ImageNet dataset demonstrate its effectiveness in solving tasks such as image classification and object localization. Chunshui Cao, Xianming Liu 0005, Yi Yang 0007, Yinan Yu, Jiang Wang 0001, Zilei Wang, Yongzhen Huang, Liang Wang 0001, Chang Huang, Wei Xu 0017, Deva Ramanan, Thomas S. Huang |
ICCV | 4 |
| 2015 | A Deep Visual Correspondence Embedding Model for Stereo Matching CostsabstractThis paper presents a data-driven matching cost for stereo matching. A novel deep visual correspondence embedding model is trained via Convolutional Neural Network on a large set of stereo images with ground truth disparities. This deep embedding model leverages appearance data to learn visual similarity relationships between corresponding image patches, and explicitly maps intensity values into an embedding feature space to measure pixel dissimilarities. Experimental results on KITTI and Middlebury data sets demonstrate the effectiveness of our model. First, we prove that the new measure of pixel dissimilarity outperforms traditional matching costs. Furthermore, when integrated with a global stereo framework, our method ranks top 3 among all two-frame algorithms on the KITTI benchmark. Finally, cross-validation results show that our model is able to make correct predictions for unseen data which are outside of its labeled training set. Zhuoyuan Chen, Liang Wang 0001, Yinan Yu, Chang Huang |
ICCV | 4 |
| 2014 | Feature reduction based on Sum-Of-SNR (SOSNR) optimizationabstractDimensionality reduction plays an important role in machine learning techniques. In classification, data transformation aims to reduce the number of feature dimensions, whereas attempts to enhance the class separability. To this end, we propose a new classifier-independent criterion called “Sum-of-Signal-to-Noise-Ratio” (SoSNR). A framework designed for maximization with respect to this criterion is presented and three types of algorithms, respectively based on (1) gradient, (2) deflation and (3) sparsity, are proposed. The techniques are conducted on standard UCI databases and compared to other related methods. Results show trade-offs between computational complexity and classification accuracy among different approaches. Yinan Yu, Tomas McKelvey, Sun-Yuan Kung |
ICASSP | 1 |
| 2014 | Learning Convolutional Nonlinear Features for K Nearest Neighbor Image ClassificationabstractLearning low-dimensional feature representations is a crucial task in machine learning and computer vision. Recently the impressive breakthrough in general object recognition made by large scale convolutional networks shows that convolutional networks are able to extract discriminative hierarchical features in large scale object classification task. However, for vision tasks other than end-to-end classification, such as K Nearest Neighbor classification, the learned intermediate features are not necessary optimal for the specific problem. In this paper, we aim to exploit the power of deep convolutional networks and optimize the output feature layer with respect to the task of K Nearest Neighbor (kNN) classification. By directly optimizing the kNN classification error on training data, we in fact learn convolutional nonlinear features in a data-driven and task-driven way. Experimental results on standard image classification benchmarks show that the proposed method is able to learn better feature representations than other general end-to-end classification methods on kNN classification task. Weiqiang Ren, Yinan Yu, Junge Zhang, Kaiqi Huang |
ICPR | 2 |
| 2014 | Early Hierarchical Contexts Learned by Convolutional Networks for Image SegmentationabstractWe propose a foreground segmentation method based on convolutional networks. To predict the label of a pixel in an image, the model takes a hierarchical context as the input, which is obtained by combining multiple context patches on different scales. Short range contexts depict the local details, while long range contexts capture the object-scene relationships in an image. Early means that we combine the context patches of a pixel into a hierarchical one before any trainable layers are learned, i.e., early-combing. In contrast, late-combing means that the combination occurs later, e.g., when the convolutional feature extractor in a network has already been learned. We find that it is vital for the whole model to jointly learn the patterns of contexts on different scales in our task. Experiments show that early-combing performs better than late-combing. On the dataset1 built up by Baidu IDL2 for a latest person segmentation contest, our method beats all the competitors with a considerable margin. Qualitative results also show that the proposed method is almost ready for practical application. Zifeng Wu, Yongzhen Huang, Yinan Yu, Liang Wang 0001, Tieniu Tan |
ICPR | 3 |
| 2013 | A classification scheme for 'high-dimensional-small-sample-size' data using soda and ridge-SVM with microwave measurement applicationsabstractThe generalization performance of SVM-type classifiers severely suffers from the `curse of dimensionality'. For some real world applications, the dimensionality of the measurement is sometimes significantly larger compared to the amount of training data samples available. In this paper, a classification scheme is proposed and compared with existing techniques for such scenarios. The proposed scheme includes two parts: (i) feature selection and transformation based on Fisher discriminant criteria and (ii) a hybrid classifier combining Kernel Ridge Regression with Support Vector Machine to predict the label of the data. The first part is named Successively Orthogonal Discriminant Analysis (SODA), which is applied after Fisher score based feature selection as a preliminary processing for dimensionality reduction. At this step, SODA maximizes the ratio of between-class-scatter and within-class-scatter to obtain an orthogonal transformation matrix which maps the features to a new low dimensional feature space where the class separability is maximized. The techniques are tested on high dimensional data from a microwave measurements system and are compared with existing techniques. Yinan Yu, Tomas McKelvey, Sun-Yuan Kung |
ICASSP | 1 |
| 2013 | Kernel SODA: A Feature Reduction Technique Using Kernel Based AnalysisabstractA feature extraction technique called Successively Orthogonal Discriminant Analysis (SODA) has been recently proposed to overcome the limitation of Linear Discriminant Analysis (LDA), whose objective is to find a projection vector such that the projected values of data from both classes have maximum class separability. However, in LDA, only one such vector can be found due to the rank deficiency for binary classification problems. On the other hand, as a feature extraction technique, the proposed algorithm SODA attempts to obtain a transformation matrix instead of a vector. In this paper, the kernel version of SODA is presented in both intrinsic space and empirical space. To obtain the solution without sacrificing numerical efficiency, we propose a relaxed formulation and data selection for large scale computations. Simulations are conducted on 5 data sets from UCI database to verify and evaluated the new approach. Yinan Yu, Tomas McKelvey, Sun-Yuan Kung |
ICMLA (1) | 1 |
| 2013 | Exploring the Power of Kernel in Feature Representation for Object Categorization
Weiqiang Ren, Yinan Yu, Junge Zhang, Kaiqi Huang |
ICONIP (3) | 2 |
| 2012 | CLUMOC: Multiple Motion Estimation by Cluster Motion ConsensusabstractIn this paper, we present techniques for robust multiple motions estimation based on dual consensus via clustering in both the image spatial space and the motion parameter space. Starting from traditional Random Samples Consensus algorithm, we novelly propose the CLUster MOtion Consensus (CLUMOC) to extract robust motions. The proposed algorithm has two advantages: (1), instead of random samples, the CLUMOC employs clustering in initial sample selection, which can remove outliers from correct pairs of motion, (2), CLUMOC automatically decides the number of motions, by employing competition among motion and samples, that each motion needs to compete for matching pairs and each pair of matching competes for motions. The experimental results show that the proposed method is effective and efficient under various situations. Yinan Yu, Weiqiang Ren, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
AVSS | 1 |
| 2012 | Feature coding via vector difference for image classificationabstractAn effective image representation is important to an image classification task. The most popular image representation framework utilizes a feature coding algorithm to encode the extracted low-level feature descriptors into a vector representation. In this paper, we analyze the recently developed feature coding methods in a general way. According to their common characteristics, we propose a new coding scheme to perform feature coding based on the vector difference in a high-dimensional space which is obtained by explicit feature maps. As we illustrate, our method has promising results with small codebook sizes and generalizes most existing coding methods in a unified form. Xin Zhao 0012, Yinan Yu, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICIP | 2 |
| 2012 | A Novel Algorithm for View and Illumination Invariant Image MatchingabstractThe challenges in local-feature-based image matching are variations of view and illumination. Many methods have been recently proposed to address these problems by using invariant feature detectors and distinctive descriptors. However, the matching performance is still unstable and inaccurate, particularly when large variation in view or illumination occurs. In this paper, we propose a view and illumination invariant image-matching method. We iteratively estimate the relationship of the relative view and illumination of the images, transform the view of one image to the other, and normalize their illumination for accurate matching. Our method does not aim to increase the invariance of the detector but to improve the accuracy, stability, and reliability of the matching results. The performance of matching is significantly improved and is not affected by the changes of view and illumination in a valid range. The proposed method would fail when the initial view and illumination method fails, which gives us a new sight to evaluate the traditional detectors. We propose two novel indicators for detector evaluation, namely, valid angle and valid illumination, which reflect the maximum allowable change in view and illumination, respectively. Extensive experimental results show that our method improves the traditional detector significantly, even in large variations, and the two indicators are much more distinctive. Yinan Yu, Kaiqi Huang, Wei Chen 0012, Tieniu Tan |
IEEE Trans. Image Process. | 1 |
| 2011 | Salient coding for image classificationabstractThe codebook based (bag-of-words) model is a widely applied model for image classification. We analyze recent coding strategies in this model, and find that saliency is the fundamental characteristic of coding. The saliency in coding means that if a visual code is much closer to a descriptor than other codes, it will obtain a very strong response. The salient representation under maximum pooling operation leads to the state-of-the-art performance on many databases and competitions. However, most current coding schemes do not recognize the role of salient representation, so that they may lead to large deviations in representing local descriptors. In this paper, we propose “salient coding”, which employs the ratio between descriptors' nearest code and other codes to describe descriptors. This approach can guarantee salient representation without deviations. We study salient coding on two sets of image classification databases (15-Scenes and PASCAL VOC2007). The experimental results demonstrate that our approach outperforms all other coding methods in image classification. Yongzhen Huang, Kaiqi Huang, Yinan Yu, Tieniu Tan |
CVPR | 3 |
| 2011 | Boosted local structured HOG-LBP for object localizationabstractObject localization is a challenging problem due to variations in object's structure and illumination. Although existing part based models have achieved impressive progress in the past several years, their improvement is still limited by low-level feature representation. Therefore, this paper mainly studies the description of object structure from both feature level and topology level. Following the bottom-up paradigm, we propose a boosted Local Structured HOG-LBP based object detector. Firstly, at feature level, we propose Local Structured Descriptor to capture the object's local structure, and develop the descriptors from shape and texture information, respectively. Secondly, at topology level, we present a boosted feature selection and fusion scheme for part based object detector. All experiments are conducted on the challenging PASCAL VOC2007 datasets. Experimental results show that our method achieves the state-of-the-art performance. Junge Zhang, Kaiqi Huang, Yinan Yu, Tieniu Tan |
CVPR | 3 |
| 2009 | A Harris-Like Scale Invariant Feature Detector
Yinan Yu, Kaiqi Huang, Tieniu Tan |
ACCV (2) | 1 |