A Flexible Supervised Term-Weighting Technique and its Application to Variable Extraction and Information Retrieval

Mariano Maisonnave

Fernando Delbianco

Fernando Abel Tohmé

Ana Gabriela Maguitman

Resumen

Successful modeling and prediction depend on effective methods for the extraction of domain-relevant variables. This paper proposes a methodology for identifying domain-specific terms. The proposed methodology relies on a collection of documents labeled as relevant or irrelevant to the domain under analysis. Based on the labeled document collection, we propose a supervised technique that weights terms based on their descriptive and discriminating power. Finally, the descriptive and discriminating values are combined into a general measure that, through the use of an adjustable parameter, allows to independently favor different aspects of retrieval such as maximizing precision or recall, or achieving a balance between both of them. The proposed technique is applied to the economic domain and is empirically evaluated through a human-subject experiment involving experts and non-experts in Economy. It is also evaluated as a term-weighting technique for query-term selection showing promising results. We finally illustrate the applicability of the proposed technique to address diverse problems such as building prediction models, supporting knowledge modeling, and achieving total recall.

Palabras claves

Term Weighting - Variable Extraction - Information Retrieval - Query-Term Selection

Acceso

PÁGINAS

pp. 61 - 80

NÚMERO

Volumen: 22 Número: 63 Parte: 0 (2019)

MATERIAS

INGENIERÍA Y CONSTRUCCIÓN CIVIL
TECNOLOGÍA

REVISTAS SIMILARES

Information
Algorithms
Aerospace

Artículos similares

Applying the Geostatistical Eigenvector Spatial Filter Approach into Regularized Regression for Improving Prediction Accuracy for Mass Appraisal

Acceso

Michael McCord, Daniel Lo, Peadar Davis, John McCord, Luc Hermans and Paul Bidanset

Prediction accuracy for mass appraisal purposes has evolved substantially over the last few decades, facilitated by the evolution in big data, data availability and open source software. Accompanying these advances, newer forms of geo-spatial approaches ... ver más

Revista: Applied Sciences

Anatomy of a Data Science Software Toolkit That Uses Machine Learning to Aid ?Bench-to-Bedside? Medical Research?With Essential Concepts of Data Mining and Analysis Explained

Acceso

László Beinrohr, Eszter Kail, Péter Piros, Erzsébet Tóth, Rita Fleiner and Krasimir Kolev

Data science and machine learning are buzzwords of the early 21st century. Now pervasive through human civilization, how do these concepts translate to use by researchers and clinicians in the life-science and medical field? Here, we describe a software ... ver más

Revista: Applied Sciences

MODC: A Pareto-Optimal Optimization Approach for Network Traffic Classification Based on the Divide and Conquer Strategy

Acceso

Zuleika Nascimento and Djamel Sadok

Network traffic classification aims to identify categories of traffic or applications of network packets or flows. It is an area that continues to gain attention by researchers due to the necessity of understanding the composition of network traffics, wh... ver más

Revista: Information

Revistas destacadas

Acceso directo a los números publicados en la revista Infrastructures

Infrastructures

Acceso directo a los números publicados en la revista Informed Infraestructure

Informed Infraestructure

Acceso directo a los números publicados en la revista BiT

Acceso directo a los números publicados en la revista Revista de la Construcción

Revista de la Construcción

Ver todas las revistas