Ad-RuLer: A novel rule-driven data synthesis technique for imbalanced classification

dc.contributor.authorZhang, Xiao
dc.contributor.authorPaz Ortiz, Alejandro Iván
dc.contributor.authorNebot Castells, M. Àngela
dc.contributor.authorMúgica Álvarez, Francisco
dc.contributor.authorRomero Merino, Enrique
dc.contributor.groupUniversitat Politècnica de Catalunya. IDEAI-UPC - Intelligent Data sciEnce and Artificial Intelligence Research Group
dc.contributor.otherUniversitat Politècnica de Catalunya. Doctorat en Intel·ligència Artificial
dc.contributor.otherUniversitat Politècnica de Catalunya. Departament de Ciències de la Computació
dc.date.accessioned2023-12-21T13:37:09Z
dc.date.available2023-12-21T13:37:09Z
dc.date.issued2023-11-23
dc.description.abstractWhen classifiers face imbalanced class distributions, they often misclassify minority class samples, consequently diminishing the predictive performance of machine learning models. Existing oversampling techniques predominantly rely on the selection of neighboring data via interpolation, with less emphasis on uncovering the intrinsic patterns and relationships within the data. In this research, we present the usefulness of an algorithm named RuLer to deal with the problem of classification with imbalanced data. RuLer is a learning algorithm initially designed to recognize new sound patterns within the context of the performative artistic practice known as live coding. This paper demonstrates that this algorithm, once adapted (Ad-RuLer), has great potential to address the problem of oversampling imbalanced data. An extensive comparison with other mainstream oversampling algorithms (SMOTE, ADASYN, Tomek-links, Borderline-SMOTE, and KmeansSMOTE), using different classifiers (logistic regression, random forest, and XGBoost) is performed on several real-world datasets with different degrees of data imbalance. The experiment results indicate that Ad-RuLer serves as an effective oversampling technique with extensive applicability.
dc.description.peerreviewedPeer Reviewed
dc.description.versionPostprint (published version)
dc.format.extent22 p.
dc.identifier.citationZhang, X. [et al.]. Ad-RuLer: A novel rule-driven data synthesis technique for imbalanced classification. "Applied sciences (Basel)", 23 Novembre 2023, vol. 13, núm. 23, article 12636.
dc.identifier.doi10.3390/app132312636
dc.identifier.issn2076-3417
dc.identifier.urihttps://hdl.handle.net/2117/398676
dc.language.isoeng
dc.publisherMultidisciplinary Digital Publishing Institute
dc.relation.publisherversionhttps://www.mdpi.com/2076-3417/13/23/12636
dc.rights.accessOpen Access
dc.rights.licensenameAttribution 4.0 International
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subjectÀrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic
dc.subject.lcshMachine learning
dc.subject.lcshAlgorithms
dc.subject.lemacAprenentatge automàtic
dc.subject.lemacAlgorismes
dc.subject.otherRule-based approach
dc.subject.otherOversampling
dc.subject.otherData synthesis
dc.subject.otherImbalanced data
dc.subject.otherClassification
dc.titleAd-RuLer: A novel rule-driven data synthesis technique for imbalanced classification
dc.typeArticle
dspace.entity.typePublication
local.citation.authorZhang, X.; Paz, A.; Nebot, A.; Mugica, F.; Romero, E.
local.citation.number23, article 12636
local.citation.publicationNameApplied sciences (Basel)
local.citation.volume13
local.identifier.drac37834461

Fitxers

Paquet original

Mostrant 1 - 1 de 1
Carregant...
Miniatura
Nom:
applsci-13-12636-v4.pdf
Mida:
6.52 MB
Format:
Adobe Portable Document Format
Descripció: