Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Ensemble approach to code smell identification: Evaluating ensemble machine learning techniques to identify code smells within a software system
Jönköping University, School of Engineering, JTH, Computer Science and Informatics.
2020 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

The need for automated methods for identifying refactoring items is prelevent in many software projects today. Symptoms of refactoring needs is the concept of code smells within a software system. Recent studies have used single model machine learning to combat this issue. This study aims to test the possibility of improving machine learning code smell detection using ensemble methods. Therefore identifying the strongest ensemble model in the context of code smells and the relative sensitivity of the strongest perfoming ensemble identified. The ensemble models performance was studied by performing experiments using WekaNose to create datasets of code smells and Weka to train and test the models on the dataset. The datasets created was based on Qualitas Corpus curated java project. Each tested ensemble method was then compared to all the other ensembles, using f-measure, accuracy and AUC ROC scores. The tested ensemble methods were stacking, voting, bagging and boosting. The models to implement the ensemble methods with were models that previous studies had identified as strongest performer for code smell identification. The models where Jrip, J48, Naive Bayes and SMO.

The findings showed, that compared to previous studies, bagging J48 improved results by 0.5%. And that the nominally implemented baggin of J48 in Weka follows best practices and the model where impacted negatively. However, due to the complexity of stacking and voting ensembles further work is needed regarding stacking and voting ensemble models in the context of code smell identification.

Place, publisher, year, edition, pages
2020. , p. 59
Keywords [en]
Ensemble machine learning, code smell, technical debt, code smell identification, automated code smell identification
National Category
Computer Systems
Identifiers
URN: urn:nbn:se:hj:diva-49319ISRN: JU-JTH-PRU-2-20200197OAI: oai:DiVA.org:hj-49319DiVA, id: diva2:1441088
Subject / course
JTH, Computer Engineering
Supervisors
Examiners
Available from: 2020-06-24 Created: 2020-06-15 Last updated: 2025-10-13Bibliographically approved

Open Access in DiVA

fulltext(859 kB)2220 downloads
File information
File name FULLTEXT01.pdfFile size 859 kBChecksum SHA-512
51a6e5d8828451e47ab84a3a8827cfb5f45e2bbf5f3197996912348ceac47244e191f24fbf4b3df8a758c355ac2727030123ca787cdeab15c7e19f3b1c60534a
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Johansson, Alfred
By organisation
JTH, Computer Science and Informatics
Computer Systems

Search outside of DiVA

GoogleGoogle Scholar
Total: 2222 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 980 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf