Repository navigation
Expand file tree
/
Copy pathtmp.R
More file actions
627 lines (397 loc) · 23.3 KB
/
Copy pathtmp.R
File metadata and controls
627 lines (397 loc) · 23.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
#### Correlation
The next step invloves checking the correlation between the features. This will help us reduce the number of features based on the strength of association. In addition, training a model on a data set with features which have little or no correlation may lead to inaccurate results. It is therefore important to identify and filter out features which are not correlated to other features, more so if measuring the similar aspects, [Yu, Lei and Liu, Huan](http://www.aaai.org/Papers/ICML/2003/ICML03-111.pdf).
```{r, echo=T}
features_df <- wdbc[, ind_vars]
cor_mat <- cor(features_df)
corrplot(cor_mat, order = "hclust", tl.cex = 1, addrect = 8)
```
## Model Fitting
### Principal Component Analysis and Linear Discriminant Analysis
Since there are many correlated features, we will use PCA to reduce the dimension of the data. Both LDA and PCA are used to classify and reduce the dimentionality of the data. One of the key difference is that PCA is unsupervised learning while LDA is supervised.
#### PCA
### Machine Learning Models
Machine learning is a science which involves the application of Artificial Intelligence (AI) that enables the computers to automatically learn and improve or get things done based on experience without being explicitly programed. This involves computer looking for patterns based on previous observations, experiences or instructions. ML algorithms are broadly classified as supervised, unsupervised and reinforcement learning. In this exercise, we will fit some of the supervised learning algorithms:
* LDA
* Neural Networks
* K - Nearest Neighbors (KNN)
* Random Forest
* Naive Bayes
* Support Vector Machine (SVM)
We then further, to some models, incorporate either PCA or LDA to assess if there could be any improvement in model performance.
#### Dataset Partitioning
We need to assess the model quality after model fitting. Datset partitioning helps us split the datset into training (used for model fitting) and test (used for model validation) datasets. In the present case, we split the datset into $80$% of which we will use to train our models and $20$% that we will hold back as a validation dataset.
```{r, echo=T}
set.seed(1000)
df_index <- createDataPartition(wdbc$diagnosis, p=0.7, list = F)
train_df <- wdbc[df_index, -1]
test_df <- wdbc[-df_index, -1]
```
Data \textbf{pre-processing} involves transforming data into a specific format to improve the performance of machine learning algorithms. We'll use \textbf{preProcess} function in package \textit{caret} to transform the data.
Our main focus is transforming our predictors (characteristics) such that we reduce their dimensionality to fewer dimension where the new characterisitcs are uncorrelated. Therefore, we'll apply \textit{pca} method with a threshold of $.99$.
```{r, echo=T}
model_control <- trainControl(method="cv",
number = 5,
preProcOptions = list(thresh = 0.99),
classProbs = TRUE,
summaryFunction = twoClassSummary)
```
#### LDA
Define training and testing data sets for LDA models.
```{r, echo=T}
train_lda_df <- lda_result_df[df_index,]
test_lda_df <- lda_result_df[-df_index,]
```
```{r, echo=T}
lda_model <- train(diagnosis~., train_lda_df, method="lda2", metric="ROC",
preProc = c("center", "scale"),
tuneLength = 10,
trControl=model_control)
```
Using \textit{predict} function, we generate the clasiffication and posterior probabilities based on the LDA model.
```{r, echo=T}
lda_predicted <- predict(lda_model, test_lda_df)
knitr::kable(as.data.frame(table(lda_predicted)))
```
We can summarize the performance of the LDA classification using $\textit{Confusion Matrix}$.
```{r, echo=T}
lda_confusion_mat <- confusionMatrix(lda_predicted, test_lda_df$diagnosis, positive = "M")
knitr::kable(lda_confusion_mat$table)
```
The overall accuracy of our model is `r round(lda_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(lda_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(lda_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
lda_predicted_prob <- predict(lda_model, test_lda_df, type="prob")
#colAUC(lda_predicted_prob, test_lda_df$diagnosis, plotROC=TRUE)
```
#### Neural Networks
Our second ML algorithm applied to the data is Neural Networks. A neural network model can be though of as a model consisting of an activation function. The activation function transforms inputto output through information processing unit. For this and other reasons, neural networks is considered complex and mostly oftenly regarded as a 'black box'.
```{r, echo=T}
nn_model <- train(diagnosis~.,
train_df,
method="nnet",
metric="ROC",
preProcess=c('center', 'scale'),
trace=FALSE,
tuneLength=10,
trControl=model_control)
```
Predicted values for Neural Network model.
```{r}
nn_predicted <- predict(nn_model, test_df)
knitr::kable(as.data.frame(table(nn_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
nn_confusion_mat <- confusionMatrix(nn_predicted, test_df$diagnosis, positive = "M")
knitr::kable(nn_confusion_mat$table)
```
The overall accuracy of our model is `r round(nn_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(nn_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(nn_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
nn_predicted_prob <- predict(nn_model, test_df, type="prob")
#colAUC(nn_predicted_prob, test_df$diagnosis, plotROC=TRUE)
```
#### K-nearest Neighbors (KNN)
KNN is a ML algorithm widely used in classification problems. The algorithm goes through the datset to find the k-nearest case and classifies new cases by a majority vote of its k neighbors. The output, in case of classification problem, is the most frequent class. The similairty measure is expressed as a distance measure. These include Euclidean, Manhattan, Minkowski and Hamming distance. Or simply, once the data is stored in the system, its similarity is calculated for any input data point coming into the system.
```{r, echo=T}
knn_model <- train(diagnosis~.,
train_df,
method="knn",
metric="ROC",
preProcess = c('center', 'scale'),
tuneLength=10,
trControl=model_control)
```
Predicted values for Neural Network model.
```{r}
knn_predicted <- predict(knn_model, test_df)
knitr::kable(as.data.frame(table(knn_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
knn_confusion_mat <- confusionMatrix(knn_predicted, test_df$diagnosis, positive = "M")
knitr::kable(knn_confusion_mat$table)
```
The overall accuracy of our model is `r round(knn_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(knn_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(knn_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
knn_predicted_prob <- predict(knn_model, test_df, type="prob")
#colAUC(knn_predicted_prob, test_df$diagnosis, plotROC=TRUE)
```
#### Random Forest
Random Forest is a supervised learning algorithm and can be understood as a bootstrapping algorithm with a Decision tree model, mostly, trained using 'bagging' method. In other words, Random forests built from several decision trees merged together to get a more stable and accurate predictions. It only important to note that the results tend to be more accurate with larger number of trees. Thi algorithm adds additonal randomness into the model with increasing number of trees.
```{r, echo=T}
rand_forest_model <- train(diagnosis~.,
train_df,
method="ranger",
metric="ROC",
preProcess = c('center', 'scale'),
tuneLength=10,
trControl=model_control)
```
Predicted values for random forest model.
```{r}
rand_forest_predicted <- predict(rand_forest_model, test_df)
knitr::kable(as.data.frame(table(rand_forest_predicted)))
```
Summary based on $textit{Confusion Matrix}$.
```{r, echo=T}
rand_forest_confusion_mat <- confusionMatrix(rand_forest_predicted, test_df$diagnosis, positive = "M")
knitr::kable(rand_forest_confusion_mat$table)
```
The overall accuracy of our model is `r round(rand_forest_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(rand_forest_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(rand_forest_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
rand_forest_confusion_prob <- predict(rand_forest_model, test_df, type="prob")
#colAUC(rand_forest_confusion_prob, test_df$diagnosis, plotROC=TRUE)
```
#### Naive Bayes
Naive Bayes is a classification problem based on Bayes' theorem. It assumes that the all the dependent variables are independent of each other, which generally not true in many real-life problems. In other words, it assumes that the presence of a given feature in a given class is completely unrelated to the presence of any other features. The application of Naive Bayes model is considered less complicated, mostly application in large datasets and provides a way to calculte the posterior probabilities.
```{r, echo=T}
naive_bayes_model <- train(diagnosis~.,
train_df,
method="nb",
metric="ROC",
preProcess = c('center', 'scale'),
trace=F,
trControl=model_control)
```
Predicted values for Naive Bayes model.
```{r}
naive_bayes_predicted <- predict(naive_bayes_model, test_df)
knitr::kable(as.data.frame(table(naive_bayes_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
naive_bayes_confusion_mat <- confusionMatrix(naive_bayes_predicted, test_df$diagnosis, positive = "M")
knitr::kable(naive_bayes_confusion_mat$table)
```
The overall accuracy of our model is `r round(naive_bayes_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(naive_bayes_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(naive_bayes_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
naive_bayes_confusion_prob <- predict(naive_bayes_model, test_df, type="prob")
#colAUC(naive_bayes_confusion_prob, test_df$diagnosis, plotROC=TRUE)
```
#### Support Vector Machine (SVM) - with radial kernel
In this ML algorithm, the input vector is initially mapped into feature space of higher dimensionality and identifies the hyperplane that seperates the data points into two classes. The aim is to maximize the marginal distance between the decision hyperplane and closest boundary data points. In other words, we plot each data point in a n-dimensional space, where n is the number of features in the dataset. We then find the line that seperates the data into two different classes. The line should be such that the distance from the closest point in each of the two classes is maximized.
```{r, echo=T}
svm_model <- train(diagnosis~.,
train_df,
method="svmRadial",
metric="ROC",
preProcess = c('center', 'scale'),
trace=F,
trControl=model_control)
```
Predicted values for Vector Machine (SVM) - with radial kernel.
```{r}
svm_predicted <- predict(svm_model, test_df)
knitr::kable(as.data.frame(table(svm_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
svm_confusion_mat <- confusionMatrix(svm_predicted, test_df$diagnosis, positive = "M")
knitr::kable(svm_confusion_mat$table)
```
The overall accuracy of our model is `r round(svm_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(svm_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(svm_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
svm_confusion_prob <- predict(svm_model, test_df, type="prob")
#colAUC(svm_confusion_prob, test_df$diagnosis, plotROC=TRUE)
```
#### Extending Models with PCA and LDA
We extend some of the models fitted above using LDA or/and PCA to assess whether there could some improvement in the models. Extended models are shown below.
##### Random Forest with PCA
We also fit random forest model with PCA.
```{r, echo=T}
rand_forest_pca_model <- train(diagnosis~.,
train_df,
method="ranger",
metric="ROC",
preProcess = c('center', 'scale', 'pca'),
trControl=model_control)
```
Predicted values for random forest model with PCA.
```{r}
rand_forest_pca_predicted <- predict(rand_forest_pca_model, test_df)
knitr::kable(as.data.frame(table(rand_forest_pca_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
rand_forest_pca_confusion_mat <- confusionMatrix(rand_forest_pca_predicted,
test_df$diagnosis, positive = "M")
knitr::kable(rand_forest_pca_confusion_mat$table)
```
The overall accuracy of our model is `r round(rand_forest_pca_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(rand_forest_pca_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(rand_forest_pca_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
rand_forest_pca_confusion_prob <- predict(rand_forest_pca_model, test_df, type="prob")
#colAUC(rand_forest_pca_confusion_prob, test_df$diagnosis, plotROC=TRUE)
```
##### Neural Networks with PCA
We also fit NW model with PCA.
```{r, echo=T}
nn_pca_model <- train(diagnosis~.,
train_df,
method="nnet",
metric="ROC",
preProcess = c('center', 'scale', 'pca'),
tuneLength=10,
trace=F,
trControl=model_control)
```
Predicted values for NW model with PCA.
```{r}
nn_pca_predicted <- predict(nn_pca_model, test_df)
knitr::kable(as.data.frame(table(nn_pca_predicted)))
```
Summary based on $textit{Confusion Matrix}$.
```{r, echo=T}
nn_pca_confusion_mat <- confusionMatrix(nn_pca_predicted,
test_df$diagnosis, positive = "M")
knitr::kable(nn_pca_confusion_mat$table)
```
The overall accuracy of our model is `r round(nn_pca_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(nn_pca_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(nn_pca_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
nn_pca_confusion_prob <- predict(nn_pca_model, test_df, type="prob")
#colAUC(nn_pca_confusion_prob, test_df$diagnosis, plotROC=TRUE)
```
##### Naive Bayes with LDA
We also fit Naive Bayes model with LDA.
```{r, echo=T}
nb_lda_model <- train(diagnosis~.,
train_lda_df,
method="nb",
metric="ROC",
preProcess = c('center', 'scale'),
trace=F,
trControl=model_control)
```
Predicted values for Naive Bayes model with LDA.
```{r}
nb_lda_predicted <- predict(nb_lda_model, test_lda_df)
knitr::kable(as.data.frame(table(nb_lda_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
nb_lda_confusion_mat <- confusionMatrix(nb_lda_predicted,
test_df$diagnosis, positive = "M")
knitr::kable(nb_lda_confusion_mat$table)
```
The overall accuracy of our model is `r round(nb_lda_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(nb_lda_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(nb_lda_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
nb_lda_confusion_prob <- predict(nb_lda_model, test_lda_df, type="prob")
#colAUC(nb_lda_confusion_prob, test_lda_df$diagnosis, plotROC=TRUE)
```
##### Neural Networks model with LDA
We also fit Neural Networks model with LDA.
```{r, echo=T}
nn_lda_model <- train(diagnosis~.,
train_lda_df,
method="nnet",
metric="ROC",
preProcess = c('center', 'scale'),
tuneLength=10,
trace=F,
trControl=model_control)
```
Predicted values for Neural Networks model with LDA.
```{r}
nn_lda_predicted <- predict(nn_lda_model, test_lda_df)
knitr::kable(as.data.frame(table(nn_lda_predicted)))
```
Summary based on $\textit{Confusion Matrix}$.
```{r, echo=T}
nn_lda_confusion_mat <- confusionMatrix(nn_lda_predicted,
test_df$diagnosis, positive = "M")
knitr::kable(nn_lda_confusion_mat$table)
```
The overall accuracy of our model is `r round(nn_lda_confusion_mat$overall[1], 4)*100`%. In addition, our classifier correctly identified $\textit{M}$ `r round(nn_lda_confusion_mat$byClass[1], 4)*100`% of the time (correctly predicting $\textit{M}$ when indeed we should). Further, the true negatives (our specificity) is `r round(nn_lda_confusion_mat$byClass[2], 4)*100`%.
```{r, echo=T}
nn_lda_confusion_prob <- predict(nn_lda_model, test_lda_df, type="prob")
#colAUC(nn_lda_confusion_prob, test_lda_df$diagnosis, plotROC=T)
```
### Assesing Model Performance
In this Section, we describe some methods we've used in evaluating ML models fitted above. Specifically, we have:
* $\textbf{Accuracy}$ - Is the proportion (percentage) of correctly classifying all the data points.
* $\textbf{Kappa}$ - This provides the classification accuracy. It important where there are imbalances in classes.
* $\textbf{Sensitivity}$ - Is the true positive rate also called the recall. It is the number instances from the positive (first) class that actually predicted correctly.
* $\textbf{Specificity}$ - Is also called the true negative rate. Is the number of instances from the negative class (second) class that were actually predicted correctly.
We start by comparing the area under curve (AUC) of ROC curves. The AUC provides the models ability to discriminate between positive and negative classes. An AUC of $1$ implies that the model provide a perfect prediction while an AUC of 0.5 implies that the model is good as random.
```{r, echo=F}
old.par <- par(mfrow=c(2, 5))
colAUC(lda_predicted_prob, test_lda_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "LDA", cex = 1.5)
colAUC(nn_predicted_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "NN", cex = 1.5)
colAUC(knn_predicted_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "KNN", cex = 1.5)
colAUC(rand_forest_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "RF", cex = 1.5)
colAUC(naive_bayes_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "NB", cex = 1.5)
colAUC(svm_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "SVM", cex = 1.5)
colAUC(rand_forest_pca_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "RF_PCA", cex = 1.5)
colAUC(nn_pca_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "NN_PCA", cex = 1.5)
colAUC(nb_lda_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "NB_LDA", cex = 1.5)
colAUC(nn_lda_confusion_prob, test_df$diagnosis, plotROC=TRUE)
text(0.5, 0.5, "NN_LDA", cex = 1.5)
par(old.par)
```
We apply resample method to comapre the models. Resampling is a method of comparing the performance of various competing machine learning algorithms. It estimates point estimate of the performances (based on a number of samples) which are then compared to obtain the best performing model \textit{(See .pdf file)}.
```{r, echo=T}
models <- list(LDA=lda_model, NN=nn_model, KNN=knn_model,
RF = rand_forest_model, NB=naive_bayes_model,
SVM=svm_model, RF_PCA=rand_forest_pca_model,
NN_PCA=nn_pca_model, NB_LDA=nb_lda_model,
NN_LDA=nn_lda_model)
models_resamples <- resamples(models)
#summary(models_resamples)
```
Let's plot the resamples summary output.
```{r}
# Draw box plots to compare models
scales <- list(x=list(relation="free"), y=list(relation="free"))
bwplot(models_resamples, scales=scales)
```
We further plot the correlation matrix between the models.
```{r}
models_correlation <- modelCor(models_resamples)
corrplot(models_correlation, method="number")
```
<!-- We further provide visual representation of ROC. We observe high variability in some models, specifically, NB and RF_PCA. On the other hand, NN_LDA and LDA provide high AUC but with some variability. -->
<!-- ```{r, error=T} -->
<!-- bwplot(models_resamples, metric="ROC") -->
<!-- ``` -->
#### Model Selection
LDA, NB_LDA and NN_LDA provide the highest accuracy and Kappa in comparisson to the other models.
```{r}
models_conf_mat <- list(LDA=lda_confusion_mat, NN=nn_confusion_mat, KNN=knn_confusion_mat,
RF = rand_forest_confusion_mat, NB=naive_bayes_confusion_mat,
SVM=svm_confusion_mat, RF_PCA=rand_forest_pca_confusion_mat,
NN_PCA=nn_confusion_mat, NB_LDA=nb_lda_confusion_mat,
NN_LDA=nn_lda_confusion_mat)
all_model_accuracy <- sapply(models_conf_mat, function(x){
round(x$overall, 2)
})
knitr::kable(data.frame(all_model_accuracy))
```
In addition, NN_LDA model provides the highest sensitivity.
```{r, echo=T}
#function(x) x$byClass
all_model_results <- sapply(models_conf_mat, function(x){
round(x$byClass, 2)
})
knitr::kable(data.frame(all_model_results))
```
To provide a summary of the best model, we pullout model which performs better in each metric. From the output below, we observe that the model with the highest sensitivity (detection of breast cancer cases) is NN_LDA.
```{r}
all_model_results_max <- apply(all_model_results, 1, which.is.max)
model_select <- data.frame(metric=names(all_model_results_max),
best_model=colnames(all_model_results)[all_model_results_max],
value=mapply(function(x,y) {all_model_results[x,y]},
names(all_model_results_max),
all_model_results_max))
rownames(model_select) <- NULL
names(model_select) <- c("Metric", "Best_model", "Value")
knitr::kable(data.frame(model_select))
```
### Conclusion
In conclusion, we have found that a model based on Neural Network and LDA provides a good classification of the data. This model has a sensitivity of 0.94.