การตรวจจับการฉ้อโกงใน R
Bart Baesens
Professor Data Science at KU Leuven



การฉ้อโกงเป็นอาชญากรรมที่เกิดขึ้นน้อย วางแผนมาอย่างดี ซ่อนเร้นแนบเนียน เปลี่ยนแปลงตามเวลา และมักถูกจัดการอย่างเป็นระบบ โดยปรากฏในรูปแบบที่หลากหลาย









หลังพายุครั้งใหญ่ บริษัทประกันภัยได้รับเคลมจำนวนมาก
สัดส่วนของเคลมฉ้อโกงในข้อมูลหาได้ด้วยฟังก์ชัน
table() และ prop.table()ใช้ prop.table(table(...)) เพื่อดูสัดส่วนของการฉ้อโกง
prop.table(table(fraud_label))
0 1
0.9911 0.0089
labels <- c("no fraud", "fraud")
labels <- paste(labels, round(100 * prop.table(table(fraud_label)), 2), "%")
pie(table(fraud_label), labels, col = c("blue", "red"),
main = "Pie chart of storm claims")

ใช้สำหรับประเมินโมเดลตรวจจับการฉ้อโกง:

predictions <- rep.int(0, times = nrow(claims))
predictions <- factor(predictions, levels = c("no fraud", "fraud"))
confusionMatrix() จากแพ็กเกจ caret:library(caret)
confusionMatrix(data = predictions, reference = fraud_label)
Reference
Prediction 0 1
0 614 14
1 0 0
Accuracy : 0.9777
> total_cost <- sum(claim_amount[fraud_label == "fraud"])
> print(total_cost)
2301508
การตรวจจับการฉ้อโกงใน R