数据分析之美：决策树R语言实现-CDA数据分析师官网

数据分析之美：决策树R语言实现

2018-01-23

数据分析之美：决策树 R语言实现

1.准备数据
[plain] view plain copy
    > install.packages("tree")
    > library(tree)
    > library(ISLR)
    > attach(Carseats)
    > High=ifelse(Sales<=8,"No","Yes") //set high values by sales data to calssify
    > Carseats=data.frame(Carseats,High) //include the high data into the data source
    > fix(Carseats)
2.生成决策树
[plain] view plain copy

    > tree.carseats=tree(High~.-Sales,Carseats)
    > summary(tree.carseats)

[plain] view plain copy
    //output training error is 9%
    Classification tree:
    tree(formula = High ~ . - Sales, data = Carseats)
    Variables actually used in tree construction:
    [1] "ShelveLoc"   "Price"       "Income"      "CompPrice"   "Population"
    [6] "Advertising" "Age"         "US"
    Number of terminal nodes: 27
    Residual mean deviance: 0.4575 = 170.7 / 373
    Misclassification error rate: 0.09 = 36 / 400
3. 显示决策树
[plain] view plain copy

    > plot(tree . carseats )
    > text(tree .carseats ,pretty =0)
4.Test Error

[plain] view plain copy

    //prepare train data and test data
    //We begin by using the sample() function to split the set of observations sample() into two halves, by selecting a random subset of 200 observations out of the original 400 observations.
    > set . seed (1)
    > train=sample(1:nrow(Carseats),200)
    > Carseats.test=Carseats[-train,]
    > High.test=High[-train]
    //get the tree model with train data
    > tree. carseats =tree (High~.-Sales , Carseats , subset =train )
    //get the test error with tree model, train data and predict method
    //predict is a generic function for predictions from the results of various model fitting functions.
    > tree.pred = predict ( tree.carseats , Carseats .test ,type =" class ")
    > table ( tree.pred ,High. test)
    High. test
    tree. pred No Yes
    No 86 27
    Yes 30 57
    > (86+57) /200
    [1] 0.715

5.决策树剪枝
[plain] view plain copy

    /**
    Next, we consider whether pruning the tree might lead to improved results. The function cv.tree() performs cross-validation in order to cv.tree() determine the optimal level of tree complexity; cost complexity pruning is used in order to select a sequence of trees for consideration.

    For regression trees, only the default, deviance, is accepted. For classification trees, the default is deviance and the alternative is misclass (number of misclassifications or total loss).
    We use the argument FUN=prune.misclass in order to indicate that we want the classification error rate to guide the cross-validation and pruning process, rather than the default for the cv.tree() function, which is deviance.

    If the tree is regression tree,
    > plot(cv. boston$size ,cv. boston$dev ,type=’b ’)
    */
    > set . seed (3)
    > cv. carseats =cv. tree(tree .carseats ,FUN = prune . misclass ,K=10)
    //The cv.tree() function reports the number of terminal nodes of each tree considered (size) as well as the corresponding error rate(dev) and the value of the cost-complexity parameter used (k, which corresponds to α.
    > names (cv. carseats )
    [1] " size" "dev " "k" " method "
    > cv. carseats
    $size //the number of terminal nodes of each tree considered
    [1] 19 17 14 13 9 7 3 2 1
    $dev //the corresponding error rate
    [1] 55 55 53 52 50 56 69 65 80
    $k // the value of the cost-complexity parameter used
    [1] -Inf 0.0000000 0.6666667 1.0000000 1.7500000
    2.0000000 4.2500000
    [8] 5.0000000 23.0000000
    $method   //miscalss for classification tree
    [1] " misclass "
    attr (," class ")
    [1] " prune " "tree. sequence "

[plain] view plain copy

    //plot the error rate with tree node size to see whcih node size is best
    > plot(cv. carseats$size ,cv. carseats$dev ,type=’b ’)

    /**
    Note that, despite the name, dev corresponds to the cross-validation error rate in this instance. The tree with 9 terminal nodes results in the lowest cross-validation error rate, with 50 cross-validation errors. We plot the error rate as a function of both size and k.
    */
    > prune . carseats = prune . misclass ( tree. carseats , best =9)
    > plot( prune . carseats )
    > text( prune .carseats , pretty =0)

    //get test error again to see whether the this pruned tree perform on the test data set
    > tree.pred = predict ( prune . carseats , Carseats .test , type =" class ")
    > table ( tree.pred ,High. test)
    High. test
    tree. pred No Yes
    No 94 24
    Yes 22 60
    > (94+60) /200
    [1] 0.77

CDA数据分析师考试相关入口一览（建议收藏）：

▷ 想报名CDA认证考试，点击>>> “CDA报名” 了解CDA考试详情；

▷ 想学习CDA考试教材，点击>>> “CDA教材” 了解CDA考试详情；

▷ 想加入CDA考试题库，点击>>> “CDA题库” 了解CDA考试详情；

▷ 想了解CDA考试含金量，点击>>> “CDA含金量” 了解CDA考试详情；

决策树 R语言决策树剪枝数据分析

数据分析咨询请扫描二维码

若不方便扫码，搜微信号：CDAshujufenxi

上一篇存储之于大数据分析

下一篇Python中最大最小赋值小技巧

数据分析之美：决策树R语言实现

CDA考试动态

CDA报考指南

热门栏目

最新资讯

单因素方差分析结果与多重比较

【CDA干货】13年国企财务：这样使用财务数据分析模 ...

Youtube百万粉丝大佬：数据分析师职业发展路径 ...

【干货】“数据又崩了”？其实是你还不会做归因分析 ...

【CDA干货】解锁企业数据价值的3大关键 ——从政策 ...

【CDA案例】基于 EAST和 FineBI 实现 AARRR 信用卡 ...

【CDA干货】互联网运营必看：私域用户质量数据分析 ...

【CDA持证人案例分享】用 Excel 精准监控电商及推广 ...

【CDA持证人干货分享】13年国企财务：如何借助DeepS ...

【CDA持证人案例分享】Excel动态报表设计：基于业务 ...

【CDA干货】字节大佬：如何通过动态分级快速提升转 ...

Windows 系统和 MacOS 系统下的 Anaconda 安装教程 ...

数据运营的工作内容、技能要求及发展前景 ...

【干货】字节大佬：教培行业销售运营全景作战地图【 ...

四大一线城市约50%人口租房，数据分析能挖出哪些 “ ...

Python 实战案例 —RFM 客户价值分析模型 ...

【案例】奥利奥坚果新品与蒙牛“数字牧场”的成功经 ...

美关税政策下的全球金融市场动荡：深度数据分析与洞 ...

【重磅】苹果捐赠3000万给浙大这专业，透露未来就业 ...

非专业，怎么才能证明自己的数据分析能力？ ...