Overview
I taught the R part of a big data training course (88 hours; R, Hadoop, and Spark) run in Incheon in summer 2020. The course was for university students and job seekers, and in the final curriculum my lectures take 24 hours over four days: July 16, 17, 23, and 24. Machine learning, Hadoop, and Spark in the later part of the course were taught by other instructors. The lecture materials are slides for 14 sessions (Days 1 to 4) gathered under the title “데이터분석과 통계 기초” [Data Analysis and Basic Statistics], together with practice scripts; I also built a textbook skeleton with the same table of contents as the curriculum.
Contents
- Day 1: installing R and RStudio and basic usage (variables, functions, packages, data types). Getting to know the data, preprocessing with dplyr, and visualization with ggplot2.
- Day 2: basic SQL syntax learned through sqldf. Foundations of analysis (the steps of scientific research, types of research, data and measurement, data structures).
- Day 3: basic statistics and sampling, hypothesis testing, the chi-square test and cross-tabulation, hypothesis tests on means (one-sample, independent two-sample), analysis of variance, correlation analysis, and regression analysis and logistic regression.
- Day 4: cluster analysis, association analysis (market basket analysis), and social network analysis (representing networks, network-level indicators and individual-level centrality, practice in R).
The curriculum and the textbook table of contents also include documentation with R Markdown, and text mining and topic modeling. Most sessions end with exercises.
Tools and materials
- R and RStudio.
- Packages: dplyr, ggplot2, sqldf.
- Slides for Sessions 1 to 14 and an R script for each session.
- A textbook skeleton built with R Markdown, “빅데이터 교육 과정” [Big Data Training Course].
- Practice data:
mtcars,iris, and an example file of exam scores.