Introduction to Statistics (Using R Studio)
Series Description: This course will cover some fundamental concepts in statistics, giving students the tools and confidence to analyse and present their data. This will include an introduction to using RStudio, a widely used software package for statistics.
Target Audience: This series is for any student who will be working with data as part of their assignments, project, or dissertation, but who has little or no statistics background. This includes UG, PGT and PGR students.
> Moodle page for this series (includes slides and any recordings)
| Date | Time | Class Title | Class Description | Venue |
| Wed 30 Sep | 13:00-14:00 | Introduction to R Studio (Part 1) | This first session introduces some of the basic functionality of R Studio. Bring your laptop with you to follow along! | Gilbert Scott Building: Room 356 |
| Fri 2 Oct | 13:00-14:00 | Introduction to R Studio (Part 1) (online repeat) | This first session introduces some of the basic functionality of R Studio. Bring your laptop with you to follow along! | Teams Link |
| Wed 7 Oct | 13:00-14:00 | Introduction to R Studio (Part 2) | In the second session of this series, we will become more comfortable with R Studio and use it to create impactful graphs and predictive models. | Gilbert Scott Building: Room 356 |
| Fri 9 Oct | 13:00-14:00 | Introduction to R Studio (Part 2) (online repeat) | In the second session of this series, we will become more comfortable with R Studio and use it to create impactful graphs and predictive models. | Teams Link |
| Wed 14 Oct | 13:00-14:00 | Descriptive Statistics | The third session in this series looks at what information we can draw immediately from our data, while still painting a more complete picture than a simple average. We will cover measures of central tendencies, dispersion, and position. | Gilbert Scott Building: Room 356 |
| Fri 16 Oct | 13:00-14:00 | Descriptive Statistics (online repeat) | The third session in this series looks at what information we can draw immediately from our data, while still painting a more complete picture than a simple average. We will cover measures of central tendencies, dispersion, and position. | Teams Link |
| Wed 21 Oct | 13:00-14:00 | Probability | To certainly give students a better chance of answering the question "how likely was that?", our fourth session covers the basic rules of probability, as well as both discrete and continuous probability distributions. | Gilbert Scott Building: Room 356 |
| Fri 23 Oct | 13:00-14:00 | Probability (online repeat) | To certainly give students a better chance of answering the question "how likely was that?", our fourth session covers the basic rules of probability, as well as both discrete and continuous probability distributions. | Teams Link |
| Wed 28 Oct | 13:00-14:00 | Hypothesis Testing | This fifth session will cover hypothesis testing, which is used to draw conclusions about a whole population from a sample of data, e.g. how can news outlets call an election with only a fraction of the votes tallied? We will discuss how to choose the null and alternative hypothesis, and which distributions to use. | Gilbert Scott Building: Room 356 |
| Fri 30 Oct | 13:00-14:00 | Hypothesis Testing (online repeat) | This fifth session will cover hypothesis testing, which is used to draw conclusions about a whole population from a sample of data, e.g. how can news outlets call an election with only a fraction of the votes tallied? We will discuss how to choose the null and alternative hypothesis, and which distributions to use. | Teams Link |
| Wed 4 Nov | 13:00-14:00 | Simple and Multiple Linear Regression | This sixth session will discuss the relationship, or more precisely the correlation, between variables, and how to describe these relationships using simple and multiple linear regression. We will use R to generate a best fit line to pairwise ordered data, and then also generate a more complex linear model. | Gilbert Scott Building: Room 356 |
| Fri 6 Nov | 13:00-14:00 | Simple and Multiple Linear Regression (online repeat) | This sixth session will discuss the relationship, or more precisely the correlation, between variables, and how to describe these relationships using simple and multiple linear regression. We will use R to generate a best fit line to pairwise ordered data, and then also generate a more complex linear model. | Teams Link |
| Wed 11 Nov | 13:00-14:00 | Logistic Regression | Does the amount of time a student spends studying increase the probability of passing their course, and if so, what’s my probability of passing if I spend x hours studying? This session will show how this can be answered using logistic regression, and how this can be implemented in R Studio. | Gilbert Scott Building: Room 356 |
| Fri 13 Nov | 13:00-14:00 | Logistic Regression (online repeat) | Does the amount of time a student spends studying increase the probability of passing their course, and if so, what’s my probability of passing if I spend x hours studying? This session will show how this can be answered using logistic regression, and how this can be implemented in R Studio. | Teams Link |
| Wed 18 Nov | 13:00-14:00 | Flexible Regression | Sometimes a linear model won’t be appropriate to model the data we have and we have to instead use a flexible yet smooth curve. The last of our sessions will show how to create a flexible regression model using the R package “mgcv”. | Gilbert Scott Building: Room 356 |
| Fri 20 Nov | 13:00-14:00 | Flexible Regression (online repeat) | Sometimes a linear model won’t be appropriate to model the data we have and we have to instead use a flexible yet smooth curve. The last of our sessions will show how to create a flexible regression model using the R package “mgcv”. | Teams Link |