Linear Regression in R
Experiment: Linear Regression in R
1. Experiment Title
Implementation of Simple Linear Regression Using R
2. Aim
To study and implement Simple Linear Regression in R to analyze the relationship between an independent variable and a dependent variable and to predict the dependent variable.
3. Objectives
After completing this experiment, students should be able to:
- Understand the concept of simple linear regression.
- Identify independent and dependent variables.
- build a linear regression model using R.
- Interpret the regression coefficients.
- Predict new values using the regression model.
- Visualize the regression line.
- Evaluate the performance of the regression model.
4. Theory
4.1 What is Linear Regression?
Linear Regression is a statistical method used to study the relationship between two numerical variables.
It helps us understand how a dependent variable () changes when an independent variable () changes.
For example:
How do student marks change based on the number of hours studied?
Here:
- Independent Variable () → Hours Studied
- Dependent Variable () → Marks Obtained
The relationship is represented using a straight-line equation:
where:
- = Intercept
- = Slope (Regression Coefficient)
- = Independent variable
- = Dependent variable
The regression line is the best-fitting straight line that represents the relationship between and .
4.2 Simple Linear Regression Model
For a dataset containing several observations:
where:
- = Actual dependent variable
- = Independent variable
- = Intercept
- = Regression coefficient
- = Error or residual
4.3 Method of Least Squares
The regression line is obtained using the Least Squares Method.
The objective is to minimize the sum of squared differences between the actual and predicted values.
The regression coefficients are calculated as:
Slope
Intercept
5. Problem Statement
A teacher wants to study the relationship between the number of hours studied by students and their marks obtained in an examination.
The following data were collected:
| Student | Hours Studied | Marks Obtained |
|---|---|---|
| 1 | 1 | 42 |
| 2 | 2 | 48 |
| 3 | 3 | 55 |
| 4 | 4 | 60 |
| 5 | 5 | 65 |
| 6 | 6 | 72 |
| 7 | 7 | 78 |
| 8 | 8 | 85 |
The task is to:
- Build a simple linear regression model.
- Study the relationship between hours studied and marks obtained.
- Obtain the regression equation.
- Predict the marks of a student who studies for 9 hours.
- Visualize the regression line.
6. R Program
# Simple Linear Regression in R # Create the dataset hours <- c(1, 2, 3, 4, 5, 6, 7, 8) marks <- c(42, 48, 55, 60, 65, 72, 78, 85) # Create a data frame student_data <- data.frame(hours, marks) print(student_data) # Create the Linear Regression Model model <- lm(marks ~ hours, data = student_data) # Display the regression model print(model) # Display detailed model summary print("Summary") print(summary(model)) # Obtain regression coefficients print(coefficients(model)) # Predict marks for a student studying 9 hours new_data <- data.frame(hours = 9) prediction <- predict(model, newdata = new_data) cat("Predicted Marks for 9 Hours of Study =", prediction) # Plot the original data plot(student_data$hours, student_data$marks, main = "Hours Studied vs Marks", xlab = "Hours Studied", ylab = "Marks Obtained", pch = 19) # Add regression line abline(model, col = "blue", lwd = 2)Output
hours marks 1 1 42 2 2 48 3 3 55 4 4 60 5 5 65 6 6 72 7 7 78 8 8 85 Call: lm(formula = marks ~ hours, data = student_data)"Summary" Call: lm(formula = marks ~ hours, data = student_data) Residuals: Min 1Q Median 3Q Max -1.14286 -0.18750 -0.07143 0.18750 0.92857 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) 35.9643 0.5343 67.31 7.23e-10 *** hours 6.0357 0.1058 57.04 1.95e-09 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 0.6857 on 6 degrees of freedom Multiple R-squared: 0.9982, Adjusted R-squared: 0.9979 F-statistic: 3254 on 1 and 6 DF, p-value: 1.95e-09Coefficients: (Intercept) hours 35.964 6.036 Predicted Marks for 9 Hours of Study = 90.28571
7. Explanation of the Program
Step 1: Create the Variables
hours <- c(1, 2, 3, 4, 5, 6, 7, 8) marks <- c(42, 48, 55, 60, 65, 72, 78, 85)
Two numerical vectors are created:
-
hours→ Independent variable -
marks→ Dependent variable
Step 2: Create a Data Frame
student_data <- data.frame(hours, marks)
The two variables are combined into a data frame.
Step 3: Create the Regression Model
model <- lm(marks ~ hours, data = student_data)
The lm() function is used to create a linear model.
The notation:
marks ~ hours
means:
Predict marks using hours.
Step 4: Display the Model
print(model)
This displays the estimated regression equation.
The output will be approximately:
Step 5: Get Detailed Results
summary(model)
The summary provides:
- Regression coefficients
- Standard errors
- t-values
- p-values
- Residual information
- R-squared value
- Adjusted R-squared value
- F-statistic
Step 6: Make a Prediction
new_data <- data.frame(hours = 9) prediction <- predict(model, newdata = new_data)
The predict() function uses the trained regression model to estimate the marks for a student studying 9 hours.
Step 7: Plot the Regression Line
plot(student_data$hours, student_data$marks)
This creates a scatter plot.
abline(model)
This adds the best-fitting regression line to the plot.
8. Expected Output
The regression model will produce an equation approximately of the form:
The exact values can be obtained by running the program in R.
Interpretation
The slope is approximately:
This means:
For every additional hour of study, the predicted marks increase by approximately 6 marks.
9. Model Evaluation
9.1 R-Squared Value
The R-squared () value indicates how well the regression model explains the variation in the dependent variable.
summary(model)$r.squared
A value closer to 1 indicates that the model explains a large proportion of the variation in the dependent variable.
9.2 Predicted Values
predict(model)
This gives the predicted marks for the existing observations.
9.3 Residuals
Residuals represent the difference between actual and predicted values.
In R:
residuals(model)
10. Result
A Simple Linear Regression model was successfully implemented using R to study the relationship between hours studied and marks obtained. The regression equation was obtained, the relationship between the variables was analyzed, and the model was used to predict marks for a new value of the independent variable. The regression line was also visualized using a scatter plot.
Important R Functions Used
| Function | Purpose |
|---|---|
lm() | Creates the linear regression model |
summary() | Displays detailed regression results |
coefficients() | Obtains regression coefficients |
predict() | Predicts new values |
plot() | Creates a scatter plot |
abline() | Adds the regression line |
residuals() | Obtains prediction errors |
Comments
Post a Comment