Multiple Linear Regression in R
Experiment: Multiple Linear Regression in R
1. Experiment Title
Implementation of Multiple Linear Regression Using R
2. Aim
To study and implement Multiple Linear Regression in R to analyze the relationship between a dependent variable and two or more independent variables and to predict the dependent variable.
3. Objectives
After completing this experiment, students should be able to:
- Understand the concept of multiple linear regression.
- Identify the dependent variable and multiple independent variables.
- Build a multiple linear regression model using R.
- Interpret regression coefficients.
- Evaluate the performance of the regression model.
- Predict the dependent variable for new observations.
- Analyze residuals and model accuracy.
4. Theory
4.1 What is Multiple Linear Regression?
Multiple Linear Regression is an extension of simple linear regression.
While simple linear regression uses one independent variable to predict a dependent variable, multiple linear regression uses two or more independent variables.
For example, a student's examination marks may depend on:
- Hours studied
- Attendance percentage
In this case:
- Dependent Variable () → Examination Marks
- Independent Variables () → Hours Studied and Attendance
The general equation of multiple linear regression is:
where:
- = Dependent or response variable
- = Independent or predictor variables
- = Intercept
- = Regression coefficients
- = Error term
4.2 Multiple Regression Model Used in This Experiment
In this experiment, we use two predictors:
- Hours Studied
- Attendance Percentage
The model is:
For this experiment:
The model estimates the relationship between these variables and predicts student marks.
4.3 Interpretation of Regression Coefficients
Suppose the regression equation is:
Then:
- 20 → Intercept
- 5 → Expected change in marks for one additional hour of study, assuming attendance remains constant.
- 0.3 → Expected change in marks for a one ശതമാനം increase in attendance, assuming study hours remain constant.
A major advantage of multiple regression is that it allows us to study the effect of one variable while controlling for the other variables.
5. Problem Statement
A teacher wants to study how hours studied and attendance percentage influence the examination marks of students.
The following data are collected from 10 students.
| Student | Hours Studied | Attendance (%) | Marks |
|---|---|---|---|
| 1 | 2 | 60 | 45 |
| 2 | 3 | 65 | 50 |
| 3 | 4 | 70 | 55 |
| 4 | 5 | 72 | 60 |
| 5 | 4 | 75 | 58 |
| 6 | 6 | 80 | 68 |
| 7 | 7 | 85 | 75 |
| 8 | 5 | 78 | 65 |
| 9 | 8 | 90 | 82 |
| 10 | 9 | 95 | 88 |
Using the given data:
- Build a multiple linear regression model.
- Study the relationship between marks, hours studied, and attendance.
- Obtain and interpret the regression equation.
- Evaluate the model using .
- Predict the marks of a student who studies for 6 hours and has 85% attendance.
6. R Program
# Multiple Linear Regression in R # Create the dataset hours <- c(2, 3, 4, 5, 4, 6, 7, 5, 8, 9) attendance <- c(60, 65, 70, 72, 75, 80, 85, 78, 90, 95) marks <- c(45, 50, 55, 60, 58, 68, 75, 65, 82, 88) # Create a data frame student_data <- data.frame( hours, attendance, marks ) cat("Student Data:\n") print(student_data) # Create Multiple Linear Regression Model model <- lm(marks ~ hours + attendance, data = student_data) # Display the regression model cat("\nRegression Model:\n") print(model) # Display detailed summary cat("\nModel Summary:\n") summary(model) # Display regression coefficients cat("\nRegression Coefficients:\n") print(coefficients(model)) # Display R-squared value cat("\nR-squared Value:\n") print(summary(model)$r.squared) # Predict marks for a new student new_student <- data.frame( hours = 6, attendance = 85 ) predicted_marks <- predict( model, newdata = new_student ) cat("\nPredicted Marks =", predicted_marks, "\n") # Obtain predicted values for existing data predicted_values <- predict(model) cat("\nPredicted Values:\n") print(predicted_values) # Calculate residuals cat("\nResiduals:\n") print(residuals(model)) # Plot Actual vs Predicted Marks plot( marks, predicted_values, main = "Actual vs Predicted Marks", xlab = "Actual Marks", ylab = "Predicted Marks", pch = 19 ) # Add reference line abline(a = 0, b = 1, col = "blue", lwd = 2)Output
Student Data: hours attendance marks 1 2 60 45 2 3 65 50 3 4 70 55 4 5 72 60 5 4 75 58 6 6 80 68 7 7 85 75 8 5 78 65 9 8 90 82 10 9 95 88 Regression Model: Call: lm(formula = marks ~ hours + attendance, data = student_data) Coefficients: (Intercept) hours attendance 2.5767 3.4669 0.5669 Model Summary: Regression Coefficients: (Intercept) hours attendance 2.5767290 3.4669113 0.5668655 R-squared Value: [1] 0.9960543 Predicted Marks = 71.56176 Predicted Values: 1 2 3 4 5 6 7 43.52248 49.82372 56.12496 60.72560 58.95928 68.72743 75.02867 8 9 10 64.12679 81.32991 87.63115 Residuals: 1 2 3 4 5 1.47752036 0.17628168 -1.12495699 -0.72559927 -0.95928432 6 7 8 9 10 -0.72743434 -0.02867301 0.87320794 0.67008831 0.36884964
7. Explanation of the Program
Step 1: Create the Variables
hours <- c(2, 3, 4, 5, 4, 6, 7, 5, 8, 9) attendance <- c(60, 65, 70, 72, 75, 80, 85, 78, 90, 95) marks <- c(45, 50, 55, 60, 58, 68, 75, 65, 82, 88)
Three variables are created:
-
hours→ Number of hours studied -
attendance→ Attendance percentage -
marks→ Examination marks
Here:
Hours and Attendance are independent variables.
Marks is the dependent variable.
Step 2: Create a Data Frame
student_data <- data.frame( hours, attendance, marks )
This combines all variables into a single data frame.
Step 3: Build the Multiple Regression Model
model <- lm(marks ~ hours + attendance, data = student_data)
The lm() function creates the regression model.
The expression:
marks ~ hours + attendance
means:
Predict marks using hours studied and attendance percentage.
Step 4: Display the Model
print(model)
This displays the estimated regression equation in the form:
The actual coefficient values are calculated from the given data.
8. Understanding summary(model)
The following command:
summary(model)
provides detailed information about the regression model.
Important outputs include:
1. Regression Coefficients
The output contains coefficients for:
- Intercept
- Hours
- Attendance
These indicate how each independent variable influences the dependent variable.
2. p-values
The p-values help determine whether each predictor has a statistically significant relationship with the dependent variable.
A commonly used significance level is:
Generally:
- p-value < 0.05 → Predictor is statistically significant.
- p-value ≥ 0.05 → Insufficient evidence that the predictor has a statistically significant effect.
3. R-Squared Value
The coefficient of determination is denoted by:
It indicates how much variation in the dependent variable is explained by the independent variables.
For example:
means approximately 90% of the variation in marks is explained by hours studied and attendance.
In R:
summary(model)$r.squared
4. Adjusted R-Squared
Multiple regression also provides Adjusted R-squared.
Unlike ordinary , Adjusted considers the number of predictors included in the model.
It is particularly useful when comparing models containing different numbers of independent variables.
9. Prediction Using the Model
To predict marks for a student with:
- Hours studied = 6
- Attendance = 85%
we create:
new_student <- data.frame( hours = 6, attendance = 85 )
Then:
predict(model, newdata = new_student)
The model uses the regression equation to estimate the student's marks.
10. Actual vs Predicted Values
The following plot compares:
- Actual marks
- Predicted marks
plot(marks, predicted_values)
If the model performs well, the points should generally lie close to the reference line:
The reference line is added using:
abline(a = 0, b = 1)
11. Residuals
A residual is the difference between the actual value and the predicted value.
In R:
residuals(model)
Residual analysis helps us understand the errors made by the regression model.
12. Important Functions Used
| Function | Purpose |
|---|---|
data.frame() | Creates a data frame |
lm() | Builds a linear regression model |
summary() | Displays detailed model information |
coefficients() | Displays regression coefficients |
predict() | Predicts new values |
residuals() | Obtains prediction errors |
plot() | Creates a graph |
abline() | Adds a reference line |
13. Result
Thus, a Multiple Linear Regression model was successfully implemented using R to study the relationship between examination marks and multiple independent variables, namely hours studied and attendance percentage. The regression equation was obtained, the regression coefficients and model performance were analyzed, and the model was used to predict the marks of a new student.
Comments
Post a Comment