Logistic Regression in R
Experiment: Logistic Regression in R
1. Experiment Title
Implementation of Binary Logistic Regression Using R
2. Aim
To study and implement Binary Logistic Regression in R to predict a binary outcome based on an independent variable.
3. Objectives
After completing this experiment, students should be able to:
- Understand the concept of Binary Logistic Regression.
- Identify independent and dependent variables.
- Build a Logistic Regression model using R.
- Predict probabilities using the model.
- Convert probabilities into class labels.
- Evaluate the model using a confusion matrix and accuracy.
- Visualize the Logistic Regression probability curve.
4. Theory
4.1 What is Logistic Regression?
Logistic Regression is a statistical technique used to model the relationship between one or more independent variables and a categorical dependent variable.
When the dependent variable has only two possible outcomes, it is called Binary Logistic Regression.
Examples of binary outcomes:
- Pass / Fail
- Yes / No
- Approved / Rejected
- Present / Absent
- Success / Failure
Although it is called regression, Logistic Regression is mainly used for classification problems.
4.2 Problem Variables
In this experiment, we predict whether a student will Pass or Fail based on the number of hours studied.
Independent Variable
Dependent Variable
The result is represented as:
4.3 Logistic Regression Model
Unlike linear regression, Logistic Regression predicts the probability that an observation belongs to a particular class.
The logistic regression model is based on the logit equation:
where:
- = Probability of passing
- = Probability of failing
- = Number of hours studied
- = Intercept
- = Regression coefficient
The probability is obtained using the sigmoid function:
The output probability always lies between:
4.4 Classification Using a Threshold
The predicted probability is converted into a class label using a threshold.
A commonly used threshold is:
The classification rule is:
Therefore:
| Probability | Predicted Result |
|---|---|
| ≥ 0.5 | Pass |
| < 0.5 | Fail |
5. Problem Statement
A teacher wants to predict whether a student will Pass or Fail based on the number of hours studied.
The following data were collected from 15 students.
| Student | Hours Studied | Result |
|---|---|---|
| 1 | 1 | Fail |
| 2 | 2 | Fail |
| 3 | 2 | Fail |
| 4 | 3 | Fail |
| 5 | 3 | Pass |
| 6 | 4 | Fail |
| 7 | 4 | Pass |
| 8 | 5 | Fail |
| 9 | 5 | Pass |
| 10 | 6 | Pass |
| 11 | 6 | Pass |
| 12 | 7 | Pass |
| 13 | 7 | Fail |
| 14 | 8 | Pass |
| 15 | 8 | Pass |
Using the given data:
- Build a Binary Logistic Regression model.
- Predict the probability of passing based on hours studied.
- Classify students as Pass or Fail using a threshold of 0.5.
- Create a confusion matrix.
- Calculate the accuracy of the model.
- Predict the result for a new student.
- Visualize the Logistic Regression probability curve.
6. R Program
# -------------------------------------------------- # Binary Logistic Regression in R # -------------------------------------------------- # Create the dataset hours <- c(1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, 7, 7, 8, 8) # 0 = Fail # 1 = Pass result <- c(0, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1) # Create a data frame student_data <- data.frame( hours, result ) cat("Student Data:\n") print(student_data) # -------------------------------------------------- # Build the Logistic Regression Model # -------------------------------------------------- model <- glm( result ~ hours, data = student_data, family = binomial ) # Display the model cat("\nLogistic Regression Model:\n") print(model) # Display detailed summary cat("\nModel Summary:\n") summary(model) # -------------------------------------------------- # Predict probabilities # -------------------------------------------------- predicted_probability <- predict( model, type = "response" ) cat("\nPredicted Probabilities:\n") print(predicted_probability) # -------------------------------------------------- # Convert probabilities into class labels # -------------------------------------------------- predicted_class <- ifelse( predicted_probability >= 0.5, 1, 0 ) cat("\nPredicted Classes:\n") print(predicted_class) # -------------------------------------------------- # Add predictions to the data frame # -------------------------------------------------- student_data$Predicted_Probability <- predicted_probability student_data$Predicted_Class <- predicted_class cat("\nActual and Predicted Results:\n") print(student_data) # -------------------------------------------------- # Create Confusion Matrix # -------------------------------------------------- confusion_matrix <- table( Actual = result, Predicted = predicted_class ) cat("\nConfusion Matrix:\n") print(confusion_matrix) # -------------------------------------------------- # Calculate Accuracy # -------------------------------------------------- accuracy <- sum(diag(confusion_matrix)) / sum(confusion_matrix) cat("\nAccuracy =", accuracy * 100, "%\n") # -------------------------------------------------- # Predict Result for a New Student # -------------------------------------------------- new_student <- data.frame( hours = 4.5 ) # Predict probability new_probability <- predict( model, newdata = new_student, type = "response" ) # Convert probability into class new_class <- ifelse( new_probability >= 0.5, "Pass", "Fail" ) cat("\nPrediction for New Student\n") cat("Hours Studied =", 4.5, "\n") cat( "Probability of Passing =", round(new_probability, 4), "\n" ) cat( "Predicted Result =", new_class, "\n" ) # -------------------------------------------------- # Plot Logistic Regression Curve # -------------------------------------------------- # Generate values for the X-axis x_values <- seq( min(hours), max(hours), length.out = 100 ) # Predict probabilities for the values probabilities <- predict( model, newdata = data.frame(hours = x_values), type = "response" ) # Plot actual observations plot( hours, result, main = "Logistic Regression: Probability of Passing", xlab = "Hours Studied", ylab = "Probability of Passing", pch = 19, ylim = c(0, 1) ) # Add Logistic Regression Curve lines( x_values, probabilities, col = "blue", lwd = 2 ) # Add threshold line abline( h = 0.5, lty = 2, col = "red" )Output
Student Data: hours result 1 1 0 2 2 0 3 2 0 4 3 0 5 3 1 6 4 0 7 4 1 8 5 0 9 5 1 10 6 1 11 6 1 12 7 1 13 7 0 14 8 1 15 8 1 Logistic Regression Model: Call: glm(formula = result ~ hours, family = binomial, data = student_data) Coefficients: (Intercept) hours -2.9414 0.6615 Degrees of Freedom: 14 Total (i.e. Null); 13 Residual Null Deviance: 20.73 Residual Deviance: 15.43 AIC: 19.43 Model Summary: Predicted Probabilities: 1 2 3 4 5 6 0.09280252 0.16542512 0.16542512 0.27749502 0.27749502 0.42667289 7 8 9 10 11 12 0.42667289 0.59050266 0.59050266 0.73643602 0.73643602 0.84409376 13 14 15 0.84409376 0.91297327 0.91297327 Predicted Classes: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 Actual and Predicted Results: hours result Predicted_Probability Predicted_Class 1 1 0 0.09280252 0 2 2 0 0.16542512 0 3 2 0 0.16542512 0 4 3 0 0.27749502 0 5 3 1 0.27749502 0 6 4 0 0.42667289 0 7 4 1 0.42667289 0 8 5 0 0.59050266 1 9 5 1 0.59050266 1 10 6 1 0.73643602 1 11 6 1 0.73643602 1 12 7 1 0.84409376 1 13 7 0 0.84409376 1 14 8 1 0.91297327 1 15 8 1 0.91297327 1 Confusion Matrix: Predicted Actual 0 1 0 5 2 1 2 6 Accuracy = 73.33333 % Prediction for New Student Hours Studied = 4.5 Probability of Passing = 0.5088 Predicted Result = Pass
7. Explanation of the Program
Step 1: Creating the Dataset
hours <- c(1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, 7, 7, 8, 8)
This vector contains the number of hours studied by each student.
result <- c(0, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1)
The result is represented as:
-
0→ Fail -
1→ Pass
The data contain some overlap between Pass and Fail, making it suitable for demonstrating Logistic Regression.
Step 2: Creating a Data Frame
student_data <- data.frame(hours, result)
This combines the variables into a structured dataset.
Step 3: Building the Logistic Regression Model
model <- glm( result ~ hours, data = student_data, family = binomial )
The glm() function creates a Generalized Linear Model.
The argument:
family = binomial
specifies that the model is a Binary Logistic Regression model.
The expression:
result ~ hours
means:
Predict the examination result using the number of hours studied.
8. Understanding summary(model)
The command:
summary(model)
displays detailed information about the Logistic Regression model.
It includes:
- Regression coefficients
- Standard errors
- z-values
- p-values
- Deviance statistics
- Model fitting information
Interpretation of the Coefficient
If the coefficient of hours is positive, it generally indicates:
As the number of hours studied increases, the probability of passing increases.
9. Predicting Probabilities
predicted_probability <- predict( model, type = "response" )
The argument:
type = "response"
returns probabilities between:
For example:
-
0.20→ 20% probability of passing -
0.75→ 75% probability of passing
10. Converting Probability into Classes
The probabilities are converted into class labels using:
predicted_class <- ifelse( predicted_probability >= 0.5, 1, 0 )
The threshold used is 0.5.
| Probability | Class |
|---|---|
| ≥ 0.5 | Pass (1) |
| < 0.5 | Fail (0) |
11. Confusion Matrix
The confusion matrix is created using:
confusion_matrix <- table( Actual = result, Predicted = predicted_class )
It compares:
- The actual result
- The predicted result
A typical confusion matrix has the following interpretation:
| Actual | Predicted | Meaning |
|---|---|---|
| Fail | Fail | Correct prediction |
| Fail | Pass | Incorrect prediction |
| Pass | Pass | Correct prediction |
| Pass | Fail | Incorrect prediction |
12. Accuracy of the Model
Accuracy is calculated as:
In the program:
accuracy <- sum(diag(confusion_matrix)) / sum(confusion_matrix)
The diagonal elements represent the correctly classified observations.
13. Predicting a New Student's Result
The program predicts the result for a student who studies for:
new_student <- data.frame(hours = 4.5)
The model first calculates the probability of passing and then converts the probability into:
- Pass, or
- Fail
using the threshold value of 0.5.
14. Logistic Regression Curve
The following part of the program creates the probability curve:
lines( x_values, probabilities )
The curve shows how the probability of passing changes as the number of hours studied increases.
Interpretation
- Lower study hours generally correspond to a lower probability of passing.
- Higher study hours generally correspond to a higher probability of passing.
- The curve has the characteristic S-shape of the Logistic Regression model.
The horizontal line:
abline(h = 0.5)
represents the classification threshold.
Above the threshold:
Predicted as Pass
Below the threshold:
Predicted as Fail
15. Important R Functions Used
| Function | Purpose |
|---|---|
data.frame() | Creates a data frame |
glm() | Creates a generalized linear model |
family = binomial | Specifies Binary Logistic Regression |
summary() | Displays detailed model results |
predict() | Predicts probabilities |
ifelse() | Converts probabilities into classes |
table() | Creates a confusion matrix |
diag() | Extracts diagonal values |
plot() | Creates a graph |
lines() | Adds the logistic curve |
abline() | Adds the classification threshold |
16. Result
A Binary Logistic Regression model was successfully implemented using R to predict whether a student would Pass or Fail based on the number of hours studied. The model was used to calculate predicted probabilities, classify students using a threshold value of 0.5, evaluate the model using a confusion matrix and accuracy, and visualize the relationship using a Logistic Regression probability curve.
Comments
Post a Comment