Logistic Regression in R

 

Experiment: Logistic Regression in R

1. Experiment Title

Implementation of Binary Logistic Regression Using R


2. Aim

To study and implement Binary Logistic Regression in R to predict a binary outcome based on an independent variable.


3. Objectives

After completing this experiment, students should be able to:

  1. Understand the concept of Binary Logistic Regression.
  2. Identify independent and dependent variables.
  3. Build a Logistic Regression model using R.
  4. Predict probabilities using the model.
  5. Convert probabilities into class labels.
  6. Evaluate the model using a confusion matrix and accuracy.
  7. Visualize the Logistic Regression probability curve.

4. Theory

4.1 What is Logistic Regression?

Logistic Regression is a statistical technique used to model the relationship between one or more independent variables and a categorical dependent variable.

When the dependent variable has only two possible outcomes, it is called Binary Logistic Regression.

Examples of binary outcomes:

  • Pass / Fail
  • Yes / No
  • Approved / Rejected
  • Present / Absent
  • Success / Failure

Although it is called regression, Logistic Regression is mainly used for classification problems.


4.2 Problem Variables

In this experiment, we predict whether a student will Pass or Fail based on the number of hours studied.

Independent Variable

X=Hours StudiedX = \text{Hours Studied}

Dependent Variable

Y=Examination ResultY = \text{Examination Result}

The result is represented as:

Y={0Fail1PassY= \begin{cases} 0 & \text{Fail}\\ 1 & \text{Pass} \end{cases}

4.3 Logistic Regression Model

Unlike linear regression, Logistic Regression predicts the probability that an observation belongs to a particular class.

The logistic regression model is based on the logit equation:

log⁡(p1−p)=β0+β1X\log\left(\frac{p}{1-p}\right) = \beta_0+\beta_1X

where:

  • pp = Probability of passing
  • 1−p1-p = Probability of failing
  • XX = Number of hours studied
  • β0\beta_0 = Intercept
  • β1\beta_1 = Regression coefficient

The probability is obtained using the sigmoid function:

p=11+e−(β0+β1X)p= \frac{1}{1+e^{-(\beta_0+\beta_1X)}}

The output probability always lies between:

0 and 10 \text{ and } 1

4.4 Classification Using a Threshold

The predicted probability is converted into a class label using a threshold.

A commonly used threshold is:

0.50.5

The classification rule is:

Predicted Class={1if p≥0.50if p<0.5\text{Predicted Class}= \begin{cases} 1 & \text{if } p\geq0.5\\ 0 & \text{if } p<0.5 \end{cases}

Therefore:

ProbabilityPredicted Result
≥ 0.5Pass
< 0.5Fail

5. Problem Statement

A teacher wants to predict whether a student will Pass or Fail based on the number of hours studied.

The following data were collected from 15 students.

StudentHours StudiedResult
11Fail
22Fail
32Fail
43Fail
53Pass
64Fail
74Pass
85Fail
95Pass
106Pass
116Pass
127Pass
137Fail
148Pass
158Pass

Using the given data:

  1. Build a Binary Logistic Regression model.
  2. Predict the probability of passing based on hours studied.
  3. Classify students as Pass or Fail using a threshold of 0.5.
  4. Create a confusion matrix.
  5. Calculate the accuracy of the model.
  6. Predict the result for a new student.
  7. Visualize the Logistic Regression probability curve.

6. R Program

# --------------------------------------------------
# Binary Logistic Regression in R
# --------------------------------------------------

# Create the dataset

hours <- c(1, 2, 2, 3, 3,
           4, 4, 5, 5, 6,
           6, 7, 7, 8, 8)

# 0 = Fail
# 1 = Pass

result <- c(0, 0, 0, 0, 1,
            0, 1, 0, 1, 1,
            1, 1, 0, 1, 1)


# Create a data frame

student_data <- data.frame(
  hours,
  result
)

cat("Student Data:\n")
print(student_data)


# --------------------------------------------------
# Build the Logistic Regression Model
# --------------------------------------------------

model <- glm(
  result ~ hours,
  data = student_data,
  family = binomial
)


# Display the model

cat("\nLogistic Regression Model:\n")
print(model)


# Display detailed summary

cat("\nModel Summary:\n")
summary(model)


# --------------------------------------------------
# Predict probabilities
# --------------------------------------------------

predicted_probability <- predict(
  model,
  type = "response"
)

cat("\nPredicted Probabilities:\n")
print(predicted_probability)


# --------------------------------------------------
# Convert probabilities into class labels
# --------------------------------------------------

predicted_class <- ifelse(
  predicted_probability >= 0.5,
  1,
  0
)

cat("\nPredicted Classes:\n")
print(predicted_class)


# --------------------------------------------------
# Add predictions to the data frame
# --------------------------------------------------

student_data$Predicted_Probability <- predicted_probability

student_data$Predicted_Class <- predicted_class


cat("\nActual and Predicted Results:\n")

print(student_data)


# --------------------------------------------------
# Create Confusion Matrix
# --------------------------------------------------

confusion_matrix <- table(
  Actual = result,
  Predicted = predicted_class
)

cat("\nConfusion Matrix:\n")

print(confusion_matrix)


# --------------------------------------------------
# Calculate Accuracy
# --------------------------------------------------

accuracy <- sum(diag(confusion_matrix)) /
            sum(confusion_matrix)

cat("\nAccuracy =", accuracy * 100, "%\n")


# --------------------------------------------------
# Predict Result for a New Student
# --------------------------------------------------

new_student <- data.frame(
  hours = 4.5
)


# Predict probability

new_probability <- predict(
  model,
  newdata = new_student,
  type = "response"
)


# Convert probability into class

new_class <- ifelse(
  new_probability >= 0.5,
  "Pass",
  "Fail"
)


cat("\nPrediction for New Student\n")

cat("Hours Studied =", 4.5, "\n")

cat(
  "Probability of Passing =",
  round(new_probability, 4),
  "\n"
)

cat(
  "Predicted Result =",
  new_class,
  "\n"
)


# --------------------------------------------------
# Plot Logistic Regression Curve
# --------------------------------------------------

# Generate values for the X-axis

x_values <- seq(
  min(hours),
  max(hours),
  length.out = 100
)


# Predict probabilities for the values

probabilities <- predict(
  model,
  newdata = data.frame(hours = x_values),
  type = "response"
)


# Plot actual observations

plot(
  hours,
  result,
  main = "Logistic Regression: Probability of Passing",
  xlab = "Hours Studied",
  ylab = "Probability of Passing",
  pch = 19,
  ylim = c(0, 1)
)


# Add Logistic Regression Curve

lines(
  x_values,
  probabilities,
  col = "blue",
  lwd = 2
)


# Add threshold line

abline(
  h = 0.5,
  lty = 2,
  col = "red"
)

Output

Student Data:
   hours result
1      1      0
2      2      0
3      2      0
4      3      0
5      3      1
6      4      0
7      4      1
8      5      0
9      5      1
10     6      1
11     6      1
12     7      1
13     7      0
14     8      1
15     8      1

Logistic Regression Model:

Call:  glm(formula = result ~ hours, family = binomial, data = student_data)

Coefficients:
(Intercept)        hours  
    -2.9414       0.6615  

Degrees of Freedom: 14 Total (i.e. Null);  13 Residual
Null Deviance:	    20.73 
Residual Deviance: 15.43 	AIC: 19.43

Model Summary:

Predicted Probabilities:
         1          2          3          4          5          6 
0.09280252 0.16542512 0.16542512 0.27749502 0.27749502 0.42667289 
         7          8          9         10         11         12 
0.42667289 0.59050266 0.59050266 0.73643602 0.73643602 0.84409376 
        13         14         15 
0.84409376 0.91297327 0.91297327 

Predicted Classes:
 1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 
 0  0  0  0  0  0  0  1  1  1  1  1  1  1  1 

Actual and Predicted Results:
   hours result Predicted_Probability Predicted_Class
1      1      0            0.09280252               0
2      2      0            0.16542512               0
3      2      0            0.16542512               0
4      3      0            0.27749502               0
5      3      1            0.27749502               0
6      4      0            0.42667289               0
7      4      1            0.42667289               0
8      5      0            0.59050266               1
9      5      1            0.59050266               1
10     6      1            0.73643602               1
11     6      1            0.73643602               1
12     7      1            0.84409376               1
13     7      0            0.84409376               1
14     8      1            0.91297327               1
15     8      1            0.91297327               1

Confusion Matrix:
      Predicted
Actual 0 1
     0 5 2
     1 2 6

Accuracy = 73.33333 %

Prediction for New Student
Hours Studied = 4.5 
Probability of Passing = 0.5088 
Predicted Result = Pass





7. Explanation of the Program

Step 1: Creating the Dataset

hours <- c(1, 2, 2, 3, 3,
           4, 4, 5, 5, 6,
           6, 7, 7, 8, 8)

This vector contains the number of hours studied by each student.

result <- c(0, 0, 0, 0, 1,
            0, 1, 0, 1, 1,
            1, 1, 0, 1, 1)

The result is represented as:

  • 0 → Fail
  • 1 → Pass

The data contain some overlap between Pass and Fail, making it suitable for demonstrating Logistic Regression.


Step 2: Creating a Data Frame

student_data <- data.frame(hours, result)

This combines the variables into a structured dataset.


Step 3: Building the Logistic Regression Model

model <- glm(
  result ~ hours,
  data = student_data,
  family = binomial
)

The glm() function creates a Generalized Linear Model.

The argument:

family = binomial

specifies that the model is a Binary Logistic Regression model.

The expression:

result ~ hours

means:

Predict the examination result using the number of hours studied.


8. Understanding summary(model)

The command:

summary(model)

displays detailed information about the Logistic Regression model.

It includes:

  • Regression coefficients
  • Standard errors
  • z-values
  • p-values
  • Deviance statistics
  • Model fitting information

Interpretation of the Coefficient

If the coefficient of hours is positive, it generally indicates:

As the number of hours studied increases, the probability of passing increases.


9. Predicting Probabilities

predicted_probability <- predict(
  model,
  type = "response"
)

The argument:

type = "response"

returns probabilities between:

0 and 10 \text{ and } 1

For example:

  • 0.20 → 20% probability of passing
  • 0.75 → 75% probability of passing

10. Converting Probability into Classes

The probabilities are converted into class labels using:

predicted_class <- ifelse(
  predicted_probability >= 0.5,
  1,
  0
)

The threshold used is 0.5.

ProbabilityClass
≥ 0.5Pass (1)
< 0.5Fail (0)

11. Confusion Matrix

The confusion matrix is created using:

confusion_matrix <- table(
  Actual = result,
  Predicted = predicted_class
)

It compares:

  • The actual result
  • The predicted result

A typical confusion matrix has the following interpretation:

Actual     PredictedMeaning
Fail    FailCorrect prediction
Fail    PassIncorrect prediction
Pass    PassCorrect prediction
Pass    FailIncorrect prediction

12. Accuracy of the Model

Accuracy is calculated as:

Accuracy=Correct PredictionsTotal Predictions×100\text{Accuracy} = \frac{\text{Correct Predictions}} {\text{Total Predictions}} \times100

In the program:

accuracy <- sum(diag(confusion_matrix)) /
            sum(confusion_matrix)

The diagonal elements represent the correctly classified observations.


13. Predicting a New Student's Result

The program predicts the result for a student who studies for:

4.5 hours4.5 \text{ hours}
new_student <- data.frame(hours = 4.5)

The model first calculates the probability of passing and then converts the probability into:

  • Pass, or
  • Fail

using the threshold value of 0.5.


14. Logistic Regression Curve

The following part of the program creates the probability curve:

lines(
  x_values,
  probabilities
)

The curve shows how the probability of passing changes as the number of hours studied increases.

Interpretation

  • Lower study hours generally correspond to a lower probability of passing.
  • Higher study hours generally correspond to a higher probability of passing.
  • The curve has the characteristic S-shape of the Logistic Regression model.

The horizontal line:

abline(h = 0.5)

represents the classification threshold.

Above the threshold:

Predicted as Pass

Below the threshold:

Predicted as Fail


15. Important R Functions Used

FunctionPurpose
data.frame()    Creates a data frame
glm()    Creates a generalized linear model
family = binomial    Specifies Binary Logistic Regression
summary()    Displays detailed model results
predict()    Predicts probabilities
ifelse()    Converts probabilities into classes
table()    Creates a confusion matrix
diag()    Extracts diagonal values
plot()    Creates a graph
lines()    Adds the logistic curve
abline()    Adds the classification threshold

16. Result

A Binary Logistic Regression model was successfully implemented using R to predict whether a student would Pass or Fail based on the number of hours studied. The model was used to calculate predicted probabilities, classify students using a threshold value of 0.5, evaluate the model using a confusion matrix and accuracy, and visualize the relationship using a Logistic Regression probability curve.

Comments

Popular posts from this blog

Statistical Methods Lab ( R Language) PCCBL308 Semester 3 KTU BTech CB and CU 2024 Scheme - Dr Binu V P

Programs in R - using control statements - Assignment 2

Programs to try using Functions in R - Assignment 3