Linear Regression in R

 

Experiment: Linear Regression in R

1. Experiment Title

Implementation of Simple Linear Regression Using R


2. Aim

To study and implement Simple Linear Regression in R to analyze the relationship between an independent variable and a dependent variable and to predict the dependent variable.


3. Objectives

After completing this experiment, students should be able to:

  1. Understand the concept of simple linear regression.
  2. Identify independent and dependent variables.
  3. build a linear regression model using R.
  4. Interpret the regression coefficients.
  5. Predict new values using the regression model.
  6. Visualize the regression line.
  7. Evaluate the performance of the regression model.

4. Theory

4.1 What is Linear Regression?

Linear Regression is a statistical method used to study the relationship between two numerical variables.

It helps us understand how a dependent variable (YY) changes when an independent variable (XX) changes.

For example:

How do student marks change based on the number of hours studied?

Here:

  • Independent Variable (XX) → Hours Studied
  • Dependent Variable (YY) → Marks Obtained

The relationship is represented using a straight-line equation:

where:

  • β0\beta_0 = Intercept
  • β1\beta_1 = Slope (Regression Coefficient)
  • XX = Independent variable
  • YY = Dependent variable

The regression line is the best-fitting straight line that represents the relationship between XX and YY.


4.2 Simple Linear Regression Model

For a dataset containing several observations:

Yi=β0+β1Xi+ϵiY_i=\beta_0+\beta_1X_i+\epsilon_i

where:

  • YiY_i = Actual dependent variable
  • XiX_i = Independent variable
  • β0\beta_0 = Intercept
  • β1\beta_1 = Regression coefficient
  • ϵi\epsilon_i = Error or residual

4.3 Method of Least Squares

The regression line is obtained using the Least Squares Method.

The objective is to minimize the sum of squared differences between the actual and predicted values.

The regression coefficients are calculated as:

Slope

β1=∑(Xi−Xˉ)(Yi−Yˉ)∑(Xi−Xˉ)2\beta_1= \frac{\sum(X_i-\bar{X})(Y_i-\bar{Y})} {\sum(X_i-\bar{X})^2}

Intercept

β0=Yˉ−β1Xˉ\beta_0= \bar{Y}-\beta_1\bar{X}

5. Problem Statement

A teacher wants to study the relationship between the number of hours studied by students and their marks obtained in an examination.

The following data were collected:

StudentHours StudiedMarks Obtained
1142
2248
3355
4460
5565
6672
7778
8885

The task is to:

  1. Build a simple linear regression model.
  2. Study the relationship between hours studied and marks obtained.
  3. Obtain the regression equation.
  4. Predict the marks of a student who studies for 9 hours.
  5. Visualize the regression line.

6. R Program

# Simple Linear Regression in R

# Create the dataset

hours <- c(1, 2, 3, 4, 5, 6, 7, 8)

marks <- c(42, 48, 55, 60, 65, 72, 78, 85)


# Create a data frame

student_data <- data.frame(hours, marks)

print(student_data)


# Create the Linear Regression Model

model <- lm(marks ~ hours, data = student_data)


# Display the regression model

print(model)


# Display detailed model summary
print("Summary")
print(summary(model))


# Obtain regression coefficients

print(coefficients(model))


# Predict marks for a student studying 9 hours

new_data <- data.frame(hours = 9)

prediction <- predict(model, newdata = new_data)

cat("Predicted Marks for 9 Hours of Study =", prediction)


# Plot the original data

plot(student_data$hours,
     student_data$marks,
     main = "Hours Studied vs Marks",
     xlab = "Hours Studied",
     ylab = "Marks Obtained",
     pch = 19)


# Add regression line

abline(model, col = "blue", lwd = 2)


 Output


7. Explanation of the Program

Step 1: Create the Variables

hours <- c(1, 2, 3, 4, 5, 6, 7, 8)

marks <- c(42, 48, 55, 60, 65, 72, 78, 85)

Two numerical vectors are created:

  • hours → Independent variable
  • marks → Dependent variable

Step 2: Create a Data Frame

student_data <- data.frame(hours, marks)

The two variables are combined into a data frame.


Step 3: Create the Regression Model

model <- lm(marks ~ hours, data = student_data)

The lm() function is used to create a linear model.

The notation:

marks ~ hours

means:

Predict marks using hours.


Step 4: Display the Model

print(model)

This displays the estimated regression equation.

The output will be approximately:

Marks=β0+β1(Hours)\text{Marks}=\beta_0+\beta_1(\text{Hours})

Step 5: Get Detailed Results

summary(model)

The summary provides:

  • Regression coefficients
  • Standard errors
  • t-values
  • p-values
  • Residual information
  • R-squared value
  • Adjusted R-squared value
  • F-statistic

Step 6: Make a Prediction

new_data <- data.frame(hours = 9)

prediction <- predict(model, newdata = new_data)

The predict() function uses the trained regression model to estimate the marks for a student studying 9 hours.


Step 7: Plot the Regression Line

plot(student_data$hours,
     student_data$marks)

This creates a scatter plot.

abline(model)

This adds the best-fitting regression line to the plot.


8. Expected Output

The regression model will produce an equation approximately of the form:

Marks=35.5+6.04(Hours)\boxed{ \text{Marks}=35.5+6.04(\text{Hours}) }

The exact values can be obtained by running the program in R.

Interpretation

The slope is approximately:

6.046.04

This means:

For every additional hour of study, the predicted marks increase by approximately 6 marks.


9. Model Evaluation

9.1 R-Squared Value

The R-squared (R2R^2) value indicates how well the regression model explains the variation in the dependent variable.

summary(model)$r.squared

A value closer to 1 indicates that the model explains a large proportion of the variation in the dependent variable.


9.2 Predicted Values

predict(model)

This gives the predicted marks for the existing observations.


9.3 Residuals

Residuals represent the difference between actual and predicted values.

Residual=Actual Value−Predicted Value\text{Residual}= \text{Actual Value}-\text{Predicted Value}

In R:

residuals(model)

10. Result

A Simple Linear Regression model was successfully implemented using R to study the relationship between hours studied and marks obtained. The regression equation was obtained, the relationship between the variables was analyzed, and the model was used to predict marks for a new value of the independent variable. The regression line was also visualized using a scatter plot.


Important R Functions Used

FunctionPurpose
lm()    Creates the linear regression model
summary()    Displays detailed regression results
coefficients()    Obtains regression coefficients
predict()    Predicts new values
plot()    Creates a scatter plot
abline()    Adds the regression line
residuals()    Obtains prediction errors

Comments

Popular posts from this blog

Statistical Methods Lab ( R Language) PCCBL308 Semester 3 KTU BTech CB and CU 2024 Scheme - Dr Binu V P

Programs in R - using control statements - Assignment 2

Programs to try using Functions in R - Assignment 3