Regression Analysis Using a CSV Dataset Downloaded from Kaggle in R

 

Experiment: Regression Analysis Using a CSV Dataset Downloaded from Kaggle in R

1. Experiment Title

Implementation of Regression Analysis Using a Real-World CSV Dataset Downloaded from Kaggle


2. Aim

To import a CSV dataset downloaded from Kaggle into R, preprocess and explore the dataset, build a regression model, evaluate its performance, and make predictions using the developed model.


3. Objectives

After completing this experiment, students should be able to:

  1. Download a real-world dataset from Kaggle.
  2. Import a CSV file into R.
  3. Explore the structure and contents of the dataset.
  4. Handle missing values.
  5. Select suitable independent and dependent variables.
  6. Build a regression model using lm().
  7. Interpret regression coefficients.
  8. Calculate predicted values and residuals.
  9. Evaluate the model using R2R^2, Adjusted R2R^2, MAE, and RMSE.
  10. Visualize actual and predicted values.
  11. Make predictions for new observations.

4. Theory

4.1 Regression Analysis

Regression analysis is a statistical technique used to study the relationship between a dependent variable and one or more independent variables.

Simple Linear Regression

When one independent variable is used:

Y=β0+β1X+ϵY = \beta_0 + \beta_1X + \epsilon

where:

  • YY = Dependent variable
  • XX = Independent variable
  • β0\beta_0 = Intercept
  • β1\beta_1 = Regression coefficient
  • ϵ\epsilon = Error term

Multiple Linear Regression

When two or more independent variables are used:

Y=β0+β1X1+β2X2+⋯+βnXn+ϵY = \beta_0+\beta_1X_1+\beta_2X_2+\cdots+\beta_nX_n+\epsilon

The regression model estimates the relationship between the predictor variables and the target variable.


5. Problem Statement

Download a suitable CSV dataset from Kaggle containing numerical variables.

Import the dataset into R and perform regression analysis to predict a numerical target variable.

Students should perform the following:

  1. Import the CSV file.
  2. Explore the dataset.
  3. Check for missing values.
  4. Select a numerical target variable.
  5. Select suitable predictor variables.
  6. Build a regression model.
  7. Analyze the regression results.
  8. Calculate predictions and residuals.
  9. Evaluate the model.
  10. Visualize the results.
  11. Predict the value for a new observation.

6. Suggested Dataset

A very suitable dataset for this experiment is a House Price Prediction dataset from Kaggle.

Download : https://www.kaggle.com/datasets/yasserh/housing-prices-dataset/data

The dataset should contain variables similar to:

VariableDescription
area    Area of the house
bedrooms    Number of bedrooms
bathrooms    Number of bathrooms
stories    Number of floors
parking    Number of parking spaces
price    Price of the house

Regression Problem

Predict the price of a house based on its area, number of bedrooms, number of bathrooms, and parking facilities.

This makes an excellent real-world Multiple Linear Regression experiment.


7. Software Requirements

  • R
  • RStudio
  • CSV dataset downloaded from Kaggle

8. R Program

The following program assumes that the downloaded file is named:

Housing.csv

and is stored in the working directory.

# -------------------------------------------------------
# Regression Analysis Using a CSV Dataset from Kaggle
# -------------------------------------------------------


# -------------------------------------------------------
# Step 1: Import the CSV File
# -------------------------------------------------------

data <- read.csv("Housing.csv")


# -------------------------------------------------------
# Step 2: Display the First Few Records
# -------------------------------------------------------

cat("First Few Records:\n")

head(data)


# -------------------------------------------------------
# Step 3: Display Dataset Structure
# -------------------------------------------------------

cat("\nStructure of Dataset:\n")

str(data)


# -------------------------------------------------------
# Step 4: Display Dataset Dimensions
# -------------------------------------------------------

cat("\nNumber of Rows and Columns:\n")

dim(data)


# -------------------------------------------------------
# Step 5: Display Column Names
# -------------------------------------------------------

cat("\nColumn Names:\n")

names(data)


# -------------------------------------------------------
# Step 6: Display Summary Statistics
# -------------------------------------------------------

cat("\nSummary Statistics:\n")

summary(data)


# -------------------------------------------------------
# Step 7: Check for Missing Values
# -------------------------------------------------------

cat("\nMissing Values in Each Column:\n")

colSums(is.na(data))


# -------------------------------------------------------
# Step 8: Remove Missing Values
# -------------------------------------------------------

data <- na.omit(data)


# -------------------------------------------------------
# Step 9: Select Required Variables
# -------------------------------------------------------

house_data <- data[, c(
  "price",
  "area",
  "bedrooms",
  "bathrooms",
  "parking"
)]


# -------------------------------------------------------
# Step 10: Build Multiple Linear Regression Model
# -------------------------------------------------------

model <- lm(
  price ~ area + bedrooms + bathrooms + parking,
  data = house_data
)


# -------------------------------------------------------
# Step 11: Display Model Summary
# -------------------------------------------------------

cat("\nRegression Model Summary:\n")

summary(model)


# -------------------------------------------------------
# Step 12: Display Regression Coefficients
# -------------------------------------------------------

cat("\nRegression Coefficients:\n")

coef(model)


# -------------------------------------------------------
# Step 13: Calculate Predicted Values
# -------------------------------------------------------

predicted_price <- predict(model)


# -------------------------------------------------------
# Step 14: Calculate Residuals
# -------------------------------------------------------

residuals_value <- residuals(model)


# -------------------------------------------------------
# Step 15: Add Predictions to Dataset
# -------------------------------------------------------

house_data$Predicted_Price <- predicted_price

house_data$Residual <- residuals_value


cat("\nActual and Predicted Values:\n")

head(house_data)


# -------------------------------------------------------
# Step 16: Calculate R-Squared Values
# -------------------------------------------------------

model_summary <- summary(model)


cat(
  "\nR-Squared =",
  model_summary$r.squared,
  "\n"
)


cat(
  "Adjusted R-Squared =",
  model_summary$adj.r.squared,
  "\n"
)


# -------------------------------------------------------
# Step 17: Calculate MAE
# -------------------------------------------------------

mae <- mean(
  abs(
    house_data$price -
    house_data$Predicted_Price
  )
)


cat("\nMAE =", mae, "\n")


# -------------------------------------------------------
# Step 18: Calculate RMSE
# -------------------------------------------------------

rmse <- sqrt(
  mean(
    (
      house_data$price -
      house_data$Predicted_Price
    )^2
  )
)


cat("RMSE =", rmse, "\n")


# -------------------------------------------------------
# Step 19: Plot Actual vs Predicted Prices
# -------------------------------------------------------

plot(
  house_data$price,
  house_data$Predicted_Price,
  main = "Actual vs Predicted House Prices",
  xlab = "Actual Price",
  ylab = "Predicted Price",
  pch = 19
)


# Add reference line

abline(
  a = 0,
  b = 1,
  col = "blue",
  lwd = 2
)


# -------------------------------------------------------
# Step 20: Plot Area vs Price
# -------------------------------------------------------

plot(
  house_data$area,
  house_data$price,
  main = "House Area vs Price",
  xlab = "Area",
  ylab = "Price",
  pch = 19
)


# -------------------------------------------------------
# Step 21: Predict Price of a New House
# -------------------------------------------------------

new_house <- data.frame(
  area = 5000,
  bedrooms = 3,
  bathrooms = 2,
  parking = 2
)


new_prediction <- predict(
  model,
  newdata = new_house
)


cat(
  "\nPredicted Price of New House =",
  new_prediction,
  "\n"
)

Output

First Few Records:

Structure of Dataset:
'data.frame':	545 obs. of  13 variables:
 $ price           : int  13300000 12250000 12250000 12215000 11410000 10850000 10150000 10150000 9870000 9800000 ...
 $ area            : int  7420 8960 9960 7500 7420 7500 8580 16200 8100 5750 ...
 $ bedrooms        : int  4 4 3 4 4 3 4 5 4 3 ...
 $ bathrooms       : int  2 4 2 2 1 3 3 3 1 2 ...
 $ stories         : int  3 4 2 2 2 1 4 2 2 4 ...
 $ mainroad        : chr  "yes" "yes" "yes" "yes" ...
 $ guestroom       : chr  "no" "no" "no" "no" ...
 $ basement        : chr  "no" "no" "yes" "yes" ...
 $ hotwaterheating : chr  "no" "no" "no" "no" ...
 $ airconditioning : chr  "yes" "yes" "no" "yes" ...
 $ parking         : int  2 3 2 3 2 2 2 0 2 1 ...
 $ prefarea        : chr  "yes" "no" "yes" "yes" ...
 $ furnishingstatus: chr  "furnished" "furnished" "semi-furnished" "furnished" ...

Number of Rows and Columns:

Column Names:

Summary Statistics:

Missing Values in Each Column:

Regression Model Summary:

Regression Coefficients:

Actual and Predicted Values:

R-Squared = 0.5101313 
Adjusted R-Squared = 0.5065026 

MAE = 988779.9 
RMSE = 1307931 

Predicted Price of New House = 6142727 
> source("~/test.R")
First Few Records:

Structure of Dataset:
'data.frame':	545 obs. of  13 variables:
 $ price           : int  13300000 12250000 12250000 12215000 11410000 10850000 10150000 10150000 9870000 9800000 ...
 $ area            : int  7420 8960 9960 7500 7420 7500 8580 16200 8100 5750 ...
 $ bedrooms        : int  4 4 3 4 4 3 4 5 4 3 ...
 $ bathrooms       : int  2 4 2 2 1 3 3 3 1 2 ...
 $ stories         : int  3 4 2 2 2 1 4 2 2 4 ...
 $ mainroad        : chr  "yes" "yes" "yes" "yes" ...
 $ guestroom       : chr  "no" "no" "no" "no" ...
 $ basement        : chr  "no" "no" "yes" "yes" ...
 $ hotwaterheating : chr  "no" "no" "no" "no" ...
 $ airconditioning : chr  "yes" "yes" "no" "yes" ...
 $ parking         : int  2 3 2 3 2 2 2 0 2 1 ...
 $ prefarea        : chr  "yes" "no" "yes" "yes" ...
 $ furnishingstatus: chr  "furnished" "furnished" "semi-furnished" "furnished" ...

Number of Rows and Columns:

Column Names:

Summary Statistics:

Missing Values in Each Column:

Regression Model Summary:

Regression Coefficients:

Actual and Predicted Values:

R-Squared = 0.5101313 
Adjusted R-Squared = 0.5065026 

MAE = 988779.9 
RMSE = 1307931 

Predicted Price of New House = 6142727






9. Explanation of the Important Steps

Step 1: Importing the CSV File

data <- read.csv("Housing.csv")

The read.csv() function imports the CSV file into R as a data frame.

Important: The file name must exactly match the downloaded CSV file name.


Step 2: Exploring the Dataset

Students should use:

head(data)

to view the first few records.

str(data)

to understand:

  • Variable names
  • Data types
  • Dataset structure
summary(data)

to obtain descriptive statistics.


Step 3: Checking Missing Values

colSums(is.na(data))

This checks the number of missing values in every column.

For this basic experiment:

data <- na.omit(data)

removes rows containing missing values.


Step 4: Building the Regression Model

model <- lm(
  price ~ area + bedrooms + bathrooms + parking,
  data = house_data
)

This means:

Predict house price using area, bedrooms, bathrooms, and parking.

The model has the general form:

Price^=β0+β1(Area)+β2(Bedrooms)+β3(Bathrooms)+β4(Parking)\widehat{Price} = \beta_0+ \beta_1(Area)+ \beta_2(Bedrooms)+ \beta_3(Bathrooms)+ \beta_4(Parking)

10. Predicted Values

The predicted house prices are obtained using:

predicted_price <- predict(model)

Each house will have:

  • Actual price
  • Predicted price

11. Residuals

The residual is:

Residual=Actual Value−Predicted ValueResidual = Actual\ Value - Predicted\ Value

In R:

residuals(model)

A good regression model should generally have residuals randomly distributed around zero.


12. Model Evaluation Measures

R-Squared

R2R^2 represents the proportion of variation in the target variable explained by the regression model.

For example:

R2=0.75R^2 = 0.75

means approximately:

75% of the variation in house prices is explained by the selected variables.


Adjusted R-Squared

Adjusted R2R^2 is particularly useful in Multiple Regression because it considers the number of predictors.


MAE — Mean Absolute Error

MAE=1n∑∣Y−Y^∣MAE= \frac{1}{n} \sum |Y-\hat{Y}|

It measures the average absolute prediction error.

Lower MAE indicates better predictions.


RMSE — Root Mean Square Error

RMSE=1n∑(Y−Y^)2RMSE= \sqrt{ \frac{1}{n} \sum(Y-\hat{Y})^2 }

RMSE gives greater importance to large errors.

Lower RMSE indicates better predictions.


13. Interpretation of the Actual vs Predicted Plot

The graph contains:

  • X-axis → Actual house prices
  • Y-axis → Predicted house prices

The line:

Y=XY=X

represents perfect prediction.

Therefore:

The closer the points are to the reference line, the better the predictions.


14. Result

Regression analysis was successfully performed using a real-world CSV dataset downloaded from Kaggle. The dataset was imported into R, explored, and preprocessed. A Multiple Linear Regression model was developed to predict house prices based on selected features. The model was evaluated using R-squared, Adjusted R-squared, MAE, and RMSE, and predictions were made for new observations.

Comments

Popular posts from this blog

Statistical Methods Lab ( R Language) PCCBL308 Semester 3 KTU BTech CB and CU 2024 Scheme - Dr Binu V P

Programs in R - using control statements - Assignment 2

Programs to try using Functions in R - Assignment 3