Regression Analysis Using a CSV Dataset Downloaded from Kaggle in R
Experiment: Regression Analysis Using a CSV Dataset Downloaded from Kaggle in R
1. Experiment Title
Implementation of Regression Analysis Using a Real-World CSV Dataset Downloaded from Kaggle
2. Aim
To import a CSV dataset downloaded from Kaggle into R, preprocess and explore the dataset, build a regression model, evaluate its performance, and make predictions using the developed model.
3. Objectives
After completing this experiment, students should be able to:
- Download a real-world dataset from Kaggle.
- Import a CSV file into R.
- Explore the structure and contents of the dataset.
- Handle missing values.
- Select suitable independent and dependent variables.
-
Build a regression model using
lm(). - Interpret regression coefficients.
- Calculate predicted values and residuals.
- Evaluate the model using , Adjusted , MAE, and RMSE.
- Visualize actual and predicted values.
- Make predictions for new observations.
4. Theory
4.1 Regression Analysis
Regression analysis is a statistical technique used to study the relationship between a dependent variable and one or more independent variables.
Simple Linear Regression
When one independent variable is used:
where:
- = Dependent variable
- = Independent variable
- = Intercept
- = Regression coefficient
- = Error term
Multiple Linear Regression
When two or more independent variables are used:
The regression model estimates the relationship between the predictor variables and the target variable.
5. Problem Statement
Download a suitable CSV dataset from Kaggle containing numerical variables.
Import the dataset into R and perform regression analysis to predict a numerical target variable.
Students should perform the following:
- Import the CSV file.
- Explore the dataset.
- Check for missing values.
- Select a numerical target variable.
- Select suitable predictor variables.
- Build a regression model.
- Analyze the regression results.
- Calculate predictions and residuals.
- Evaluate the model.
- Visualize the results.
- Predict the value for a new observation.
6. Suggested Dataset
A very suitable dataset for this experiment is a House Price Prediction dataset from Kaggle.
Download : https://www.kaggle.com/datasets/yasserh/housing-prices-dataset/data
The dataset should contain variables similar to:
| Variable | Description |
|---|---|
area | Area of the house |
bedrooms | Number of bedrooms |
bathrooms | Number of bathrooms |
stories | Number of floors |
parking | Number of parking spaces |
price | Price of the house |
Regression Problem
Predict the price of a house based on its area, number of bedrooms, number of bathrooms, and parking facilities.
This makes an excellent real-world Multiple Linear Regression experiment.
7. Software Requirements
- R
- RStudio
- CSV dataset downloaded from Kaggle
8. R Program
The following program assumes that the downloaded file is named:
Housing.csv
and is stored in the working directory.
# ------------------------------------------------------- # Regression Analysis Using a CSV Dataset from Kaggle # ------------------------------------------------------- # ------------------------------------------------------- # Step 1: Import the CSV File # ------------------------------------------------------- data <- read.csv("Housing.csv") # ------------------------------------------------------- # Step 2: Display the First Few Records # ------------------------------------------------------- cat("First Few Records:\n") head(data) # ------------------------------------------------------- # Step 3: Display Dataset Structure # ------------------------------------------------------- cat("\nStructure of Dataset:\n") str(data) # ------------------------------------------------------- # Step 4: Display Dataset Dimensions # ------------------------------------------------------- cat("\nNumber of Rows and Columns:\n") dim(data) # ------------------------------------------------------- # Step 5: Display Column Names # ------------------------------------------------------- cat("\nColumn Names:\n") names(data) # ------------------------------------------------------- # Step 6: Display Summary Statistics # ------------------------------------------------------- cat("\nSummary Statistics:\n") summary(data) # ------------------------------------------------------- # Step 7: Check for Missing Values # ------------------------------------------------------- cat("\nMissing Values in Each Column:\n") colSums(is.na(data)) # ------------------------------------------------------- # Step 8: Remove Missing Values # ------------------------------------------------------- data <- na.omit(data) # ------------------------------------------------------- # Step 9: Select Required Variables # ------------------------------------------------------- house_data <- data[, c( "price", "area", "bedrooms", "bathrooms", "parking" )] # ------------------------------------------------------- # Step 10: Build Multiple Linear Regression Model # ------------------------------------------------------- model <- lm( price ~ area + bedrooms + bathrooms + parking, data = house_data ) # ------------------------------------------------------- # Step 11: Display Model Summary # ------------------------------------------------------- cat("\nRegression Model Summary:\n") summary(model) # ------------------------------------------------------- # Step 12: Display Regression Coefficients # ------------------------------------------------------- cat("\nRegression Coefficients:\n") coef(model) # ------------------------------------------------------- # Step 13: Calculate Predicted Values # ------------------------------------------------------- predicted_price <- predict(model) # ------------------------------------------------------- # Step 14: Calculate Residuals # ------------------------------------------------------- residuals_value <- residuals(model) # ------------------------------------------------------- # Step 15: Add Predictions to Dataset # ------------------------------------------------------- house_data$Predicted_Price <- predicted_price house_data$Residual <- residuals_value cat("\nActual and Predicted Values:\n") head(house_data) # ------------------------------------------------------- # Step 16: Calculate R-Squared Values # ------------------------------------------------------- model_summary <- summary(model) cat( "\nR-Squared =", model_summary$r.squared, "\n" ) cat( "Adjusted R-Squared =", model_summary$adj.r.squared, "\n" ) # ------------------------------------------------------- # Step 17: Calculate MAE # ------------------------------------------------------- mae <- mean( abs( house_data$price - house_data$Predicted_Price ) ) cat("\nMAE =", mae, "\n") # ------------------------------------------------------- # Step 18: Calculate RMSE # ------------------------------------------------------- rmse <- sqrt( mean( ( house_data$price - house_data$Predicted_Price )^2 ) ) cat("RMSE =", rmse, "\n") # ------------------------------------------------------- # Step 19: Plot Actual vs Predicted Prices # ------------------------------------------------------- plot( house_data$price, house_data$Predicted_Price, main = "Actual vs Predicted House Prices", xlab = "Actual Price", ylab = "Predicted Price", pch = 19 ) # Add reference line abline( a = 0, b = 1, col = "blue", lwd = 2 ) # ------------------------------------------------------- # Step 20: Plot Area vs Price # ------------------------------------------------------- plot( house_data$area, house_data$price, main = "House Area vs Price", xlab = "Area", ylab = "Price", pch = 19 ) # ------------------------------------------------------- # Step 21: Predict Price of a New House # ------------------------------------------------------- new_house <- data.frame( area = 5000, bedrooms = 3, bathrooms = 2, parking = 2 ) new_prediction <- predict( model, newdata = new_house ) cat( "\nPredicted Price of New House =", new_prediction, "\n" )Output
First Few Records: Structure of Dataset: 'data.frame': 545 obs. of 13 variables: $ price : int 13300000 12250000 12250000 12215000 11410000 10850000 10150000 10150000 9870000 9800000 ... $ area : int 7420 8960 9960 7500 7420 7500 8580 16200 8100 5750 ... $ bedrooms : int 4 4 3 4 4 3 4 5 4 3 ... $ bathrooms : int 2 4 2 2 1 3 3 3 1 2 ... $ stories : int 3 4 2 2 2 1 4 2 2 4 ... $ mainroad : chr "yes" "yes" "yes" "yes" ... $ guestroom : chr "no" "no" "no" "no" ... $ basement : chr "no" "no" "yes" "yes" ... $ hotwaterheating : chr "no" "no" "no" "no" ... $ airconditioning : chr "yes" "yes" "no" "yes" ... $ parking : int 2 3 2 3 2 2 2 0 2 1 ... $ prefarea : chr "yes" "no" "yes" "yes" ... $ furnishingstatus: chr "furnished" "furnished" "semi-furnished" "furnished" ... Number of Rows and Columns: Column Names: Summary Statistics: Missing Values in Each Column: Regression Model Summary: Regression Coefficients: Actual and Predicted Values: R-Squared = 0.5101313 Adjusted R-Squared = 0.5065026 MAE = 988779.9 RMSE = 1307931 Predicted Price of New House = 6142727 > source("~/test.R") First Few Records: Structure of Dataset: 'data.frame': 545 obs. of 13 variables: $ price : int 13300000 12250000 12250000 12215000 11410000 10850000 10150000 10150000 9870000 9800000 ... $ area : int 7420 8960 9960 7500 7420 7500 8580 16200 8100 5750 ... $ bedrooms : int 4 4 3 4 4 3 4 5 4 3 ... $ bathrooms : int 2 4 2 2 1 3 3 3 1 2 ... $ stories : int 3 4 2 2 2 1 4 2 2 4 ... $ mainroad : chr "yes" "yes" "yes" "yes" ... $ guestroom : chr "no" "no" "no" "no" ... $ basement : chr "no" "no" "yes" "yes" ... $ hotwaterheating : chr "no" "no" "no" "no" ... $ airconditioning : chr "yes" "yes" "no" "yes" ... $ parking : int 2 3 2 3 2 2 2 0 2 1 ... $ prefarea : chr "yes" "no" "yes" "yes" ... $ furnishingstatus: chr "furnished" "furnished" "semi-furnished" "furnished" ... Number of Rows and Columns: Column Names: Summary Statistics: Missing Values in Each Column: Regression Model Summary: Regression Coefficients: Actual and Predicted Values: R-Squared = 0.5101313 Adjusted R-Squared = 0.5065026 MAE = 988779.9 RMSE = 1307931 Predicted Price of New House = 6142727
9. Explanation of the Important Steps
Step 1: Importing the CSV File
data <- read.csv("Housing.csv")
The read.csv() function imports the CSV file into R as a data frame.
Important: The file name must exactly match the downloaded CSV file name.
Step 2: Exploring the Dataset
Students should use:
head(data)
to view the first few records.
str(data)
to understand:
- Variable names
- Data types
- Dataset structure
summary(data)
to obtain descriptive statistics.
Step 3: Checking Missing Values
colSums(is.na(data))
This checks the number of missing values in every column.
For this basic experiment:
data <- na.omit(data)
removes rows containing missing values.
Step 4: Building the Regression Model
model <- lm( price ~ area + bedrooms + bathrooms + parking, data = house_data )
This means:
Predict house price using area, bedrooms, bathrooms, and parking.
The model has the general form:
10. Predicted Values
The predicted house prices are obtained using:
predicted_price <- predict(model)
Each house will have:
- Actual price
- Predicted price
11. Residuals
The residual is:
In R:
residuals(model)
A good regression model should generally have residuals randomly distributed around zero.
12. Model Evaluation Measures
R-Squared
represents the proportion of variation in the target variable explained by the regression model.
For example:
means approximately:
75% of the variation in house prices is explained by the selected variables.
Adjusted R-Squared
Adjusted is particularly useful in Multiple Regression because it considers the number of predictors.
MAE — Mean Absolute Error
It measures the average absolute prediction error.
Lower MAE indicates better predictions.
RMSE — Root Mean Square Error
RMSE gives greater importance to large errors.
Lower RMSE indicates better predictions.
13. Interpretation of the Actual vs Predicted Plot
The graph contains:
- X-axis → Actual house prices
- Y-axis → Predicted house prices
The line:
represents perfect prediction.
Therefore:
The closer the points are to the reference line, the better the predictions.
14. Result
Regression analysis was successfully performed using a real-world CSV dataset downloaded from Kaggle. The dataset was imported into R, explored, and preprocessed. A Multiple Linear Regression model was developed to predict house prices based on selected features. The model was evaluated using R-squared, Adjusted R-squared, MAE, and RMSE, and predictions were made for new observations.
Comments
Post a Comment