Covariance Analysis Using R
Experiment
Covariance Analysis Using R
1. Aim
To study the concept of covariance and implement the calculation of covariance between two numerical variables using R.
2. Objectives
After completing this experiment, students should be able to:
- Understand the concept of covariance.
- Distinguish between positive, negative, and near-zero covariance.
- Calculate covariance manually using the mathematical formula.
- Implement covariance calculation using R.
-
Compare the manually calculated result with R's built-in
cov()function. - Interpret the covariance between two variables.
3. Problem Statement
The following data represents the number of hours studied and the corresponding marks obtained by 8 students.
| Student | Hours Studied (X) | Marks (Y) |
|---|---|---|
| 1 | 2 | 45 |
| 2 | 3 | 50 |
| 3 | 4 | 55 |
| 4 | 5 | 60 |
| 5 | 6 | 65 |
| 6 | 7 | 70 |
| 7 | 8 | 75 |
| 8 | 9 | 80 |
Determine the covariance between Hours Studied and Marks and interpret the result.
4. Theory
What is covariance?
Covariance is a statistical measure that describes the direction of the relationship between two variables.
Suppose we have two variables:
- = Hours Studied
- = Marks
Covariance tells us how the two variables vary together.
Interpretation
| Covariance | Interpretation |
|---|---|
| Positive | Both variables tend to increase together |
| Negative | One variable tends to increase when the other decreases |
| Near zero | No clear linear co-movement |
For example, if students who study more hours generally obtain higher marks, we expect the covariance between study hours and marks to be positive.
5. Mathematical Formula
For a sample of observations, covariance is calculated as:
where:
- = ith value of X
- = ith value of Y
- = mean of X
- = mean of Y
- = number of observations
The denominator is because we are calculating sample covariance.
6. Algorithm
Algorithm: Calculate Covariance
Step 1: Start.
Step 2: Read the two vectors and .
Step 3: Find the number of observations .
Step 4: Calculate the mean of .
Step 5: Calculate the mean of .
Step 6: For each observation, calculate:
Step 7: Add all the products.
Step 8: Divide the sum by .
Step 9: Display the covariance.
Step 10: Interpret whether the covariance is positive, negative, or approximately zero.
Step 11: Stop.
7. R Program – Creating the Data
# Hours studied hours <- c(2, 3, 4, 5, 6, 7, 8, 9) # Marks obtained marks <- c(45, 50, 55, 60, 65, 70, 75, 80) # Create data frame students <- data.frame( Hours = hours, Marks = marks ) print(students)
Output
Hours Marks 1 2 45 2 3 50 3 4 55 4 5 60 5 6 65 6 7 70 7 8 75 8 9 80
8. Covariance Using R Built-in Function
R provides the cov() function to calculate covariance.
covariance <- cov(hours, marks) cat("Covariance =", covariance, "\n")
For this dataset, the output is:
Covariance = 15
Therefore,
Covariance = 15
Since the covariance is positive, hours studied and marks tend to increase together.
9. Manual Implementation of Covariance
For the lab, it is useful to make students implement the formula themselves instead of directly using cov().
hours <- c(2, 3, 4, 5, 6, 7, 8, 9) marks <- c(45, 50, 55, 60, 65, 70, 75, 80) n <- length(hours) mean_x <- sum(hours) / n mean_y <- sum(marks) / n sum_product <- 0 for(i in 1:n) { sum_product <- sum_product + (hours[i] - mean_x) * (marks[i] - mean_y) } covariance <- sum_product / (n - 1) cat("Mean of Hours =", mean_x, "\n") cat("Mean of Marks =", mean_y, "\n") cat("Covariance =", covariance, "\n")
Output
Mean of Hours = 5.5 Mean of Marks = 62.5 Covariance = 15
10. Covariance Using a Data Frame
Once students understand vectors, covariance can also be calculated directly from data-frame columns.
students <- data.frame( Hours = c(2,3,4,5,6,7,8,9), Marks = c(45,50,55,60,65,70,75,80) ) covariance <- cov(students$Hours, students$Marks) cat("Covariance =", covariance)
Output:
Covariance = 15
11. Verify the Result
Students should compare their own implementation with the built-in function.
manual_covariance <- sum_product / (n - 1) r_covariance <- cov(hours, marks) cat("Manual Covariance =", manual_covariance, "\n") cat("R Covariance =", r_covariance, "\n") if(manual_covariance == r_covariance) { cat("Both results are equal.\n") } else { cat("Results are different.\n") }
Output
Manual Covariance = 15 R Covariance = 15 Both results are equal.
12. Interpretation of Covariance
For the given data:
Covariance = 15
The covariance is positive.
Therefore:
Hours studied and marks obtained tend to increase together. Students who study more hours tend to obtain higher marks in this dataset.
However, covariance mainly indicates the direction of the relationship. Its numerical magnitude depends on the units of the variables.
For example, changing marks from marks to percentage or hours to minutes would change the numerical value of covariance.
13. Additional Experiment – Understanding Different Types of Covariance
To help students understand the concept better, you can give three small datasets.
Dataset 1 – Positive Covariance
x <- c(1, 2, 3, 4, 5) y <- c(10, 20, 30, 40, 50) cov(x, y)
Expected result:
25
Interpretation: As X increases, Y also increases.
Dataset 2 – Negative Covariance
x <- c(1, 2, 3, 4, 5) y <- c(50, 40, 30, 20, 10) cov(x, y)
Expected result:
-25
Interpretation: As X increases, Y decreases.
Dataset 3 – Zero Covariance
x <- c(1, 2, 3, 4, 5) y <- c(10, 20, 10, 20, 10) cov(x, y)
Here the covariance is close to zero.
Interpretation: There is no clear linear co-movement between X and Y.
14. Optional Visualization
Although covariance itself is a numerical measure, a scatter plot can help students understand its meaning visually.
plot(hours, marks, main = "Hours Studied vs Marks", xlab = "Hours Studied", ylab = "Marks", pch = 19)
Students should observe that the points generally move upward from left to right, which is consistent with the positive covariance.
15. Important Observation
Students should understand that:
Covariance indicates the direction of the relationship, but its magnitude is not easily comparable across datasets because it depends on the units of measurement.
For measuring both direction and standardized strength of linear relationship, the correlation coefficient is more appropriate.
20. Result
The concept of covariance was studied and the covariance between two variables was calculated using R. The covariance was also implemented manually using an algorithm and verified using R's built-in
cov()function. For the given student dataset, the covariance between hours studied and marks obtained was 15, indicating a positive relationship between the two variables.
Comments
Post a Comment