Brief Introduction to Statistical Tests

 

Brief Introduction to Statistical Tests

A statistical test is a method used to make a decision or draw a conclusion about a population based on data collected from a sample.

In many real-world situations, we want to answer questions such as:

  • Is there a significant difference between two groups?
  • Does a new teaching method improve student performance?
  • Are two variables related?
  • Is an observed difference due to chance or a real effect?

Statistical testing helps us answer such questions scientifically.

Basic Idea

A statistical test usually begins with two hypotheses:

  • Null Hypothesis (H0H_0): Assumes that there is no significant effect, difference, or relationship.
  • Alternative Hypothesis (H1H_1): Assumes that there is a significant effect, difference, or relationship.

The test analyzes the sample data and produces a test statistic and usually a p-value.

Decision Rule

  • If the p-value is less than the significance level (commonly α=0.05\alpha = 0.05), we reject the null hypothesis.
  • If the p-value is greater than or equal to 0.05, we fail to reject the null hypothesis.

In simple terms, statistical tests help us determine whether the patterns observed in sample data are likely to represent a real effect or could have occurred simply by random chance.

Statistical tests are broadly classified into parametric tests and non-parametric tests, which students can explore further through experiments in R.


Parametric and Non-Parametric Tests

Statistical tests are commonly classified into parametric and non-parametric tests based on the assumptions they make about the data.

1. Parametric Tests

Parametric tests assume that the data follow a particular probability distribution, usually the Normal distribution. They typically work with population parameters such as the mean and standard deviation.

Common assumptions

Parametric tests often assume:

  • Data are approximately normally distributed.
  • Observations are independent.
  • The data are measured on an interval or ratio scale.
  • Variances are equal between groups in some tests.

Examples

  • Z-test
  • One-sample t-test
  • Independent samples t-test
  • Paired t-test
  • ANOVA
  • Pearson correlation

Example

Suppose we want to compare the average marks of two groups of students. If the data satisfy the required assumptions, an independent samples t-test can be used.


2. Non-Parametric Tests

Non-parametric tests make fewer assumptions about the underlying distribution of the data. They are particularly useful when data are:

  • Not normally distributed
  • Ordinal or ranked
  • Contain significant outliers
  • Based on small samples where normality assumptions are difficult to justify

Many non-parametric tests use ranks rather than the original numerical values.

Examples

  • Mann–Whitney U test
  • Wilcoxon signed-rank test
  • Kruskal–Wallis test
  • Friedman test
  • Spearman rank correlation
  • Chi-square test (for categorical frequency data)

Key Difference

FeatureParametric TestsNon-Parametric Tests
Assumption about distributionUsually assumes a specific distributionMakes fewer distributional assumptions
Common data typeInterval/RatioOrdinal, ranked, or non-normal data
UsesOriginal numerical valuesOften ranks or frequencies
Sensitivity to outliersUsually more sensitiveOften more robust
Examplest-test, ANOVA, Pearson correlationWilcoxon, Mann–Whitney, Kruskal–Wallis, Spearman

Simple Rule for Students

Use a parametric test when its assumptions are reasonably satisfied. Use an appropriate non-parametric alternative when those assumptions are not satisfied or when the data type calls for it.

Comments

Popular posts from this blog

Statistical Methods Lab ( R Language) PCCBL308 Semester 3 KTU BTech CB and CU 2024 Scheme - Dr Binu V P

Programs in R - using control statements - Assignment 2

Programs to try using Functions in R - Assignment 3