π Doing statistics involves three stages:
Data require six contextual questions: who the cases are, what variables were measured, why the data were collected, how they were collected, when they were collected, and where they were collected.
π Descriptive statistics describe sample data quantitatively, visually, and verbally, whereas inferential statistics use sample data to draw conclusions and make reliable forecasts about a population.
Think β Show β Tell
β Must-know
π Quantitative data are measured by numbers and usually have units, whereas qualitative data identify a group or category.
π A numerical variable is quantitative only when calculations with its values make sense, so student ID numbers and classroom numbers are numerical but categorical.
Further detail
Common sampling techniques include:
The four stated sources of data are:
Quantitative measures amount; qualitative identifies category
β Must-know
π Bar charts and pie charts are adapted for qualitative variables, whereas quantitative data are better represented with summaries or displays such as stem-and-leaf plots and histograms.
Further detail
A frequency table lists each category and its number of observations, and it may also show percentages.
A Pareto diagram is a bar graph whose categories are arranged by decreasing height from left to right.
For the European teenage-smoking data, the bin frequencies for [10β20), [20β30), [30β40), and [40β50) were 6, 15, 12, and 3, respectively. β European School Survey Project on Alcohol and Other Drugs, 2007
Bars have gaps; histograms do not
β Must-know
π Formula β The range equals the largest observation minus the smallest observation: .
π When a distribution is not unimodal or symmetric or contains outliers, gaps, or clusters, the median and IQR are generally preferred to the mean and standard deviation.
π For a bell-shaped, unimodal, symmetric distribution, the empirical rule states that approximately 68%, 95%, and 99.7% of observations lie within one, two, and three standard deviations of the mean, respectively.
π Formula β A sample z-score measures the number of sample standard deviations a value lies from the sample mean: .
Further detail
π Formula β The coefficient of variation compares standard deviation with the mean: .
Mean and SD are sensitive; median and IQR are resistant
π Formula β Under independence, the expected count in row i and column j is , where Ri is the row total, Cj is the column total, and n is the sample size.
π Formula β The chi-squared statistic compares observed and expected counts: .
Condition on a variable β compare distributions β assess independence
β Must-know
π Correlation does not prove causation because an association may result from coincidence, causality, or a common underlying cause called a lurking variable.
π Formula β A linear regression model has the form , where b0 is the y-intercept and b1 is the slope.
π Formula β A residual is the observed response minus the predicted response: .
π Interpolation predicts within the domain used to construct a model, whereas extrapolation predicts outside that domain and requires caution.
Further detail
π The least-squares regression line is the unique line that minimizes the sum of squared residuals.
π A useful residual plot should show roughly equal scatter across the range, no bends, and few or no outliers.
Scatterplot β correlation β regression β residuals
β Must-know
π Formula β The coefficient of determination is .
For the manatee-deaths and powerboat-registrations example, , so the model accounts for 89.5% of the variation in manatee deaths and leaves 10.5% unexplained.
A useful residual plot should have:
Further detail
Better fit β smaller residual variation β larger rΒ²
π Formula β For an event A, the empirical probability is .
π Formula β For any event A, the complement rule is .
π Formula β For disjoint events A and B, the addition rule is .
π Formula β For any two events A and B, the general addition rule is .
π Formula β The conditional probability of B given A is , provided that .
π Formula β For any two events A and B, the general multiplication rule is .
π Events A and B are independent when , equivalently when .
Disjoint events cannot occur together; independent events do not influence one another
π A discrete random variable has a countable number of distinct outcomes, whereas a continuous random variable can take any numerical value within a range.
π A probability model must assign nonnegative probabilities to all values and satisfy and .
π Formula β The expected value of a discrete random variable is .
π Formula β The variance and standard deviation of a random variable are and .
Model β expected value β variance β standard deviation
β Must-know
π Formula β For a binomial random variable with n trials and success probability p, the expected value is and the standard deviation is , where .
Further detail
π Formula β The number of combinations of r successes among n trials is .
Binomial counts successes in fixed trials; Poisson counts rare events in an interval
π The 68-95-99.7 rule states that approximately 68%, 95%, and 99.7% of Normal-model values lie within one, two, and three standard deviations of the mean, respectively.
π Formula β A value is standardized with the z-score formula , converting it to the standard Normal model N(0,1).
π The Normal approximation to a binomial model is appropriate when the expected numbers of successes and failures satisfy and .
A bell-shaped curve with probability represented by shaded area
β Must-know
π A population parameter describes an entire population, whereas a sample statistic describes a sample and is calculated from its observations.
π The Central Limit Theorem states that the sampling distribution of a random sample mean becomes approximately Normal as sample size increases, usually with nβ₯30 being sufficient depending on the original distribution.
Further detail
Larger samples β smaller standard error β more nearly normal sample means
β Must-know
π Formula β For a quantitative variable with population mean ΞΌ and standard deviation Ο, the sampling distribution of the mean has mean and standard deviation , called the standard error.
π If the population distribution is normal, the sampling distribution of the mean is normal; even when the population is not normal, it tends toward normality as n increases.
π The Central Limit Theorem states that the mean of a random sample of size n can be approximated by a normal model, with the approximation improving as n increases and usually being adequate when nβ₯30.
Further detail
Larger samples β smaller standard error β more precise sample means
β Must-know
π Formula β For a qualitative variable with population proportion p, the sampling distribution of the proportion has mean and standard deviation , called the standard error.
π The sampling distribution of a proportion is approximately normal when and .
Further detail
Means describe quantitative variables, whereas proportions describe qualitative properties
π Formula β The general confidence-interval formula is .
π A 95% confidence level means that 95% of confidence intervals constructed from all possible random samples by the same method would contain the true population parameter.
Point estimates give one value, whereas confidence intervals give a range with uncertainty
β Must-know
π Formula β When Ο is known, a confidence interval for a population mean is .
π A z-interval for a population mean requires a random sample, known population standard deviation Ο, and a large sample size such as nβ₯30 to support an approximately normal sampling distribution.
π Formula β When Ο is unknown, a confidence interval for a population mean is , with degrees of freedom .
Further detail
π The Student t distribution is symmetric and bell-shaped like the normal distribution but has heavier tails, and it approaches the normal distribution as sample size increases.
Known Ο uses z, whereas unknown Ο uses the Student t distribution
π Formula β A confidence interval for a population proportion is .
π A confidence interval for a population proportion requires a random or representative sample and sufficiently large counts satisfying and .
π Increasing the confidence level requires a larger interval, which increases certainty but decreases precision.
Higher confidence β larger margin of error β lower precision
π Rejecting Hβ provides evidence supporting Hβ, whereas failing to reject Hβ means only that there is insufficient evidence to support Hβ.
A Type I error occurs when a true null hypothesis is rejected, and its probability is the significance level Ξ±; a Type II error occurs when a false null hypothesis is not rejected, and its probability is Ξ².
π A hypothesis test proceeds by: stating Hβ and Hβ, checking conditions, calculating a test statistic, finding critical values or a p-value, concluding by rejecting or failing to reject Hβ in context
π The p-value is the conditional probability of obtaining a test statistic at least as extreme as the observed one, assuming Hβ is true; reject Hβ when p-value<Ξ± and fail to reject Hβ when p-value>Ξ±.
π Valid large-sample inference for the difference between two population means requires two independent random samples and sample sizes of at least 30 in both populations.
π Formula β The large-sample confidence interval for the difference between two population means is .
π Formula β The large-sample test statistic for two population means is approximately .
π Valid small-sample inference for the difference between two population means requires independent random samples, approximately normal populations, and equal population variances.
π Formula β The pooled variance for two independent small samples is .
π Formula β The small-sample confidence interval for two population means is with degrees of freedom.
π Valid large-sample inference for the difference between two population proportions requires independent random samples and, for each sample, at least 15 expected successes and 15 expected failures based on the sample proportion.
π Formula β The large-sample confidence interval for two population proportions is .
π Formula β For testing equality of two population proportions, the pooled proportion is and the test statistic is , where .
Hypotheses β conditions β statistic β critical value or p-value β conclusion
| Measure | What it describes | Sensitivity to outliers |
|---|---|---|
| Mean | Balance point of the data | Sensitive |
| Median | Middle of ordered data | Resistant |
| Range | Maximum minus minimum | Sensitive |
| Standard deviation | Spread around the mean | Sensitive |
| Interquartile range | Spread of the middle 50% | Resistant |
| Feature | Discrete | Continuous |
|---|---|---|
| Possible values | Countable distinct outcomes | Any numerical value in an interval |
| Probability focus | Probability of individual values | Probability over intervals |
| Typical examples | Number of successes or errors | Height, distance, or time |
Test your knowledge on Statistics for Management with 64 multiple-choice questions with detailed corrections.
1. Which statement best distinguishes statistics as a discipline from statistics as numerical results?
2. Why does an isolated value provide limited statistical meaning?
Memorize the key concepts of Statistics for Management with 88 interactive flashcards.
What is statistics as a way of reasoning and tools?
Statistics is a way of reasoning and a collection of tools to analyze data.
What are statistics in the plural?
Statistics in plural are results of calculations made with data.
What are data?
Data are collections of numbers, characters, images, or items providing information.
Import your course and AI generates sheets, quizzes and flashcards in 30 seconds.
Sheet generator