Quiz: Statistics Foundations and Probability — 70 questions

Detailed questions and answers

1. What does statistical inference primarily do when analyzing sample data?

It records every unit belonging to a target population
It proves mathematical laws about probability models
It draws conclusions about a population from a sample
It organizes observations into tables and graphs

It draws conclusions about a population from a sample

Explanation

Inferential statistics uses sample information to draw conclusions about the broader population. Summarizing observed data in tables or graphs is descriptive statistics rather than statistical inference.

2. Which activity best illustrates applied statistics?

Establishing general laws without analyzing practical observations
Proving a new theorem about the behavior of an estimator
Deriving a mathematical formula for a probability distribution
Using statistical methods to solve a real-world health problem

Using statistical methods to solve a real-world health problem

Explanation

Applied statistics uses established methods and formulas to address practical problems. Developing or proving statistical methods belongs to theoretical statistics.

3. In a study of customer satisfaction among all subscribers of a service, what is the population?

All subscribers whose satisfaction is being studied
The subscribers selected to complete the questionnaire
The average satisfaction score calculated from the responses
The satisfaction scores recorded from respondents

All subscribers whose satisfaction is being studied

Explanation

The population consists of all observation units whose characteristics are under study, here all subscribers of the service. The selected respondents form a sample, not the population.

4. What distinguishes a census from a sample survey?

A sample survey gives every element a known selection probability
A census collects information from every population element
A census collects information from a selected part of the population
A sample survey excludes all units that lack measured characteristics

A census collects information from every population element

Explanation

A census surveys every element in the population. A sample survey gathers information from only part of that population.

5. Which example represents a variable rather than a constant?

The fixed conversion rate used in a calculation
The unchanged identification code assigned to one record
The number of children reported by different households
The predetermined number of questions on a questionnaire

The number of children reported by different households

Explanation

A variable can take different values across observation units, such as household child counts. A constant has a fixed value within the relevant study or calculation.

6. Which characteristic is quantitative rather than qualitative?

The preferred method of communication reported by each employee
The measured income of each employee
The type of contract held by each employee
The department assigned to each employee

The measured income of each employee

Explanation

Income is expressed numerically and is therefore quantitative. Departments, contract types, and communication preferences are categorical qualitative variables.

7. Which variable is discrete?

The exact temperature recorded in a storage room
The number of defects found in a manufactured item
The duration required to complete a task
The amount of fuel consumed during a journey

The number of defects found in a manufactured item

Explanation

The number of defects is a count with isolated, countable values, making it discrete. Temperature, fuel amount, and duration can vary continuously within intervals.

8. A category has frequency f=18f=18 in a dataset with total frequency f=120\sum f=120. What is its relative frequency?

0.180.18
6.676.67
0.150.15
102102

$$0.15$$

Explanation

Relative frequency is calculated as fr=ff=18120=0.15f_r=\frac{f}{\sum f}=\frac{18}{120}=0.15. The value 0.180.18 confuses the category frequency with its proportion of the total.

9. Which notation represents the mean of a population?

xˉ\bar{x}
μ\mu
ss
Q1Q_1

$$\mu$$

Explanation

The population arithmetic mean is denoted by μ\mu. The symbol xˉ\bar{x} denotes a sample mean, while Q1Q_1 and ss refer to a quartile and sample standard deviation.

10. What is the median of the ordered values 3, 5, 8, 12, 20, and 25?

10
12
12.5
8

10

Explanation

With six observations, the median is the mean of the two middle values: 8+122=10\frac{8+12}{2}=10. Choosing 8 or 12 uses one middle observation instead of averaging both.

11. A data series contains the values 2, 4, 4, 5, 7, and 7. Which statement correctly describes its mode?

It is bimodal, with modes 4 and 7
It has one mode, equal to 7
It has no mode because two values are tied
It has one mode, equal to 4

It is bimodal, with modes 4 and 7

Explanation

Both 4 and 7 occur twice, which is the highest frequency, so the series has two modes. A tie at the greatest frequency produces multiple modes rather than no mode.

12. Two datasets have the same arithmetic mean, but Dataset A has a smaller standard deviation than Dataset B. What does this imply?

Dataset B must have a higher arithmetic mean
Dataset B contains fewer observations than Dataset A
Dataset A is more tightly grouped around its mean
Dataset A has a larger difference between its quartiles

Dataset A is more tightly grouped around its mean

Explanation

A smaller standard deviation indicates that values are less dispersed around the arithmetic mean. It does not determine the number of observations, quartile difference, or mean when the means are stated to be equal.

13. What does the sample space of a probability experiment contain?

The outcomes selected for a particular event
The trials that produce identical results
All possible outcomes of the experiment
The probabilities assigned to each event

All possible outcomes of the experiment

Explanation

The sample space is the set containing every possible outcome of an experiment. An event is a selected subset of those outcomes, not the entire sample space.

14. A fair die is rolled once. What is the probability of obtaining an even number?

23\frac{2}{3}
12\frac{1}{2}
13\frac{1}{3}
16\frac{1}{6}

$$\frac{1}{2}$$

Explanation

There are three favorable outcomes, 2, 4, and 6, among six equally likely outcomes, so the probability is 36=12\frac{3}{6}=\frac{1}{2}. The calculation uses classical probability because the die outcomes are equally likely.

15. In 200 trials, an event occurs 46 times. What relative-frequency estimate should be used for its probability?

1200=0.005\frac{1}{200}=0.005
461540.30\frac{46}{154}\approx0.30
46200=0.23\frac{46}{200}=0.23
200464.35\frac{200}{46}\approx4.35

$$\frac{46}{200}=0.23$$

Explanation

Relative-frequency probability is estimated by dividing the number of realizations by the number of trials, giving P(A)=fn=46200=0.23P(A)=\frac{f}{n}=\frac{46}{200}=0.23. The other ratios do not represent occurrences divided by total trials.

16. Which situation satisfies the requirements of a binomial experiment?

Ten independent shots, each resulting in a hit or miss with a constant hit probability
Ten measurements whose values vary continuously across a range
A sequence of draws where the chance of success changes after every draw
One trial with four possible outcomes having unequal probabilities

Ten independent shots, each resulting in a hit or miss with a constant hit probability

Explanation

A binomial experiment has a fixed number of identical independent trials, two outcomes per trial, and constant outcome probabilities. Continuous measurements, changing probabilities, and more than two outcomes violate these conditions.

17. What determines the value of a random variable?

The average of all measured characteristics
The outcome of a random experiment
The fixed setting chosen before observation
The label assigned to an observation unit

The outcome of a random experiment

Explanation

A random variable receives its value from the outcome of a random experiment. An ordinary variable instead describes a measured characteristic of observation units.

18. Which variable is discrete rather than continuous?

The exact mass of a randomly selected package
The number of customers entering a store
The time required to complete a task
The temperature recorded at noon

The number of customers entering a store

Explanation

The number of customers is a countable quantity, so it is a discrete random variable. Mass, temperature, and completion time can take values across intervals and are continuous variables.

19. Which condition must a discrete probability distribution satisfy?

The probabilities sum to 0 when the variable has many values
Each probability is positive, and the largest value sums to 1
Each probability is between 0 and 1, and all probabilities sum to 1
The probabilities may exceed 1 if their average equals 1

Each probability is between 0 and 1, and all probabilities sum to 1

Explanation

A valid discrete distribution requires 0P(x)10\leq P(x)\leq1 for every value and P(x)=1\sum P(x)=1. Probabilities cannot exceed 1, and their total must represent the complete probability mass.

20. For a discrete random variable with values 1 and 3 having probabilities 0.25 and 0.75, what is its expected value?

3.03.0
1.751.75
2.52.5
2.02.0

$$2.5$$

Explanation

The expected value is the probability-weighted sum, E(X)=1(0.25)+3(0.75)=2.5E(X)=1(0.25)+3(0.75)=2.5. It is not found by taking an unweighted average of the possible values.

21. What percentage of a normal distribution lies within two standard deviations of its mean under the empirical rule?

50.00%
95.44%
99.74%
68.26%

95.44%

Explanation

The empirical rule places 95.44% of the total area within two standard deviations of the mean. The 68.26% value applies to one standard deviation, while 99.74% applies to three.

22. Which description identifies the standardized normal distribution?

A normal distribution with mean 0 and standard deviation 1
A normal distribution whose mean and spread vary by observation
A normal distribution containing only positive standardized values
A normal distribution with mean 1 and standard deviation 0

A normal distribution with mean 0 and standard deviation 1

Explanation

The standardized normal distribution has mean 0 and standard deviation 1, and its random variable is denoted by Z. A general normal distribution may have different mean and standard deviation values.

23. If F(a)F(a) is the area to the left of aa under the standardized normal curve, how is P(Za)P(Z\geq a) calculated?

F(a)F(a)
F(a)F(a)F(-a)-F(a)
1F(a)1-F(a)
F(a)1F(a)-1

$$1-F(a)$$

Explanation

The area at or above aa is the complement of the area below aa, so P(Za)=1F(a)P(Z\geq a)=1-F(a). The expression F(a)F(a) represents the area to the left, not the right.

24. A normally distributed measurement has mean 80, standard deviation 5, and observed value 90. What is its standardized value?

z=0.5z=0.5
z=85z=85
z=2z=2
z=18z=18

$$z=2$$

Explanation

Standardization uses z=xμσz=\frac{x-\mu}{\sigma}, so z=90805=2z=\frac{90-80}{5}=2. The value is two standard deviations above the mean, not the raw difference or the ratio of the observation to the mean.

25. What does a sampling distribution describe?

The individual measurements observed throughout the entire population
The systematic errors introduced during data collection and recording
The possible values of a sample statistic and their probabilities across samples
The parameter values that would result from measuring every population member

The possible values of a sample statistic and their probabilities across samples

Explanation

A sampling distribution gives the possible values of a statistic over repeated samples and the probability of each value. A population distribution instead describes individual values in the population, not a statistic computed from samples.

26. When a random sample is selected without nonsampling error, what explains the difference between a sample statistic and its corresponding population parameter?

A calculation error made when summarizing the sample data
A population change occurring after the sample was collected
Random error caused by selecting one sample rather than another
Nonsampling error caused by inaccurate measurement procedures

Random error caused by selecting one sample rather than another

Explanation

Random sampling can produce different statistics from different samples, creating random error relative to the population parameter. Nonsampling error arises from problems such as data collection or data entry rather than sample selection.

27. A population has standard deviation σ\sigma, and a random sample of size n=100n=100 is at most 5% of the population. What is the standard error of the sample mean?

σ100\frac{\sigma}{100}
σ10\frac{\sigma}{10}
10σ10\sigma
σ99\frac{\sigma}{\sqrt{99}}

$$\frac{\sigma}{10}$$

Explanation

The standard error is σXˉ=σn\sigma_{\bar X}=\frac{\sigma}{\sqrt n}, so with n=100n=100 it becomes σ10\frac{\sigma}{10}. The finite population correction is not needed because the sample is no more than 5% of the population.

28. Why can the sampling distribution of sample means be treated as approximately normal when the sample size is sufficiently large, even if the population is not normal?

A large sample removes all random error from the sample mean
The sample mean has the same distribution as every individual population value
The central limit theorem produces approximate normality for large samples
The population distribution becomes normal after enough observations are collected

The central limit theorem produces approximate normality for large samples

Explanation

The central limit theorem states that sample means are approximately normally distributed for large samples, generally taken as n30n\geq30, regardless of the population distribution. A large sample does not make the individual population values normal or eliminate random error.

29. What is the purpose of statistical estimation?

To describe the distribution of individual observations without estimating a parameter
To use sample information to assign a numerical value or values to a population parameter
To measure every member of a population and avoid drawing conclusions from a sample
To identify data-entry problems before calculating any sample statistic

To use sample information to assign a numerical value or values to a population parameter

Explanation

Statistical estimation uses information from a sample to infer a numerical value or range for a population parameter. Measuring every member is a census, which differs from estimation based on sample information.

30. Which sequence correctly describes the estimation procedure?

Select a sample, collect information, calculate a statistic, and assign it to the parameter
Select a parameter, calculate a statistic, collect information, and then choose a sample
Collect population data, assign a parameter, select a sample, and calculate a statistic
Calculate a parameter, select a sample, collect information, and verify the population census

Select a sample, collect information, calculate a statistic, and assign it to the parameter

Explanation

Estimation proceeds from sample selection to information collection, calculation of a sample statistic, and assignment of that statistic to the corresponding population parameter. The other sequences place the sample, data, or parameter steps in an incorrect order.

31. Which statement correctly distinguishes a point estimate from an interval estimate?

A point estimate measures a census, whereas an interval estimate summarizes a sampling distribution
A point estimate is one sample-based value, whereas an interval estimate is a range for the parameter
A point estimate describes sampling error, whereas an interval estimate describes nonsampling error
A point estimate is a range of sample values, whereas an interval estimate is one population value

A point estimate is one sample-based value, whereas an interval estimate is a range for the parameter

Explanation

A point estimate provides one value calculated from sample information, while an interval estimate provides a range believed to contain the population parameter. The other choices confuse estimation forms with data-collection methods or error types.

32. A sample mean is 7272 and the margin of error is 44. What confidence interval is formed using the general confidence-interval structure?

68 to 7668\text{ to }76
64 to 8064\text{ to }80
68 to 7268\text{ to }72
72 to 7672\text{ to }76

$$68\text{ to }76$$

Explanation

A confidence interval is calculated as the point estimate plus or minus the margin of error, so 72±472\pm4 gives 68 to 7668\text{ to }76. Adding the margin only on one side or doubling it does not follow the interval formula.

33. What role does the null hypothesis play in hypothesis testing?

It represents the conclusion that must be reported whenever a sample statistic changes
It is treated as true initially and rejected only when sufficient evidence supports rejection
It is accepted as false initially and retained when the alternative lacks evidence
It describes the alternative claim that becomes the testing standard before data are collected

It is treated as true initially and rejected only when sufficient evidence supports rejection

Explanation

The null hypothesis is the population-parameter statement treated as true until the evidence is strong enough to reject it. The alternative hypothesis is the competing statement supported when the null hypothesis is rejected.

34. A statistical test rejects a null hypothesis that is actually true. What type of error has occurred?

A Type II error with probability β\beta
A measurement error with probability 1β1-\beta
A sampling error with probability 1α1-\alpha
A Type I error with probability α\alpha

A Type I error with probability $$\alpha$$

Explanation

A Type I error occurs when a true null hypothesis is incorrectly rejected, and its probability is α\alpha. A Type II error instead occurs when a false null hypothesis is not rejected.

35. A false null hypothesis is not rejected in a test. Which statement describes the result?

A two-tailed error occurred, with probability α+β\alpha+\beta, and power is αβ\alpha\beta
A sampling error occurred, with probability 1β1-\beta, and power is β\beta
A Type II error occurred, with probability β\beta, and test power is 1β1-\beta
A Type I error occurred, with probability α\alpha, and test power is 1α1-\alpha

A Type II error occurred, with probability $$\beta$$, and test power is $$1-\beta$$

Explanation

Failing to reject a false null hypothesis is a Type II error, whose probability is β\beta; the test power is 1β1-\beta. A Type I error involves rejecting a true null hypothesis instead.

36. A researcher wants to test whether a population parameter differs from a hypothesized value in either direction. Where should the rejection regions be placed?

At both ends of the sampling distribution in a two-tailed test
Around the center of the sampling distribution in a two-tailed test
At the left end of the sampling distribution in a left-tailed test
At the right end of the sampling distribution in a right-tailed test

At both ends of the sampling distribution in a two-tailed test

Explanation

A two-tailed test examines departures in both directions, so its rejection regions lie at both ends of the distribution. A left-tailed or right-tailed test places a single rejection region at the corresponding end.

37. Which sequence correctly describes the critical-value approach to hypothesis testing?

Formulate hypotheses, choose a distribution, define rejection regions, calculate the statistic, and decide
Choose a distribution, formulate hypotheses, calculate the statistic, define rejection regions, and decide
Calculate the statistic, formulate hypotheses, choose a distribution, define rejection regions, and decide
Define rejection regions, calculate the statistic, formulate hypotheses, choose a distribution, and decide

Formulate hypotheses, choose a distribution, define rejection regions, calculate the statistic, and decide

Explanation

The critical-value procedure begins with the hypotheses and sampling distribution, then establishes rejection and nonrejection regions before calculating the statistic and deciding. The realized statistic does not define the rejection region; the critical value does.

38. A population has known standard deviation σ=12\sigma=12, and a sample of n=36n=36 has mean xˉ=54\bar{x}=54. To test H0:μ=50H_0:\mu=50, which test statistic should be used?

z=1.50z=1.50
z=6.00z=6.00
z=0.33z=0.33
z=2.00z=2.00

$$z=2.00$$

Explanation

The standard error is σXˉ=12/36=2\sigma_{\bar X}=12/\sqrt{36}=2, so the statistic is z=(5450)/2=2.00z=(54-50)/2=2.00. The standard error measures sampling variability, while the numerator compares the sample mean with the hypothesized mean.

39. What does a p-value represent in a hypothesis test?

The probability of obtaining a result inside the nonrejection region under H1H_1
The probability that the sample statistic equals the hypothesized population parameter
The probability that H0H_0 is true after observing the sample statistic
The probability, assuming H0H_0 is true, of obtaining a result at least as extreme in the direction of H1H_1

The probability, assuming $$H_0$$ is true, of obtaining a result at least as extreme in the direction of $$H_1$$

Explanation

A p-value measures how likely a result at least as extreme as the observed one would be if the null hypothesis were true, using the direction specified by the alternative. It is not the posterior probability that the null hypothesis itself is true.

40. If a hypothesis test produces a p-value of 0.032 and uses α=0.05\alpha=0.05, what decision should be made?

Reject H0H_0 because the p-value is greater than or equal to α\alpha
Reject H0H_0 because the p-value is less than α\alpha
Do not reject H0H_0 because the p-value is greater than or equal to α\alpha
Do not reject H0H_0 because the p-value is less than α\alpha

Reject $$H_0$$ because the p-value is less than $$\alpha$$

Explanation

The p-value approach rejects the null hypothesis when the p-value is less than the significance level. Here, 0.032 is below 0.05, so the result falls in the rejection decision.

41. Which four stages form the p-value testing procedure?

Formulate hypotheses, calculate the statistic, choose a sample, and construct a frequency table
Formulate hypotheses, choose a distribution, calculate the p-value, and make the decision
Choose a distribution, calculate the p-value, formulate hypotheses, and estimate the parameter
Calculate the p-value, formulate hypotheses, define confidence limits, and make the decision

Formulate hypotheses, choose a distribution, calculate the p-value, and make the decision

Explanation

The p-value procedure consists of stating the null and alternative hypotheses, selecting the appropriate distribution, finding the p-value, and making the decision. A confidence interval or frequency table is not one of its four required stages.

42. A sample of size n=18n=18 comes from a normal population, and the population standard deviation is known. Which distribution is appropriate for testing the population mean?

The normal distribution, because the population is normal and σ\sigma is known
The normal distribution, because sample size has no role when testing a mean
The chi-square distribution, because the sample concerns a population mean
The t distribution, because every sample smaller than 30 requires t

The normal distribution, because the population is normal and $$\sigma$$ is known

Explanation

For a small sample from a normal population, the normal distribution is appropriate when the population standard deviation is known. The t distribution is associated with an unknown population standard deviation, not merely with a small sample.

43. A sample mean is used to estimate a population mean with known σ\sigma. Which expression gives the normal-theory confidence interval?

xˉ±zσn\bar{x}\pm z\frac{\sigma}{\sqrt n}
xˉ±zσ2n\bar{x}\pm z\frac{\sigma^2}{n}
xˉ±tsn\bar{x}\pm t\frac{s}{\sqrt n}
xˉ±σzn\bar{x}\pm\frac{\sigma}{z\sqrt n}

$$\bar{x}\pm z\frac{\sigma}{\sqrt n}$$

Explanation

When the population standard deviation is known, the interval is xˉ±zσxˉ\bar{x}\pm z\sigma_{\bar{x}} with σxˉ=σ/n\sigma_{\bar{x}}=\sigma/\sqrt n. The t-based expression uses an estimated standard deviation and therefore represents a different setting.

44. A population standard deviation is unknown, the population is normal, and the sample size is 12. Which distribution should be used for inference about the population mean?

The t distribution, because known standard deviations require t for small samples
The normal distribution, because any sample from a normal population uses z
The chi-square distribution, because the sample size is below 30
The t distribution, because σ\sigma is unknown and the small-sample population is normal

The t distribution, because $$\sigma$$ is unknown and the small-sample population is normal

Explanation

The t distribution is used for a population mean when the population standard deviation is unknown, and normality is required for a small sample. The z distribution applies when the population standard deviation is known.

45. What question does a chi-square goodness-of-fit test address?

Whether observed category frequencies conform to a specified theoretical distribution
Whether two population means differ under a normal sampling model
Whether a population variance equals a specified numerical value
Whether two categorical variables are associated within a contingency table

Whether observed category frequencies conform to a specified theoretical distribution

Explanation

A goodness-of-fit test compares observed frequencies with frequencies expected from a specified theoretical distribution. Testing association between two categorical variables is the purpose of a chi-square independence test.

46. A goodness-of-fit experiment has n=200n=200 observations, and one category has hypothesized probability p=0.15p=0.15. What is its expected frequency?

E=15E=15
E=1,333.33E=1{,}333.33
E=30E=30
E=170E=170

$$E=30$$

Explanation

The expected frequency is calculated as E=np=200×0.15=30E=np=200\times0.15=30. The degrees-of-freedom rule for a goodness-of-fit test is a separate calculation, df=k1df=k-1.

47. For a goodness-of-fit test, the observed frequency in a category is 18 and the expected frequency is 15. What contribution does this category make to the chi-square statistic?

3.003.00
0.600.60
0.200.20
6.676.67

$$0.60$$

Explanation

The category contribution is (OE)2E=(1815)215=915=0.60\frac{(O-E)^2}{E}=\frac{(18-15)^2}{15}=\frac{9}{15}=0.60. The raw difference of 3 is not itself the chi-square contribution because the squared difference must be scaled by the expected frequency.

48. A contingency table has 4 rows and 3 columns. What are the degrees of freedom for its chi-square test of independence?

df=7df=7
df=5df=5
df=6df=6
df=12df=12

$$df=6$$

Explanation

For an independence test, the degrees of freedom are df=(R1)(K1)=(41)(31)=6df=(R-1)(K-1)=(4-1)(3-1)=6. The product of the row and column counts, 12, does not account for the constraints created by the marginal totals.

49. What does a one-way ANOVA test when comparing several populations?

Whether observations within a single sample have identical values
Whether two sample means differ using several explanatory factors
Whether more than two population means are equal using one explanatory factor
Whether population variances are equal without comparing population means

Whether more than two population means are equal using one explanatory factor

Explanation

One-way ANOVA evaluates equality of more than two population means through one factor or explanatory variable. Pairwise z or t tests instead compare two means at a time, making them less suitable for a simultaneous multi-group comparison.

50. Which set of conditions describes the assumptions of a one-way ANOVA?

Normal populations, unequal variances, and fixed dependent samples
Normal populations, equal population variances, and random independent samples
Skewed populations, unequal variances, and dependent observations
Uniform populations, equal variances, and systematically ordered samples

Normal populations, equal population variances, and random independent samples

Explanation

One-way ANOVA assumes normally distributed populations, equal population variances, and random independent samples. Unequal variances violate the stated equal-variance assumption rather than satisfying it.

51. What does the between-sample variance VAV_A estimate in one-way ANOVA?

Population variance from differences among observations within samples
Population variance from differences between the sample means
Sampling error from the number of observations in each sample
Total variation from differences between every pair of observations

Population variance from differences between the sample means

Explanation

VAV_A estimates population variance using differences between sample means, so it reflects variation between groups. Within-sample variance VRV_R instead uses differences among observations inside the samples.

52. If an ANOVA has between-sample variance VA=18V_A=18 and within-sample variance VR=6V_R=6, what is the test statistic?

F=24F=24
F=3F=3
F=12F=12
F=108F=108

$$F=3$$

Explanation

The ANOVA statistic is computed as F=VAVR=186=3F=\frac{V_A}{V_R}=\frac{18}{6}=3. Adding or multiplying the two variance estimates would not follow the ANOVA test-statistic formula.

53. Which equation represents the deterministic simple linear regression model?

Y=β0+εX+β1Y=\beta_0+\varepsilon X+\beta_1
Y=β0+β1X+εY=\beta_0+\beta_1X+\varepsilon
Y=β0X+β1+εY=\beta_0X+\beta_1+\varepsilon
Y=β0+β1XY=\beta_0+\beta_1X

$$Y=\beta_0+\beta_1X$$

Explanation

A deterministic linear model assigns one exact value of YY to each value of XX through Y=β0+β1XY=\beta_0+\beta_1X. The error term ε\varepsilon belongs to the stochastic model, which allows unexplained variation.

54. In the stochastic simple linear regression model, what does ε\varepsilon represent?

Omitted variables and random variations affecting YY
The measured value of the explanatory variable before fitting
The fixed change in YY caused by one unit of XX
The expected value of YY when XX equals zero

Omitted variables and random variations affecting $$Y$$

Explanation

In Y=β0+β1X+εY=\beta_0+\beta_1X+\varepsilon, ε\varepsilon captures omitted influences and random variation in the response. The slope represents the unit change in YY, while the intercept gives the value at X=0X=0.

55. Which statement correctly distinguishes homoskedasticity from autocorrelation in regression assumptions?

Homoskedasticity means fixed explanatory values, whereas autocorrelation concerns zero mean errors
Homoskedasticity means independent errors, whereas autocorrelation concerns unequal error variances
Homoskedasticity means equal error variances, whereas autocorrelation concerns dependence between errors
Homoskedasticity means linear effects, whereas autocorrelation concerns normally distributed errors

Homoskedasticity means equal error variances, whereas autocorrelation concerns dependence between errors

Explanation

Homoskedasticity requires the error variance to remain equal across observations, while autocorrelation concerns relationships among errors. These are distinct assumptions rather than two names for the same condition.

56. What does a Spearman rank correlation of rS=0.7r_S=-0.7 indicate?

A fairly strong inverse monotonic association between the ranked variables
A fairly strong direct monotonic association between the ranked variables
A correlation value outside the possible Spearman range
No monotonic association between the ranked variables

A fairly strong inverse monotonic association between the ranked variables

Explanation

Spearman's coefficient ranges from −1 to 1, and a negative value indicates an inverse monotonic association. A value of zero would indicate no monotonic association, while positive values indicate a direct association.

57. What is the distinction between estimating a mean response and predicting an individual response in regression?

Mean-response estimation concerns the slope, whereas individual prediction concerns the intercept
Mean-response estimation concerns an average YY, whereas individual prediction concerns one outcome
Mean-response estimation concerns the error variance, whereas individual prediction concerns the sample size
Mean-response estimation concerns one outcome, whereas individual prediction concerns an average YY

Mean-response estimation concerns an average $$Y$$, whereas individual prediction concerns one outcome

Explanation

Regression can estimate the average value of YY for a given XX or predict one individual outcome at that XX. Individual outcomes include random variation, so they are not the same target as the population mean response.

58. What expression gives the conditional mean of YY for a specified value of XX?

μYX=β0+β1X\mu_{Y|X}=\frac{\beta_0+\beta_1}{X}
μYX=β0X+β1ε\mu_{Y|X}=\beta_0X+\beta_1\varepsilon
μYX=β0+β1X+ε\mu_{Y|X}=\beta_0+\beta_1X+\varepsilon
μYX=β0+β1X\mu_{Y|X}=\beta_0+\beta_1X

$$\mu_{Y|X}=\beta_0+\beta_1X$$

Explanation

The conditional mean is the regression population line, μYX=β0+β1X\mu_{Y|X}=\beta_0+\beta_1X. The random error term is not included because the conditional mean describes the average response rather than an individual deviation.

59. Which interval formula is appropriate for predicting an individual response at X=XpX=X_p?

Y^p±tSYp\hat{Y}_p\pm tS_{Y_p}, where SYp=S1+1n+(XpXˉ)2SKXXS_{Y_p}=S\sqrt{1+\frac{1}{n}+\frac{(X_p-\bar{X})^2}{SK_{XX}}}
μYX±zS\mu_{Y|X}\pm zS, where SS excludes the individual-error component
Y^p±SYp\hat{Y}_p\pm S_{Y_p}, where the interval uses no degrees of freedom
Y^p±tSY^p\hat{Y}_p\pm tS_{\hat{Y}_p}, where SY^p=S1n+(XpXˉ)2SKXXS_{\hat{Y}_p}=S\sqrt{\frac{1}{n}+\frac{(X_p-\bar{X})^2}{SK_{XX}}}

$$\hat{Y}_p\pm tS_{Y_p}$$, where $$S_{Y_p}=S\sqrt{1+\frac{1}{n}+\frac{(X_p-\bar{X})^2}{SK_{XX}}}$$

Explanation

An individual prediction interval includes the extra individual-outcome variation through the leading term 1+1n1+\frac{1}{n} in SYpS_{Y_p}. The mean-response interval lacks that leading 1 and therefore describes an average response rather than one observation.

60. Why do regression confidence and prediction intervals use the t distribution?

The unknown error standard deviation is replaced by its sample estimate SS
The regression slope is calculated from ranks rather than measured values
The response variable has equal variance at every explanatory value
The explanatory variable is required to have a normal distribution

The unknown error standard deviation is replaced by its sample estimate $$S$$

Explanation

Regression intervals use the t distribution because the unknown error standard deviation is estimated by the sample regression standard error SS. Normality of errors is a regression assumption, but it does not by itself explain the use of the t distribution.

61. What does the sample Pearson correlation coefficient rr measure?

The causal effect of one variable on another variable
The degree of linear quantitative agreement between two variables
The proportion of variation explained by a nonlinear model
The average difference between the values of two variables

The degree of linear quantitative agreement between two variables

Explanation

Pearson’s rr measures the degree of linear quantitative agreement and ranges from 1-1 to 11. A zero value does not exclude every possible nonlinear association, and correlation itself does not establish causation.

62. If XX increases while YY tends to decrease, what type of correlation describes their relationship?

Zero correlation
Curvilinear correlation
Positive correlation
Negative correlation

Negative correlation

Explanation

A negative correlation indicates that YY tends to decrease as XX increases. Positive correlation describes the opposite directional tendency, with both variables generally increasing together.

63. When calculating Spearman’s rank correlation coefficient, what does dd represent for each paired observation?

The difference between the rank of XX and the rank of YY
The difference between the sample means of XX and YY
The product of the original values of XX and YY
The ratio of the two variables’ standard deviations

The difference between the rank of $$X$$ and the rank of $$Y$$

Explanation

For Spearman’s coefficient, the variables are ranked separately and d=uvd=u-v is the difference between their corresponding ranks. The formula then uses the sum of the squared rank differences.

64. Which expression represents a multiple linear regression model with one dependent variable and kk explanatory variables?

Y=β0+β1X1X2Xk+εY=\beta_0+\beta_1X_1X_2\cdots X_k+\varepsilon
Y=β0+β1X1+εY=\beta_0+\beta_1X_1+\varepsilon
X1=β0+β1Y+β2X2++βkXk+εX_1=\beta_0+\beta_1Y+\beta_2X_2+\cdots+\beta_kX_k+\varepsilon
Y=β0+β1X1+β2X2++βkXk+εY=\beta_0+\beta_1X_1+\beta_2X_2+\cdots+\beta_kX_k+\varepsilon

$$Y=\beta_0+\beta_1X_1+\beta_2X_2+\cdots+\beta_kX_k+\varepsilon$$

Explanation

Multiple regression models one dependent variable as a linear function of several explanatory variables plus an error term. A model with just one explanatory variable is simple regression, not multiple regression.

65. In a multiple regression model, how should a partial regression coefficient βi\beta_i be interpreted?

The average change in YY from a one-unit increase in XiX_i while other explanatory variables remain unchanged
The average change in every explanatory variable from a one-unit increase in YY
The total change in YY when all explanatory variables increase by one unit together
The correlation between XiX_i and YY without considering any other variable

The average change in $$Y$$ from a one-unit increase in $$X_i$$ while other explanatory variables remain unchanged

Explanation

A partial coefficient measures the average change in YY associated with a one-unit increase in XiX_i while the other explanatory variables are held constant. This control for the remaining predictors distinguishes it from a simple regression coefficient.

66. Which condition is an assumption of multiple linear regression?

The explanatory variables must all have identical sample means
The explanatory variables have no linear dependence among themselves
The dependent variable must have a uniform distribution
The regression coefficients must be equal to one another

The explanatory variables have no linear dependence among themselves

Explanation

Multiple regression assumes that the explanatory variables do not exhibit linear dependence that prevents their effects from being separated. Identical means, a uniform dependent-variable distribution, and equal coefficients are not required assumptions.

67. A fitted model predicts YY for an XX value outside the range observed in the sample. What is this procedure called?

Standardization
Extrapolation
Interpolation
Aggregation

Extrapolation

Explanation

Extrapolation predicts or estimates YY for an XX value outside the sample’s observed range. Interpolation instead concerns an XX value within the observed range.

68. How should a prediction be interpreted when its target XX value lies far beyond the sample-data range?

As evidence that the original sample range was sufficiently broad
As an interpolation because the model supplies a fitted value
With greater caution because uncertainty increases farther from the observed range
With less caution because the fitted trend becomes more established

With greater caution because uncertainty increases farther from the observed range

Explanation

Extrapolated conclusions require increasing caution as the target moves farther from the sample-data range. A fitted value does not make a distant prediction an interpolation or demonstrate that the original range was broad enough.

69. A product costs 80 units in the base period and 100 units in the current period. What is its index number?

125
180
80
100

125

Explanation

The index number is the current level divided by the base-period level and multiplied by 100: 10080×100=125\frac{100}{80}\times100=125. An index of 125 therefore represents a 25 percent increase over the base period.

70. Which statement correctly distinguishes base indices from chain indices?

Base indices describe groups of phenomena, whereas chain indices describe individual phenomena
Base indices compare each period with the preceding period, whereas chain indices use a fixed reference period
Base indices measure decreases, whereas chain indices measure increases
Base indices use a fixed reference period, whereas chain indices compare each period with the preceding period

Base indices use a fixed reference period, whereas chain indices compare each period with the preceding period

Explanation

A base index compares every period with one fixed reference period, while a chain index compares each period with the immediately preceding period. The distinction concerns the reference period, not whether the index measures increases or decreases.

Review with flashcards

Memorize the answers with 94 flashcards on Statistics Foundations and Probability.

What is Statistics as a scientific method?

It is used to collect, present, analyze, interpret data, and draw conclusions.

What does theoretical statistics develop and prove?

Statistical theorems, methods, formulas, rules, and laws.

How does applied statistics use statistical methods?

It uses them to solve real problems.

See flashcards →

Read the study sheet

Read the complete study sheet on Statistics Foundations and Probability.

See study sheet →

Similar courses

Create your own quizzes

Import your course and AI generates quizzes with corrections in 30 seconds.

Quiz generator