Name:
Each question is worth 1 point for a total of 20 points. Partial credit will be given
where appropriate.
You own a bakery and decide to compare your weekly flour consumption in
pounds (x- variable) and the sales you make in dollars (y-variable) each week. You
enter your raw data into Excel and run a simple linear regression. Below are your
summary output results.
1)
SUMMARY OUTPUT
Regression Statistics
Multiple R
R Square
Adjusted R
Square
Standard
Error
0.7529
0.5669
0.5128
24.1940
Observ ations
10
ANOVA
df
SS
MS
F
10.472
Regression
1
6129.716
6129.716
Residual
8
4682.784
585.348
Total
Significance
F
9 10812.500
-214.45
Standard
Error
Stat
113.65
5.53
1.71
Coefficients
Intercept
X Variable 1
-1.89
Pvalue
95%
0.10
3.24
0.01
t
0.012
Upper
95%
-476.54
47.63
Lower
95.1%
476.54
1.59
9.48
1.59
Lower
Upper 95.0%
a)
Using the above data, identify the simple linear regression model
(equation) that could be used to predict sales based on flour consumption.
ŷ=
b)
Identify the percent of variability in sales (y) that is explained by y’s
relationship with flour consumption (x).
percent variability =
Assume you are in the analysis phase of a project. Name three statistical
tools that you could apply to your data in order to drive your next steps in the DMAIC
process.
a)
b)
2)
Page 1 of 5
47.63
9.48
c)
Page 2 of 5
Describe the following data five different ways (include numbers in your answer):
3)
Data: 13, 10, 15, 11, 12, 12, 7, 8, 16, 14
a)
b)
c)
d)
e)
In a normal distribution of measurements having a mean of 500 feet and
a standard deviation of 80 feet, what percent of the distribution falls between
300 and 450 feet?
4)
%
You have just performed a linear regression analysis on successive values of a
2
time series and you see autocorrelation. What might your r be equal to?
5)
2
r =
The null hypothesis is Ho: µ = 10, and the alternative hypothesis is Ha: µ ≠
10. Assume alpha = 0.01. If the null hypothesis was rejected, what would the 99%
confidence interval for µ look like?
6)
a) (12.1, 15.3)
b) (8.5, 12.1)
c) (5.3, 15.5)
d) (9.8, 10.5)
7)
Given the above range chart, what can you conclude?
a) Process variation is unstable and unpredictable.
b) Measurement variation is declining.
c) Within-subgroup variation is stable and predictable.
d) Discrimination is a problem.
8)
Describe two ways to determine whether your measurement system is
repeatable and reproducible:
a)
b)
9)
You are interested in developing a control chart. Your dimension of concern is the
diameter of a cylindrical part. Every hour you take two measurements. Choose
the most appropriate chart.
a) np chart
10)
b) IMR chart
c) c chart
d) x-bar/R chart
You are interested in developing a second control chart. However, you have
collected data on the number of visible scratches on the part. You take a small
constant subgroup size of 2 every day. Choose the most appropriate chart.
a) np chart
b) IMR chart
c) c chart
d) x-bar/R chart
A hypothesis is being tested at alpha = 0.05. At which of the following p-values
would the null hypothesis be rejected?
11)
a) 0.150
b) 0.005
c) 0.055
d) 0.350
Page 3 of 5
12) It was reported in USA Today that from 1999 through 2003 the number of daily spam
messages sent worldwide was:
(x)Year Number
1
2
3
4
5
Year (y) Spam Messages Sent (billions)
1999
1.0
2000
2.3
2001
4.0
2002
5.6
2003
7.3
The regression equation was determined to be: y = –0.73 + 1.59 x
where y is the number of spam messages sent in billions and x is the year number.
Using the model, what is the predicted number of spam messages sent in 2004?
a) 9.00 billion
b) 7.85 billion
13) If we want to detect a
a) continuous
c) 3,185 billion
d) 8.81 billion
change in the process, we increase our sample size.
b) larger
c) smaller
d) normal
14) Specific models have been developed to aid in the analysis of time series data when
usual regression methods are not appropriate. What model uses the average of the
last several values of a time series to forecast the next value?
Name of the model:
15) A strong correlation does not mean a cause-and-effect relationship. Causation is only
one explanation of an observed association. What else could produce a strong
correlation?
a) Confounding factor
b) Coincidence
c) Common cause
d) All of the above
16)
A correlation coefficient r = –0.72 would indicate:
a) There is a strong positive correlation between two factors
b) There is a moderate negative correlation between two factors
c) There is no correlation between four factors
d) There is a moderate positive correlation between four factors
17) When the variability in x decreases (for example outliers are removed from the
data), the correlation coefficient gets closer to
.
Page 4 of 5
18)You enter your data into Excel and run a multiple regression. Below are your
summary output results. What variables are significant and should be included in
your model?
Name the variables:
Regression Statistics
Multiple R
0.94898
R Square
0.90056
Adjusted R
Square
0.82101
Standard Error
6.24921
Observations
10
ANOVA
df
4
5
9
Coefficients
48.3628
-21.3770
-12.4787
-4.2240
0.2849
Regression
Residual
Total
Intercept
weight
height
power
speed
SS
1768.3366
195.2634
1963.6000
Standard
Error
14.1772
9.7575
7.3110
8.8810
0.3917
MS
442.0841
39.0527
F
11.3202
Significance
F
0.0101
t Stat
3.4113
-2.1908
-1.7068
-0.4756
0.7272
P-value
0.0190
0.0300
0.1486
0.6544
0.4997
Lower 95%
11.9192
-46.4595
-31.2723
-27.0534
-0.7221
19) Using the above Excel summary output results, answer the following questions.
a) How many samples were collected to generate this data?
b) What is the correlation for this multiple regression, and what does it
indicate?
20) Certain data is inappropriate for a regression analysis such as:
a)
b)
c)
d)
Residuals that form a pattern when plotted
There aren’t any outliers
The correlation coefficient is less than 1
All of the above

