Name:
Each question is worth 1 point for a total of 20 points. Partial credit will be given
where appropriate.
You own a bakery and decide to compare your weekly flour consumption in
pounds (x- variable) and the sales you make in dollars (y-variable) each week. You
enter your raw data into Excel and run a simple linear regression. Below are your
summary output results.
1)

SUMMARY OUTPUT
Regression Statistics
Multiple R
R Square
Adjusted R
Square
Standard
Error

0.7529
0.5669
0.5128
24.1940

Observ ations

10

ANOVA
df

SS

MS

F
10.472

Regression

1

6129.716

6129.716

Residual

8

4682.784

585.348

Total

Significance
F

9 10812.500

-214.45

Standard
Error
Stat
113.65

5.53

1.71

Coefficients
Intercept
X Variable 1

-1.89

Pvalue
95%
0.10

3.24

0.01

t

0.012

Upper
95%

-476.54

47.63

Lower
95.1%
476.54

1.59

9.48

1.59

Lower

Upper 95.0%

a)
Using the above data, identify the simple linear regression model
(equation) that could be used to predict sales based on flour consumption.
ŷ=
b)
Identify the percent of variability in sales (y) that is explained by y’s
relationship with flour consumption (x).
percent variability =
Assume you are in the analysis phase of a project. Name three statistical
tools that you could apply to your data in order to drive your next steps in the DMAIC
process.
a)
b)
2)

Page 1 of 5

47.63
9.48

c)

Page 2 of 5

Describe the following data five different ways (include numbers in your answer):

3)

Data: 13, 10, 15, 11, 12, 12, 7, 8, 16, 14
a)
b)
c)
d)
e)

In a normal distribution of measurements having a mean of 500 feet and
a standard deviation of 80 feet, what percent of the distribution falls between
300 and 450 feet?
4)

%

You have just performed a linear regression analysis on successive values of a
2
time series and you see autocorrelation. What might your r be equal to?
5)

2

r =

The null hypothesis is Ho: µ = 10, and the alternative hypothesis is Ha: µ ≠
10. Assume alpha = 0.01. If the null hypothesis was rejected, what would the 99%
confidence interval for µ look like?
6)

a) (12.1, 15.3)

b) (8.5, 12.1)

c) (5.3, 15.5)

d) (9.8, 10.5)

7)

Given the above range chart, what can you conclude?
a) Process variation is unstable and unpredictable.
b) Measurement variation is declining.
c) Within-subgroup variation is stable and predictable.
d) Discrimination is a problem.
8)

Describe two ways to determine whether your measurement system is
repeatable and reproducible:
a)

b)

9)

You are interested in developing a control chart. Your dimension of concern is the
diameter of a cylindrical part. Every hour you take two measurements. Choose
the most appropriate chart.
a) np chart

10)

b) IMR chart

c) c chart

d) x-bar/R chart

You are interested in developing a second control chart. However, you have
collected data on the number of visible scratches on the part. You take a small
constant subgroup size of 2 every day. Choose the most appropriate chart.
a) np chart

b) IMR chart

c) c chart

d) x-bar/R chart

A hypothesis is being tested at alpha = 0.05. At which of the following p-values
would the null hypothesis be rejected?
11)

a) 0.150

b) 0.005

c) 0.055

d) 0.350
Page 3 of 5

12) It was reported in USA Today that from 1999 through 2003 the number of daily spam
messages sent worldwide was:
(x)Year Number
1
2
3
4
5

Year (y) Spam Messages Sent (billions)
1999
1.0
2000
2.3
2001
4.0
2002
5.6
2003
7.3

The regression equation was determined to be: y = –0.73 + 1.59 x
where y is the number of spam messages sent in billions and x is the year number.
Using the model, what is the predicted number of spam messages sent in 2004?
a) 9.00 billion

b) 7.85 billion

13) If we want to detect a
a) continuous

c) 3,185 billion

d) 8.81 billion

change in the process, we increase our sample size.
b) larger

c) smaller

d) normal

14) Specific models have been developed to aid in the analysis of time series data when
usual regression methods are not appropriate. What model uses the average of the
last several values of a time series to forecast the next value?
Name of the model:
15) A strong correlation does not mean a cause-and-effect relationship. Causation is only
one explanation of an observed association. What else could produce a strong
correlation?
a) Confounding factor
b) Coincidence
c) Common cause
d) All of the above

16)

A correlation coefficient r = –0.72 would indicate:

a) There is a strong positive correlation between two factors
b) There is a moderate negative correlation between two factors
c) There is no correlation between four factors
d) There is a moderate positive correlation between four factors
17) When the variability in x decreases (for example outliers are removed from the
data), the correlation coefficient gets closer to
.
Page 4 of 5

18)You enter your data into Excel and run a multiple regression. Below are your
summary output results. What variables are significant and should be included in
your model?
Name the variables:
Regression Statistics
Multiple R
0.94898
R Square
0.90056
Adjusted R
Square
0.82101
Standard Error
6.24921
Observations
10
ANOVA
df
4
5
9
Coefficients
48.3628
-21.3770
-12.4787
-4.2240
0.2849

Regression
Residual
Total

Intercept
weight
height
power
speed

SS
1768.3366
195.2634
1963.6000
Standard
Error
14.1772
9.7575
7.3110
8.8810
0.3917

MS
442.0841
39.0527

F
11.3202

Significance
F
0.0101

t Stat
3.4113
-2.1908
-1.7068
-0.4756
0.7272

P-value
0.0190
0.0300
0.1486
0.6544
0.4997

Lower 95%
11.9192
-46.4595
-31.2723
-27.0534
-0.7221

19) Using the above Excel summary output results, answer the following questions.
a) How many samples were collected to generate this data?
b) What is the correlation for this multiple regression, and what does it
indicate?
20) Certain data is inappropriate for a regression analysis such as:
a)
b)
c)
d)

Residuals that form a pattern when plotted
There aren’t any outliers
The correlation coefficient is less than 1
All of the above