Test #1
Question 2:
Experimental data is data collected through test methods, measurement, and experimental designs. Like in clinical research, the data produced through experiments are results of clinical trials. Hence, experimental data can be quantitative or qualitative, depending on the investigations. Qualitative data are subjective and descriptive in comparison to getting a continuous measurement scale producing numbers. Quantitative data is gathered in an experimentally repetitive manner (“OpenStax CNX,” 2020). The qualitative data is more related to event meaning hence, subjective to interpretation by observers. Different investigators can generate the experimental data, and a variety of mathematical analysis techniques may be performed on the data.
Question 3:
Observation is an example of a primary data collection method. It involves observation, recording, and use of data by the researcher making observations. The data is always firsthand and has never been used anywhere. Secondary data collection is data that has not been collected by the researcher. It involves the use of data that has not been personally collected by the researcher. Examples of secondary sources of data include books, journals, magazines, and newspapers (“Ch. 12 Introduction – Introductory Statistics | OpenStax”, 2020).
Question 4:
Time series data is a set of observations collected at usually discrete and equally spaced time intervals. Cross-sectional data are observations that come from different individuals or groups at a single point in time (Kazmier, 2004). The table data consists of three different categories: Dow Jones Industrial, NASDAQ, and S&P, at an exact point in time, July 29, 2009. Hence, the given data is of a cross-sectional type.
Question 5:
Discrete data is unique in the sense that they can only take specific values. The numbers may potentially be infinite of the values; however, each is distinct with no grey area in between. The data can be numeric, such as many oranges and can be categorical like male or female, good or bad, red or blue.
Continuous data has no restrictions to defined separate values; however, it can occupy any value form over a given continuous range. Notably, between any two continuous values, there can be an infinite number of other values. The data is always in a numeric form. In the process, it makes sense to treat continuous data as discrete and vice versa. Categorical data may also be considered as continuous data.
Question 6:
Data comes in different sizes and shapes. The data measurements are the basis of different things at various times. In most situations, financial analysts concentrate on a specific type of data, such as cross-sectional data and time-series data (Kazmier, 2004).
Time series data refers to observations recorded at equal and discretely spaced time intervals (Kazmier, 2004). Cross-sectional data refers to the observations made on different groups or individuals at a single moment in time.
Question 7:
Outliers are numbers within a set of data that are smaller or larger in comparison to other set values. The median, mode, and the mean are measures of central tendency. However, it is only the mean that is always affected by an outlier. The mean, which is sometimes known as the average, is the most popular measure of central tendency.
An outlier can also be an extreme value within a data set, much lower or higher than the other values or numbers.
The mean is the standard set of data whose computation is through the division of the sum of data and the number of items within the set. The median is the mid-value when data is arranged numerically and may also be the average of the two mid values if the set is of an even number of data. Lastly, the model is the number with the most appearances or repetition within the data set. The outliers only affect the mean but have little impact on the mode and median of a data set.
Example:
A college student receives following scores in a test of 10 units: 0, 60, 60, 70, 75, 80, 80, 80, 85, and 90.
Outlier = 0
Mean = (0 + 60 + 60 + 70 + 75 + 80 + 80 +80 + 85 + 90)/10
= 680/10
= 68
Median = (75+80)/2
=155/2
= 77.5
Mode = 80 since it occurs three times and more in comparison with the other values.
Getting a zero on a test affected the students’ mean.
Question 8:
- Mothers not employed = 1-0.78
= 0.22
=22%
- Working full time = 0.64*78%
= 49.92%
= 50%
Employed part-time = 78% -50%
= 28%
Question 10:
The coefficient of variation refers to a statistical measure of dispersion where data points within a data series surround the mean. It is a representation of the ratio between the standard deviation and means. Moreover, it is useful in comparison to the level of variation from data series to the other, without consideration of the means.
Question 11:
C.I- Class means Interval
f- Denotes frequency
x- Denotes frequency
C.I
X
F
xf
X^2f
20 – 30
25
67
1,675
91,875
30 – 40
35
111
3,885
135,975
40 – 50
45
125
5,625
253,125
50 – 60
55
21
1,155
63,525
60 – 70
65
38
2,470
160,550
TOTAL
–
362
14,810
655,050
Note that class mark of a class is;
X = (LCL + UCL)/2
LCL = Lower class limit
UCL = Upper class limit
The average income;
= the total frequency
(n) = the number of classes
N=362, n = 5
= 40.9116
Average income of the respondents = $ (40.9116 *1,000)
= $ 40,911.60
= 136.1473
Standard deviation is
=
= 11.6682
The standard deviation for the respondent’s income is;
= $ (11.6682 * 1,000)
= $ 11,668.20
Test #2
Question 1:
E(x) = 23,200
Sd(x) = 7,500
N = 30
P(X < 24,000) = P (Z< [(24,000 – 23,200)/ (7,500/√30)])
= P (Z< [(800/ (7,500/√30)])
= P (Z< [(800/ (7,500/5.47723)])
= P (Z< [800/1,369.3064])
= P (Z< [800/1,369.3064])
= P (Z< 0.5842)
= 0.7190
Question 2:
A significance level of 0.10 shows a 10% probability result due to opportunity inference used to reject or support claims based on sample data.
Question 3:
Systematic random sampling creates a list of each member of the population. The values are randomly selected from the first sample elements in the population list. After which any given element can be selected from the list. The method is distinct from simple random sampling because every possible sample of given elements is not equal. It involves the selection of sample values from the ordered frame. The sampling frame is a list of participants where a sample is taken (“OpenStax CNX,” 2020). The items are selected completely randomly; hence, each element has an equal opportunity of being chosen.
In stratified sampling, the population under study is divided into groups depending on some given characteristics. A probability sample, usually simple random sampling, is chosen within each group. The groups are known as strata. An example is the division of a population into groups depending on geographical locations during a national survey. After that, respondents can be randomly selected within each group. Each subpopulation is sampled independently.
The first step is dividing the population into similar subgroups before sampling. Population members belong to single groups.
Question 4:
A sampling error refers to a statistical error which occurs when an analyst fails to select a representative sample of entire population data and results from the sample does not represent that of the entire population.
The sampling error includes; sample frame error, which occurs when the wrong sub-population is used in a sample. The result leads to wrong outcomes, and the only solution is to repeat sampling using the correct sub-population. The population error occurs when researchers fail to understand the people they should survey (Tiemann, n.d.).
Selection error occurs when respondents select their participants in the study, and only the ones interested respond. The error can be regulated by going extra miles to get participants. None response errors occur in situations where respondents are different from the ones who failed to respond. The error can be corrected through follow-ups on surveys using alternatives. Sampling errors can be reduced through the development of careful sample designs, the use of large samples, and the establishment of many contacts to affirm representative response.
Question 6:
Under probability sampling, all the population members have pre-specified and equal chances of being part of the sample. In contrast, in non-probability sampling, all the items of the universe have no equal opportunity to be part of the sample. The difference is that nonprobability samplings do not involve random selection, while probability sampling each population member has a fair chance of selection.
The target population in cluster sampling is divided into several clusters. The clusters are randomly selected for sampling or multiple stages sampling, which is done to form a target sample, while in stratified sampling, a large population is divided into homogeneous, unique strata and members randomly selected, forming a sample. The sampling elements selected are giving the entire population equal chances of being part of the samples.
The difference is that in cluster sampling, the clusters are treated as sampling units; thus, sampling depends on a population cluster. While, in stratified sampling, the samples are taken from the elements in each stratum.
Question 7:
A 90% confidence level means 90% of the interval estimates should be expected to be included in the population parameter (“OpenStax CNX,” 2020). Similarly, a 99% confidence level shows that 95% of the intervals should be in the parameter.
Question 8:
Null hypothesis means, no difference (“OpenStax CNX,” 2020)
Hence, null hypothesis Ho: u = 60,000
Alternate hypothesis is the research hypothesis
Here claim is that average is changed
So, alternate hypothesis Ha: u not equal to 60,000
Ho: u = 60,000
Ha: u not equal to 60,000
Question 9:
P= x/n
(x) =274
(n)= 330
P= 274/330
= 0.8303
At 90% confidence interval in estimation of these proportion;
Margin of error α = 1-0.90
= 0.10
ME =
= 1.6449 *
= 1.6449 *
= 1.6449 *
= 1.6449 *
= 1.6449 * 0.02066
= 0.03399
Lower confidence level is;
LCL = (p – ME)
= (0.8303 – 0.03399)
= 0.79631
Upper confidence limit is;
UCL = (P – ME)
= (0.8303 + 0.03399)
= 0.86429
By observing the confidence interval calculated, it is concluded that the claim of the GMAC report is validated.
Question 10:
Type-II error is also referred to as the consumer’s risk, which occurs as a result of not rejecting maybe a worthless service or product shown by the null hypothesis. Furthermore, type II error; occurs when one fails to reject the null hypothesis, and it is false. The error shows that a condition has failed though it was successful. Type-II error probability is denoted by β.
Question 11:
Hypothesis Testing must meet then following situations; simple random sampling is the sampling method used; every sample has a possibility of two outcomes known as success or failure; the sample includes at least ten failures and ten successes; and the population size is 20 times as big as the sample size(“OpenStax CNX,” 2020).
Hypothesis test process entails four steps: stating the hypotheses, formulating an analysis plan, analyzing sample data, and interpreting results.
When stating the hypothesis, every test requires the analysis to state an alternative hypothesis and null hypothesis. The process must ensure that they are stated mutually exclusive, with either true or false.
Formulating an analysis plan highlights the use of sample data to reject or accept the null hypothesis. The significance level and test method should be specified. The significance level chosen in most scenarios equals 0.01, 0.05, or 0.10 through any vale from 0 to 1 can be used. The test method used is the z-test in the determination of whether the hypothesis population proportion significantly differs from sample proportions observed.
In the sample analysis, the test statistic is calculated, and the association with P-Value established. The computing of the standard deviation of sample distribution using the formula σ = √ [P * (1 – P) / n]. The n is the sample size, and P is the null hypothesis value within-population proportion. The test statistic is a z-score whose computation uses the equation, z = (p – P) / σ; P is the null hypothesis value, p is proportion sample, and σ is the sampling distribution standard deviation.
The results are interpreted based on the findings. In the situation of the null hypothesis, the researcher is to reject the null hypothesis. It involves a comparison of the P-value and the significance level. Further rejecting the null hypothesis in case the significance level is higher than P-value (“OpenStax CNX,” 2020).