Skip to ContentGo to accessibility page
Principles of Data Science

Quantitative Problems

Principles of Data ScienceQuantitative Problems

Quantitative Problems

1 .
An automotive designer wants to create a 90% confidence interval for the mean length of a body panel on a new model electric vehicle. From a sample of 35 cars, the sample mean length is 37 cm and the sample standard deviation is 2.8 cm.
a.
Create a 90% confidence interval for the population mean length of the body panel using Python.
b.
Create a 90% confidence interval for the population mean length of the body panel using Excel.
c.
Create a 95% confidence interval for the population mean length of the body panel using Python.
d.
Create a 95% confidence interval for the population mean length of the body panel using Excel.
e.
Which confidence interval is narrower, the 90% or 95% confidence interval?
2 .
A political candidate wants to conduct a survey to generate a 95% confidence interval for the proportion of voters planning to vote for this candidate. Use a margin of error of 3% and assume a prior estimate for the proportion of voters planning to vote for this candidate is 58%. What sample size should be used for this survey?
3 .
A medical researcher wants to test the claim that the mean time for ibuprofen to relieve body pain is 12 minutes. A sample of 25 participants are put in a medical study, and data is collected on the time it takes for ibuprofen to relieve body pain. The sample mean time is 13.2 minutes with a sample standard deviation of 3.7 minutes.
a.
Calculate the p-value for this hypothesis test using Python.
b.
Calculate the p-value for this hypothesis test using Excel.
c.
Test the researcher’s claim at a 5% level of significance.
4 .
A consumer group is studying the percentage of drivers who wear seat belts. A local police agency states that the percentage of drivers wearing seat belts is at least 70%. To test the claim, the consumer group selects a random sample of 150 drivers and finds that 117 are wearing seat belts.
a.
Calculate the p-value for this hypothesis test using Python.
b.
Calculate the p-value for this hypothesis test using Excel.
c.
Use a level of significance of 0.10 to test the claim made by the local police agency.
5 .
A local newspaper is comparing the costs of a meal at two local fast-food restaurants. The newspaper claims that there is no difference in the average price of a meal at the two restaurants. A sample of diners at each of the two restaurants is taken with the following results:

Restaurant A: Sample mean price of a meal = $ 17.55 , sample standard deviation = $ 7.45 , sample size = 15
Restaurant B: Sample mean price of a meal = $ 22.75 , sample standard deviation = $ 8.10 , sample size = 18
a.
Calculate the p-value for this hypothesis test using Python.
b.
Calculate the p-value for this hypothesis test using Excel.
c.
Use a level of significance of 0.10 to test the claim made by the newspaper. Assume population variances are not equal.
6 .
A health official claims the proportion of men smokers is greater than the proportion of women smokers. To test the claim, the following data was collected:

Men: Sample size = 150 , Number of smokers = 41
Women: Sample size = 140 , Number of smokers = 31
a.
Calculate the p-value for this hypothesis test using Python
b.
Calculate the p-value for this hypothesis test using Excel.
c.
Use a level of significance of 0.05 to test the claim made by the health official.
7 .
A software company has been tracking revenues versus cash flow in recent years, and the data is shown in the table below. Consider cash flow to be the dependent variable.
Revenues ($000s) Cash Flow ($000s)
212 81
239 84
218 81
297 99
287 91
Table 4.19 Software Co. Revenue vs. Cash Flow
a.
Calculate the correlation coefficient for this data (all dollar amounts are in thousands) using Python.
b.
Calculate the correlation coefficient using Excel.
c.
Determine if the correlation is significant using a significance level of 0.05.
d.
Construct the best-fit linear equation for this dataset.
e.
Predict the cash flow when revenue is 250.
f.
Calculate the residual for the revenue value of 239.
8 .
A human resources administrator calculates the correlation coefficient for salary versus years of experience as 0.68. The dataset contains 10 ( x , y ) data points. Determine if the correlation is significant or not significant at the 0.05 level of significance.
9 .
A biologist is comparing three treatments of light intensity on plant growth and is interested to know if the average plant growth is different for any of the three treatments.
Sample data is collected on plant growth (in centimeters) for the three treatments, and data is summarized in the table below.
Use an ANOVA hypothesis test method with 0.05 level of significance.
Treatment A Treatment B Treatment C
4.3 3.9 2.8
3.7 4.1 2.9
3.1 3.6 3.5
3.4 3.5 3.9
4.1 4.3 4.4
4.0 4.4 4.3
N/A n/a 4.1
Table 4.20
a.
Calculate the p-value for this hypothesis test using Python.
b.
Determine if the average plant growth is different for any of the three treatments.
Citation/Attribution
Reuse and redistribution of this content in digital or print format:
  • This book may not be used in the training of large language models or otherwise be ingested into large language models or generative AI offerings without OpenStax's prior written permission.
  • This book uses the Creative Commons Attribution-NonCommercial-ShareAlike License, which means that you can reuse and modify the material only for noncommercial purposes, must attribute OpenStax, and must distribute any derivative works under the same license.
  • Any commercial printing of this textbook, including using a local or custom printer, must be approved by OpenStax, and proper citation provided.
  • OpenStax-copyrighted images, activities, assessments, and similar components of this book are subject to the same licensing – CC-BY-NC-SA. They can be used for noncommercial purposes with attribution. Commercial use requires permission.
  • Permission requests: Anyone who intends to incorporate this content (including text, images, and other components) into large language models, use it in AI offerings, use it commercially (including in print), and/or has questions about another use case is welcome to complete our reuse request form.
Attribution information
  • If you are redistributing all or part of this book in a noncommercial print format, then you must include on every physical page the following attribution:

    Access for free at https://openstax.org/books/principles-data-science/pages/1-introduction

  • If you are redistributing all or part of this book in a noncommercial digital format, then for every page that includes OpenStax content, you must license the derivative work under the same CC-BY-NC-SA license as the original, and include on every digital page view the following attribution:

    Access for free at https://openstax.org/books/principles-data-science/pages/1-introduction

Citation information

The information below includes the information needed to generate citations in most major styles (APA, MLA, etc.); you must reformat and organize the information as needed to fit the requirements of the style. Use the information below to generate a citation. We recommend using a citation tool such as this one.

© Apr 23, 2026 OpenStax. Textbook content produced by OpenStax is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike License. The OpenStax name, OpenStax logo, OpenStax book covers, OpenStax CNX name, and OpenStax CNX logo, and Rice University name, and Rice University logo trademarks, or wordmarks are not subject to the Creative Commons license and may not be reproduced without the prior and express written consent of Rice University.