Skip to ContentGo to accessibility page

Figure 4.1 Inferential statistics is used extensively in data science to draw conclusions about a larger population—and drive decision-making. (credit: modification of work “Compiling Health Statistics” by Ernie Branson/Wikimedia Commons, Public Domain)

Introduction

Inferential statistics plays a key role in data science applications, as its techniques allow researchers to infer or generalize observations from samples to the larger population from which they were selected. If the researcher had access to a full set of population data, then these methods would not be needed. But in most real-world scenarios, population data cannot be obtained or is impractical to obtain, making inferential analysis essential. This chapter will explore the techniques of inferential statistics and their applications in data science.

Confidence intervals and hypothesis testing allow a data scientist to formulate conclusions regarding population parameters based on sample data.

One technique, correlation analysis, allows the determination of a statistical relationship between two numeric quantities, often referred to as variables. A variable is a characteristic or attribute that can be measured or observed. A correlation between two variables is said to exist where there is an association between them. Finance professionals often use correlation analysis to predict future trends and mitigate risk in a stock portfolio. For example, if two investments are strongly correlated, an investor might not want to have both investments in a certain portfolio since the two investments would tend to move in the same direction as market prices rose or fell. To diversify a portfolio, an investor might seek investments that are not strongly correlated with one another.

Regression analysis takes correlation analysis one step further by modeling the relationship between the two numeric quantities or variables when a correlation exists. In statistics, modeling refers specifically to the process of creating a mathematical representation that describes the relationship between different variables in a dataset. The model is then used to understand, explain, and predict the behavior of the data.

This chapter focuses on linear regression, which is analysis of the relationship between one dependent variable and one independent variable, where the relationship can be modeled using a linear equation. The foundations of regression analysis have many applications in data science, including in machine learning models where a mathematical model is created to determine a relationship between input and output variables of a dataset. Several such applications of regression analysis in machine learning are further explored in Decision-Making Using Machine Learning Basics. In Time Series and Forecasting, we will use time series models to analyze and predict data points for data collected at different points in time.

Citation/Attribution
Reuse and redistribution of this content in digital or print format:
  • This book may not be used in the training of large language models or otherwise be ingested into large language models or generative AI offerings without OpenStax's prior written permission.
  • This book uses the Creative Commons Attribution-NonCommercial-ShareAlike License, which means that you can reuse and modify the material only for noncommercial purposes, must attribute OpenStax, and must distribute any derivative works under the same license.
  • Any commercial printing of this textbook, including using a local or custom printer, must be approved by OpenStax, and proper citation provided.
  • OpenStax-copyrighted images, activities, assessments, and similar components of this book are subject to the same licensing – CC-BY-NC-SA. They can be used for noncommercial purposes with attribution. Commercial use requires permission.
  • Permission requests: Anyone who intends to incorporate this content (including text, images, and other components) into large language models, use it in AI offerings, use it commercially (including in print), and/or has questions about another use case is welcome to complete our reuse request form.
Attribution information
  • If you are redistributing all or part of this book in a noncommercial print format, then you must include on every physical page the following attribution:

    Access for free at https://openstax.org/books/principles-data-science/pages/1-introduction

  • If you are redistributing all or part of this book in a noncommercial digital format, then for every page that includes OpenStax content, you must license the derivative work under the same CC-BY-NC-SA license as the original, and include on every digital page view the following attribution:

    Access for free at https://openstax.org/books/principles-data-science/pages/1-introduction

Citation information

The information below includes the information needed to generate citations in most major styles (APA, MLA, etc.); you must reformat and organize the information as needed to fit the requirements of the style. Use the information below to generate a citation. We recommend using a citation tool such as this one.

© Apr 23, 2026 OpenStax. Textbook content produced by OpenStax is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike License. The OpenStax name, OpenStax logo, OpenStax book covers, OpenStax CNX name, and OpenStax CNX logo, and Rice University name, and Rice University logo trademarks, or wordmarks are not subject to the Creative Commons license and may not be reproduced without the prior and express written consent of Rice University.