Skip to ContentGo to accessibility page

Figure 8.1 Big data taxonomy includes data providers, data consumers, data owners, and data viewers while providing the data flow between them. (credit: modification of “DARPA Big Data” by Defense Advanced Research Projects Agency (DARPA)/Wikimedia Commons, Public Domain)

Introduction

Managing data today requires an end-to-end perspective and the ability to combine the use of various types of data. For example, Amazon needs to run its day-to-day businesses and sell products to customers, handle returns, pay commissions to retailers and wholesalers, and develop and price new products. This is all part of day-to-day operations, and most retail companies strive to achieve operational excellence as it is a key driver of customer satisfaction. For that purpose, retailers typically use traditional relational databases and structured data to run their operations. At the same time, retailers need to remain competitive and develop with pricing strategies that attract and retain customers. Doing so requires being able to analyze market prices accordingly before advertising products to their customers. For that part, businesses generally rely on datasets and big data analytics to predict competitive prices in order to optimize sales. Overall, businesses and organizations that deliver products to customers need to leverage insights coming from metadata to transform as they perform and remain competitive while sustaining their operations. It is no longer possible to rely on operational excellence to ensure continued success.

In this chapter, we will cover the spectrum of data management activities, including the storage and retrieval of various types of data. In other words, the chapter is about how big data are managed today from an end-to-end standpoint.

As an example, TechWorks is a start-up company that is 100% committed to leveraging innovative technologies as part of its repeatable business model and as a business growth facilitator. TechWorks has multiple departments including human resources, finance, sales, marketing, operation management, and information technology. Each department has its own information system and database; they do not share any resources. TechWorks has many success stories in the market. However, they have a huge problem with integrating their reports to aid in decision-making. TechWorks has many challenges that impact their competitiveness and survival in the market, including their traditional database not working properly; their practices in collecting, managing, and analyzing data not promoting better business decision-making; and their servers being old and sometimes unable to sufficiently handle the company data. The chief executive officer (CEO) of TechWorks decides to develop a new database management system with the following objectives:

  • Integrate all of the departments’ practices into one system.
  • Use the cloud to store, manage, analyze, and maintain the data.
  • Hire a data scientist, computer scientist, and information architect to work as a team with the database designer and the database administrator.
  • Apply the extraction, transformation, and loading process to the data.
  • Use business intelligence and machine learning tools to make decisions based on the data.
Citation/Attribution
Reuse and redistribution of this content in digital or print format:
  • This book may not be used in the training of large language models or otherwise be ingested into large language models or generative AI offerings without OpenStax's prior written permission.
  • This book uses the Creative Commons Attribution-NonCommercial-ShareAlike License, which means that you can reuse and modify the material only for noncommercial purposes, must attribute OpenStax, and must distribute any derivative works under the same license.
  • Any commercial printing of this textbook, including using a local or custom printer, must be approved by OpenStax, and proper citation provided.
  • OpenStax-copyrighted images, activities, assessments, and similar components of this book are subject to the same licensing – CC-BY-NC-SA. They can be used for noncommercial purposes with attribution. Commercial use requires permission.
  • Permission requests: Anyone who intends to incorporate this content (including text, images, and other components) into large language models, use it in AI offerings, use it commercially (including in print), and/or has questions about another use case is welcome to complete our reuse request form.
Attribution information
  • If you are redistributing all or part of this book in a noncommercial print format, then you must include on every physical page the following attribution:

    Access for free at https://openstax.org/books/introduction-computer-science/pages/1-introduction

  • If you are redistributing all or part of this book in a noncommercial digital format, then for every page that includes OpenStax content, you must license the derivative work under the same CC-BY-NC-SA license as the original, and include on every digital page view the following attribution:

    Access for free at https://openstax.org/books/introduction-computer-science/pages/1-introduction

Citation information

The information below includes the information needed to generate citations in most major styles (APA, MLA, etc.); you must reformat and organize the information as needed to fit the requirements of the style. Use the information below to generate a citation. We recommend using a citation tool such as this one.

© Apr 23, 2026 OpenStax. Textbook content produced by OpenStax is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike License. The OpenStax name, OpenStax logo, OpenStax book covers, OpenStax CNX name, and OpenStax CNX logo, and Rice University name, and Rice University logo trademarks, or wordmarks are not subject to the Creative Commons license and may not be reproduced without the prior and express written consent of Rice University.