Skip to ContentGo to accessibility page

Labs

1 .
Search online to learn what a virtual machine is. You are setting up a virtual machine (VM) on Microsoft Azure and would like to perform data science experiments. Research the best way to gain access to all the tooling you need without having to research and install the individual tools on your own.
2 .
Select three examples of commercial or open-source DBMSs that use different data models. Install the trial versions of each one of these DBMSs and illustrate their use via a simple tutorial example. Document your work and evaluate the benefits and drawbacks of each system based on your experience.
3 .
Explore MySQL and experiment with MySQL Workbench to build a simple website using Django. Refer to the instructions and tutorial for more information.
4 .
Build a simple Django application that implements a social media website and uses a cloud-based data management service for data management. (Hint: You can use this article from Medium that contains some guidance.)
5 .
Explore how to use AWS service areas when solutioning use cases for a data lake. Data are stored in a raw state initially, and some use cases will use raw data as is. More often, solutions require varying degrees of data preparedness based on a collection of query usage profiles that correlate to actual use cases. Based on the solution, data may be refined and staged with the intent to promote modularity and reuse. The goal is to not overprocess the dataset because it is intended for multiple purposes downstream, such as AWS RedShift for relational analytics, AWS Elasticsearch for text search, or an optimized distributed file system for low-cost active archive storage, which can be queried with an MPP SQL engine.
6 .
Investigate how to put together an end-to-end data management infrastructure for a recommender application being built by a start-up. The application is expected to collect hundreds of gigabytes of both structured (customer profiles, temperatures, prices, and transaction records) and unstructured (customers’ posts/comments and image files) data from users daily. Predictive models will need to be retrained with new data weekly and make recommendations instantaneously on demand. Data collection, storage, and analytics capacity would have to be extremely scalable. The questions at hand are: How can you design a scalable data science process and productionize the models? What are the tools needed to get the job done? You will need to explain how to set up a data pipeline,
7 .
Leverage the types of choices suggested in the associated diagram, decide between on-premises and cloud services, choose a cloud service provider if applicable (in particular, investigate the cloud service provider’s ML/DL capabilities and build your solution to avoid cloud vendor lock-in), and develop robust cloud management practices.
8 .
Search the Internet for available informatics platforms and experiment with any of the ones you find.
Citation/Attribution
Reuse and redistribution of this content in digital or print format:
  • This book may not be used in the training of large language models or otherwise be ingested into large language models or generative AI offerings without OpenStax's prior written permission.
  • This book uses the Creative Commons Attribution-NonCommercial-ShareAlike License, which means that you can reuse and modify the material only for noncommercial purposes, must attribute OpenStax, and must distribute any derivative works under the same license.
  • Any commercial printing of this textbook, including using a local or custom printer, must be approved by OpenStax, and proper citation provided.
  • OpenStax-copyrighted images, activities, assessments, and similar components of this book are subject to the same licensing – CC-BY-NC-SA. They can be used for noncommercial purposes with attribution. Commercial use requires permission.
  • Permission requests: Anyone who intends to incorporate this content (including text, images, and other components) into large language models, use it in AI offerings, use it commercially (including in print), and/or has questions about another use case is welcome to complete our reuse request form.
Attribution information
  • If you are redistributing all or part of this book in a noncommercial print format, then you must include on every physical page the following attribution:

    Access for free at https://openstax.org/books/introduction-computer-science/pages/1-introduction

  • If you are redistributing all or part of this book in a noncommercial digital format, then for every page that includes OpenStax content, you must license the derivative work under the same CC-BY-NC-SA license as the original, and include on every digital page view the following attribution:

    Access for free at https://openstax.org/books/introduction-computer-science/pages/1-introduction

Citation information

The information below includes the information needed to generate citations in most major styles (APA, MLA, etc.); you must reformat and organize the information as needed to fit the requirements of the style. Use the information below to generate a citation. We recommend using a citation tool such as this one.

© Apr 23, 2026 OpenStax. Textbook content produced by OpenStax is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike License. The OpenStax name, OpenStax logo, OpenStax book covers, OpenStax CNX name, and OpenStax CNX logo, and Rice University name, and Rice University logo trademarks, or wordmarks are not subject to the Creative Commons license and may not be reproduced without the prior and express written consent of Rice University.