A new social media platform is looking for ways to improve its content-delivery algorithm so that it appeals to a broader audience. Currently, the platform is being used primarily by college-aged students. Which of these datasets would be the most appropriate data to collect and analyze to improve the algorithm and reduce potential bias in the content-delivery model?
Dataset 1, consisting of all the past browsing data on the platform
Dataset 2, the results of an extensive survey of people from various backgrounds and demographic segments who may or may not be familiar with the new social media platform
Dataset 3, the results of an internal company questionnaire
The admissions team of a large university would like to conduct research on which factors contribute the most to student success, thereby improving their selection process for new students. The dataset that they plan to use has numerical features such as high school GPA and scores on standardized tests as well as categorical features such as name, address, email address, and whether they played sports or did other extracurricular activities. The analysis of the data will be done by a team that involves student workers. Which of these features should be anonymized before the student workers can get to work?
Applying universal design principles in data visualization and data source attribution helps ensure that people in which group are able to view and use the data?
There was a growing concern at the Environmental and Agriculture Department about the rising levels of CO₂ in the atmosphere. It was clear that human activities, especially the burning of fossil fuels, were the primary cause of this increase. The department assigned a team of data scientists to research the potential consequences of climate patterns in the United States. They studied the trends of CO₂ emissions over the years and noticed a significant uptick in recent times. The team simulated various scenarios using climate models and predicted that if the current trend continues, there will be disastrous consequences for the climate. The four graphs displayed in Figure 8.11 depict the identical dataset from four different researchers. Out of the four graphs, which one accurately represents the CO₂ data without any type of distortion or misrepresentation?
Figure 8.11 Source: Hannah Ritchie, Max Roser and Pablo Rosado (2020) - "CO₂ and Greenhouse Gas Emissions." Published online at OurWorldInData.org. Retrieved from: 'https://ourworldindata.org/co2-and-greenhouse-gas-emissions' [Online Resource]
Reuse and redistribution of this content in digital or print format:
This book may not be used in the training of large language models or otherwise be ingested into large language models or generative AI offerings without OpenStax's prior written permission.
This book uses the
Creative Commons Attribution-NonCommercial-ShareAlike License, which means that you can reuse and modify the material only for noncommercial purposes, must attribute OpenStax, and must distribute any derivative works under the same license.
Any commercial printing of this textbook, including using a local or custom printer, must be approved by OpenStax, and proper citation provided.
OpenStax-copyrighted images, activities, assessments, and similar components of this book are subject to the same licensing – CC-BY-NC-SA. They can be used for noncommercial purposes with attribution. Commercial use requires permission.
Permission requests: Anyone who intends to incorporate this content (including text, images, and other components) into large language models, use it in AI offerings, use it commercially (including in print), and/or has questions about another use case is welcome to complete our
reuse request form.
Attribution information
If you are redistributing all or part of this book in a noncommercial print format,
then you must include on every physical page the following attribution:
Access for free at https://openstax.org/books/principles-data-science/pages/1-introduction
If you are redistributing all or part of this book in a noncommercial digital format,
then for every page that includes OpenStax content, you must license the derivative work
under the same CC-BY-NC-SA license as the original, and include on every digital page
view the following attribution:
The information below includes the information needed to generate citations in most
major styles (APA, MLA, etc.); you must reformat and organize the information as needed
to fit the requirements of the style. Use the information below to generate a citation.
We recommend using a citation tool such as
this one.
Authors: Dr. Shaun V. Ault, Dr. Soohyun Nam Liao, Larry Musolino