OpenStax Research
Studying how students learn in real learning environments
For more than a decade, OpenStax has conducted research in learning science, OER efficacy, and AI and machine learning for education. We work with students and researchers at Rice, connect with a national research community through SafeInsights, and partner with school districts and university faculty.
Learn with us https://openstax.org/blog?collection=Education+research orange 70% 0 0 center 0 2 #f7fcff bright blue divider center 100% 3pxHow we work
Our interdisciplinary approach to learning research
OpenStax research draws on learning science, cognitive psychology, and AI and machine learning. What we find informs how we design materials and how we measure whether learning is actually improving, particularly for students farthest from opportunity.
80% 0 0 center 3 5 #ffffff research icon center 100pxOpenStax Kinetic
A learning lab where students discover how they learn
Kinetic is OpenStax's research platform, where U.S. higher education students participate in short, voluntary studies alongside their coursework. Studies explore questions like what kinds of practice make learning stick and how motivation shapes performance. Students gain insight into their own learning and how to improve their study skills, and their participation shapes what OpenStax creates for the students who follow.
100% 0 0 kinetic decorative right 0 2 #F7FCFF #F5FEF8 to bottomResearch-practice partnerships
Research alongside educators and learners
Some questions can only be answered inside real classrooms. OpenStax partners with school districts and teachers to run studies where the teaching happens, from a design sprint with 11 Spring ISD high school students on how struggling algebra learners experience math, to a classroom impact study of EduStax, our AI lesson-planning tool, now in final data collection with results expected this year.
Findings from this work, including a randomized blind comparison of AI-generated and expert-created lessons, are in the publications below.
100% 0 0 research partnerships decorative left 0 18 #F5FEF8 #F5FFEC to bottomNational infrastructure for learning research
SafeInsights is OpenStax's sister organization at Rice University, a national education research hub backed by the largest investment NSF has made in education R&D infrastructure. It rests on one idea. Bring the research questions to the data. Researchers' analysis code runs inside secure enclaves where the data already lives, so studies reach population scale while individual student records never move.
SafeInsights is built by the field, for the field. Digital platform partners, including OpenStax, connect researchers to authentic learning data across K-12 and higher education. OpenStax-SafeInsights studies begin this year.
80% 3 #ffffff #00C1DE center 1 22 0 12 2 0 #ffffff research_siAI/ML RESEARCH
What we're working on now
This year, our researchers introduced the Learning Context, a framework for what an AI system needs to understand about a learner to genuinely personalize instruction, and presented new work at AIED 2026 and the Educational Data Mining conference in Seoul, from measuring where AI instruction falls short to generating privacy-preserving synthetic student data. The findings and the decade of work behind them are in the publications below.
80% 0 0 0 #FFFFFF center 0 #FFFFFF 2Publications
Explore the research
OpenStax and Rice University researchers publish peer-reviewed work on learning analytics, AI in education, and OER efficacy in venues including PLOS ONE, AAAI, EMNLP, ACM Learning at Scale, EDM, and AIED. Independent scholars study OpenStax materials and open educational resources in their own work. Both bodies of research are collected here. Select any study to read what it found, what it means, and the full citation.
center 70% 0 0 #F0FEF9 to bottomOpenStax and Rice research
AIED
publication_1For more than a decade, the Rice DSP lab and OpenStax research team have studied how AI can serve learning. The research here examines AI systems themselves: whether they can generate and classify educational content, whether they can be aligned to teach rather than answer, what they miss about the learner, and how they fail. The most recent work turns to the learner side of that question: what a model of the learner needs to contain, how to measure where AI instruction falls short, and what happens when models are trained on student data.
PublicationsProviding learner context measurably changes AI instructional decisions, but does not yet produce pedagogically appropriate personalization
This study presents a framework for measuring how Learning Context influences instructional strategy selection in LLM-based tutoring systems. Comparing context-blind and context-aware conditions against the judgments of subject matter experts, the researchers found that providing learner context moves AI instructional decisions measurably closer to expert judgment, but substantial misalignment remains. A relevance-impact analysis diagnoses which learner characteristics the models attend to, ignore, or weight spuriously.
What this means. Simply telling an AI about the learner is not enough. The instruction changes, but not reliably in the ways an expert educator would choose. Genuine personalization is still an open problem, and this work is a way to measure progress on it.
Learning Context Matters: Measuring and Diagnosing Personalization Gaps in LLM-Based Instructional Design — Hatchett, Basu Mallick, Bradford, and Baraniuk, 2026, AIED
Teachers rated AI-generated active learning lessons on par with an expert's, and could not reliably tell which was which
In this study, 26 teachers scored active learning tasks created by an LLM and by a human subject matter expert in a randomized, blind comparison, using the same rubric. Scores did not differ significantly on any category, teachers slightly favored the LLM lesson, and they could not consistently identify which was AI-generated.
What this means. Active learning works, but designing those lessons is time-consuming enough that many teachers cannot use them regularly. If AI-generated lessons hold up to expert quality, the time barrier starts to come down.
A Randomized Blind Comparison of SME and LLM-Generated Active Learning Tasks for Math Classes — Bainbridge, Strelich, Basu Mallick, and Baraniuk, 2026, AIED (Springer)
Synthetic student data can stay statistically reliable across regeneration cycles while protecting the students it describes
This paper introduces a Non-Parametric Gaussian Copula method for generating synthetic student data that preserves the statistical structure of the original, integrates differential privacy, and stays stable where deep learning generators degrade. It was validated through deployment in a real-world online learning platform.
What this means. A big blocker in education research is access: student data is locked down for good reasons. Synthetic stand-ins that behave like the real data mean more researchers can work on real problems without a single student record changing hands.
Stable and Privacy-Preserving Synthetic Educational Data with Empirical Marginals: A Copula-Based Approach — Díaz-Ramos, Luzi, Basu Mallick, and Baraniuk, 2026, EDM
A roadmap for AI education systems that adapt to the whole learner over time
The Learning Context framework (2025) proposes an approach for context-aware AI in education that integrates cognitive, affective, and sociocultural factors over time rather than responding only to a student's most recent input. The paper describes an intended implementation path through SafeInsights and the OpenStax platform.
What this means. Today, a student's learning history dies every time they close a chatbot. This roadmap proposes letting a student carry their progress into any AI tool they use. A proposal and a research direction, not a shipped product.
Learning Context: A Unified Framework and Roadmap for Context-Aware AI in Education — Liu, Bradford, Hatchett, Diaz, Luzi, Wang, Basu Mallick, and Baraniuk, 2025, arXiv preprint
Prompted with text they were likely trained on, some LLMs reproduce open textbooks word for word
This study developed a prompting method showing that some large language models reproduce open textbook passages verbatim when prompted with text from their likely training data, at rates far higher than for text published after their training cutoff.
What this means. When a student asks a chatbot for homework help, there is a decent chance the answer is textbook content with the serial numbers filed off. The knowledge made it into the models. The peer review, the context, and the credit did not.
Many-Shot Regurgitation Prompting — Sonkar, Liu, and Baraniuk, 2025, AIED (Springer LNCS 15881)
Training LLMs for personalized learning can produce regressive equity effects
This study found that fine-tuning large language models on student data to improve personalization creates unintended side effects, including reduced performance for students from underrepresented groups. The paper connects findings from foundational research on self-consuming AI systems to the specific context of educational AI.
What this means. Training AI on student data can make the tools worse, and worse specifically for students who most need them. The case for evaluating AI on every group it serves, not on average.
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning — Sonkar, Liu, and Baraniuk, 2024, EMNLP Findings
Large language models can be evaluated for counterfactual reasoning using a pedagogical approach
MalAlgoQA is a dataset designed to test whether LLMs can reason about why incorrect answers are wrong, not just identify correct ones. The approach uses mathematics education as the evaluation domain.
What this means. A good tutor does not just know the right answer. They can look at a wrong answer and see the reasoning that produced it. This research shows today's AI is much weaker at that second skill, which is the one tutoring actually depends on.
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education — Sonkar, Liu, Le, and Baraniuk, 2024, EMNLP Findings
Large language models can be aligned to follow pedagogical principles
This paper describes methods for prompting and training LLMs to follow effective pedagogical principles in tutoring dialogues, rather than defaulting to answer-provision.
What this means. Out of the box, a chatbot hands students the answer. This work shows that the habit can be trained out; the difference between an answer engine and a tutor is a training choice.
Pedagogical Alignment of Large Language Models — Sonkar, Ni, Chaudhary, and Baraniuk, 2024, EMNLP Findings
A design framework for intelligent tutoring systems grounded in learning science
The CLASS framework describes design principles for building AI tutors that guide students step-by-step through problems. A prototype called SPOCK was tested on college biology content and evaluated by biology experts for its ability to scaffold problem-solving without providing direct answers.
What this means. Before you can build an AI tutor that guides instead of answers, someone has to specify what guiding looks like, step by step. This framework did that.
CLASS: A Design Framework for Building Intelligent Tutoring Systems Based on Learning Science Principles — Sonkar, Liu, Basu Mallick, and Baraniuk, 2023, EMNLP Findings
Human-like question generation from educational content using large language models
This study investigated the impact of various prompting strategies on the quality of questions generated by large pretrained language models and identified strategies that produced the best outcomes for educational use.
What this means. Practice questions are the workhorse of learning, and writing good ones is slow, expert work. If experts cannot tell AI-written questions from human ones, question supply stops being the constraint on how much practice students get.
Towards Human-like Educational Question Generation with Large Language Models — Wang, Valdez, Basu Mallick, and Baraniuk, 2022, AIED (Springer)
Bloom’s taxonomy classification can be automated without labeled training data
This paper adapted large language models to classify educational content by Bloom's taxonomy level without requiring hand-labeled training examples, using OpenStax content as the development domain.
What this means. "Does this question test recall or reasoning?" is the kind of tagging that takes experts hours per chapter. Automating it means a library the size of OpenStax's can be organized by what each item asks of a student's thinking.
Towards Human-like Educational Question Generation with Large Language Models— Wang, Manning, Basu Mallick, and Baraniuk, 2022, AIED/Springer
Question mining at scale enables personalization across large learner populations
This study combined natural language processing with machine learning to predict question difficulty, analyze question content, and personalize question selection across large educational datasets.
What this means. Knowing which question to give which student, out of tens of thousands, is the quiet machinery behind any personalized practice product.
Question Mining At Scale: Prediction, Analysis and Personalization — Wang, Tschiatschek, Woodhead, Hernandez-Lobato, Jones, Baraniuk, and Zhang, 2021, AAAI
A hybrid AI and crowdsourcing approach can bootstrap ontology graphs from textbooks
This study adapted the BERT language model and a crowdsourced relationship selection task to build ontology graphs from OpenStax textbook content, enabling structured knowledge representations at scale.
What this means. A textbook knows its concepts connect; software does not, unless someone maps the connections. This work showed the map can be built mostly automatically, with people checking the joints.
A Case Study on Bootstrapping Ontology Graphs from Textbooks — Chaudhri, Boggess, Aung, Basu Mallick, Waters, and Baraniuk, 2021, AKBC
Math word problem generation can be automated while preserving mathematical consistency
This paper presents a method for generating new math word problems from existing ones while maintaining mathematical correctness and appropriate context, using OpenStax content as a development domain.
What this means. An AI can write a story problem about anything a student cares about. The hard part is keeping the math correct while the story changes. That is the difference between novelty and useful practice.
Math Word Problem Generation with Mathematical Consistency and Problem Context Constraints — Wang, Lan, and Baraniuk, 2021, EMNLP
Short-answer grading can be automated using transformer modelsShort-answer grading can be automated using transformer models
A meta-learning augmented bidirectional transformer model was applied to the automatic grading of short-answer responses to STEM exercises, achieving competitive performance with limited labeled data.
What this means. Multiple choice scales; written answers reveal thinking. Automated grading of written answers is a path to both.
A Meta-Learning Augmented Bidirectional Transformer Model for Automatic Short Answer Grading — Wang, Lan, Waters, Grimaldi, and Baraniuk, 2019, EDM
A data-driven model can generate educationally valid questions from text at scale
QG-Net is a deep learning model trained to generate questions from educational passages. It was developed and evaluated using OpenStax content and demonstrated competitive performance against human-authored questions.
What this means. An early proof that machines could write usable study questions straight from a textbook.
QG-net: a data-driven question generation model for educational content — Wang, Lan, Nie, Waters, Grimaldi, and Baraniuk, 2018, ACM Learning at Scale
Instructor content preferences can be modeled to improve recommendations
This paper presents a latent factor model for analyzing instructor preferences across educational content, enabling more targeted content recommendations and curation support.
What this means. Instructors, not algorithms, decide what students see. Modeling what instructors actually choose is how a platform recommends without overriding.
A Latent Factor Model For Instructor Content Preference Analysis — Wang, Lan, Grimaldi, and Baraniuk, 2017, EDM
Mining textual responses can uncover systematic misconception patterns
This paper applied natural language processing to open-ended student responses in STEM courses to identify recurring misconception patterns that would not be visible from graded responses alone.
What this means. A student can get the right answer for the wrong reason, and a grade will never show it. Reading what students write, at scale, surfaces the misunderstandings a score sheet hides.
Data-Mining Textual Responses to Uncover Misconception Patterns — Michalenko, Lan, Waters, Grimaldi, and Baraniuk, 2017, EDM
Response validity in open-ended STEM exercises affects measured learning outcomes
This study examined how the validity of student responses to short-answer STEM exercises affected estimates of learning gains, using OpenStax content as the evaluation domain.
What this means. If a study counts "idk" and blank guesses as real attempts, it will misjudge what students learned. Housekeeping, but the kind that decides whether education research can be trusted.
Short-Answer Responses to STEM Exercises: Measuring Response Validity and Its Impact on Learning — Waters, Grimaldi, Lan, and Baraniuk, 2017, EDM
on off 2 #F8FCFF,#ffffffLearning Sciences
publication_2The research here studies people: how students learn and persist, what shapes motivation, and how teachers work with new materials in real classrooms. It spans the early experiences that shape whether students stay in STEM, the psychosocial factors behind motivation and financial capability, and teacher practice with AI-generated materials. It also includes the foundational cognitive science, like retrieval practice, that informs how OpenStax designs everything else.
PublicationsMathematics teachers can use generative AI to create active learning experiences
Research with Candace Walkington at Southern Methodist University examines how mathematics teachers engage with generative AI to design active learning experiences, what quality conditions matter, and what teachers need to evaluate AI-generated materials well.
What this means. The AI adoption question in schools is not whether the tools work. It is whether teachers can tell when they do. This research studies the teachers and what they need to judge AI materials with confidence.
Mathematics Teachers' Use of Generative AI to Create Active Learning Experiences — Walkington and Bainbridge, 2025, Social Innovations Journal, 30(2)
AI can generate mathematical tasks for personalized learning when appropriately reframed
This study examined how generative AI can reframe mathematical tasks to support personalized learning in K-12 mathematics, investigating what reframing strategies produce pedagogically appropriate variations.
What this means. A word problem about a student's own interests lands better than one about widgets, but only if the math underneath stays rigorous. This research maps where personalization helps and where it starts to cost precision.
Using Generative AI to Reframe Mathematical Tasks for Personalized Learning — Beauchamp, Walkington, and Bainbridge, 2025, Ohio Journal of School Mathematics
Research infrastructure for studying learning at population scale, in authentic academic settings
The Kinetic platform enables large-sample, privacy-preserving studies on OpenStax's student population, supporting correlational, longitudinal, and interventional research designs that classroom-based studies cannot replicate.
What this means. Most learning studies run on a few dozen volunteers in a lab. Kinetic runs them with real students doing real coursework at a large scale, and it was the proof-of-concept for the national SafeInsights infrastructure.
Secure Education and Learning Research at Scale with OpenStax Kinetic — Basu Mallick, Bradford, and Baraniuk, 2023, ACM Learning at Scale
Early academic experiences in science and math shape whether students stay in STEM
This work-in-progress examined childhood and early academic experiences through qualitative analysis and identified seven biographical themes among higher education students, building toward a model for understanding which early experiences predict STEM persistence.
What this means. By the time a student drops a STEM major, the decision has usually been decades in the making. Knowing which early experiences matter tells schools and families where support actually moves the needle.
Does STEM Success Start Young? Exploring Higher Ed Students' Early Academic Experiences in Science and Math at Scale — Bradford, 2023, ACM Learning at Scale
Financial literacy interventions for young adults require design grounded in psychosocial mechanisms
Through interviews with subject matter experts and students, this study identified the highest-priority financial literacy learning objectives for young adults and developed a brief online intervention deployed through Kinetic, exploring what psychosocial factors predict whether financial education actually changes behavior.
What this means. Money stress is a quiet reason students leave college. This work is testing whether a short, well-designed financial intervention can change behavior. Early stage; outcome findings ongoing.
Unlocking Financial Success: Empowering Higher Ed Students and Developing Financial Literacy Interventions at Scale — Bradford, Basu Mallick, and Baraniuk, 2023, ACM Learning at Scale
A mixed-methods needs assessment of U.S. young adults' financial needs
This study presents a mixed-methods investigation of financial literacy needs among young adults, combining survey data with qualitative interviews to establish the most critical knowledge and skill gaps prior to developing interventions.
What this means. Before building a financial literacy course, ask young adults what they actually do not know. This study did the asking.
Money Matters: A Mixed-Methods Needs Assessment of US Young Adults' Financial Needs — Bradford, Basu Mallick, and Baraniuk, 2023, preprint
Integrating cognitive science principles with technology improved learning outcomes in a STEM classroom
This study examined a classroom intervention that combined retrieval practice, spaced repetition, and interleaving with digital tools in an undergraduate STEM course, finding significant improvements in exam performance compared to a control group.
What this means. Retrieval practice and spacing are not lab curiosities. Redesign a real course around them and exam scores move. This study, co-authored by J.P. Slavinsky, now at OpenStax, is one of the cleanest classroom demonstrations.
Integrating Cognitive Science and Technology Improves Learning in a STEM Classroom — Butler, Marsh, Slavinsky, and Baraniuk, 2014, Educational Psychology Review
on 3 #F8FCFF,#ffffffInstructional Design
publication_3The research here builds the model of the learner that adaptive instruction depends on. Every paper in this group produces or improves a way of estimating what a student knows, starting with SPARFA, which estimates concept mastery from graded responses alone, and extending those methods to open-ended answers and to the scale OpenStax operates at. Papers that build AI teaching capabilities appear under AIED; the work here models the student those capabilities have to serve.
PublicationsOpen-ended knowledge tracing extends learning analytics to free-text responses
This paper extends knowledge tracing methods beyond multiple-choice formats to open-ended responses, using natural language processing to model student knowledge from text answers.
What this means. Multiple choice tells you what a student picked. A written answer tells you how they think. Extending learner models to written work means adaptive systems can finally learn from the richer signal.
GPT-based Open-Ended Knowledge Tracing— Liu, Wang, Baraniuk, and Lan, 2022, EMNLP
A variational factor analysis framework improves the efficiency of Bayesian learning analytics
This paper presents a variational inference approach to learning analytics that achieves the accuracy of full Bayesian methods at reduced computational cost, enabling more scalable deployment.
What this means. An analysis that works for one classroom often collapses at a million students. This work is the unglamorous engineering that lets learner models run at the scale OpenStax actually serves.
A Variational Factor Analysis Framework for Efficient Bayesian Learning Analytics — Wang, Gu, Lan, and Baraniuk, 2020, arXiv preprint
Boolean logic analysis enables graded response data to reveal latent knowledge structures
BLAh is a framework for applying Boolean logic analysis to graded student response data, enabling finer-grained mapping of what students know beyond standard item response models.
What this means. The difference between "this student is struggling" and "this student is missing this specific idea" is the difference between more homework and the right homework.
BLAh: Boolean Logic Analysis for Graded Student Response Data — Lan, Waters, Studer, and Baraniuk, 2017, IEEE Journal of Selected Topics in Signal Processing
A machine-learning framework for jointly estimating student knowledge, question difficulty, and concept relationships from graded responses
SPARFA (Sparse Factor Analysis) jointly estimates learner concept mastery, question difficulty, and the relationships between questions and underlying concepts, using only graded response data. It was the analytical foundation for personalization work in earlier OpenStax products.
What this means. Adaptive learning products promise to know what a student knows. This is the math that makes such a claim honest, published openly rather than kept as a trade secret. Foundational research, not a current product.
Sparse Factor Analysis for Learning and Content Analytics — Lan, Waters, Studer, and Baraniuk, 2014, JMLR
Tracking how a student's knowledge changes as they learn, not just where it stands
This paper extends SPARFA from a single snapshot to a time-varying model, estimating how a learner's concept mastery evolves across a course from the sequence of their graded responses.
What this means. A snapshot tells you where a student is. A moving picture tells you whether the teaching is working. This extension is what makes learner models useful across a semester instead of at one moment.
Time-Varying Learning and Content Analytics via Sparse Factor Analysis — Lan, Studer, and Baraniuk, 2014, ACM Learning at Scale
on 5 #F8FCFF,#ffffffEfficacy Research
publication_4The research here measures outcomes: whether free access to course materials changes what students achieve. OpenStax research contributed a methodological insight the field now uses. The benefit of open resources concentrates among students who would otherwise go without materials entirely, so studies that miss that population miss the effect. Independent efficacy studies of OpenStax and OER appear under Independent Research below.
PublicationsOpen educational resources primarily benefit students who would otherwise go without course materials
This PLOS ONE study contributed a methodological insight to the OER field: the benefit of open resources concentrates among students who would otherwise go without course materials entirely. Standard research designs that study populations where most students would have bought the textbook anyway will underestimate or miss the effect.
What this means. For some students the choice was never free book versus paid book. It was book versus no book. Studies that miss those students miss the effect entirely.
Do Open Educational Resources Improve Student Learning? Implications of the Access Hypothesis — Grimaldi, Basu Mallick, Waters, and Baraniuk, 2019, PLOS ONE
on 2 #F8FCFF,#ffffffIndependent research
OER outcomes and adoption
publication_5More than a decade of independent studies on what happens when courses switch to free, open materials: student outcomes, completion, faculty perceptions, and adoption across U.S. higher education.
PublicationsOpenStax has become a mainstream choice in higher education
Bay View Analytics has tracked OER awareness and adoption in U.S. higher education since 2009. By 2022, OpenStax had become a viable alternative to commercial publishers. By 2024-25, 33% of faculty required OER as course materials, up from 5% in 2015-16. Large-enrollment introductory courses adopted OpenStax at twice the rate of OER generally.
What this means. Choosing OpenStax is no longer a pioneering decision. The quality concerns that once gave faculty pause have largely not held up.
OER in U.S. Higher Education: Annual Survey Series — Bay View Analytics, Seaman and Seaman, ongoing
Students in OER course pathways complete more credits, faster
Across 11 community colleges, enrollment in OER courses was associated with greater credit accumulation over multiple terms. Students in OER pathways made faster progress toward their degrees.
What this means. Free materials may accelerate the path to completion, particularly at institutions where cost is a persistent barrier.
Encouraging Impacts of an OER Degree Initiative on College Students' Progress to Degree — Griffiths et al., 2022, Higher Education
Students learn as well or better with OER, and most would choose it again
A synthesis of 16 efficacy and 20 perceptions studies covering 121,168 students and faculty found students using OER achieved the same or better learning outcomes as students using commercial textbooks, while spending significantly less money.
What this means for practice. The quality concern about free materials has been tested across the literature. The evidence now spans over a hundred thousand students.
Open Educational Resources, Student Efficacy, and User Perceptions — Hilton, 2020, Educational Technology Research and Development
Who is represented in a textbook affects whether students feel they belong in a course
Crowd-sourced revisions to the OpenStax Psychology textbook enhanced belonging for first-generation students. OpenStax's open license makes genuine revision possible, and that revision produced measurable effects.
What this means. Representation is a design variable with outcomes attached to it.
Who Gets to Wield Academic Mjolnir? — Nusbaum, 2020, Journal of Interactive Media in Education
The students who benefit most from open materials are the ones who needed them most
A study of 21,822 students found OER adoption improved course grades and reduced withdrawal rates. The largest gains went to Pell-eligible students, part-time students, and students historically underserved by higher education.
What this means for practice. The effect concentrates where access was the barrier.
The Impact of Open Educational Resources on Various Student Success Metrics — Mann, 2018
Adopting OpenStax often prompts faculty to redesign their courses
In a biology course of 1,299 students, 64% rated OpenStax quality as comparable to other textbooks and 22% as higher. Faculty reported clearer learning outcomes and improved course structure following adoption.
What this means. The free textbook functions as a catalyst. The act of switching tends to open up other changes in how faculty teach.
Student and Faculty Perceptions of OpenStax in High Enrollment Courses — Watson, Domizi, and Clouser, 2017, International Review of Research in Open and Distributed Learning
Faculty who adopt OpenStax shift how they teach, not just what they assign
Among 137 faculty using OpenStax, 62% rated quality as comparable to traditional textbooks and 19% as better. Most reported pedagogical shifts in course structure, student engagement, and how they used class time.
What this means. OpenStax adoption tends to be a teaching decision as much as a materials decision. The open license appears to encourage adaptation and experimentation.
Higher Education Faculty Perceptions of Open Textbook Adoption — Jung, Bauer, and Heaps, 2017, International Review of Research in Open and Distributed Learning
Educators who switch to OpenStax tend to rethink how they teach
Educator surveys found quality perceptions equal to or higher than commercial alternatives, alongside meaningful shifts in teaching practice following adoption.
What this means for practice. The transition goes beyond cost. Teachers often come out the other side with a course they feel better about.
Mainstreaming Open Textbooks: Educator Perspectives on the Impact of OpenStax College Open Textbooks — Pitt, 2015, International Review of Research in Open and Distributed Learning
on 2 #F8FCFF,#ffffffAI built on OpenStax content
publication_6Independent research using OpenStax's open library as the content foundation for AI learning tools.
PublicationsPersonalized AI built on OpenStax content produces meaningfully better retention than a static digital reader
Google built Learn Your Way on OpenStax chapters, using generative AI to create personalized, multimodal learning experiences. An independent randomized controlled trial found students scored 9 percentage points higher on immediate recall and 11 points higher on retention three to five days later (78% vs. 67%).
What this means for practice. How content is delivered matters as much as the content itself. The free library is what makes this kind of research possible.
Towards an AI-Augmented Textbook — LearnLM Team, Google, 2025 · Independent RCT, 2025
on 2 #F8FCFF,#ffffffFoundational learning science
publication_7Foundational studies from the wider field, from active learning meta-analyses to retrieval practice, that inform how OpenStax designs its materials and research.
PublicationsActive learning closes achievement gaps; the gains are largest for the students furthest behind
Across 26 studies covering 44,606 students, high-intensity active learning reduced exam score gaps between underrepresented and well-represented students by 33% and narrowed passing rate gaps by 45%.
What this means. How a course is taught matters as much as what materials are used. Active learning's equity effects are as important as its average effects.
Active Learning Narrows Achievement Gaps for Underrepresented Students in Undergraduate STEM — Theobald et al., 2020, PNAS
Replacing lecture with active practice makes students 1.5 times less likely to fail
A meta-analysis of 225 STEM studies found students in active learning courses scored 0.47 standard deviations higher on exams than students in traditional lecture. The effect held across class sizes and disciplines.
What this means. Lecture is not a neutral default. Every hour of passive instruction is a choice with a measurable cost.
Active Learning Increases Student Performance in Science, Engineering, and Mathematics — Freeman et al., 2014, PNAS
Locus of control in academic contexts can be measured reliably in college students
This paper presents a revised version of the Academic Locus of Control scale validated for college student populations, providing a reliable instrument for research on student motivation and self-efficacy.
What this means. You cannot study motivation without a reliable way to measure it. This is one of the measuring sticks Kinetic studies use.
A Revision of the Academic Locus of Control Scale for College Students — Curtis and Trice, 2013, Perceptual and Motor Skills
Retrieval practice produces more durable learning than elaborative study strategies
A controlled experiment found that students who used retrieval practice to study retained significantly more material one week later than students who used concept mapping or re-reading, even when initial performance was equivalent.
What this means. The study habit most students trust, re-reading, loses to the one most avoid, self-testing. Findings like this one are why OpenStax materials emphasize practice.
Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping — Karpicke and Blunt, 2011, Science
on 2 #F8FCFF,#ffffff 70% 4 0 #ffffff publications_list 2 #FFFFFF #F5FFEC research_publicationsOur researchers
Who does this work
100% 0 0 0 2 0 #FFFFFFCore team
OpenStax research is led by Richard Baraniuk, founder and director of OpenStax and C. Sidney Burrus Professor of Electrical and Computer Engineering at Rice University, who serves as principal investigator on many of the major studies and directs the Rice DSP lab where foundational work on learning analytics and AI in education was developed. Debshila Basu Mallick serves as Director of Research and Scientific Director of SafeInsights. Katie Bainbridge leads K-12 research and Brittany Bradford leads higher education research.
The AI and machine learning research team includes Zichao (Jack) Wang, Naiming (Lucy) Liu, Hossein Babaei, and Shashank Sonkar. Zihan Liu serves as a postdoctoral research associate. Additional team members include Karyssa Courey (psychometrics), Makai Ruffin (adult skills and knowledge research), and a broader product, engineering, and knowledge translation team.
left 100% 0 0Collaborating researchers
OpenStax research is conducted in collaboration with scholars across disciplines and institutions. Current collaborators include Fred Oswald, Professor of Psychology at UC-Irvine; Margaret Beier, Professor of Psychological Sciences at Rice University; Phil Kortum, Associate Professor of Psychological Sciences at Rice University; Danielle McNamara, Professor of Psychology at Arizona State University; Andrew Lan, Assistant Professor at the University of Massachusetts Amherst; and Candace Walkington, Associate Professor of Mathematics Education at Southern Methodist University.
Alumni
Many researchers who contributed to OpenStax have gone on to roles across industry and academia, including Phillip Grimaldi (Senior Efficacy and Research Scientist, Khan Academy), Andrew Waters (Director of Data Science, Pinnacle), and others. Their work remains part of the published record.
100% 0 0 flex #FFFFFF 4 0 0 research_peopleRice Digital Signal Processing lab
0 100% 0 2 0 #ffffffGet involved
Study learning at scale, with us
OpenStax welcomes researchers who want to study learning at scale. Kinetic connects vetted researchers with U.S. higher education students in authentic academic environments, supporting correlational, longitudinal, and interventional studies within a privacy-preserving infrastructure. A Library of Learner Characteristics, validated measures of sociodemographic, cognitive, motivational, and psychosocial factors, is available as focal or control variables. Researchers can also join the SafeInsights research beta or collaborate directly with our team.
Design a study https://riceuniversity.co1.qualtrics.com/jfe/form/SV_6EbRsmpDb2Hs69w orange 0 0 60% center 2 #ffffff bright blue divider 6px -3px 100% center