Advanced Data Science:
a graduate course built from open educational resources
Case study · AI Tools & Strategy · Learning Design
A complete graduate data-science course. I curated vetted open educational resources for the readings and interactive models, and built the professional simulations that anchor every module myself, assessed by performance, not exams.

Goal
Advanced Data Science tests whether the synthetic-SME method can produce a rigorous graduate course in a field the designer does not teach, so the quality has to come from the design rather than from subject-matter command. It is a full 15-week, three-credit, MS-level course, designed backward from its outcomes (Wiggins & McTighe), combining vetted open educational resources I found for its readings and models with original simulations I built for every assessment, and with every activity and assessment aligned to an outcome (constructive alignment, Biggs).
Audience
The learners are graduate, MS-level data-science students. As a portfolio piece, the page is written for learning designers and academic leaders judging whether AI-assisted course production can hold to real quality standards.
Process
The course is designed backward, outcomes first, then the evidence, then instruction (Wiggins & McTighe; Biggs). A panel of AI agents drafted each module against a fixed checklist of online-quality standards; I directed the panel and own every decision.
Outcomes
Every module, reading, and assessment is designed back from seven observable outcomes. On completion a learner can:
- Acquire data from structured and unstructured sources (files, databases, APIs, the web), documenting provenance.
- Preprocess messy, dynamic data: cleaning, missing values, outliers, and feature engineering.
- Choose among regression, classification, and clustering for a given problem and its data.
- Evaluate models with resampling, appropriate metrics, and the bias-variance trade-off.
- Communicate findings to a non-technical decision-maker through sound visualization and a defensible narrative.
- Assess the ethical dimensions of a data problem with a recognized governance framework.
- Build a complete data-science project end to end and defend its decisions.
How success is measured
Success is a performance, not a test score. Each of the eight modules ends in a summative simulation of a real professional task, engineered so a generative model cannot complete it for the student. Core datasets are seeded per student from a data-generating process only the instructor knows, so there is no answer key on the web. Incidents (leakage, bias, missingness, a mid-project data-drift event) are planted, so the student has to detect, diagnose, and adapt. Each major deliverable is defended in a simulated stakeholder meeting where a role-played executive, compliance officer, and engineer ask unscripted questions, and the capstone is a three-week consulting engagement with a live defense. This is authentic, performance-based assessment (Wiggins): competence is inferred from what a student does with a real, private artifact, which preserves construct validity where a conventional exam no longer can. When the graded object is live reasoning about a private artifact, the durable answer to AI is to change what students are asked to do, not to police it.
Pacing and interaction
Every module follows the same rhythm, so the structure never becomes the obstacle (cognitive load theory, Sweller): a short recorded introduction, a vetted OER reading, a low-stakes interactive practice, the summative simulation, and a Regular and Substantive Interaction discussion. The course is sized to the credit-hour standard the Higher Learning Commission applies (34 CFR 600.2), 135 hours of student work across 15 weeks, about nine hours a week, with an estimated-time note on every task to support planning and self-regulated learning (Zimmerman). Instructor presence is designed in from a Getting Started module and a weekly, instructor-initiated interaction plan rather than bolted on, building the teaching presence a Community of Inquiry depends on (Garrison, Anderson & Archer).
The learning arc, week by week:
| Weeks | Module and focus | Summative simulation |
|---|---|---|
| 1–2 | M1 · Acquisition and the reproducible workflow | The Provenance Incident |
| 3–4 | M2 · Preprocessing messy and dynamic data | Recover the Signal |
| 5–6 | M3 · EDA, visualization and the stakeholder story | The 10-Minute Executive Brief |
| 7–8 | M4 · Supervised learning | Private-Leaderboard Competition |
| 9–10 | M5 · Rigorous evaluation | Ridgeline Renewables model audit |
| 11 | M6 · Unsupervised learning | How Many Segments, Really? |
| 12 | M7 · Data and AI ethics and governance | The Model-Governance Review Board |
| 13–15 | M8 · Capstone consulting engagement | The RiverCity Engagement + live defense |
Access
Universal Design for Learning (CAST) and WCAG 2.1 AA are built into the design rather than audited on afterward. All three UDL principles are addressed: multiple means of engagement, of representation (recorded introductions, readings, and interactive explorables), and of action and expression (the simulations and defended project let students demonstrate mastery in more than one way). Videos are captioned with slide alternatives, and time-on-task is signaled in words and a bold label, not styling alone (WCAG 1.3.1). The Getting Started module orients every learner to the online environment before content begins.
Technology
- Build and orchestration. A reusable Claude skill (the Gemini equivalent is a Gem) orchestrates an agent panel: a subject-matter-expert agent, an instructional-designer agent, and a synthetic student modeled on the incoming learner. Built with the synthetic-SME method (AI-assisted course design →).
- Design frameworks. Backward design (Wiggins & McTighe), constructive alignment (Biggs), and the revised Bloom’s taxonomy for outcome verbs.
- Quality and compliance standards, enforced on every pass. OSCQR, Regular and Substantive Interaction (34 CFR 600.2), Universal Design for Learning (CAST), and WCAG 2.1 AA.
- Open educational resources and explorables I found and license-verified. MLU-Explain, TensorFlow Playground, R2D3, Setosa, Distill, Seeing Theory, Facets, Data-to-Viz, the FT Visual Vocabulary, StatKey, Kaggle Data Cleaning, missingno, Select Star SQL, and Google’s PAIR fairness tools, Aequitas, and the What-If Tool.
- Simulations I built for this course. The eight professional simulations are my own, made with Synthetic Data Vault and Faker for per-student seeded datasets, Evidently for data-drift detection, and Docker for reproducible environments.
- Course-grounded AI tutor. A public Gemini Notebook that answers only from the course and its vetted sources, with citations.
- Language and delivery. Python throughout; the course itself delivered as static HTML and CSS.
- Instructor materials. Eight slide decks with speaker-note scripts and a full instructor preparation kit.
Status
Currently in review. A complete, standards-aligned draft. The expertise built into it is synthetic, which is exactly what the test is probing, so it is now with a data science subject-matter expert who is verifying the content. Not yet delivered to students.
The full modules, readings, schedule, and setup live in the course; the hand-off notes are in instructor prep. The summative simulations are mapped in their own tab and run in full in the Authentic Assessment section.
Adopt this course
Download the whole course as a Common Cartridge and import it into Canvas, Blackboard, Moodle, or D2L (in Canvas: Settings, then Import Course Content, then Common Cartridge). It brings in a Start Here page, the Getting Started module, and all eight modules as pages, with the four simulations and the defense tool linked, and an editable PowerPoint lecture deck bundled inside each module. Free to reuse and adapt for noncommercial purposes under the license below.
This course is an open educational resource by Michelle Blomberg, released under CC BY-NC-SA 4.0. Reuse, adapt, and share for noncommercial purposes with attribution, under the same license. To credit it, use: “Advanced Data Science by Michelle Blomberg, licensed under CC BY-NC-SA 4.0.”