Michelle Blomberg

Student Journey Gap Analysis

Case study · AI Tools & Strategy · UX Design

Fifty synthetic students, each built from real data on who the district enrolls and what stands in their way, try a real task on every college’s live website and report where they get stuck. People confirm every finding before it counts.

The study builds a fixed set of 50 personas, and each one is a whole student, never a single trait on its own. Race and ethnicity, first-generation status, food security, disability, language, academic preparedness, and enrollment path all live together in one person, in the proportions the district actually enrolls, so the barriers the set surfaces are the barriers real students hit, weighted by how common each kind of student is. Every number behind that mix is sourced at the bottom of this page.

How each persona is built from the data

  1. Start from real numbers. Gather what is known about students in the district and the barriers they face, from district figures and national research, with every figure cited in the sources at the bottom of this page.
  2. Set the mix. Turn those rates into targets for the set of 50: about 44% food-insecure, 44% carrying a disability, 38% who would place into a developmental course, 16% LGBTQ+, in the district’s real proportions.
  3. Build whole students. Distribute the attributes across 50 people so each persona is one coherent student, for example a first-generation, food-insecure nursing hopeful with undiagnosed ADHD, not a trait in isolation.
  4. Write the limits, not just the labels. Fill each persona into one fixed template: what she does not know, her device, her reading speed, her language, and the exact moment she would give up, so the agent behaves like her, not like a capable AI.
  5. Set it loose on the site. Each persona becomes an AI agent that browses a college’s live public pages from a felt need, thinks aloud, and records where it breaks. How a run works and how people validate it is spelled out below.

Above the persona agents sits one orchestrator agent that sequences the journey, from finding the college through to work, and routes each run to the right persona at the right college, so the whole matrix stays organized and comparable. The personas stay whole; the orchestrator coordinates them, it does not flatten them.

Each persona agent carries one whole persona at a time: one persona, one college, one task that starts from a felt need, never from an office, because a student is never told where to go. Every task is run at every college by enough different personas to reach saturation, the point where new personas stop surfacing new barriers, a reach only synthetic agents make practical. How the runs are scored, validated against real people, and kept clear of real student data is covered below and on the Ethics page.

Modality is written into each persona (all online, some online, or all in person), because a fully online student cannot reach a service offered only on campus, and that unreachability is itself a finding, not an error.

The students behind the personas

Every persona is built from what we know about students in the district: real district figures, what national research shows about the online population, and the attributes the study wrote in on purpose. The charts below hold all of it, the district totals first, then the 50-persona sample the study actually built.

Race and ethnicity, district totals only


Study modality, the 50-persona sample



What each persona also carries, counted across the 50

District totals from the district published data (Fast Facts and the college Student Right-To-Know pages); the study modality split follows a national public two-year baseline pending the district Institutional Research. District totals only, never per-college.

Neurotype, disability, and disclosure

Disability is written into the set the way it shows up in real life, unevenly and mostly out of sight. Twenty-two of the fifty carry a condition or are still working out whether they have one, but only six are known to their college. The rest are diagnosed and private, or never formally identified. That mirrors the national pattern: about 37% of students with a disability disclose it, and roughly half are not diagnosed until college. First-generation and low-income students are the least likely to have ever been evaluated, so their conditions are written as suspected, not registered. Veterans carry PTSD, not ADHD by default; the one homeschooled persona carries suspected autism; and foster, refugee, and DACA students carry trauma-linked anxiety. The undisclosed majority is the point: a student who never finds Disability Resource Services never gets the accommodation, and that gap is itself a barrier the study looks for.

Disability and disclosure across all 50 personas


Of the 22 who carry something, by type

This is a prototype and fieldwork has not started, so these are attributes written into the personas, not measured outcomes. Shares are anchored to national data (about 20% of undergraduates report a disability, with ADHD, autism, learning, and mental-health conditions among them) and tuned to a community-college, Hispanic-Serving, first-generation, low-income population where under-identification runs high. the district does not track a disability a student never discloses, and does not track whether a student was homeschooled, so the undisclosed and homeschooled shares here are modeled from national figures, not district counts.

The sample by age and language

Age, the 50-persona sample


Home language, the 50-persona sample

Academic preparedness and placement

’s K-12 system is among the weakest in the country, so many students arrive underprepared, and the personas carry that. On the 2024 Nation’s Report Card only about 26% of fourth graders and 25% of eighth graders were proficient in reading, meaning roughly three in four scored below proficient, and the state sits below the national average. The high-school graduation rate is about 78%, but only around half of graduates enroll in college within a year, and a student can leave an high school without solid reading or math skills. That is not a knock on the students, it is the reality the services have to meet.

In the district, a course numbered below 100 is developmental and does not count toward a degree: real courses like RDG091 (College Reading Skills), ENG091, and MAT091 (Introductory Algebra). Placement runs on multiple measures, EdReady, the Placement Coach, and prior scores, much of it self-directed. Nationally about 40% of community-college students take at least one developmental course, and the sequence is a known attrition point: only about 20% of students referred to developmental math and 37% referred to developmental reading pass the college-level gatekeeper course within three years. So the study writes preparedness in, and it writes in the specific barrier it creates: a student who self-advises or skips placement can land in a college-level course they are not ready for, or enroll in a below-100 course without realizing it will not count toward the degree.

Academic preparedness across the 50 personas

Of the 19, six self-advise or skip placement and risk landing in the wrong course, the placement-and-advising barrier the study tests directly. Share anchored to the national developmental-education rate (about 40%) and tuned to ’s below-average K-12 outcomes; assigned by each student’s story (low literacy, long out of school, first-generation with no one to explain placement), kept distinct from disability and from language alone.

How the test runs and is validated

How one run works, end to end

  1. Orchestrate. One orchestrator agent sequences the journey, from finding the college through to work, and routes each run to the right persona at the right college.
  2. Walk. A persona agent walks one real task on the live public site, constrained to that student’s knowledge and habits, thinking aloud, and stops where a real student would give up.
  3. Return a candidate. It returns a structured finding: outcome, the path and clicks, where it broke, and a severity guess from 0 to 4.
  4. Run to saturation. Each task repeats at every college with enough personas that new ones stop surfacing new barriers, the saturation stopping rule, rather than a fixed quota.
  5. Validate with people. A human sample re-runs a share of the tasks and confirms severity (Cohen’s kappa on the agreement). A finding counts only when a person validates it, because agents are least reliable at judging severity.

The method pairs expert heuristic evaluation (trained human evaluators check each college’s pages against Nielsen’s usability heuristics) with task-based usability testing (a persona walks a real task from a felt need, thinking aloud), and findability adds first-click testing and, for human runs only, the Single Ease Question, since an ease self-report needs a person’s felt experience. The walk runs in three parts by access: Part 1, the public tasks that need no login, runs now; Part 2, the signed-in tasks, runs on one sanctioned OIT test account once Domain 1 clears it; Part 3, the Salesforce-dependent tasks, waits for that platform to launch. Between parts the method is audited on the safest data before it escalates.

Synthetic agents carry the reach; people are the check. A small human sample re-runs some of the same tasks as a spot-check, about one tester per campus, compared on whether a barrier is present and a kappa on severity, so synthetic runs scale only if that agreement holds, and if it is weak the study leans on the human sample. Separately, 80 in-person office walk-ins, one per office, are single-visit probes rather than a representative sample. This follows current research: AI agents can run structured usability testing and catch most known issues but are least reliable at judging severity, so trained human raters set it, and every finding stays a candidate until a human validates it.

The evidence behind the personas

These fifty are not blind statistics. Every attribute written into them, disability, food insecurity, language, academic preparedness, enrollment path, traces to real data on who the district enrolls and what gets in their way, then to each student’s own story. The study is trying to account for everything that could stand between a student and their success, so the barriers it surfaces are the ones real students hit. The sources are here in the open.

This is a prototype and fieldwork has not started, so the persona attributes are grounded estimates, not measured outcomes. Where the district does not track something, disability a student never discloses, whether a student was homeschooled, the share is modeled from national data and labeled as such.