UX Case Study · INTERNAL ENTERPRISE SUPPORT TOOL
Designing a Scalable, Human-in-the-Loop Configuration System (Create, Read, Update, and Delete) for AI-Generated Assessments
Team
1 PM, 2 Devs, 1 PD
/
Role
Lead PD (IC Lead)
/
Timeline
~30 Days
/
Platforms
Web
/
Audience
Internal Staff

Udacity's enterprise business relied on third-party assessments and manual, engineering or content-driven workflows to create and manage in-house assessments. As demand increased and AI-generated content became viable, that process stopped scaling.
I led design on an AI-enhanced assessment creation and management system inside Udacity's enterprise platform — giving internal, non-technical teams the ability to create, review, edit, delete, and maintain assessments using both manual and AI-generated questions, in 30 days.
Four issues were limiting Udacity's ability to scale assessments:
Engineering and Content dependency — assessments and questions were generated or manually created or configured by engineers and the content team, creating bottlenecks
Non-scalable workflows — non-technical teams (Content, Customer Success, Ops) had no tools to create or manage assessments independently
No quality control framework — no consistent way to evaluate question quality or performance
AI without governance — as AI-generated questions became viable, there was no structured human-review process to ensure they were accurate, relevant, and trustworthy before reaching a learner
All of this needed a working solution in weeks, not months, to meet enterprise contract commitments already in motion for Q3–Q4.
The project started as a narrower brief: design the AI-generated question flow. Early discovery changed that.
Auditing the existing system, I found the foundational assessment CRUD experience itself was incomplete — there wasn't yet a reliable way to create, edit, or manage assessments at all, AI-generated or not. Layering an AI generation flow on top of that would have meant building trust and governance on a foundation that couldn't support it.
I made the call to expand scope to include the foundational assessment CRUD system, not just the AI layer — even though it meant less time available for visual polish inside an aggressive 30-day window. The reasoning: an AI review workflow is only as trustworthy as the system it lives inside. Shipping a well-designed AI review flow bolted onto a broken underlying system would have solved the visible problem while leaving the real one in place.
This was the core of the project, so it's worth walking through the mechanism rather than describing it abstractly.
1. Guided generation. A content team member inputs specific criteria — skill area, difficulty, format — and references specific learning content the questions should be grounded in, rather than generating from an open-ended prompt. This keeps AI-generated questions tied to real Udacity curriculum instead of generic or ungrounded content.
2. Human review before anything goes live. Once questions are generated, the same person or another content team member reviews the output — each question and its multiple-choice answers — before it's approved. Reviewers can approve, edit, or reject each item, and can add notes explaining their decision. Nothing reaches a learner without passing through this step.
3. Visibility into the process itself. The system tracks approval rate as a metric, and reviewer notes create a running record of why content was approved, edited, or rejected — not just a pass/fail log. This was a deliberate design choice: it turns review from a one-time gate into a growing source of insight about where AI generation is strong and where it consistently needs human correction.
Designing this meant making decisions that don't show up in a screenshot — like what criteria inputs actually needed to be to keep generation results high quality and grounded, what a reviewer needs to see at the moment of decision to review efficiently without rubber-stamping, and how to surface approval-rate data usefully rather than as a vanity number no one acts on.
Rapid system understanding — audited the existing assessment experience, reviewed an engineering-built prototype for AI-generated questions, partnered closely with engineers to understand backend constraints and data models
Experience principles — AI should accelerate creation, not bypass human judgment; non-technical users should operate independently; CRUD workflows must support iteration, not just creation
AI-assisted ideation — used AI tools (including Lovable for rapid interactive prototyping) to explore workflow variations and stress-test edge cases in the review flow faster than static mockups would have allowed, while keeping design decisions deliberate rather than AI-generated wholesale
Assessment CRUD — create, edit, publish, and manage assessments without engineering support
Question management — add, edit, replace, and remove questions across manual and AI-generated content
AI-generated question flow — guided, criteria- and content-referenced generation with clear visual distinction between AI-generated and approved content
Human-in-the-loop review — approve/edit/reject workflow with reviewer notes and approval-rate tracking (see above)












Before: Assessments manually created by engineers, no scalable path to update or iterate, no visibility or control for non-technical teams.
After: Fully functional assessment and question CRUD, supporting both AI-generated and manual workflows, with non-technical teams operating independently.
Business impact:
Unblocked enterprise assessment contracts dependent on Q3–Q4 delivery
Reduced engineering load for day-to-day assessment creation and maintenance
Cut question generation time by roughly 10x — work that previously required engineering involvement could be generated by a content team member directly, in a fraction of the time
Established a governed foundation for AI-generated content that can extend to future assessment types
I left the company shortly after this shipped, so I can't speak to how the content team's review workload evolved as more AI-generated questions moved through the queue, or whether the review process itself needed further iteration at higher volume — that's the natural next thing I'd want to measure if I'd stayed on the project.
What I'd track next: review queue volume and turnaround time as generation scaled, approval/edit/reject rate over time (to see if AI output quality improved or plateaued), and reviewer time-per-question — to know whether the review step itself needed its own efficiency pass once generation was no longer the bottleneck.
This project required real trade-offs: speed over visual refinement, system integrity over shipping the narrower feature first, and governance over unchecked AI automation. The result wasn't a polished surface feature — it was a durable internal system built to keep working as AI generation capabilities mature, with a review mechanism designed to get smarter over time rather than stay a static gate.
More than anything, it's a case study in what "responsible AI integration" looks like in practice: not a disclaimer, but a specific, designed point where a human makes the final call — and a system that keeps a record of why.