Invented trial records with assignment method, pre/post scores, sample sizes, dropouts, and an overstated draft conclusion.
What your agent must check
Keep a fully documented fixture descriptive without certifying causation.
Expose a post-only mistake, self-selection, an invented sample size, and a causal overclaim.
Retain an unreported dropout count as an unresolved method question.
Keep the scope clear
Label every trial as invented. No claims about real app effectiveness, real grades, diagnoses, judgments of student ability, or universal product recommendations.
Research behind this project path
Original explanations and authored practice records draw on these research and engineering ideas. The source organizations do not endorse this course or supply its fictional results.
NIH · Understanding Clinical Studies
Observational associations and randomized intervention designs support different kinds of conclusions.
This explains study design; it does not provide evidence for the fictional study cards in this lab. Direct page access was blocked during review; its primary indexed text was available.
Anthropic · Demystifying evals for AI agents
Define tasks, trials, and graders; inspect both execution records and final outcomes; repeat trials when model behavior varies.
A score depends on its cases and grading rules. Repeating a deterministic classroom case does not measure the variability of a live model.
Your JavaScript really runs. The model decisions and school data are authored simulations, so you can learn without an API key. Every workspace also includes a separate real SDK example to explore next. Passing the lab’s cases is practice, not proof that an agent is ready for real-world use.