Objective 4.2 · Domain 4 · 16% of the exam
Design evaluation datasets and test frameworks using mixed methodologies
Building evaluation datasets and test frameworks with mixed methods: exact-match and programmatic checks where the answer is checkable, model-graded evaluation where it is not, and human review where the stakes justify it. Dataset design matters as much as grading here. An eval set drawn only from happy-path traffic will report a system as healthy right up until it is not.
5 questions · answers and explanations shown as you go · free, no sign-up
The rest of domain 4: Evaluation, Testing & Optimization
This domain is 16% of a scored form, about 10 of the 63 questions, spread across 6 objectives. Drill the whole domain or pick another objective below.
- 4.1Define evaluation metrics (accuracy, latency, cost, safety, security)(6)
- 4.3Conduct A/B testing and iterative improvements(5)
- 4.4Diagnose system issues (prompt failure, hallucinations, model mismatch)(5)
- 4.5Optimize token usage, latency, and cost-performance trade-offs(5)
- 4.6Monitor system performance using logging and observability tools(6)
One objective is a narrow slice. A full-length timed mock is what tells you whether the whole thing holds together.
Sit a timed mock →