# AI Quality Tasks

Module scope: the evaluation harness for AI components across every other module — test datasets, decision-quality scoring, hallucination/tool-use checks, regression tests, and prompt versioning. See `docs/TASKS.md` for the full-program index.

| Tasks | Hours |
|---|---|
| 1 | 36 |

---

## Id: 31
- **Phase:** AI Quality
- **Priority:** P1
- **Area:** AI Evaluation
- **Task:** Test datasets, decision-quality evaluation, hallucination/tool-use checks, regression tests and prompt versioning
- **Description:** Build the evaluation harness for the AI components — curated test datasets, scoring for decision quality, checks for hallucinated data or incorrect tool use, regression tests to catch prompt/model changes that degrade behavior, and version tracking for prompts over time.
- **Primary Role/Assignee:** Unassigned
- **Estimated Hours:** 36

---

## Current implementation status

Not started — only ordinary PHPUnit fakes (`Image::fake()`, agent fakes) exist for the creative-generation jobs, not a dedicated AI evaluation harness.
