Open Bug 2046605 Opened 2 months ago Updated 27 days ago

[Monitor Agent][Models] Build evals for the monitor agent

Categories

(Core :: Machine Learning: General, task)

task

Tracking

()

People

(Reporter: tetchart, Unassigned)

References

(Blocks 1 open bug)

Details

(Whiteboard: [aiact])

We need to make sure that the monitor is working well and need evaluations built around that. It will also be the first task to help determine whether we can build on standard_chat evaluations from assistant or if we need a new spec for agents

Blocks: 2054541
You need to log in before you can comment on or make changes to this bug.