Open
Bug 2046605
Opened 2 months ago
Updated 27 days ago
[Monitor Agent][Models] Build evals for the monitor agent
Categories
(Core :: Machine Learning: General, task)
Core
Machine Learning: General
Tracking
()
NEW
People
(Reporter: tetchart, Unassigned)
References
(Blocks 1 open bug)
Details
(Whiteboard: [aiact])
We need to make sure that the monitor is working well and need evaluations built around that. It will also be the first task to help determine whether we can build on standard_chat evaluations from assistant or if we need a new spec for agents
Updated•2 months ago
|
You need to log in
before you can comment on or make changes to this bug.
Description
•