Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion apps/docs/content/docs/workflows/blocks/evaluator.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,9 @@ The content to score. Usually an earlier output like `<agent.content>`. Structur

### Model

The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you.
The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. On hosted Sim the API key is supplied for you.

Temperature and the system prompt aren't configurable. Scoring always runs at temperature 0.1, and the system prompt is generated from your metrics. To steer how the model scores, put the guidance in the metric descriptions or in the content.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Temperature claim is too absolute. The Evaluator requests temperature 0.1, but the provider removes that setting for models that do not support it. The default claude-sonnet-5-5 does not receive a temperature setting, so readers may incorrectly expect every score to be produced at 0.1. Describe 0.1 as the requested value where supported.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Guidance mixed with scored content. The content field is the material being scored, usually an earlier agent output. Advising users to put scoring guidance there changes what the model evaluates and can bias the scores. Metric descriptions are the appropriate place for scoring instructions.


### Fallback models

Expand Down