docs(evaluator): note that temperature and system prompt are fixed - #8448
Rajkumar2002-Rk wants to merge 1 commit into
Conversation
|
@Rajkumar2002-Rk is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
|
| The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you. | ||
| The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. On hosted Sim the API key is supplied for you. | ||
|
|
||
| Temperature and the system prompt aren't configurable. Scoring always runs at temperature 0.1, and the system prompt is generated from your metrics. To steer how the model scores, put the guidance in the metric descriptions or in the content. |
There was a problem hiding this comment.
Temperature claim is too absolute. The Evaluator requests temperature 0.1, but the provider removes that setting for models that do not support it. The default
claude-sonnet-5-5 does not receive a temperature setting, so readers may incorrectly expect every score to be produced at 0.1. Describe 0.1 as the requested value where supported.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you. | ||
| The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. On hosted Sim the API key is supplied for you. | ||
|
|
||
| Temperature and the system prompt aren't configurable. Scoring always runs at temperature 0.1, and the system prompt is generated from your metrics. To steer how the model scores, put the guidance in the metric descriptions or in the content. |
There was a problem hiding this comment.
Summary
The Evaluator page says Temperature and a System Prompt are available under advanced, but neither can be set. Both sub-blocks are
hidden: trueinapps/sim/blocks/blocks/evaluator.ts, andevaluator-handler.tsalways sendsEVALUATOR.DEFAULT_TEMPERATURE(0.1) with a system prompt built from the metrics.This updates the Model section to say that, and points to the metric descriptions and the content as the places to steer scoring. I ran into it while evaluating a triage workflow with the Evaluator block on a self-hosted install.
Type of Change
Testing
Docs only. Checked the wording against
evaluator.tsandevaluator-handler.tsonstaging.Checklist