Skip to content

docs(evaluator): note that temperature and system prompt are fixed - #8448

Open
Rajkumar2002-Rk wants to merge 1 commit into
simstudioai:stagingfrom
Rajkumar2002-Rk:docs/evaluator-fixed-settings
Open

Rajkumar2002-Rk wants to merge 1 commit into
simstudioai:stagingfrom
Rajkumar2002-Rk:docs/evaluator-fixed-settings

Conversation

@Rajkumar2002-Rk

Copy link
Copy Markdown

Summary

The Evaluator page says Temperature and a System Prompt are available under advanced, but neither can be set. Both sub-blocks are hidden: true in apps/sim/blocks/blocks/evaluator.ts, and evaluator-handler.ts always sends EVALUATOR.DEFAULT_TEMPERATURE (0.1) with a system prompt built from the metrics.

This updates the Model section to say that, and points to the metric descriptions and the content as the places to steer scoring. I ran into it while evaluating a triage workflow with the Evaluator block on a self-hosted install.

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation
  • Other: ___________

Testing

Docs only. Checked the wording against evaluator.ts and evaluator-handler.ts on staging.

Checklist

  • Code follows project style guidelines
  • Self-reviewed my changes
  • Tests added/updated and passing (docs only, no tests affected)
  • No new warnings introduced
  • I confirm that I have read and agree to the terms outlined in the Contributor License Agreement (CLA)

@vercel

vercel Bot commented Sep 30, 2026

Copy link
Copy Markdown

@Rajkumar2002-Rk is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Low risk] Updates documentation about evaluator block configuration.

The documentation change appears safe to merge, though two non-blocking statements should be corrected.

Findings

  1. P2 Temperature claim is too absolute ▶
  2. P2 Guidance mixed with scored content ▶

Summary

The PR corrects the Evaluator documentation to say that temperature and the system prompt cannot be configured. Two parts of the new guidance need refinement:

  • The default model does not receive a temperature setting, despite the “always 0.1” claim.
  • Scoring instructions should be placed in metric descriptions rather than in the content being evaluated.

Reviews (1) · Last reviewed commit: "docs(evaluator): note that temperature a..."

The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you.
The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. On hosted Sim the API key is supplied for you.

Temperature and the system prompt aren't configurable. Scoring always runs at temperature 0.1, and the system prompt is generated from your metrics. To steer how the model scores, put the guidance in the metric descriptions or in the content.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Temperature claim is too absolute. The Evaluator requests temperature 0.1, but the provider removes that setting for models that do not support it. The default claude-sonnet-5-5 does not receive a temperature setting, so readers may incorrectly expect every score to be produced at 0.1. Describe 0.1 as the requested value where supported.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you.
The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. On hosted Sim the API key is supplied for you.

Temperature and the system prompt aren't configurable. Scoring always runs at temperature 0.1, and the system prompt is generated from your metrics. To steer how the model scores, put the guidance in the metric descriptions or in the content.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Guidance mixed with scored content. The content field is the material being scored, usually an earlier agent output. Advising users to put scoring guidance there changes what the model evaluates and can bias the scores. Metric descriptions are the appropriate place for scoring instructions.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant