Repository navigation
Conversation
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
ruff EXE001: the file has a shebang but shipped as mode 100644. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01Q7W9hWwuZU8Mkwo3ZCrR7m Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
- Layout lists the real paths (src/assetops_harbor, template/, overlays/) and marks datasets/ as generated. - Run steps work from the repo root: correct Dockerfile path, adapter defaults instead of a --template pointing at a generated task, no PYTHONPATH=agent, no 'push'. - Drop the pin to main@81265cb; the branch targets aafeedback_changes. - Replace 'Still unverified' with the Docker-host results reported in #558. Signed-off-by: Shuxin Lin <[email protected]>
generate_tasks.py gains three options for corpora that do not ship in the repo, such as AssetOpsBenchScenarioGeneration/scenarios_data: - --runtime-image: the image each task builds FROM. - --data-dir: where the corpus lives inside that image. The per-task layer copies the scenario there and the healthcheck runs init_data.py with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy, as scenario_suite_runner does. - --skip-missing: warn and skip profile scenarios absent from the corpus (all.yaml lists wosr-62, which the corpus lacks). corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into one layer over the runtime image, so tasks share it instead of each carrying it in their build context. Also stop copying scenario 1's description and the wosr keyword into every generated task. Signed-off-by: Shuxin Lin <[email protected]>
Signed-off-by: Shuxin Lin <[email protected]>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.
- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.
Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.
Signed-off-by: Shuxin Lin <[email protected]>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, and the generated dataset assetopsbench-corpus -> assetopsbench-suite, with matching wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <[email protected]>
Completes 533bc5a, which only moved corpus-image/ to suite-image/: assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <[email protected]>
A dropped VPN made the LiteLLM proxy unreachable mid-session. Every trial still built its containers, loaded its data, retried the model for four minutes and exited 1, and the job recorded each as done, so a resume would not have re-run them. - Probe LITELLM_BASE_URL / TOKENROUTER_BASE_URL (per model prefix) before starting or resuming a job, and skip the model with a clear message when the router does not answer. - Resume with --filter-error-type NonZeroAgentExitCodeError, so trials whose agent crashed run again while scored trials are kept. Signed-off-by: Shuxin Lin <[email protected]>
Add the generate_tasks.py flags a non-open profile needs, and record that those tasks load empty CouchDB collections: manifests reference shared/ files that exist only in the full corpus, not the runtime image, and the loader only warns when a data file is missing. Signed-off-by: Shuxin Lin <[email protected]>
StirrupAgent now loads the nearest .env above the working directory before checking router credentials, so tokenrouter/ and litellm_proxy/ models run without exporting their variables. override=False keeps exported variables and --ae ahead of the file, and only CREDENTIAL_ENV_VARS reach the container. Tests chdir into tmp_path so a developer's .env cannot leak in, and the credential fixture now restores variables that load_dotenv writes. Signed-off-by: Shuxin Lin <[email protected]>
The earlier note said the three FMEA mini scenarios load cleanly. They do not: their manifest lists files as an array, which the check that produced the note skipped. All 35 mini scenarios reference corpus-only files. Also record that the referenced files are 1.5 MB (mini) to 58 MB (all), not the 2.1 GB of the corpus shared/ directory, and how the failure shows up. Signed-off-by: Shuxin Lin <[email protected]>
feat(harbor): load .env for StirrupAgent credentials
Resolve the README conflict by keeping both new sections. The base's hand-generation section now uses this branch's --runtime-image, --data-dir and --skip-missing flags, and its known-gap note becomes what happens without the suite image. Signed-off-by: Shuxin Lin <[email protected]>
Add overlays/private-data.yaml, which bind-mounts the directory in AOB_PRIVATE_DIR read-only into main at /opt/suite/scenarios_data, the path the suite image uses. Tasks generated with --data-dir then load their manifests' shared/ files from the mount, with no suite image rebuild when the data changes. The variable is required with :? so an unset value fails naming it rather than as an invalid volume spec. Signed-off-by: Shuxin Lin <[email protected]>
refactor: drop FMSR generation, withhold LLM keys from MCP servers
Collaborator
|
@ShuxinLin merge already happen on main. |
Collaborator
Author
some edit go to main and rebase to aa branch. |
Restores the ten FMEA IDs dropped in 3f788d5. Co-Authored-By: Claude <[email protected]> Signed-off-by: Shuxin Lin <[email protected]>
DhavalRepo18
force-pushed
the
aafeedback_changes
branch
from
October 2, 2026 22:54
04dffef to
7c4b27d
Compare
It handles the boundary case where the category field is empty, and the agent keeps trying with different options. Co-Authored-By: Claude Opus 5.5 <[email protected]> Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
DhavalRepo18
force-pushed
the
aafeedback_changes
branch
from
October 5, 2026 13:39
b4d8e41 to
7c4b27d
Compare
Standarize the System Promot and Coed Exec tool for AA Study
This PR is based on v1 run
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Updated car list with additional values.
Signed-off-by: Dhaval Patel <[email protected]>
into add-tsfm-and-car- Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
Add tsfm and car
Bring TSFM and IoT server updates
run.sh gains -k (repeat each task n times, the basis for pass@k and pass^k), -t (sampling temperature, default 0.6) and -u (agent turn budget, default 100), with validation on each. passk.py computes pass@k and pass^k from a leaderboard directory using the existing "passed" reward key. Taken from #618 without the to_reward.py change. Co-Authored-By: Claude Opus 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01PcbP743uYAdkJiV9UbiTjy Signed-off-by: Dhaval Patel <[email protected]>
Add pass@k/pass^k, temperature and turn-budget controls to harbor runs
Signed-off-by: Dhaval Patel <[email protected]>
Signed-off-by: Dhaval Patel <[email protected]>
A CAR answer now passes when it is a single-key object whose key matches the gold mode (response / clarification / abstain). The value under the key no longer affects pass or score; required-term coverage stays in the mode_* fields as a diagnostic. Replaces the CAR tests that targeted the removed groundtruth_eval.json term scoring and updates docs/static-json-evaluation.md to match. Signed-off-by: Shuxin Lin <[email protected]>
ShuxinLin
force-pushed
the
aafeedback_changes
branch
from
October 8, 2026 13:21
25ed13f to
8c59cd9
Compare
fix(evaluation): score CAR scenarios on the mode key only
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
placeholder for now