RAVEL reinterprets published LHC searches. You ask a physics question in plain language, a coding agent plans the
calculation and stops for your approval, and the ravel package runs and checks MadGraph5_aMC@NLO, Pythia 8,
Delphes with SimpleAnalysis or Rivet, and pyhf. The result is a 95% CLs exclusion limit with the evidence behind it.
Install · Quick start · How RAVEL works · Commands · Capabilities · Validation · Limitations · Documentation
Note
RAVEL is research software, version 0.4.0. Install it from this repository: the package is not on PyPI, and
pip install ravel installs an unrelated project. Support and validation differ from analysis to analysis, so
read Limitations before you start a new study.
RAVEL does:
- turn a request into a task contract and a run ledger that record what was asked and what happened;
- survey what has been published (HEPData records, analysis routines, the paper's figures) and show you a plan at CHECK-IN 1 before any events are generated;
- run each stage under supervision, so a long run can be resumed;
- compute 95% CLs limits for one model point, and scan a grid of points into an exclusion contour;
- record every step, and keep failed or missing points visible;
- offer smaller routes: a cached replay, likelihood-only and generation-only runs, studies of events or models you already have, and a catalogue of analysis routines.
RAVEL does not:
- claim discoveries or quote p-values for new physics;
- generate events without a recorded approval;
- treat a passing software check as physics certification.
You need Git and Python 3.10 to 3.12; the hash-locked environment below uses Python 3.12. The Python package runs on macOS and Linux. The HEP simulation tools are installed separately, and only for event generation (see Toolchain and platforms).
git clone https://github.com/ammarphp/ravel.git
cd ravel
python3.12 -m venv .venv-replay
.venv-replay/bin/python -m pip install --require-hashes -r requirements-replay.lock
.venv-replay/bin/python -m pip install --no-deps .
.venv-replay/bin/ravel --helpThe lock file pins and verifies every dependency; --no-deps stops the package installation from changing them. The
package is ravel-hep; the command and the Python import are ravel. Commands with uv, and the development setup,
are in the installation guide.
Optional extras
| Extra | Adds | Needed for |
|---|---|---|
replay |
pyhf, NumPy, SciPy, Matplotlib | Replay and statistics (already in the lock file) |
science |
mplhep, uproot | Figures and ROOT files in supplied-data studies |
measurement |
Spey | Measurement inference |
classifier |
ONNX Runtime | Studies that use a trained classifier |
test |
pytest | Running the test suite |
The examples below write their outputs under local-runs/, which Git ignores.
.venv-replay/bin/ravel replay --out local-runs/replay-exampleThis takes under a minute, needs no network access after installation and works from any directory if you call ravel
by its full path. It refits a bundled benchmark, the ATLAS 2016 squark search (arXiv:1605.03814) at a squark mass of
800 GeV and a neutralino mass of 100 GeV, through pyhf from cached inputs. It does not generate events, and its
acceptance check is taken from the recorded baseline and labelled cached_replay: the replay is a regression check,
not a new analysis. It prints GATE: OK and writes environment.json (Python, dependencies, platform and bundle
fingerprint), results.json (fresh checks and their labelled scope) and work/ (statistical outputs and logs).
Use a new output directory for each attempt; RAVEL refuses to overwrite one, and keeps failed attempts.
This runs the plan, approve and run cycle without an agent, on a counting model bundled with RAVEL: signal region 2jl of the same ATLAS 2016 squark search, with 263 observed events and 283 ± 24 expected background events at 3.2 fb⁻¹. It takes a few seconds and runs offline. First draft the request and make the plan:
.venv-replay/bin/ravel initiate \
--prompt "Compute a likelihood-only check from the supplied counting model." \
--out local-runs/first-limit
.venv-replay/bin/ravel plan --rundir local-runs/first-limit \
--spec benchmarks/scoped/atlas-2jl-counting.jsonRead the plan before you approve it: CHECKIN1.md in the run directory, and the inputs/checkin1.json it points to,
give the recipe, the source, the assumptions and the 120-second ceiling. Then approve, in your own words, and run:
.venv-replay/bin/ravel approve --rundir local-runs/first-limit \
--quote "I have read the plan and approve this likelihood-only check."
.venv-replay/bin/ravel run --rundir local-runs/first-limit
.venv-replay/bin/ravel status --rundir local-runs/first-limit --writeYour quote is recorded as the approval, and status reports the lifecycle verdict PASS. The result is in
outputs/RESULT.md. It starts with an observed 95% CLs limit of 43.651372 and a median expected limit of 54.884401
signal events. This check uses a single-bin counting approximation, not the experiment's full likelihood.
Open this repository in your coding agent and type your request, for example "Reproduce Figure 16a of
arXiv:1911.12606 for a slepton-bino model". The agent drafts the task contract in its own run directory under
trial-runs/, where the repository's guardrails look, surveys the published inputs and stops at CHECK-IN 1. Nothing
is generated until you approve. See Working with a coding agent.
Colours: amber is you and the check-ins, pink the agent, red the guardrails and required checks, blue RAVEL, green external software and public data, grey the records.
| Part | What it does |
|---|---|
| You | Ask the question, and approve the plan at CHECK-IN 1 and the scale-up at CHECK-IN 2. You can also run ravel yourself. |
| Coding agent | Follows the workflow guide and drives RAVEL. |
| Workflow guide | The nine steps, each with a written procedure, and skills for recurring tasks. |
| Guardrails | Check the agent's actions as they happen: for example, no event generation without a recorded approval. |
| Public inputs | HEPData records, papers and public analysis code that RAVEL reads. |
| RAVEL | The ravel package: the command line, the workflow and run ledger, the physics and statistics code, figures and validation. |
| External HEP software | MadGraph5_aMC@NLO, Pythia 8, Delphes, SimpleAnalysis, Rivet and pyhf, run and supervised by RAVEL. |
| Run directory and result | Holds the task contract, approvals, logs, outputs and the result, so every number can be traced. |
Terms used in this README
- Task contract: your request in structured form: the analysis, model, target figure and masses.
- Check-in: a message to you at a set point in a run. CHECK-IN 1 and CHECK-IN 2 are gates: heavy compute waits for your answer. A deviation check-in reports a change of course; the run continues.
- Gate and hook: a gate blocks an action until its condition is met; hooks are scripts that a coding agent runs automatically at set moments, and they apply the gates.
- Lifecycle validator: checks that a run's required stages, approvals and records exist and are in order. Its
verdict is
PASSorFAIL. - Scoped run: one fixed calculation described in a JSON spec, planned with
ravel planand run once withravel run, without an agent. - Native and container fallback: native tools are installed directly on your Mac; the container fallback is a prebuilt x86 container for SimpleAnalysis routines without a native port, about 9 hours per point.
- Result pack: the machine-readable result files from which
RESULT.mdis written. - Sensitivity only: the expected reach of a search, not an exclusion.
Each step, and the check-ins, guardrails, records and execution backends, has its own page in the architecture pages.
First comes intake: your request becomes a task contract (analysis, model, figure, masses). Nothing runs. Then:
- Environment: check the native toolchain, or set it up.
- Inputs and route: find the published inputs and choose the analysis route. Then CHECK-IN 1: you approve the plan, budget and assumptions.
- Generate: generate events and check them before the shower.
- Analyze: apply the detector response and the analysis selection.
- Visualize: compare the signal with the published distributions.
- Acquire data: fetch the published statistical model: the likelihood, or per-region counts.
- Exclude: compute the 95% CLs limit.
- Scan: repeat steps 3 to 7 over a grid and draw the contour.
- Verify: trace every number and review the physics. Then the final check-in delivers the results, for you to accept or send back.
CHECK-IN 2 comes once, at the waypoint agreed at CHECK-IN 1: an early, cheap comparison with part of a published
figure, before the bulk of the compute. A change of course is reported at once in a deviation check-in and recorded;
the run continues. The replay and likelihood-only routes run only the statistics part of this sequence; a
generation-only run stops at parton-level events. To resume a run, rebuild its state with ravel status --rundir <run> --write; never start it again.
Open this repository in your coding agent and describe the task in plain language. The agent follows the workflow start guide: it surveys the published inputs, proposes a target figure, a budget and assumptions, and stops at CHECK-IN 1. No events are generated, not even a small test run, before you approve.
Coding agents that support the repository's hook settings enforce the gates automatically. Other agents read
AGENTS.md and the skills and follow the same gates from the written instructions; on those hosts the
agent runs each check itself, including the lifecycle validator before delivery, and nothing enforces this
automatically. Agent runs are stored under trial-runs/, one directory per run; see
the run-directory guide.
| Route | Use it when | Needs |
|---|---|---|
| Agent workflow | You want a full reproduction, reinterpretation or scan | A coding agent; the native toolchain for event generation |
| Scoped likelihood-only | You have a published likelihood or a counting model | Python only |
| Scoped generation-only | You want parton-level events for a declared process | An installed native MadGraph5_aMC@NLO |
| Supplied events or models | You already have events, weights or a model to study | Python, with the matching extras |
| Cached replay | You want to check an installation | Python only |
| Analysis discovery | You want to know which routines exist for an analysis | Python only |
The scoped and supplied-data routes are described in scoped workflows and scientific studies.
| Command | What it does |
|---|---|
ravel initiate |
Draft a task contract and an empty run ledger from a request; runs no compute |
ravel plan |
Prepare a concrete CHECK-IN 1 for a scoped or supplied-data run; runs no compute |
ravel approve |
Record your approval, bound to that CHECK-IN 1 |
ravel run |
Run one approved scoped attempt; --resume checks a finished attempt without rerunning it |
ravel status |
Rebuild the current state of a run from its records |
ravel validate |
Check a task contract; never authorises compute |
ravel compare-recipes |
Compare runs only when their executed recipes are verified |
ravel replay |
Refit the bundled benchmark from cached inputs |
ravel audit |
Inspect a checkout against the readiness criteria |
ravel analyses |
Search the catalogue of analysis routines (list, show, check-source, summary) |
Exit code 0 means the command completed (a status or audit report can still list failures), 1 means a check failed
or a request is unsupported, and 2 means a usage or input error. The CLI reference has every option
and output file.
ravel initiate runs no simulation and calls no AI model; a coding agent can add its own reading of the request with
--interpretation. A fresh draft's status reports FAIL and lists the stages still to run; it passes only when
every required stage has run and been checked.
RAVEL reinterprets a search when it can run the search's selection, certify that selection's acceptance against the experiment's published values, and use the published likelihood or per-region counts.
Events are generated on the native toolchain: MadGraph5_aMC@NLO and Pythia 8, with Delphes for the detector, installed directly on a Mac, at about 30 to 50 minutes per model point. Two kinds of selection code read them:
- Rivet routines that apply their own detector smearing, for ATLAS and CMS searches (22 in the Rivet release RAVEL uses);
- three ATLAS SimpleAnalysis routines, rewritten in Python:
| Routine | ATLAS search | Models | Output |
|---|---|---|---|
EwkCompressed2018 |
Compressed sleptons, arXiv:1911.12606 | Slepton–bino | Yields; limits once acceptance is certified |
EwkThreeLeptonERJR2018 |
Three leptons via WZ, arXiv:1912.08479 | Chargino–neutralino | Yields |
ZeroLeptonDiscovery2018 |
Zero leptons, arXiv:2010.14293 | Squark– and gluino–neutralino | Yields |
Limits use the experiment's published likelihood from HEPData when one exists. Otherwise RAVEL uses counting mode:
each signal region becomes one Poisson count of observed events, expected background with its uncertainty, and
signal, and the region with the best expected sensitivity sets the 95% CLs limit. The
capabilities page covers every route, and ravel analyses list searches a catalogue
of 633 public routines.
| Tool | Tested version |
|---|---|
| MadGraph5_aMC@NLO | 2.9.27 |
| Pythia 8 | 8.312 |
| Rivet, with YODA and HepMC3 | 4.1.3 |
| Delphes, with ROOT and FastJet | Not pinned; see the recorded environment |
| pyhf | 0.7.6 |
| SModelS, as a cross-check | 3.1.1 |
| Container fallback | mapyde 0.5.0 on podman |
Event generation needs the native toolchain, on macOS 11 or later; it is verified on macOS 15.5 with Apple Silicon. Check the computer with the read-only health check:
.venv-replay/bin/python -B -m ravel.validation.native_doctor --jsonAdd --require-rjr for routines that use RestFrames (recursive jigsaw reconstruction), such as EwkCompressed2018.
If tools are missing, follow the step 1 procedure, or let your coding agent
do it at step 1. On Intel Macs and Linux, use the Python-only routes. A native point takes about 30 to 50 minutes and
a scan from hours to overnight; the estimate at CHECK-IN 1 is the one to trust. Event generation also needs network
access and free disk: about 6 GB for each point running at once during a scan. See also native
portability.
An agent-workflow run directory, trial-runs/<date>_<label>/, holds:
<run>/
├── RESULT.md # The result, in words, with its figures
├── DEVIATIONS.md # Every change of course
├── RESUME.md # Current state and the commands to resume
├── run_state.json # The run ledger
├── result.json # The result pack (scan.json for a scan)
├── inputs/ # Task contract, cards, check-ins and approvals
├── outputs/ # Statistical results and validation records
├── plots/ # Figures and their index
└── logs/ # Logs and execution receipts
Scoped runs from ravel run use a flatter layout: RESULT.md, result.json and the figure are in outputs/.
Limits carry observed and expected values and say whether each crossing was resolved or hit a bound. A scan records which points finished, failed or are missing. A resumed stage reuses its output only when its recorded inputs, code and runtime still match. See scientific results, durable execution and the run-directory guide.
On a full agent-workflow run that generates events, RAVEL's validators check:
- the task contract and the order of the lifecycle stages;
- that approvals are bound to the inputs you reviewed;
- the generated events, before the shower;
- acceptance times efficiency against published values, before a limit is quoted;
- that a likelihood matches the selection that produced its yields;
- numerical safety: non-finite inputs, missing reference data, scan limits, bracketed CLs crossings and unsupported interpolation;
- figures drawn by RAVEL's plotting tools, when they are saved.
At step 9, the verification panel (the agent, with a fresh reviewer) checks every quoted number against the file it came from. Scoped, replay and supplied-data runs apply only the checks within their scope: approval binding, inputs, numerical status and receipts.
These checks test the software and the statistics; detector fidelity and inference rest on their own evidence. Each number comes from the analyses named in its row. S95 is the 95% CL upper limit on the number of signal events.
| Check | Analyses | Recorded evidence | What it establishes |
|---|---|---|---|
| Statistical recovery | Four 2016 ATLAS searches: arXiv:1605.03814, arXiv:1604.07773, ATLAS-CONF-2016-054 and ATLAS-CONF-2016-037 | 7 observed S95 comparisons within 8.6% (statistical layer) | Published S95 values recovered from the published counts |
| Implementation comparison | Compressed sleptons, arXiv:1911.12606 (EwkCompressed2018), at a 200 GeV slepton and a 150 GeV neutralino |
141/141 signal regions identical; final limit delta 0.51% | RAVEL's Python port matches ATLAS's code on the same detector-level events, and independent runs agree on the observed limit |
| Selection fidelity | arXiv:1605.03814, arXiv:1806.02293 and arXiv:2211.08028 | Six scorable cases: four pass, one warning, one fail | Agreement with published acceptance times efficiency |
| Workflow guards | Constructed fixtures, one per guardrail | 30 constructed gate cases | Each guardrail responds to the state it targets |
CI checks the marked values against the claim registry. The validation results give the provenance of each check.
- Detector response is fast simulation (Delphes, or Rivet's smearing), not Geant4.
- Certified runs use Pythia 8's default Monash tune, not the ATLAS A14 tune.
- Higher-order cross sections enter as a flat k-factor; PDF and scale uncertainties are not propagated.
- Decay spin correlations are not modelled in decay-table decays.
- For analyses without a published likelihood, the counting model is an approximation.
See the full list of known limitations, and what RAVEL claims.
| Goal | Start here |
|---|---|
| Get started | Documentation index · Installation · CLI reference |
| Understand the design | Architecture · Figures and their sources |
| Run a physics task | Start here · Workflow guide · Scoped workflows · Scientific studies |
| Understand results | Scientific results · Validation results · Evidence index |
| Reference | Capabilities · Scope · Task contract · Limitations · Failure modes · Third-party software |
| Contribute | Contributing · Directory map · Changelog |
Report problems in GitHub Issues. Include the command, its exit code,
and either the replay's environment.json or the output of the health check above. Development setup and the checks
a change must pass are in CONTRIBUTING.md.
Use GitHub's "Cite this repository" button or CITATION.cff. Please also cite the tools you used (listed in third-party software), the mapyde reference pipeline (arXiv:2306.11055) that RAVEL reproduces and extends, and the experimental publication and HEPData record behind your result.
RAVEL is released under the Apache License 2.0; see NOTICE. The external tools keep their own licences.
Appendix: detailed architecture
| Panel | Contains | Page |
|---|---|---|
| You | Physicist, and the four check-ins | Check-ins |
| Agent and guidance | Coding agent, instructions, workflow guide, skills, physics reviewer | Workflow guide |
| Guardrails | Session hooks, compute gate, deviation guard, stop checks, lifecycle validator | Guardrails |
| RAVEL | Command line, intake and routing, run ledger, resource census, supervised execution, scans and contours, statistics, cross sections, plotting, validation | The nine steps |
| Execution | Native toolchain, container fallback, Python only | Execution backends |
| Public inputs | HEPData, papers and code, routine catalogue | Inputs and route |
| External HEP software | MadGraph5_aMC@NLO, Pythia 8, Delphes, SimpleAnalysis, Rivet, pyhf, SModelS (cross-check) | Toolchain |
| Records and evidence | Run directory, result pack, curated evidence, benchmarks, continuous integration | Records |
Open the detailed architecture as a PDF. The check-ins, guardrails, execution and records panels, and every step, are explained with their own figures in the architecture pages.