You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Runs ~20 options strategies against live market data in shadow mode, records every hypothetical fill under worst/base/optimistic assumptions, and grades each with anytime-valid e-processes. Places no orders.
Which model is really behind your API relay or agent IDE? Behavioral fingerprinting + anytime-valid sequential tests (FPR<=1%). LLMs can't be random - measured on 9 frontier models.
Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.
A/B testing and causal inference scored against known answers: simulations where I set the effect, and a randomised benchmark the observational methods have to recover.
Bayesian multi-armed bandits for continuous prompt experimentation: Thompson sampling routes traffic to the best prompt variant and a stopping rule promotes a winner without a fixed-N A/B test. Zero dependencies, TypeScript-first.
Ships ML models in stages - shadow, then 1/5/25/50% of traffic - and rolls a bad one back automatically. The guardrails stay valid under constant checking: 0.6% false rollback with two identical models, where a repeatedly-read A/B test acts wrongly 36.7% of the time.
A/B testing toolkit: power analysis, SRM, CUPED, BH-FDR, and a Monte Carlo calibration harness that checks whether naive peeking rules actually control the false-positive rate they claim to.
Typed toolkit for A/B experiment analysis: power analysis, CUPED variance reduction, always-valid sequential testing (mSPRT), Welch/z analysis with CIs - property-based tests (hypothesis) and a worked forecasting-rollout case study.
A/B testing framework with Frequentist, Bayesian, and Sequential testing — includes a peeking problem simulator that visually proves why standard testing inflates false positive rates by 3-4x.