Benchmarks & research

padflow-jev-evals

@zsavage81PythonMITupdated 2026-09-17

Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.

zsavage8/padflow-jev-evals

Where it calls Jev

python scripts/run_baseline.py --model jev --base-url https://api.typesafe.ai/v1 --api-key-env TYPESAFE_API_KEY

scripts/run_baseline.py:7

The link points at the commit we read, so the line number still holds.

This one cannot run here

Its question set is assembled at runtime, or never written out literally in the code, so there is nothing to lift. The source link above will show you.

Other projects in this category