Benchmarks & research

jev-bbq-experiment

@simonmesmith1Rupdated 2026-09-19

Reproducible evaluation of TypeSafe Jev on all 58,492 BBQ questions: accuracy, stereotype bias, uncertainty, cost and latency.

simonmesmith/jev-bbq-experiment

Where it calls Jev

req=urllib.request.Request('https://api.typesafe.ai/v1/systemone',data=json.dumps(p).encode(),headers={'Authorization':'Bearer '+self.key,'Content-Type':'applic

src/run.py:32

The link points at the commit we read, so the line number still holds.

This one cannot run here

Its question set is assembled at runtime, or never written out literally in the code, so there is nothing to lift. The source link above will show you.

Other projects in this category