JEV RECIPES
Classifying, moderating, grading, routing, gating — these produce an enum and a probability, not a paragraph. Below are 5 jobs and 15 ready-made question sets, each runnable on this page with your own text.
We put Jev, GPT-4.1-mini and GPT-5.6 Sol / Terra / Luna through the same 2,390 labelled questions. True/false accuracy 92.5%, level with the newest reasoning models; median latency 456 ms against 1.3–1.7 s; $0.049 against $4.27 for the whole run. The costs: about 5 points behind on multiple choice, and clearly weak at counting.
The full data, a claim-by-claim check of the official numbers, and the limits of the run are in the benchmark report (written in Chinese for now).