Benchmarks & research

jev-search-rerank-eval

@zhuyansen4PythonMITupdated 2026-09-18

Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.

zhuyansen/jev-search-rerank-eval

Where it calls Jev

"jev": jev[k], "llm": llm[k], "label": None}

src/jse/cli.py:131

The link points at the commit we read, so the line number still holds.

This one cannot run here

Its question set is assembled at runtime, or never written out literally in the code, so there is nothing to lift. The source link above will show you.

Other projects in this category