EXPLORE JEV
评测与研究
查找评测、开源实现和研究资料;比较结果时注意数据、模型版本与测试条件。
THE RESOURCE INDEX
浏览本分类资源。
项目、工具、教程与文章
原始链接,按用途整理。
支持中文与多关键词,例如:浏览器 DOM。
全部资源
jevlike
开源选项评分实现,不是 TypeSafe 的模型权重。
Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.
jev-eval-agent
Personal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.
typesafe-ai-benchmark
This is a LLM Gateway that mimics typesafe ai structured output. Like an imposter Jev.
open-jev (daseinlabs)
One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.
jevmlx
Jev-style parallel constrained decisions for any MLX model on Apple Silicon. Typed, schema-valid JSON in one forward pass.
jev-on-a-laptop
Unofficial study: Jev-style parallel typed decisions on stock 1.5B-8B models on an Apple Silicon laptop. Benchmarks, research notes, and a Hugging Face Space demo.
open-jev (JoshuaSP)
Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results
jevbetter
A stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling, benchmarked against jevlike.
decider
One-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B
TypeAR
Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.
jev-benchmarks
Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks.
jev-lm
A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval
jev-korean-benchmark
Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence
jevfire
JEV-inspired parallel decisions for CUDA LLMs. One context, many decisions. vLLM API, game-agent examples, and reproducible benchmarks.
openjev (zhihz)
Local bilingual probability decisions from context, questions, and candidate answers. Independent research preview inspired by TypeSafe Jev.
system-one-open
Open replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal)
jev-behavior-study
Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.
jev-chat
A chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.
jev_stock
An experimental JEV-powered framework for forecasting short-term stock price direction from structured market data.
jev-for-engineers
Eight minimal working examples of TypeSafe's Jev (a System One model) applied to mechanical and electrical engineering: CAD/CAE/CAM routing, FEM result triage, DFM screening, BOM alignment, hallucination-proof extraction. Zero dependencies.
jev-benchmark
Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs
calibre
在自己的数据上测量置信度阈值,再据此在大小模型之间路由。
Calibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions
jevgpt
A chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively)
system-one
Batched single-token choice inference for open language models, compatible with TypeSafe
jev-little-airways
A show-and-tell capability study for Jev, TypeSafe's System One decision model.
kyotsu-ai-bench
AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard)
decisionbridge
A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.
padflow-jev-evals
Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.
mcts-agent
Discriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini
jev-rerank-bench
重排序评测,使用其结论前需查看数据集与测试方法。
Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.
jev-sec-bench
Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go
jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
jev-synergy-screening
Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels
jev-freeform
An observable raw-character chat experiment powered entirely by TypeSafe Jev Choice
jev-playground (hegargarcia)
Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.
jev-report
中文报告与复现材料;2026-09-19 已确认仓库存在,本清单未复现其中结果。
发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表
jev-research-eval
Reproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator). Not investment advice.
jev-secret-detection
Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets
jev-phishing-bench
钓鱼内容分类对比,包含 Jev 表现不及对照模型的结果。
Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.
system-one-gemma
Open-source Jev-style System One decision model. Gemma 3 270M with a scoring head — fast, calibrated decisions in a single forward pass. No text generation. Inspired by TypeSafe.ai's Jev.
jev-routing-experiment
Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena
jev-jp-address
Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証
shade-arena-jev-monitor
Evaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.
thaiexam-jev-charts
Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models
jev-column-race
Jev vs Gemini 3.8 Flash: labelling 1,000 app reviews, 4.1× faster and 7× cheaper
jev-trace-classifier
Application of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local Qwen3.8-Flash-Next
jev-shadcn-lint-eval
A small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.
jev-deferred-crispification
Position paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models. Two lemmas, one principle (Deferred Crispification), one architecture (BSF-S1).
jev-pick-and-place-study
A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.
jev-spam-eval
Zero-shot spam filtering with TypeSafe Jev Noul questions, compared with TF-IDF baselines
jev-anotacao-sentencas
Jev (TypeSafe) vs. Gemini 3.8 Flash vs. GPT-5.6 Luna na anotação estruturada de sentenças do TJSP: qualidade, tempo e custo
jev-headline-bench
Can Jev pick the winner of a real headline A/B test? 64.5% across 10,984 Upworthy randomized experiments, 74.7% when the difference was decisive.
jev-alpha-bench
Does Jev predict stock returns from news? It reads the news well; there is no tradeable alpha. Three arms separate reading from recall.
FinancialPredictionJev
Using Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good)
Typesafe_chess_eval
An evaluation of typesafe AI chess. As it turns out, the AI isn't doing really well even though chess is not a particularly open-ended game. Still, it's only a prototype and this probably wasn't optimzied for games.
typesafe-oracles
Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call
system-one-adapter-rust
Rust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations)
misereru-slide-jev
Ongoing Japanese research deck on Jev and System One models, maintained as Markdown slides.
Parallel Constrained Decoding (Qwen2.5-1B-RLCD)
Hugging Face Space exploring open-source parallel constrained decoding as an alternative to Jev.
carlaiau/jev-reranking
Search engine experimentation on the TREC collections. Currently focused on zero-shot reranking implementations with typesafe.ai's JEV model
OmniJev/awesome-jev
Papers, open reproductions and independent evaluations behind System One models and Jev.
abhixhek/jevcal
Stop guessing confidence thresholds: calibrate, threshold, and drift-check typed decision models (TypeSafe Jev) against an LLM teacher.
zhuyansen/jev-search-rerank-eval
Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.
AntonioCoppe/jev-harness
Decision harness for TypeSafe Jev — confidence gates, shadow mode, recipes, and evals. Claude CLI 48.9s → Jev 1.3s on the same row-filter job.
0xnairb/research_desk
TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers
AkashPriyadarshii/jev-curate
High-throughput synthetic & pretraining dataset sifter powered by TypeSafe AI Jev (api.typesafe.ai). Stream, filter, and score Parquet & JSONL datasets at 1,500+ rows/sec using System One typed decisions (Choice, Score, Noul).
AnshChoudhary/typesafe-ai-firewall
Shadow-mode validation harness for a pre-execution firewall on AI agent tool calls (TypeSafe/Jev). Real run, findings in report.md.
hamakyo/jev-starter
Typed, policy-driven decision workflows on top of TypeSafe AI Jev: confidence routing, fallbacks, evaluation, and RAG patterns for TypeScript apps.
memovai/openevals
Affordable platform for parallel agent evals and observability. Powered by JEV.
rongxinzy/LightJev
Train lightweight language backbones for typed decisions and candidate probabilities. CE/Brier training, evaluation, and an offline end-to-end demo.
24601/rh-guard
Reward-hack radar for coding agents: structural denies + TypeSafe Jev System One sidecar for Claude Code & Cursor hooks
4esv/jev-eval
Independent eval of TypeSafe Jev vs GPT-5.6 Terra: accuracy, calibration, latency, cost
AIPI-mvoronovych/JEVBenchmark-Contradiction-Detection
Checking JEV's Contradiction detection (Model by TypeSafe.AI)
avshalomd/longjev
Long inputs for TypeSafe AI's Jev decision model. An experiment, published with its evals.
BrendanH18/jev-lab
Six small apps and a workbench that show what TypeSafe's Jev (System One) model can do
choxos/jev-reviewer
Ask a trial report and its supplements for systematic review data by voice, text or a questions file. Jev (TypeSafe System One) points at the lines; every answer is a verbatim quote with its file and page. PDF, Word and text files; CSV export.
cmartinez9/jev-judge-bench
Binary LLM-judge bench — compare Jev (TypeSafe System One) against a frontier LLM judge on speed, cost, and agreement with human labels.
DECRUX9812/openjev-lm
open-Jev LM arm: Qwen2.5-0.5B + LoRA reproducing a hosted decision model's judgment at 92.9% on hand-labelled gold - trained overnight on a 6-vCPU CPU-only host, $0/call. Paper, corpora, harnesses, receipts.
eggmasonvalue/jev-takes-mauboussin
Evaluating TypeSafe's Jev on Michael Mauboussin's 50-question decision calibration test
FFatTiger/new-api-plugin-typesafe
TypeSafe AI System One (Jev) task plugin for QuantumNous/new-api — native /v1/systemone, synchronous evaluation, token billing
heaven-hm/jev-system-one
A polished OpenAI + TypeSafe Jev terminal interface for answers with transparent decision reports
Jabbslad/pi-jev-tools
TypeSafe ranking, classification, retrieval and structured-decision tools for Pi coding agents
javiergradiche/ruby_llm-providers-typesafe
TypeSafe System One models (Jev) for RubyLLM: typed judgments, evaluations and reranking.
jeiel85/jevscope
Local-first visual decision debugger and regression testbench for TypeSafe AI Jev
jmanhype/jev-dspy-lab
Reproducible calibration and selective-risk benchmarks for Jev/TypeSafe decisions in DSPy workflows
misaalya/snbt-jev-bench
Jev on Indonesia's SNBT 2025 university entrance test: 159 questions, seven subtests, audited answer keys.
NicolasMontone/jev-evals
Rubric-based eval harness cheap enough to run on every PR, powered by typesafe-ai/jev
robipop22/Jev-is-odd
Ask Jev by TypeSafe AI whether a number is odd. TypeScript, real token usage, and latency benchmarks.
samoweb3/jev-x-posts
Sortable 48-hour X post report for Jev, TypeSafe AI and Diogo Almeida, with Jev sentiment labels
TyrellD1/typesafe-ai_smoke-test
Smoke test: route prompts to a work or life database with TypeSafe AI (Jev), 30-case eval
暂时没有找到。
试试项目英文名、相关用途,或重置筛选。
部分项目补充了中文用途说明,原始描述保留供对照。在 GitHub 阅读完整目录 ↗