← 返回完整目录

EXPLORE JEV

评测与研究

查找评测、开源实现和研究资料;比较结果时注意数据、模型版本与测试条件。

THE RESOURCE INDEX

浏览本分类资源。

项目、工具、教程与文章
原始链接,按用途整理。

支持中文与多关键词,例如:浏览器 DOM。

全部资源

评测与研究

jevlike

开源选项评分实现,不是 TypeSafe 的模型权重。

Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.

评测与研究

jev-eval-agent

Personal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.

评测与研究

open-jev (daseinlabs)

One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.

评测与研究

jevmlx

Jev-style parallel constrained decisions for any MLX model on Apple Silicon. Typed, schema-valid JSON in one forward pass.

评测与研究

jev-on-a-laptop

Unofficial study: Jev-style parallel typed decisions on stock 1.5B-8B models on an Apple Silicon laptop. Benchmarks, research notes, and a Hugging Face Space demo.

评测与研究

open-jev (JoshuaSP)

Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results

评测与研究

jevbetter

A stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling, benchmarked against jevlike.

评测与研究

decider

One-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B

评测与研究

TypeAR

Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.

评测与研究

openvons

openvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド

评测与研究

jev-benchmarks

Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks.

评测与研究

jev-lm

A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval

评测与研究

jev-korean-benchmark

Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence

评测与研究

jevfire

JEV-inspired parallel decisions for CUDA LLMs. One context, many decisions. vLLM API, game-agent examples, and reproducible benchmarks.

评测与研究

openjev (zhihz)

Local bilingual probability decisions from context, questions, and candidate answers. Independent research preview inspired by TypeSafe Jev.

评测与研究

system-one-open

Open replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal)

评测与研究

jev-behavior-study

Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.

评测与研究

jev-chat

A chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.

评测与研究

jev_stock

An experimental JEV-powered framework for forecasting short-term stock price direction from structured market data.

评测与研究

jev-for-engineers

Eight minimal working examples of TypeSafe's Jev (a System One model) applied to mechanical and electrical engineering: CAD/CAE/CAM routing, FEM result triage, DFM screening, BOM alignment, hallucination-proof extraction. Zero dependencies.

评测与研究

trade-jev

Backtest Jev (TypeSafe) as a BUY/SELL/HOLD trader on NQ L10 order-book data

评测与研究

jev-benchmark

Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs

评测与研究

calibre

在自己的数据上测量置信度阈值,再据此在大小模型之间路由。

Calibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions

评测与研究

jevgpt

A chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively)

评测与研究

system-one

Batched single-token choice inference for open language models, compatible with TypeSafe

评测与研究

jev-little-airways

A show-and-tell capability study for Jev, TypeSafe's System One decision model.

评测与研究

kyotsu-ai-bench

AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard)

评测与研究

decisionbridge

A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.

评测与研究

padflow-jev-evals

Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.

评测与研究

mcts-agent

Discriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini

评测与研究

jev-rerank-bench

重排序评测,使用其结论前需查看数据集与测试方法。

Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.

评测与研究

jev-sec-bench

Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go

评测与研究

jev-agent-failure-benchmark

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).

评测与研究

qwen-rlcd

Jev-style calibrated decision model (Choice/Score/Noul) on Qwen3.5-0.8B

评测与研究

RISC-jeV

I tortured Jev into being a RISC-V CPU.

评测与研究

jev-synergy-screening

Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels

评测与研究

jev-freeform

An observable raw-character chat experiment powered entirely by TypeSafe Jev Choice

评测与研究

jev-playground (hegargarcia)

Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.

评测与研究

jev-report

中文报告与复现材料;2026-09-19 已确认仓库存在,本清单未复现其中结果。

发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表

评测与研究

jev-research-eval

Reproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator). Not investment advice.

评测与研究

jev-exploration

Jev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code

评测与研究

jev-phishing-bench

钓鱼内容分类对比,包含 Jev 表现不及对照模型的结果。

Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.

评测与研究

system-one-gemma

Open-source Jev-style System One decision model. Gemma 3 270M with a scoring head — fast, calibrated decisions in a single forward pass. No text generation. Inspired by TypeSafe.ai's Jev.

评测与研究

jev-routing-experiment

Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena

评测与研究

jev-jp-address

Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証

评测与研究

shade-arena-jev-monitor

Evaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.

评测与研究

thaiexam-jev-charts

Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models

评测与研究

jev-column-race

Jev vs Gemini 3.8 Flash: labelling 1,000 app reviews, 4.1× faster and 7× cheaper

评测与研究

jev-trace-classifier

Application of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local Qwen3.8-Flash-Next

评测与研究

jev-shadcn-lint-eval

A small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.

评测与研究

jev-deferred-crispification

Position paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models. Two lemmas, one principle (Deferred Crispification), one architecture (BSF-S1).

评测与研究

jev-pick-and-place-study

A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.

评测与研究

jev-spam-eval

Zero-shot spam filtering with TypeSafe Jev Noul questions, compared with TF-IDF baselines

评测与研究

jev-anotacao-sentencas

Jev (TypeSafe) vs. Gemini 3.8 Flash vs. GPT-5.6 Luna na anotação estruturada de sentenças do TJSP: qualidade, tempo e custo

评测与研究

jev-lab

TypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model

评测与研究

jev-headline-bench

Can Jev pick the winner of a real headline A/B test? 64.5% across 10,984 Upworthy randomized experiments, 74.7% when the difference was decisive.

评测与研究

jev-alpha-bench

Does Jev predict stock returns from news? It reads the news well; there is no tradeable alpha. Three arms separate reading from recall.

评测与研究

FinancialPredictionJev

Using Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good)

评测与研究

Typesafe_chess_eval

An evaluation of typesafe AI chess. As it turns out, the AI isn't doing really well even though chess is not a particularly open-ended game. Still, it's only a prototype and this probably wasn't optimzied for games.

评测与研究

typesafe-oracles

Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call

评测与研究

PocketJev

On-device iPhone visual decision tool using MLX and Qwen3-VL direct option logits.

评测与研究

misereru-slide-jev

Ongoing Japanese research deck on Jev and System One models, maintained as Markdown slides.

评测与研究

carlaiau/jev-reranking

Search engine experimentation on the TREC collections. Currently focused on zero-shot reranking implementations with typesafe.ai's JEV model

评测与研究

abhixhek/jevcal

Stop guessing confidence thresholds: calibrate, threshold, and drift-check typed decision models (TypeSafe Jev) against an LLM teacher.

评测与研究

zhuyansen/jev-search-rerank-eval

Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.

评测与研究

AntonioCoppe/jev-harness

Decision harness for TypeSafe Jev — confidence gates, shadow mode, recipes, and evals. Claude CLI 48.9s → Jev 1.3s on the same row-filter job.

评测与研究

0xnairb/research_desk

TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers

评测与研究

AkashPriyadarshii/jev-curate

High-throughput synthetic & pretraining dataset sifter powered by TypeSafe AI Jev (api.typesafe.ai). Stream, filter, and score Parquet & JSONL datasets at 1,500+ rows/sec using System One typed decisions (Choice, Score, Noul).

评测与研究

hamakyo/jev-starter

Typed, policy-driven decision workflows on top of TypeSafe AI Jev: confidence routing, fallbacks, evaluation, and RAG patterns for TypeScript apps.

评测与研究

memovai/openevals

Affordable platform for parallel agent evals and observability. Powered by JEV.

评测与研究

rongxinzy/LightJev

Train lightweight language backbones for typed decisions and candidate probabilities. CE/Brier training, evaluation, and an offline end-to-end demo.

评测与研究

24601/rh-guard

Reward-hack radar for coding agents: structural denies + TypeSafe Jev System One sidecar for Claude Code & Cursor hooks

评测与研究

4esv/jev-eval

Independent eval of TypeSafe Jev vs GPT-5.6 Terra: accuracy, calibration, latency, cost

评测与研究

avshalomd/longjev

Long inputs for TypeSafe AI's Jev decision model. An experiment, published with its evals.

评测与研究

BrendanH18/jev-lab

Six small apps and a workbench that show what TypeSafe's Jev (System One) model can do

评测与研究

choxos/jev-reviewer

Ask a trial report and its supplements for systematic review data by voice, text or a questions file. Jev (TypeSafe System One) points at the lines; every answer is a verbatim quote with its file and page. PDF, Word and text files; CSV export.

评测与研究

cmartinez9/jev-judge-bench

Binary LLM-judge bench — compare Jev (TypeSafe System One) against a frontier LLM judge on speed, cost, and agreement with human labels.

评测与研究

DECRUX9812/openjev-lm

open-Jev LM arm: Qwen2.5-0.5B + LoRA reproducing a hosted decision model's judgment at 92.9% on hand-labelled gold - trained overnight on a 6-vCPU CPU-only host, $0/call. Paper, corpora, harnesses, receipts.

评测与研究

heaven-hm/jev-system-one

A polished OpenAI + TypeSafe Jev terminal interface for answers with transparent decision reports

评测与研究

Jabbslad/pi-jev-tools

TypeSafe ranking, classification, retrieval and structured-decision tools for Pi coding agents

评测与研究

jmanhype/jev-dspy-lab

Reproducible calibration and selective-risk benchmarks for Jev/TypeSafe decisions in DSPy workflows

评测与研究

misaalya/snbt-jev-bench

Jev on Indonesia's SNBT 2025 university entrance test: 159 questions, seven subtests, audited answer keys.

评测与研究

robipop22/Jev-is-odd

Ask Jev by TypeSafe AI whether a number is odd. TypeScript, real token usage, and latency benchmarks.

评测与研究

samoweb3/jev-x-posts

Sortable 48-hour X post report for Jev, TypeSafe AI and Diogo Almeida, with Jev sentiment labels

部分项目补充了中文用途说明,原始描述保留供对照。在 GitHub 阅读完整目录 ↗