Awesome Decision Models

Awesome Decision Models

A curated list of decision models (also called System One models or typed decision models) and the APIs, runtimes, tools, applications, benchmarks, and research around them.

精选的决策模型(decision model,也称 System One 模型、类型化决策模型)资源,以及围绕它们的 API、运行时、工具、应用、评测与研究。

A decision model reads a state (text, JSON, and for some models images) plus questions whose answers you declare up front, and returns a probability for every answer instead of generated text. Questions come in three shapes: Noul (the probability that a statement is true), Choice (one of your options), and Score (a level on an ordered rubric). TypeSafe AI introduced the category in September 2026 with Jev and its /v1/systemone API, and many of the models and runtimes below accept the same request shape.

决策模型读入一段状态(文本、JSON,部分模型也支持图片),再加上答案事先声明好的问题,为每个答案返回一个概率,而不是生成文本。问题有三种形态:Noul(某个陈述为真的概率)、Choice(从你给的选项里选一个)和 Score(在有序量表上的等级)。TypeSafe AI 在 2026 年 9 月用 Jev 和它的 /v1/systemone API 开创了这一类别,下面许多模型和运行时都接受同样的请求格式。

Community-maintained and not affiliated with any model provider. Pull requests welcome.

由社区维护,与任何模型厂商无隶属关系。欢迎 PR。

Hosted APIs托管 API 7

API-only models, oldest first. Open-weight models that their publishers also host, such as Clef and pplx-decider, are under Open Models.

只提供 API 的模型,按发布时间排序。厂商同时托管的开源权重模型(如 Clef、pplx-decider)放在开源模型。

  • Jevdocs.typesafe.ai

    TypeSafe AI's first System One model and the origin of the /v1/systemone API: Noul, Choice, and Score questions over a text state in one request. Keys from the console; also served through Vercel AI Gateway, Cloudflare Workers AI, and OpenRouter.

    TypeSafe AI 的第一个 System One 模型,也是 /v1/systemone API 的源头:一次请求里对文本状态回答 Noul、Choice 和 Score 问题。密钥在控制台获取;也可以通过 Vercel AI Gateway、Cloudflare Workers AI 和 OpenRouter 调用。

  • meraGPT Decider 1meragpt.com

    Hosted decision model (state-decider-1, alias sd-1) answering Noul, Choice, and Score through /v1/systemone, compatible with the TypeSafe SDK; requests are limited to 4,096 tokens and Choice to ten labels

    托管决策模型(state-decider-1,别名 sd-1),通过 /v1/systemone 回答 Noul、Choice 和 Score,可用 TypeSafe SDK 调用;请求上限为 4,096 token,Choice 最多十个标签

  • Solar Decideopenrouter.ai

    Upstage's decision model on Solar Mini 4, in beta, with a 512K context and the same request format as Jev

    Upstage 基于 Solar Mini 4 的决策模型,目前为 beta,512K 上下文,请求格式与 Jev 相同

  • Span-01respan.ai

    Respan's behavior classifier for AI traces: for each plain-language behavior you define, such as prompt injection, hallucination, or agent loops, it returns present, absent, or not observable. Span-01 and a Lite tier are on OpenRouter.

    Respan 面向 AI trace 的行为分类模型:对你用自然语言定义的每种行为(提示注入、幻觉、Agent 死循环等)判断为出现、未出现或无法观察。Span-01 及其 Lite 版已上架 OpenRouter。

  • d1docs.liquid.ai

    Liquid AI's first decision model, served at a /v1/systemone endpoint that the TypeSafe SDKs can call, and on OpenRouter. Model size not disclosed.

    Liquid AI 的第一个决策模型,提供可用 TypeSafe SDK 调用的 /v1/systemone 端点,也上架了 OpenRouter。未公开模型规模。

  • OpenAI Decisions APIopenai.com

    Announced at DevDay 2026 and built on Luna: context, a question, and a closed list of answers in; an answer with a confidence score out. Limited preview, pricing not yet published.

    在 DevDay 2026 发布,基于 Luna:输入上下文、问题和一组封闭的候选答案,返回一个答案及其置信度。目前为 limited preview,尚未公布定价。

  • Mercury Decideopenrouter.ai

    Inception's decision model, with a free route on OpenRouter; Inception cites up to 14 decisions per second.

    Inception 的决策模型,OpenRouter 上有免费线路;Inception 称每秒最多可做 14 次决策。

Open Models开源模型 27

Open-weight models you can download and run. Scores are as reported by each project on the benchmark it chose, so numbers are not comparable across entries.

可以下载并自行运行的开源权重模型。分数均为各项目在自选 benchmark 上的自报结果,不同条目之间不可直接比较。

From companies and labs公司与机构发布

  • Clefhuggingface.co/Cloudflare

    Cloudflare's Apache-2.0 decision models post-trained from Qwen3.8-27B: Clef reads text, JSON, images, or video and scores every option of every question jointly, and Clef-flash is the smaller, faster variant. Both are hosted on Workers AI with a Jev-compatible API. Launch post with Jev Decision Index results: blog.

    Cloudflare 基于 Qwen3.8-27B 后训练的 Apache-2.0 决策模型:Clef 可读文本、JSON、图片或视频,对所有问题的所有选项联合打分;Clef-flash 是更小更快的版本。两者都托管在 Workers AI 上,API 与 Jev 兼容。发布文章附 Jev Decision Index 成绩:blog。

  • pplx-decider-v1-27bhuggingface.co/perplexity-ai

    Perplexity's Apache-2.0 multimodal decision model fine-tuned from Qwen3.8-27B; its card reports a mean of 85.71% over 11 benchmarks against 84.51% for Jev. Also served through the Perplexity Decisions API.

    Perplexity 基于 Qwen3.8-27B 微调的 Apache-2.0 多模态决策模型;模型卡报告 11 项 benchmark 平均 85.71%,Jev 为 84.51%。也可通过 Perplexity Decisions API 调用。

  • Strands Decider 2Bgithub.com/strands-labs

    AWS Strands Labs' decision model: Qwen3.5-2B with the LM head replaced by a pointer head of about a million parameters that scores each option, plus a rank-16 LoRA. Weights, training data, and scripts are public; the launch post uses it to check a Strands agent's tool calls before they run.

    AWS Strands Labs 的决策模型:把 Qwen3.5-2B 的 LM head 换成约一百万参数的 pointer head 为每个选项打分,并加 rank-16 LoRA。权重、训练数据和脚本全部公开;发布文章用它在 Strands Agent 调用工具前做检查。

  • Nimblegithub.com/bespokelabsai

    Bespoke Labs' 9B LoRA fine-tune of Qwen3.5-9B on contrastively curated synthetic data, with a public 13-dataset benchmark suite. In Ollama as nimble.

    Bespoke Labs 用对比式筛选的合成数据对 Qwen3.5-9B 做的 9B LoRA 微调,附公开的 13 个数据集评测套件。Ollama 中名为 nimble。

  • Tev1huggingface.co/togethercomputer

    Together AI's experimental 4B and 0.8B supervised fine-tunes of Qwen3.5, released with a walkthrough on training your own decision model for $17. In Ollama as tev1.

    Together AI 基于 Qwen3.5 的实验性监督微调模型(4B 和 0.8B),同时发布了花 17 美元训练自己的决策模型的教程。Ollama 中名为 tev1。

  • Layagithub.com/NandhaKishorM

    Convai Innovations' multilingual non-autoregressive decision model: Choice, Score, and Noul in one forward pass, with published weights, a PyPI package, and a router that picks a checkpoint per request

    Convai Innovations 的多语言非自回归决策模型:一次前向完成 Choice、Score 和 Noul,权重和 PyPI 包已公开,并按请求选择检查点

  • GLiNER2.5-Decidehuggingface.co/fastino

    Fastino's Apache-2.0 DeBERTa-v3-large decision classifier: caller-defined tasks and labels get probabilities in one forward pass through gliner2 on CPU or GPU, for operational classification, routing, and ordinal scoring

    Fastino 基于 DeBERTa-v3-large 的 Apache-2.0 决策分类模型:调用方定义任务和标签,通过 gliner2 在 CPU 或 GPU 上一次前向返回概率,面向业务分类、路由与有序评分

  • Standard Onehuggingface.co/StandardThinking

    Standard Thinking's Apache-2.0 3B and 8B decision models on Ministral 3, with a retained Pixtral vision tower: text or images in, option probabilities out through /v1/systemone, with merged weights, adapters, GGUF builds, and server code

    Standard Thinking 基于 Ministral 3 的 Apache-2.0 决策模型(3B、8B),保留 Pixtral 视觉编码器:输入文本或图片,通过 /v1/systemone 返回选项概率,附合并权重、适配器、GGUF 和服务端代码

  • OpenJevhuggingface.co/openjev

    Independent open-weight decision model that scores caller-defined options in one forward pass, with calibration and serving code; weights are CC BY-NC 4.0 for noncommercial use and helper/server code is Apache-2.0. Separate from SemIf, formerly named OpenJev.

    独立开放权重决策模型,一次前向为调用方定义的选项打分,附校准与服务端代码;权重采用 CC BY-NC 4.0,仅限非商业用途,辅助与服务端代码采用 Apache-2.0。与曾名为 OpenJev 的 SemIf 是不同项目。

Community models社区模型

  • Kevgithub.com/jaredpalmer

    Qwen3.5 decision models (0.8B, 4B, 9B) you can train and serve yourself. Choice, Score, and Noul in one forward pass, with published weights and frozen eval suites, and a local server that speaks /v1/systemone.

    Qwen3.5 决策模型(0.8B、4B、9B),可以自己训练和部署。一次前向完成 Choice、Score 和 Noul,权重和固定评测集已公开,本地服务实现 /v1/systemone。

  • Vongithub.com/wfzyx

    Local non-autoregressive System One model with a /v1/systemone-compatible server and a Doom demo where each move is one forward pass

    本地非自回归 System One 模型,服务端兼容 /v1/systemone,带 Doom 演示:每步动作是一次前向

  • NanoJevgithub.com/TianyuCodings

    0.6B parallel decision model: state and questions in, a full distribution out, no decoded text. Published weights and one checkpoint for ViZDoom, Maze, and Snake; on the author's ViZDoom Basic split, 128/128 against 56/128 for Jev.

    0.6B 并行决策模型:输入状态和问题,一次前向给出完整分布,不解码文本。权重已公开,同一检查点可玩 ViZDoom、迷宫和贪吃蛇;作者自己的 ViZDoom Basic 划分上是 128/128,对照 Jev 为 56/128。

  • jevlikegithub.com/vinnylarouge

    Train a small one-pass scorer that maps context + N text options to a probability per option. Includes Doom / chess vision demos and a Wikispeedia next-click example. Explicitly not a reproduction of TypeSafe's architecture or RLCD.

    训练一个小的单次 scorer:上下文 + N 个文本选项 → 每个选项一个概率。含 Doom / 国际象棋视觉 demo,以及 Wikispeedia 下一跳例子。明确*不是* TypeSafe 架构或 RLCD 的复现。

  • decidergithub.com/Mapika

    Qwen3.5-2B fine-tune that emits typed decisions with calibrated probabilities in one pass

    基于 Qwen3.5-2B 的微调:一次前向就给出类型化决策和校准概率

  • RSI-Jevgithub.com/Shanghua-Gao

    Jev-like model trained by a recursively self-improving (RSI) AI research system that publishes every experiment, failures included. 2B Qwen3.5, with Choice, Score, and Noul in one forward pass over text or up to four images (v4.0-VL), open weights, and a /v1/systemone-compatible server.

    由递归自我改进(RSI)的 AI 研究系统训练的 Jev-like 模型,公开每一次实验(包括失败的)。2B Qwen3.5,一次前向给出 Choice、Score、Noul,可读文本或最多四张图片(v4.0-VL),开放权重,提供兼容 /v1/systemone 的服务。

  • jev-stylegithub.com/lawrence3699

    0.8B Qwen3.5 decision model you install with pip install "jev-style[torch]" (or [mlx] on Apple silicon) and run locally (PyTorch, MLX, or llama.cpp with a separately built scorer) behind a /v1/systemone-compatible server: Choice, Score, and Noul in one pass, plus an MCP server and a Claude Code guard hook

    0.8B 决策模型(Qwen3.5 微调),pip install "jev-style[torch]"(Apple 芯片用 [mlx])即可在本地运行(PyTorch、MLX,或配合单独编译的打分程序用 llama.cpp),提供兼容 /v1/systemone 的服务:一次前向回答 Choice、Score、Noul,并附带 MCP 服务和 Claude Code 守门钩子

  • PlayJevgithub.com/OmniJev

    Qwen3.5-0.8B-Base fine-tuned to play ten browser games from 448 px frames: one forward pass per move, a probability over the game's option list read off the option letters, no generated text. Open weights and a demo of all ten in the browser.

    Qwen3.5-0.8B-Base 微调后从 448 px 画面玩十款浏览器小游戏:每步一次前向,概率直接从选项字母上读出,不生成任何文本。权重和十款游戏的浏览器 demo 都已公开。

  • OneJevgithub.com/OmniJev

    Open multimodal System One model from the OmniJev team, in four sizes (0.8B to 27B): Choice, Score, and Noul questions about a screenshot, photo, video, or text get a calibrated probability for every option in one forward pass. Weights on Hugging Face.

    OmniJev 团队的开源多模态 System One 模型,四种尺寸(0.8B 到 27B):对截图、照片、视频或文本提出 Choice、Score、Noul,一次前向为每个选项给出校准概率。权重在 Hugging Face。

  • jevosgithub.com/feder-cr

    1B decision model for CPU-only laptops: MiniCPM5 cut to 17 layers with a one-logit head, GGUF q4_k_m at 619 MB, running on llama.cpp with no GPU, about 54 ms per short request. Speaks Jev's /v1/systemone wire format for Noul (yes/no) questions only; Choice and Score return 422.

    1B 决策模型,面向纯 CPU 笔记本:把 MiniCPM5 裁剪到 17 层并接一个单 logit 输出头,GGUF q4_k_m 量化后 619 MB,跑在 llama.cpp 上不需要 GPU,短请求约 54 ms。只支持 Jev /v1/systemone 协议里的 Noul(是/否)问题,Choice 和 Score 会返回 422。

  • WebJevgithub.com/lexmount

    Jev-compatible decision model for browser agents: a Qwen3.5-35B-A3B fine-tune that picks the next operation and target element in the Jev Ultrafast loop, behind a /v1/systemone-compatible server. On 125 real-website tasks graded by deterministic verifiers, it completes 38.5% vs 16.7% for Jev 1.13 in the same agent. Apache-2.0 weights, training data and recipe, and a demo app.

    浏览器 agent 决策模型,接口与 Jev 兼容:基于 Qwen3.5-35B-A3B 微调,在 Jev Ultrafast 循环中选择下一步操作和目标元素,提供兼容 /v1/systemone 的服务。在 125 个由确定性判据评分的真实网站任务上,同一 agent 中完成率为 38.5%,Jev 1.13 为 16.7%。公开 Apache-2.0 权重、训练数据、训练方法和演示应用。

  • Vevgithub.com/Xiaooolong

    Open-source Jev implementation with vision input, fine-tuned from Qwen3.5 4B / 9B. Screenshots and photos go straight into the state, and Choice, Score, and Noul answers draw on both text and image, read from label-token probabilities with no generated text. Serves /v1/systemone; open code and weights, weights for non-commercial use only.

    支持视觉输入的 Jev 开源实现,基于 Qwen3.5 4B / 9B 微调。截图、照片可以直接放进 state,Choice、Score、Noul 结合文本与图像作答,概率直接从标签 token 上读出,不生成文本。提供兼容 /v1/systemone 的服务,代码与权重开源,权重仅限非商用。

  • JEV-27Bhuggingface.co/autotrust

    AutoTrust's Apache-2.0 model distilled from Jev 1.13 onto Qwen3.8-27B; one vLLM engine serves both the decisions and the untouched Qwen model for ordinary generation. The authors report 84.07 against 83.85 for Jev over six public decision benchmarks in their own runs. Smaller sibling: JEV-9B.

    AutoTrust 以 Jev 1.13 为教师、在 Qwen3.8-27B 上蒸馏的 Apache-2.0 模型;同一个 vLLM 引擎既提供决策,也提供未改动的 Qwen 用于普通生成。作者自测六个公开决策 benchmark 平均 84.07,Jev 为 83.85。较小版本:JEV-9B。

  • Winnowhuggingface.co/EldanRing

    Apache-2.0 Gemma 4 fine-tunes (12B and E4B) for typed decisions, served from a llama.cpp-based server that answers both /v1/systemone and /v1/chat/completions.

    面向类型化决策的 Apache-2.0 Gemma 4 微调(12B 和 E4B),基于 llama.cpp 的服务同时提供 /v1/systemone 和 /v1/chat/completions。

  • JevK5github.com/allebee

    Apache-2.0 decision models on Qwen3.5 (2B, 4B, 9B) behind a /v1/systemone server, with GGUF builds for llama.cpp on CPUs and GPUs

    基于 Qwen3.5 的 Apache-2.0 决策模型(2B、4B、9B),提供 /v1/systemone 服务,并有可在 CPU 和 GPU 上用 llama.cpp 运行的 GGUF 版本

  • reflexgithub.com/kshetrajna12

    Small decision model on a frozen Qwen3.5-4B: one /v1/systemone endpoint answered by a single forward pass in about 200 ms, compared against Jev on JevBench's public items

    基于冻结 Qwen3.5-4B 的小型决策模型:单个 /v1/systemone 端点,一次前向约 200 ms 作答,并在 JevBench 公开题上与 Jev 对比

  • imajevhuggingface.co/mohit67890

    Apache-2.0 4B LoRA on Qwen3.5-4B for decisions about photos: checks a photo against your record or compares two photos, with a trained unknown probability so the app can stop instead of guessing. Jev's request shape plus images; runs with MLX or PyTorch.

    基于 Qwen3.5-4B 的 Apache-2.0 4B LoRA,用于针对照片的决策:把照片和你的记录核对,或比较两张照片;训练了 unknown 概率,应用可以停下而不是瞎猜。请求格式为 Jev 的格式外加 images,可用 MLX 或 PyTorch 运行。

  • NeoHorse-Jev-4Bhuggingface.co/TokenRhythm

    TokenRhythm's Apache-2.0 4B decision model on NeoHorse-1-4B: text or a single image with text, prefill-only inference, served at /v1/systemone

    TokenRhythm 基于 NeoHorse-1-4B 的 Apache-2.0 4B 决策模型:输入文本或单张图片加文本,只做 prefill 推理,提供 /v1/systemone 服务

Inference Techniques推理方法 8

Ways to get decision-model behavior from existing models, usually by reading option probabilities in one forward pass, with little or no training.

不训练或少量训练、直接让现有模型表现得像决策模型的方法,通常是在一次前向中读出各选项的概率。

  • SemIfgithub.com/TheoLeeCJ

    Formerly OpenJev. Typed decisions from open models on a home RTX 3090 and in the browser: reads option logits instead of generating text.

    原名 OpenJev。用开源模型在家用 RTX 3090 和浏览器里做类型化决策,直接读选项 logits,不生成文本。

  • LitJevgithub.com/zhengxuyu

    A reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.

    Jev 的复现:把任意 Qwen 模型变成快速决策模型,提供与 Jev 相同的 /v1/systemone schema(Choice、Score、Noul),不训练、不生成回答文本

  • TetraJevgithub.com/FeiLiuEM

    Locally-deployed decision layer for complex decision problems across domains: two frozen open-weight readers give four readings per item — letter and per-candidate yes/no — fused fit-free and routed by agreement with calibrated release gates; evaluated on eight decision suites and the RAG reranking pass, including DecisionBench's 35 real-world task categories. No training of any kind.

    可本地部署的决策层,处理跨领域的复杂决策:两个冻结的开源权重读取器对每条样本给出四次读数(选项字母,以及逐候选项的是/否),无拟合地融合,再按一致性路由并设有校准后的放行门;已在八个决策套件和 RAG 重排上评测,包括 DecisionBench 的 35 类真实任务。全程不训练。

  • jevmlxgithub.com/bnsd55

    Jev-style parallel constrained decisions for any MLX model on Apple Silicon: schema-valid JSON in one forward pass

    给任意 MLX 模型做 Jev 风格并行约束决策:一次前向得到带概率的、按 schema 合法的 JSON

  • JEVfiregithub.com/kikoncuo

    Jev-inspired parallel decisions for CUDA LLMs via vLLM, with a browser Mario demo (~71 ms/action locally)

    CUDA LLM 上的 Jev 风格并行决策(vLLM),带浏览器马里奥 demo(本地约 71 ms/步)

  • PocketJevgithub.com/NullPo-jp

    On-device iPhone visual decisions with MLX + Qwen3-VL option logits. Camera + 3-choice, no text generation, ~1s, no photo saved.

    iPhone 端侧视觉判断:MLX + Qwen3-VL 选项 logits。相机 + 三选一,不生成文字,约 1 秒,不存照片。

  • jev-visualgithub.com/hr98w

    Educational Jev-like visual inference on Apple Silicon: shared multimodal context, candidate scoring, sorting-factory / Breakout / gesture demos

    Apple Silicon 上的教学向 Jev 风格视觉推理:共享多模态上下文、候选打分,含分拣厂 / Breakout / 手势 demo

  • DiffusionGemma decision endpointDiffusionGemma 决策端点huggingface.co/spaces/victor

    Free Hugging Face Space that reads a probability for every answer from DiffusionGemma, without fine-tuning, behind the /v1/systemone wire format

    免费的 Hugging Face Space:不做微调,直接从 DiffusionGemma 读出每个答案的概率,对外使用 /v1/systemone 协议

Runtimes & Platforms运行时与平台 6

Servers, gateways, and native runtimes for running decision models.

运行决策模型的服务端、网关与原生运行时。

  • Ollamaollama.com

    Serves /v1/systemone locally since 0.35; the first decision models in its library are nimble and tev1.

    从 0.35 起在本地提供 /v1/systemone;模型库里首批决策模型是 nimble 和 tev1。

  • Ollayagithub.com/ollaya-dev

    Ollama-style runtime for decision models: pulls and serves open encoders and decoders (Laya, Von, Kev, Decider, Nimble, Winnow, and more) with each author's calibration, behind /v1/systemone; the official TypeSafe Python SDK works against it unchanged. Site: ollaya.dev.

    面向决策模型的 Ollama 式运行时:拉取并运行开源 encoder 与 decoder 模型(Laya、Von、Kev、Decider、Nimble、Winnow 等),沿用各作者的校准,对外提供 /v1/systemone;官方 TypeSafe Python SDK 无需修改即可连接。官网:ollaya.dev。

  • Laya-MLXgithub.com/mizorewww

    Independent native MLX port of Laya for Apple Silicon: Choice, Score, and Noul without text generation or a cloud API, retaining upstream question formatting and calibration, with published port-fidelity checks and performance measurements

    面向 Apple Silicon 的独立 Laya 原生 MLX 移植:本地完成 Choice、Score 和 Noul,无文本生成或云 API,沿用上游问题格式与校准,附公开的移植一致性检查和性能测量

  • OpenRouter decision modelsOpenRouter 决策模型openrouter.ai

    Decision models from several publishers behind OpenRouter's alpha Decisions API, which chat-completions SDKs cannot call. Usage guide as an agent skill: openrouter-decisions.

    OpenRouter alpha 版 Decisions API 上多家发布方的决策模型;chat completions SDK 无法调用该 API。使用指南以 Agent skill 形式提供:openrouter-decisions。

  • Vercel AI Gatewayvercel.com

    Hosts Jev as typesafe-ai/jev

    以 typesafe-ai/jev 托管 Jev

  • stuntdgithub.com/bladedevoff

    Local proxy on the open Laya model that speaks the Jev System One API, records the app's Choice, Score and Noul answers from a Jev upstream, trains a per-question head, and serves it with a calibrated confidence threshold and fallback to the upstream

    基于开放 Laya 模型的本地代理,实现 Jev System One API;记录来自 Jev 上游的 Choice、Score 和 Noul 答案,为每个问题训练一个决策头,并以校准过的置信度阈值提供服务,低于阈值时回退到上游

SDKs & ClientsSDK 与客户端 31

Clients for the /v1/systemone API, official first. Most were written for TypeSafe's hosted Jev; the official TypeSafe SDKs also work against compatible servers such as Ollaya and Liquid d1. Community packages are not affiliated with any provider unless noted.

调用 /v1/systemone API 的客户端,官方在前。大多数是为 TypeSafe 托管的 Jev 编写的;官方 TypeSafe SDK 也能连接 Ollaya、Liquid d1 等兼容服务。除非另行说明,社区包与任何厂商都无隶属关系。

  • Python SDKgithub.com/typesafe-ai

    Official client. pip install typesafe-sdk. Docs: Python SDK. The community typesafe-ai package is a defensive redirect shim; install typesafe-sdk directly.

    官方客户端。pip install typesafe-sdk。文档:Python SDK。社区包 typesafe-ai 是为防占名而注册的重定向包;请直接安装 typesafe-sdk。

  • JavaScript / TypeScript SDKgithub.com/typesafe-ai

    Official client. npm install @typesafe-ai/sdk. Docs: JavaScript SDK.

    官方客户端。npm install @typesafe-ai/sdk。文档:JavaScript SDK。

  • System One adapter (Python)System One adapter(Python)github.com/typesafe-ai

    Official drop-in TypeSafeClient replacement backed by LLM APIs, for comparing Jev against chat models on the same questions. pip install system-one-adapter.

    官方提供的 TypeSafeClient 替身,后端走 LLM API,方便用同一套问题对比 Jev 与聊天模型。pip install system-one-adapter。

  • Vercel AI SDK providerai-sdk.dev

    @ai-sdk/typesafe-ai plus experimental_evaluate. Use typeSafeAi.evaluationModel('jev-latest') or the Gateway id typesafe-ai/jev.

    @ai-sdk/typesafe-ai + experimental_evaluate。可用 typeSafeAi.evaluationModel('jev-latest'),或 Gateway id typesafe-ai/jev。

  • Milvus Modelgithub.com/milvus-io

    Python reranker adapter that sends candidate documents as Jev Noul questions in one request, then sorts the returned scores and preserves original document indices

    社区 Python 重排适配器,一次请求用 Jev Noul 判断候选文档,再按分数排序并保留原始文档索引

  • Elixir SDKgithub.com/nshkrdotcom

    Community Hex package typesafe_sdk for system_one and model listing. Docs: HexDocs.

    社区 Hex 包 typesafe_sdk,支持 system_one 与模型列表。文档:HexDocs。

  • Jev (Elixir OTP)Jev(Elixir OTP)github.com/dannote

    Hex package jev: Jev as a peer GenServer; answers arrive as messages you pattern-match, with network-free tests

    Hex 包 jev:把 Jev 当成对等 GenServer,答案以消息到达再 pattern match,测试可以不碰网络

  • Ruby SDKgithub.com/joshmn

    Community Ruby 3.1+ client: Noul / Choice / Score, retries, model listing, thread-safe pooled HTTP. No async client.

    社区 Ruby 3.1+ 客户端:Noul / Choice / Score、重试、模型列表、线程安全连接池。没有异步客户端。

  • RubyLLM TypeSafegithub.com/kieranklaassen

    TypeSafe provider for RubyLLM 2 with offline model metadata and typed responses.

    RubyLLM 2 的 TypeSafe provider,带离线模型元数据和类型化响应。

  • typesafe-ai-railsgithub.com/GenieRobot

    Unofficial Rails integration on the community typesafe-sdk Ruby gem: configuration, usage/cost telemetry, and opt-in confidence policies

    非官方 Rails 集成,基于社区 Ruby gem typesafe-sdk:配置、用量/成本遥测,以及可选的置信度策略

  • Rust SDK (typesafe-ai-rs)github.com/gilljon

    Independent async and blocking client for System One.

    独立的异步 / 阻塞 System One 客户端。

  • TypeSafe AI for Rustgithub.com/Twister915

    Another Rust client: async + blocking transports, typed responses, observable retries.

    另一个 Rust 客户端:异步 + 阻塞传输、类型化响应、可观测重试。

  • typesafe-rsgithub.com/AbdelStark

    Latency-focused Rust transport SDK aiming for behavioral parity with the official clients.

    偏延迟的 Rust 传输 SDK,目标对齐官方客户端行为。

  • s1-rsgithub.com/AbdelStark

    Rust derive layer for Choice / Score / Noul, typed question sets, confidence gates, and network-free tests.

    Rust derive 层:Choice / Score / Noul、类型化问题集、置信度门控、无网络测试。

  • Advocaatgithub.com/pithings

    Small TypeScript client with tagged helpers for chances, choices, and scores.

    小型 TypeScript 客户端,给 chance / choice / score 打了 tagged helper。

  • Scala / ZIO SDKgithub.com/jamesward

    Community ZIO client with a small DSL for noul / choice / score.

    社区 ZIO 客户端,带 noul / choice / score 的小 DSL。

  • .NET SDKgithub.com/saibimajdi

    Community client for typed questions and confidence-scored answers.

    社区客户端,类型化问题 + 带置信度的答案。

  • PHP SDKgithub.com/Butochnikov

    Unofficial PHP client: typed DTOs, promises, and exceptions. Used by the Laravel package below.

    非官方 PHP 客户端:类型化 DTO、Promise 与异常。下面的 Laravel 包基于它。

  • Laravel TypeSafe Jevgithub.com/Butochnikov

    Unofficial Laravel 12/13 integration: config, facade, scoped DI, and a recording fake on the PHP SDK.

    非官方 Laravel 12/13 集成:配置、Facade、scoped DI,以及基于 PHP SDK 的 recording fake。

  • jev-gogithub.com/Gaurav-Gosain

    Unofficial Go client for typed judgments and calibrated probabilities. go get github.com/Gaurav-Gosain/jev-go.

    非官方 Go 客户端,返回类型化判断与校准概率。go get github.com/Gaurav-Gosain/jev-go。

  • Stumble/jev-gogithub.com/Stumble

    Unofficial dependency-free Go SDK for TypeSafe direct and Vercel AI Gateway, with typed questions, retries, an interactive CLI, and an installable agent skill

    非官方零依赖 Go SDK,支持 TypeSafe 直连和 Vercel AI Gateway,并提供类型化问题、重试、交互式 CLI 和可安装的 agent skill

  • jevclientgithub.com/AboveColin

    Unofficial async Python client (pip install jevclient). Typed Noul / Choice / Score helpers, separate from the official typesafe-sdk.

    非官方异步 Python 客户端(pip install jevclient)。带 Noul / Choice / Score helper,与官方 typesafe-sdk 不是同一个包。

  • LlamaIndex Jevgithub.com/WiktorB2004

    Unofficial LlamaIndex reranker (JevRerank) and router (JevSingleSelector / JevMultiSelector) on the official Python SDK

    非官方 LlamaIndex 重排序(JevRerank)与路由(JevSingleSelector / JevMultiSelector),基于官方 Python SDK

  • Swift SDKgithub.com/ainame

    Unofficial Swift 6.4 client aligned with the Python SDK 0.6.0 API, including Linux

    非官方 Swift 6.4 客户端,对齐 Python SDK 0.6.0 API,含 Linux

  • TypeSafe AI Swift SDKgithub.com/alterhq

    Unofficial dependency-free Swift 6 client for Choice / Score / Noul, with strict concurrency, configurable authentication and retries, and network-free tests

    非官方零依赖 Swift 6 客户端,支持 Choice / Score / Noul、严格并发、可配置鉴权与重试,以及无网络测试

  • System One Foundation Modelsgithub.com/peterfriese

    Unofficial Swift 6 bridge mapping Apple's @Generable types to Noul, Choice, and Score, with hosted Jev, HTTP Laya, and on-device Core ML Laya backends plus confidence routing

    非官方 Swift 6 桥接库,将 Apple 的 @Generable 类型映射为 Noul、Choice 和 Score,支持托管 Jev、HTTP Laya 和设备端 Core ML Laya,并提供置信度路由

  • discerngithub.com/doeixd

    Unofficial Effect library: Choice / Noul / Score answers become typed patterns with an explicit Uncertain branch you must handle, plus routable procedures, with recording, replay, caching and call budgets as DecisionModel middleware. Provider-neutral; reaches Jev through @effect/ai-typesafe

    非官方 Effect 库:把 Choice / Noul / Score 答案变成带类型的模式匹配,Uncertain 是必须显式处理的分支,并支持可路由的 procedure;录制、回放、缓存与调用预算都做成 DecisionModel 中间件。不绑定供应商,通过 @effect/ai-typesafe 接入 Jev

  • kojev (Kotlin Multiplatform)kojev(Kotlin Multiplatform)github.com/ItisNoMatter

    Community client for JVM, Android, and iOS. Choice and Score answers come back as your own enums; one typed way to read them, no default thresholds. Maven Central: io.github.itisnomatter:kojev:0.1.0.

    社区客户端,支持 JVM、Android 和 iOS。Choice 与 Score 的答案直接回到你自己的 enum;只有一种带类型的读取方式,不设默认阈值。Maven Central:io.github.itisnomatter:kojev:0.1.0。

  • jev4kgithub.com/pambrose

    Unofficial JVM Kotlin client: Choice, Score, and Noul as a DSL, with answers read back as typed values including enums. Maven Central: com.pambrose:jev4k

    非官方 JVM Kotlin 客户端:用 DSL 写 Choice、Score 和 Noul,答案以类型化的值读回,包括 enum。Maven Central:com.pambrose:jev4k

  • hunchgithub.com/steven-shoemaker

    Unofficial Python library, with a TypeScript port, that turns Choice / Score / Noul into functions over lists and DataFrames (classify, score, check, where, extract, pick, rank, verify), with deduplication, caching, and optional escalation of unsure rows to an LLM that must pick from the same labels

    非官方 Python 库(另有 TypeScript 版本),把 Choice / Score / Noul 变成作用于列表和 DataFrame 的函数(classify、score、check、where、extract、pick、rank、verify),支持请求去重、缓存,并可把不确定的行交给 LLM 在同一组标签中复核

  • JevT++github.com/wiatrM

    Unofficial C++20 library with compile-time enum schemas, typed decisions and abstention, local Laya backends, and an optional TypeSafe System One HTTP client; remote tests use mocks and loopback HTTP, not live-provider validation

    非官方 C++20 库,提供编译期枚举模式、类型化决策与弃权机制、本地 Laya 后端及可选的 TypeSafe System One HTTP 客户端;远程测试使用模拟响应和本地 HTTP 服务,尚未验证真实服务兼容性

Applications应用 64

Open-source products and demos that put a decision model in a real loop. Most use Jev today.

把决策模型放进真实循环里的开源产品与 demo。目前大多使用 Jev。

  • MemSearchgithub.com/zilliztech

    Markdown memory for coding agents with an optional Jev Noul reranker and a published English/Chinese retrieval evaluation; community integration, not an official TypeSafe SDK

    面向编程 Agent 的 Markdown 记忆系统,提供可选的 Jev Noul 重排器与公开的中英文检索评测;属于社区集成,并非 TypeSafe 官方 SDK

  • Jev-Memgithub.com/libingzheren

    Unofficial agent memory system where Jev or local Laya controls memory organization, query routing, retrieval budgets, candidate scoring, and stopping while an LLM writes answers; its paper reports LoCoMo results

    非官方 Agent 记忆系统:Jev 或本地 Laya 负责记忆组织、查询路由、检索预算、候选评分与停止判断,LLM 负责生成答案;论文报告了作者在 LoCoMo 上的评测结果

  • Jev RAGgithub.com/aifabrice

    Unofficial local-first knowledge search app using SQLite BM25 or agent-planned lexical retrieval, Jev Noul judgments for evidence reranking, and a published reproducible NFCorpus evaluation; no vector database is required by default

    非官方本地优先知识检索应用:使用 SQLite BM25 或 Agent 规划的关键词检索召回候选,以 Jev Noul 判断重排证据,并公开可复现的 NFCorpus 评测;默认不需要向量数据库

  • Jev Deep Researchgithub.com/sunyasheng

    Unofficial experimental research agent: Jev Choice locates evidence and Noul checks its presence across document regions in parallel; Pi-Serini returns original passages for GPT to verify and continue

    非官方实验性研究 Agent:Jev 通过 Choice 定位证据、Noul 判断证据是否存在,并行处理文档区域,再由 Pi-Serini 返回原文供 GPT 核查并继续研究

  • Jev Ultrafastgithub.com/browser-use

    Browser agent from Browser Use. Jev picks an operation and a DOM element in one request; a small LLM writes text only for TYPE_TEXT. Zürich → London on Google Flights in ~7s. Library, local inspector, and measurements included.

    Browser Use 的浏览器 Agent。一次请求里由 Jev 选出操作和 DOM 元素;只有 TYPE_TEXT 才让小模型写字。Google Flights 苏黎世 → 伦敦约 7 秒。含库、本地 inspector 与测时。

  • Jev Socialgithub.com/socai-io

    Browser-grounded social research: Jev selects bounded Instagram, TikTok, and LinkedIn search/read operations, socai executes them in the user's Chrome, and reports cite the captured posts, comments, and video evidence; unofficial community project

    浏览器实证社媒调研:Jev 选择受限的 Instagram、TikTok 与 LinkedIn 搜索/读取操作,socai 在用户 Chrome 中执行,报告仅引用捕获的帖子、评论与视频证据;非官方社区项目

  • Jev Web Analyzergithub.com/replynodes

    Community project that analyzes a public SaaS landing page as clean Markdown and asks Jev ten bounded Choice questions about first-visit understanding, including the first change to make.

    社区项目:把公开 SaaS 落地页提取为干净 Markdown,再让 Jev 提出十个有界的 Choice 问题,判断首次访问者能理解什么,包括最先要改的地方。

  • jev-align (Sutro)github.com/sutro-sh

    Unofficial active-learning CLI that evaluates CSV, Parquet, and JSONL rows with Jev, asks people to label uncertain and audit samples, and uses GEPA to propose improved definitions

    非官方主动学习 CLI:用 Jev 评估 CSV、Parquet 和 JSONL 数据,让人工标注不确定样本与审计样本,并用 GEPA 提议改进后的定义

  • JevSpangithub.com/lzq-0529

    Unofficial zero-shot named entity recognition for Chinese and English: code enumerates candidate spans at punctuation, Jev Choice questions nominate, verify, and fix the boundary of each entity, and every entity keeps its probability and a decision trace

    非官方中英文零样本命名实体识别:代码按标点列出候选片段,由 Jev 的 Choice 问题提名、核验并确定每个实体的边界,每个实体都保留概率和完整的决策过程

  • Jev for Chromegithub.com/chy4pro

    Unofficial Chrome extension (Manifest V3) port of Jev Ultrafast: Jev picks the operation and DOM element in one request, a small text model writes typed values, and it runs in the user's own tabs through OpenRouter, TypeSafe or Cloudflare; includes a 17-task headless-Chromium suite with recorded traces.

    Jev Ultrafast 的非官方 Chrome 扩展(Manifest V3)移植:Jev 一次请求同时选出操作和 DOM 元素,只有打字时才调用小文本模型,直接跑在用户自己的标签页里(OpenRouter / TypeSafe / Cloudflare 三种渠道);附 17 个任务的 headless Chromium 测试套件和完整轨迹(仓库 docs/ 目录,同一套任务多轮 13–14/17)。

  • jev-egogithub.com/romaluev

    Browser agent on ego lite: one TypeSafe request picks operation + indexed element; agent-facing observe/act/suggest/step CLI

    ego lite 上的浏览器 Agent:一次 TypeSafe 请求选出操作和编号元素;面向 Agent 的 observe/act/suggest/step CLI

  • jev-browsergithub.com/Ying-Kai-Liao

    Unofficial browser automation: an LLM plans the outcome, Jev decides each click/type on a Playwright snapshot (~300 ms/call). Ships as a library, CLI, and MCP server (npx -y -p jev-browser jev-browser-mcp).

    非官方浏览器自动化:LLM 规划目标,Jev 在 Playwright 快照上决定每次点击/输入(约 300 ms/次)。提供库、CLI 与 MCP 服务(npx -y -p jev-browser jev-browser-mcp)。

  • Sedumgithub.com/sedum-dev

    Unofficial open-source end-to-end AI testing tool for Playwright: in goal mode Jev picks each next action and its target, plain-English steps use a Jev Choice to find their element, and verify claims are two Nouls (holds, contradicted), while clicks, waits, verdicts and exit codes stay deterministic

    非官方的开源 Playwright 端到端 AI 测试工具:目标模式下由 Jev 选出每一步操作及其目标,纯英文步骤用 Jev Choice 找到对应元素,验证断言是两个 Noul(holds、contradicted),点击、等待、判定和退出码均由确定性代码完成

  • typesafe-computer-usegithub.com/awlevin

    macOS computer-use loop: OCR the screen, Jev classifies the next action, then click. About $0.0002/step.

    macOS computer-use:OCR 屏幕,Jev 分类下一步动作再点击。约 $0.0002/步。

  • Yappyyappy.biz

    macOS voice agent (closed source, public write-up with measurements). On its hosted plan Jev picks the operation and target control from the window's accessibility table each step; a chat model writes text only for typing, and the full agent takes over when confidence drops. Author-reported: 275–690 ms per decision, $0.003 for five.

    macOS 语音 Agent(闭源,附公开测量数据)。在其托管方案上,Jev 每一步从窗口的无障碍控件表中选择操作与目标控件;只有输入文本时才调用聊天模型,置信度下降时交回完整 Agent。作者报告:每次决策 275–690 ms,五次共 $0.003。

  • Mobile Jevgithub.com/droidrun

    Android agent on Mobilerun: Jev decides each tap. Opens Uber, SFO → Golden Gate, payment screen in ~21s / 9 actions. Live studio, CLI, and traces. No ADB.

    Mobilerun 上的 Android Agent:每次点击由 Jev 决定。打开 Uber,旧金山机场 → 金门大桥,约 21 秒 / 9 步到支付页。含实时 studio、CLI 与 traces。不需要 ADB。

  • Uncluttergithub.com/kitze

    Chrome / Firefox extension: Jev classifies nonessential page elements; local template rules hide them on later visits.

    Chrome / Firefox 扩展:Jev 标出页面上不重要的元素,本地按页面模板记住并在下次访问时藏起来。

  • jevMailgithub.com/ilyamk

    Unofficial open-source Gmail AI spam filter, auto-labeler, and inbox organizer: Jev understands each email's intent to apply custom labels and optionally archive high-confidence unwanted mail

    非官方开源 Gmail AI 垃圾邮件过滤、自动标签与收件箱整理工具:Jev 理解每封邮件的意图,应用自定义标签,并可自动归档高置信度的无用邮件

  • HA-Jevgithub.com/AboveColin

    Unofficial Home Assistant integration: typed questions about entity state become sensors and automation actions, with a target picker that builds the state from the user's own entities and usage, cost, and daily-budget entities alongside the answers

    非官方 Home Assistant 集成:把关于实体状态的类型化提问变成传感器与自动化动作;可直接选取实体、设备或区域来构造 state,并附带用量、成本与每日 token 预算实体

  • Everygithub.com/sufianetaouil

    Semantic code-search CLI: a yes/no question against every function, ranked by Noul probability.

    语义代码搜索 CLI:对每个函数问 yes/no,按 Noul 概率排序。

  • JevPDFgithub.com/kylemclaren

    Unofficial Ctrl+F by meaning for PDFs: pdf.js extracts lines in the browser, Jev answers one Noul per line on whether it answers the query, and matching lines light up ranked by probability

    非官方的 PDF 语义版 Ctrl+F:pdf.js 在浏览器中逐行提取文本,Jev 对每一行问一个 Noul(这一行是否回答了问题),命中的行按概率排序并高亮

  • blinkgithub.com/ellipsis-dev

    Codebase search: an ensemble of walkers asks Jev which file answers a natural-language query

    代码库搜索:一组 walker 并行走文件系统,由 Jev 判断哪个文件能回答自然语言查询

  • Jev Searchgithub.com/superagents-lab

    Unofficial web search app using Jev's Choice and Noul judgments to select sources, time ranges, and query candidates, then rank results retrieved through Search1API

    非官方网页搜索应用:用 Jev 的 Choice 和 Noul 判断选择来源、时间范围和候选查询词,再对 Search1API 返回的结果进行相关性排序

  • Jev Reranker (Rust CLI)github.com/shinpr

    Unofficial JSON-in/JSON-out CLI that uses Jev Noul judgments to rerank search results, filter documents without usable evidence, or extract query-specific passages

    非官方 JSON 输入/输出 CLI:使用 Jev 的 Noul 判断重排搜索结果、过滤不含可用证据的文档,或提取与查询相关的原文片段

  • jevsearchgithub.com/kylemclaren

    Unofficial shadcn/ui site-search block: keyword hits appear on the first keystroke, then one Jev request re-ranks the top 20 with a Noul per page, a Choice for the best answer, and a Noul for whether any page answers

    非官方 shadcn/ui 站内搜索组件:首次按键即显示关键词结果,随后用一次 Jev 请求对前 20 条重排:每个页面一个 Noul,一个 Choice 选出最佳答案,再用一个 Noul 判断是否有页面能回答

  • jev-research-pipelinegithub.com/shimo4228

    Unofficial experimental research monitor: Jev screens papers and other sources against each open research question (Noul gates, Score dimensions), code applies thresholds, and Qwen writes question-centric notes into an Obsidian vault

    非官方实验性研究监测流水线:Jev 按研究问题筛选论文等来源(Noul 门控 + Score 评分),阈值由代码判定,Qwen 为通过筛选的来源撰写按问题组织的 Obsidian 笔记

  • neo4jevgithub.com/jexp

    Neo4j graph navigation: at each node Jev chooses which relationship to follow, with beam search over log-probabilities

    Neo4j 图导航:每个节点上由 Jev 选择跟哪条关系走,并对 log 概率做 beam search

  • hono-jev-routergithub.com/yusukebe

    Experimental Hono router: Jev matches an incoming request to a plain-language route description

    实验性 Hono 路由器:用自然语言描述路由,由 Jev 匹配进来的请求

  • sqlite3-jevgithub.com/mattn

    SQLite C extension: jev_noul / jev_choice / jev_score as SQL functions via libcurl

    SQLite C 扩展:把 jev_noul / jev_choice / jev_score 做成 SQL 函数,只依赖 libcurl

  • jevqlgithub.com/kylemclaren

    Unofficial psql-shaped CLI and Go/TypeScript/Python SDKs for vanilla Postgres: jev() / jev_prob / jev_choice / jev_score in plain SQL with no extension, the SQL runs on the server and Jev judges the surviving rows in batches

    非官方类 psql 命令行工具与 Go/TS/Python SDK:无需扩展即可在原生 Postgres 中使用 jev() / jev_prob / jev_choice / jev_score,SQL 在服务端执行,剩余行由 Jev 批量判断,结果会缓存

  • jev-resiliencegithub.com/Vicente-MD

    Unofficial Spring WebFlux starter: a semantic circuit breaker that uses Jev to catch silent HTTP 200 failures

    非官方 Spring WebFlux starter:语义熔断器,用 Jev 抓 HTTP 200 里的静默失败

  • tripwiregithub.com/noelzappy

    Unofficial AI SDK middleware and OpenAI-compatible proxy: seven Jev checks on every LLM response in ~100 ms, confidence-gated

    非官方 AI SDK middleware 与 OpenAI 兼容代理:约 100 ms 内对每条 LLM 回复做七项 Jev 检查,按置信度门控

  • ProgressGategithub.com/AshutoshVJTI

    Detects semantic stagnation in agent loops: Jev judges the trajectory; code returns CONTINUE / WARN / REPLAN / HALT

    检测 Agent 循环里的语义停滞:Jev 评判轨迹,代码返回 CONTINUE / WARN / REPLAN / HALT

  • jev-harnessgithub.com/AntonioCoppe

    Unofficial production layer around Jev: policy, confidence gate, shadow mode, recipes, and an eval CLI

    非官方生产层:策略、置信度门控、影子模式、配方和 eval CLI

  • jev-treegithub.com/reachjalil

    Recursive Choice over a taxonomy so catalogs larger than Jev's 255-option cap still fit

    在分类树上递归做 Choice,突破 Jev 单次最多 255 个选项的上限

  • jev-shell-historygithub.com/mrnugget

    Fish-style zsh autosuggestions: Jev ranks recent history as you type

    Fish 风格的 zsh 自动补全:输入时由 Jev 给近期历史排序

  • Supercovgithub.com/supercorp-ai

    Code quality and test coverage for coding agents: Jev scores each source file so the agent knows what to fix first

    面向编码 agent 的代码质量与测试覆盖率工具:Jev 给每个源文件打分,agent 就知道该先修哪里

  • jev-lintgithub.com/ckorhonen

    Unofficial fuzzy linter for Claude Code and Codex that uses Jev to flag team-rule violations at edit time, with configurable rule packs and repository-specific rules so agents can fix issues before code review

    非官方 Claude Code 和 Codex 模糊代码检查工具,使用 Jev 在编辑时发现违反团队规则的代码,并支持可配置的规则包和仓库专属规则,让智能体能在代码审查前修正问题

  • Jev Reviewgithub.com/devagrawal09

    Staged code-review workflow and local dashboard driven by focused Jev calls.

    分阶段代码审查工作流 + 本地 dashboard,由聚焦的 Jev 调用驱动。

  • Foremangithub.com/thruwire

    Software-factory loop: Codex implements; Jev independently judges completeness, tests, and whether a human is needed.

    软件工厂循环:Codex 写实现,Jev 独立判断是否做完、测试够不够、要不要人来看。

  • Jev Dronegithub.com/RomanSlack

    MuJoCo quadrotor: control and safety stay in code; Jev handles slower tactical judgments.

    MuJoCo 四旋翼:控制和安全留在代码里,Jev 做较慢的战术判断。

  • Jev Plays StarCraftgithub.com/phyous

    Structured-state harness for the original StarCraft shareware campaign, with verified run and probability traces.

    原版星际争霸共享战役的结构化 state harness,带验证跑次和概率轨迹。

  • Jev × Civilization IIgithub.com/phyous

    Original Civ II in a browser; Jev chooses empire, city, research, and unit actions. Experimental; no verified win yet

    浏览器里跑原版文明 II;Jev 选帝国、城市、科研和单位动作。实验性,尚未验证通关

  • Jev Tradegithub.com/aowang-ai

    Live Hyperliquid desk: each tick Jev answers Choice questions for long/short, open/close/hold, and leverage; code places or pulls the quote. Dry-run by default; a live key sends real orders. Demo: jev-trade.com.

    Hyperliquid 实盘桌面:每个 tick 由 Jev 用 Choice 回答多空、开平或 hold、以及杠杆;下单和撤单由代码执行。默认 dry-run;配置私钥后会真下单。在线:jev-trade.com。

  • Jev Tradergithub.com/jarrodwatts

    One buy/sell decision per Monad block on Kuru's MON-USDC book. Live demo: jev-trader.vercel.app.

    每个 Monad 区块对 Kuru 的 MON-USDC 下一笔买卖。在线 demo:jev-trader.vercel.app。

  • Human Compilergithub.com/asfarsadewa

    Paste corporate prose; Jev scores passive-aggression, urgency, and information density, then code emits rustc-style diagnostics. Live: human-compiler.asfarlab.fun.

    粘贴职场废话,Jev 打被动攻击 / 紧急感 / 信息密度,代码按 rustc 风格报诊断。在线:human-compiler.asfarlab.fun。

  • Jev Wrappedgithub.com/gaborishka

    Telegram channel X-ray: Jev judges up to 1,500 public posts from a channel's last year with one Choice over ten kinds of post and three Noul checks for paid ad, clickbait and emotional pressure; code draws the monthly mix on a shareable card and links the highest-scoring posts. Live: wrapped.ivanhabor.com.

    Telegram 频道透视:Jev 逐条判断公开频道近一年最多 1,500 条帖子,用一个 Choice 从十种帖子类型中选一种,再用三个 Noul 判断是否为付费广告、标题党和情绪施压;代码把每月构成画成可分享的卡片,并附上得分最高的帖子链接。在线:wrapped.ivanhabor.com。

  • JEVMETERgithub.com/ChetasLua

    Live Jev meter on any video: every sentence scored, rendered as a 16:9 edit. Demo: Chetaslua.

    给任意视频挂上实时 Jev 仪表:逐句打分,导出 16:9 成片。演示:Chetaslua。

  • jev-audio-beepergithub.com/santos-sanz

    Low-latency audio insult detector: Jev decides, ffmpeg beeps in ~466 ms without rewriting the rest of the track.

    低延迟音频脏话检测:Jev 判定后 ffmpeg 在约 466 ms 内叠一声 beep,不改其余音轨。

  • jev-askable-armgithub.com/TarunTomar122

    Zero-shot English goals on a simulated Franka. Jev chains hardcoded primitives.

    仿真 Franka 上用英文目标做 zero-shot;Jev 把硬编码原语串起来。

  • jev-codex-routergithub.com/0xNatoshi

    Per-turn Codex routing: Jev picks model, thinking depth, and speed mode.

    Codex 每轮路由:Jev 选模型、思考深度和速度模式。

  • Codex Jev Routergithub.com/suenot

    Codex subagent routing: Jev chooses a model and reasoning effort from typed Choice and Noul answers; code applies confidence gates and falls back to Sol.

    Codex 子代理路由:Jev 用 Choice 和 Noul 选择模型与推理档位,代码检查置信度,不确定时回退到 Sol。

  • Jev Auto Routergithub.com/miniLV

    Unofficial model-routing prototype that uses Jev to choose a model and reasoning effort for each call, with a local Responses proxy and independent task verification

    非官方模型路由原型:用 Jev 为每次调用选择模型与推理档位,并通过本地 Responses 代理转发、独立验收任务

  • jev-routergithub.com/gargpratyush

    Per-turn routing for Claude Code and Codex: Jev sends simple work to the fast tier and hard work to the strong tier. npm i -g jev-router.

    Claude Code 与 Codex 的每轮路由:简单活走快档,难活走强档。npm i -g jev-router。

  • jev-secret-detectiongithub.com/teyhouse

    Secret-in-diff detector with repeatable Jev verdicts.

    用 Jev 扫 diff 里的密钥,结果可复现。

  • commit-minergithub.com/devanshbatham

    Rust CLI that classifies commit diffs with Jev: bug fixes, security/CWEs, and change types. HTML/CSV reports.

    用 Jev 给 commit diff 分类的 Rust CLI:修 bug、安全/CWE、变更类型。可出 HTML/CSV 报告。

  • jev-eval-agentgithub.com/vinilana

    Public eval harness for early Jev tests.

    早期 Jev 测试的公开评测 harness。

  • Jev Logsgithub.com/reachjalil

    OpenTelemetry log triage: Jev scores diagnostic value and priority before an expensive LLM looks at the archive.

    OpenTelemetry 日志分流:先让 Jev 打诊断价值和优先级,再决定要不要花 LLM。

  • Smart home assistant demodocs.typesafe.ai

    Official interactive demo of speculative fan-out

    官方互动 demo,演示 speculative fan-out

  • jev.nvimgithub.com/valentynkit

    Neovim plugin that splits the buffer into functions with Treesitter, scores each against a plain-language question with Jev, and ranks answers by probability in the quickfix window.

    Neovim 插件:用 Treesitter 把缓冲区拆成函数,向每个函数提出一个自然语言问题让 Jev 打分,结果按概率排进 quickfix 列表。

  • jev-skipgithub.com/valentynkit

    Browser extension that reads the YouTube caption track and paints a per-segment sponsor probability on the seek bar before the intro ends, with no crowd database, reporting catching 77% of SponsorBlock's sponsor seconds across 23 videos at $0.0008 a video.

    浏览器扩展:读取 YouTube 字幕轨道,在片头结束前就把每段视频的赞助概率画到进度条上,不依赖众包数据库,据报告在 23 个视频上抓住了 SponsorBlock 77% 的赞助时长,每个视频约 0.0008 美元。

  • JevBystandergithub.com/Nisaka520

    Android accessibility app that reads the visible WeChat chat screen and sends one batched Jev request (10-way intent Choice, 9-way emotion distribution, 0-3 urgency Score, 11-way reply-posture Choice) to show exactly three toasts - no generated reply text, no input injection, no screenshot or OCR; a local contact table supplies relation aliases as state context.

    安卓无障碍应用:读取微信当前可见的聊天文字,一次批量 Jev 请求(10 类意图 Choice、9 类情绪分布、0–3 着急程度 Score、11 类回复姿态 Choice)后只弹三条 Toast;不生成回复文案、不注入输入、不截屏也不做 OCR;本地联系人表把关系别名放进 state

  • Jev Chat Assistantgithub.com/jev-chat

    Unofficial Android chat copilot: an accessibility service reads the visible QQ, X, or Feishu/Lark chat (Feishu text via on-device OCR), one batched Jev request asks intent, needs, and next-action Choice questions, a danger Score, and three Noul checks, a separate chat model drafts three replies that one more Choice ranks, and code fills the picked reply into the input box without sending

    非官方安卓聊天副驾:用无障碍服务读取 QQ、X、飞书当前可见的对话(飞书正文走本机离线 OCR),一次批量 Jev 请求问意图、对方需要什么、下一步动作 3 个 Choice,危险等级 Score 和 3 个 Noul;另一个聊天模型起草 3 条候选,再用一个 Choice 排序;代码只把选中的回复填进输入框,不代发

  • Paper Radargithub.com/Eliot5566

    Unofficial daily arXiv and bioRxiv radar. Jev answers one Noul per plain-English interest for every new paper; code applies the thresholds and publishes a page and RSS feed from GitHub Actions. Live demo needs no key.

    非官方每日 arXiv 和 bioRxiv 论文雷达:Jev 对每篇新论文按每条自然语言兴趣给出一个 Noul 判断,代码应用阈值,再由 GitHub Actions 发布网页和 RSS 订阅源。公开运行页面 无需密钥即可查看

Demos & GamesDemo 与游戏 30

Toys, live sites, and realtime agents.

玩具、小站和实时 Agent。

  • Yes / Noyesno.coderai.dev

    Free no-signup Noul demo. Ask a question, get yes / no / maybe, with web search when needed.

    免登录 Noul demo。问一句,得到 yes / no / maybe,必要时联网检索。

  • TypeSafe AdBlockgithub.com/realZachi

    Chrome extension demo: Jev judges candidate DOM elements and removes likely ads, with BYOK and no backend; each page consumes API tokens and the author documents missed ads and mistaken removals

    Chrome 扩展 demo:Jev 判断候选 DOM 元素并移除疑似广告,BYOK、无后端;每页消耗 API token,作者说明了漏删广告和误删元素的局限

  • Jev Tetrisjev-omega.vercel.app

    Jev picks rotation and column from holes, stack height, and bumpiness.

    Jev 根据空洞、堆高、起伏选旋转和落点列。

  • Jev Pac-Manjev-pacman.ephraimduncan.com

    Maze as JSON; Jev picks the turn at each junction in realtime.

    迷宫做成 JSON,每个路口由 Jev 选转向,实时玩。

  • Jev Chessjevchess.com

    One shared board, the internet vs Jev; every legal move is one Choice question, probabilities shade the pieces, live calibration panel scores every move.

    全网对 Jev 的一盘共享棋;每个合法着法都是一个 Choice 问题,概率给棋子上色,实时校准面板为每一步打分。

  • Chess with Jevchriswijnia.com

    Chess and Chess960 in the browser: code works out each legal move's facts and Jev picks one per turn as a single Choice, with its candidates drawn as arrows (source)

    浏览器里的国际象棋与 Chess960:代码算出每个合法着法的事实,Jev 每回合以一次 Choice 选一步,候选着法画成箭头(源码)。

  • typesafe-mariogithub.com/fhshaik

    Super Mario Bros. from structured emulator state.

    从结构化模拟器状态玩超级马里奥。

  • jev-doom-agentgithub.com/lukaske

    Browser-native Doom with Chocolate Doom WASM, spatial state, and live decision telemetry.

    浏览器里的 Doom(Chocolate Doom WASM),空间状态 + 实时决策遥测。

  • jev-gomokugithub.com/mizchi

    MoonBit client plus Jev-vs-Jev gomoku; write-up: jev 同士に五目並べで対戦させた.

    MoonBit 客户端 + Jev 对打五子棋。文章:jev 同士に五目並べで対戦させた。

  • jev-t-rex-runnergithub.com/joshlarsen

    Chrome dinosaur game played by Jev.

    Chrome 小恐龙由 Jev 来跳。

  • snake-jevgithub.com/siroccomask

    Snake: hundreds of typed direction decisions per run.

    贪吃蛇:每局几百次类型化转向决策。

  • Jev Guardguard-jev.vercel.app

    Comment-moderation playground.

    评论审核 playground。

  • jev-fitjev-fit.com

    Paste a software idea; Jev answers a fixed typed rubric in one call and the page says plain code, Jev, or a reasoning LLM, with probabilities. Unofficial, closed source, free page and API.

    粘贴一个软件想法;Jev 在一次调用中回答一套固定的类型化问题,页面给出结论:普通代码、Jev 或推理型 LLM,并附概率。非官方,闭源,页面和 API 免费。

  • Hollow Creekhollow-creek-sigma.vercel.app

    Village NPCs that judge you each tick (what to do, how they feel) instead of chatting.

    村庄 NPC 每个 tick *评判*你(做什么、对你什么感觉),而不是聊天。

  • Jev mood demojev-demo.vercel.app

    Talk nicely or nastily over time; structured state tracks mood.

    长时间对它好或坏,结构化 state 跟踪心情。

  • Jev Roomjev-room.moe136231.chatgpt.site

    One sentence → six room settings. Jev chooses, the app renders.

    一句话 → 六个房间设定。Jev 选,应用渲染。

  • 1 Million Emojischriswijnia.com

    A shared 1000 × 1000 emoji canvas, live for everyone; after each stroke Jev picks a square next to it and its emoji as one Choice (source)

    一块人人实时共享的 1000 × 1000 emoji 画布;每一笔之后,Jev 用一次 Choice 选定旁边的一格及其 emoji(源码)。

  • Jevviechriswijnia.com

    Page companion: the page offers its actions as WebMCP tools, and one Jev Choice picks the action a visitor's request means (with a Choice per argument asked alongside), asking back when the top two are close; a voxel character then hops to the button and does it (source).

    页面小助手:页面把操作暴露为 WebMCP 工具,一次 Jev Choice 判定访客请求对应哪个操作(参数也各用一次 Choice),前两名接近时会再问一句;体素角色再跳到按钮上执行(源码)。

  • TypeSafe Typewritertypesafe-demo.val.run

    Live Val Town demo: 16 typed judgments update as you type. Launch post: Steve Krouse.

    Val Town 在线 demo:打字时 16 条类型化判断实时更新。发布帖:Steve Krouse。

  • got-jevgithub.com/phureewat29

    Game of Thrones roleplay as Jon Snow. A story model writes the scene; Jev answers where he is, how much danger, and what should play under it.

    权力的游戏角色扮演:你是琼恩·雪诺。故事模型写下一场,Jev 回答他在哪、有多危险、该配什么音乐。

  • Little Airwaysgithub.com/lbotinelly

    Toy archipelago ATC: Jev judges divert / emergency / who lands first from each plane's local state, ~150 ms.

    玩具群岛空管:每架飞机只看见自己附近,Jev 判断备降 / 紧急 / 谁先落地,约 150 ms。

  • jev-plays-pokemon-redgithub.com/valentynkit

    Pokemon Red on PyBoy where deterministic code owns the route and arithmetic and Jev picks only at branches, with every battle turn's faint prediction scored by Brier against the emulator's RAM state.

    基于 PyBoy 的精灵宝可梦红版:路线和数值运算都由代码掌控,Jev 只在分支点做选择,每回合战斗都会记录一次用 Brier 分数对照 RAM 状态检验的濒死预测。

  • jev-canvasgithub.com/gaborishka

    Draw on a tldraw canvas with your voice and a webcam-tracked finger; Jev decides action, target and place on every partial transcript. English and Ukrainian commands.

    用语音和摄像头追踪的手指在 tldraw 画布上绘图;Jev 在每段实时转写上决定动作、目标和位置。支持英语和乌克兰语指令。

  • Jevtowngithub.com/gaborishka

    A town of 10,000 computed personas reads your post, listing, product or headline. Jev scores who the text is for to pick the first 600 readers and answers one Choice per persona for its reaction; code sends the text to the next wave only while glad readers outnumber annoyed ones by at least a tenth of the wave. Live: jevtown.ivanhabor.com.

    由 10,000 个计算生成的人物组成的小镇,阅读你的帖子、分类广告、产品或标题。Jev 判断文本适合哪些人,选出最先的 600 位读者,并为每个人物回答一个 Choice 给出反应;只有高兴的读者比反感的读者至少多出这一波人数的十分之一,代码才把文本送往下一波。在线:jevtown.ivanhabor.com。

  • sudoku-vs-jevgithub.com/zebedelu

    Terminal Sudoku where Python owns the rules and Jev picks one move per turn, steady while forced moves exist and shaky once it has to guess.

    终端数独:Python 掌握规则,Jev 每回合选择一步,在存在必走步时表现稳健,一旦需要猜测则表现不稳。

  • chess-vs-jevgithub.com/zebedelu

    Pygame chess where python-chess owns the rules and Jev picks one legal move per turn, playable Human vs Human, Human vs Jev, or Jev vs Jev.

    Pygame 国际象棋:python-chess 掌握规则,Jev 每回合选择一个合法走法,支持人 vs 人、人 vs Jev 和 Jev vs Jev。

  • JevsBistrogithub.com/andrewsilber

    Deterministic 3D restaurant sim that replays the same dinner service to compare rule-based, camera-assisted, and Jev-planned waiters, logging each decision's state, options, confidence, and latency.

    确定性的 3D 餐厅模拟:重放同一场晚餐服务,对比规则驱动、摄像头辅助和由 Jev 规划的服务员,并记录每次决策的状态、选项、置信度和延迟。

  • jev-asks-until-suregithub.com/mintannn

    Twenty questions where confidence sets the stopping rule: Jev commits, hedges, or refuses to guess, and the UI narrates every judgment. Live: jev.mintan.org.

    用置信度决定还要问几题的二十问游戏:Jev 会断言、含糊其辞,或者干脆拒绝作答,界面同步播报每一次判定。在线:jev.mintan.org。

  • Jev × 2048jev-2048-ultra.vercel.app

    A web lab where Jev is the 2048 decision engine, showing each move's probability distribution, confidence, latency, and token cost so you can watch how context design shapes the decision model.

    一个把 Jev 当作 2048 决策引擎的网页实验台,展示每一步的概率分布、置信度、延迟与 token 消耗,观察上下文设计如何影响决策模型。

  • Book Auroragithub.com/dani1005

    Jev reads a whole novel in seconds: each passage gets nine emotion scores plus intensity in one call, and every passage becomes a feathered row of colour. Frankenstein is 601 passages, 6,010 typed decisions, about 25 s and 3 cents; exports a poster.

    让 Jev 几十秒读完一整本小说:每段文字一次调用返回九种情绪打分和强度,每段变成一行羽化的色带,整本书就是一幅极光。《弗兰肯斯坦》601 段、6010 次类型化判断,约 25 秒、3 美分,可导出海报。

Agent ToolsAgent 工具 38

Tools that expose decision models to coding agents and MCP clients.

把决策模型接到编程 Agent 与 MCP 客户端上的工具。

  • TypeSafe agent skillgithub.com/typesafe-ai

    Official skill: primitives, patterns, and how to structure evaluations. Claude Code: claude plugin marketplace add typesafe-ai/skills then claude plugin install typesafe@typesafe-ai. Other agents: npx skills add typesafe-ai/skills --skill typesafe-ai.

    官方技能包:原语、模式、如何组织 evaluation。Claude Code:claude plugin marketplace add typesafe-ai/skills,再 claude plugin install typesafe@typesafe-ai。其他 Agent:npx skills add typesafe-ai/skills --skill typesafe-ai。

  • fast-jev-compactiongithub.com/tamaratran

    Claude Code plugin and npm library: Jev scores tool calls and drops stale ones instead of summarizing context

    Claude Code 插件 + npm 库:用 Jev 给工具调用打分并丢掉过时的,而不是把上下文摘要掉

  • SkillRankergithub.com/Dicklesworthstone

    Rust CLI: Jev ranks which agent skill fits the next step from live session context, with Claude Code hooks

    Rust CLI:根据当前会话上下文,让 Jev 给下一步该用哪个 agent skill 排序,带 Claude Code hook

  • langchain-skill-routergithub.com/deyna256

    LangChain deepagents middleware for per-turn skill routing: Jev ranks and verifies which SKILL.md skills each turn needs from a catalog of hundreds, splitting the ranking to fit Jev's limits and falling back to the full catalog on failure. The judge is pluggable. pip install "langchain-skill-router[jev]".

    LangChain deepagents 中间件,按轮路由 skill:Jev 从数百个 SKILL.md skill 中排序并核验本轮需要哪些,排序会拆分以适应 Jev 的调用上限,出错时回退到完整目录。判定器可替换。pip install "langchain-skill-router[jev]"。

  • JevRoutergithub.com/BillionsBobby

    Unofficial router that puts models, subagents, skills, MCP tools, and CLIs in one candidate set: Jev answers one Choice, and code enforces availability, permissions, risk, and confirmation. On 10 Toolathlon tasks, position-wise hits were 38–44% for Jev against 24% for DeepSeek V4.1 Flash

    非官方路由器:模型、子 agent、skill、MCP 工具和 CLI 放进同一个候选集,Jev 做一次 Choice,代码负责可用性、权限、风险和确认。10 个 Toolathlon 任务上,Jev 的位置命中率是 38–44%,DeepSeek V4.1 Flash 是 24%

  • JevLoopgithub.com/zjunlp

    Unofficial agent loop that sends each fork (tool, risk, done) to Jev 1.13.0 and keeps the LLM for writing; with no key it falls back to local Laya, then rules. npm run demo runs offline

    非官方 Agent 循环:每个分叉(选工具、风险、是否做完)交给 Jev 1.13.0,写字仍留给 LLM;没有 key 时退到本地 Laya,再退到规则。npm run demo 可以离线跑

  • Jevbridgegithub.com/gamesonrblx

    Unofficial ACP/MCP adapter: typed Jev decisions and computer use beside Codex, Claude, Grok, and OpenCode

    非官方 ACP/MCP 适配器:把 Jev 的类型化判断和 computer use 接到 Codex、Claude、Grok、OpenCode 旁边

  • evegithub.com/vercel

    Vercel's agent framework. Experimental autoModel defaults to Gateway typesafe-ai/jev to pick a language model from an allowlist.

    Vercel 的 Agent 框架。实验性 autoModel 默认用 Gateway 上的 typesafe-ai/jev,从白名单里挑语言模型。

  • jev-mcpgithub.com/jkudish

    Node MCP wrapping three cookbook patterns: jev_verify (citation check), jev_screen (prompt-injection / guardrails), jev_find (semantic ranking without embeddings). npx -y github:jkudish/jev-mcp.

    Node MCP,封装三条 cookbook:jev_verify(引文核验)、jev_screen(注入/护栏)、jev_find(无需 embedding 的语义排序)。npx -y github:jkudish/jev-mcp。

  • Jev MCP (Python)Jev MCP(Python)github.com/blakestone-x

    Python MCP server: classify, score, check, match, and screen tools.

    Python MCP:classify、score、check、match、screen。

  • Jev Review MCPgithub.com/NiazMorshed2007

    Local-first MCP: Claude Code, Codex, Cursor, and OpenCode get structured quality review from Jev while they write. Not the same project as Jev Review above.

    本地优先的 MCP:Claude Code、Codex、Cursor、OpenCode 边写边拿 Jev 的结构化质量审查。与上面应用里的 Jev Review 不是同一个项目。

  • typesafe-mcpgithub.com/itsmostafa

    Go CLI and single-binary MCP for Claude Desktop, Claude Code, and Codex.

    Go CLI + 单二进制 MCP,适配 Claude Desktop、Claude Code、Codex。

  • pi-typesafegithub.com/DevMortimer

    Pi extension: one consented, key-managed TypeSafe client, batched typesafe_evaluate, offline-testable transport.

    Pi 扩展:一份经同意的、密钥托管的 TypeSafe 客户端,批量 typesafe_evaluate,可离线测传输。

  • pi-jevgithub.com/y0usaf

    Pi extension with a shadow-mode tool-call gate, output judge, and typed jev_ask.

    Pi 扩展:影子模式工具调用门控、输出评判、类型化 jev_ask。

  • pi-wardengithub.com/DevMortimer

    Pi guardrails on pi-typesafe: held tool results instead of a dialog; write checks against a project rules file.

    基于 pi-typesafe 的 Pi 护栏:把判决当成 held tool result 而不是对话框;对照项目规则文件检查写入。

  • pi-jev-auto-modegithub.com/jomatsu

    Pi auto mode: Jev semantically approves bash / write / edit, and fails closed when it cannot decide.

    Pi 自动模式:Jev 按语义批准 bash / write / edit,判断不了就拒绝。

  • Bicameralgithub.com/AbdelStark

    Pi coding harness: LLM writes, Jev supplies typed reflexes for policy, loop detection, and review. Explicitly not a sandbox.

    Pi 编程 harness:LLM 写代码,Jev 提供策略、循环检测和 review 的类型化反射。明确不是沙箱。

  • jev-prefgithub.com/doeixd

    Turn AGENTS.md preferences into a Jev-powered AI linter: project-specific semantic review rules in jev-pref.json, checked against hunks, staged files, or PRs, with findings fed back to your coding agent. npx jev-pref setup.

    把 AGENTS.md 里的偏好变成 Jev 驱动的 AI linter:在 jev-pref.json 定义项目语义审查规则,对 diff hunk、暂存文件或 PR 求值,并把结果反馈给编程 Agent。npx jev-pref setup。

  • ask-jev-skillgithub.com/shantanugoel

    Hermes skill: ask Jev whenever the agent needs a bounded decision.

    Hermes skill:Agent 需要有界决策时去问 Jev。

  • hermes-jev-skillsgithub.com/kerpopule

    Hermes pack, also for Claude Code and Codex: Jev handles model routing, skill choice, search, memory, and compaction. About 0.4 s per routing turn, and about 2.8 s to pick among 377 skills. A measured handoff digest recalled less than the plain transcript, so handoffs keep the dialogue

    Hermes(也覆盖 Claude Code 和 Codex)的一组技能:模型路由、技能选择、检索、记忆和压缩交给 Jev。路由大约 0.4 秒,在 377 个技能里挑选大约 2.8 秒。交接摘要的实测召回不如原文,所以交接默认仍保留完整对话

  • jev-system-architectgithub.com/samtay32

    Skill that hunts for brittle semantic logic and turns it into Choice / Score / Noul boundaries.

    专门找脆弱语义逻辑、改写成 Choice / Score / Noul 边界的 skill。

  • augustusgithub.com/24601

    Unofficial augustus and augustus-train agent skills for application-specific decision models: primitive/base-model/method selection, data assembly, fitting, export/reload, bounded improvement and independent evaluation; TypeSafe Jev is the default hosted exemplar

    非官方的 augustus 和 augustus-train 智能体技能,面向特定应用的决策模型,涵盖原语、基础模型与方法选择、数据组装、拟合、导出与重新加载、有界改进及独立评估;默认以 TypeSafe Jev 为托管模型示例

  • jev-axigithub.com/shiftynick

    CLI plus Claude Code and Codex hooks: Jev scores each shell command for hazards before it runs and screens fetched text for prompt injection, with routine commands decided locally so nothing is sent

    CLI 加 Claude Code、Codex hook:命令执行前先用 Jev 给危险性打分,并筛查抓取到的文本是否含提示注入,常规命令在本地判定、不发送任何内容

  • jev-engineeringgithub.com/eugeniughelbur

    Decision layer for coding agents: deterministic rules before any model call, then one Jev request, as a Claude Code hook, an MCP server, a loopback service and a shared team policy. Ships the 300-call injection test behind its own numbers.

    编程 Agent 的决策层:先走确定性规则再发一次 Jev 请求,可作为 Claude Code hook、MCP 服务、本地回环服务,并带共享团队策略。附带支撑其数字的 300 次注入测试。

  • toolgategithub.com/RiskAverseTech

    Unofficial tool-call firewall for Claude Code hooks and MCP servers: Jev answers seven risk questions per call and deterministic code maps them to allow/ask/deny, fails closed, and tracks write-then-execute across a session

    非官方:面向 Claude Code 钩子与 MCP 服务器的工具调用防火墙,由 Jev 回答七个风险问题,确定性代码将其映射为 allow/ask/deny,默认失败即拦截,并跟踪同一会话中先写入后执行的文件

  • Jevoniangithub.com/xinyao27

    Local OpenAI / Anthropic / Responses-compatible proxy where one Jev call answers both the model route and the thinking level for jevonian/auto, from session state (recent messages and tool results, consecutive errors, context headroom, quota, candidate capabilities, cache-switch penalties); deterministic code filters candidates and owns every threshold first, a pinned model or explicit jevonian/<route> skips Jev entirely, and each decision is recorded with the serving model, reason, token usage, and estimated cost.

    本地 OpenAI / Anthropic / Responses 兼容代理:jevonian/auto 用一次 Jev 请求同时决定走哪个模型和用多深的思考,状态来自会话(近期消息与工具结果、连续报错次数、上下文余量、配额、候选能力、切换模型的缓存代价);候选筛选和全部阈值由确定性代码负责,指定具体模型或显式 jevonian/<route> 时完全不调用 Jev,每次决策都会记录实际服务的模型、理由、真实 token 用量和估算成本。

  • jev-opusgithub.com/WXK-AI

    Unofficial CLI and Claude Code gateway: runs Claude Code on Opus 5.5 and asks Jev typed Choice / Score / Noul questions on each prompt and after every tool batch to pick the next call's reasoning effort, with code enforcing floors and ceilings and sending the result as a per-message effort statement so the prompt cache is kept

    非官方 CLI 和 Claude Code 网关:在 Opus 5.5 上运行 Claude Code,每次收到提示词和每批工具调用之后,用 Jev 的 Choice / Score / Noul 类型化问题决定下一次调用的推理强度(effort),由代码负责上下限,并以逐条消息的 effort 声明发送,因此不会破坏提示缓存

  • jev-belaygithub.com/valentynkit

    Claude Code Stop hook that checks the transcript for evidence before trusting a "done" claim, spending one four-question Jev call only when files changed with no passing check since, and failing open on every error path.

    Claude Code 的 Stop 钩子:先从对话记录里找证据,只有在文件改动且之后没有通过检查时才发起一次四问的 Jev 调用来核实"完成",任何出错都放行。

  • jev-commitgithub.com/valentynkit

    Pre-commit hook where one Jev call judges whether the commit message matches the staged diff, flags debug leftovers and unmentioned work, and blocks only when it detects a credential.

    Git 预提交钩子:用一次 Jev 调用判断提交信息是否匹配暂存的改动,并检查调试残留、未提及的改动和凭据泄露,只有检测到凭据才会阻止提交。

  • jev-usegithub.com/shitianfang

    Claude Code, Codex and pi plugin: Jev answers the batched typed questions an agent loop needs, and a typed escalation contract hands writing and low-confidence steps back to the LLM

    Claude Code、Codex 和 pi 插件:把 Agent 循环里不需要输出文本的判断批量交给 Jev,需要写字或置信度不足的步骤按类型化契约退回 LLM

  • dsh-jev-toolsgithub.com/HorusJiang

    DeepSeek Harness plugin: Jev prunes oversized tool output, screens fetched pages for injected instructions, and picks which skill fits the next step, plus the jev_ask and jev_gate tools

    DeepSeek Harness 插件:用 Jev 精简超长工具输出、筛查抓取页面里的注入指令、挑选下一步该用的 skill,并提供 jev_ask 与 jev_gate 两个工具

  • slop-gradergithub.com/lukstei

    Rule-based CLI and agent skill that grades text against custom rulesets for AI slop, grammar, and technical documentation quality, and guides an AI agent to auto-fix violations

    基于规则的命令行与 Agent skill:按自定义规则集(custom rulesets)用 Jev 评分和行级标志检查文本的 AI 废话、语法和技术文档质量,并引导 AI Agent 自动修复违规

  • pytest-jevgithub.com/allebee

    pytest plugin for semantic assertions on LLM output: each plain-English claim about a reply becomes a Jev Noul in one request, a claim passes at p ≥ 0.8, and failures print every claim's probability; choice and score cover routing and rubric checks

    pytest 插件,为 LLM 输出做语义断言:关于回复的每条自然语言断言都作为 Jev Noul 问题在一次请求中提出,p ≥ 0.8 才算通过,失败时列出每条断言的概率;choice 和 score 用于路由和评分检查

  • jgrep (kyu1204)github.com/kyu1204

    Semantic grep for code, git diffs and CSV rows: one Noul per 5-60 line chunk, 16 chunks per Jev request, grep-style file:line output and exit codes for CI lint rules written in English

    面向代码、git diff 和 CSV 行的语义 grep:每个 5-60 行代码块一个 Noul,每次 Jev 请求打包 16 个块,输出 grep 风格的 file:line 和退出码,可在 CI 中用英文句子做规则检查

  • jevgrep (allebee)github.com/allebee

    Streaming grep by meaning for logs: asks Jev one Noul per line against a plain-English question and prints the lines at or above a threshold, including from tail -f

    面向日志的流式语义 grep:对每一行向 Jev 提出一个 Noul 问题(用自然语言描述条件),打印概率不低于阈值的行,也可接在 tail -f 后使用

  • wellposedgithub.com/suraj-phanindra

    Offline linter and agent skill for Jev requests: 40 structural checks with no model call (missing none-of-the-above options, broken state paths, wrong criteria shapes), plus Jev-on-Jev checks for what structure cannot decide, with labelled corpora that score both layers.

    面向 Jev 请求的离线 linter 与 agent skill:40 条结构检查完全不调用模型(缺少「以上都不是」选项、state 路径失效、criteria 形状错误),再用 Jev 自身检查结构无法判定的部分,并附带为两层分别打分的标注语料。

  • jev-auto-approvegithub.com/BasmaAbouzied0

    Claude Code PreToolUse hook: one Jev Noul per Bash command on whether it is strictly read-only; auto-approves at p ≥ 0.95, otherwise falls back to the normal permission prompt and never denies. A local hard-no list and injection filter keep risky commands away from Jev; 0 of 8 state-changing commands approved in its published calibration

    Claude Code PreToolUse hook:每条 Bash 命令向 Jev 提一个 Noul,判断是否严格只读;p ≥ 0.95 自动批准,否则回退到正常的权限确认,从不拒绝。本地黑名单和注入过滤让高风险命令不会发给 Jev;公开校准中 8 条会改变状态的命令无一被批准

  • jev-secret-guardgithub.com/BasmaAbouzied0

    Claude Code PreToolUse hook that stops agents writing or sending secrets: known key formats are blocked locally, unknown high-entropy strings go to Jev as a Noul only in masked form so the check never leaks the value; p ≥ 0.80 blocks, 0.30 to 0.80 or any Jev error asks the human. 6 of 6 secrets and 0 of 6 benign strings blocked in its published calibration

    阻止 Agent 写入或发送密钥的 Claude Code PreToolUse hook:已知格式的密钥在本地直接拦截,未知的高熵字符串只以脱敏形式作为 Noul 发给 Jev,检查过程本身不会泄露密钥;p ≥ 0.80 拦截,0.30 到 0.80 或 Jev 出错时交给人确认。公开校准中 6 个密钥全部拦截,6 个无害字符串无一被拦截

Benchmarks & Evaluations评测与排行榜 18

Leaderboards first, then single-task studies and evaluation tooling. Most studies so far measure Jev.

先列排行榜,再列单项评测和评测工具。目前大多数评测针对的是 Jev。

Leaderboards排行榜

  • Jev Decision Indexhuggingface.co/spaces/multimodalart

    Hugging Face Space comparing Jev with open decision models on a versioned benchmark suite, with calibration metrics and documented methods; open-model inference timings and Jev's hosted HTTPS latency are not directly comparable

    Hugging Face Space 排行榜:用版本化 benchmark 套件对比 Jev 与开源决策模型,含校准指标和评测方法;开源模型推理耗时与 Jev 托管 HTTPS API 的延迟不可直接比较

  • JevBenchgithub.com/fstandhartinger

    MIT-licensed harness and leaderboard of 534 English decisions with public and sealed tiers, reporting accuracy, latency, and price together. Discussion: Show HN.

    MIT 协议的评测框架与排行榜:534 道英文决策题,分公开题和封存题,同时报告准确率、延迟和价格。讨论:Show HN。

  • Jevals.comjevals.com

    Independent benchmark of hosted Jev and six LLMs on the same Noul, Choice and Score questions, graded against human labels (PubMedQA, Banking77, HelpSteer2), with per-decision logs as open data

    独立评测:托管 Jev 与六个 LLM 回答同样的 Noul、Choice、Score 问题,按人工标签打分(PubMedQA、Banking77、HelpSteer2),每次决策的日志公开

  • LangWatch Jev benchmarklangwatch.ai

    Jev against seven open models under 1B on 15 decision tasks, with 95% intervals, latency, and contamination checks

    在 15 个决策任务上对比 Jev 与 7 个 1B 以下开源模型,附 95% 区间、延迟和数据污染检查

Studies单项评测

  • typesafe-ai-benchmarkgithub.com/iammrduncan

    Side-by-side of Jev vs Qwen 3.8 27B on Cerebras for the same System One questions. Video: Shannon.

    同一套 System One 问题,对比 Jev 与 Cerebras 上的 Qwen 3.8 27B。视频:Shannon。

  • Jev Rerank Benchgithub.com/anessbelbati

    Reranking comparison with raw provider responses, scoring code, uncertainty intervals, and documented limits.

    重排序对比:原始 provider 响应、打分代码、不确定区间、写明的局限。

  • Jev Spam Evalgithub.com/bitnovus

    Exploratory zero-shot spam study vs trained TF-IDF baselines, with post-hoc-tuning caveats.

    探索性零样本垃圾邮件研究,对照训练过的 TF-IDF 基线,并写了事后调参的 caveat。

  • Jev × NASA Keplergist.github.com

    Independent retrospective test of Jev 1.13 on 8,054 historical Kepler Objects of Interest with NASA Exoplanet Archive dispositions hidden during prediction; 72.5% archive-disposition match vs 64.4% for a fixed 3-rule baseline, with exact requests, metrics, baseline, and caveats

    对 8,054 个历史 Kepler 关注目标(Kepler Objects of Interest)进行的独立回顾性 Jev 1.13 测试;预测期间隐藏 NASA 系外行星档案库分类,档案分类匹配率为 72.5%,固定三规则基线为 64.4%,并公开了完整请求、指标、基线和局限说明

  • Jev Phishing Benchgithub.com/anisselbd

    2,000 emails: Jev vs Claude Haiku 4.5 on click-or-not, with calibration, latency, and cost. Haiku wins accuracy here.

    2000 封邮件:Jev 对 Claude Haiku 4.5 做点不点链接,带校准、延迟和成本。这里准确率是 Haiku 更高。

  • jev-agent-failure-benchmarkgithub.com/TokenTrim

    Who&When Pro (injected agent failures): Jev vs a strong LLM on who / which step / error category.

    Who&When Pro(注入的 Agent 故障):Jev 对强 LLM,预测是谁 / 哪一步 / 哪类错误。

  • jev-sec-benchgithub.com/Gaurav-Gosain

    Blind prompt-injection and vulnerable-code detection benches on public corpora, built on jev-go.

    公开语料上的盲测:提示注入和漏洞代码检测,基于 jev-go。

  • ASSAY-001github.com/jourdanlabs

    Independent pre-registered check of Jev calibration and type safety on Banking77 / CLINC150. Split verdict, full logs. Write-up: donttrustme.ai

    独立预注册核验:Banking77 / CLINC150 上测 Jev 校准与类型安全。结论分裂,日志全公开。文章:donttrustme.ai

  • Jev search rerank evalgithub.com/zhuyansen

    9,831 labelled pairs: Jev rerank vs BM25 / bge-m3, with judge-circularity measured. Fusion wins; Jev alone does not beat embeddings

    9831 对标注:Jev rerank 对照 BM25 / bge-m3,并量化评委循环偏差。融合最好;Jev 单独打不过 embedding

  • Smoking-history extraction benchmark吸烟史抽取评测github.com/vclic

    1,000 synthetic notes: Jev vs OpenAI structured outputs on accuracy, cost, and latency

    1000 条合成病历:Jev 对 OpenAI structured outputs,比准确率、成本和延迟

  • jev-fanout-benchgithub.com/blowxian

    Compares batched and separate Jev calls in 2,976 requests through OpenRouter, reporting approximately 261 fixed input tokens per request, charges matching the published token rate, and answer differences comparable to repeat-request noise.

    用 2976 次经 OpenRouter 的请求比较一次批量提问和拆开提问:每次请求大约有 261 个固定输入 token,费用与公布的 token 单价一致,答案差异和重复请求的噪声相当

Tooling评测工具

  • System One Playgroundgithub.com/goodboybeau

    Local Apple Silicon workbench comparing Laya, Decider, Kev, Jev, and other engines side by side, with public-dataset benchmarks for accuracy and calibration, input-truncation diagnostics, latency, memory, and load tests

    本地 Apple Silicon 决策模型工作台,可并排比较 Laya、Decider、Kev、Jev 等引擎,附公开数据集上的准确率与校准评测、输入截断诊断,以及延迟、内存和负载测试

  • Jev DSPy Labgithub.com/jmanhype

    Unofficial DSPy companion that records and replays TypeSafe calls while measuring calibration, selective risk, confidence-gated abstention, latency, tokens, and modeled cost.

    非官方 DSPy 配套评测:录制并重放 TypeSafe 调用,测量校准、选择性风险、置信度弃权、延迟、token 和建模成本。

  • jevcalgithub.com/abhixhek

    Unofficial CLI that fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, shows how much traffic still needs an LLM fallback, and fails CI when a Jev update breaks the locked thresholds

    非官方命令行工具:用你自己的标注数据按目标准确率为每个问题拟合置信度阈值,在留出集上验证,给出仍需回退到 LLM 的流量比例,并在 Jev 更新导致已锁定阈值失效时让 CI 失败

Papers论文 9

  • Typed Decision Models: An Early Evidence Audit and Evaluation Checklistarxiv.org

    Reviews 28 papers from the first days after Jev's release: the typed readout shows no independent accuracy advantage over comparable label-probability readouts, the clearest gains are latency and cost, and the paper derives a 14-item evaluation checklist

    综述 Jev 发布后头几天的 28 篇论文:类型化读出相比同类标签概率读出没有显示出独立的准确率优势,最明确的收益是延迟和成本;并据此给出 14 条评测检查清单

  • Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystemarxiv.org

    First data-driven survey and analysis of Jev's application ecosystem across 2,170 public GitHub projects, covering rapid early growth, application domains, and decision-use patterns

    首个 Jev 应用生态综述与分析:覆盖 2,170 个公开 GitHub 项目,记录早期快速增长、应用领域与决策用途分布

  • Evaluating and Benchmarking the System One Model Jevarxiv.org

    Zero-shot Jev 1.13 on 37 datasets (346,009 requests for under $10) against Qwen3.8-27B and Gemma-4-E4B option probabilities: Jev beats Qwen on 27 of 37 and Gemma on all 37, with well-calibrated choice probabilities but poorly placed binary thresholds. Code and raw responses released.

    在 37 个数据集上零样本评测 Jev 1.13(346,009 次请求,花费不到 10 美元),对照 Qwen3.8-27B 和 Gemma-4-E4B 的选项概率:Jev 在 27/37 个数据集上胜过 Qwen,37 个全部胜过 Gemma;Choice 概率校准良好,但二元概率相对 0.5 阈值的位置偏差较大。代码和原始响应已公开。

  • Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?arxiv.org

    Label-free coherence checks: Jev's probabilities for "X" and "not X" miss summing to one by 0.064 on average against 0.293 for Qwen3.8-27B's first-token readout, and its single-label probabilities over-sum to 1.14

    无需标签的一致性检验:Jev 对「是 X」与「不是 X」的概率之和平均偏离 1 达 0.064,Qwen3.8-27B 首 token 读出为 0.293;Jev 三个单标签概率之和平均为 1.14

  • Evaluating System One Models for Agent Security Decisionsarxiv.org

    Jev, Laya, Decider, and Nimble against specialized classifiers and LLM judges on prompt injection and risk screening: good aggregate calibration can hide confident failures concentrated in particular attack groups

    在提示注入和风险筛查上,把 Jev、Laya、Decider、Nimble 与专用分类器和 LLM 评委对比:整体校准良好,也可能掩盖集中在特定攻击类别上的高置信错误

  • JevAdvBencharxiv.org

    First adversarial benchmark for decision models: 812 typed questions and 9,744 single-edit attacks. On Jev 1.13, one unverified opinion appended to the state flips 12.1% of decisions, so the state should be treated as untrusted input.

    第一个面向决策模型的对抗 benchmark:812 道类型化问题、9,744 个单点修改攻击。在 Jev 1.13 上,往状态里追加一句未经核实的观点就能翻转 12.1% 的决策,因此状态应被视为不可信输入。

  • From Text Decisions to Pixelsarxiv.org

    PixelJev, a native-image decision interface on small open multimodal models; 64-shot adaptation lifts Pets accuracy from 60.13% to 92.40%, while calibration does not follow accuracy gains

    PixelJev:基于小型开源多模态模型的原生图像决策接口;64-shot 适配把 Pets 准确率从 60.13% 提升到 92.40%,但校准并不随准确率一起提升

  • Chinese-Jevarxiv.org

    Encoder-only System One model for Chinese, pretrained on 10 million examples and fine-tuned for medicine, law, and finance, with the CJ-Bench benchmark; the authors report higher general-domain accuracy than Jev at a 20x speedup

    面向中文的 encoder-only System One 模型:在 1000 万条样本上预训练,再分别针对医疗、法律、金融微调,并发布 CJ-Bench;作者报告通用领域准确率高于 Jev,速度快 20 倍

  • Calibrated Decision Models for Autonomous Penetration-Testing Harnessesarxiv.org

    Where Jev and Laya fit in LLM pentest agents: finding adjudication, severity recalibration, agent pruning, and confirmation loops, with an exploratory case study

    探讨 Jev 与 Laya 在 LLM 渗透测试 Agent 中的位置:漏洞判定、严重度校正、Agent 剪枝和确认循环,并附一个探索性案例

Articles文章 10

Independent measurements and experiments.

独立实测与实验。

  • Jev in Search: Three Practical EvaluationsJev 搜索场景:三项实测zc277584121.github.io

    Independent experiments on search stopping, memory reranking, and multi-hop relation selection, with implementation links and limitations including private data, unequal sample counts, and a simulated speed illustration

    停搜、记忆重排与多跳关系筛选的独立实测,附实现链接,并说明私有数据、样本数差异与速度动画为模拟等限制

  • Mini-Vibe Check: TypeSafe's Jev Judged Everything I’ve Written in 0.7 Secondsevery.to

    Every's Mike Taylor runs Jev over a writing corpus.

    Every 的 Mike Taylor 用 Jev 扫过自己的写作语料。

  • TypeSafeのJevを正しく驚く、それってLLMでできませんか?zenn.dev

    Reproduces the JSON-vs-logit shortcut on Gemma and compares Jev with LLMs on the public Mario harness.

    用 Gemma 的 logit 并行复现 JSON 捷径,并在公开 Mario harness 上对比 Jev 与 LLM。

  • Jev: one judge call, or twelve dimension scores? I measured both on three tasksagentjournal.dev

    Independent measurement on three classification tasks: one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights, with token costs, confidence intervals, and false-positive rates.

    独立实测:三个分类任务上,每行一次直接提问 vs 12–14 个 Jev 维度加本地拟合权重,附 token 成本、置信区间与误报率。

  • Testing Jev on public and private data: classifier or filter?amankumar.ai

    16,000 calls vs gpt-5.4-mini and gpt-5.6-luna; where it wins, where it breaks, and a threshold procedure

    16000 次调用对照 gpt-5.4-mini 与 gpt-5.6-luna:哪里赢、哪里崩、阈值怎么定

  • Is Jev as Accurate as Frontier Models at Classification?openrouter.ai

    OpenRouter runs all 3,080 Banking77 test utterances through Jev 1.13 and Claude Opus 5: 81.0% vs 84.4% accuracy, 175 ms vs 2,266 ms median, about $0.11 vs $2.42 per 1,000

    OpenRouter 用全部 3080 条 Banking77 测试集对比 Jev 1.13 与 Claude Opus 5:准确率 81.0% 对 84.4%,中位延迟 175 ms 对 2266 ms,每千次约 $0.11 对 $2.42

  • We Tested Jev on 791 Labeled Decisions Against Four LLMsayautomate.com

    Independent OpenRouter run on 8-way and 77-way Banking77 routing plus prompt-injection detection: Jev matches the small models, trails GPT-5.6 Terra by about 5 points on 77-way routing, and a 0.80 confidence gate that escalates the rest to Terra matches Terra's accuracy at about a quarter of the cost

    独立评测,经 OpenRouter 跑 8 类和 77 类 Banking77 路由以及提示注入检测:Jev 与中小模型接近,77 类路由上落后 GPT-5.6 Terra 约 5 个点;置信度不低于 0.80 才采用、其余交给 Terra 时,准确率与 Terra 单独跑对齐,成本大约是其四分之一

  • Jev × LexGLUEgithub.com/chepyle

    Reproducible zero-shot run of Jev 1.13 (typesafe/jev-1.13-20260917) on all seven LexGLUE tasks, 23,607 test examples: mean micro-F1 69.9 at $4.02, against 71.3 at $16.45 for GPT-5.6 Luna via chat JSON

    可复现的零样本评测:Jev 1.13(typesafe/jev-1.13-20260917)跑完全部七个 LexGLUE 任务、23607 条测试样本,平均 micro-F1 69.9、花费 $4.02;对照 GPT-5.6 Luna 的对话 JSON 为 71.3、$16.45

  • Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die rollkantahayashiai.github.io

    Independent calibration check on fair dice, coins and spinners, where the true probability is known exactly, and on synthetic forecast documents; Jev selects face 1 on all 400 die rolls with 82.9% mean reported probability against 19.0% accuracy. Code and raw responses on GitHub.

    独立校准实测:用真实概率已知的公平骰子、硬币和转盘,以及合成预测文档测试 Jev;在 400 次隐藏六面骰实验中,Jev 每次都选择 1,平均报告概率为 82.9%,实际命中率为 19.0%。代码和原始响应见 GitHub。

Community cookbooks.

社区实践教程。

  • Milvus Search with Jevgithub.com/milvus-io

    Nine runnable Python notebooks combining Gemini embeddings, Milvus retrieval, and Jev decisions for reranking, filtering, search stopping, routing, cache reuse, curation, guardrails, and evaluation

    9 篇可运行的社区 Notebook,结合 Gemini 嵌入、Milvus 检索与 Jev 判断,覆盖重排、过滤、停搜、路由、缓存复用、数据筛选、护栏和评估