1 · Correct answers
More is better · 0–20Bonsai’s 20/20 uses its final answers; two initial parser reads were corrected against the saved responses.
Classification field test · 19 September 2026
On the same 20 evidence questions, Bonsai answered all 20 correctly. Five other workflows answered 19; Jev answered 18. DeepSeek matched 19/20 locally but took 10.52 seconds per answer and matched only 2/4 screenshot claims. The graphs show accuracy, speed, and provider-call cost.
Main result · same 20 source-grounded claims
Nine model variants saw the same evidence and claims. Read the three measures together: correct decisions, median answer time, and provider-call charge.
Bonsai’s 20/20 uses its final answers; two initial parser reads were corrected against the saved responses.
Logarithmic axis, 0.03–30 seconds. Labels give the measured seconds. Model loading is excluded; Jev includes its gateway round trip. Local models ran on different Mac/GPU hardware, so these are observed waits, not a hardware-controlled speed ranking.
| Model | First answer |
|---|---|
| Laya Typed | 0.05 s |
| Raw Gemma | 0.17 s |
| Simple Jev · Gemma | 0.42 s |
| Jev | 0.65 s |
| Kev 4B | 1.66 s |
| Nimble 9B | 2.59 s |
| Bonsai 2 27B | 5.00 s |
| Laya English | 6.01 s |
| DeepSeek V4.1 Flash | 13.32 s |
Jev: 9,143 input tokens × $0.042 per million; 937 output tokens × $0. Price snapshot ↗. This is an estimate, not an invoice. Local GPU/CPU electricity, hardware, and storage were not metered, so $0 API is not $0 total cost.
Separate jobs remain below: the 24-post bookmark tagging comparison, four field extractions, and four native screenshot questions use different inputs or answer types. Their scores are not mixed into these 20-case graphs.
For each claim, the decision has three choices: supported, contradicted, or not established by the evidence. Six original variants, two Gemma workflows, and local DeepSeek V4.1 Flash saw the same 20 source excerpts and claims. Luna checked the answer key against saved public records.
A GitHub release, World Bank population row, or USGS earthquake record.
Every variant got the same source excerpt and exact claim. Its API wrapper differed by model.
“Yes,” “no,” or “not enough information.” Luna checked all 20 answers we used for grading.
One actual case from the scored quiz
World Bank source ↗: 333,996,304 in 2022 → 336,755,052 in 2023.
Real increase: 2,758,748. So the claim is wrong.
Source usa (World Bank US population (2022–2023)): indicator=SP.POP.TOTL (Population, total); country=United States; 2022=333996304; 2023=336755052. Claim: The World Bank US population increased by 3,758,748 from 2022 to 2023.
The source excerpt and claim bytes were shared; each model needed its own instruction or typed-question wrapper.
| Model | Answer | Result |
|---|---|---|
| Bonsai 2 27B | No | Correct |
| Nimble 9B | Yes | Missed |
| Kev 4B | Yes | Missed |
| Jev | Yes | Missed |
| Laya English | Yes | Missed |
| Laya Typed | Not enough information | Missed |
| DeepSeek V4.1 Flash | Yes | Missed |
4 identical source-grounded claims per model
4 identical source-grounded claims per model
4 identical source-grounded claims per model
4 identical source-grounded claims per model
4 identical source-grounded claims per model
Same four field source excerpts, but GLiNER extracted values rather than returning claim labels. All four target values appeared; only two fields were clean. It duplicated the Node tag and returned both US 2022 and 2023 populations for the 2023 field. DeepSeek used a generative JSON field prompt on the same four excerpts and returned four clean values; its interface differs from GLiNER’s native extractor. Neither field score is mixed into the 20 verdicts. Median answer wait: GLiNER 0.92 s; DeepSeek 8.21 s. Both had $0 provider API charge; local compute was not metered.
The same Node.js release used in the text suite was captured as actual page pixels. Bonsai matched 4/4; DeepSeek matched 2/4 with its matching vision encoder. It called the false release-date and LTS-status claims not established instead of contradicted. Luna inspected the screenshot and labels. Median answer wait: Bonsai 7.87 s; DeepSeek 19.29 s. Both had $0 provider API charge; local compute was not metered. N/A marks variants without native image input; no OCR or text surrogate was scored as pixel reading.
GitHub source page ↗ · Screenshot SHA-256: f23ca6084fa1008f3a81dfcb793be036b56fd9a36c2cb434540e0151f570985a
Swipe horizontally to compare every model.
| Job / source-grounded claim | Bonsai | Nimble | Kev | Jev | Laya EN | Laya Typed | DS V4.1 |
|---|---|---|---|---|---|---|---|
| Field extractionnode tagThe Node.js release tag is v22.0.0. Gold: supported · node | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| numpy dateThe NumPy v2.0.0 release was published on 2024-04-24. Gold: contradicted · numpy | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| usa populationThe World Bank reports US population in 2023 as 336,755,052. Gold: supported · usa | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| noto magnitudeUSGS lists the Noto event magnitude as 7.0. Gold: contradicted · noto | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| Claim supportnode draftThe GitHub Node.js v22.0.0 release is marked draft. Gold: contradicted · node | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| numpy prereleaseThe GitHub NumPy v2.0.0 release is marked prerelease. Gold: contradicted · numpy | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| noto reviewedUSGS marks the Noto event as reviewed. Gold: supported · noto | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| aykol placeUSGS places the Aykol event in China. Gold: supported · aykol | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Source attributionnode numpy dateThe Node.js release has the NumPy release publication date, 2024-06-16. Gold: contradicted · node, numpy | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| usa canada valueThe World Bank 2023 US population is 40,049,088. Gold: contradicted · usa, canada | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| noto aykol magUSGS lists the Noto event magnitude as the Aykol magnitude, 7.0. Gold: contradicted · noto, aykol | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| aykol depthUSGS lists the Aykol event depth as 13 km. Gold: supported · noto, aykol | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| Arithmetic and unitsus growthThe World Bank US population increased by 2,758,748 from 2022 to 2023. Gold: supported · usa | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| us growth wrongThe World Bank US population increased by 3,758,748 from 2022 to 2023. Gold: contradicted · usa | ✓ | × | × | × | × | × | × |
| quake mag differenceThe Noto event magnitude is 0.5 higher than the Aykol event magnitude. Gold: supported · noto, aykol | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| noto depth mThe Noto event depth is 1,000 meters. Gold: contradicted · noto | ✓ | ✓ | ✓ | × | × | × | ✓ |
| Missing evidencenode downloadsThe Node.js release had more than one million downloads in its first week. Gold: not_established · node | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| numpy adoptionMost NumPy users upgraded to v2.0.0 within a month. Gold: not_established · numpy | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| usa median ageThe US median age in 2023 was 38.9 years. Gold: not_established · usa | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ |
| noto fatalitiesThe Noto earthquake caused 200 fatalities. Gold: not_established · noto | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
Open a model to see every failed question, its answer, the expected answer, and the decisive source fact. “Why” describes the visible mismatch, not the model’s hidden reasoning. Each question links to the full case matrix.
Jev: both misses were in arithmetic and units: US population growth and the earthquake depth conversion. Kev: one miss, on that same US growth claim; it matched the other 19 text answers. Nimble: also missed the US growth claim.
All 20 final text verdicts matched the answer key.
336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.
336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.
336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.
The source gives 10 km, which is 10,000 meters, not 1,000.
The NumPy release record says 16 June 2024, not 24 April 2024.
The USGS record says magnitude 7.5, not 7.0.
The Node.js release record explicitly says draft=false.
24 April is the Node.js publication date; 16 June belongs to NumPy.
40,049,088 is Canada’s 2023 population. The US figure is 336,755,052.
The Aykol record explicitly gives depth_km=13.
336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.
Noto 7.5 − Aykol 7.0 = 0.5, so the claim is supported.
The source gives 10 km, which is 10,000 meters, not 1,000.
The supplied release excerpt has no user adoption data.
The supplied World Bank excerpt has population totals, not median age.
The NumPy release record says 16 June 2024, not 24 April 2024.
The USGS record says magnitude 7.5, not 7.0.
The Node.js release record explicitly says draft=false.
24 April is the Node.js publication date; 16 June belongs to NumPy.
40,049,088 is Canada’s 2023 population. The US figure is 336,755,052.
The Noto record says 7.5; 7.0 belongs to Aykol.
The Aykol record explicitly gives depth_km=13.
336,755,052 − 333,996,304 = 2,758,748, so the claim is supported.
336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.
Noto 7.5 − Aykol 7.0 = 0.5, so the claim is supported.
The source gives 10 km, which is 10,000 meters, not 1,000.
The supplied release excerpt has no first-week download count.
The supplied release excerpt has no user adoption data.
336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.
This separate extraction task did not ask for a yes/no/not-enough verdict.
GitHub release records, World Bank population rows, and USGS event records were fetched on 19 September 2026. The SHA-256 prefixes below identify the saved raw JSON. Claims about absent evidence refer only to the bounded excerpts shown to models.
Suite SHA-256: 51edf178f4fcfb5f22007113f581970a0d4e57558566e6bfb3ba8c08144ae640. Raw prompts, responses, and receipt hashes are retained in the lab score ledger.
Direct links to what we ran, the publisher-listed base or backbone where available, and implementation code. These links explain provenance; only the cases above contribute to the scores.
TypeSafe’s hosted System One model, called as typesafe-ai/jev through Vercel AI Gateway. No local Jev checkpoint ran.
The Mac test used Prism’s MLX 2-bit package, including its vision tower.
The Mac test used the merged local form of the published LoRA adapter.
The Mac test combined the published adapter and decision head with its stated base model.
The English checkpoint was loaded from the root of the publisher’s model repository.
The typed-decisions checkpoint was loaded from the subfolder in that same repository.
The MacBook Pro served the local 341 GiB Q2 GGUF through ds4.c Metal SSD streaming, with its matching vision encoder for native image input.
This was scored on native field extraction, not on the shared three-choice verdict task.
Publisher pages and repositories checked 19 September 2026. “Backbone” names an architecture reference, not a second model tested in this comparison. Jev was accessed through a hosted API; its model guide does not identify a downloadable upstream checkpoint.
We ran the existing Gemma 4 12B Q4 bookmark model both directly and with Simple Jev’s repository prompt and scoring code ↗. The shared 20 claims and 24 saved public X posts were frozen before this run.
Four of four on field-value checks, claim support, source attribution, and missing evidence; three of four on arithmetic and units. Median GPU answer time after load: raw 0.17 s; Simple Jev 0.29 s. Both missed the same population-growth calculation. API charge: $0; GPU electricity was not metered.
Luna set acceptable primary tags from post text before seeing model answers. This is agreement with those sets, not a unique ground-truth accuracy score. Raw Gemma produced one to three tags: its first tag matched on 23/24 posts, and at least one tag matched on 24/24. Median GPU answer time: raw 0.18 s; Simple Jev 0.51 s. A single primary choice is narrower than the live bookmark workflow’s one to three tags.
Model weights: Gemma 4 12B GGUF ↗. Simple Jev used the repository’s shared v1 decision template and scorer through a local llama.cpp next-token adapter; raw Gemma used the production bookmark prompt or a direct evidence-verdict prompt. Both used the same GGUF and GPU, but the prompts and answer formats differed. The Simple Jev run was not the repository’s Hugging Face server or a fine-tuned Jev checkpoint. The bookmark score tests topic classification; the 20-claim score tests evidence verification. These denominators must be read separately.
The MacBook Pro served the downloaded Q2 GGUF with Metal SSD streaming. These are local model calls, not an OpenRouter or gateway run. The exact 20 evidence claims and 24 public X posts were reused.
DeepSeek returned one to three tags. Its first matched 19/24; at least one tag matched 22/24. Raw Gemma reached 24/24 with any tag. DeepSeek’s median local answer wait was 17.22 s. These are agreement with Luna’s subjective acceptable sets, not unique ground truth.
DeepSeek returned the four requested values cleanly as JSON; GLiNER returned two clean native extractions. They used the same record excerpts and requested fields but different model interfaces. DeepSeek also matched 19/20 evidence verdicts, missing the false US population-growth claim. On four native screenshot claims, it matched 2/4 at a median 19.29 s.
Local weights: DeepSeek V4.1 Flash Q2 GGUF ↗; runtime: ds4.c ↗. The 341 GiB file was served from SSD on a 128 GiB MacBook Pro. No provider API charge was incurred; electricity, SSD wear, and machine cost were not measured.
Method & scope
What we tested: classification for verification: nine variants selected one of three evidence verdicts for the same 20 claims built from GitHub release, World Bank, and USGS records. We also gave Bonsai and DeepSeek the same captured release-page image for four pixel questions, and asked GLiNER and DeepSeek to extract four fields from the same source excerpts. The pixel and extraction jobs use different answer types, so their scores are shown separately.
What the counts mean: exact semantic final-label matches on one frozen set of 20 constructed claims drawn from real primary records. Each of the five text jobs has four identical cases across all nine variants. Bonsai returned two correct final labels after reasoning that confused the initial parser, so its 20/20 is final-answer performance; the initial parser recorded 18/20. No model inspected pixels in the shared text suite. A separate native-pixel job used a captured Node.js release page: Bonsai matched 4/4 visible claims and DeepSeek matched 2/4; variants without native image input were N/A. GLiNER’s separate 2/4 and DeepSeek’s 4/4 measure clean field extraction on the same four excerpts through different model interfaces, not claim classification. The results describe this small fixed suite, not accuracy across all future claims.
Verification: Luna independently checked all 20 gold labels and model receipts against saved raw primary records. Luna checks rendered charts and remains the lead visual and data verifier. The full prompts, raw answers, source hashes, and scoring rule are retained in the lab.