Classification field notes.Mac GPU + gateway trials · 19 September 2026

Classification field test · 19 September 2026

Evidence in.
Decision out.

On the same 20 evidence questions, Bonsai answered all 20 correctly. Five other workflows answered 19; Jev answered 18. DeepSeek matched 19/20 locally but took 10.52 seconds per answer and matched only 2/4 screenshot claims. The graphs show accuracy, speed, and provider-call cost.

20 shared text cases5 text jobs + pixelsMac GPU + Vercel GatewayLuna verifies the evidence
See performance graphs ↓Explore model repos and checkpoints ↓See every failed question ↓

Main result · same 20 source-grounded claims

Performance, at a glance

Nine model variants saw the same evidence and claims. Read the three measures together: correct decisions, median answer time, and provider-call charge.

See one scored question ↓

1 · Correct answers

More is better · 0–20
Bonsai 2 27B20/20
Nimble 9B19/20
Raw Gemma19/20
Kev 4B19/20
Simple Jev · Gemma19/20
DeepSeek V4.1 Flash19/20
Jev18/20
Laya English9/20
Laya Typed7/20

Bonsai’s 20/20 uses its final answers; two initial parser reads were corrected against the saved responses.

2 · Seconds per answer

Less is faster · median of 20
0.030.10.3131030 s
Laya Typed0.04 s
Laya English0.10 s
Nimble 9B0.17 s
Raw Gemma0.17 s
Kev 4B0.24 s
Simple Jev · Gemma0.29 s
Jev0.36 s
Bonsai 2 27B5.87 s
DeepSeek V4.1 Flash10.52 s

Logarithmic axis, 0.03–30 seconds. Labels give the measured seconds. Model loading is excluded; Jev includes its gateway round trip. Local models ran on different Mac/GPU hardware, so these are observed waits, not a hardware-controlled speed ranking.

First observed answer times
ModelFirst answer
Laya Typed0.05 s
Raw Gemma0.17 s
Simple Jev · Gemma0.42 s
Jev0.65 s
Kev 4B1.66 s
Nimble 9B2.59 s
Bonsai 2 27B5.00 s
Laya English6.01 s
DeepSeek V4.1 Flash13.32 s

3 · Provider-call cost

Same 20 calls · USD
Jev via gateway$0.000384
Eight local variants$0 API

Jev: 9,143 input tokens × $0.042 per million; 937 output tokens × $0. Price snapshot ↗. This is an estimate, not an invoice. Local GPU/CPU electricity, hardware, and storage were not metered, so $0 API is not $0 total cost.

Separate jobs remain below: the 24-post bookmark tagging comparison, four field extractions, and four native screenshot questions use different inputs or answer types. Their scores are not mixed into these 20-case graphs.

How we tested the decisions

For each claim, the decision has three choices: supported, contradicted, or not established by the evidence. Six original variants, two Gemma workflows, and local DeepSeek V4.1 Flash saw the same 20 source excerpts and claims. Luna checked the answer key against saved public records.

1Show a real record

A GitHub release, World Bank population row, or USGS earthquake record.

2Ask one claim

Every variant got the same source excerpt and exact claim. Its API wrapper differed by model.

3Check the answer

“Yes,” “no,” or “not enough information.” Luna checked all 20 answers we used for grading.

One actual case from the scored quiz

Did US population grow by 3,758,748?

World Bank source ↗: 333,996,304 in 2022336,755,052 in 2023.

Real increase: 2,758,748. So the claim is wrong.

Exact evidence and claim sent to every classifier
Source usa (World Bank US population (2022–2023)): indicator=SP.POP.TOTL (Population, total); country=United States; 2022=333996304; 2023=336755052.
Claim: The World Bank US population increased by 3,758,748 from 2022 to 2023.

The source excerpt and claim bytes were shared; each model needed its own instruction or typed-question wrapper.

How they answered

ModelAnswerResult
Bonsai 2 27BNoCorrect
Nimble 9BYesMissed
Kev 4BYesMissed
JevYesMissed
Laya EnglishYesMissed
Laya TypedNot enough informationMissed
DeepSeek V4.1 FlashYesMissed

Accuracy by verification job

Each bar uses the same four cases within its job · 0–4 scale

Field extraction

4 identical source-grounded claims per model

Bonsai 2 27B4/4
Nimble 9B4/4
Kev 4B4/4
Jev4/4
Laya English2/4
Laya Typed2/4
Raw Gemma4/4
Simple Jev · Gemma4/4
DeepSeek V4.1 Flash4/4

Claim support

4 identical source-grounded claims per model

Bonsai 2 27B4/4
Nimble 9B4/4
Kev 4B4/4
Jev4/4
Laya English3/4
Laya Typed3/4
Raw Gemma4/4
Simple Jev · Gemma4/4
DeepSeek V4.1 Flash4/4

Source attribution

4 identical source-grounded claims per model

Bonsai 2 27B4/4
Nimble 9B4/4
Kev 4B4/4
Jev4/4
Laya English1/4
Laya Typed0/4
Raw Gemma4/4
Simple Jev · Gemma4/4
DeepSeek V4.1 Flash4/4

Arithmetic and units

4 identical source-grounded claims per model

Bonsai 2 27B4/4
Nimble 9B3/4
Kev 4B3/4
Jev2/4
Laya English1/4
Laya Typed0/4
Raw Gemma3/4
Simple Jev · Gemma3/4
DeepSeek V4.1 Flash3/4

Missing evidence

4 identical source-grounded claims per model

Bonsai 2 27B4/4
Nimble 9B4/4
Kev 4B4/4
Jev4/4
Laya English2/4
Laya Typed2/4
Raw Gemma4/4
Simple Jev · Gemma4/4
DeepSeek V4.1 Flash4/4

GLiNER’s native field extraction

Separate task · 0–4 scale
GLiNER 2.52/4
DeepSeek V4.1 Flash4/4

Same four field source excerpts, but GLiNER extracted values rather than returning claim labels. All four target values appeared; only two fields were clean. It duplicated the Node tag and returned both US 2022 and 2023 populations for the 2023 field. DeepSeek used a generative JSON field prompt on the same four excerpts and returned four clean values; its interface differs from GLiNER’s native extractor. Neither field score is mixed into the 20 verdicts. Median answer wait: GLiNER 0.92 s; DeepSeek 8.21 s. Both had $0 provider API charge; local compute was not metered.

Pixel reading · captured GitHub release

Native image input · 0–4 scale
Bonsai 2 27B4/4
DeepSeek V4.1 Flash2/4
Nimble 9BNo native pixel inputN/A
Kev 4BNo native pixel inputN/A
JevNo native pixel inputN/A
Laya EnglishNo native pixel inputN/A
Laya TypedNo native pixel inputN/A

The same Node.js release used in the text suite was captured as actual page pixels. Bonsai matched 4/4; DeepSeek matched 2/4 with its matching vision encoder. It called the false release-date and LTS-status claims not established instead of contradicted. Luna inspected the screenshot and labels. Median answer wait: Bonsai 7.87 s; DeepSeek 19.29 s. Both had $0 provider API charge; local compute was not metered. N/A marks variants without native image input; no OCR or text surrogate was scored as pixel reading.

DeepSeek’s four pixel answers
  • The visible release page shows Node.js Version 22.0.0. DeepSeek: supported; expected supported.
  • × The visible release page dates this release June 16, 2024. DeepSeek: not_established; expected contradicted.
  • × The visible page says Node.js 22 was already in long-term support when released. DeepSeek: not_established; expected contradicted.
  • The visible release text mentions a WebSocket client. DeepSeek: supported; expected supported.
See the four pixel cases and screenshot
  • The visible release page shows Node.js Version 22.0.0. supported
  • The visible release page dates this release June 16, 2024. contradicted
  • The visible page says Node.js 22 was already in long-term support when released. contradicted
  • The visible release text mentions a WebSocket client. supported

GitHub source page ↗ · Screenshot SHA-256: f23ca6084fa1008f3a81dfcb793be036b56fd9a36c2cb434540e0151f570985a

Captured GitHub Node.js v22.0.0 release page showing date, title, LTS timing, and WebSocket client text

Every case, every classifier

✓ matched gold · × missed gold · focus a mark for its predicted label

Swipe horizontally to compare every model.

Job / source-grounded claimBonsaiNimbleKevJevLaya ENLaya TypedDS V4.1
Field extractionnode tagThe Node.js release tag is v22.0.0.
Gold: supported · node
numpy dateThe NumPy v2.0.0 release was published on 2024-04-24.
Gold: contradicted · numpy
××
usa populationThe World Bank reports US population in 2023 as 336,755,052.
Gold: supported · usa
noto magnitudeUSGS lists the Noto event magnitude as 7.0.
Gold: contradicted · noto
××
Claim supportnode draftThe GitHub Node.js v22.0.0 release is marked draft.
Gold: contradicted · node
××
numpy prereleaseThe GitHub NumPy v2.0.0 release is marked prerelease.
Gold: contradicted · numpy
noto reviewedUSGS marks the Noto event as reviewed.
Gold: supported · noto
aykol placeUSGS places the Aykol event in China.
Gold: supported · aykol
Source attributionnode numpy dateThe Node.js release has the NumPy release publication date, 2024-06-16.
Gold: contradicted · node, numpy
××
usa canada valueThe World Bank 2023 US population is 40,049,088.
Gold: contradicted · usa, canada
××
noto aykol magUSGS lists the Noto event magnitude as the Aykol magnitude, 7.0.
Gold: contradicted · noto, aykol
×
aykol depthUSGS lists the Aykol event depth as 13 km.
Gold: supported · noto, aykol
××
Arithmetic and unitsus growthThe World Bank US population increased by 2,758,748 from 2022 to 2023.
Gold: supported · usa
×
us growth wrongThe World Bank US population increased by 3,758,748 from 2022 to 2023.
Gold: contradicted · usa
××××××
quake mag differenceThe Noto event magnitude is 0.5 higher than the Aykol event magnitude.
Gold: supported · noto, aykol
××
noto depth mThe Noto event depth is 1,000 meters.
Gold: contradicted · noto
×××
Missing evidencenode downloadsThe Node.js release had more than one million downloads in its first week.
Gold: not_established · node
×
numpy adoptionMost NumPy users upgraded to v2.0.0 within a month.
Gold: not_established · numpy
××
usa median ageThe US median age in 2023 was 38.9 years.
Gold: not_established · usa
×
noto fatalitiesThe Noto earthquake caused 200 fatalities.
Gold: not_established · noto

Which questions did they miss?

Open a model to see every failed question, its answer, the expected answer, and the decisive source fact. “Why” describes the visible mismatch, not the model’s hidden reasoning. Each question links to the full case matrix.

Jev: both misses were in arithmetic and units: US population growth and the earthquake depth conversion. Kev: one miss, on that same US growth claim; it matched the other 19 text answers. Nimble: also missed the US growth claim.

Bonsai 2 27B0 missed / 20

All 20 final text verdicts matched the answer key.

Nimble 9B1 missed / 20
  1. The World Bank US population increased by 3,758,748 from 2022 to 2023.Answered Yes; expected No.

    336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.

Kev 4B1 missed / 20
  1. The World Bank US population increased by 3,758,748 from 2022 to 2023.Answered Yes; expected No.

    336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.

Jev2 missed / 20
  1. The World Bank US population increased by 3,758,748 from 2022 to 2023.Answered Yes; expected No.

    336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.

  2. The Noto event depth is 1,000 meters.Answered Yes; expected No.

    The source gives 10 km, which is 10,000 meters, not 1,000.

Laya English11 missed / 20
  1. The NumPy v2.0.0 release was published on 2024-04-24.Answered Not enough information; expected No.

    The NumPy release record says 16 June 2024, not 24 April 2024.

  2. USGS lists the Noto event magnitude as 7.0.Answered Yes; expected No.

    The USGS record says magnitude 7.5, not 7.0.

  3. The GitHub Node.js v22.0.0 release is marked draft.Answered Yes; expected No.

    The Node.js release record explicitly says draft=false.

  4. The Node.js release has the NumPy release publication date, 2024-06-16.Answered Yes; expected No.

    24 April is the Node.js publication date; 16 June belongs to NumPy.

  5. The World Bank 2023 US population is 40,049,088.Answered Yes; expected No.

    40,049,088 is Canada’s 2023 population. The US figure is 336,755,052.

  6. USGS lists the Aykol event depth as 13 km.Answered Not enough information; expected Yes.

    The Aykol record explicitly gives depth_km=13.

  7. The World Bank US population increased by 3,758,748 from 2022 to 2023.Answered Yes; expected No.

    336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.

  8. The Noto event magnitude is 0.5 higher than the Aykol event magnitude.Answered Not enough information; expected Yes.

    Noto 7.5 − Aykol 7.0 = 0.5, so the claim is supported.

  9. The Noto event depth is 1,000 meters.Answered Not enough information; expected No.

    The source gives 10 km, which is 10,000 meters, not 1,000.

  10. Most NumPy users upgraded to v2.0.0 within a month.Answered Yes; expected Not enough information.

    The supplied release excerpt has no user adoption data.

  11. The US median age in 2023 was 38.9 years.Answered Yes; expected Not enough information.

    The supplied World Bank excerpt has population totals, not median age.

Laya Typed13 missed / 20
  1. The NumPy v2.0.0 release was published on 2024-04-24.Answered Not enough information; expected No.

    The NumPy release record says 16 June 2024, not 24 April 2024.

  2. USGS lists the Noto event magnitude as 7.0.Answered Not enough information; expected No.

    The USGS record says magnitude 7.5, not 7.0.

  3. The GitHub Node.js v22.0.0 release is marked draft.Answered Yes; expected No.

    The Node.js release record explicitly says draft=false.

  4. The Node.js release has the NumPy release publication date, 2024-06-16.Answered Yes; expected No.

    24 April is the Node.js publication date; 16 June belongs to NumPy.

  5. The World Bank 2023 US population is 40,049,088.Answered Not enough information; expected No.

    40,049,088 is Canada’s 2023 population. The US figure is 336,755,052.

  6. USGS lists the Noto event magnitude as the Aykol magnitude, 7.0.Answered Not enough information; expected No.

    The Noto record says 7.5; 7.0 belongs to Aykol.

  7. USGS lists the Aykol event depth as 13 km.Answered Not enough information; expected Yes.

    The Aykol record explicitly gives depth_km=13.

  8. The World Bank US population increased by 2,758,748 from 2022 to 2023.Answered Not enough information; expected Yes.

    336,755,052 − 333,996,304 = 2,758,748, so the claim is supported.

  9. The World Bank US population increased by 3,758,748 from 2022 to 2023.Answered Not enough information; expected No.

    336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.

  10. The Noto event magnitude is 0.5 higher than the Aykol event magnitude.Answered Not enough information; expected Yes.

    Noto 7.5 − Aykol 7.0 = 0.5, so the claim is supported.

  11. The Noto event depth is 1,000 meters.Answered Not enough information; expected No.

    The source gives 10 km, which is 10,000 meters, not 1,000.

  12. The Node.js release had more than one million downloads in its first week.Answered Yes; expected Not enough information.

    The supplied release excerpt has no first-week download count.

  13. Most NumPy users upgraded to v2.0.0 within a month.Answered Yes; expected Not enough information.

    The supplied release excerpt has no user adoption data.

DeepSeek V4.1 Flash1 missed / 20
  1. The World Bank US population increased by 3,758,748 from 2022 to 2023.Answered Yes; expected No.

    336,755,052 − 333,996,304 = 2,758,748, not 3,758,748.

GLiNER 2.5 extraction2 unclean / 4
  • Node release tag: extracted v22.0.0 twice; the clean output should contain it once.
  • US 2023 population: returned both 2022 and 2023 values; the requested field was 336,755,052 alone.

This separate extraction task did not ask for a yes/no/not-enough verdict.

Primary sources and frozen evidence

GitHub release records, World Bank population rows, and USGS event records were fetched on 19 September 2026. The SHA-256 prefixes below identify the saved raw JSON. Claims about absent evidence refer only to the bounded excerpts shown to models.

Suite SHA-256: 51edf178f4fcfb5f22007113f581970a0d4e57558566e6bfb3ba8c08144ae640. Raw prompts, responses, and receipt hashes are retained in the lab score ledger.

Model repos & source checkpoints

Direct links to what we ran, the publisher-listed base or backbone where available, and implementation code. These links explain provenance; only the cases above contribute to the scores.

Publisher pages and repositories checked 19 September 2026. “Backbone” names an architecture reference, not a second model tested in this comparison. Jev was accessed through a hosted API; its model guide does not identify a downloadable upstream checkpoint.

Raw Gemma vs Simple Jev on the same GPU

We ran the existing Gemma 4 12B Q4 bookmark model both directly and with Simple Jev’s repository prompt and scoring code ↗. The shared 20 claims and 24 saved public X posts were frozen before this run.

Same 20 evidence claims

Identical source text and answer key
Raw Gemma19/20
Simple Jev · Gemma19/20

Four of four on field-value checks, claim support, source attribution, and missing evidence; three of four on arithmetic and units. Median GPU answer time after load: raw 0.17 s; Simple Jev 0.29 s. Both missed the same population-growth calculation. API charge: $0; GPU electricity was not metered.

Show the missed question
  • us growth wrong: chose supported; expected contradicted. The World Bank rows give a 2,758,748 increase, not 3,758,748.

Bookmark primary topics

24 same posts · compare the first tag
Raw Gemma · first tag23/24
Simple Jev · one tag22/24

Luna set acceptable primary tags from post text before seeing model answers. This is agreement with those sets, not a unique ground-truth accuracy score. Raw Gemma produced one to three tags: its first tag matched on 23/24 posts, and at least one tag matched on 24/24. Median GPU answer time: raw 0.18 s; Simple Jev 0.51 s. A single primary choice is narrower than the live bookmark workflow’s one to three tags.

Show the two mismatches
Show the raw Gemma first-tag mismatch

Model weights: Gemma 4 12B GGUF ↗. Simple Jev used the repository’s shared v1 decision template and scorer through a local llama.cpp next-token adapter; raw Gemma used the production bookmark prompt or a direct evidence-verdict prompt. Both used the same GGUF and GPU, but the prompts and answer formats differed. The Simple Jev run was not the repository’s Hugging Face server or a fine-tuned Jev checkpoint. The bookmark score tests topic classification; the 20-claim score tests evidence verification. These denominators must be read separately.

DeepSeek V4.1 Flash 

The MacBook Pro served the downloaded Q2 GGUF with Metal SSD streaming. These are local model calls, not an OpenRouter or gateway run. The exact 20 evidence claims and 24 public X posts were reused.

Bookmark first tags

Same 24 posts · Luna acceptable sets
Raw Gemma23/24
Simple Jev · Gemma22/24
DeepSeek V4.1 Flash19/24

DeepSeek returned one to three tags. Its first matched 19/24; at least one tag matched 22/24. Raw Gemma reached 24/24 with any tag. DeepSeek’s median local answer wait was 17.22 s. These are agreement with Luna’s subjective acceptable sets, not unique ground truth.

Show DeepSeek’s five first-tag misses

Fields from real records

Same four excerpts · separate task
GLiNER 2.52/4
DeepSeek V4.1 Flash4/4

DeepSeek returned the four requested values cleanly as JSON; GLiNER returned two clean native extractions. They used the same record excerpts and requested fields but different model interfaces. DeepSeek also matched 19/20 evidence verdicts, missing the false US population-growth claim. On four native screenshot claims, it matched 2/4 at a median 19.29 s.

Local weights: DeepSeek V4.1 Flash Q2 GGUF ↗; runtime: ds4.c ↗. The 341 GiB file was served from SSD on a 128 GiB MacBook Pro. No provider API charge was incurred; electricity, SSD wear, and machine cost were not measured.

Method & scope

What we tested: classification for verification: nine variants selected one of three evidence verdicts for the same 20 claims built from GitHub release, World Bank, and USGS records. We also gave Bonsai and DeepSeek the same captured release-page image for four pixel questions, and asked GLiNER and DeepSeek to extract four fields from the same source excerpts. The pixel and extraction jobs use different answer types, so their scores are shown separately.

What the counts mean: exact semantic final-label matches on one frozen set of 20 constructed claims drawn from real primary records. Each of the five text jobs has four identical cases across all nine variants. Bonsai returned two correct final labels after reasoning that confused the initial parser, so its 20/20 is final-answer performance; the initial parser recorded 18/20. No model inspected pixels in the shared text suite. A separate native-pixel job used a captured Node.js release page: Bonsai matched 4/4 visible claims and DeepSeek matched 2/4; variants without native image input were N/A. GLiNER’s separate 2/4 and DeepSeek’s 4/4 measure clean field extraction on the same four excerpts through different model interfaces, not claim classification. The results describe this small fixed suite, not accuracy across all future claims.

Verification: Luna independently checked all 20 gold labels and model receipts against saved raw primary records. Luna checks rendered charts and remains the lead visual and data verifier. The full prompts, raw answers, source hashes, and scoring rule are retained in the lab.