DecisionTune

DecisionTune 1.0: a 395M decision model. $20 of compute. Clean data.

29.57 on Decision Index 0.2.1

Measured on one full run. On Decision Index, the highest score we can see under 500M parameters (results). On JevBench, it is not (JevBench).

The run is complete. Nothing was truncated. No options were removed. Requests over the context limit were refused, not cut.

Try it

Coming soon: pip package and live demo.

Updates will land here and at hotin.ai.

Results

Under 500M parameters, the highest other entry we can see is Dinah-0 at 27.63 (150M, pending, not merged). DecisionTune 1.0 is 1.94 points above it. Decision 2.0 Sol (2B) scores 29.53 (pending): level with DecisionTune 1.0, at five times the size. Board as of 2026-10-03.

Scatter chart of Decision Index against parameter count for entrants up to 2B. DecisionTune 1.0 sits at 395M parameters with 29.57. Dinah-0 is 27.63 at 150M. Decision 2.0 Sol is 29.53 at 2B. Filled points are merged entries; hollow points are pending entries.
Bar chart of skill by area, DecisionTune 1.0 against DecisionTune 0.9 Preview. Tools and Automation rises from 28.1 to 46.5. Other areas move by 0.2 to 2.5 points.

Per area (skill)

Source: our own full runs on Decision Index 0.2.1. DecisionTune 0.9 Preview is our earlier clean model; it scores 25.12.
Area1.00.9 PreviewChangeDinah-0
Knowledge & Reasoning13.312.2+1.119.1
Language Understanding31.529.0+2.522.7
Retrieval & Classification45.044.8+0.247.8
Tools & Automation46.528.1+18.437.9
Arts & Human Taste4.83.7+1.13.5

Dinah-0 is ahead in two areas (Knowledge & Reasoning, Retrieval & Classification). Nine benchmarks score 0.0, including ANLI, GPQA Diamond, ChessBench and HLE.

JevBench

JevBench is a separate public benchmark of typed decisions. We ran its public set of 231 tasks on our own machine. DecisionTune 1.0 answers 55.0% correctly (127 of 231, 95% CI 48.1 to 61.5). DecisionTune 0.9 Preview answers 51.5% (119 of 231). The gap is 3.5 points, within noise.

On JevBench, DecisionTune 1.0 is not at the top of its size class. Public results for models under 500M parameters range from 22.1% to 62.8%. At least seven of them are above DecisionTune 1.0.

Bar chart of JevBench public-set accuracy. DecisionTune 1.0 scores 55.0 percent, DecisionTune 0.9 Preview 51.5 percent, each with its 95 percent interval. A shaded band shows the range of public results for models under 500M parameters, 22.1 to 62.8 percent.

18 of the 231 tasks are rating questions. DecisionTune never trained on rating questions. It answers them zero-shot, by choosing among the level texts: 9 of 18 correct, where chance is about 4.4. This is the public set only, not a JevBench board score. We copied the peer results from the benchmark's published files. We did not run them again.

Cost

All cloud compute for the project cost $20.34 on rented A100s. Training DecisionTune 1.0 itself cost $1.69.

Horizontal bar chart of where the 20.34 dollars of cloud compute went. The largest line is the final experiments at 9.47 dollars. Training DecisionTune 1.0 itself cost 1.69 dollars.

How we built it

  1. We started from ModernBERT-large and added a 4 KB scoring head.
  2. We checked every license first. We kept a dataset only when its publisher tag is permissive. We dropped anything marked non-commercial, share-alike or research-only.
  3. A $0.14 speed test moved all training off our laptop. We trained on rented A100s at $0.47 to $0.54 per hour.
  4. Part of its training learned from a larger open model, Clef-Flash (Cloudflare/clef-flash, Apache-2.0). It labeled 93,303 training units for about $0.76 (estimate). We used its answers on 8 datasets.
  5. We found shortcuts in our own training data and our own scorer. We fixed them and measured again.

The full build story comes in a launch write-up.

Honesty

The full list of our mistakes comes with the launch write-up.