Nimbus Models · Available now

Small models.
Serious work.

An open family of local coding and reasoning models, shaped for the memory people actually have—not a datacenter they rent by the token.

Public BF16 and GGUF releases with model cards, manifests, and checksums.

Official EvalPlus
Native thinking · pass@1
ResultHumanEvalMBPP
Nimbus-9B v2.1Canonical tests
89.0%87.3%
EvalPlusExpanded tests
82.3%73.3%
542 tasks · every task countedReleased Q5_K_M
One family · Four operating points

Choose the model your machine can keep close.

Nimbus-2B, Nimbus-4B, and Nimbus-9B v2.1 are public now in Transformers and compact GGUF editions. Nimbus-12B remains planned for higher-memory machines later.

02BFastest

Nimbus-2B

A compact coding partner for quick edits, local automation, and lightweight agent loops.

Device target
8 GB class
Local edition
Q5_K_M
Reference edition
BF16
04BBalanced

Nimbus-4B

The middle path: more planning room and code context while remaining practical on everyday hardware.

Device target
8–16 GB
Local edition
Q5_K_M
Reference edition
BF16
09BDeepest

Nimbus-9B v2.1

The most capable first-wave Nimbus model, aimed at deeper reasoning and longer multi-step coding work.

Device target
16 GB class
Local edition
Q5_K_M
Smaller option
Q4_K_M
12BIn the lab

Nimbus-12B

The next weather-named Nimbus model, planned for deeper agentic work on high-memory consumer machines.

Status
Coming later
Device target
16–32 GB
Scores
Not evaluated

Device classes are early guidance. Final recommendations will include real memory measurements, because the model download is only part of what an AI workload uses while it runs.

Measured model downloads · July 2026

Built for memory you can own.

These are the actual model download sizes. A running model needs additional memory for its conversation, working space, and the app itself.

Nimbus-2B8 GB target
Q5_K_M1.31 GiB
BF163.52 GiB
Nimbus-4B8–16 GB target
Q5_K_M2.86 GiB
BF167.85 GiB
Nimbus-9B16 GB Q5 target
Q5_K_M6.02 GiB
BF1616.69 GiB

Practical fit: Nimbus-9B Q5_K_M is aimed at 16 GB devices. The much larger BF16 reference is better suited to a desktop or high-memory laptop with at least 24 GB, and preferably 32 GB.

Nimbus-2B and Nimbus-4B · Complete direct-mode results

Fast answers, measured end to end.

Direct mode asks each model for an immediate final solution. We ran every HumanEval and MBPP problem once, then tested the same answers against both the standard and expanded EvalPlus test suites.

Nimbus-2B · Direct

HumanEval

82 standard passes · 75 HumanEval+ passes.

50.0% / 45.7%
Nimbus-2B · Direct

MBPP

202 standard passes · 167 MBPP+ passes.

53.4% / 44.2%
Nimbus-4B · Direct

HumanEval

121 standard passes · 112 HumanEval+ passes.

73.8% / 68.3%
Nimbus-4B · Direct

MBPP

285 standard passes · 232 MBPP+ passes.

75.4% / 61.4%

Mode matters. Direct mode measures low-latency final answers. Nimbus-9B v2.1 below uses native thinking for deeper problems, so we keep the two result sets clearly labeled instead of presenting them as a controlled size comparison.

Nimbus-9B v2.1 Q5 · Native thinking · Official EvalPlus scores

Strong results across both coding suites.

Nimbus-9B v2.1 completed every HumanEval and MBPP task under the frozen native-thinking protocol. Base and Plus scores are reported separately, with one answer counted for every task.

01

HumanEval

146 of 164 tasks passed the canonical tests.

89.0%
02

HumanEval+

135 of 164 tasks passed both the base and expanded tests.

82.3%
03

MBPP

330 of 378 tasks passed the canonical tests.

87.3%
04

MBPP+

277 of 378 tasks passed both the base and expanded tests.

73.3%
Four complete benchmarks · pass@1

The full result, not a cherry-picked subset.

All 542 benchmark tasks are represented once in each applicable score. HumanEval and MBPP show standard-test correctness; the Plus bars require solutions to survive the expanded test suites too.

Official local EvalPlus scoring of the Nimbus-9B v2.1 Q5_K_M artifact (SHA-256 84c8604a77bcccf850e2a89bf2f3a28d2d846bf11e5f8dbec3094fe31be6e4f6). The protocol used llama.cpp b10007, native thinking, pass@1, and one predefined 60,000-token recovery attempt for each original length-limited nonresponse. Five unrecovered tasks remain misses. No task was sampled repeatedly or selected from multiple answers.

Released Nimbus Q5_K_M · Publisher-reported peer context

Four benchmarks, with format and protocol differences made explicit.

Every Nimbus value below was reproduced from the released Q5_K_M GGUF. The peer values are first-party published results, but their evaluation precision or quantization is not consistently disclosed; they must not be read as Q5_K_M-matched measurements. Prompting, shot count, runtime, harness revision, and reasoning policy also differ.

HumanEval · Full suite

Nimbus-9B v2.1 Think · Q5_K_M 89.0%

Nimbus-9B v2.1 Think · Q5_K_M89.0
Qwen2.5-Coder 7B · reported88.4
Qwen2.5-Coder 3B · reported84.1
Granite 4.1 3B · reported81.71
Gemma 3n E4B · reported75.0
Gemma 3 4B · reported71.3
CodeGemma 7B · reported60.4
HumanEval+ · Full suite

Nimbus-9B v2.1 Think · Q5_K_M 82.3%

Qwen2.5-Coder 7B · reported84.1
Nimbus-9B v2.1 Think · Q5_K_M82.3
Qwen2.5-Coder 3B · reported80.5
Granite 4.1 3B · reported76.83
MBPP · Full suite

Nimbus-9B v2.1 Think · Q5_K_M 87.3%

Nimbus-9B v2.1 Think · Q5_K_M87.3
Qwen2.5-Coder 7B · reported83.5
Qwen2.5-Coder 3B · reported73.6
Granite 4.1 3B · reported71.16
Gemma 3n E4B · reported63.6
Gemma 3 4B · reported63.2
CodeGemma 7B · reported55.6
MBPP+ · Full suite

Nimbus-9B v2.1 Think · Q5_K_M 73.3%

Nimbus-9B v2.1 Think · Q5_K_M73.3
Qwen2.5-Coder 7B · reported71.7
Qwen2.5-Coder 3B · reported62.4
Granite 4.1 3B · reported62.17
LiveCodeBench v6 · Publisher context

Gemma 4 E4B 52.0%

Gemma 4 E4B · reported52.0
Gemma 4 E2B · reported44.0

Nimbus-9B has not yet completed this benchmark under a release-grade protocol, so no Nimbus bar is shown.

Nimbus values are locally reproduced from the released Q5_K_M artifact with the disclosed EvalPlus protocols. Peer values are reported by their publishers: Qwen2.5-Coder technical report, Table 16, IBM Granite 4.1 model card, Google Gemma 4 model card, Google Gemma 3 model card, Google Gemma 3n model card, and Google CodeGemma model card. Peer quantization is not consistently stated and is not assumed to be Q5_K_M. Prompting, shot count, decoding, reasoning budgets, harness revisions, and contamination controls may also differ. These charts provide size context, not a controlled head-to-head ranking.

Inside Nimbus Labs · Research in progress

Rime explores more than one train of thought.

Rime is our experimental architecture for letting several small, shared-parameter workspaces propose, exchange, and refine ideas before producing an answer. The goal is to test whether structured collaboration can deliver more useful reasoning for the same practical compute budget.

The questionCan structure improve reasoning?

We are testing whether compact collaborating workspaces can use parameters and computation more effectively than familiar dense designs.

The study10 designs · 3 starts each

Thirty controlled runs compare recurrent, shared-workspace, and dense Transformer designs under matched conditions.

ProgressCore engineering gates passed

Deterministic execution, exact pause and resume, memory accounting, and reproducible run records are working.

StatusBlinded training underway

The first candidate is learning cleanly. Identities and comparisons remain sealed until every scheduled run is complete.

Rime is research, not part of the first Nimbus model release. These are preliminary engineering findings—not a performance claim. We will share the controlled results whether they support or falsify the idea.

How the official scores were measured

One problem. One answer. Every task counted.

The headline results use native thinking. Each score uses one response per problem—without selecting the best of several attempts—and incomplete answers remain misses.

Reasoning modeNative thinking

The model can reason before returning its final code answer; hidden reasoning text is not scored as the answer.

Model formatQ5_K_M GGUF

The published headline result uses the compact format intended for local devices.

AccountingPass@1

Exactly one final answer is counted for every task; misses remain in the denominator.

Local runtimellama.cpp

Runtime versions, generated answers, manifests, and checksums are preserved with the evaluation.

HumanEvalAll 164 tasks

Every problem is included in the score, including incomplete or incorrect answers.

MBPPAll 378 tasks

The complete task set is scored once for the released Nimbus-9B v2.1 result.

Nimbus Labs · First model family

Available now.

Nimbus-2B, Nimbus-4B, and Nimbus-9B v2.1 are public on Hugging Face with model cards, BF16 weights, compact GGUF editions, manifests, and checksums. Nimbus-12B follows later.

Open the collection