Skip to content

Benchmarking models

lcode bench answers the most common question about local models: which one works best for coding on my machine? It runs a fixed set of small coding tasks against one or more models and reports how many they solved, how fast they ran and how much memory they used.

lcode bench                           # the configured model
lcode bench qwen3.6-35b qwen3.5-9b    # compare models
lcode bench --tasks fix-bug,rename    # run some of the tasks
lcode bench --json results.json       # also save machine-readable results
lcode bench --markdown                # also print a table to paste into a test report
lcode bench · NVIDIA GeForce RTX 4080 Laptop GPU (12 GB VRAM), 31 GB RAM · Ollama 0.32.15

lcode-qwen3.6-35b · 32K context · reasoning on · 8 tasks
  Loaded in 16s · 22.3 GB, 45% on the GPU · reads prompts at 855 tok/s
  ✓ fix-bug       Fix a bug so the tests pass         22s  the tests pass
  ✓ find-code     Answer with file:line                9s  next_wait at backoff.py:13
  ✓ write-script  Write a script from a spec          13s  correct output, also on data it hadn't seen
  ✓ rename        Rename a function across files      26s  its tests and extra checks pass
  ✓ precise-edit  One-line edit in a long file        10s  changed only service_087's timeout
  ✓ recover       Recover from failing commands       30s  used --source and dropped the invalid line
  ✓ spec-fix      Fix a function to match its spec    55s  follows the whole docstring
  ✓ feature       Add a feature across files          71s  priorities work, including old todos

At the end it prints a table comparing the models (see results below).

A model usually takes 4 to 12 minutes; each task stops after 5 minutes (--timeout).

The tasks

Every task runs in a new temporary folder, with every permission granted (yolo mode) and web access off, and ends with an automatic check. The folders are deleted afterwards unless you pass --keep, which also keeps each task's transcript (lcode-bench.log).

Task The model has to Passes when
fix-bug Find and fix a bug so the failing unit tests pass The tests pass and the test file is unchanged
find-code Answer a question about a small codebase It names the right function and cites its file:line
write-script Write a script from a spec and run it The script prints the right output, also for data it never saw
rename Rename a function across several files and tests The old name is gone and extra hidden tests pass
precise-edit Change one value in a 900-line file of near-identical entries Exactly that one line changed
recover Run a command that fails twice (a removed option, a bad input line) and work around it The report is correct and only the bad line was dropped
spec-fix Fix a function from a one-line bug report so it does everything its docstring says Hidden tests of every rule in the docstring pass, edge cases included
feature Add priorities to a small todo app (a new option, sorting, a marker in the output) Its tests pass, and a scripted session gives the exact expected output, including for todos saved before the change

The first six tasks are everyday work that good models get right; the last two need careful reading and are where smaller models tend to slip. The checks are deterministic and can't be passed by hard-coding: for example, write-script reruns the script on new data, and rename, spec-fix and feature are checked with tests the model never saw.

What's measured

Passed Tasks solved within the time limit
Time Total time for the tasks (loading the model is reported separately)
Generation Output tokens per second, across all the model's responses
Prompt reading Prompt tokens per second: how fast the model reads files and tool output. Measured per model with fresh 4,000-token prompts (the faster of two readings, since the first one after loading includes warm-up), because during the tasks Ollama serves most of each request from its prompt cache
Tool-call errors Tool calls that failed (wrong arguments, editing without reading first, unmatched edits) or that the model wrote as invalid JSON
Memory The model's memory use and how much of it is on the GPU, from Ollama

Results depend on the context window, so lcode bench uses 32K by default on every machine, which makes results comparable. When comparing models, each one is unloaded before the next starts so every model gets the whole GPU; lcode warns if other models are already loaded. Use --context 128k to measure a larger window, and --no-think to measure a model without reasoning.

Results on a 12 GB laptop

RTX 4080 Laptop GPU (12 GB) with 31 GB of RAM, Ollama 0.32.15, 32K context, reasoning on:

qwen3.6-35b qwen3.5-9b qwen3.5-4b
fix-bug ✓ 22s ✓ 16s ✓ 11s
find-code ✓ 9s ✓ 6s ✓ 5s
write-script ✓ 13s ✓ 44s ✗ 18s
rename ✓ 26s ✓ 28s ✓ 30s
precise-edit ✓ 10s ✓ 12s ✓ 8s
recover ✓ 30s ✗ 29s ✓ 24s
spec-fix ✓ 55s ✓ 4m 40s ✗ 2m 57s
feature ✓ 1m 11s ✗ 5m 00s ✓ 47s
Passed 8/8 6/8 6/8
Time 3m 57s 11m 55s 5m 20s
Generation 71.4 tok/s 60.1 tok/s 91.3 tok/s
Prompt reading 855 tok/s 3,507 tok/s 5,186 tok/s
Tool-call errors 0 9 1
Memory 22.3 GB, 45% on GPU 6.6 GB, 100% on GPU 4.2 GB, 100% on GPU

The small models fit entirely on the GPU and read prompts much faster, but the larger mixture-of- experts model solves more and needs fewer attempts. The small models' failures were real mistakes: copying read_file's line numbers into a data file, ignoring negative numbers, missing an edge case of the spec, and running out of time.

Models sample their answers, so a task near the edge of a model's ability can pass in one run and fail in the next (qwen3.6-35b scored 7/8 in an earlier run). Compare totals rather than single tasks, and run twice before drawing conclusions.

Sharing results

Results from other machines, especially Apple Silicon Macs and other GPUs, decide which models lcode recommends. Run

lcode bench qwen3.6-35b --markdown

and paste the table into a model test report.

JSON format

--json FILE writes one document per run. The schema is versioned: fields may be added, but existing fields keep their meaning until schema changes.

{
  "schema": 1,
  "lcode": "0.3.1",
  "ollama": "0.32.15",
  "date": "2026-09-30T19:09:09+02:00",
  "hardware": {
    "os": "linux", "cpu": "13th Gen Intel(R) Core(TM) i9-13980HX", "ram_gib": 30.96,
    "gpu": "NVIDIA GeForce RTX 4080 Laptop GPU", "vram_gib": 11.99, "unified": false,
    "description": "NVIDIA GeForce RTX 4080 Laptop GPU (12 GB VRAM), 31 GB RAM"
  },
  "tasks": [{"id": "fix-bug", "title": "Fix a bug so the tests pass"}, …],
  "runs": [
    {
      "model": "qwen3.6-35b",
      "ollama_model": "lcode-qwen3.6-35b",
      "context": 32768,
      "num_batch": null,
      "think": true,
      "load_seconds": 16.2,
      "prompt_tps": 854.6,
      "memory_gb": 22.3,
      "gpu_percent": 45,
      "passed": 8,
      "total": 8,
      "seconds": 237.3,
      "generation_tps": 71.4,
      "tool_errors": 0,
      "usage": {"requests": 56, "prompt_tokens": 221522, "prompt_ns": 73377987000,
                "output_tokens": 11490, "output_ns": 161017657999},
      "tasks": [
        {"id": "fix-bug", "passed": true, "seconds": 21.9, "detail": "the tests pass",
         "steps": 5, "tool_errors": 0, "error": null},
        …
      ],
      "error": null,
      "interrupted": false
    }
  ]
}
Field Meaning
runs[].model The name you gave (catalog key or Ollama tag); ollama_model is the model that ran
runs[].context, num_batch, think The settings used. lcode lowers the batch size or context if the GPU runs out of memory, and reports what it ended up using
runs[].memory_gb, gpu_percent From Ollama's /api/ps; null if unavailable
runs[].generation_tps Output tokens per second across the tasks; null if Ollama reported no timings
runs[].prompt_tps Prompt tokens per second for a fresh 4,000-token prompt (the faster of two readings); null if it couldn't be measured
runs[].usage Totals over the tasks as Ollama reports them. prompt_tokens includes tokens served from the prompt cache
runs[].tasks[].steps Model responses in the task
runs[].tasks[].error Why the task stopped early (timed out after 300s, an Ollama error), otherwise null
runs[].error Why the model didn't run at all (for example, not installed), otherwise null
runs[].interrupted You pressed Ctrl+C; the results cover the tasks that finished