Shane Burrell
AI-era leadership 14 min read

One Brain on the Desk: Strix Halo, Flash-Next, and OpenCode Web

A simple non-cloud coding setup that finishes real work: one Flash-Next model on a Strix Halo class machine, an OpenAI-compatible API, and OpenCode Web. Measured throughput, the hardware constraints that matter, and a short howto.

The useful question is smaller than the catalog of local-AI projects. What is a simple setup, off the cloud meter, that can actually do coding work at a pace you will wait for?

I keep one. A single model on a high-memory AMD APU, one OpenAI-compatible API, and OpenCode Web as the client. The browser is a window. The session lives on a machine that is not the GPU. Closing the tab does not kill the job.

That is the whole product. One brain. Decent performance. A diff a person still owns.

A person at a quiet wooden desk with a laptop and one small, plain mini PC beside it, working in soft daylight.
The whole setup fits on a desk: a browser session, and one box that holds the model.

What “Decent” Means

Decent is a number you measured on a task you would actually run, then dated so you can throw it out when the engine changes.

On this class of box, a Flash-Next serve on 2026-09-15 held prefill at about 1.3–1.5k tokens a second from a short prompt out through 131k tokens. Short generation was about 47 tokens a second. A later 32k turn that hit the prompt cache finished in about half a second. The first follow-up with a new prefix still paid the full prefill. Cache is a gift on the second turn, and it is honest about that.

The engine I am serving now is halogen-flash-server 0.16.1, checked on 2026-10-02 and still the process on the box. Same model line, the v2 checkpoint. Footprint about 85 GiB of a 125 GiB unified pool, which leaves on the order of 37 GiB. That leftover is headroom for the OS, the cache, and a spike. It is room. Spend it on a second model and you no longer have this brain.

Tool calls work. A forced echo tool returned finish_reason: tool_calls, with a whole-number argument staying a whole number. The engine context is 262144 tokens. OpenCode is told 258048 in and 16384 out, so a turn has room to finish inside the window. Vision is on: a screenshot is an input, text is the output.

Forty-seven tokens a second is a pace you can watch. A multi-file edit is a few minutes, and you can read the diff while it lands. Prefill that stays in the thousands out past a hundred thousand tokens is what makes a long agent turn possible. A model that answers a one-line prompt and then falls over at repo scale is a demo.

This is the coding-runtime sibling of the leadership-grade homelab. That practice is restore, blast radius, and GitOps. This one is a session you can restart, a model you can name, and a token bill you are not renting for the afternoon.

What You Are Running

Three pieces, and a rule about where each one sits.

Piece What I run Why it is the piece
Brain Qwen3.8 Flash-Next via halogen-flash-server 0.16.1 One model ID, 262k context, tool calls, vision
Contract OpenAI-compatible /v1 on the LAN OpenCode speaks it. So can anything else you trust on that network
Client OpenCode Web 2 on a separate host The UI. Basic auth. The session survives the tab
Stack diagram: a browser on the desk, OpenCode Web with LAN basic auth, an OpenAI-compatible API, and one Flash-Next model on Strix Halo.
Own the middle. The laptop can change. The model ID should not, until you mean to change it.

The GPU host is a Strix Halo class APU: an AMD Ryzen AI MAX+ with on the order of 128 GiB of memory shared by the CPU and the iGPU. The interesting property is the pool. Weights and context live in memory the GPU can actually use, which is why a model this size fits in a small box.

OpenCode Web does not run on that box. It runs next to your repos, on a machine whose job is the agent: files, shell, and the session store. The GPU machine’s job is inference. Mixing a desktop session onto the same iGPU spends the pool you just measured.

Hardware Constraints That Matter

A few of these were settled the hard way. They transfer to any high-memory APU in this shape.

Leave the BIOS VRAM carveout low. The Windows habit is to gift the GPU a huge static slab. On Linux the GPU takes memory from GTT at runtime. Raising the carveout shrinks the pool. A 512 MB carveout is the setting that leaves the DIMMs available.

Raise GTT, then reboot. On this class of machine a ceiling on the order of 110 GiB is what makes an 85 GiB brain fit with room left over. Confirm free reports something in the neighborhood of the DIMM size. If you skip the reboot, you will debug the wrong layer for an hour.

The device nodes are the GPU. This engine wants /dev/kfd and /dev/dri/renderD128. The user that starts the container needs render and video. Headless, there is no seat ACL handing those out. If either node is missing, stop. You are about to serve on the CPU and call it a result.

Wired Ethernet. A long prefill is a steady stream. Wi-Fi will teach you a lesson about timeouts that has nothing to do with the model.

Disk is a weight budget. The v2 checkpoint plus its lookup table want on the order of 110 GiB free. Keep the previous checkpoint on disk until the new one has passed a real turn. A brain you cannot roll back is a demo with a reboot.

The GPU host stays headless once the install is done. A graphical session competes for the same iGPU.

You can run a smaller model on less memory. Do not keep this context window and hope.

Phase How-To

Each phase ends with a check. If the check fails, stop. Later phases assume it passed.

Phase 1 — Weights

Install the Hugging Face CLI in a user tool, not a system Python that will fight you.

pipx install huggingface-hub

mkdir -p ~/halogen-models
hf download peonist-ai/halogen-qwen3.8-flash-next \
  --local-dir ~/halogen-models \
  qwen38-flash-next-v2.hgn \
  qwen38-flash-next-ngram.hgn

The checkpoint is the model. The ngram file is the paged lookup table the server expects beside it. Vision is a separate file (qwen38-flash-next-vision.hgn). Download it when you want screenshots in the turn. Leave it out and the brain is text-only, which is a fine first day.

Check: both required files are on disk, and df still shows room for the previous checkpoint during a swap.

Phase 2 — One engine

Pin a halogen tag you have smoked. 0.16.1 is the one running here. “Latest” is not a release process.

The container needs the GPU devices and a high memlock. Publish the API on the LAN address your other machines can reach. Do not publish it on the public internet.

mkdir -p ~/halogen-cache

podman run --rm --name halogen-flash \
  --device /dev/kfd --device /dev/dri \
  --group-add keep-groups \
  --ipc=host \
  --ulimit memlock=-1:-1 \
  -p 8080:8080 \
  -e HALOGEN_API_PORT=8080 \
  -e HALOGEN_MODEL_ID=halogen-qwen3.8-flash-next \
  -e HALOGEN_CHECKPOINT=/models/qwen38-flash-next-v2.hgn \
  -e HALOGEN_CTX=262144 \
  -e HALOGEN_MAX_TOK=16384 \
  -e HALOGEN_MAX_TOKENS_DEFAULT=16384 \
  -e HALOGEN_TOP_K=20 \
  -e HALOGEN_TOP_P=0.95 \
  -e HALOGEN_PROMPT_CACHE=2 \
  -e HALOGEN_CACHE_DIR=/cache \
  -e HALOGEN_CACHE_DISK_GIB=16 \
  -v "$HOME/halogen-models:/models:ro" \
  -v "$HOME/halogen-cache:/cache" \
  ghcr.io/peonist-ai/halogen-flash-server:0.16.1

Add -e HALOGEN_VISION_TOWER=1 only when the vision file is in that models directory.

Docker works in place of Podman. Drop --group-add keep-groups and add --group-add video --group-add render instead.

HALOGEN_CTX is the window one request can use. HALOGEN_MAX_TOK is how large a generation the fit will budget. The default reply cap is already 16384 tokens. Leave the KV slot count at the image default. Slots are generator state. The pool is the memory knob, and the server lowers it on its own when the box cannot spare it. Believe the startup line about memory left over, not a hopeful reading of free while the model is still loading.

curl -sf http://127.0.0.1:8080/v1/models
curl -sf http://127.0.0.1:8080/health

You want one model ID, halogen-qwen3.8-flash-next, a context of 262144, and a health document whose capability probe says ok. Then one forced tool call. The response you care about is finish_reason of tool_calls, not a paragraph that describes a tool it did not call.

Check: /v1/models lists that single ID, health is ok, and a tool call comes back as a tool call.

Phase 3 — OpenCode Web, on a different machine

Install OpenCode 2 on the machine that holds the repos. I am on 2.0.22. Pin a build you have smoked the same way you pinned the engine.

Run it as a server, not as a terminal you have to keep open. opencode web is convenient on a laptop because it opens a browser. On the host that should stay up, opencode serve is the process.

opencode serve --hostname 0.0.0.0 --port 4096

Put a password on it. Without OPENCODE_SERVER_PASSWORD the server is open to anyone who can reach the port.

# ~/.config/opencode/web.env  (mode 0600)
OPENCODE_SERVER_USERNAME=opencode
OPENCODE_SERVER_PASSWORD=choose-a-long-password

A user systemd unit keeps it up after you log out. BROWSER=/bin/true stops a headless xdg-open from failing the unit. loginctl enable-linger for that user, or the unit dies with the session.

[Service]
WorkingDirectory=%h/src
Environment=BROWSER=/bin/true
EnvironmentFile=%h/.config/opencode/web.env
ExecStart=%h/.opencode/bin/opencode serve --hostname 0.0.0.0 --port 4096
Restart=on-failure

Restrict TCP 4096 to the LAN. The HTML shell at / returns 200 without credentials. That is the page shell. The API is the door.

curl -sS -o /dev/null -w "%{http_code}\n" http://127.0.0.1:4096/
curl -sS -o /dev/null -w "%{http_code}\n" http://127.0.0.1:4096/api/session

Check: / is 200, and /api/session without the username and password is 401. Authenticated, it is 2xx. Closing the browser leaves the process and the session store running.

Phase 4 — Point every agent at the one model

OpenCode will happily ask a second model for titles, summaries, and compaction. There is one model. Tell it that, in ~/.config/opencode/opencode.json on the web host.

The provider id (local) is your name for the endpoint. The model id has to match the engine character for character.

{
  "$schema": "https://opencode.ai/config.json",
  "model": "local/halogen-qwen3.8-flash-next",
  "small_model": "local/halogen-qwen3.8-flash-next",
  "provider": {
    "local": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Local brain",
      "options": {
        "baseURL": "http://brain.lan:8080/v1",
        "apiKey": "YOUR_API_KEY"
      },
      "models": {
        "halogen-qwen3.8-flash-next": {
          "name": "Qwen3.8 Flash-Next",
          "attachment": true,
          "modalities": {
            "input": ["text", "image"],
            "output": ["text"]
          },
          "limit": {
            "context": 258048,
            "output": 16384
          }
        }
      }
    }
  },
  "agent": {
    "build": {
      "model": "local/halogen-qwen3.8-flash-next",
      "temperature": 1.0,
      "top_p": 0.95
    },
    "plan": {
      "model": "local/halogen-qwen3.8-flash-next",
      "temperature": 1.0,
      "top_p": 0.95
    },
    "explore": {
      "model": "local/halogen-qwen3.8-flash-next",
      "permission": {
        "bash": "deny",
        "edit": "deny"
      }
    },
    "title": { "model": "local/halogen-qwen3.8-flash-next" },
    "summary": { "model": "local/halogen-qwen3.8-flash-next" }
  }
}

baseURL is a placeholder. Use the LAN name of the GPU host. The apiKey field is part of the OpenAI-compatible client. Treat the network and the web password as the access control: the API stays on the LAN, and the browser requires basic auth. A string in a config file is not a substitute for either one.

explore stays read-only. A search that can also edit will start a second write loop in the same window. build is the agent that changes the tree. On a private UI you may allow bash and edit for that agent so a long job is not stuck on an approval prompt while the tab is closed. That is a convenience. It is not permission to skip the diff.

Turn compaction on when sessions get long. Reserve on the order of 20k tokens so the turn can finish after the summary. A window that dies at the edge feels like a model failure. It is a budget failure.

Sampling above matches the engine defaults I serve: temperature 1.0, top-p 0.95, top-k 20.

Check: a title, a plan, and a build turn all hit halogen-qwen3.8-flash-next. Nothing in the log is asking for a model id you did not start.

Phase 5 — A real edit

A smoke prompt lies in a friendly way. The brain is proven when OpenCode finishes a turn that looks like work, in a repo you own:

  1. Ask for a plan of a small change.
  2. Let it read the files. explore should be unable to edit them.
  3. Let build implement a thin slice.
  4. Read the diff. You merge it, or you do not.
Four steps in one session: ask in OpenCode Web, one Flash-Next context, tool calls, then a human reviews the diff. A note under the steps records prefill around 1.3 to 1.5k tokens a second and short generation around 47.
The loop is short on purpose. Generation is the middle. Ownership is the end.

If step 4 feels optional, you built a toy. Cognitive debt still applies on hardware you own: the human keeps intent, boundaries, and outcomes. Local generation is still generation. Review capacity is still the constraint that matters once the tokens are free.

Check: you can point at the diff and say what the model changed, and what you refused.

What to Refuse

A public bind. The API and the web UI belong on the LAN, with a password on the UI. “It is only my house network” is how a lab becomes an incident report the week a guest joins the Wi-Fi. If you need to reach it from elsewhere, that is a VPN, not a port forward.

A second model in the leftover RAM. Thirty-odd GiB looks like room for another brain. Loading one spends the headroom the fit was counting on, and then both sessions get slow in a way that is miserable to debug. One engine. One checkpoint. Headroom stays headroom.

Skipping review because the tokens were free. The token bill is a reason to own the runtime. It is not a reason to merge what you did not read. Faster unowned code is still unowned.

Prompt archives as a shared drive. Local does not mean free of retention rules. Treat traces like production logs. If you would not paste the prompt into the company ticket system, do not write it to a world-readable file.

Gear theater. A louder box that never completes a scored edit is the LED rack of this era. Constraints beat capacity. That is the same rule as the homelab.

You Need a Brain You Can Restart

You do not need this exact box. You need one named model, a contract other clients can speak, a UI that keeps the session when you walk away, and a habit of reading the diff.

A leadership-grade homelab keeps infrastructure judgment current. A single local brain keeps coding-agent judgment current: context as a budget, tool calls as a contract, and the difference between a fast reply and a change you will own. I use both for the same reason. Abstraction without practice produces decks.

Skin in the Runtime is the longer argument: review capacity, cognitive debt, measurement, and what happens when generation gets cheap. If you would rather start with a conversation than a model server, get in touch.

I take on selective fractional and advisory work where AI-stakes platform decisions need battle-tested judgment. For how that engagement is scoped, see The Fractional CTO Playbook.


Standing up a local coding brain, or trying to get a workflow off a vendor meter? Connect with me on LinkedIn to continue the conversation.